Prediction method, device and equipment for hydrolysis rate constant of plasticizer

By using the coefficient of determination index to select hyperparameters in QSAR modeling, the problem of insufficient transparency and interpretability in the prediction of plasticizer hydrolysis rate constant is solved, realizing the transparency and reproducibility of the model, improving the prediction accuracy, and conforming to international QSAR model specifications.

CN121724167APending Publication Date: 2026-03-24HEBEI UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing QSAR modeling techniques lack transparency and interpretability in the selection process of hyperparameters when determining the hydrolysis rate constant of plasticizers, making it difficult to guarantee the reproducibility of the model and failing to meet the requirements of international QSAR model construction specifications.

Method used

By using the coefficient of determination index of different candidate QSAR models to determine the hyperparameters of the target QSAR model, the XGBoost algorithm is used to construct the QSAR model, and the optimal combination of hyperparameters is selected based on the coefficient of determination index to ensure the transparency and interpretability of the model.

Benefits of technology

The process of hyperparameter selection is made transparent and interpretable, ensuring the reproducibility of the model, meeting the requirements of the international QSAR model guidelines, and improving the prediction accuracy and applicability of the plasticizer hydrolysis rate constant.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121724167A_ABST
    Figure CN121724167A_ABST
Patent Text Reader

Abstract

The invention provides a prediction method, device and equipment for a hydrolysis rate constant of a plasticizer, and relates to the technical field of plasticizers. The method comprises the following steps: acquiring a molecular descriptor value corresponding to a plasticizer to be predicted; inputting the molecular descriptor value into the target QSAR model to obtain a prediction result output by the target QSAR model, and determining a predicted hydrolysis rate constant of the plasticizer to be predicted based on the prediction result; wherein the value of each hyper-parameter in the target QSAR model is determined based on decision coefficient indexes of different candidate QSAR models; the values of the hyper-parameters in different candidate QSAR models are different. According to the method, the transparency and the interpretability of the QSAR model can be improved when the hydrolysis rate constant is predicted by using the QSAR modeling technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of plasticizer technology, and in particular to a method, apparatus and equipment for predicting the hydrolysis rate constant of plasticizers. Background Technology

[0002] Plasticizers, as typical high-volume, high-concern chemicals, are a key target for the control of emerging pollutants. Accurate assessment of their environmental risks relies heavily on the core parameter of the hydrolysis rate constant, which directly reflects the persistence of plasticizers in the environment.

[0003] Currently, traditional methods for obtaining the hydrolysis rate constant mainly rely on experimental determination and quantum chemical calculations. However, experimental methods are time-consuming, labor-intensive, and costly; quantum chemical calculations are difficult to accurately simulate in complex environmental media, and neither can meet the regulatory requirements for high-throughput and rapid screening.

[0004] To overcome the aforementioned limitations, the quantitative structure-activity relationship (QSAR) model has become an effective alternative strategy. This method achieves efficient prediction of the hydrolysis rate constant by establishing a mathematical relationship between molecular structural features and the hydrolysis rate constant.

[0005] In recent years, the development of machine learning technology has provided powerful tools for building high-performance QSAR models. However, when applying machine learning algorithms to QSAR modeling, the determination of model hyperparameters often relies on empirical or black-box search strategies (such as grid search combined with 5-fold cross-validation). While these methods can optimize model performance to some extent, their selection process lacks transparency and interpretability: that is, it cannot explain why a certain set of hyperparameters is more suitable for the current dataset and prediction endpoint, nor can it systematically evaluate the combined impact of hyperparameter combinations on model robustness and generalization ability.

[0006] It is worth noting that internationally recognized QSAR model building standards (such as the OECD guidelines) explicitly require that modeling algorithms and processes should be clear, reproducible, and interpretable. However, existing black-box hyperparameter optimization methods are inherently in conflict with this principle, resulting in models with low transparency, poor interpretability, and difficulty in guaranteeing reproducibility. Summary of the Invention

[0007] This invention provides a method, apparatus, and device for predicting the hydrolysis rate constant of plasticizers, in order to solve the problems of low transparency, poor interpretability, and difficulty in ensuring reproducibility in the process of determining model hyperparameters when using QSAR modeling technology to predict the hydrolysis rate constant.

[0008] In a first aspect, embodiments of the present invention provide a method for predicting the hydrolysis rate constant of a plasticizer, comprising: Obtain the molecular descriptor value corresponding to the plasticizer to be predicted; The molecular descriptor value is input into the target QSAR model to obtain the prediction result output by the target QSAR model, and the predicted hydrolysis rate constant of the plasticizer to be predicted is determined based on the prediction result. The values ​​of each hyperparameter in the target QSAR model are determined based on the coefficient of determination index of different candidate QSAR models; the values ​​of each hyperparameter are different in different candidate QSAR models.

[0009] In one possible implementation, the method for determining the values ​​of each hyperparameter in the target QSAR model includes: Obtain multiple sets of hyperparameter combinations; each hyperparameter combination contains the values ​​of each hyperparameter required by the QSAR model, and the values ​​of each hyperparameter are different in different hyperparameter combinations; For each hyperparameter combination, the hyperparameter combination is applied to the QSAR model to obtain a candidate QSAR model, and the training set is input into the candidate QSAR model to obtain the prediction result output by the candidate QSAR model; Based on the actual hydrolysis rate constant corresponding to the training set and the prediction results, the determination coefficient index corresponding to the hyperparameter combination is determined. Based on the aforementioned coefficient of determination index, the overall fit of the hyperparameter combination is determined; From the multiple sets of hyperparameter combinations, determine the set of hyperparameter combinations with the highest overall fit, and determine the values ​​of each hyperparameter in this set as the values ​​of each hyperparameter in the target QSAR model.

[0010] In one possible implementation, for each combination of hyperparameters, the overall fit of the hyperparameter combination is determined based on the determination coefficient index, including: For each combination of hyperparameters, according to Determine the overall fit of this hyperparameter combination; in, This indicates the overall fitness level corresponding to the combination of hyperparameters. This indicates the number of hyperparameters in the hyperparameter combination. In the hyperparameter combination, the first... The weights of each hyperparameter In the hyperparameter combination, the first... The fit of each hyperparameter, The coefficient of determination index represents the candidate QSAR model corresponding to the hyperparameter combination. This represents the penalty factor.

[0011] In one possible implementation, the step of inputting the training set into the candidate QSAR model to obtain the prediction result output by the candidate QSAR model includes: Obtain the value of each training sample in the training set under various molecular descriptors, as well as the logarithm of the true hydrolysis rate constant corresponding to each training sample; For each type of molecular descriptor, based on the value of each training sample under that type of molecular descriptor and the logarithm of the true hydrolysis rate constant corresponding to each training sample, the correlation coefficient between that type of molecular descriptor and the logarithm of the true hydrolysis rate constant is calculated. From various molecular descriptors, identify target molecular descriptors whose correlation coefficient is less than or equal to a first set value and greater than a second set value; The values ​​of each training sample under the target class molecule descriptor are input into the candidate QSAR model to obtain the logarithmic value of the predicted hydrolysis rate constant corresponding to each training sample output by the candidate QSAR model.

[0012] In one possible implementation, the step involves obtaining the value of the molecular descriptor corresponding to the plasticizer to be predicted; Obtain the value of the plasticizer to be predicted under the target class molecule descriptor; The value of the plasticizer to be predicted under the target class molecular descriptor is determined as the molecular descriptor value corresponding to the plasticizer to be predicted.

[0013] In one possible implementation, before inputting the molecular descriptor values ​​into the target QSAR model to obtain the prediction results output by the target QSAR model, the method further includes: Based on the value of the molecular descriptor, it is determined whether the plasticizer to be predicted conforms to the application domain of the target QSAR model; If the plasticizer to be predicted conforms to the application domain of the target QSAR model, then the process jumps to the step of inputting the molecular descriptor value into the target QSAR model to obtain the prediction result output by the target QSAR model.

[0014] In one possible implementation, detecting whether the plasticizer to be predicted conforms to the application domain of the target QSAR model based on the value of the molecular descriptor includes: Based on the value of the molecular descriptor, the leverage value corresponding to the plasticizer to be predicted is calculated; Based on the training set of the target QSAR model, determine the leverage threshold and standardized residual threshold corresponding to the target QSAR model; If the leverage value is less than the absolute value of the standardized residual threshold, and the leverage value is less than the leverage value threshold, then the plasticizer to be predicted is determined to conform to the application domain of the target QSAR model.

[0015] In one possible implementation, calculating the leverage value corresponding to the plasticizer to be predicted based on the value of the molecular descriptor includes: according to Calculate the leverage value corresponding to the plasticizer to be predicted; in, This represents the leverage value corresponding to the plasticizer to be predicted. This indicates the value of the molecular descriptor corresponding to the plasticizer to be predicted. Indicates the transpose factor. This represents the molecular descriptor matrix corresponding to the training set.

[0016] Secondly, embodiments of the present invention provide a device for predicting the hydrolysis rate constant of a plasticizer, comprising: The acquisition module is used to obtain the molecular descriptor value corresponding to the plasticizer to be predicted; The prediction module is used to input the values ​​of the molecular descriptor into the target QSAR model, obtain the prediction results output by the target QSAR model, and determine the predicted hydrolysis rate constant of the plasticizer to be predicted based on the prediction results. The values ​​of each hyperparameter in the target QSAR model are determined based on the coefficient of determination index of different candidate QSAR models; the values ​​of each hyperparameter are different in different candidate QSAR models.

[0017] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect or any possible implementation thereof.

[0018] In this embodiment of the invention, by utilizing the coefficient of determination (COD) index of different candidate QSAR models to determine the hyperparameters of the target QSAR model, the selection of hyperparameters can be directly linked to the COD index—an objective and quantifiable statistical indicator of model fit. This transforms the hyperparameter optimization process from an experience-dependent "black box search" into a process with clear standards and transparent decision-making criteria. This fundamentally solves the problems of lack of transparency and interpretability in the hyperparameter determination process.

[0019] Furthermore, this transparent and rule-based hyperparameter determination method lays the foundation for the reproducibility of the entire model, effectively solving the problem of "difficulty in guaranteeing reproducibility" in existing technologies. Moreover, it ensures that the entire modeling process meets the core requirements of international QSAR model guidelines (such as OECD principles) for "clear algorithms and well-defined processes." Attached Figure Description

[0020] Figure 1 This is an application scenario diagram of the method for predicting the hydrolysis rate constant of plasticizers provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating the implementation of a method for determining the values ​​of various hyperparameters in a target QSAR model according to an embodiment of the present invention. Figure 3 This is a flowchart illustrating the implementation of a method for predicting the hydrolysis rate constant of plasticizers according to another embodiment of the present invention. Figure 4 This is a schematic diagram of the structure of a device for predicting the hydrolysis rate constant of plasticizers provided in an embodiment of the present invention. Detailed Implementation

[0021] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0022] When applying machine learning algorithms to QSAR modeling, the determination of model hyperparameters often relies on empirical or black-box search strategies (such as grid search combined with 5-fold cross-validation). However, these black-box hyperparameter determination methods lack transparency and interpretability. Internationally recognized QSAR model building standards (such as OECD guidelines) explicitly require that modeling algorithms and processes be clear, reproducible, and interpretable. Existing black-box hyperparameter determination methods inherently conflict with this principle, resulting in models with low transparency, poor interpretability, and difficulty in guaranteeing reproducibility.

[0023] To increase the transparency and interpretability of QSAR modeling (especially the determination of model hyperparameters) and ensure the reproducibility of the modeling process, this invention constructs different candidate QSAR models using different hyperparameters. Based on the coefficient of determination (COD) during the training of each candidate QSAR model, the optimal hyperparameters are calculated and determined. Applying the optimal hyperparameters to the QSAR model determines the target QSAR model, and the hydrolysis rate constant is predicted based on the target QSAR model. This invention directly links the selection of hyperparameters with the COD, an objective and quantifiable statistical indicator of model fit. This transforms the hyperparameter optimization process from an experience-dependent "black box search" to a process with clear standards and transparent decision-making criteria, fundamentally solving the problems of lack of transparency and difficulty in ensuring reproducibility in the hyperparameter determination process.

[0024] See Figure 1 The flowchart illustrates the implementation of the method for predicting the hydrolysis rate constant of plasticizers provided in this embodiment of the invention, and is described in detail below: Step 101: Obtain the molecular descriptor value corresponding to the plasticizer to be predicted.

[0025] Each plasticizer can correspond to multiple molecular descriptors, including categories such as molecular topology, electronic distribution, and physicochemical properties.

[0026] Considering that too many molecular descriptors can lead to model overfitting, this embodiment of the invention filters various molecular descriptors to determine target-class molecular descriptors, and identifies the values ​​of the plasticizer to be predicted under the target-class molecular descriptors as the corresponding molecular descriptor values. The process of determining the target-class molecular descriptors is not detailed here, but will be explained in the following sections. Here, the number of molecular descriptor values ​​corresponding to the plasticizer to be predicted is greater than or equal to one.

[0027] Step 102: Input the molecular descriptor values ​​into the target QSAR model to obtain the prediction results output by the target QSAR model, and determine the predicted hydrolysis rate constant of the plasticizer to be predicted based on the prediction results.

[0028] Here, the target QSAR model receives the molecular descriptor value corresponding to the plasticizer to be predicted and outputs the logarithmic value of the predicted hydrolysis rate constant corresponding to the plasticizer. Based on the logarithmic value of the predicted hydrolysis rate constant output by the target QSAR model, this embodiment of the invention can inversely calculate the predicted hydrolysis rate constant corresponding to the plasticizer to be predicted.

[0029] The values ​​of each hyperparameter in the target QSAR model are determined based on the coefficient of determination index of different candidate QSAR models; the values ​​of each hyperparameter are different in different candidate QSAR models.

[0030] This invention embodiment can use the XGBoost algorithm to construct a QSAR model. Essentially, the hyperparameters in the XGBoost algorithm are the same as the hyperparameters of the QSAR model. The QSAR model is essentially an XGBoost model. The hyperparameters to be optimized and determined in the QSAR model mainly include: the learning rate (…). learning_rate ), the maximum depth of the tree ( max_depth ), column sampling ratio for each tree ( colsample_bytree ), sample sampling ratio ( subsample ) and the number of classifiers ( n_estimators ).

[0031] The embodiments of the present invention can set a variety of different hyperparameter combinations. Each hyperparameter combination includes the above 5 hyperparameters, but the values ​​of each hyperparameter are different in different hyperparameter combinations.

[0032] Here, different combinations of hyperparameters can be applied to the QSAR model to obtain different candidate QSAR models. Then, based on the coefficient of determination (COD) of each candidate QSAR model, the optimal hyperparameter combination can be determined. The COD reflects the model's goodness of fit and predictive ability.

[0033] This invention allows the application of the optimal hyperparameter combination to a QSAR model to obtain a target QSAR model, which is then used to predict the plasticizer to be predicted. Essentially, the candidate QSAR model corresponding to the optimal hyperparameter combination is the target QSAR model.

[0034] Compared to existing technologies, this invention determines the hyperparameters of the target QSAR model by utilizing the coefficient of determination (COD) index of different candidate QSAR models. This directly links the selection of hyperparameters to the COD index—an objective and quantifiable statistical indicator of model fit—transforming the hyperparameter optimization process from an experience-dependent "black box search" to a process with clearly defined standards and transparent decision-making criteria. This fundamentally solves the problems of lack of transparency and interpretability in the hyperparameter determination process.

[0035] Furthermore, this transparent and rule-based hyperparameter determination method lays the foundation for the reproducibility of the entire model, effectively solving the problem of "difficulty in guaranteeing reproducibility" in existing technologies. Moreover, it ensures that the entire modeling process meets the core requirements of international QSAR model guidelines (such as OECD principles) for "clear algorithms and well-defined processes."

[0036] The following is combined Figure 2 The method for determining the hyperparameters of the target QSAR model, i.e., the method for determining the optimal combination of hyperparameters, is described in detail below: Step 201: Obtain multiple sets of hyperparameter combinations; each set of hyperparameter combinations contains the values ​​of each hyperparameter required by the QSAR model, and the values ​​of each hyperparameter are different in different hyperparameter combinations.

[0037] Here, different values ​​can be assigned to each hyperparameter to construct multiple sets of hyperparameter combinations. For example, the learning rate ( learning_rate The value of can be (0.1, 0.01, 0.001), and the maximum depth of the tree is ( max_depth The value of ) can be (3, 6, 9), and the column sampling ratio for each tree is ( colsample_bytree The value of ) can be (0.5, 0.8, 1.0), and the sample sampling ratio ( subsampleThe value of ) can be (0.5, 0.8, 1.0), and the number of classifiers ( n_ estimators The value of ) can be (50, 100, 150). In this embodiment of the invention, multiple hyperparameter combinations can be constructed based on the values ​​of each hyperparameter, wherein the values ​​of each hyperparameter differ in different hyperparameter combinations.

[0038] Step 202: For each hyperparameter combination, apply the hyperparameter combination to the QSAR model to obtain a candidate QSAR model, and input the training set into the candidate QSAR model to obtain the prediction result output by the candidate QSAR model.

[0039] In this embodiment of the invention, various hyperparameter combinations can be applied to the QSAR model to obtain multiple candidate QSAR models.

[0040] For each candidate QSAR model, training samples from the training set can be input into the QSAR model to obtain the prediction results for each training sample corresponding to the candidate QSAR model. Here, various plasticizer samples and their corresponding true hydrolysis rate constants can be obtained from relevant literature and software databases, with each plasticizer sample serving as a training sample to construct a training set. For example, this embodiment of the invention collected 86 plasticizer samples.

[0041] In some embodiments, specific implementations of inputting the training set into a candidate QSAR model to obtain the prediction results output by the candidate QSAR model may include: Step A: Obtain the value of each training sample in the training set under various molecular descriptors, and the logarithm of the true hydrolysis rate constant corresponding to each training sample.

[0042] For each training sample, the RDKit package in Python software can be used to calculate the value of the training sample under various molecular descriptors. For example, for each training sample, the embodiments of the present invention can use the RDKit package to calculate the value of the training sample under 209 molecular descriptors.

[0043] In addition, the logarithm of the training sample can be calculated based on the actual hydrolysis rate constant corresponding to the training sample.

[0044] Step B: For each type of molecular descriptor, based on the value of each training sample under that type of molecular descriptor and the logarithm of the true hydrolysis rate constant corresponding to each training sample, calculate the correlation coefficient between that type of molecular descriptor and the logarithm of the true hydrolysis rate constant.

[0045] Here, for each type of molecular descriptor, the values ​​of each training sample under that type of molecular descriptor, and the logarithm of the true hydrolysis rate constant corresponding to each training sample, can be substituted into the formula for calculating the Pearson correlation coefficient to obtain the correlation coefficient between that type of molecular descriptor and the logarithm of the true hydrolysis rate constant, thus obtaining the correlation coefficients corresponding to each type of molecular descriptor.

[0046] Step C: From various molecular descriptors, determine the target class molecular descriptors whose correlation coefficient is less than or equal to a first set value and greater than a second set value.

[0047] In this embodiment of the invention, molecular descriptors with a correlation coefficient less than or equal to a first set value and greater than a second set value are determined as target class molecular descriptors.

[0048] For example, the first set value can be 0.9, and the second set value can be 0.55. Here, by removing molecular descriptors with a correlation coefficient greater than 0.9, high collinearity can be avoided; by removing molecular descriptors with a correlation coefficient less than 0.55, a large number of weakly correlated or noisy features can be removed, avoiding model overfitting. In this embodiment of the invention, by removing molecular descriptors with a correlation coefficient greater than 0.9 and molecular descriptors with a correlation coefficient less than 0.55, the above 209 molecular descriptors can be reduced to 7 target molecular descriptors.

[0049] Furthermore, for the plasticizer to be predicted, embodiments of the present invention can use the RDKit package to calculate the value of the plasticizer to be predicted under the above-mentioned target class molecular descriptor, and determine the value of the plasticizer under the target class molecular descriptor as the molecular descriptor value of the plasticizer to be predicted, which is then input into the target QSAR model for prediction.

[0050] Step D: Input the values ​​of each training sample under the target class molecule descriptor into the candidate QSAR model to obtain the logarithmic value of the predicted hydrolysis rate constant corresponding to each training sample output by the candidate QSAR model.

[0051] For each training sample, the value of that training sample under the target class molecule descriptor can be input into the candidate QSAR model to obtain the prediction result output by the candidate QSAR model. It should be noted that the prediction result refers to the logarithmic value of the predicted hydrolysis rate constant.

[0052] Step 203: Based on the actual hydrolysis rate constant and prediction results corresponding to the training set, determine the coefficient of determination index corresponding to the hyperparameter combination.

[0053] Based on the logarithmic value of the predicted hydrolysis rate constant output by the candidate QSAR model, the embodiments of the present invention can back-calculate the predicted hydrolysis rate constant of each training sample.

[0054] The embodiments of the present invention can calculate the coefficient of determination index of the candidate QSAR model, that is, the coefficient of determination index corresponding to the hyperparameter combination, based on the true hydrolysis rate constant and the predicted hydrolysis rate constant corresponding to each training sample in the training set.

[0055] The formula for calculating the coefficient of determination can be expressed as:

[0056] in, This represents the coefficient of determination index. Indicates the number of training samples. Indicates the first The true hydrolysis rate constant of each training sample Indicates the first The predicted hydrolysis rate constant for each training sample. This represents the average of the true hydrolysis rate constants for all training samples.

[0057] Step 204: Determine the overall fitness of the hyperparameter combination based on the coefficient of determination index.

[0058] For each combination of hyperparameters, according to Determine the overall fit of this hyperparameter combination; in, This indicates the overall fitness level corresponding to the combination of hyperparameters. This indicates the number of hyperparameters in the hyperparameter combination. In the hyperparameter combination, the first... The weights of each hyperparameter In the hyperparameter combination, the first... The fit of each hyperparameter, The coefficient of determination index represents the candidate QSAR model corresponding to the hyperparameter combination. This represents the penalty factor, whose value is set based on how well the hyperparameter combination satisfies and deviates from specific constraints, typically between 0 and 1. It is used to quantify the degree to which the hyperparameter combination meets these conditions and to calculate the final overall fitness.

[0059] Specifically, a constraint is set for each hyperparameter in the hyperparameter combination. If the hyperparameter satisfies the constraint, its penalty coefficient is set to 1. If the hyperparameter does not satisfy the constraint, the penalty coefficient is gradually reduced based on the deviation between the hyperparameter and the constraint. For each hyperparameter combination, the penalty coefficients of each hyperparameter in the combination are multiplied together to obtain the penalty factor.

[0060] For example, targeting max_depthThis hyperparameter has the constraint that max_depth > 6. For each combination of hyperparameters... max_depth, like max_depth If the value is less than 6, then the hyperparameter combination is determined to be in this combination. max_depth The penalty coefficient is 1. If max_depth If the value is greater than 6, then it is determined. max_depth For every increase of 1 in the value of , the penalty coefficient is reduced by 0.1. For example, if max_depth If the value is 7, then the penalty coefficient is 0.9; if max_depth If the value of is 8, then its penalty coefficient is 0.8. In essence, the constraint conditions corresponding to each hyperparameter are the boundary conditions for the values ​​of each hyperparameter.

[0061] Here, the weights and fitness of each hyperparameter in the hyperparameter combination can be determined according to the actual situation, and the embodiments of the present invention do not impose specific limitations on this.

[0062] Based on the exemplary values ​​of the hyperparameters mentioned above, the optimal combination of hyperparameters can be calculated and determined as {' learning_rat ': 0.1,' max_depth ': 3,' colsample_bytree ': 0.5,' subsample ': 1.0,' n_estimators ': 150}. The overall fit of this optimal hyperparameter combination. It is 0.985.

[0063] Step 205: From multiple sets of hyperparameter combinations, determine the set of hyperparameter combinations with the highest comprehensive fit, and determine the values ​​of each hyperparameter in this set of hyperparameters as the values ​​of each hyperparameter in the target QSAR model.

[0064] Based on the coefficient of determination index of the candidate QSAR model corresponding to each hyperparameter combination, the overall fitness of each hyperparameter combination can be calculated. In this embodiment of the invention, the hyperparameter combination with the highest overall fitness is determined as the optimal hyperparameter combination, and the values ​​of each hyperparameter in the optimal hyperparameter combination are determined as the hyperparameter values ​​in the target QSAR model.

[0065] For the target QSAR model, this embodiment of the invention can input each training sample in the training set into the target QSAR model to obtain the prediction result output by the target QSAR model, and calculate the coefficient of determination index of the target QSAR model based on the prediction result of the training samples and the actual hydrolysis rate constant. Mean Absolute Error Index MAE Mean square error index RMSE and root mean square error index MSE The coefficient of determination is used to reflect the goodness of fit of the target QSAR model. The closer it is to 1, the higher the mean absolute error index. MAE Mean square error index RMSE and root mean square error index MSE The closer the value is to 0, the better the model fit.

[0066] In addition, a test set can be set up, and each sample in the test set can be input into the target QSAR model to calculate the coefficient of determination index. Mean Absolute Error Index MAE Mean square error index ​ and root mean square error index ​ This is used to reflect the predictive power of the target QSAR model. In addition, the cross-validation coefficients of the target QSAR model can be calculated to reflect its robustness.

[0067] For example, with the above optimal hyperparameter combination {' ​ ': 0.1,' ​ ': 3,' ​ ': 0.5,' ​ ': 1.0,' ​ Taking the target QSAR model with a defined target of ': 150' as an example, based on the training set, the coefficient of determination index of the target QSAR model can be calculated to be 0.992, and the mean absolute error index is... ​ The mean square error index is 0.095. ​ The root mean square error index is 0.128. ​ The coefficient of determination (COD) for this target on the QSAR model is 0.908, based on the test set. The mean absolute error (MAE) is 0.016. ​ The mean square error index is 0.278. ​ The root mean square error index is 0.504. ​ It is 0.254.

[0068] See ​ Based on the above embodiments, this invention provides another method for predicting the hydrolysis rate constant of plasticizers, comprising the following steps: Step 301: Obtain the molecular descriptor value corresponding to the plasticizer to be predicted.

[0069] For details on the specific implementation of step 301, please refer to [link / reference]. ​ The corresponding implementation examples will not be described in detail here.

[0070] Step 302: Based on the value of the molecular descriptor, detect whether the plasticizer to be predicted conforms to the application domain of the target QSAR model.

[0071] Here, the Williams diagram can be used to characterize the application domain of the model. When detecting whether the plasticizer to be predicted conforms to the application domain of the target QSAR model, the leverage value corresponding to the plasticizer to be predicted can be calculated first based on the value of the molecular descriptor. Then, based on the training set of the target QSAR model, the leverage value threshold and the standardized residual threshold corresponding to the target QSAR model are determined. If the leverage value is less than the absolute value of the standardized residual threshold and the leverage value is less than the leverage value threshold, then it is determined that the plasticizer to be predicted conforms to the application domain of the target QSAR model.

[0072] Specifically, according to The leverage value corresponding to the plasticizer to be predicted can be calculated; in, This represents the leverage value corresponding to the plasticizer to be predicted. This indicates the value of the molecular descriptor corresponding to the plasticizer to be predicted. Indicates the transpose factor. This represents the molecular descriptor matrix corresponding to the training set. The molecular descriptor matrix contains the values ​​of the molecular descriptors corresponding to each training sample in the training set.

[0073] according to The leverage threshold corresponding to the target QSAR model can be calculated; in, Indicates the leverage threshold. Indicates the number of values ​​that the molecule descriptor can take.

[0074] according to The standardized residual threshold corresponding to the target QSAR model can be calculated; in, This represents the standardized residual threshold.

[0075] Step 303: If the plasticizer to be predicted conforms to the application domain of the target QSAR model, then input the molecular descriptor value into the target QSAR model to obtain the prediction result output by the target QSAR model, and determine the predicted hydrolysis rate constant of the plasticizer to be predicted based on the prediction result.

[0076] For details on the specific implementation of step 303, please refer to [link / reference]. ​ The corresponding implementation examples will not be described in detail here.

[0077] The embodiments of this invention can rapidly and accurately predict the hydrolysis rate constant of plasticizers, a typical new pollutant. The prediction accuracy and applicability of the target QSAR model significantly surpass those of traditional linear models, and the target QSAR model can effectively reveal the key molecular structural features and mechanisms of action that affect the degradation of plasticizers in the aquatic environment.

[0078] In this embodiment of the invention, the modeling process of the target QSAR model strictly follows OECD standards, ensuring the scientific reliability of the results, thereby providing an effective management tool for the environmental persistence assessment and ecological risk assessment of new pollutants.

[0079] Here, we illustrate the effectiveness of the target QSAR model provided in the embodiments of the present invention with several examples.

[0080] Given a plasticizer molecule: di(2-ethylhexyl) phthalate (DEHP, CAS Registry No.: 117-81-7), predict the logarithm of its alkaline hydrolysis rate constant (log k H First, based on the DEHP SMILES code, the molecular descriptor values ​​are calculated using the RDKit software package. Then, the leverage value is calculated to confirm that it falls within the application domain of the target QSAR model. Finally, the logarithmic value of its hydrolysis rate constant is predicted using the target QSAR model. The experimental value (i.e., the logarithmic value of the true hydrolysis rate constant determined experimentally) is: log k H = -2.100. The model's predicted value is: log k H = -2.154. The model prediction is very close to the experimental value, proving the effectiveness of the target QSAR model.

[0081] Given a plasticizer molecule: dibutyl phthalate (DBP, CAS Registry No.: 84-74-2), predict the logarithm of its alkaline hydrolysis rate constant (log k H First, based on the DBP's SMILES code, the molecular descriptor values ​​were calculated using the RDKit software package, confirming that they fall within the application domain of the target QSAR model. Then, the logarithm of its hydrolysis rate constant was predicted using the target QSAR model. The experimental value is: log k H = -1.760. The model's predicted value is: log k H = -1.812. The model prediction is in high agreement with the experimental value.

[0082] Given a plasticizer molecule: acetyl tributyl citrate (ATBC, CAS Registry No.: 77-90-7), predict the logarithm of its alkaline hydrolysis rate constant (log k HFirst, based on the ATBC SMILES code, its molecular descriptor was calculated using the RDKit software package, and application domain evaluation confirmed that it lies within the application domain of the target QSAR model. Then, the logarithm of its hydrolysis rate constant was predicted using the target QSAR model. The experimental value is: log k H = -0.700, the model prediction value is: log k H = -0.6. The model predictions are in excellent agreement with the experimental values.

[0083] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0084] The following are device embodiments of the present invention. For details not described in detail, please refer to the corresponding method embodiments described above.

[0085] ​ A schematic diagram of the structure of the device for predicting the hydrolysis rate constant of plasticizers provided in an embodiment of the present invention is shown. For ease of explanation, only the parts related to the embodiments of the present invention are shown, and are described in detail below: like ​ As shown, the device 4 for predicting the hydrolysis rate constant of plasticizers includes an acquisition module 41 and a prediction module 42.

[0086] The acquisition module 41 is used to acquire the value of the molecular descriptor corresponding to the plasticizer to be predicted; The prediction module 42 is used to input the molecular descriptor values ​​into the target QSAR model, obtain the prediction results output by the target QSAR model, and determine the predicted hydrolysis rate constant of the plasticizer to be predicted based on the prediction results. The values ​​of each hyperparameter in the target QSAR model are determined based on the coefficient of determination index of different candidate QSAR models; the values ​​of each hyperparameter are different in different candidate QSAR models.

[0087] Optionally, methods for determining the values ​​of various hyperparameters in the target QSAR model include: Obtain multiple sets of hyperparameter combinations; each hyperparameter combination contains the values ​​of each hyperparameter required by the QSAR model, and the values ​​of each hyperparameter are different in different hyperparameter combinations; For each hyperparameter combination, the hyperparameter combination is applied to the QSAR model to obtain a candidate QSAR model, and the training set is input into the candidate QSAR model to obtain the prediction result output by the candidate QSAR model; Based on the actual hydrolysis rate constant and prediction results corresponding to the training set, the determination coefficient index corresponding to this hyperparameter combination is determined. The overall fitness of the hyperparameter combination is determined based on the coefficient of determination index. From multiple sets of hyperparameter combinations, the hyperparameter combination with the highest overall fit is determined, and the values ​​of each hyperparameter in this hyperparameter combination are determined as the values ​​of each hyperparameter in the target QSAR model.

[0088] Optionally, for each hyperparameter combination, the overall fitness of that hyperparameter combination is determined based on the coefficient of determination index, including: For each combination of hyperparameters, according to Determine the overall fit of this hyperparameter combination; in, This indicates the overall fitness level corresponding to the combination of hyperparameters. This indicates the number of hyperparameters in the hyperparameter combination. In the hyperparameter combination, the first... The weights of each hyperparameter In the hyperparameter combination, the first... The fit of each hyperparameter, The coefficient of determination index represents the candidate QSAR model corresponding to the hyperparameter combination. This represents the penalty factor.

[0089] Optionally, the training set can be input into the candidate QSAR model to obtain the prediction results output by the candidate QSAR model, including: Obtain the value of each training sample in the training set under various molecular descriptors, as well as the logarithm of the true hydrolysis rate constant corresponding to each training sample; For each type of molecular descriptor, based on the value of each training sample under that type of molecular descriptor and the logarithm of the true hydrolysis rate constant corresponding to each training sample, the correlation coefficient between that type of molecular descriptor and the logarithm of the true hydrolysis rate constant is calculated. From various molecular descriptors, identify target molecular descriptors whose correlation coefficient is less than or equal to a first set value and greater than a second set value; The values ​​of each training sample under the target class molecule descriptor are input into the candidate QSAR model to obtain the logarithmic value of the predicted hydrolysis rate constant corresponding to each training sample output by the candidate QSAR model.

[0090] Optionally, module 41 is used for: Obtain the value of the plasticizer to be predicted under the target class molecule descriptor; The value of the plasticizer to be predicted under the target class molecular descriptor is determined as the value of the molecular descriptor corresponding to the plasticizer to be predicted.

[0091] Optionally, prediction module 42 is also used for: Based on the values ​​of molecular descriptors, it is determined whether the plasticizer to be predicted conforms to the application domain of the target QSAR model; If the plasticizer to be predicted conforms to the application domain of the target QSAR model, then proceed to the step of inputting the molecular descriptor value into the target QSAR model to obtain the prediction result output by the target QSAR model.

[0092] Optional, prediction module 42, specifically used for: Based on the values ​​of the molecular descriptors, the leverage value corresponding to the plasticizer to be predicted is calculated. Based on the training set of the target QSAR model, determine the leverage threshold and standardized residual threshold corresponding to the target QSAR model; If the leverage value is less than the absolute value of the standardized residual threshold, and the leverage value is less than the leverage value threshold, then the plasticizer to be predicted is determined to conform to the application domain of the target QSAR model.

[0093] Optional, prediction module 42, specifically used for: according to Calculate the leverage value corresponding to the plasticizer to be predicted; in, This represents the leverage value corresponding to the plasticizer to be predicted. This indicates the value of the molecular descriptor corresponding to the plasticizer to be predicted. Indicates the transpose factor. This represents the molecular descriptor matrix corresponding to the training set.

[0094] This device embodiment can be used to implement the above method embodiment, and its technical principle and implementation effect are the same as those of the above method embodiment.

[0095] This invention also provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the above method embodiments.

[0096] In the above embodiments, the descriptions of each embodiment have their own emphasis. Parts not detailed or described in a particular embodiment can be referred to in the relevant descriptions of other embodiments. Unless otherwise specified or in conflict with logic, the terminology and / or descriptions between different embodiments are consistent and can be referenced interchangeably. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.

[0097] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for predicting the hydrolysis rate constant of a plasticizer, characterized in that, include: Obtain the molecular descriptor value corresponding to the plasticizer to be predicted; The molecular descriptor value is input into the target QSAR model to obtain the prediction result output by the target QSAR model, and the predicted hydrolysis rate constant of the plasticizer to be predicted is determined based on the prediction result. The values ​​of each hyperparameter in the target QSAR model are determined based on the coefficient of determination index of different candidate QSAR models; the values ​​of each hyperparameter are different in different candidate QSAR models.

2. The method for predicting the hydrolysis rate constant of plasticizers according to claim 1, characterized in that, The method for determining the values ​​of each hyperparameter in the target QSAR model includes: Obtain multiple sets of hyperparameter combinations; each hyperparameter combination contains the values ​​of each hyperparameter required by the QSAR model, and the values ​​of each hyperparameter are different in different hyperparameter combinations; For each hyperparameter combination, the hyperparameter combination is applied to the QSAR model to obtain a candidate QSAR model, and the training set is input into the candidate QSAR model to obtain the prediction result output by the candidate QSAR model; Based on the actual hydrolysis rate constant corresponding to the training set and the prediction results, the determination coefficient index corresponding to the hyperparameter combination is determined. Based on the aforementioned coefficient of determination index, the overall fit of the hyperparameter combination is determined; From the multiple sets of hyperparameter combinations, determine the set of hyperparameter combinations with the highest overall fit, and determine the values ​​of each hyperparameter in this set as the values ​​of each hyperparameter in the target QSAR model.

3. The method for predicting the hydrolysis rate constant of plasticizers according to claim 2, characterized in that, For each combination of hyperparameters, the overall fitness of that combination is determined based on the coefficient of determination index, including: For each combination of hyperparameters, according to Determine the overall fit of this hyperparameter combination; in, This indicates the overall fitness level corresponding to the combination of hyperparameters. This indicates the number of hyperparameters in the hyperparameter combination. In the hyperparameter combination, the first... The weights of each hyperparameter In the hyperparameter combination, the first... The fit of each hyperparameter, The coefficient of determination index represents the candidate QSAR model corresponding to the hyperparameter combination. This represents the penalty factor.

4. The method for predicting the hydrolysis rate constant of a plasticizer according to claim 2 or 3, characterized in that, The step of inputting the training set into the candidate QSAR model to obtain the prediction result output by the candidate QSAR model includes: Obtain the value of each training sample in the training set under various molecular descriptors, as well as the logarithm of the true hydrolysis rate constant corresponding to each training sample; For each type of molecular descriptor, based on the value of each training sample under that type of molecular descriptor and the logarithm of the true hydrolysis rate constant corresponding to each training sample, the correlation coefficient between that type of molecular descriptor and the logarithm of the true hydrolysis rate constant is calculated. From various molecular descriptors, identify target molecular descriptors whose correlation coefficient is less than or equal to a first set value and greater than a second set value; The values ​​of each training sample under the target class molecule descriptor are input into the candidate QSAR model to obtain the logarithmic value of the predicted hydrolysis rate constant corresponding to each training sample output by the candidate QSAR model.

5. The method for predicting the hydrolysis rate constant of plasticizers according to claim 4, characterized in that, The value of the molecular descriptor corresponding to the plasticizer to be predicted is obtained; Obtain the value of the plasticizer to be predicted under the target class molecule descriptor; The value of the plasticizer to be predicted under the target class molecular descriptor is determined as the molecular descriptor value corresponding to the plasticizer to be predicted.

6. The method for predicting the hydrolysis rate constant of a plasticizer according to any one of claims 1-3, characterized in that, Before inputting the molecular descriptor values ​​into the target QSAR model to obtain the prediction results output by the target QSAR model, the process further includes: Based on the value of the molecular descriptor, it is determined whether the plasticizer to be predicted conforms to the application domain of the target QSAR model; If the plasticizer to be predicted conforms to the application domain of the target QSAR model, then the process jumps to the step of inputting the molecular descriptor value into the target QSAR model to obtain the prediction result output by the target QSAR model.

7. The method for predicting the hydrolysis rate constant of plasticizers according to claim 6, characterized in that, The step of detecting whether the plasticizer to be predicted conforms to the application domain of the target QSAR model based on the value of the molecular descriptor includes: Based on the value of the molecular descriptor, the leverage value corresponding to the plasticizer to be predicted is calculated; Based on the training set of the target QSAR model, determine the leverage threshold and standardized residual threshold corresponding to the target QSAR model; If the leverage value is less than the absolute value of the standardized residual threshold, and the leverage value is less than the leverage value threshold, then the plasticizer to be predicted is determined to conform to the application domain of the target QSAR model.

8. The method for predicting the hydrolysis rate constant of a plasticizer according to claim 7, characterized in that, The step of calculating the leverage value corresponding to the plasticizer to be predicted based on the value of the molecular descriptor includes: according to Calculate the leverage value corresponding to the plasticizer to be predicted; in, This represents the leverage value corresponding to the plasticizer to be predicted. This indicates the value of the molecular descriptor corresponding to the plasticizer to be predicted. Indicates the transpose factor. This represents the molecular descriptor matrix corresponding to the training set.

9. A device for predicting the hydrolysis rate constant of a plasticizer, characterized in that, include: The acquisition module is used to obtain the molecular descriptor value corresponding to the plasticizer to be predicted; The prediction module is used to input the values ​​of the molecular descriptor into the target QSAR model, obtain the prediction results output by the target QSAR model, and determine the predicted hydrolysis rate constant of the plasticizer to be predicted based on the prediction results. The values ​​of each hyperparameter in the target QSAR model are determined based on the coefficient of determination index of different candidate QSAR models; the values ​​of each hyperparameter are different in different candidate QSAR models.

10. An electronic device, characterized in that, It includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 8.