An Optimal Synthesis Method, System, Device and Medium of Hydrogel

By establishing a dataset of hydrogel adsorption of heavy metals, using the random forest algorithm to fill in missing values ​​and training the XGBoost model, important features were screened out, solving the problems of low prediction accuracy and environmental pollution in the hydrogel synthesis process, and realizing rapid screening of the optimal synthesis conditions.

CN116525038BActive Publication Date: 2025-08-01SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310416801.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-19
Publication Date
2025-08-01
Estimated Expiration
2043-04-19

AI Technical Summary

Technical Problem

Existing technologies for hydrogel synthesis suffer from problems such as time-consuming screening of synthesis conditions, environmental pollution, and low prediction accuracy, especially the lack of effective machine learning models for the adsorption process.

Method used

By establishing a dataset of heavy metal adsorption by hydrogels, missing values ​​were filled using the random forest algorithm, various machine learning models were trained, the best-performing model such as XGBoost was selected, the SHAP value was calculated, important features were screened out, and the optimal preparation conditions for hydrogels were constructed.

Benefits of technology

This improves the accuracy of predicting hydrogel adsorption coefficients, enables rapid screening of optimal synthesis conditions, and reduces the environmental impact and cost of the synthesis process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116525038B_ABST
    Figure CN116525038B_ABST
Patent Text Reader

Abstract

The present invention discloses an optimal synthesis method, system, device and medium of hydrogel, belonging to the field of hydrogel synthesis. First, the input features of the data set are preprocessed to solve the potential data imbalance problem of the source data. The data set obtained by using the random forest algorithm to fill in the missing values is used to construct a machine learning model, which helps to further improve the prediction accuracy of the given modeling algorithm. Then, multiple machine learning models are constructed, and the machine learning model with the best performance is selected from them to predict the adsorption coefficient. The SHAP method is used to select the input features most relevant to the prediction target to form different hydrogel preparation conditions, and then the optimal machine learning model is input to obtain the best preparation conditions of the hydrogel. The present invention improves the prediction accuracy of the hydrogel adsorption coefficient and can quickly screen the best synthesis conditions of the hydrogel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of hydrogel synthesis, and particularly to an optimal hydrogel synthesis method, system, device and medium. Background Art

[0002] Hydrogel composites, due to their large surface area, can be designed to have a high density of metal ion coordination groups and recycled, showing excellent performance and potential applications in adsorbing and removing heavy metals in wastewater. A better understanding of the adsorption mechanism of hydrogels and a full understanding of the properties of hydrogels can effectively guide the design and preparation of new and efficient adsorbents. The selection of functional monomers is crucial for the hydrogel polymerization process, and the best adsorption performance of hydrogels is determined by the matching degree between the target ions and the functional monomers. In addition, the dosage of the initiator determines the polymerization rate and molecular distribution of the hydrogel. As mentioned above, the selection of suitable functional monomers, crosslinking agents and initiators for hydrogel preparation will determine the final adsorption efficiency of the hydrogel. In order to obtain high-quality hydrogel absorbents, researchers must invest a lot of time, energy and cost in screening out the best synthesis conditions. And the preparation of hydrogels usually takes a long time, and the remaining reactants in the synthesis process will pollute the environment. For example, the commonly used initiators (ammonium persulfate, potassium persulfate) are harmful to the human body. Therefore, there is an urgent need for a theoretical system to help quickly screen the best synthesis conditions by operating specific procedures for simulating and guiding practical investigations.

[0003] Machine learning algorithms have received a great deal of attention for their powerful fitting ability to handle complex and multi-dimensional data sets. In recent years, machine learning (ML) methods have demonstrated powerful capabilities in the research of environmental functional materials due to their low cost, high prediction accuracy, strong robustness, etc. Compared with traditional methods from scratch, machine learning methods based on large-scale databases constructed from experimental or computational means can provide more reliable results. The data-driven mode of ML is widely used in the design of new materials by providing some guiding rules, such as using the XGBoost model to guide the controllable synthesis of high-quality one-dimensional few-layer WTe2 nanoribbons. In issues related to adsorption prediction, a large number of studies have also applied machine learning for parameter optimization and constructing prediction models. However, currently, it is lacking to incorporate the contributions of relevant influencing factors involved in the adsorption process, such as adsorbent properties (specific surface area, pore volume, swelling ratio, etc.), equilibrium time, and temperature, to adsorption prediction into ML modeling research. Therefore, it may affect the performance of the ML model and change the feature importance. In addition to the algorithm, the richness of the source data plays a crucial role in improving the prediction accuracy. For experimental studies like adsorption, the available data is almost always limited, and fewer data samples are not sufficient to support the wide application of the entire ML model in other scenarios. Therefore, choosing different algorithms or optimizing hyperparameters will be beneficial for improving the prediction. However, the issue of how to solve the prediction accuracy problem of source data processing brought about by data scarcity has rarely been mentioned. Summary of the Invention

[0004] The purpose of the present invention is to provide a method, system, device, and medium for the optimal synthesis of hydrogels, which can improve the prediction accuracy of the adsorption coefficient of hydrogels and quickly screen the optimal synthesis conditions of hydrogels.

[0005] To achieve the above object, the present invention provides the following solutions:

[0006] A method for the optimal synthesis of hydrogels includes:

[0007] Establish a data set for a hydrogel to adsorb a heavy metal; the input features of the data set include adsorbent properties, synthesis conditions, heavy metal properties, and adsorption conditions, and the label of the data set is the logarithm of the adsorption coefficient of the hydrogel to adsorb the heavy metal;

[0008] Preprocess the input features of the data set; the preprocessing includes using the random forest algorithm to fill in the missing values in the input features;

[0009] Use the preprocessed data set to train multiple machine learning models respectively, and select the trained machine learning model with the optimal performance as the optimal adsorption prediction model; the multiple machine learning models include a random forest model, a support vector machine model, a gradient boosting decision tree model, and an XGBoost model;

[0010] Calculate the SHAP values of each input feature for the adsorption coefficient using the optimal adsorption prediction model, and select the input features corresponding to the preset number of SHAP values from largest to smallest;

[0011] Set multiple values for each selected input feature, and combine the values of all selected input features arbitrarily to form different hydrogel preparation conditions;

[0012] Input each hydrogel preparation condition into the optimal adsorption prediction model, output the logarithm of the adsorption coefficient of the hydrogel for adsorbing the heavy metal under each hydrogel preparation condition, and determine the hydrogel preparation condition corresponding to the maximum logarithm of the adsorption coefficient as the optimal preparation condition for the hydrogel used to adsorb the heavy metal.

[0013] Optionally, the properties of the adsorbent include: specific surface area, pore volume, pore diameter, swelling ratio, and mechanical strength;

[0014] The synthesis conditions include: monomer functional group, type of nanomaterial, synthesis time, synthesis temperature, type of initiator, dosage of initiator, type of crosslinker, and dosage of crosslinker;

[0015] The properties of the heavy metal include: first ionization energy, ionic radius, and hydrated ionic radius;

[0016] The adsorption conditions include: initial concentration of the heavy metal, equilibrium concentration of the heavy metal, dosage of the hydrogel, adsorption oscillation frequency, adsorption time, equilibrium time, solution pH, and temperature.

[0017] Optionally, the process of establishing the label is as follows:

[0018] Collect the adsorption isotherm of the hydrogel for adsorbing the heavy metal; the abscissa of the adsorption isotherm is the equilibrium concentration C of the heavy metal e , and the ordinate is the adsorption capacity Q of the hydrogel for adsorbing the heavy metal e ;

[0019] Obtain the data points corresponding to the equilibrium concentration values of the heavy metal in the input features from the adsorption isotherm;

[0020] Use the formula logK d = log(Q e / C e ), calculate the logarithm of the adsorption coefficient corresponding to each data point; in the formula, K d represents the adsorption coefficient, and logK d represents the logarithm of the adsorption coefficient.

[0021] Optionally, preprocess the input features of the data set, specifically including:

[0022] Perform dummy variable processing on the input features to convert the text variables in the dataset into numerical variables;

[0023] Use the Spearman coefficient to perform correlation analysis on the input features after dummy variable processing, and eliminate redundant features;

[0024] Use the random forest algorithm to fill in the missing values in the input features after eliminating redundant features.

[0025] Optionally, train multiple machine learning models using the preprocessed dataset, specifically including:

[0026] Split the preprocessed dataset into a training set and a test set in a ratio of 8:2, and perform ten-fold cross-validation on the training set, so that the training set is divided into a pre-training set and a validation set in a ratio of 9:1;

[0027] Use the pre-training set, validation set, and test set to train each machine learning model, and select the Bayesian optimization algorithm to optimize and adjust the hyperparameters of the machine learning model to obtain multiple trained machine learning models.

[0028] Optionally, calculate the SHAP values of each input feature for the adsorption coefficient using the optimal adsorption prediction model, and select the input features corresponding to the preset number of SHAP values from largest to smallest, specifically including:

[0029] According to the preprocessed dataset, use the optimal adsorption prediction model to calculate the average absolute Shapley value of each input feature for the adsorption coefficient;

[0030] Sort the average absolute Shapley values of each input feature for the adsorption coefficient in descending order, and select the input features corresponding to the preset number of average absolute Shapley values from largest to smallest.

[0031] A hydrogel optimal synthesis system, comprising:

[0032] A dataset establishment module for establishing a dataset of a hydrogel adsorbing a heavy metal; the input features of the dataset include adsorbent properties, synthesis conditions, heavy metal properties, and adsorption conditions, and the label of the dataset is the adsorption coefficient of the hydrogel adsorbing the heavy metal;

[0033] A preprocessing module for preprocessing the input features of the dataset; the preprocessing includes using the random forest algorithm to fill in the missing values in the input features;

[0034] A training module, which is used to train multiple machine learning models respectively using the preprocessed dataset, and select the trained machine learning model with the best performance as the optimal adsorption prediction model; the multiple machine learning models include a random forest model, a support vector machine model, a gradient boosting decision tree model, and an XGBoost model;

[0035] A feature selection module, which is used to calculate the SHAP values of each input feature on the adsorption coefficient using the optimal adsorption prediction model, and select the input features corresponding to a preset number of SHAP values from large to small;

[0036] A preparation condition composition module, which is used to set multiple values for each selected input feature, and combine the values of all selected input features arbitrarily to form different hydrogel preparation conditions;

[0037] A prediction module, which is used to input each hydrogel preparation condition into the optimal adsorption prediction model, output the logarithm of the adsorption coefficient of the hydrogel adsorbing the heavy metal under each hydrogel preparation condition, and determine the hydrogel preparation condition corresponding to the maximum logarithm of the adsorption coefficient as the best preparation condition of the hydrogel for adsorbing the heavy metal.

[0038] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the optimal hydrogel synthesis method as described above is implemented.

[0039] A computer-readable storage medium stores a computer program, and when the computer program is executed, the optimal hydrogel synthesis method as described above is implemented.

[0040] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0041] The present invention discloses an optimal hydrogel synthesis method, system, device, and medium. First, the input features of the dataset are preprocessed to solve the potential data imbalance problem of the source data. The dataset obtained by using the random forest algorithm to fill in the missing values is used to construct a machine learning model, which helps to further improve the prediction accuracy of the given modeling algorithm. Then, multiple machine learning models are constructed, and the machine learning model with the best performance is selected from them to predict the adsorption coefficient. The SHAP method is used to select the input features most relevant to the prediction target to form different hydrogel preparation conditions, and then input into the optimal machine learning model to obtain the best preparation conditions of the hydrogel. The present invention improves the prediction accuracy of the hydrogel adsorption coefficient and can quickly screen the best synthesis conditions of the hydrogel. Description of the Drawings

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 A flow chart of an optimal hydrogel synthesis method provided by an embodiment of the present invention;

[0044] Figure 2 A schematic diagram of correlation analysis between input features of the nanocomposite hydrogel provided in an embodiment of the present invention;

[0045] Figure 3 A schematic diagram of correlation analysis between input features of a synthetic hydrogel provided by an embodiment of the present invention;

[0046] Figure 4 A schematic diagram of correlation analysis between input features of a dual-network hydrogel provided in an embodiment of the present invention;

[0047] Figure 5 Schematic diagram of the performance of different machine learning algorithms for nanocomposite hydrogels provided in embodiments of the present invention in predicting hydrogel adsorption;

[0048] Figure 6 Schematic diagram of the performance of different machine learning algorithms for synthesizing hydrogels provided in embodiments of the present invention in predicting hydrogel adsorption;

[0049] Figure 7 Schematic diagram of the performance of different machine learning algorithms for predicting hydrogel adsorption for dual-network hydrogels provided in an embodiment of the present invention;

[0050] Figure 8 Schematic diagram of the evaluation of the XGBoost model performance for heavy metal adsorption by synthetic hydrogels before missing value filling provided by an embodiment of the present invention;

[0051] Figure 9 Schematic diagram of evaluating the performance of the XGBoost model for heavy metal adsorption by synthetic hydrogels after missing value filling provided by an embodiment of the present invention;

[0052] Figure 10 A schematic diagram of a visual analysis of the SHAP values of heavy metal adsorption by the synthetic hydrogel before missing value filling provided in an embodiment of the present invention;

[0053] Figure 11 A schematic diagram of the visualization analysis of the SHAP values of heavy metal adsorption by the synthetic hydrogel after missing value filling provided by an embodiment of the present invention;

[0054] Figure 12 Schematic diagram of the positive and negative correlations and magnitude rankings of the top 20 important features of the machine learning model for the adsorption of heavy metals by the synthetic hydrogel provided in the embodiment of the present invention;

[0055] Figure 13 Schematic diagram of the top ten important features affecting the adsorption effect of the synthetic hydrogel on heavy metals provided in the embodiment of the present invention;

[0056] Figure 14 Schematic diagram of the correlation analysis of the input features of the synthetic hydrogel provided in the embodiment of the present invention;

[0057] Figure 15 Schematic diagram of the visual analysis of the SHAP values of the bivariate superposition effect in the process of the synthetic hydrogel adsorbing heavy metals provided in the embodiment of the present invention. Detailed implementation manners

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] The object of the present invention is to provide an optimal synthetic method, system, device and medium for hydrogels, which can improve the prediction accuracy of the adsorption coefficient of hydrogels and quickly screen the optimal synthesis conditions of hydrogels.

[0060] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0061] As Figure 1 shown, an optimal synthetic method for hydrogels provided in the embodiment of the present invention includes:

[0062] Step 1: Establish a data set for the hydrogel to adsorb a heavy metal; the input features of the data set include adsorbent properties, synthesis conditions, heavy metal properties and adsorption conditions, and the label of the data set is the logarithm of the adsorption coefficient of the hydrogel to adsorb the heavy metal.

[0063] Taking heavy metal lead (II) as an example, first, a data set based on the literature related to the adsorption of lead (II) by hydrogels is mined and established, where the input features include: 1) Properties of the adsorbent: specific surface area (BET, m 2 / g), pore volume (V t, cm3 / g), pore size (Size, μm), swelling ratio (SR, %), mechanical strength (MS, MPa); 2) Synthesis conditions: monomer functional groups, types of nanomaterials, synthesis time (R_time, h) and temperature (R_tem, °C), types and dosages of initiators (M in , mol%), types and dosages of crosslinkers (M Cross , mol%); 3) Properties of heavy metals: first ionization energy (IE, kJ / mol), ionic radius (Radius, nm) and hydrated ionic radius (Hydra_R, nm); 4) Adsorption conditions: initial concentration of heavy metals (C0, ppm) and equilibrium concentration (C e , ppm), dosage of hydrogel (M hydro , g / L), adsorption oscillation frequency (Stirring, rpm), types of nanomaterials, adsorption time (A_time, h), equilibrium time (CT, min), solution pH (pH) and temperature (A_Tem, °C).

[0064] The output (label) is the logarithm of the ratio of the adsorption capacity of the hydrogel for heavy metals to the equilibrium concentration of heavy metals (logK d , L / g). K d is the adsorption coefficient. The determination method of the output is as follows: Conduct a comprehensive literature search to obtain relevant data on the adsorption of heavy metal lead (II) by hydrogels. First, collect adsorption isotherms from the literature, and then obtain the corresponding adsorption capacity (Qe) and equilibrium concentration (C e ) data points of heavy metals within the reported equilibrium concentration range from the isotherms. Among them, the abscissa of the adsorption isotherm is the equilibrium concentration C e of metal ions, and the ordinate is the adsorption capacity Q e of the hydrogel for adsorbing heavy metals. K d is calculated as shown in formulas (1) and (2):

[0065] Q e = (C0 - C e )V / m (1)

[0066] K d = (C0 - C e ) / C e × V / m = Q e / C e (2)

[0067] V is the volume (mL) of the metal solution, and m is the mass (g) of the adsorbent.

[0068] Step 2: Preprocess the input features of the dataset; the preprocessing includes filling missing values in the input features using the random forest algorithm.

[0069] Dummy variable processing is performed on the input features for ML modeling to unify the units of each input feature variable, and the text variables in the dataset are converted into numerical variables; Spearman's coefficient is used to analyze the correlation of features, and features are eliminated to avoid variable redundancy; data preprocessing methods such as using the random forest algorithm to fill in the missing samples in the features are used to improve the performance of the ML model.

[0070] Step 3: Use the preprocessed dataset to train multiple machine learning models respectively, and select the trained machine learning model with the best performance as the optimal adsorption prediction model; the multiple machine learning models include a random forest model, a support vector machine model, a gradient boosting decision tree model, and an XGBoost model.

[0071] The preprocessed dataset is split into a training set and a test set in a ratio of 8:2, and 10-fold cross-validation is performed on the training dataset, where the training set is further divided into a pre-training set and a validation set in a ratio of 9:1.

[0072] The pre-training set, the validation set, and the test set are used to construct ML mathematical models based on the random forest (RF), support vector machine (SVM), gradient boosting decision tree (GBDT), and XGBoost algorithms, and the Bayesian optimization algorithm is selected to optimize and adjust the hyperparameters of the ML model. By determining the coefficient (R 2 ), root mean square error (RMSE), and mean absolute error (MAE), the three evaluation indicators are used to screen out the model with the best performance.

[0073] Step 4: Use the optimal adsorption prediction model to calculate the SHAP values of each input feature for the adsorption coefficient, and select the input features corresponding to the preset number of SHAP values from largest to smallest.

[0074] After the model is developed, the SHAP method is used to calculate the Shapley values of each feature to quantify the contribution of the input features to the hydrogel adsorption performance. The working principle of the SHAP method is to check the prediction differences before and after deleting the features. By including all possible ways of deleting features, the function-to-function interaction information is considered. The mean absolute Shapley (MAS) value of each input feature is calculated based on all data points. The MAS values of each feature are sorted by size to obtain the six more important features among the input features: synthesis time, synthesis temperature, initiator dosage, crosslinker dosage, hydrogel dosage, and oscillation frequency.

[0075] Step 5: Set multiple numerical values for each selected input feature, and arbitrarily combine the numerical values of all selected input features to form different hydrogel preparation conditions.

[0076] Set reasonable value ranges (Mhydro : 0.1 - 2 g / L; R_tem: 35 - 60 °C; R_time: 1 - 10 h; M cross : 0.5 - 2 mol%; M in : 0.5 - 1.5 mol%; Stirring: 140 - 160 rpm), and set multiple values for each selected input feature within a reasonable range of values.

[0077] Step 6: Input each hydrogel preparation condition into the optimal adsorption prediction model, output the logarithm of the adsorption coefficient of the hydrogel for adsorbing the heavy metal under each hydrogel preparation condition, and determine the hydrogel preparation condition corresponding to the maximum logarithm of the adsorption coefficient as the best preparation condition for the hydrogel used to adsorb the heavy metal.

[0078] Screen out the most important synthesis features according to the model results and input them into the optimal adsorption prediction model to predict the adsorption coefficient of the hydrogel, and obtain the corresponding optimal conditions for hydrogel synthesis.

[0079] Input the hydrogel preparation conditions in Step 5 into the optimal adsorption prediction model to predict the adsorption coefficient, and obtain the best preparation conditions for synthesizing the hydrogel (M hydro : 0.1 g / L; R_tem: 55 °C; R_time: 20 h; M in : 0.5 mol%; Stirring: 120 rpm; M cross : 2 mol%).

[0080] The following combines Figures 2 to 15 to further clarify the synthesis method of the present invention.

[0081] Figure 2 and Figure 3 and Figure 4 show the use of RF to fill in missing data values and perform correlation screening on features, adjust the model performance, and obtain an optimized model. Figure 5 and Figure 6 and Figure 7 show three evaluation indicators (R 2 , MSE, and RMSE) of the ML model under four algorithms. By comparison, it is found that the XGBoost model achieves the best performance and is therefore selected for further discussion. Figures 5 to 7 The abscissa Predict model of

[0082] Figure 8 and Figure 9The comparison of the model performance before and after filling the dataset of the synthetic hydrogel's adsorption of heavy metals shows that using RF to fill in the missing values does not affect the model performance, and the results show that the XGB model has high reliability and robustness, where the R 2 , RMSE, and MAE of the test dataset are 0.956, 0.179, and 0.097 respectively. As Figure 10 and 11 show, filling in the missing values with RF is beneficial to improving the rationality of the importance feature ranking. For adsorption time, pH, -COOH, Reactiontime, Radius, the points with low feature values are mainly on the left side, while the points with high feature values are mainly on the right side, indicating that these feature values are positively correlated with the predicted value - higher feature values are beneficial for adsorption. On the contrary, for logC e , synthesis temperature, oscillation frequency, -RCONH2, and -OH, the points with higher values are usually distributed on the left side, showing a negative correlation with the predicted value for adsorption. Figure 8 and Figure 9 The abscissa y_ture represents the true value of logK d , and the ordinate y_predict represents the predicted value of logK d . Figure 10 and Figure 11 The SHAP value(miss) in Figure 10 and Figure 11 represents the SHAP values corresponding to different input features under the condition that the calculated dataset has missing values. Feature value_miss represents the samples of different input features under the condition that the calculated dataset has missing values. SHAPvalue(fill) represents the SHAP values corresponding to different input features after the missing values in the calculated dataset are filled. Feature value_fill represents the samples of different input features after the missing values in the calculated dataset are filled. Low represents low, and high represents high.

[0083] Subsequently, the average absolute Shapley (MAS) value of each input feature was calculated based on all data points to quantify the overall importance of each descriptor (as Figure 12 shown), and it was found that the changing trend of the SHAP values of the samples of 20 features ( Figure 15 ) was consistent with the Pearson correlation coefficient between each descriptor predicted by Figure 14 and the adsorption coefficient. The feature types were divided into four parts: adsorption system, properties of the adsorbent, heavy metal properties, and types of monomer functional groups, and it was found that the overall feature ranking follows the order of adsorption system > types of monomer functional groups > properties of the adsorbent > heavy metal properties. As Figure 13 shown, visual analysis of the top ten features in the overall ranking gives the ranking result as logC e > -COOH > M hydro > C0 > SR > Vt >R_tem>M Cross >A_time>Stirring. The above results show that the monomer functional groups, adsorbent dosage, swelling ratio of the material, pore volume, synthesis temperature, synthesis time, amounts of initiator and crosslinking agent are the main important factors determining the heavy metal adsorption of the hydrogel. Based on the above conclusions, reasonable value ranges are set for five characteristics including adsorbent dosage, synthesis temperature, synthesis time, amounts of initiator and crosslinking agent, and they are input into the pre-established XGBoost model to predict the adsorption coefficients of the material under different corresponding preparation conditions. Finally, the optimal preparation conditions for synthesizing the hydrogel for heavy metal adsorption are obtained (M hydro : 0.1 g / L; R_tem: 55 °C; R_time: 20 h; M in : 0.5 mol%; M cross : 2 mol%). Figure 12 In it, Absolute SHAP value represents the absolute SHAP value, Adsorption system represents the adsorption system, Synthesis of adsorbent represents the synthesis of the adsorbent, Heavy metal represents heavy metals, and Functional groups of monomers represents the monomer functional groups. Figure 15 In it, Feature value represents the feature value.

[0084] The method of the present invention first mines and establishes a data set based on the literature related to the hydrogel adsorption of Pb(II), preprocesses the obtained data to solve the potential data imbalance problem of the source data; selects the influencing factors and data samples most relevant to the prediction target to construct an ML model; finally, quantitatively analyzes the important factors based on the results of multivariable interactions to obtain the corresponding optimal synthesis method for the hydrogel to adsorb lead (II). The regression model established based on ML in the method of the present invention effectively predicts the heavy metal adsorption of the nanocomposite hydrogel, realizing the rational design and efficient functional preparation of the hydrogel. Compared with the traditional experimental exploration, introducing ML into the guided synthesis of materials can speed up the exploration process and reduce costs, providing an effective solution for the efficient selective adsorption and green recycling of heavy metals. This method does not involve toxic and costly chemical reagents. The data set obtained by filling in the missing values through the random forest (RF) algorithm is used to construct the ML model, which helps to further improve the prediction accuracy of the given modeling algorithm.

[0085] The embodiment of the present invention also provides a hydrogel optimal synthesis system, including:[[]]

[0086] A dataset establishment module for establishing a dataset of a hydrogel adsorbing a heavy metal; the input features of the dataset include adsorbent properties, synthesis conditions, heavy metal properties, and adsorption conditions, and the label of the dataset is the adsorption coefficient of the hydrogel adsorbing the heavy metal.

[0087] A preprocessing module for preprocessing the input features of the dataset; the preprocessing includes using a random forest algorithm to fill in the missing values in the input features.

[0088] A training module for training multiple machine learning models respectively using the preprocessed dataset, and selecting the trained machine learning model with the best performance as the optimal adsorption prediction model; the multiple machine learning models include a random forest model, a support vector machine model, a gradient boosting decision tree model, and an XGBoost model.

[0089] A feature selection module for calculating the SHAP values of each input feature for the adsorption coefficient using the optimal adsorption prediction model, and selecting the input features corresponding to a preset number of SHAP values from large to small.

[0090] A preparation condition composition module for setting multiple values for each selected input feature, and combining the values of all selected input features arbitrarily to form different hydrogel preparation conditions.

[0091] A prediction module for inputting each hydrogel preparation condition into the optimal adsorption prediction model, outputting the logarithm of the adsorption coefficient of the hydrogel adsorbing the heavy metal under each hydrogel preparation condition, and determining the hydrogel preparation condition corresponding to the maximum logarithm of the adsorption coefficient as the best preparation condition of the hydrogel for adsorbing the heavy metal.

[0092] The hydrogel optimal synthesis system provided by the embodiments of the present invention is similar to the hydrogel optimal synthesis method described in the above embodiments in terms of working principle and beneficial effects, so it will not be elaborated here. For specific content, reference can be made to the introduction of the above method embodiments.

[0093] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the hydrogel optimal synthesis method as described above.

[0094] In addition, when the computer program in the above-mentioned memory is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs that can store program codes.

[0095] Further, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, it implements the optimal hydrogel synthesis method as described above.

[0096] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0097] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. An optimal synthesis method of hydrogel, characterized in that, Including: Establishing a dataset of a hydrogel adsorbing a heavy metal; the input features of the dataset include adsorbent properties, synthesis conditions, heavy metal properties, and adsorption conditions, and the label of the dataset is the logarithm of the adsorption coefficient of the hydrogel adsorbing the heavy metal; Preprocessing the input features of the dataset; The preprocessing includes using a random forest algorithm to fill in the missing values in the input features; Training multiple machine learning models respectively with the preprocessed dataset, and selecting the trained machine learning model with the best performance as the optimal adsorption prediction model; the multiple machine learning models include a random forest model, a support vector machine model, a gradient boosting decision tree model, and an XGBoost model; Calculating the SHAP values of each input feature on the adsorption coefficient using the optimal adsorption prediction model, and selecting the input features corresponding to a preset number of SHAP values from largest to smallest; Setting multiple values for each selected input feature, and making any combination of the values of all selected input features to form different hydrogel preparation conditions; Inputting each hydrogel preparation condition into the optimal adsorption prediction model, outputting the logarithm of the adsorption coefficient of the hydrogel adsorbing the heavy metal under each hydrogel preparation condition, and determining the hydrogel preparation condition corresponding to the largest logarithm of the adsorption coefficient as the best preparation condition of the hydrogel for adsorbing the heavy metal.

2. The optimal synthesis method of the hydrogel according to claim 1, characterized in that, The adsorbent properties include: specific surface area, pore volume, pore diameter, swelling ratio, and mechanical strength; The synthesis conditions include: monomer functional group, type of nanomaterial, synthesis time, synthesis temperature, type and dosage of initiator, type of crosslinker, and dosage of crosslinker; The heavy metal properties include: first ionization energy, ionic radius, and hydrated ionic radius; The adsorption conditions include: initial concentration of heavy metal, equilibrium concentration of heavy metal, hydrogel dosage, adsorption oscillation frequency, adsorption time, equilibrium time, solution pH, and temperature.

3. The optimal synthesis method of the hydrogel according to claim 2, wherein The process of establishing the label is: Collect the adsorption isotherm of the hydrogel for adsorbing the heavy metal; the abscissa of the adsorption isotherm is the equilibrium concentration C of the heavy metal e , and the ordinate is the adsorption capacity Q of the hydrogel for adsorbing the heavy metal e ; Obtaining the data points corresponding to the equilibrium concentration values of the heavy metal in the input features from the adsorption isotherm; Using the formula logK d = log(Q e / C e ), calculate the logarithm of the adsorption coefficient corresponding to each data point; where K d represents the adsorption coefficient, and logK d represents the logarithm of the adsorption coefficient.

4. The optimal synthesis method of the hydrogel according to claim 1, characterized in that, Preprocessing the input features of the dataset, specifically including: Performing dummy variable processing on the input features to convert the text variables in the dataset into numerical variables; Using the Spearman coefficient to perform correlation analysis on the input features after dummy variable processing, and removing redundant features; Using a random forest algorithm to fill in the missing values in the input features after removing redundant features.

5. The optimal synthesis method of the hydrogel according to claim 1, characterized in that, The process of training multiple machine learning models respectively with the preprocessed dataset specifically includes: Splitting the preprocessed dataset into a training set and a test set in a ratio of 8:2, and performing ten-fold cross-validation on the training set, so that the training set is divided into a pre-training set and a validation set in a ratio of 9:1; Training each machine learning model using the pre-training set, validation set, and test set, and selecting the Bayesian optimization algorithm to optimize and adjust the hyperparameters of the machine learning model to obtain multiple trained machine learning models.

6. The optimal synthesis method of the hydrogel according to claim 1, wherein, Calculating the SHAP values of each input feature on the adsorption coefficient using the optimal adsorption prediction model, and selecting the input features corresponding to a preset number of SHAP values from largest to smallest, specifically including: According to the preprocessed dataset, calculate the average absolute Shapley value of each input feature for the adsorption coefficient using the optimal adsorption prediction model; Arrange the average absolute Shapley values of each input feature for the adsorption coefficient in descending order, and select the input features corresponding to the preset number of average absolute Shapley values from largest to smallest.

7. A hydrogel optimal synthesis system, characterized in that, Including: A dataset establishment module for establishing a dataset of a hydrogel adsorbing a heavy metal; the input features of the dataset include adsorbent properties, synthesis conditions, heavy metal properties, and adsorption conditions, and the label of the dataset is the logarithm of the adsorption coefficient of the hydrogel adsorbing the heavy metal; A preprocessing module for preprocessing the input features of the dataset; The preprocessing includes using a random forest algorithm to fill in the missing values in the input features; A training module for training multiple machine learning models respectively using the preprocessed dataset, and selecting the best-performing trained machine learning model as the optimal adsorption prediction model; the multiple machine learning models include a random forest model, a support vector machine model, a gradient boosting decision tree model, and an XGBoost model; A feature selection module for calculating the SHAP value of each input feature for the adsorption coefficient using the optimal adsorption prediction model, and selecting the input features corresponding to the preset number of SHAP values from largest to smallest; A preparation condition composition module for setting multiple values for each selected input feature and arbitrarily combining the values of all selected input features to form different hydrogel preparation conditions; A prediction module for inputting each hydrogel preparation condition into the optimal adsorption prediction model, outputting the logarithm of the adsorption coefficient of the hydrogel adsorbing the heavy metal under each hydrogel preparation condition, and determining the hydrogel preparation condition corresponding to the largest logarithm of the adsorption coefficient as the best preparation condition of the hydrogel for adsorbing the heavy metal.

8. An electronic device, characterized in that, Including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the hydrogel optimal synthesis method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, Stored thereon is a computer program, and when the computer program is executed, it implements the hydrogel optimal synthesis method according to any one of claims 1 to 6.