Method for predicting antibiotic degradation rate in mushroom dreg hydrothermal process based on data driving
By constructing a data-driven machine learning model and optimizing hydrothermal parameters using dependency graphs, the problems of low efficiency and high cost in traditional methods were solved. This enabled accurate prediction of antibiotic degradation rates and parameter optimization, improving the efficiency and precision of bacterial residue treatment.
Patent Information
- Application Number
- CN202511694655.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-10
AI Technical Summary
Traditional methods for optimizing antibiotic degradation are inefficient and costly. Changes in hydrothermal process parameters affect degradation performance, and there is a lack of rapid and accurate prediction methods.
Using a data-driven approach, we constructed extreme gradient boosting, random forest, and K-nearest neighbor models, and optimized hydrothermal process parameters through single-factor and two-factor dependency graphs to predict antibiotic degradation rates.
This method enables accurate prediction and parameter optimization of antibiotic degradation rate during the hydrothermal process of bacterial residue, reducing experimental costs and improving process optimization efficiency.
Smart Images

Figure CN121506259A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of solid waste disposal and resource utilization technology, specifically relating to a data-driven method for predicting antibiotic degradation rate in the hydrothermal process of bacterial residue. Background Technology
[0002] Antibiotic fermentation residue (commonly known as "microbial residue") is solid biomass separated from the filtration and extraction processes during antibiotic production using microbial fermentation. For every ton of antibiotics produced, typically 8-10 tons of wet residue are generated. This residue contains a large amount of mycelium, residual culture medium, sugars, proteins, and 0.03%-0.1% active antibiotics. Therefore, timely and effective treatment of the generated wet residue is necessary. Traditional disposal methods include incineration or landfill, but the high moisture content and high organic load result in high energy consumption for incineration and large land requirements for landfill, while also posing risks of antibiotic leakage and the spread of antibiotic resistance genes. In recent years, hydrothermal technology has been used to simultaneously degrade residual antibiotics and recover organic matter. However, parameter changes during the hydrothermal process can affect the final degradation efficiency of antibiotics, posing risks to the recycling and treatment of microbial residue and hindering related research. Traditional methods typically involve adjusting parameters based on experience and trial-and-error experiments to optimize antibiotic degradation. However, this approach is inefficient and incurs high experimental costs due to trial and error. Therefore, rapidly and accurately predicting the relationship between parameters and antibiotic degradation rates during hydrothermal processes is of great significance for real-time monitoring and process optimization of bacterial residue treatment. Summary of the Invention
[0003] The purpose of this invention is to address the problems of low efficiency and high cost in traditional antibiotic degradation optimization methods. It proposes a data-driven method for predicting antibiotic degradation rate in the hydrothermal process of bacterial residue, and optimizes the parameters of the hydrothermal process in reverse based on the predicted antibiotic degradation rate.
[0004] The technical solution adopted by this invention to solve the above-mentioned technical problems is: a data-driven method for predicting the antibiotic degradation rate in the hydrothermal process of bacterial residue, the method specifically including the following steps:
[0005] Step 1: Collect raw material parameters of different fungal residues, as well as process operation parameters and degradation rate data of different fungal residues under different hydrothermal conditions, and divide the collected data into two parts: training set and test set;
[0006] Step 2: Train each prediction model using the training set, then evaluate the prediction performance of each prediction model on the test set, and select the best prediction model from among the models based on the prediction performance.
[0007] Step 3: Generate single-factor dependency plots and two-factor dependency plots based on the best prediction model;
[0008] Step 4: Adjust the input parameters of the optimal prediction model based on the single-factor dependency plot and two-factor dependency plot generated in Step 3. Then, determine the optimal raw material parameters and process operation parameters in the hydrothermal process based on the prediction results of the optimal prediction model for antibiotic degradation rate under different input parameters.
[0009] Furthermore, the raw material parameters of the bacterial residue include bacterial residue quality, moisture content, type of antibiotic, and pH value of the initial environment.
[0010] Furthermore, the process operating parameters include the volume of the reaction apparatus, residence time, reaction temperature, heating rate, and acid concentration.
[0011] Furthermore, the various prediction models include the extreme gradient boosting model, the random forest model, and the K-nearest neighbor model;
[0012] The inputs to the prediction model include raw material parameters of the bacterial residue and process operation parameters under hydrothermal conditions. The output of the prediction model is degradation rate data.
[0013] Furthermore, the evaluation of the prediction performance of each prediction model on the test set, and the selection of the best prediction model from among the models based on the prediction performance, specifically involves:
[0014] Evaluate the coefficient of determination of each prediction model on the test set, and select the prediction model with the largest coefficient of determination as the best prediction model.
[0015] Furthermore, the prediction results of the extreme gradient boosting model are as follows:
[0016]
[0017] in, Indicates the first decision trees for the first Predicted degradation rate for each sample Indicates the total number of decision trees. This indicates that the extreme gradient boosting model applies to the first... Degradation rate prediction results for each sample.
[0018] Furthermore, the prediction results of the random forest model are as follows:
[0019]
[0020] in, Indicates the random forest model for the th Predicted degradation rate for each sample The total number of trees, For the first tree to the first Predicted degradation rate for each sample.
[0021] Furthermore, the prediction result of the K-nearest neighbor model is as follows:
[0022]
[0023] in, Indicates the first A set consisting of the K neighbors of a sample. Indicates the first Predicted degradation rate of each neighboring unit. The K-nearest neighbor model represents the first... Predicted degradation rate for each sample.
[0024] Furthermore, the specific process of step four is as follows:
[0025] Step 4.1. For the parameter residence time, first obtain the maximum dependency value in the single-factor PDP plot corresponding to the residence time, and record the residence time corresponding to the maximum dependency value as . The first interval of stay is then determined as ;
[0026] in, This indicates half the length of the set dwell time interval;
[0027] For the parameter reaction temperature, first obtain the maximum dependency value in the single-factor PDP plot corresponding to the reaction temperature, and record the residence time corresponding to the maximum dependency value as . The first reaction temperature range determined is: ;
[0028] in, This indicates half the length of the set reaction temperature range;
[0029] For the initial pH value of the parameter environment, first obtain the maximum dependence value in the single-factor PDP plot corresponding to the initial pH value, and record the pH value corresponding to the maximum dependence value as . The first pH range determined is: ;
[0030] in, This indicates half the length of the set pH range;
[0031] For the parameter of mushroom residue quality, firstly, obtain the maximum dependency value in the single-factor PDP plot corresponding to the mushroom residue quality, and record the mushroom residue quality corresponding to the maximum dependency value as . The first quality range of the fungal residue is then determined as follows: ;
[0032] in, This indicates half the length of the set range of mushroom residue quality.
[0033] For the heating rate parameter, first obtain the maximum dependency value in the single-factor PDP plot corresponding to the heating rate, and denote the heating rate corresponding to the maximum dependency value as . The first heating rate range determined is: ;
[0034] in, This indicates half the length of the set heating rate range;
[0035] For the parameter acid concentration, first obtain the maximum dependence value in the single-factor PDP plot corresponding to the acid concentration, and denote the acid concentration corresponding to the maximum dependence value as . The first acid concentration range determined is: ;
[0036] in, This indicates half the length of the set acid concentration range;
[0037] For the parameter moisture content, if the maximum dependence value in the single-factor PDP plot of the parameter moisture content corresponds to a moisture content of 0, then the first moisture content interval is... ;
[0038] in, Indicates the length of the set moisture content range;
[0039] For the parameter reactor volume, first obtain the maximum dependency value in the single-factor PDP plot corresponding to the reactor volume, and denote the reactor volume corresponding to the maximum dependency value as . The determined volume range of the first reaction device is: ;
[0040] in, This indicates half the length of the defined reaction device volume range;
[0041] Step 4.2: Determine the second range of parameters such as residence time, reaction temperature, initial ambient pH, residue mass, heating rate, acid concentration, moisture content, and reaction device volume based on the two-factor PDP diagram.
[0042] Taking residence time as an example, this invention first obtains two-factor PDP plots for residence time in relation to reaction temperature, initial ambient pH, substrate mass, heating rate, acid concentration, moisture content, and reactor volume. Then, it finds the maximum value in each two-factor PDP plot and obtains the corresponding parameter values for each maximum value in its own two-factor PDP plot, i.e., obtaining the following sets of parameter values: (residence time, reaction temperature), (residence time, initial ambient pH), (residence time, substrate mass), (residence time, heating rate), (residence time, acid concentration), (residence time, moisture content), and (residence time, reactor volume).
[0043] Taking the dwell time value in each set of parameter values as the center, similar to step 4.1, expand outwards to the left and right of the center point. The length is then calculated, and the intersection of the 7 extended results is taken as the second dwell time interval.
[0044] Step 43: Find the intersection of the first and second intervals of the dwell time, and use the result of the intersection as the adjustment range of the dwell time;
[0045] Similarly, the adjustment ranges for reaction temperature, initial ambient pH, bacterial residue mass, heating rate, acid concentration, moisture content, and reaction device volume were obtained respectively.
[0046] Step 4: Set the adjustment step size for each parameter, select parameters within the determined adjustment range, input each parameter combination into the best prediction model, take the maximum degradation rate value output by the best prediction model as the optimal solution rate value, and take the parameter combination corresponding to the optimal solution rate value as the optimal parameter combination.
[0047] The beneficial effects of this invention are:
[0048] This invention provides a machine learning-based method for predicting antibiotic degradation rate during the hydrothermal treatment of fungal residue. Based on the predicted antibiotic degradation rate, parameters in the fungal residue hydrothermal treatment process can be optimized and controlled. First, experimental parameters and results of the fungal residue hydrothermal process are collected to construct a numerical raw experimental dataset. Then, extreme boosting, random forest, and K-nearest neighbor methods are used to construct a machine learning prediction model for the antibiotic degradation rate during the fungal residue hydrothermal process. The optimal model is selected based on evaluation indicators, thereby achieving accurate prediction of the antibiotic degradation rate under different parameters. A single-factor PDP plot represents the effect of each parameter on the degradation rate, and a two-factor PDP plot characterizes the synergistic effect between the parameters. Finally, the parameters of the fungal residue hydrothermal treatment process are optimized based on the prediction results and the PDP plot. Attached Figure Description
[0049] Figure 1This is a flowchart of a data-driven method for predicting antibiotic degradation rate in a bacterial residue hydrothermal process according to the present invention.
[0050] Figure 2(a) shows the effect of the XGB model during the training process;
[0051] Degradation real represents the actual degradation rate, and Degradation Predictions represent the predicted degradation rate.
[0052] Figure 2(b) shows the effect of the RF model during the training process;
[0053] Figure 2(c) shows the effect of the KNN model during the training process;
[0054] Figure 2(d) shows the effect of the XGB model during the testing process;
[0055] Figure 2(e) shows the effect of the RF model during the testing process;
[0056] Figure 2(f) shows the performance of the KNN model during the testing process;
[0057] Figure 3 This is a diagram illustrating the importance of variables in the Shap method analysis and prediction process;
[0058] In the figure, t represents residence time, T represents reaction temperature, Class represents antibiotic type, PH represents initial pH value, m represents residue mass, HR represents heating rate, Conv represents acid concentration, Water content represents moisture content, and V represents reactor volume. Taking residence time as an example, each point in the row of residence time in the left figure represents the importance of residence time in each sample of the dataset. Taking residence time as an example, the absolute value of the importance of residence time in each sample is taken, and then the average of all absolute values corresponding to residence time is calculated. The right figure shows that the average value corresponding to residence time is 0.16.
[0059] Figure 4(a) is a graph showing the dependence of antibiotic degradation rate on residence time;
[0060] partial dependencies represent dependencies;
[0061] Figure 4(b) shows the dependence of antibiotic degradation rate on reaction temperature;
[0062] Figure 4(c) shows the dependence of antibiotic degradation rate on different antibiotic structures;
[0063] Figure 4(d) shows the dependence of antibiotic degradation rate on the initial pH value;
[0064] Figure 4(e) shows the dependence of antibiotic degradation rate on the quality of the bacterial residue;
[0065] Figure 4(f) shows the dependence of antibiotic degradation rate on heating rate;
[0066] Figure 4(g) shows the dependence of antibiotic degradation rate on acid concentration;
[0067] Figure 4(h) is a graph showing the dependence of antibiotic degradation rate on water content;
[0068] Figure 4(i) is a graph showing the dependence of antibiotic degradation rate on the volume of the reaction apparatus;
[0069] Figure 5(a) is a two-factor partial dependence plot of reaction temperature and residence time;
[0070] Figure 5(b) is a two-factor partial dependence diagram of the pH value of the initial environment and the residence time;
[0071] Figure 5(c) is a two-factor partial dependence diagram of the relationship between the quality of the fungal residue and the retention time;
[0072] Figure 5(d) is a two-factor partial dependence plot of heating rate and residence time;
[0073] Figure 5(e) is a two-factor partial dependence plot of the initial environment pH value versus the reaction temperature;
[0074] Figure 5(f) is a two-factor partial dependence diagram of the quality of the fungal residue on the reaction temperature;
[0075] Figure 5(g) is a two-factor partial dependence plot of heating rate on reaction temperature;
[0076] Figure 5(h) is a two-factor partial dependence diagram of the quality of the substrate residue and the pH value of the initial environment;
[0077] Figure 5(i) is a two-factor partial dependence plot of heating rate on the initial pH value;
[0078] Figure 5(j) is a two-factor partial dependence diagram of heating rate and substrate quality;
[0079] Figure 6(a) shows the comparison between the predicted and actual values of the degradation rate of spiramycin bacterial residue;
[0080] Figure 6(b) shows the comparison between the predicted and actual values of the degradation rate of oxytetracycline residue. Detailed Implementation
[0081] Specific implementation method one: Combining Figure 1 This embodiment describes a data-driven method for predicting antibiotic degradation rate in a bacterial residue hydrothermal process. The method specifically includes the following steps:
[0082] Step 1: Conduct a literature search on the ScienceNet website, and collect raw material parameters of different fungal residues, as well as process operation parameters and degradation rate data of different fungal residues under different hydrothermal conditions based on the literature search results.
[0083] The raw material parameters of the mushroom residue include the quality of the mushroom residue, moisture content, type of antibiotics, and pH value of the initial environment;
[0084] Process operating parameters include reactor volume, residence time, reaction temperature, heating rate, and acid concentration;
[0085] For example, for a type of fungal residue, the raw material parameters of the fungal residue and the process parameters of the fungal residue in a single hydrothermal process are collected, and the degradation rate data of the fungal residue after the current hydrothermal process are also collected. The collected raw material parameters, process parameters and degradation rate data can be combined into a set of data. By repeating the above process, all the sets of data required for model training and testing can be obtained, wherein the ratio of the number of data sets in the training set to the number of data sets in the testing set is 8:2.
[0086] Step 2: Train each prediction model using the training set. The prediction models include Extreme Gradient Boosting (XGB), Random Forest (RF), and K Nearest Neighbors (KNN). The model training is based on hyperparameters. For example, the hyperparameters of the Extreme Gradient Boosting model include learning rate, maximum depth, number of subtrees, subsampling, column sampling, and random state. The input of the prediction model includes the raw material parameters of the fungal residue and the process operation parameters under hydrothermal conditions. The output of the prediction model is the degradation rate data.
[0087] The prediction results of the extreme gradient boosting model are as follows:
[0088]
[0089] in, Indicates the first decision trees for the first Predicted degradation rate for each sample (each data set is considered as one sample). Indicates the total number of decision trees. This indicates that the extreme gradient boosting model applies to the first... Degradation rate prediction results for each sample.
[0090] The loss function during the training of the extreme gradient boosting model is:
[0091]
[0092]
[0093] in, This represents the loss function, used to measure the performance of the model before... Predicted value of step Compared with the true value The differences between them It can be seen as the former Step model for the first The predicted value for each sample, It is in the The new decision tree added in step 1 is for the first step The predicted value for each sample. is the regularization term, used to control the complexity of the model and prevent overfitting. T is the number of leaf nodes; the regularization term is proportional to the number of leaf nodes and is used to control the intensity of the penalty. The more leaf nodes, the more complex the model. It is a leaf node The weights are used to further limit the complexity of the model by penalizing the sum of squares of the weights of the leaf nodes, making the model smoother and reducing the risk of overfitting.
[0094] The prediction results of the random forest model are as follows:
[0095]
[0096] in, Indicates the random forest model for the th Predicted degradation rate for each sample The total number of trees, For the first tree to the first Predicted degradation rate for each sample.
[0097] The prediction results of the K-nearest neighbor model are as follows:
[0098]
[0099] in, Indicates the first The set consisting of the K neighbors of the nth sample (using the nth sample) The raw material parameters and process parameters of the first sample are combined to form a vector. Then, the vector corresponding to each of the other samples is calculated and combined with the vector of the first sample. The Euclidean distance between the vectors corresponding to the n samples is used to select the K samples with the smallest Euclidean distance as the nth sample. (neighbors of each sample) Indicates the first Predicted degradation rate of each neighboring unit. The K-nearest neighbor model represents the first... Predicted degradation rate for each sample.
[0100] The predictive performance of each prediction model on the test set is then evaluated. Based on the predictive performance, the best prediction model is selected from the models. The main evaluation factor is the coefficient of determination of the prediction results of each prediction model on the test set. The prediction model with the largest coefficient of determination is selected as the best prediction model.
[0101] Figures 2(a) to 2(f) The predictive performance of three models—XGB, RF, and KNN—on antibiotic degradation rates is presented. XGB and KNN performed best during training, with each point concentrated on the 1:1 line (as shown in Figures 2(a), 2(b), and 2(c)). During testing, XGB performed best, with each point evenly distributed on both sides of the 1:1 line (test results are shown in Figures 2(d), 2(e), and 2(f)). This demonstrates that the XGB model has good performance in predicting antibiotic degradation rates, as can be seen from the R-squared values of the three models in Table 1. 2 The MAER and MSE values show that the XGB model has good test performance. 2 The value is 0.923, which is much higher than the 0.877 and 0.852 of the RF and KNN models, respectively.
[0102] Table 1 Performance metrics of three different machine learning models (XGB, RF, KNN)
[0103]
[0104] Meanwhile, to further explore the marginal impact mechanism of key parameters on antibiotic degradation rate, this invention employs Partial Dependence Diagrams (PDPs) for visualization analysis. PDPs, while keeping other parameter values constant, can investigate the marginal effect of target parameters on model prediction results, providing a clear theoretical basis for optimizing parameters in the hydrothermal process of bacterial residue.
[0105] Step 3: Generate single-factor dependency plots and two-factor dependency plots based on the best prediction model. Specifically, you can select a set of data corresponding to the maximum degradation rate obtained in Step 1, fix the other parameter values in this set, and change only the single parameter value to obtain a single-factor PDP plot. Similarly, fix the other parameter values in this set and change only the two parameter values to obtain a two-factor PDP plot.
[0106] Figures 4(a) to 4(i)The influence of key parameters on the antibiotic degradation rate in the hydrothermal process was illustrated using single-factor PDP plots. Figures 4(a) and 4(b) show that the antibiotic degradation rate increases with increasing reaction temperature and residence time. Figure 4(c) shows the nonlinear relationship between different antibiotic structures and the antibiotic degradation rate. Figure 4(d) shows the relationship between the initial pH value and the antibiotic degradation rate; it can be seen that the closer the initial pH value is to neutral, the more favorable it is for antibiotic degradation. Figure 4(e) shows the relationship between the mass of the bacterial residue and the antibiotic degradation rate; it can be seen that more bacterial residue is not necessarily better, and on the contrary, excessive mass inhibits antibiotic degradation. Figure 4(f) shows the relationship between the heating rate and the antibiotic degradation rate; it can be seen that when the heating rate reaches a certain level, further increasing the heating rate is detrimental to antibiotic degradation. Figure 4(g) shows the nonlinear relationship between acid concentration and the antibiotic degradation rate. Figure 4(h) shows the nonlinear relationship between moisture content and the antibiotic degradation rate; it can be seen that lower moisture content is more favorable for antibiotic degradation. Figure 4(i) shows the nonlinear relationship between the volume of the reaction device and the antibiotic degradation rate. It can be seen that the larger the volume of the reaction device, the more beneficial it is to antibiotic degradation.
[0107] Figures 5(a) to 5(j) Two-factor PDP plots illustrate the synergistic effects of pairwise parameters on antibiotic degradation. It should be noted that only partial results of the synergistic effects between pairwise parameters are shown here, not the entire picture. As shown in Figure 5(a), the degradation capacity is maximized when the reaction time exceeds approximately 75 min and the temperature is 150–250 °C. Similar phenomena were observed in Figures 5(b) to 5(d) under the combined influence of reaction time and other key parameters. Furthermore, as shown in Figures 5(e) to 5(g), the changes in degradation capacity related to the heating rate are more pronounced when the reaction temperature and substrate quality are increased. As shown in Figures 5(h) and 5(i), the dependence of degradation capacity on reaction temperature shows a consistent trend across different initial environmental pH, heating rate, and substrate quality: a slow initial increase followed by a rapid increase. As shown in Figure 5(j), the optimal reaction temperature for improving antibiotic degradation kinetics is approximately 200 °C. Finally, controlling the initial environmental pH in the reaction environment to neutral conditions is beneficial for large-scale treatment of substrate and can effectively improve the degradation capacity of various antibiotics.
[0108] Step 4: Adjust the input parameters of the optimal prediction model based on the single-factor dependency plot and two-factor dependency plot generated in Step 3. Then, determine the optimal raw material parameters and process operation parameters in the hydrothermal process based on the prediction results of the optimal prediction model for antibiotic degradation rate under different input parameters.
[0109] The specific process of step four is as follows:
[0110] Step 4.1. For the parameter residence time, first obtain the maximum dependency value in the single-factor PDP plot corresponding to the residence time, and record the residence time corresponding to the maximum dependency value as . The first interval of stay is then determined as ;
[0111] in, This indicates half the length of the set dwell time interval;
[0112] For the parameter reaction temperature, first obtain the maximum dependency value in the single-factor PDP plot corresponding to the reaction temperature, and record the residence time corresponding to the maximum dependency value as . The first reaction temperature range determined is: ;
[0113] in, This indicates half the length of the set reaction temperature range;
[0114] For the initial pH value of the parameter environment, first obtain the maximum dependence value in the single-factor PDP plot corresponding to the initial pH value, and record the pH value corresponding to the maximum dependence value as . The first pH range determined is: ;
[0115] in, This indicates half the length of the set pH range;
[0116] For the parameter of mushroom residue quality, firstly, obtain the maximum dependency value in the single-factor PDP plot corresponding to the mushroom residue quality, and record the mushroom residue quality corresponding to the maximum dependency value as . The first quality range of the fungal residue is then determined as follows: ;
[0117] in, This indicates half the length of the set range of mushroom residue quality.
[0118] For the heating rate parameter, first obtain the maximum dependency value in the single-factor PDP plot corresponding to the heating rate, and denote the heating rate corresponding to the maximum dependency value as . The first heating rate range determined is: ;
[0119] in, This indicates half the length of the set heating rate range;
[0120] For the parameter acid concentration, first obtain the maximum dependence value in the single-factor PDP plot corresponding to the acid concentration, and denote the acid concentration corresponding to the maximum dependence value as . The first acid concentration range determined is: ;
[0121] in, This indicates half the length of the set acid concentration range;
[0122] For the parameter moisture content, if the maximum dependence value in the single-factor PDP plot of the parameter moisture content corresponds to a moisture content of 0, then the first moisture content interval is... ;
[0123] in, Indicates the length of the set moisture content range;
[0124] For the parameter reactor volume, first obtain the maximum dependency value in the single-factor PDP plot corresponding to the reactor volume, and denote the reactor volume corresponding to the maximum dependency value as . The determined volume range of the first reaction device is: ;
[0125] in, This indicates half the length of the defined reaction device volume range;
[0126] It should be noted that since the antibiotic type needs to be identified first during the optimization process, there is no need to consider optimizing the antibiotic type parameter.
[0127] Step 4.2: Determine the second range of parameters such as residence time, reaction temperature, initial ambient pH, residue mass, heating rate, acid concentration, moisture content, and reaction device volume based on the two-factor PDP diagram.
[0128] Taking residence time as an example, this invention first obtains two-factor PDP plots for residence time in relation to reaction temperature, initial ambient pH, substrate mass, heating rate, acid concentration, moisture content, and reactor volume. Then, it finds the maximum value in each two-factor PDP plot and obtains the corresponding parameter values for each maximum value in its own two-factor PDP plot, i.e., obtaining the following sets of parameter values: (residence time, reaction temperature), (residence time, initial ambient pH), (residence time, substrate mass), (residence time, heating rate), (residence time, acid concentration), (residence time, moisture content), and (residence time, reactor volume).
[0129] Taking the dwell time value in each set of parameter values as the center, similar to step 4.1, expand outwards to the left and right of the center point. The length is then calculated, and the intersection of the 7 extended results is taken as the second dwell time interval.
[0130] Step 43: Find the intersection of the first and second intervals of the dwell time, and use the result of the intersection as the adjustment range of the dwell time;
[0131] Similarly, the adjustment ranges for reaction temperature, initial ambient pH, bacterial residue mass, heating rate, acid concentration, moisture content, and reaction device volume were obtained respectively.
[0132] Step 4: Set the adjustment step size for each parameter, select parameters within the determined adjustment range, input each parameter combination into the best prediction model, take the maximum degradation rate value output by the best prediction model as the optimal solution rate value, and take the parameter combination corresponding to the optimal solution rate value as the optimal parameter combination. This realizes the parameter optimization method based on degradation rate prediction.
[0133] The specific process of step four is further explained in detail below. Taking the reaction temperature as an example, the left endpoint of the adjustment range is taken as the first feasible parameter value. Then, an adjustment step size is added to obtain the second feasible parameter value, and so on until the right endpoint of the interval is reached, thus obtaining all feasible parameter values corresponding to the reaction temperature. Similarly, the feasible parameter values corresponding to each parameter are obtained. By combining the feasible values of each parameter, multiple sets of input parameters for the model can be obtained. Each set of input parameters is then input into the optimal prediction model to select the best parameter combination.
[0134] To further verify the applicability of the method of the present invention in real-world scenarios, the actual degradation data and predicted results of spiramycin and oxytetracycline bacterial residues were compared, as shown in Figures 6(a) and 6(b). The predicted degradation rates were close to the actual values, with errors within an acceptable range, indicating that the model has good predictive ability. For tetracycline bacterial residues, after optimizing the parameters according to the method of the present invention, the actual degradation rate of tetracycline bacterial residues can reach 98.2%, and the model's predicted result is 98.4%, which is close to the actual degradation rate of 98.2%, proving the effectiveness of parameter optimization based on the method of the present invention.
[0135] In summary, due to the complex environment of hydrothermal systems and the existence of multicollinearity among variables, model construction requires meticulous design in feature selection, variable representation, and algorithm matching; otherwise, it is prone to overfitting or prediction distortion. Therefore, compared with existing methods, this invention has made unique designs in variable construction and model design, specifically as follows:
[0136] (1) In selecting variables, this invention does not simply use conventional parameters, but combines the reaction characteristics of the hydrothermal system to screen out representative variables that can comprehensively reflect the energy input, reaction environment and medium properties of the reaction system, thereby improving the scientific nature and predictive stability of the model.
[0137] (2) Through multi-algorithm comparison and cross-validation mechanism, the effects of different input variables on antibiotic degradation rate were systematically evaluated, and the interaction relationship of variable factors was clarified.
[0138] (3) The prediction model constructed in this invention not only shows excellent fitting accuracy and generalization ability in the hydrothermal system of bacterial residue, but also has strong generalizability and transferability.
[0139] (4) The input variables on which the model depends are measurable or calculable general physicochemical parameters, which can be applied to the hydrothermal degradation process of other antibiotic-containing organic solid wastes;
[0140] (5) Through model retraining and feature recalibration, rapid transfer and adaptive optimization can be achieved between different reaction conditions, reaction systems or antibiotic types, thereby significantly reducing experimental costs and parameter screening difficulty.
[0141] The optimal prediction model of this invention can identify complex nonlinear relationships between variables from data without requiring a clear understanding of the reaction mechanism. This invention not only achieves high accuracy and interpretability in predicting antibiotic degradation rates during the hydrothermal process of bacterial residue, but also makes the model generalizable and transferable, providing a universal and forward-looking technical approach for the intelligent treatment of complex solid waste systems.
[0142] Furthermore, to interpret key experimental parameters related to antibiotic degradation, Shapley additive explanatory values were used to assess the relative contribution of each influencing factor to the antibiotic degradation rate and to visualize the internal workings of the best prediction model. The magnitude of the Shapley value indicates the importance of the parameter, with a larger value indicating greater importance. In addition, Shapley values can reveal the positive and negative impacts of these parameters on degradation rate prediction.
[0143] Calculate the Shapley value for each parameter in the raw material parameters and process operation parameters respectively;
[0144]
[0145] in, It is a collection of all raw material parameters and process operation parameters. It does not include parameters Any subset of parameters (i.e., first set the parameters) From the set After removing from the middle, we get the set of remaining parameters. , For set (a subset of) Representative set The number of parameters in Represents a set The number of parameters in; Is it only the subset When the parameters in the model are used as input to the optimal prediction model, the output of the optimal prediction model is the predicted value of the degradation rate. Indicates only subset Parameters and parameters When used as input to the optimal prediction model, the optimal prediction model for the first... Degradation rate prediction results for each sample; This represents the total number of samples. Indicates parameters The Shapley value.
[0146] like Figure 3 As shown, the importance of various parameters in the hydrothermal treatment process is demonstrated. The results show that the reaction temperature and residence time have the greatest impact on the hydrothermal process, followed by the type of antibiotic, the pH value of the initial environment, the quality of the bacterial residue, the heating rate, and the moisture content. This invention can reveal the key control factors and their interactions in the hydrothermal system through methods such as characteristic importance analysis.
[0147] In summary, this invention embeds an antibiotic degradation prediction model into the entire process of hydrothermal process development, realizing a digital R&D model of "calculate first, then do".
[0148] (1) Laboratory stage: Only the relevant parameters such as reaction temperature, pH, moisture content, and residence time need to be input. The degradation rate is output by the model, which replaces the traditional fixed gradient trial and error process. The process window screening can be completed in a short period of time, reducing the amount of experimentation.
[0149] (2) Pilot-scale: The average temperature inside the device is obtained by CFD thermal gradient simulation coupling. The average temperature is used as the input of the prediction model to correct the temperature deviation caused by the amplification effect in advance and ensure the success rate of the pilot-scale process.
[0150] (3) Raw material switching: When the source of bacterial residue or the type of antibiotic changes, historical data can be used for retraining without conducting a complete experiment, thus shortening the process conversion cycle;
[0151] (4) Process design: The model directly derives the recommended operating point of "high degradation rate and low energy consumption" to reduce resource waste caused by conservative design.
[0152] The predictive model is deployed as a field-level "soft instrument + feedback controller" to achieve closed-loop intelligent control of the hydrothermal reaction process.
[0153] (1) Online detection: Temperature, pH and other sensors are installed at the inlet and outlet of the reactor respectively. The model outputs the current degradation rate prediction value every 20 seconds, which replaces offline HPLC and reduces the analysis cost;
[0154] (2) Control strategy: When the predicted degradation rate is lower than the set value (e.g., 90%) and continues for 3 cycles, the system will automatically trigger the "heating + acid replenishment" operation to bring the degradation rate back to the target range;
[0155] (3) Remote operation and maintenance: All real-time data and control records are uploaded to the cloud in real time. Engineers can remotely view the degradation rate trend, modify the target value, or switch between the two operating modes of "energy saving / high removal" with one click through the mobile APP, so as to achieve high-quality, low-cost and low-carbon emission operation under unattended operation.
[0156] The above examples of this invention are merely illustrative of the computational model and process of this invention, and are not intended to limit the implementation of this invention. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is impossible to exhaustively list all possible implementations here. Any obvious variations or modifications derived from the technical solutions of this invention are still within the scope of protection of this invention.
Claims
1. A data-driven method for predicting antibiotic degradation rate in the hydrothermal process of bacterial residue, characterized in that, The method specifically includes the following steps: Step 1: Collect raw material parameters of different fungal residues, as well as process operation parameters and degradation rate data of different fungal residues under different hydrothermal conditions, and divide the collected data into two parts: training set and test set; Step 2: Train each prediction model using the training set, then evaluate the prediction performance of each prediction model on the test set, and select the best prediction model from among the models based on the prediction performance. Step 3: Generate single-factor dependency plots and two-factor dependency plots based on the best prediction model; Step 4: Adjust the input parameters of the optimal prediction model based on the single-factor dependency plot and two-factor dependency plot generated in Step 3. Then, determine the optimal raw material parameters and process operation parameters in the hydrothermal process based on the prediction results of the optimal prediction model for antibiotic degradation rate under different input parameters.
2. The data-driven method for predicting antibiotic degradation rate in bacterial residue hydrothermal processes according to claim 1, characterized in that, The raw material parameters of the bacterial residue include bacterial residue quality, moisture content, type of antibiotics, and pH value of the initial environment.
3. The data-driven method for predicting antibiotic degradation rate in the hydrothermal process of bacterial residue according to claim 2, characterized in that, The process operating parameters include the volume of the reaction apparatus, residence time, reaction temperature, heating rate, and acid concentration.
4. The data-driven method for predicting antibiotic degradation rate in the hydrothermal process of bacterial residue according to claim 3, characterized in that, The prediction models include the extreme gradient boosting model, the random forest model, and the K-nearest neighbor model; The inputs to the prediction model include raw material parameters of the bacterial residue and process operation parameters under hydrothermal conditions. The output of the prediction model is degradation rate data.
5. The data-driven method for predicting antibiotic degradation rate in the hydrothermal process of bacterial residue according to claim 4, characterized in that, The evaluation of the prediction performance of each prediction model on the test set, and the selection of the best prediction model from among them based on the prediction performance, specifically involves: Evaluate the coefficient of determination of each prediction model on the test set, and select the prediction model with the largest coefficient of determination as the best prediction model.
6. The data-driven method for predicting antibiotic degradation rate in the hydrothermal process of bacterial residue according to claim 5, characterized in that, The prediction results of the extreme gradient boosting model are as follows: in, Indicates the first decision trees for the first Predicted degradation rate for each sample Indicates the total number of decision trees. This indicates that the extreme gradient boosting model applies to the first... Degradation rate prediction results for each sample.
7. The data-driven method for predicting antibiotic degradation rate in bacterial residue hydrothermal processes according to claim 5, characterized in that, The prediction results of the random forest model are as follows: in, Indicates the random forest model for the th Predicted degradation rate for each sample The total number of trees, For the first tree to the first Predicted degradation rate for each sample.
8. The data-driven method for predicting antibiotic degradation rate in the hydrothermal process of bacterial residue according to claim 5, characterized in that, The prediction results of the K-nearest neighbor model are as follows: in, Indicates the first A set consisting of the K neighbors of a sample. Indicates the first Predicted degradation rate of one neighboring unit. The K-nearest neighbor model represents the first... Predicted degradation rate for each sample.
9. The data-driven method for predicting antibiotic degradation rate in the hydrothermal process of bacterial residue according to claim 5, characterized in that, The specific process of step four is as follows: Step 4.
1. For the parameter residence time, first obtain the maximum dependency value in the single-factor PDP plot corresponding to the residence time, and record the residence time corresponding to the maximum dependency value as . The first interval of stay is then determined as ; in, This indicates half the length of the set dwell time interval; For the parameter reaction temperature, first obtain the maximum dependency value in the single-factor PDP plot corresponding to the reaction temperature, and record the residence time corresponding to the maximum dependency value as . The first reaction temperature range determined is: ; in, This indicates half the length of the set reaction temperature range; For the initial pH value of the parameter environment, first obtain the maximum dependence value in the single-factor PDP plot corresponding to the initial pH value, and record the pH value corresponding to the maximum dependence value as . The first pH range determined is: ; in, This indicates half the length of the set pH range; For the parameter of mushroom residue quality, firstly, obtain the maximum dependency value in the single-factor PDP plot corresponding to the mushroom residue quality, and record the mushroom residue quality corresponding to the maximum dependency value as . The first quality range of the fungal residue is then determined as follows: ; in, This indicates half the length of the set range of mushroom residue quality. For the heating rate parameter, first obtain the maximum dependency value in the single-factor PDP plot corresponding to the heating rate, and denote the heating rate corresponding to the maximum dependency value as . The first heating rate range determined is: ; in, This indicates half the length of the set heating rate range; For the parameter acid concentration, first obtain the maximum dependence value in the single-factor PDP plot corresponding to the acid concentration, and denote the acid concentration corresponding to the maximum dependence value as . The first acid concentration range determined is: ; in, This indicates half the length of the set acid concentration range; For the parameter moisture content, if the maximum dependence value in the single-factor PDP plot of the parameter moisture content corresponds to a moisture content of 0, then the first moisture content interval is... ; in, Indicates the length of the set moisture content range; For the parameter reactor volume, first obtain the maximum dependency value in the single-factor PDP plot corresponding to the reactor volume, and denote the reactor volume corresponding to the maximum dependency value as . The determined volume range of the first reaction device is: ; in, This indicates half the length of the defined reaction device volume range; Step 4.2: Determine the second range of parameters such as residence time, reaction temperature, initial ambient pH, residue mass, heating rate, acid concentration, moisture content, and reaction device volume based on the two-factor PDP diagram. Taking residence time as an example, this invention first obtains two-factor PDP plots for residence time in relation to reaction temperature, initial ambient pH, substrate mass, heating rate, acid concentration, moisture content, and reactor volume. Then, it finds the maximum value in each two-factor PDP plot and obtains the corresponding parameter values for each maximum value in its own two-factor PDP plot, i.e., obtaining the following sets of parameter values: (residence time, reaction temperature), (residence time, initial ambient pH), (residence time, substrate mass), (residence time, heating rate), (residence time, acid concentration), (residence time, moisture content), and (residence time, reactor volume). Taking the dwell time value in each set of parameter values as the center, extend to the left and right of the center point respectively. The length is then calculated, and the intersection of the 7 extended results is taken as the second dwell time interval. Step 43: Find the intersection of the first and second intervals of the dwell time, and use the result of the intersection as the adjustment range of the dwell time; Similarly, the adjustment ranges for reaction temperature, initial ambient pH, bacterial residue mass, heating rate, acid concentration, moisture content, and reaction device volume were obtained respectively. Step 4: Set the adjustment step size for each parameter, select parameters within the determined adjustment range, input each parameter combination into the best prediction model, take the maximum degradation rate value output by the best prediction model as the optimal solution rate value, and take the parameter combination corresponding to the optimal solution rate value as the optimal parameter combination.