Machine learning-based methods, systems, storage media, and equipment for predicting water toxicity in ozone advanced oxidation systems.

CN122575555APending Publication Date: 2026-08-14TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-20
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

该方法通过整合催化剂结构特性、反应条件和污染物的本征属性的关键参数,构建高精度的机器学习模型,实现对中间产物综合毒性的快速预测,从而筛选出毒性较低的降解路径,解决实验评估效率低、成本高、难以应用于复杂中间产物体系的技术难题

Benefits of technology

(1)本发明选取特定特征数据训练得到的机器学习模型,能够快速预测臭氧高级氧化过程中间产物的综合毒性,避免了耗时长、成本高的传统生物学实验。经测试,本发明所构建的随机森林模型在测试集上的预测准确性达到了0.8581,能够为实际应用提供可靠的参考。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575555A_ABST
    Figure CN122575555A_ABST
Patent Text Reader

Abstract

This invention relates to the field of water toxicity prediction technology, specifically to a method, system, storage medium, and device for predicting the water toxicity of ozone advanced oxidation systems based on machine learning. This method integrates key parameters of catalyst structural characteristics, reaction conditions, and intrinsic properties of pollutants to construct a high-precision machine learning model, enabling rapid prediction of the comprehensive toxicity of intermediate products, thereby screening out degradation pathways with lower toxicity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of water toxicity prediction technology, specifically to a method, system, storage medium, and device for predicting water toxicity in an ozone advanced oxidation system based on machine learning. Background Technology

[0002] Metal-catalyzed ozone advanced oxidation technology, due to its strong oxidizing properties and broad spectrum, has been widely used for the degradation of common organic pollutants in water bodies (such as phenols, dyes, and pharmaceutical intermediates). However, in practical applications, ozone oxidation often fails to completely mineralize organic pollutants into carbon dioxide and water, instead generating a series of more complex intermediate products. These intermediate products may have stronger biotoxicity than the parent pollutants, leading to an increase in the toxicity of the treated water instead of a decrease. Therefore, accurately assessing and predicting the comprehensive toxicity of intermediate products during ozone oxidation is crucial for optimizing reaction conditions and screening environmentally friendly, low-toxicity degradation pathways.

[0003] Currently, the assessment of the toxicity of intermediate products mainly relies on two types of methods. One type is traditional biological testing methods, such as acute toxicity testing using models like luminescent bacteria, Daphnia magna, or zebrafish embryos. These methods are reliable but involve long experimental cycles, complex operations, low throughput, and high costs. The other type uses toxicity prediction software such as TEST and EPI Suite to estimate toxicity based on the molecular structure of the compound. While these methods are fast and convenient, their prediction accuracy is highly dependent on the accurate identification of the compound structure. For the large quantities of complex and constantly evolving intermediate product mixtures generated during ozone oxidation, both their applicability and accuracy face significant challenges.

[0004] In recent years, machine learning, as an efficient data-driven method, has provided new ideas for solving the above problems. Existing studies have attempted to apply machine learning to the field of environmental catalysis. For example, Gao et al. (GAO W, XU Y, CHANG X, et al. Machine Learning-Driven Global Optimization of Single-Atom Catalyst-Mediated Advanced Oxidation Processes[J]. Environmental Science & Technology,2025,59(44). https: / / doi.org / 10.1021 / acs.est.5c07237) constructed a global optimization strategy for single-atom catalyst-driven advanced oxidation processes using machine learning, predicted pollutant degradation performance using a random forest model, revealed that the number of d electrons in the central metal and the average electronegativity of the coordination environment are key descriptors that determine catalytic performance, and verified the linear relationship between these descriptors and the activation energy of persulfate through theoretical calculations. Furthermore, He et al. (HE X, GAO W, XUJ, et al. Machine learning-assisted construction of C=O and pyridinic N active sites in sludge-based catalysts[J]. Chinese Chemical Letters,2025:111019. https: / / doi.org / 10.1016 / j.cclet.2025.111019.) applied machine learning to the construction of active sites in sludge-based catalysts. They developed an extreme gradient enhancement model to predict C=O active sites and constructed an integrated model to predict pyridine N active sites. Through SHAP analysis and partial dependence graphs, they analyzed the influence of pyrolysis parameters and elemental composition on the formation of active sites, achieving accurate prediction of active site content.

[0005] However, the aforementioned studies mainly focus on catalyst performance prediction and active site construction, and research on the prediction of the comprehensive toxicity of intermediate products in ozone advanced oxidation processes remains relatively scarce. Therefore, there is an urgent need to develop a method that can effectively address the above challenges and achieve efficient and accurate prediction of the comprehensive toxicity of intermediate products in ozone advanced oxidation processes. Summary of the Invention

[0006] This invention provides a method, system, storage medium, and device for predicting the water toxicity of ozone advanced oxidation systems based on machine learning. This method integrates key parameters of catalyst structural characteristics, reaction conditions, and intrinsic properties of pollutants to construct a high-precision machine learning model, enabling rapid prediction of the comprehensive toxicity of intermediate products. This allows for the screening of degradation pathways with lower toxicity, solving the technical challenges of low efficiency, high cost, and difficulty in applying experimental evaluation to complex intermediate product systems.

[0007] To achieve the above objectives, the present invention provides the following technical solution: This invention provides a machine learning-based method for predicting the toxicity of water bodies in ozone advanced oxidation systems, comprising: Sample data of the ozone advanced oxidation system were obtained, and the catalyst structure parameters, reaction parameters and pollutant parameters of each sample were extracted and encoded as feature data. The comprehensive toxicity index of the intermediate products of each sample was obtained as label data, and a training set was constructed together. The random forest model is trained using the training set to establish a mapping relationship between the feature data and the label data, thereby obtaining a comprehensive toxicity prediction model. The catalyst structure parameters, reaction parameters, and pollutant parameters of the ozone advanced oxidation system to be tested are obtained and input into the comprehensive toxicity prediction model, which outputs the comprehensive toxicity index prediction value of the intermediate products of the ozone advanced oxidation system to be tested.

[0008] The inventors faced the following challenges when applying machine learning to predict the toxicity of ozone oxidation intermediates.

[0009] One challenge is model selection. The toxicity of intermediate products in ozone oxidation systems is influenced by a combination of factors, and the data often exhibits complex nonlinear characteristics. There is a lack of systematic theoretical and empirical guidance on how to select the most suitable model from numerous machine learning algorithms that strikes a balance between prediction accuracy and interpretability for such complex systems.

[0010] In order to accurately select the model in the complex nonlinear ozone oxidation system and to overcome the black box nature of high-precision models, the inventors selected four models for comparison: K-Nearest Neighbors (KNN), Support Vector Regression (SVR), Extreme Gradient Boosting (XG Boost), and Random Forest, and selected the best model.

[0011] Secondly, determining the input characteristics is difficult. Numerous factors influence the toxicity of intermediate products during ozone oxidation, and these factors are interconnected. These include reaction conditions such as ozone dosage, pH, and reaction time, as well as catalyst characteristics such as metal type and dosage. Furthermore, the physicochemical properties of the pollutants themselves, such as molecular structure and electronic characteristics, are also involved. The core challenge in determining the upper limit of the model's predictive performance is how to scientifically screen and construct a set of input characteristics that can comprehensively and accurately describe the reaction system from these multi-source and heterogeneous parameters.

[0012] To comprehensively and accurately describe the reaction system and avoid feature redundancy or omission of key information, the inventors ultimately selected specific feature data from a literature sample: catalyst structural parameters, reaction parameters, and pollutant parameters. This ensures that the selected feature data reflects the toxicity evolution pathway during ozone oxidation. Furthermore, preprocessing or feature sorting and screening can be performed to further improve the accuracy of the selected feature descriptions.

[0013] Finally, by utilizing specific feature data in conjunction with a random forest model, the accuracy of the comprehensive toxicity prediction model is maximized, enabling the prediction of the toxicity of intermediate products in the ozone advanced oxidation system.

[0014] Preferably, the catalyst structural parameters include whether the catalyst has a support, whether the support has a metal, the atomic number of the first metal in the catalyst, and the atomic number of the second metal in the catalyst. The reaction parameters include ozone dosage, catalyst dosage, pollutant concentration, pH value, reaction time, and reaction temperature; The pollutant parameters include the initial comprehensive toxicity of the parent pollutant, the number of aromatic rings in the pollutant molecule, the number of carbon atoms, the highest occupied molecular orbital energy, and the lowest unoccupied molecular orbital energy.

[0015] This invention is the first to use "initial comprehensive toxicity of the parent pollutant" as an input feature, allowing the model to directly learn the toxicity evolution path, which greatly improves the accuracy of feature description.

[0016] Preferably, the method for obtaining the label data includes: obtaining multiple toxicity indicators of intermediate products of each sample, determining the weight of each toxicity indicator through principal component analysis, and calculating the weighted comprehensive toxicity index of the intermediate product as the label data. The toxicity indicators include the 96-hour median lethal concentration (LD50) of blackhead carp, the 48-hour median lethal concentration of Daphnia magna, the 48-hour median inhibitory concentration (LD50) of Tetrahymena piriformis, the oral median lethal dose in rats, bioaccumulation factor, developmental toxicity, and Ames mutagenicity.

[0017] Preferably, before constructing the training set, the extracted feature data is preprocessed; the preprocessing includes removing outliers, filling missing values, numerically encoding non-numerical features, and calculating Pearson correlation coefficients to remove redundant features.

[0018] Preferably, the LASSO coefficient sorting method is used to filter the feature data, remove redundant features, and retain the optimal feature combination as the final feature data for training the random forest model.

[0019] Preferably, during the training of the random forest model using the training set, the hyperparameters of the random forest model are tuned using a Bayesian optimization algorithm.

[0020] Preferably, the method for predicting the toxicity of ozone advanced oxidation system water further includes: using the SHAP algorithm to analyze the comprehensive toxicity prediction model, quantitatively calculating the contribution of each feature data to the comprehensive toxicity index of intermediate products, so as to verify and identify the key factors affecting the toxicity of water.

[0021] To address the common problem of insufficient interpretability in machine learning models, the SHAP algorithm was employed for model analysis. This algorithm can quantitatively calculate the contribution of each feature data to the predicted value of the comprehensive toxicity index, not only verifying the rationality of the comprehensive toxicity prediction model, but also clearly identifying key factors affecting water toxicity (such as the number of carbon atoms and initial comprehensive toxicity), perfectly balancing high accuracy and interpretability.

[0022] This invention provides a machine learning-based water toxicity prediction system for ozone advanced oxidation systems, comprising: The data acquisition module is used to acquire the catalyst structure parameters, reaction parameters, and pollutant parameters of each sample as feature data, and the corresponding comprehensive toxicity index of intermediate products as label data to obtain the training set. The training module is used to establish a mapping relationship between the feature data and the label data on a random forest model using the training set, so as to obtain a comprehensive toxicity prediction model; The prediction module is used to input the catalyst structure parameters, reaction parameters, and pollutant parameters of the ozone advanced oxidation system to be tested, and to obtain the predicted value of the comprehensive toxicity index of the intermediate products of the ozone advanced oxidation system to be tested using the comprehensive toxicity prediction model.

[0023] The present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described method for predicting the toxicity of water bodies in an ozone advanced oxidation system.

[0024] The present invention provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the above-described method for predicting the toxicity of water bodies in an ozone advanced oxidation system.

[0025] Therefore, the present invention has the following beneficial effects: (1) The machine learning model trained by the present invention using specific feature data can quickly predict the comprehensive toxicity of intermediate products in the advanced oxidation process of ozone, avoiding the time-consuming and costly traditional biological experiments. Tests show that the random forest model constructed in this invention achieves a prediction accuracy of 0.8581 on the test set, providing a reliable reference for practical applications.

[0026] (2) The specific feature data selected in this invention covers the structure and physicochemical properties of catalysts, reaction conditions and pollutants themselves, and constructs a multi-dimensional feature training set, which helps to improve the accuracy and generalization ability of model prediction.

[0027] (3) When processing the selected specific feature data, this invention uses Pearson correlation coefficient to remove redundant features to ensure feature independence, and uses logarithmic transformation to represent the initial comprehensive toxicity of the parent pollutant; it uses the atomic number of the catalyst metal as a continuous numerical feature rather than a categorical variable, so that the model can directly learn the relationship between the metal electronic structure and catalytic activity. The above-mentioned special feature processing strategy breaks through the conventional noise reduction and normalization operations.

[0028] (4) The present invention introduces the SHAP algorithm to interpret the model, which can not only provide prediction results, but also quantitatively analyze the direction and importance of the influence of each reaction condition and pollutant structural parameters on the final toxicity, reveal the formation mechanism of toxicity, and enhance the transparency and credibility of the model.

[0029] (5) The method of the present invention can be used as an efficient screening tool. By inputting different combinations of process parameters, the comprehensive toxicity of intermediate products under the corresponding path can be predicted, thereby guiding researchers and engineers to select the degradation path with the lowest toxicity and the optimal process parameters, thereby reducing the environmental risks in the water treatment process from the source. Attached Figure Description

[0030] Figure 1 This is a flowchart of a machine learning-based method for predicting water toxicity in ozone advanced oxidation systems. Figure 2 The prediction results are from the comprehensive toxicity prediction model; Figure 3 The overall toxicity prediction error is verified experimentally; Figure 4 This is another flowchart of a machine learning-based method for predicting water toxicity in ozone advanced oxidation systems. Figure 5 This is a hardware structure block diagram of a computer terminal. Figure 6 This is a schematic diagram of an electronic device; Figure 7This is a schematic diagram of the system. Detailed Implementation

[0031] The present invention will be further described below with reference to specific embodiments. Those skilled in the art will be able to implement the present invention based on these descriptions. Furthermore, the embodiments of the present invention described below are generally only some, not all, of the embodiments of the present invention. Therefore, all other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort should fall within the scope of protection of the present invention.

[0032] Example 1 Prediction Method S1. Data Acquisition S1.1 Sample Collection A search was conducted on Web of Science for information related to the advanced oxidation system of ozone, using the keyword "catalytic ozonation". The retrieved literature was reviewed; from 2020 to 2026, a total of 244 peer-reviewed articles meeting the requirements were reviewed, and 244 sample data points were retrieved and collected to obtain sample data on the advanced oxidation system of ozone.

[0033] S1.2 Data Extraction ①Comprehensive toxicity index of intermediate products ChemDraw software was used to plot all pollutants and their intermediates in 244 samples of the ozone advanced oxidation system. TEST software was used to predict the toxicity of each intermediate. To objectively integrate seven toxicity indicators (96-hour median lethal concentration of blackhead carp, 48-hour median lethal concentration of Daphnia magna, 48-hour median inhibitory concentration of Tetrahymena piriformis, oral median lethal dose in rats, bioaccumulation factor, developmental toxicity, and Ames mutagenicity) to construct a comprehensive toxicity index, principal component analysis (PCA) was used to determine the weights of each toxicity indicator, and the weighted comprehensive toxicity index of the intermediates was calculated.

[0034] PCA was performed using standardized data from all samples and implemented using the scikit-learn library in Python. The results showed that the first principal component explained 34.21% of the total variance, indicating that it effectively represents the main variation patterns of overall toxicity. The bioaccumulation factor had the highest weight (27.9%), followed by the 96-hour LC50 of the blackhead carp. 50 (17.0%) and 48h large Daphnia LC 50 (14.8%), with the lowest weighting for developmental toxicity (7.5%), and the 48h half-maximal inhibitory concentration of Tetrahymena piraceae, the oral half-maximal lethal dose in rats, developmental toxicity, and Ames mutagenicity were 12.7%, 8.9%, 7.5%, and 12.3%, respectively.

[0035] ②Highest occupied molecular orbital energy and lowest unoccupied molecular orbital energy Using Gaussian 16 software, the structures of all pollutants and their intermediates involved in 244 sample data of the ozone advanced oxidation system were optimized using the 6-31G(d,p) basis set, and the highest occupied molecular orbital energy (HOMO) and lowest unoccupied molecular orbital energy (LUMO) values ​​were calculated.

[0036] ③ Other feature data Information was extracted from sample data of the ozone advanced oxidation system, including whether the catalyst has a support, whether the support contains metals, the atomic number of the first metal in the catalyst, the atomic number of the second metal in the catalyst, the ozone dosage, the catalyst dosage, the pollutant concentration, pH value, reaction time, reaction temperature, the initial comprehensive toxicity of the parent pollutant, and the number of aromatic rings and carbon atoms in the pollutant molecules.

[0037] Following steps ① through ③, the following feature data is generated: catalyst structural parameters (whether the catalyst has a support, whether the support contains metal, the atomic number of the first metal in the catalyst, and the atomic number of the second metal in the catalyst); reaction parameters (ozone dosage, catalyst dosage, pollutant concentration, pH value, reaction time, and reaction temperature); and pollutant parameters (initial comprehensive toxicity of the parent pollutant, number of aromatic rings in the pollutant molecule, number of carbon atoms, highest occupied molecular orbital energy, and lowest unoccupied molecular orbital energy). These 15 parameters are then One-Hot encoded to obtain the corresponding feature data. The categorical features "whether the catalyst has a support" and "whether the support contains metal" are encoded as 0 or 1, and after encoding, each is expanded into two binary features. The final feature data contains 17 encoded features.

[0038] Additionally, the label data includes the comprehensive toxicity index of the intermediate products. The combined feature data and label data will be used as the dataset for subsequent processing.

[0039] S1.3 Preprocessing The feature data in the dataset underwent preprocessing: outliers were removed, and missing values ​​were imputed using the median. Categorical features "whether the catalyst has a carrier" and "whether the carrier has metal" were encoded as 0 or 1. The Pearson correlation coefficients between all features were calculated, and all were less than 0.6, confirming the absence of redundant features. Next, the importance of the 17 encoded features was evaluated using LASSO coefficient ranking, and the top 15 features with the highest absolute coefficient values ​​were selected as the preprocessed feature data (i.e., two redundant encoded dimensions were removed, while retaining the complete 15 original feature information).

[0040] The preprocessed dataset was randomly divided into a training set and a test set at a ratio of 70% and 30%, respectively. The training set contained 170 data points, and the test set contained 74 data points.

[0041] S2. Model Training Triple-fold cross-validation was performed on the training set, and the hyperparameters of the random forest model were tuned using a Bayesian optimization algorithm to find the optimal combination of hyperparameters, thus training a comprehensive toxicity prediction model. The performance of the comprehensive toxicity prediction model was evaluated using the test set. The final comprehensive toxicity prediction model had a test set determination coefficient of 0.8581 and a root mean square error of 0.1478. Figure 2 ).

[0042] S3. Actual Result Prediction Four different catalyst types and reaction conditions were selected to validate the ozone oxidation model. For example... Figure 3 For comprehensive toxicity prediction, the prediction error is less than 6%.

[0043] Example 2 This embodiment is basically the same as Embodiment 1, except that after the S2 model is trained, the SHAP algorithm is used to analyze the comprehensive toxicity prediction model in order to verify and identify the key factors affecting the toxicity of water bodies.

[0044] SHAP values ​​for each feature data were calculated, and feature importance ranking plots and SHAP scatter plots were plotted. SHAP analysis results showed that the initial comprehensive toxicity of the parent compound and the number of carbon atoms contributed the most to the prediction of the comprehensive toxicity of intermediate products. The SHAP importance ranking of each input feature was as follows: number of carbon atoms > initial comprehensive toxicity > HOMO level > LUMO level > reaction time > first metal atomic number > pH value > ozone dosage > catalyst dosage > initial pollutant concentration > whether the catalyst has a support > second metal atomic number > number of aromatic rings > whether the support has a metal > reaction temperature. Pollutant parameters accounted for the largest share of importance (67%), followed by reaction conditions (21%), and catalyst conditions accounted for 12%.

[0045] Comparative Example 1 Model Screening Refer to Example 1, except that: Figure 4 As shown, four machine learning algorithms—K-nearest neighbors, support vector regression, extreme gradient boosting, and random forest—were used for model training.

[0046] The predictive performance of the four models was evaluated using a test set, and the coefficient of determination (COP) and root mean square error (RMSE) of each model were calculated. The COP of the Random Forest model was 0.8581, and the RMSE was 0.1478; the COP of the Extreme Gradient Boosting model was 0.61, and the RMSE was 0.24; the COP of the Support Vector Regression model was 0.54, and the RMSE was 0.27; and the COP of the K-Nearest Neighbors model was 0.16, and the RMSE was 0.36.

[0047] In summary, the random forest model exhibits the best predictive performance and is therefore selected as the final comprehensive toxicity prediction model.

[0048] Comparative Example 2 Feature Filtering Method The procedure is the same as in Example 1, except that in S1.3 preprocessing, LASSO coefficient sorting is replaced by LASSO feature selection or Boruta feature selection. The original 15 features include two categorical features (whether the catalyst has a support, and whether the support has a metal), which are expanded into two binary features after One-Hot encoding, resulting in a total of 17 features. Both LASSO and Boruta selection are performed on these 17 encoded features.

[0049] (1) LASSO feature selection: L1 regularization was used to compress the coefficients of irrelevant features to zero. Experimental results show that LASSO selected three non-zero coefficient features, corresponding to the number of carbon atoms, the initial comprehensive toxicity of the parent pollutant, and the amount of catalyst added in the original features. The three selected features were input into the random forest model, and the test set R²=0.7329, RMSE=0.2028, which is significantly lower than the accuracy of the full feature model.

[0050] (2) Boruta Feature Selection: By creating shadow features as a random baseline, features significantly related to the target variable are identified. Experimental results show that Boruta selects 6 features, corresponding to the number of carbon atoms, the initial comprehensive toxicity of the parent pollutant, the amount of catalyst added, and one encoding dimension of the category feature in the original features. When the selected 6 features are input into the random forest model, the test set R²=0.8032 and RMSE=0.1741, which are still lower than the full feature model.

[0051] Table 1 Impact of Feature Filtering Methods

[0052] Analysis of the above results shows that Boruta and LASSO screening alone both lead to a decrease in accuracy, proving that toxicity prediction requires the synergistic effect of multiple features; while the accuracy of the top 15 features in LASSO coefficient ranking (i.e., the 15 features after removing 2 redundant coding dimensions) is better than that of all features, proving that the 15 original features of this invention are not simplistic and are the verified optimal feature combination.

[0053] Feature selection in Comparative Example 3 The same procedure applies as in Example 1, except that the number of features in the dataset is different during the training of the S2 model, and no hyperparameter tuning is performed; see Table 2 for details.

[0054] Table 2 Effects of Characteristic Ablation

[0055] Analysis of Table 2 above shows that the addition of pollutant structural parameters (number of aromatic rings, number of carbon atoms) increases R. 2 The increase from 0.1754 to 0.6064 (an increase of 0.4310) indicates that the molecular skeleton of the pollutant has a decisive influence on its toxicity; the addition of pollutant electronic parameters (HOMO, LUMO) improves R... 2 The R² value was further increased to 0.7152 (an increase of 0.1088), demonstrating that molecular orbital energy levels are a key factor affecting the toxicity of intermediate products; the addition of catalyst structural parameters (metal atomic number, support characteristics) did not show a monotonically increasing R² value in this dataset. 2 The R value dropped to 0.6570, which may be due to the complex interaction between catalyst structural features and pollutant features. However, ultimately, in the full-feature model (including initial toxicity), through nonlinear modeling capabilities, R... 2 Reaching a maximum of 0.8391; the addition of the initial comprehensive toxicity of the parent contaminant caused R to... 2 The accuracy improved from 0.6570 to 0.8391 (an increase of 0.2010), and the RMSE decreased from 0.2298 to 0.1574 (a reduction of 35.6%), demonstrating that the initial toxicity features significantly contribute to the model's accuracy. In summary, the 15 feature combinations selected in this invention can effectively improve the prediction accuracy of the overall toxicity of intermediate products, and each feature plays an irreplaceable role in the model.

[0056] The same procedure applies as in Example 1, except that the number of features in the dataset is different during the training of the S2 model; see Table 3 for details.

[0057] Table 3. Effects of Characteristic Ablation

[0058] Comparative Example 4: Manifestations of the Initial Composite Toxicity of the Parent Pollutant The procedure was carried out in accordance with Example 1, with the difference that a logarithmic transformation (Tox_log = ln(Tox + 0.01)) was applied to the "initial comprehensive toxicity of the parent pollutant" in the feature data to examine the impact of the nonlinear transformation on the model's prediction accuracy. This is because small changes in toxicity values ​​in the low concentration range have a far greater impact on the ecological environment than in the high concentration range; the logarithmic transformation can amplify the differences in the low-value range, theoretically potentially improving the model's sensitivity to low-toxicity samples.

[0059] Using the same training / test set split and the same random forest parameters (n_estimators=100, max_depth=10), the result is R. 2The positivity rate was 0.8454, and the RMSE was 0.1543. Experimental results show that the model performance slightly decreased after logarithmic transformation. Therefore, this invention ultimately uses the original linear value of the initial toxicity as the input feature. This control experiment demonstrates that this invention systematically optimizes feature processing, rather than simply accepting conventional transformations.

[0060] Example 3: Computer-readable storage medium The method provided in this embodiment 2 can be executed in a mobile terminal, computer terminal or similar computing device. Figure 5 A hardware block diagram of a computer terminal (or mobile device) for a machine learning-based method of predicting water toxicity in an ozone advanced oxidation system is shown. Figure 5 As shown, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 5 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 5 The more or fewer components shown, or having the same Figure 5 The different configurations shown.

[0061] Example 4 Electronic device Figure 6 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 6 As shown, the electronic device may include: one or more ( Figure 6 (Only one is shown) processor 202, memory 204, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.

[0062] The memory 204 can be used to store software programs and modules, such as the program instructions / modules corresponding to the machine learning-based ozone advanced oxidation system water toxicity prediction method in Example 1. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned machine learning-based ozone advanced oxidation system water toxicity prediction method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory.

[0063] Example 5 Prediction System A machine learning-based water toxicity prediction system 300 for ozone advanced oxidation systems, such as... Figure 7 As shown, it includes: The data acquisition module 302 is used to acquire the catalyst structure parameters, reaction parameters, and pollutant parameters of each sample as feature data, and the corresponding comprehensive toxicity index of intermediate products as label data to obtain a training set. Training module 304 is used to establish a mapping relationship between feature data and label data on a random forest model using the training set, so as to obtain a comprehensive toxicity prediction model; The prediction module 306 is used to input the catalyst structure parameters, reaction parameters, and pollutant parameters of the ozone advanced oxidation system to be tested, and to obtain the predicted value of the comprehensive toxicity index of the intermediate products of the ozone advanced oxidation system to be tested using a comprehensive toxicity prediction model.

Claims

1. A method for predicting water toxicity in ozone advanced oxidation systems based on machine learning, characterized in that, include: Sample data of the ozone advanced oxidation system were obtained, and the catalyst structure parameters, reaction parameters and pollutant parameters of each sample were extracted and encoded as feature data. The comprehensive toxicity index of the intermediate products of each sample was obtained as label data, and a training set was constructed together. The random forest model is trained using the training set to establish a mapping relationship between the feature data and the label data, thereby obtaining a comprehensive toxicity prediction model. The catalyst structure parameters, reaction parameters, and pollutant parameters of the ozone advanced oxidation system to be tested are obtained and input into the comprehensive toxicity prediction model, and the comprehensive toxicity index prediction value of the intermediate products of the ozone advanced oxidation system to be tested is output.

2. The method for predicting water toxicity in an ozone advanced oxidation system as described in claim 1, characterized in that, The catalyst structural parameters include whether the catalyst has a support, whether the support has a metal, the atomic number of the first metal in the catalyst, and the atomic number of the second metal in the catalyst. The reaction parameters include ozone dosage, catalyst dosage, pollutant concentration, pH value, reaction time, and reaction temperature; The pollutant parameters include the initial comprehensive toxicity of the parent pollutant, the number of aromatic rings in the pollutant molecule, the number of carbon atoms, the highest occupied molecular orbital energy, and the lowest unoccupied molecular orbital energy.

3. The method for predicting water toxicity in an ozone advanced oxidation system as described in claim 1, characterized in that, The method for obtaining the label data includes: obtaining multiple toxicity indicators of intermediate products of each sample, determining the weight of each toxicity indicator through principal component analysis, and calculating the weighted comprehensive toxicity index of the intermediate product as the label data. The toxicity indicators include the 96-hour median lethal concentration (LD50) of blackhead carp, the 48-hour median lethal concentration of Daphnia magna, the 48-hour median inhibitory concentration (LD50) of Tetrahymena piriformis, the oral median lethal dose in rats, bioaccumulation factor, developmental toxicity, and Ames mutagenicity.

4. The method for predicting water toxicity in an ozone advanced oxidation system as described in claim 1, characterized in that, Before constructing the training set, the extracted feature data is preprocessed; the preprocessing includes removing outliers, filling missing values, numerically encoding non-numerical features, and calculating Pearson correlation coefficients to remove redundant features.

5. The method for predicting water toxicity in an ozone advanced oxidation system as described in claim 1 or 4, characterized in that, The LASSO coefficient ranking method is used to filter the features in the data, remove redundant features, and retain the optimal feature combination as the final feature data for training the random forest model.

6. The method for predicting water toxicity in an ozone advanced oxidation system as described in claim 1, characterized in that, During the training of the random forest model using the training set, the hyperparameters of the random forest model are tuned using the Bayesian optimization algorithm.

7. The method for predicting water toxicity in ozone advanced oxidation systems as described in any one of claims 1 to 4, characterized in that, The method for predicting the water toxicity of the ozone advanced oxidation system also includes: using the SHAP algorithm to analyze the comprehensive toxicity prediction model, quantitatively calculating the contribution of each feature data to the comprehensive toxicity index of intermediate products, so as to verify and identify the key factors affecting the toxicity of water.

8. A machine learning-based water toxicity prediction system for ozone advanced oxidation systems, characterized in that, include: The data acquisition module is used to acquire the catalyst structure parameters, reaction parameters, and pollutant parameters of each sample as feature data, and the corresponding comprehensive toxicity index of intermediate products as label data to obtain the training set. The training module is used to establish a mapping relationship between the feature data and the label data on a random forest model using the training set, so as to obtain a comprehensive toxicity prediction model; The prediction module is used to input the catalyst structure parameters, reaction parameters, and pollutant parameters of the ozone advanced oxidation system to be tested, and to obtain the predicted value of the comprehensive toxicity index of the intermediate products of the ozone advanced oxidation system to be tested using the comprehensive toxicity prediction model.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method for predicting water toxicity in an ozone advanced oxidation system as described in any one of claims 1 to 7.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the ozone advanced oxidation system water toxicity prediction method as described in any one of claims 1 to 7.