Water treatment catalyst design method and device based on machine learning

By generating and screening catalytic reaction systems using machine learning-based methods, the problem of complex catalytic reaction design in existing technologies has been solved, enabling rapid and efficient catalyst design and experimental verification.

CN121601062APending Publication Date: 2026-03-03HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511602956.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing catalytic reaction systems are complex to design, rely on experience and repeated experiments, resulting in long cycles, high costs, and difficulty in quickly adapting to complex and changing water quality and process conditions.

Method used

A machine learning-based approach was used to generate multiple candidate catalytic reaction systems, and then a highly efficient target catalytic reaction system was selected by using preset screening rules and a catalytic performance prediction model.

Benefits of technology

It reduces experimental design work and verification time, lowers economic costs, and enables rapid response to complex and changing water quality and process conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121601062A_ABST
    Figure CN121601062A_ABST
Patent Text Reader

Abstract

The invention relates to a water treatment catalyst design method and device based on machine learning, and belongs to the technical field of catalysts.The method comprises the steps that multiple sets of first candidate catalytic reaction systems are generated according to design requirement information of a water treatment catalyst; screening the plurality of groups of first candidate catalytic reaction systems according to a preset screening rule to obtain at least one second candidate catalytic reaction system; wherein the screening rule comprises at least one of screening based on similarity requirements of an actually measured catalytic reaction system, screening based on reaction condition constraints, screening based on catalyst structure constraints and / or active free radical species constraints matched with pollutant species, and screening based on catalyst structure feasibility; and determining a target catalytic reaction system according to the prediction result of the catalytic performance of the second candidate catalytic reaction system. The design cost, the experimental verification time and the economic cost can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of catalyst technology, and in particular to a method and apparatus for designing water treatment catalysts based on machine learning. Background Technology

[0002] In recent years, emerging pollutants have been widely detected in various aquatic media, posing risks to drinking water safety and ecological health. Advanced oxidation technologies (AEOs) utilize oxidants to generate highly reactive oxidizing species in situ, such as hydroxyl radicals (HO•) and sulfate radicals (SO4•). - Oxidants such as hydrogen peroxide (H₂O₂), peroxymonosulfate (PMS), persulfate (PS), and peracetic acid (PAA) demonstrate high efficiency and versatility in the removal of recalcitrant pollutants. Commonly used oxidants include hydrogen peroxide (H₂O₂), peroxymonosulfate (PMS), persulfate (PS), and peracetic acid (PAA). Transition metal spinel oxides (general formula AB₂O₄) are particularly prominent in Fenton-like catalysis due to their abundant resources, stable structure, and tunable bimetallic sites. Cobalt-based spinels, in particular, often exhibit superior activity and reusability compared to monometallic oxides.

[0003] However, the electronic structure coupling, valence state distribution, and metal-oxygen (MO) bond characteristics of spinel A-site (tetrahedral) / B-site (octahedral) dual sites jointly determine its activity, resulting in a complex material design with high dimensionality and strong coupling. In addition, there is a variety of choices for free radical types and reaction conditions. This makes it necessary to rely on experience and repeated experiments to screen catalytic reaction systems in engineering, which has the problems of long cycle and high cost, and it is difficult to quickly adapt to complex and changing water quality and process conditions. Summary of the Invention

[0004] In view of this, it is necessary to provide a machine learning-based method and apparatus for designing water treatment catalysts to solve the problem of complex design of existing catalytic reaction systems.

[0005] To address the aforementioned problems, in a first aspect, the present invention provides a machine learning-based method for designing water treatment catalysts, comprising: Based on the design requirements of water treatment catalysts, multiple first-candidate catalytic reaction systems were generated. The multiple first candidate catalytic reaction systems are screened according to preset screening rules to obtain at least one second candidate catalytic reaction system; wherein, the screening rules include at least one of the following: screening based on similarity requirements with the measured catalytic reaction system, screening based on reaction condition constraints, screening based on catalyst structure constraints and / or active free radical type constraints that are compatible with pollutant types, and screening based on catalyst structure feasibility. The catalytic performance of the second candidate catalytic reaction system is predicted based on the trained catalytic performance prediction model, and the target catalytic reaction system is determined based on the prediction results.

[0006] In one possible implementation, the screening of the multiple groups of first candidate catalytic reaction systems according to a preset screening rule includes: When the screening rule is based on the similarity requirement with the measured catalytic reaction system, the similarity between the first candidate catalytic reaction system and each measured catalytic reaction system is calculated; Determine whether the maximum similarity is less than a preset maximum similarity threshold, and whether the average similarity is less than a preset average similarity threshold; If the maximum similarity is greater than or equal to the preset maximum similarity threshold, and the average similarity is greater than or equal to the preset average similarity threshold, then the first candidate catalytic reaction system is retained. If the maximum similarity is less than the preset maximum similarity threshold, or the average similarity is less than the preset average similarity threshold, then the catalytic performance of the first candidate catalytic reaction system is predicted by the catalytic performance prediction model to obtain the confidence interval of the prediction result. The first candidate catalytic reaction system is screened based on the confidence interval of the prediction results.

[0007] In one possible implementation, the screening of the first candidate catalytic reaction system based on the confidence interval of the prediction result includes: The confidence interval is widened according to a preset rule. If the lower limit of the widened confidence interval is less than the preset catalytic activity threshold, the first candidate catalytic reaction system is screened out. If the lower limit of the widened confidence interval is greater than or equal to the preset catalytic activity threshold, then the first candidate catalytic reaction system is retained.

[0008] In one possible implementation, the reaction condition constraints include a pH value range constraint, and the input features of the catalytic performance prediction model include pH value and other features; the method further includes: At preset pH value quantiles, using pH value as a quantifier and the other characteristics as variables, multiple sets of third candidate catalytic reaction systems are generated. The catalytic performance of the multiple third candidate catalytic reaction systems is predicted using the catalytic performance prediction model. Based on the catalytic performance corresponding to each pH value quantile, a marginal effect curve for pH value is plotted, and a confidence band is superimposed on the marginal effect curve; The pH value range constraint is determined based on the marginal effect curve and the confidence band.

[0009] In one possible implementation, the water treatment catalyst is a transition metal spinel oxide; the method further includes: Using elemental combinations of elemental types at the tetrahedral and octahedral sites of spinel, the types of active free radicals, and the types of pollutants as variables, multiple sets of fourth candidate catalytic reaction systems are constructed; the catalytic performance of these multiple sets of fourth candidate catalytic reaction systems is predicted using the catalytic performance prediction model. Based on the prediction results of the multiple sets of fourth candidate catalytic reaction systems, determine the element combinations and / or types of active free radicals suitable for each pollutant type.

[0010] In one possible implementation, the screening of the multiple groups of first candidate catalytic reaction systems according to a preset screening rule includes: When the screening rule is based on the feasibility of the catalyst structure, it is determined whether the catalyst structure in the first candidate catalytic reaction system is the same as the known and confirmed stable catalyst structure. If the catalyst structure in the first candidate catalytic reaction system is the same as the structure of a known and confirmed stable catalyst, then the first candidate catalytic reaction system is retained. If the catalyst structure in the first candidate catalytic reaction system is different from the known and confirmed stable catalyst structure, then the first candidate catalytic reaction system is eliminated.

[0011] In one possible implementation, the method further includes: Obtain the measured results of the target catalytic reaction system; Determine whether the deviation between the predicted result and the measured result of the target catalytic reaction system is less than a preset deviation threshold; If the deviation is less than a preset deviation threshold, the catalytic performance and result characterization of the catalyst in the target catalytic reaction system in various application scenarios are obtained, and the correlation between the catalyst and the catalytic performance and result characterization in the various application scenarios is established.

[0012] In one possible implementation, the water treatment catalyst is a transition metal spinel oxide; the catalytic performance prediction model is trained as follows: A training sample containing input and output features was constructed based on the measured catalytic reaction system. The input features include intrinsic catalyst features, types of active free radicals, reaction condition features, and types of contaminants. The intrinsic catalyst features include at least one of the following: element type, covalence, d-band center, number of valence electrons, electronegativity, ionic radius, ionization energy, and cutoff radius at the spinel tetrahedral and octahedral sites. The reaction condition features include at least one of the following: pH value, oxidant dosage, and catalyst dosage. The output feature is the catalytic reaction rate constant. The catalytic performance prediction model is trained using the training samples until the catalytic performance prediction model converges.

[0013] In one possible implementation, the method further includes: Generate SHAP values ​​or partial dependency graphs of the input features and visualize them.

[0014] Secondly, the present invention also provides a water treatment catalyst design device based on machine learning, comprising: The first candidate catalytic reaction system generation module is used to generate multiple first candidate catalytic reaction systems based on the design requirements of water treatment catalysts. The first screening module is used to screen the multiple groups of first candidate catalytic reaction systems according to preset screening rules to obtain at least one second candidate catalytic reaction system; wherein, the screening rules include at least one of the following: screening based on similarity requirements with the measured catalytic reaction system, screening based on reaction condition constraints, screening based on catalyst structure constraints and / or active free radical type constraints that are compatible with pollutant types, and screening based on catalyst structure feasibility. The second screening module is used to predict the catalytic performance of the second candidate catalytic reaction system based on the trained catalytic performance prediction model, and to determine the target catalytic reaction system based on the prediction results.

[0015] The beneficial effects of this invention are: This invention first generates multiple sets of first candidate catalytic reaction systems based on the design requirements of water treatment catalysts. Since there may be a large number of these first candidate catalytic reaction systems, this invention further screens them based on preset screening rules. These screening rules include at least one of the following: screening based on similarity to known measured catalytic reaction systems; screening based on reaction condition constraints; screening based on catalyst structure constraints and / or active free radical constraints that are compatible with pollutant types; and screening based on catalyst structural feasibility. Based on these screening rules, first candidate catalytic reaction systems with low similarity to known measured catalytic reaction systems, or those that do not meet reaction condition constraints, or those containing catalysts and / or active free radicals with low compatibility with pollutant types, or those whose catalyst structural feasibility does not meet the requirements, can be eliminated. This results in second candidate catalytic reaction systems with higher feasibility and superior catalytic performance.

[0016] Then, the catalytic performance of the second candidate catalytic reaction system is predicted based on the trained catalytic performance prediction model, and the target catalytic reaction system is obtained by further narrowing down the scope based on the prediction results, so as to facilitate subsequent experimental verification by experimental personnel based on the target catalytic reaction system.

[0017] In summary, this invention can rapidly generate the target catalytic reaction system based on the design requirements of water treatment catalysts, reducing the design work of experimental personnel, reducing experimental verification time and economic costs, and better cope with complex and ever-changing water quality and process conditions. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A schematic flowchart of an embodiment of the machine learning-based water treatment catalyst design method provided by the present invention; Figure 2 For the present invention Figure 1 A flowchart illustrating an embodiment of S102; Figure 3 This invention provides a prediction of cobalt-based spinel catalysts with different combinations of A-site and B-site elements under the same experimental conditions. k obs Heatmap of values; Figure 4 A flowchart for model construction and analysis provided by the present invention; Figure 5 An overall flowchart of a solution provided by the present invention; Figure 6 This invention provides a comparison chart of model prediction performance and experimental measurement performance. Figure 7 This is a schematic diagram of an embodiment of the machine learning-based water treatment catalyst design device provided by the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0021] In the description of the embodiments of this invention, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following associated objects have an "or" relationship.

[0022] In the embodiments of this invention, the terms "first," "second," etc., are used to distinguish similar objects, and are not used to describe a specific order or sequence, nor to indicate or imply their relative importance or implicitly specify the number of technical features indicated. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, and the number of objects is not limited; for example, the first object can be one or more.

[0023] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0024] Reference Figure 1 The diagram illustrates a flow chart of an embodiment of the machine learning-based water treatment catalyst design method provided by the present invention. The method includes: S101, based on the design requirements of the water treatment catalyst, generates multiple first candidate catalytic reaction systems.

[0025] The design requirements for water treatment catalysts can be input by the user. For example, the user can input the query text "Which catalyst is suitable for drinking water treatment scenarios?", "Which water treatment scenarios is catalyst A suitable for?", or "Is catalyst A suitable for dyeing and printing wastewater treatment scenarios?", etc. The user-input query text can be converted into water treatment catalyst design requirements information in computer language through a large language model. Then, multiple sets of first-line candidate catalytic reaction systems can be generated based on the design requirements information.

[0026] The first candidate catalytic reaction system may include a catalyst, active free radicals, reaction conditions, contaminants, and an oxidant. The catalyst may be a transition metal spinel oxide. Active free radicals may include hydroxyl radicals, sulfate radicals, peroxide radicals, and non-free radicals. Reaction conditions may include temperature, pH value in water, catalyst dosage, and oxidant dosage. Contaminants may include sulfonamides, halogenated phenols, dyes, fluoroquinolones, pesticides, and tetracyclines. Oxidants may include H₂O₂, PMS, PS, and perPAA.

[0027] S102, according to the preset screening rules, multiple first candidate catalytic reaction systems are screened to obtain at least one second candidate catalytic reaction system; wherein, the screening rules include at least one of the following: screening based on the similarity requirement with the measured catalytic reaction system, screening based on reaction condition constraints, screening based on catalyst structure constraints and / or active free radical type constraints that are compatible with the pollutant type, and screening based on catalyst structure feasibility.

[0028] Based on the screening based on the similarity requirement with the measured catalytic reaction system, the first candidate catalytic reaction system with low similarity can be eliminated.

[0029] Screening based on reaction condition constraints can eliminate first-candidate catalytic reaction systems that do not meet the reaction condition constraints.

[0030] Screening based on catalyst structure constraints that are compatible with pollutant types can eliminate first-candidate catalytic reaction systems that contain catalysts with low compatibility with pollutant types.

[0031] Screening based on the constraint of active free radical types that are compatible with pollutant types can eliminate first-candidate catalytic reaction systems that contain active free radicals with low compatibility with pollutant types.

[0032] Based on the screening of catalyst structural feasibility, the first candidate catalytic reaction system that does not meet the requirements of catalyst structural feasibility can be eliminated.

[0033] All of the above filtering methods can be used, or selective methods can be used based on the specific content of the query text entered by the user.

[0034] S103, predict the catalytic performance of the second candidate catalytic reaction system based on the trained catalytic performance prediction model, and determine the target catalytic reaction system based on the prediction results.

[0035] The catalytic performance prediction model can be an XGBoost model, a neural network model, etc. This embodiment does not impose specific limitations on the catalytic performance prediction model. The catalytic performance prediction model can be trained using measured catalytic reaction system data.

[0036] After predicting the catalytic performance of the second candidate catalytic reaction system using the catalytic performance prediction model, the second candidate catalytic reaction systems can be ranked according to their catalytic performance, and the second candidate catalytic reaction system ranked higher can be selected as the target catalytic reaction system.

[0037] The machine learning-based water treatment catalyst design method provided in this embodiment can be applied to a machine learning-based water treatment catalyst design system. This system can be a software system running on a terminal device. The terminal device can be a tablet computer, in-vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), mobile phone, etc. This embodiment does not impose any restrictions on the specific type of terminal device.

[0038] In summary, this invention can rapidly generate the target catalytic reaction system based on the design requirements of water treatment catalysts, reducing the design work of experimental personnel, reducing experimental verification time and economic costs, and better cope with complex and ever-changing water quality and process conditions.

[0039] In some embodiments of the present invention, S101 includes: A variety of catalysts are generated based on a pre-defined catalyst element library, stoichiometry, and valence state constraints; Based on a variety of catalysts, and by combining multiple factors such as the type of active free radicals, reaction conditions, and types of pollutants, multiple first-candidate catalytic reaction systems were obtained.

[0040] Specifically, the element library is defined as follows: A-site (tetrahedral site) screening for elements with closed-shell electronic configurations (such as Li). + [He], Zn 2 + [Ar]3d 10 Metal ions with weak orbital hybridization characteristics (Li, Zn, Mg, Fe) are screened for transition metal ions (Co, Ni, Mn) with tunable electronic structures at the B site (octahedral site).

[0041] Stoichiometry and valence state constraints: Satisfy the stoichiometric ratio and charge balance of AB₂O₄ (A preferred). 2+ / B 3+ Price state combinations, such as Zn 2+ Co2 3+ O4); Multi-factor combination: Based on multiple catalysts, combined with experimental conditions (pH: 6–8, oxidant dosage: 50–200 mg / L), free radical types (mainly sulfate free radicals), and pollutant types (mainly sulfonamides), 20-dimensional features are automatically generated for each A / B element combination, ultimately forming a design space containing 1000+ candidates (first candidate catalytic reaction system).

[0042] In some embodiments of the present invention, such as Figure 2 As shown, step S102 includes: S201, when the screening rule is based on the similarity requirement with the measured catalytic reaction system, calculate the similarity between the first candidate catalytic reaction system and each measured catalytic reaction system.

[0043] S202, determine whether the maximum similarity is less than the preset maximum similarity threshold, and whether the average similarity is less than the preset average similarity threshold.

[0044] S203, if the maximum similarity is greater than or equal to the preset maximum similarity threshold, and the average similarity is greater than or equal to the preset average similarity threshold, then the first candidate catalytic reaction system is retained.

[0045] S204. If the maximum similarity is less than the preset maximum similarity threshold, or the average similarity is less than the preset average similarity threshold, then the catalytic performance of the first candidate catalytic reaction system is predicted by the catalytic performance prediction model to obtain the confidence interval of the prediction result.

[0046] S205, the confidence interval is widened according to the preset rules. If the lower limit of the widened confidence interval is less than the preset catalytic activity threshold, the first candidate catalytic reaction system is screened out.

[0047] S206. If the lower limit of the widened confidence interval is greater than or equal to the preset catalytic activity threshold, then the first candidate catalytic reaction system is retained.

[0048] The similarity between the first candidate catalytic reaction system and each measured catalytic reaction system can be represented by the Tanimoto similarity. The maximum similarity threshold and the mean similarity threshold can be 0.3620 and 0.3335, respectively.

[0049] This embodiment essentially defines the "range of feature combinations that have been fully learned" by the catalytic performance prediction model. When the Tanimoto similarity of the first candidate catalytic reaction system is lower than the threshold, it indicates that its feature combination differs significantly from the training samples of the catalytic performance prediction model, and the model has not fully learned the "feature-performance" correlation of the first candidate catalytic reaction system. In this case, the confidence interval is widened, such as from the original [0.28, 0.32] min -1 Widen to [0.22, 0.38] min -1The uncertainty is quantified by using a widened "confidence interval"—the wider the interval, the lower the reliability—providing a clear quantitative basis for subsequent decision-making. Furthermore, if the lower limit of the widened confidence interval is less than the preset catalytic activity threshold, such as a confidence interval of [0.08, 0.35] min... -1 Covering the low activity range <0.1 min -1 If the model cannot determine whether the first candidate catalytic reaction system is a highly active catalyst, it is judged as "not recommended to enter experimental verification" and screened out to avoid waste of catalyst synthesis (such as raw materials and calcination costs of sol-gel method) and performance testing (such as EPR and HPLC detection costs) costs due to unreliable prediction.

[0050] In some embodiments of the present invention, the reaction condition constraints may include pH range constraints, oxidant dosage range constraints, and catalyst dosage range constraints. For example, constraints may include pH 6–8, catalyst dosage ≤200 mg / L, and oxidant dosage ≤7 g / L.

[0051] In some embodiments of the present invention, the pH range constraint can be determined based on a marginal effect curve plotted for the pH value. Specific methods include: At preset pH value quantiles, using pH value as the quantifier and other characteristics as variables, multiple sets of third candidate catalytic reaction systems were generated. The catalytic performance of multiple third-candidate catalytic reaction systems was predicted using a catalytic performance prediction model. Based on the catalytic performance at each pH value quantile, a marginal effect curve for pH value was plotted, and a confidence band was superimposed on the marginal effect curve. The pH range constraints are determined based on the marginal effect curve and confidence band.

[0052] For example, the pH value ranges from 3 to 11, which can be used to divide the model into 15 quantiles. Then, a single core feature is fixed as the target quantile, and the remaining 19 input features (the catalytic performance prediction model includes a total of 20 input features, and the output feature is the Cui Ahu reaction rate constant) are used. k obs Monte Carlo random sampling (repeated 1000 times) is performed according to the "sample empirical distribution." The sampling process strictly follows the probability distribution of each feature in the model's training samples to avoid analytical errors caused by sampling bias. Then, the 20-dimensional feature combination of each sampling is input into the XGBoost model (catalytic performance prediction model) to obtain... k obs The predicted value is taken as the average of 1000 predicted values ​​as the marginal value of the target quantile. k obsThe value is used to eliminate the interference of other feature fluctuations on the single feature. Then, based on the pH value quantile and the corresponding margin of the quantile, k obs The marginal effect curve was plotted based on 1000 Monte Carlo samplings. k obs Predicted values, calculate standard deviation, and use "marginal" values. k obs The 95% confidence band was determined by "value ± 1.96 × standard deviation", and this confidence band was superimposed on the marginal effect curve plotted above.

[0053] Analysis of the marginal effect curve based on pH value showed that when pH was 6-8... k obs The highest pH value has the narrowest confidence band, so the pH range constraint can be set to pH values ​​of 6–8.

[0054] In some embodiments of the present invention, the input features of the catalytic performance prediction model can also be analyzed using the SHAP method to obtain the SHAP value of each input feature. Based on the SHAP values, the core input features are determined, and a Partial Dependence Plot (PDP) of the core input features is plotted (the curve in the PDP plot represents the marginal effect curve). This facilitates further analysis of the core input features by experimental personnel.

[0055] For example, the catalytic performance prediction model includes a total of 20 input features, including 7 A-site features, 7 B-site features, 3 experimental condition features, as well as pollutant type, active free radical type, and oxidant type. SHAP analysis is used to calculate the impact of all features on the catalytic activity index (…). k obs The contribution weights of SHAP are used to obtain the top 10 key features globally and their ranking (SHAP contribution percentage): No. 1 catalyst dosage (approximately 12%) NO.2 Covalentity of B site (electronic structure characteristics of B site, approximately 10%) NO.3B site d-band center (B site electronic structure characteristics, approximately 9%) NO.4 Oxidant dosage (experimental conditions, approximately 8%) NO.5 pH value (experimental condition characteristics, approximately 7%) No. 6 Free radicals (approximately 6%) NO.7A-B site covalent difference (electronic structure interaction feature, approximately 5%) No. 8 pollutant type (approximately 4%) Ionization energy of site NO.9A (A-site characteristic, approximately 3%) NO.10B site ionic radius (B site structural characteristics, approximately 2%) The core input features can be determined by SHAP sorting of the 20-dimensional input features.

[0056] For example, the core input features are identified as the B-site d-band center, pH value, free radical type, and pollutant type. The B-site d-band center and pH value are both numerical features, and the marginal effect curve for the B-site d-band center is plotted following the same process as for the pH value. Free radical type and pollutant type are both categorical features. When plotting the marginal effect curve, a single core feature can be fixed as the target category, while the remaining 19 input features are randomly sampled. The subsequent process can also refer to the marginal effect curve plotting process for the pH value.

[0057] Marginal effect curves provide a direct visual observation of the impact of a single input feature on catalytic performance. For example, when the fixed free radical type is sulfate free radical, the marginal effect curve can be observed after 1000 samplings. k obs The mean is 0.32 min. -1 When the fixed type of free radical is equal to non-free radical, the marginal... k obs The mean is 0.24 min. -1 This directly reflects the differences in the impact of different types of free radicals on activity.

[0058] In some embodiments of the present invention, the SHAP values ​​or partial dependency graphs of the input features can be visualized. By revealing the contribution and trend of the input features to catalytic performance through the SHAP and PDP methods, interpretability of catalyst performance prediction is achieved.

[0059] In some embodiments of the present invention, based on the SHAP values ​​of each input feature, it can be determined that the B-site electronic structure controls the catalytic activity and the free radical type / pollutant type assists in the tuning.

[0060] In some embodiments of the present invention, the machine learning-based water treatment catalyst design method further includes: Using elemental combinations of elemental types at tetrahedral and octahedral sites in spinel, types of active free radicals, and types of pollutants as variables, multiple sets of fourth candidate catalytic reaction systems were constructed; the catalytic performance of these multiple fourth candidate catalytic reaction systems was predicted using a catalytic performance prediction model. Based on the prediction results of multiple fourth candidate catalytic reaction systems, determine the elemental combinations and / or types of active free radicals suitable for each pollutant type.

[0061] For example, refer to Figure 3For combinations of spinel A-site (preferably Li, Zn, Mg, Fe, Co, etc.), B-site (preferably Co, Ni, Mn, Al, Fe, etc.), free radical types (such as sulfate radicals / hydroxyl radicals), and contaminant types (such as sulfamethoxazole / rhodamine B), multiple candidate systems were constructed and predictions were made. k obs According to predictions k obs Drawing Figure 3 .Depend on Figure 3 It can be seen that when the pollutant type = sulfonamides / rhodamine B, the suitable A / B site combination (such as Co at A site and Li / Zn at B site) is appropriate.

[0062] Similarly, for industrial wastewater treatment scenarios (including halogenated phenols and pesticides), the analysis yields suitable A / B site combinations (e.g., Mg at A site and Ni at B site) when the pollutant type = halogenated phenols / pesticides.

[0063] In some embodiments of the present invention, S102 includes: When the screening rule is based on the feasibility of catalyst structure, it is determined whether the catalyst structure in the first candidate catalytic reaction system is the same as the known and confirmed stable catalyst structure. If the catalyst structure in the first candidate catalytic reaction system is the same as the structure of a known and confirmed stable catalyst, then the first candidate catalytic reaction system is retained. If the catalyst structure in the first candidate catalytic reaction system is different from the known and confirmed stable catalyst structure, then the first candidate catalytic reaction system is eliminated.

[0064] In some embodiments of the present invention, the machine learning-based water treatment catalyst design method further includes: Obtain the experimental results of the target catalytic reaction system; Determine whether the deviation between the predicted and measured results of the target catalytic reaction system is less than a preset deviation threshold; If the deviation is less than the preset deviation threshold, the catalytic performance and result characterization of the catalyst in the target catalytic reaction system in various application scenarios are obtained, and the correlation between the catalyst and the catalytic performance and result characterization in various application scenarios is established.

[0065] In this embodiment, from the perspective of the experimenter, the verification of the target catalytic reaction system includes three layers: ① Layer 1 (Rapid Validation): Select the Top 5–10 highly predictive catalytic reaction systems (target catalytic reaction systems) and determine their activity through a simplified batch experiment (pollutant + dominant free radical system). k obs Preliminary verification of activity trends; ② Layer 2 (Extended Validation): For layers 1 that satisfy "prediction and actual measurement"... k obs For candidates with a deviation of <20%, conduct multi-condition (different pH, pollutant type, free radical type) verification to confirm cross-scenario stability; ③ Layer 3 (Core formulation establishment): Select 1–3 core formulations with “high activity + high stability” (such as ZnCo2O4) from Layer 2, perform structural characterization such as XRD and XPS, deepen the “structure-performance-scenario” mechanism correlation, and provide a basis for subsequent mass production of catalysts.

[0066] In some embodiments of the present invention, the catalytic performance prediction model is trained in the following manner: Training samples containing input and output features were constructed based on the experimentally measured catalytic reaction system. Input features include intrinsic catalyst characteristics, types of active free radicals, reaction condition characteristics, and types of contaminants. Intrinsic catalyst characteristics include at least one of the following: element type, covalence, d-band center, number of valence electrons, electronegativity, ionic radius, ionization energy, and cutoff radius at the spinel tetrahedral and octahedral sites. Reaction condition characteristics include at least one of the following: pH value, oxidant dosage, and catalyst dosage. The output feature is the catalytic reaction rate constant. The catalytic performance prediction model is trained using training samples until it converges.

[0067] Specifically, the training sample construction process can include data collection, data cleaning, and data preprocessing. The data collection process for the experimental catalytic reaction system can include: focusing on the catalytic activity of cobalt-based spinel catalysts (general formula AB₂O₄) in advanced oxidation systems (PAA / PMS / PS), with the target variable defined as the apparent reaction rate constant. k obs (unit: min) -1 The data sources consist of two parts: first, validated experimental data from publicly available peer-reviewed literature from 2015–2024 (including some catalyst structural parameters, experimental conditions, and kobs); and second, cobalt-based spinel catalytic data obtained through our own experiments (recording key information such as free radical types and contaminant types). Simultaneously, density functional theory (DFT) calculations were used to supplement electronic structure features (such as A / B site d-band centers and MO covalentity), ultimately forming a raw dataset covering "catalyst-reaction conditions-contaminants-free radicals-performance".

[0068] The data cleaning process may include: handling missing and outlier values ​​and removing collinearity. Features with a missing rate >30% are deleted; those ≤10% are handled using the median; those with a missing rate between 10% and 30% are imputed using quantiles according to the data distribution; strongly skewed variables are first subjected to a logarithmic Box-Cox transformation, and then outliers are identified and deleted if |z|>3; when the Pearson correlation between two features |r|>0.90, only those more correlated with Kobs are retained. After cleaning, an effective data matrix of n×20 (n is the number of samples, 20 is the number of input features) is retained.

[0069] Data preprocessing may include: standardizing continuous features (such as oxidant dosage, A-position d-band center) using Z-score (mean=0, variance=1); normalizing categorical features (oxidant type, free radical type, pollutant type) to the [0,1] interval using Min-Max normalization; and applying the target variable... k obs First, perform Min-Max normalization and then log(1+y) transformation to reduce the impact of extreme values; finally, randomly divide the training set and test set into an 85%:15% ratio (fixed random seed = 21 to ensure reproducibility).

[0070] The model building process can include: model selection, hyperparameter optimization, and training. Model selection includes choosing models that have advantages in handling nonlinear interactions of high-dimensional features and suppressing overfitting, such as the XGBoost model, which is suitable for multi-scale correlation modeling of "electronic structure-elemental properties-experimental conditions-pollutants-free radicals-catalytic performance".

[0071] Hyperparameter optimization includes: using Bayesian optimization for global hyperparameter search, selecting a tree-structured Parzen estimation surrogate model, and minimizing the test set RMSE and R0. 2 The objective function is "maximum". The hyperparameter search space is set as follows: learning rate 0.01~0.30, maximum tree depth 3~15, minimum child node weight 1~10, and regularization parameter gamma 0~1. The optimal parameter combination is obtained after 1000 iterations.

[0072] Model training includes: training the XGBoost model with the training set based on the optimal hyperparameters, introducing 5-fold cross-validation and an early stopping strategy (stopping training when the RMSE of the validation set does not decrease for 10 consecutive rounds) to avoid overfitting the model to the 20-dimensional features, and finally obtaining a stable Kobs prediction model.

[0073] The XGBoost model trained in this embodiment is subjected to the coefficient of determination R. 2 Evaluate with root mean square error (RMSE).

[0074]

[0075]

[0076] In the formula, Indicating reality , Indicates the model prediction .

[0077] Experiments show that the trained XGBoost model has high prediction accuracy, with the test set R... 2 With an accuracy of approximately 0.87 and an RMSE of only 0.19, the predicted and experimental values ​​are distributed along y=x in the scatter plot, effectively fitting the experimental results.

[0078] The model prediction includes: standardizing the 20-dimensional features of the catalytic reaction system to be predicted (A / B site elemental properties, experimental conditions, types of free radicals, types of pollutants, etc.) according to the above preprocessing steps, and then inputting them into the XGBoost model to obtain the normalized data. k obs Predicted values ​​are restored to the original scale Kobs by inverse transformation (first inverse log(1+y) then inverse Min-Max), and a 95% confidence interval is output.

[0079] Reference Figure 4 This diagram illustrates a model construction and analysis flowchart provided by the present invention. Model construction includes training sample construction, training sample partitioning, model training, crossover severity + early stopping training, parameter optimization, SHAP-based input feature analysis, PDP graph drawing, domain determination (similarity mean threshold and maximum similarity threshold determination), and model output.

[0080] Reference Figure 5 The diagram illustrates the overall flowchart of the solution provided by this invention. The solution includes: 1. Data acquisition and feature construction; 2. Model construction and optimization; 3. Model evaluation and interpretation; 4. Catalyst performance prediction in multiple scenarios (determining the A / B site combination and free radical types compatible with pollutants); 5. Catalyst design and screening; 6. Catalyst design experimental verification; and 7. High-performance catalyst output. Each step is described in the above embodiments.

[0081] The present invention will be illustrated below using a specific catalytic reaction system as an example.

[0082] 1. Catalytic reaction system design.

[0083] ① Reaction system: PAA is used as the main oxidant (parallel PMS / PS systems are used as controls). Two typical scenarios are set up: pollutant: sulfamethoxazole 5 mg / L, free radical: sulfate free radical, generated by adding PMS; pH 3–11 (preferably 6–8), catalyst dosage 50–200 mg / L, oxidant concentration 0.5–5 mmol / L; ② Dynamic fitting: A pseudo-first-order dynamic model is adopted, ln(C / C0) = - k obs t Take the linear fit R 2 Interval calculations >0.98 k obs High-frequency sampling within 0–10 min (once every 2 min); ③Characteristics of free radicals / pollutants: Free radical types were identified by EPR (electron paramagnetic resonance), and pollutant concentration changes were determined by HPLC (high performance liquid chromatography).

[0084] 2. Catalyst synthesis.

[0085] The core candidate catalysts (such as ZnCo2O4 and LiCoO2) and the reference catalyst Co3O4 were synthesized using the sol-gel method: metal nitrates (Zn(NO3)2·6H2O, Co(NO3)2·6H2O, etc.) were dissolved in ethylene glycol in stoichiometric proportions, citric acid was added as a complexing agent, and the mixture was stirred at 80°C until it reached a gel state. The mixture was then calcined at 500°C for 4 hours, ground, and passed through a 200-mesh sieve. The spinel crystal phase was verified by XRD (characteristic peaks matched the standard card JCPDS No. 23-1390), and the valence states of A / B site elements (such as Zn) were analyzed by XPS. 2+ Co 3+ ).

[0086] 3. Performance comparison and analysis.

[0087] 4. For example Figure 6 As shown, , , , , Prediction k obs The values ​​are: 0.03762542, 0.04399073, 0.08736625, 0.02916183, and 0.04003407, respectively; Actual measurements... k obs The values ​​are: 0.0215, 0.05162, 0.103192, 0.296, and 0.361, respectively. The unit is min. -1 The deviations between the predictions and the actual measurements are both within an acceptable range.

[0088] Reference Figure 7 The diagram shows a structural schematic of an embodiment of the machine learning-based water treatment catalyst design device provided by the present invention. The device includes: The first candidate catalytic reaction system generation module 701 is used to generate multiple first candidate catalytic reaction systems based on the design requirements of the water treatment catalyst. The first screening module 702 is used to screen multiple first candidate catalytic reaction systems according to preset screening rules to obtain at least one second candidate catalytic reaction system; wherein, the screening rules include at least one of the following: screening based on similarity requirements with the measured catalytic reaction system, screening based on reaction condition constraints, screening based on catalyst structure constraints and / or active free radical type constraints that are compatible with pollutant types, and screening based on catalyst structure feasibility. The second screening module 703 is used to predict the catalytic performance of the second candidate catalytic reaction system based on the trained catalytic performance prediction model, and to determine the target catalytic reaction system based on the prediction results.

[0089] It should be noted that the implementation principles or processes of the above modules can be referred to the aforementioned embodiments of the water treatment catalyst design method based on machine learning, and will not be elaborated here.

[0090] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.

[0091] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A machine learning-based method for designing water treatment catalysts, characterized in that, include: Based on the design requirements of water treatment catalysts, multiple first-candidate catalytic reaction systems were generated. The multiple first candidate catalytic reaction systems are screened according to preset screening rules to obtain at least one second candidate catalytic reaction system; wherein, the screening rules include at least one of the following: screening based on similarity requirements with the measured catalytic reaction system, screening based on reaction condition constraints, screening based on catalyst structure constraints and / or active free radical type constraints that are compatible with pollutant types, and screening based on catalyst structure feasibility. The catalytic performance of the second candidate catalytic reaction system is predicted based on the trained catalytic performance prediction model, and the target catalytic reaction system is determined based on the prediction results.

2. The water treatment catalyst design method based on machine learning according to claim 1, characterized in that, The step of screening the multiple groups of first candidate catalytic reaction systems according to preset screening rules includes: When the screening rule is based on the similarity requirement with the measured catalytic reaction system, the similarity between the first candidate catalytic reaction system and each measured catalytic reaction system is calculated; Determine whether the maximum similarity is less than a preset maximum similarity threshold, and whether the average similarity is less than a preset average similarity threshold; If the maximum similarity is greater than or equal to the preset maximum similarity threshold, and the average similarity is greater than or equal to the preset average similarity threshold, then the first candidate catalytic reaction system is retained. If the maximum similarity is less than the preset maximum similarity threshold, or the average similarity is less than the preset average similarity threshold, then the catalytic performance of the first candidate catalytic reaction system is predicted by the catalytic performance prediction model to obtain the confidence interval of the prediction result. The first candidate catalytic reaction system is screened based on the confidence interval of the prediction results.

3. The water treatment catalyst design method based on machine learning according to claim 2, characterized in that, The step of screening the first candidate catalytic reaction system based on the confidence interval of the prediction results includes: The confidence interval is widened according to a preset rule. If the lower limit of the widened confidence interval is less than the preset catalytic activity threshold, the first candidate catalytic reaction system is screened out. If the lower limit of the widened confidence interval is greater than or equal to the preset catalytic activity threshold, then the first candidate catalytic reaction system is retained.

4. The water treatment catalyst design method based on machine learning according to claim 1, characterized in that, The reaction condition constraints include pH value range constraints, and the input features of the catalytic performance prediction model include pH value and other features; The method further includes: At preset pH value quantiles, using pH value as a quantifier and the other characteristics as variables, multiple sets of third candidate catalytic reaction systems are generated. The catalytic performance of the multiple third candidate catalytic reaction systems is predicted using the catalytic performance prediction model. Based on the catalytic performance corresponding to each pH value quantile, a marginal effect curve for pH value is plotted, and a confidence band is superimposed on the marginal effect curve; The pH value range constraint is determined based on the marginal effect curve and the confidence band.

5. The water treatment catalyst design method based on machine learning according to claim 1, characterized in that, The water treatment catalyst is a transition metal spinel oxide; the method further includes: Using the elemental combinations of elemental types at the tetrahedral and octahedral sites of spinel, the types of active free radicals, and the types of pollutants as variables, multiple sets of fourth candidate catalytic reaction systems were constructed. The catalytic performance of the multiple fourth candidate catalytic reaction systems is predicted using the catalytic performance prediction model. Based on the prediction results of the multiple sets of fourth candidate catalytic reaction systems, determine the element combinations and / or types of active free radicals suitable for each pollutant type.

6. The machine learning-based water treatment catalyst design method according to claim 1, characterized in that, The step of screening the multiple groups of first candidate catalytic reaction systems according to preset screening rules includes: When the screening rule is based on the feasibility of the catalyst structure, it is determined whether the catalyst structure in the first candidate catalytic reaction system is the same as the known and confirmed stable catalyst structure. If the catalyst structure in the first candidate catalytic reaction system is the same as the structure of a known and confirmed stable catalyst, then the first candidate catalytic reaction system is retained. If the catalyst structure in the first candidate catalytic reaction system is different from the known and confirmed stable catalyst structure, then the first candidate catalytic reaction system is eliminated.

7. The machine learning-based water treatment catalyst design method according to claim 1, characterized in that, The method further includes: Obtain the measured results of the target catalytic reaction system; Determine whether the deviation between the predicted result and the measured result of the target catalytic reaction system is less than a preset deviation threshold; If the deviation is less than a preset deviation threshold, the catalytic performance and result characterization of the catalyst in the target catalytic reaction system in various application scenarios are obtained, and the correlation between the catalyst and the catalytic performance and result characterization in the various application scenarios is established.

8. The machine learning-based water treatment catalyst design method according to claim 1, characterized in that, The water treatment catalyst is a transition metal spinel oxide; the catalytic performance prediction model is trained in the following manner: A training sample containing input and output features was constructed based on the measured catalytic reaction system. The input features include intrinsic catalyst features, types of active free radicals, reaction condition features, and types of contaminants. The intrinsic catalyst features include at least one of the following: element type, covalence, d-band center, number of valence electrons, electronegativity, ionic radius, ionization energy, and cutoff radius at the spinel tetrahedral and octahedral sites. The reaction condition features include at least one of the following: pH value, oxidant dosage, and catalyst dosage. The output feature is the catalytic reaction rate constant. The catalytic performance prediction model is trained using the training samples until the catalytic performance prediction model converges.

9. The machine learning-based water treatment catalyst design method according to claim 8, characterized in that, The method further includes: Generate SHAP values ​​or partial dependency graphs of the input features and visualize them.

10. A machine learning-based water treatment catalyst design device, characterized in that, include: The first candidate catalytic reaction system generation module is used to generate multiple first candidate catalytic reaction systems based on the design requirements of water treatment catalysts. The first screening module is used to screen the multiple groups of first candidate catalytic reaction systems according to preset screening rules to obtain at least one second candidate catalytic reaction system; wherein, the screening rules include at least one of the following: screening based on similarity requirements with the measured catalytic reaction system, screening based on reaction condition constraints, screening based on catalyst structure constraints and / or active free radical type constraints that are compatible with pollutant types, and screening based on catalyst structure feasibility. The second screening module is used to predict the catalytic performance of the second candidate catalytic reaction system based on the trained catalytic performance prediction model, and to determine the target catalytic reaction system based on the prediction results.