An anaerobic ammonia oxidation activity prediction method and system based on a tabPFN model
Patent Information
- Application Number
- CN202610788936.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-03
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-06-03
AI Technical Summary
传统SAA检测依赖实验室测定,周期长、成本高、无法实时反馈;常规机器学习模型在处理连续分类异构表格数据时适配性差,对复杂交互关系学习能力不足,预测精度与泛化能力有限,难以满足工程实时调控需求
本发明针对SAA的异构表格数据特征,能够同时处理连续变量与分类变量,提高对复杂运行数据的适配能力;
Smart Images

Figure CN122324983B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of TabPFN model technology, and in particular to a method and system for predicting anaerobic ammonia oxidation activity based on the TabPFN model. Background Technology
[0002] Anaerobic ammonia oxidation (Anammox) is a highly efficient and low-carbon biological nitrogen removal technology for wastewater. It boasts advantages such as low aeration energy consumption, low sludge production, and no need for external carbon sources, making it widely applicable in the treatment of high-ammonia-nitrogen wastewater. Anaerobic ammonia oxidation activity (SAA) is a core indicator characterizing the metabolic capacity of anaerobic ammonia-oxidizing bacteria and the system's operational status, and is of great significance for reactor performance evaluation, fault diagnosis, and process optimization.
[0003] In actual operation, SAA is affected by temperature, pH, HRT, and Fe. 2+ Fe 3+ Continuous parameters such as VSS and nitrogen conversion ratio, as well as categorical variables such as sludge morphology, process type, and reactor type, interact with each other, exhibiting strong nonlinearity and high coupling characteristics. Traditional SAA detection relies on laboratory testing, which is time-consuming, costly, and cannot provide real-time feedback. Conventional machine learning models have poor adaptability when processing continuous categorical heterogeneous tabular data, lack the ability to learn complex interaction relationships, and have limited prediction accuracy and generalization ability, making it difficult to meet the real-time control needs of engineering.
[0004] Therefore, developing an intelligent method that can efficiently process heterogeneous data, accurately predict SAA, and explain key influencing factors is of great value for the stable operation and optimized regulation of anaerobic ammonia oxidation systems.
[0005] Therefore, we designed an anaerobic ammonia oxidation activity prediction method based on the TabPFN (Tabular Prior-data Fitted Network, TabPFN) model to solve the above problems. Summary of the Invention
[0006] The purpose of this invention is to address the shortcomings of existing technologies by proposing a method for predicting anaerobic ammonia oxidation activity based on the TabPFN model.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: A method for predicting anaerobic ammonia oxidation activity based on the TabPFN model includes the following steps: Step S1: Collect the operating data of the anaerobic ammonia oxidation system and construct an original dataset with the anaerobic ammonia oxidation activity (SAA) as the prediction target. The operating data includes heterogeneous tabular data composed of numerical continuous operating parameters and categorical discrete system information. Step S2: Preprocess the original dataset by completing missing value completion, outlier identification and cleaning, and classification feature encoding conversion to obtain a standardized heterogeneous table dataset. Step S3: Divide the standardized heterogeneous tabular dataset into training and testing sets, input them into the TabPFN regression model for training and validation, and use the coefficient of determination R as the basis for the results. 2 The root mean square error (RMSE) and mean absolute error (MAE) are used to evaluate the model performance and determine the optimal prediction model. Step S4: Input the operational data to be predicted into the optimal prediction model and output the predicted value of anaerobic ammonia oxidation activity (SAA). Step S5: Use the SHAP interpretability method to analyze the trained TabPFN regression model, quantify the contribution of each input feature to the SAA prediction result, and identify key driving factors.
[0008] Preferably, in step S1, the numerical continuous operating parameters include: hydraulic retention time (HRT), temperature, pH, and Fe. 2+ Fe 3+ Nitrogen Removal Rate (NRR), Volatile Suspended Solids (VSS), ΔNO2 - / ΔNH4 + ΔNO3 - / ΔNH4 + One or more of the following; the classified discrete system information includes one or more of the following: sludge morphology, process type, and reactor type.
[0009] Preferably, in step S2, missing value completion is implemented using the K-Nearest Neighbors (KNN) algorithm; outlier identification and cleaning are jointly performed using the ECOD anomaly detection algorithm and the Isolation Forest algorithm, including a two-stage joint anomaly identification and cleaning process: Step s1: Use the ECOD anomaly detection algorithm to perform preliminary anomaly identification on the SAA dataset after missing value completion and feature encoding, calculate the degree of anomaly of each sample, and remove samples that deviate significantly from the overall distribution. Step s2: Input the data after ECOD initial screening into the Isolation Forest algorithm for secondary anomaly detection to further identify samples that exhibit abnormal behavior under multivariate combination relationships; Step s3: After the second Isolation Forest detection, the samples judged as abnormal are removed, and the cleaned SAA dataset is finally obtained. This dataset is then used for the subsequent training and validation of the TabPFN regression model. Through this two-stage combined approach, ECOD is primarily used for initial global anomaly screening based on feature distribution, while Isolation Forest is primarily used for multivariate anomaly verification based on sample isolation mechanisms. The sequential execution of these two methods, with their complementary functions, can reduce the impact of misjudgments from a single anomaly detection algorithm on subsequent model training, thereby improving the stability and reliability of the SAA dataset cleaning results.
[0010] The outlier identification and cleaning process employs the ECOD anomaly detection algorithm and the Isolation Forest algorithm in sequence. First, the ECOD anomaly detection algorithm is used to initially screen the dataset for outliers, removing outliers that significantly deviate from the overall distribution. Then, the Isolation Forest algorithm is used to perform a second anomaly detection on the dataset after the ECOD initial screening, removing samples that exhibit abnormal behavior under multivariate combinations, resulting in a cleaned dataset.
[0011] Preferably, in step S2, the classification feature encoding conversion converts the discrete text features of sludge morphology, process type, and reactor type into numerical codes, which are then integrated with continuous operating parameters into heterogeneous tabular data in a unified input format.
[0012] Preferably, in step S3, the ratio of the training set to the test set is 3:1, with the training set used to fit the TabPFN model and the test set used to verify the model's generalization ability.
[0013] Preferably, in step S5, the key driving factors include one or more of VSS, HRT, temperature, pH, and nitrogen conversion ratio, and the nonlinear influence law and sensitive range of each factor on SAA are revealed by the SHAP response curve.
[0014] An anaerobic ammonia oxidation activity prediction system based on the TabPFN model includes a processor, a memory, a data input interface, and a result output interface. The memory stores a computer program executable by the processor and a weight file of the TabPFN-v2.5 pre-trained regression model. When the processor executes the computer program, it implements the following modules: The data acquisition module is used to acquire the operating data of the anaerobic ammonia oxidation system through the data input interface and to construct the original dataset with SAA as the prediction target. The data preprocessing module is used to perform missing value completion, outlier identification and cleaning, continuous variable standardization and categorical variable encoding on the original dataset to obtain a heterogeneous tabular dataset with a unified format. The model training and prediction module is used to load the weight file of the TabPFN-v2.5 pre-trained regression model, input the heterogeneous table dataset into the TabPFN regression model for supervised fitting, and output the SAA prediction value of the sample to be predicted. The interpretability analysis module is used to call the SHAP method to interpret and analyze the trained TabPFN regression model, and output the contribution of each input feature to the SAA prediction result and the identification results of key driving factors. The data acquisition module, data preprocessing module, model training and prediction module, and interpretability analysis module sequentially transmit and process data, and the result output interface is used to output SAA prediction values, model evaluation indicators, and interpretability analysis results.
[0015] This prediction system is a hardware and software integrated system based on computer equipment. The system includes a processor, memory, a data input interface, and a result output interface. The memory stores a computer program executable by the processor, as well as the weight file of the TabPFN-v2.5 pre-trained regression model. When the processor executes the computer program, it sequentially calls the data acquisition module, data preprocessing module, model training and prediction module, and interpretability analysis module to predict anaerobic ammonia oxidation (SAA) activity and identify key driving factors.
[0016] Specifically, the data acquisition module obtains operational data of the anaerobic ammonia oxidation system through the data input interface and constructs an original dataset with SAA as the prediction target; the data preprocessing module performs missing value completion, outlier identification and cleaning, continuous variable standardization, and categorical variable encoding on the original dataset to obtain a heterogeneous tabular dataset in a unified format; the model training and prediction module loads the weight file of the TabPFN-v2.5 pre-trained regression model, inputs the preprocessed heterogeneous tabular data into the TabPFN regression model for supervised fitting, and outputs the SAA prediction value for the sample to be predicted; the interpretability analysis module calls the SHAP method to interpret and analyze the trained TabPFN regression model, outputting the contribution of each input feature to the SAA prediction result and the identification results of key driving factors.
[0017] In the above system, each module is executed in the order of "data input - data preprocessing - model training and prediction - interpretability analysis - result output", and is implemented by the processor executing the computer program in memory.
[0018] The anaerobic ammonia oxidation activity prediction system based on the TabPFN model is implemented by a computer device, which includes a processor, a memory, a data input interface, and a result output interface. The memory stores the computer program and the weight file of the TabPFN-v2.5 pre-trained regression model. The processor is used to execute the computer program and sequentially call the data acquisition module, the data preprocessing module, the model training and prediction module, and the interpretability analysis module.
[0019] The data acquisition module obtains operational data of the anaerobic ammonia oxidation system through the data input interface and constructs a raw dataset with SAA as the prediction target. The data preprocessing module performs missing value completion, outlier identification and cleaning, continuous variable standardization, and categorical variable encoding on the raw dataset to obtain a heterogeneous tabular dataset with a unified format. The model training and prediction module loads the weight file of the TabPFN-v2.5 pre-trained regression model and performs supervised fitting and SAA prediction based on the preprocessed SAA dataset. The interpretability analysis module calls the SHAP method to interpret and analyze the trained TabPFN regression model, outputting the contribution of each input feature to the SAA prediction result and the identification results of key driving factors. The result output interface outputs the SAA prediction value, model evaluation index, and interpretability analysis results.
[0020] Preferably, it includes a readable storage medium capable of storing a computer program, which, when executed by a processor, implements the method steps described above.
[0021] Compared with the prior art, the beneficial effects of the present invention are: This invention addresses the heterogeneous tabular data characteristics of SAA, enabling it to simultaneously process continuous and categorical variables, thereby improving its adaptability to complex operational data. This invention can learn the nonlinear mapping relationship and interaction between SAA and multiple factors, thereby improving prediction accuracy and generalization ability; This invention can identify key driving factors through SHAP, and help reveal the influence of VSS, HRT, temperature, pH and nitrogen conversion ratio on SAA, providing a reference for system operation optimization; This invention can be applied to the activity evaluation, process optimization decision support, and operation management of anaerobic ammonia oxidation systems, and has good engineering application prospects.
[0022] Therefore, by constructing an anaerobic ammonia oxidation activity prediction method based on the TabPFN model, this invention can improve the prediction accuracy and evaluation efficiency of anaerobic ammonia oxidation activity (SAA), identify key driving factors affecting system activity, and thus provide technical support for the identification of the operating status of anaerobic ammonia oxidation systems, process optimization and control, and engineering applications. Attached Figure Description
[0023] Figure 1 This is an overall flowchart of the anaerobic ammonia oxidation activity prediction method based on the TabPFN model of the present invention; Figure 2 This is a schematic diagram comparing the prediction results of TabPFN and other prediction models for anaerobic ammonia oxidation activity (SAA) in an embodiment of the present invention. Figure 3 This is a schematic diagram of input feature importance analysis based on SHAP in an embodiment of the present invention; Figure 4 This is a diagram illustrating the influence of key operating parameters on the prediction results of anaerobic ammonia oxidation (SAA) activity in an embodiment of the present invention. Figure 5 A schematic diagram comparing the predictive performance of different prediction models for anaerobic ammonia oxidation activity (SAA). Figure 6 A schematic diagram illustrating the distribution characteristics of input variables for the SAA dataset; Figure 7 The graph shows the optimization results of the calling parameters and training configuration parameters of the TabPFN regression model using the TPE hyperparameter optimization method. Detailed Implementation
[0024] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0025] Reference Figures 1-7 A method for predicting anaerobic ammonia oxidation (SAA) activity based on the TabPFN model is proposed. By collecting operational data of the anaerobic ammonia oxidation system, a dataset with SAA as the prediction target is constructed. After preprocessing, the dataset is input into the TabPFN regression model for training and prediction to obtain SAA prediction results. The SHAP method is then combined to identify key driving factors affecting anaerobic ammonia oxidation activity.
[0026] It should be noted that the technical contribution of this application does not lie in structural modification or fine-tuning of the underlying weights of the TabPFN-v2.5 pre-trained network itself, but in combining the native prior fitting capability of TabPFN-v2.5 with the specific heterogeneous data system of anaerobic ammonia oxidation activity (SAA) to form a joint process for data preprocessing, feature encoding, hyperparameter optimization, model fitting, performance verification, and interpretability analysis suitable for SAA prediction tasks.
[0027] Specifically, the TabPFN model used in this application is the TabPFN-v2.5 pre-trained regression model. The program directly loads the locally native pre-trained weight file tabpfn-v2.5-regressor-v2.5_default.ckpt to perform the calculation.
[0028] This application does not retrain or fine-tune the underlying prior fitting mechanism, meta-learning inherent network structure, or pre-trained weight parameters of the TabPFN-v2.5 model, nor does it change the model structure, loss function, or core network parameters.
[0029] The training described in this application does not refer to secondary training or fine-tuning of the TabPFN-v2.5 pre-trained model. Rather, it refers to inputting the pre-processed SAA test dataset into the TabPFN regression model while keeping the native pre-trained weights of TabPFN-v2.5 unchanged, and completing the task adaptation for the anaerobic ammonia oxidation activity prediction scenario through a supervised fitting process. Subsequently, the test set or the data to be predicted is input into the fitted model, and the corresponding SAA predicted values are output.
[0030] This application employs a Bayesian optimization method based on a tree-structured Parzen estimator, namely the TPE hyperparameter optimization method, during model construction to optimize the adjustable hyperparameters and training configuration parameters at the model invocation level. TPE optimization uses the root mean square error (RMSE) obtained from cross-validation as the objective function, searching for optimal parameter combinations through multiple rounds of trials, and then using the obtained optimal parameters for subsequent model training and independent test set validation.
[0031] It should be noted that TPE hyperparameter optimization is not a fine-tuning of the underlying pre-trained weights of TabPFN-v2.5. The native pre-trained weights of TabPFN-v2.5 remain unchanged throughout this application. TPE optimization is only used to determine external configuration parameters during model invocation and training / validation processes, such as model invocation parameters, random seed, training / test split ratio, and whether to enable training set subsampling, thereby improving the model's fitting stability and prediction performance on the SAA dataset.
[0032] The SAA dataset processed in this application is not typical tabular data, but rather heterogeneous tabular data containing multiple types of continuous operating parameters and categorical system information. Continuous input variables include ΔNO2. - / ΔNH4 + ΔNO3 - / ΔNH4 + Fe 2+ Fe 3 +Input variables include pH, hydraulic retention time (HRT), temperature, nitrogen removal rate (NRR), and volatile suspended solids (VSS); categorized input variables include sludge morphology, process type, and reactor type.
[0033] The sludge forms include flocculent sludge, granular sludge, and biofilm; the process types include Anammox, PNA, and PDA; and the reactor types include SBR, EGSB, and UASB.
[0034] For the aforementioned SAA dataset, this application performs missing value completion and standardization on continuous variables, and one-hot encoding transformation on categorical variables to form a unified feature matrix suitable for input to the TabPFN-v2.5 regression model. (Refer to...) Figure 6 The diagram illustrates the distribution characteristics of the input variables in the SAA dataset, further demonstrating that the data processed in this application exhibits characteristics such as multi-scale, long-tailed distribution, discrete point distribution, and the coexistence of continuous and categorical variables. Based on these data characteristics, this application constructs a specific data preprocessing, feature encoding, and TabPFN-v2.5 supervised fitting process, rather than a simple replacement of publicly available models.
[0035] In a specific embodiment, the TPE hyperparameter optimization method is used to optimize the calling parameters and training configuration parameters of the TabPFN regression model. The optimization results are as follows: Figure 7 As shown. Cross-validation RMSE was used as the optimization objective, with asterisks marking preferred parameter combinations. When n_estimators=14, fit_mode=low_memory, and memory_saving=True, the model achieved a lower RMSE, thus this was chosen as the preferred configuration for subsequent training and validation. After optimization, the TabPFN regression model was implemented using TabPFNRegressor, loading the local pre-trained weight file tabpfn-v2.5-regressor-v2.5_default.ckpt, and setting ignore_pretraining_limits=True, random seed to 42, and test set ratio to 0.20. After model training, R², RMSE, and MAE were calculated using the measured SAA values and predicted values on the test set. The above TPE optimization was only used to determine the external configuration parameters at the model invocation level and did not involve fine-tuning the underlying pre-trained weights or network structure of TabPFN-v2.5.
[0036] Based on the aforementioned fixed model version, fixed pre-trained weight file, fixed TPE optimization process, fixed random seed, fixed data processing process, and fixed evaluation method, the TabPFN model in this embodiment achieves prediction results of R²=0.94, RMSE=51.58, and MAE=25.15. This technical performance does not stem from a simple replacement of the TabPFN model, but rather from the synergistic effect of the SAA-specific heterogeneous feature system, data preprocessing method, TPE hyperparameter optimization, TabPFN-v2.5's native prior fitting capability, supervised task adaptation process, and independent test set validation method.
[0037] Therefore, while existing technologies may publicly disclose machine learning for predicting SAA or wastewater treatment parameters, they do not publicly disclose the process of unifying heterogeneous features such as SAA continuous operation parameters, nitrogen conversion ratio, metal ion parameters, sludge morphology, process type, and reactor type, combining them with TPE hyperparameter optimization, inputting the TabPFN-v2.5 pre-trained regression model, and completing supervised regression fitting without fine-tuning the underlying weights, thereby obtaining a prediction effect of R²=0.94 and SHAP interpretability results.
[0038] The operational data includes continuous operating parameters and system type information, among which continuous operating parameters include HRT, temperature, pH, and Fe. 2+ Fe 3+ NRR, VSS, ΔNO2 - / ΔNH4 + and ΔNO3 - / ΔNH4 + The classification system information includes sludge morphology, process type, and reactor type.
[0039] Preprocessing included using the KNN method for missing value completion, employing a combination of ECOD and Isolation Forest for outlier identification and cleaning, and encoding categorical features. After model training, the SHAP method was used to interpret and analyze the model's prediction results.
[0040] Example 1:
[0041] Based on the collected and organized operational data of the anaerobic ammonia oxidation system, a raw dataset with SAA (Self-Activated Ammonium Oxidation) activity as the prediction target was constructed. The raw dataset underwent preprocessing, including KNN imputation for missing values, ECOD and IsolationForest methods for outlier identification and cleaning, and encoding and transformation of classification features such as sludge morphology, process type, and reactor type. Subsequently, the processed data was divided into training and testing sets. A TabPFN regression model was constructed based on the training set, and the model's predictive performance was validated using the testing set. The predicted SAA activity of the anaerobic ammonia oxidation system was then output, as shown below. Figure 2 As shown.
[0042] To further verify the applicability and effectiveness of the method of this invention, the TabPFN model was compared and analyzed with prediction models, including Feature-Labeled Self-Attention Network (FT-Transformer), Multilayer Perceptron (MLP), Extreme Gradient Boosting (XGBoost), TabNet, and Lightweight Gradient Boosting Machine (LightGBM). A unified data processing workflow and evaluation metrics were used to compare the prediction results of different models. The evaluation metrics included the coefficient of determination R. 2 The root mean square error (RMSE) and mean absolute error (MAE) of different models are compared in Table 1.
[0043] Table 1 Comparison of prediction performance of different prediction models for anaerobic ammonia oxidation activity (SAA)
[0044] As shown in Table 1, the TabPFN model exhibits the best overall performance in predicting anaerobic ammonia oxidation (SAA) activity, with a determination coefficient R0. 2 The highest values and lowest RMSE and MAE indicate that the method of this invention can adapt well to heterogeneous tabular data containing both continuous and categorical variables, and has high prediction accuracy and good stability. Therefore, it is feasible and effective to use the TabPFN model to predict anaerobic ammonia oxidation activity.
[0045] In this embodiment, the TabPFN regression model is implemented using the TabPFN-v2.5 pre-trained regression model. The program directly loads the local native pre-trained weight file tabpfn-v2.5-regressor-v2.5_default.ckpt for computation. This embodiment does not retrain or fine-tune the underlying prior fitting mechanism, meta-learning inherent network structure, or pre-trained weight parameters of the TabPFN model, nor does it modify the model structure, core network parameters, or loss function. Instead, while retaining the native prior data fitting capability of TabPFN-v2.5, the preprocessed SAA test dataset is input into the TabPFN regression model, and its task adaptation in the anaerobic ammonia oxidation activity prediction scenario is completed through a supervised fitting process.
[0046] In this embodiment, training refers to completing supervised regression fitting using SAA test data while keeping the native pre-trained weights of TabPFN-v2.5 unchanged, rather than performing secondary training or fine-tuning on the pre-trained weights of TabPFN-v2.5.
[0047] In this embodiment, a Bayesian optimization method based on the tree-structured Parzen estimator, namely the TPE hyperparameter optimization method, is employed to optimize the call-level adjustable hyperparameters and training configuration parameters of the TabPFN regression model. TPE optimization uses the root mean square error (RMSE) obtained from cross-validation as the objective function, searches for optimal parameter combinations through multiple rounds of trials, and uses these optimal parameters for subsequent model fitting and independent test set validation.
[0048] It should be noted that TPE hyperparameter optimization does not involve retraining or fine-tuning the underlying pre-trained weights, meta-learning prior network structure, or core network parameters of TabPFN-v2.5. The native pre-trained weights of TabPFN-v2.5 remain unchanged in this embodiment. TPE optimization is only used to determine external configuration parameters during model invocation and training validation processes to improve the model's fitting stability and prediction performance on the SAA dataset.
[0049] In this embodiment, anaerobic ammonium oxidation activity (SAA) is used as the prediction target variable, and ΔNO2 is used as the prediction target variable. - / ΔNH4 + ΔNO3 - / ΔNH4 + Fe 2+ Fe 3+pH, hydraulic retention time (HRT), temperature, nitrogen removal rate (NRR), and volatile suspended solids (VSS) are used as continuous input characteristics; sludge morphology, process type, and reactor type are used as categorized input characteristics. Sludge morphology includes flocculent sludge, granular sludge, and biofilm; process types include Anammox, PNA, and PDA; and reactor types include SBR, EGSB, and UASB.
[0050] For continuous input variables, missing values are first filled in, followed by standardization. For categorical input variables, one-hot encoding is used to convert them into 0 / 1 numerical features. The processed continuous and categorical variables together form the input feature matrix of the TabPFN regression model.
[0051] Furthermore, this embodiment can be achieved through... Figure 6 The diagram illustrating the distribution characteristics of the input variables in the SAA dataset shows the distribution of each input variable. This diagram indicates significant scale and distribution differences between different operating parameters, with some variables exhibiting long-tailed and discrete point distribution characteristics. This suggests that the SAA dataset possesses multi-scale, imbalanced, and heterogeneous tabular data characteristics. Therefore, in this embodiment, before inputting the data into the TabPFN regression model, missing value completion and standardization are performed on continuous variables, and one-hot encoding transformation is performed on categorical variables to form a unified input feature matrix suitable for the TabPFN-v2.5 regression model.
[0052] After TPE hyperparameter optimization, the specific parameters of the TabPFN regression model in this embodiment are set as follows: model type is TabPFNRegressor, model weight file is tabpfn-v2.5-regressor-v2.5_default.ckpt, ignore_pretraining_limits is set to True, random seed is set to 42, test set ratio is set to 0.20, random subsampling of training set is disabled, and all test set samples are used for model performance evaluation.
[0053] After the model training is completed, the coefficient of determination R is calculated using the measured SAA values and the predicted SAA values on the test set. 2 The root mean square error (RMSE) and mean absolute error (MAE) are calculated. Based on the above fixed model version, fixed weight file, fixed TPE optimization process, fixed random seed, fixed data partitioning method, and fixed evaluation metrics, this embodiment obtains the test set R of the TabPFN model. 2 The value was 0.94, the RMSE was 51.58, and the MAE was 25.15.
[0054] Example 2:
[0055] Based on the results of Example 1, the model was further interpreted and analyzed using the SHAP method to identify key driving factors affecting the prediction results of anaerobic ammonia oxidation (SAA) activity. The results of the input feature importance analysis are as follows: Figure 3 As shown, the results of the explanatory analysis of the influence relationship of key operating parameters are as follows: Figure 4 As shown in the figure. Analysis indicates that characteristics such as VSS, HRT, temperature, pH, and nitrogen conversion stoichiometry contribute significantly to the prediction results of anaerobic ammonia oxidation activity (SAA).
[0056] Furthermore, by Figure 4 It is evident that the influence of key operating parameters on the anammox activity (SAA) is not a simple linear relationship, but rather exhibits significant nonlinear response characteristics and interval differences. Some parameters show a positive contribution to the SAA prediction results within specific value ranges, while showing a negative contribution in other ranges, indicating that anammox activity is influenced by multiple factors and involves complex coupling relationships. By analyzing the SHAP response curves and confidence intervals corresponding to each parameter, the direction of influence, degree of contribution, and relative sensitivity range of different operating parameters on SAA can be identified, thus providing a basis for assessing the operating status of the anammox system and optimizing process control.
[0057] Depend on Figures 1-4 As shown in Table 1, this invention, by constructing an anaerobic ammonia oxidation (SAA) activity prediction method based on the TabPFN model, can effectively predict AAA activity and identify key driving factors affecting system activity. This method is suitable for modeling heterogeneous tabular data containing continuous operating parameters and categorized system information, which helps improve the accuracy and efficiency of AAA activity assessment, thus providing technical support for the operation management and optimized control of AAA systems.
[0058] Reference Figure 5 As shown in the figure, the table prior fitting network model exhibits the best overall prediction performance, with a determination coefficient of 0.94, higher than the feature-labeled self-attention neural network model (0.92), the multilayer perceptron model (0.91), the extreme gradient boosting model (0.90), and the table attention network model and lightweight gradient boosting machine model (0.89). Furthermore, the root mean square error (RMSE) and mean absolute error (MAE) of the table prior fitting network model are 51.58 and 25.15, respectively, both lower than the other models. In contrast, the RMS of the feature-labeled self-attention neural network model, the multilayer perceptron model, the extreme gradient boosting model, the table attention network model, and the lightweight gradient boosting machine model are 65.98, 65.92, 67.15, 70.11, and 71.70, respectively, and the MAEs are 27.29, 41.49, 64.72, 90.62, and 105.50, respectively.
[0059] The above results show that the tabular prior fitting network model used in this invention can more accurately fit the relationship between anaerobic ammonia oxidation activity (SAA) and various operating parameters, and has better effects in improving prediction accuracy and reducing prediction error.
[0060] In this invention, the English abbreviations appearing for the first time in the specification are uniformly interpreted as follows: Tabular Prior-data Fitted Network (TabPFN); Specific Anammox Activity (SAA); Hydraulic Retention Time (HRT); Nitrogen Removal Rate (NRR); Volatile Suspended Solids (VSS); K-Nearest Neighbor (KNN) algorithm Empirical Cumulative Distribution-based Outlier Detection (ECOD) algorithm; Isolation Forest algorithm; Tree-structured Parzen Estimator (TPE); Shapley Additive Explanations (SHAP) Coefficient of Determination (R²) Root Mean Square Error (RMSE); Mean Absolute Error (MAE); Feature Tokenizer Transformer (FT-Transformer) is a self-attention network that uses feature tokenization. Multilayer Perceptron (MLP); Extreme Gradient Boosting (XGBoost); TabNet (Table Attention Network); Lightweight Gradient Boosting Machine (LightGBM).
[0061] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for predicting anaerobic ammonia oxidation activity based on the TabPFN model, characterized in that, The work includes the following steps: Step S1: Collect the operating data of the anaerobic ammonia oxidation system and construct an original dataset with the anaerobic ammonia oxidation activity (SAA) as the prediction target. The operating data includes heterogeneous tabular data composed of numerical continuous operating parameters and categorical discrete system information. Step S2: Preprocess the original dataset by completing missing value completion, outlier identification and cleaning, and classification feature encoding conversion to obtain a standardized heterogeneous table dataset. Step S3: Divide the standardized heterogeneous tabular dataset into training and testing sets, input them into the TabPFN regression model for training and validation, and use the coefficient of determination R as the basis for the results. 2 The root mean square error (RMSE) and mean absolute error (MAE) are used to evaluate the model performance and determine the optimal prediction model. Step S4: Input the operational data to be predicted into the optimal prediction model and output the predicted value of anaerobic ammonia oxidation activity (SAA). Step S5: Use the SHAP interpretability method to analyze the trained TabPFN regression model, quantify the contribution of each input feature to the SAA prediction result, and identify key driving factors. In step S1, the numerical continuous operating parameters include: hydraulic retention time (HRT), temperature, pH, and Fe. 2+ Fe 3+ Nitrogen Removal Rate (NRR), Volatile Suspended Solids (VSS), ΔNO2 - / ΔNH4 + ΔNO3 - / ΔNH4 + One or more of the following; The categorized discrete system information includes one or more of the following: sludge morphology, process type, and reactor type; In step S2, missing value completion is achieved using the K-nearest neighbor missing value imputation algorithm; outlier identification and cleaning are jointly performed using the ECOD anomaly detection algorithm and the Isolation Forest algorithm, including a two-stage joint anomaly identification and cleaning process. Step s1: Use the ECOD anomaly detection algorithm to perform preliminary anomaly identification on the SAA dataset after missing value completion and feature encoding, calculate the degree of anomaly of each sample, and remove samples that deviate significantly from the overall distribution. Step s2: Input the data after ECOD initial screening into the Isolation Forest algorithm for secondary anomaly detection to further identify samples that exhibit abnormal behavior under multivariate combination relationships; Step s3: After the Isolation Forest secondary detection, samples judged as abnormal are removed, and the cleaned SAA dataset is finally obtained. This dataset is then used for subsequent TabPFN regression model training and validation.
2. The method for predicting anaerobic ammonia oxidation activity based on the TabPFN model according to claim 1, characterized in that, In step S2, the classification feature encoding conversion converts the discrete text features of sludge morphology, process type, and reactor type into numerical codes, which are then integrated with continuous operating parameters into heterogeneous tabular data in a unified input format.
3. The method for predicting anaerobic ammonia oxidation activity based on the TabPFN model according to claim 1, characterized in that, In step S3, the ratio of training set to test set is 3:
1. The training set is used to fit the TabPFN model, and the test set is used to verify the model's generalization ability.
4. The method for predicting anaerobic ammonia oxidation activity based on the TabPFN model according to claim 1, characterized in that, In step S5, the key driving factors include one or more of the following: VSS, HRT, temperature, pH, and nitrogen conversion ratio. The nonlinear influence of each factor on SAA and its sensitive range are revealed by the SHAP response curve.
5. A predictive system for anaerobic ammonia oxidation activity based on the TabPFN model, characterized in that, The system includes a processor, a memory, a data input interface, and a result output interface. The memory stores a computer program executable by the processor and a TabPFN-v2.5 pre-trained regression model weight file. When the processor executes the computer program, it implements the following modules: The data acquisition module is used to acquire the operating data of the anaerobic ammonia oxidation system through the data input interface and to construct the original dataset with SAA as the prediction target. The operational data includes heterogeneous tabular data consisting of numerical continuous operational parameters and categorized discrete system information; Numerical continuous operating parameters include: hydraulic retention time (HRT), temperature, pH, and Fe. 2+ Fe 3+ Nitrogen Removal Rate (NRR), Volatile Suspended Solids (VSS), ΔNO2 - / ΔNH4 + ΔNO3 - / ΔNH4 + One or more of the following; The categorized discrete system information includes one or more of the following: sludge morphology, process type, and reactor type; The data preprocessing module is used to perform missing value completion, outlier identification and cleaning, continuous variable standardization and categorical variable encoding on the original dataset to obtain a heterogeneous tabular dataset with a unified format. Missing value completion is achieved using the K-nearest neighbor missing value imputation algorithm; outlier identification and cleaning are jointly performed using the ECOD anomaly detection algorithm and the Isolation Forest algorithm, including a two-stage joint anomaly identification and cleaning process: Step s1: Use the ECOD anomaly detection algorithm to perform preliminary anomaly identification on the SAA dataset after missing value completion and feature encoding, calculate the degree of anomaly of each sample, and remove samples that deviate significantly from the overall distribution. Step s2: Input the data after ECOD initial screening into the Isolation Forest algorithm for secondary anomaly detection to further identify samples that exhibit abnormal behavior under multivariate combination relationships; Step s3: After the second Isolation Forest detection, the samples judged as abnormal are removed, and the cleaned SAA dataset is finally obtained. This dataset is then used for the subsequent training and validation of the TabPFN regression model. The model training and prediction module is used to load the weight file of the TabPFN-v2.5 pre-trained regression model, input the heterogeneous tabular dataset into the TabPFN regression model for supervised fitting, and output the SAA prediction value of the sample to be predicted. The interpretability analysis module is used to call the SHAP method to interpret and analyze the trained TabPFN regression model, and output the contribution of each input feature to the SAA prediction result and the identification results of key driving factors. The data acquisition module, data preprocessing module, model training and prediction module, and interpretability analysis module sequentially transmit and process data, and the result output interface is used to output SAA prediction values, model evaluation indicators, and interpretability analysis results.
6. A computer-readable storage medium, characterized in that, The medium contains a computer program, and the computer program can be executed by a processor to implement the steps of the method described in any one of claims 1-4.
Citation Information
Patent Citations
Short-cut denitrification-anaerobic ammonia oxidation system and intelligent regulation and control method thereof
CN119828482A
Foamed light soil proportion intelligent test decision-making method and device
CN122135850A