Rural organic waste compost maturity prediction and key factor analysis method based on machine learning
By constructing a machine learning-based composting maturity prediction model, introducing ventilation variables, optimizing data input and model interpretability, the problem of low prediction accuracy of existing composting models under complex operating conditions is solved, and efficient prediction and process optimization of composting maturity are achieved.
Patent Information
- Application Number
- CN202510353877.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-08-01
AI Technical Summary
The existing composting models have low prediction accuracy under complex operating conditions, ignoring ventilation variables, making it difficult to stabilize control and predict the compost effect, and lacks widespread applicability under a variety of compost conditions.
A composting maturity prediction model based on machine learning is constructed. By collecting a variety of experimental data, introducing ventilation variables, using random forest, artificial neural network and support vector machine algorithms, combined with SHAP analysis, the data input and model interpretability are optimized, and prediction accuracy and applicability are improved.
It improves the accuracy of the prediction of compost maturity and the applicability of the model, shortens the experimental time cost, provides a scientific basis for optimization of compost process, and supports the resource utilization of rural organic waste.
Smart Images

Figure CN120409889A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of resource utilization of solid waste, and specifically relates to a method for predicting the maturity of rural organic waste composting and analyzing key factors based on machine learning, which solves the problems of complex composting process and difficult prediction of fertilizer efficiency. By using a machine learning model, it accurately predicts the indicators of the compost product such as C / N ratio, GI, and IF-C / N, optimizes the composting process, and improves the efficiency and accuracy of compost maturity evaluation. Background Art
[0002] With the development of rural economy, the generation of rural organic waste, including livestock and poultry manure, straw, and kitchen waste, has been continuously increasing. If not properly treated, it will not only cause environmental problems such as water eutrophication and soil pollution, but also lead to the waste of organic resources. Therefore, realizing the efficient and harmless treatment and resource utilization of organic waste has become an important topic for the development of sustainable agriculture and circular economy. Aerobic composting is a widely used biological treatment method that can use microorganisms to degrade organic matter and convert waste into organic fertilizer rich in nutrients. However, traditional composting experimental studies have problems such as long time consumption, high cost, and being greatly affected by various environmental factors, resulting in difficult stable control and prediction of composting effects, which limits its popularization and application in actual production.
[0003] In order to improve the prediction ability and optimization efficiency of the composting process, in recent years, researchers have tried to optimize the composting process through modeling methods. Currently, composting modeling methods are mainly divided into two categories: theoretical models and empirical models. Theoretical models are mainly based on thermodynamics, kinetics, or statistical principles, and describe the physical, chemical, and microbial reactions in the composting process through mathematical equations. Thermodynamic models rely on theories such as material balance and energy conservation to calculate parameters such as temperature and gas exchange rate in the composting process. Kinetic models use first-order or second-order kinetic equations to describe the degradation rate of organic matter, while statistical regression models establish linear or nonlinear regression equations based on experimental data to analyze the relationship between variables. However, these models often rely on strict assumptions, including uniform temperature distribution, stable ventilation rate, etc. The actual composting process is affected by multiple factors such as raw material composition, environmental conditions, and microbial community succession, and has strong nonlinear and dynamic characteristics, making it difficult for theoretical models to maintain high prediction accuracy under complex working conditions. In addition, existing theoretical models often ignore key variables such as ventilation mode, ventilation frequency, and ventilation volume, and these factors directly affect the oxygen supply, temperature regulation, and microbial activity in the composting process, and play a crucial role in compost maturity.
[0004] In contrast, empirical models are based on experimental data and establish empirical relationships between variables through mathematical methods or statistical analysis to optimize the composting process. Among them, machine learning methods can automatically learn patterns in data, handle multivariable non-linear relationships, and improve prediction accuracy. However, existing machine learning research has mainly focused on modeling physicochemical parameters of composting, including temperature, pH, and moisture content, and paid less attention to the role of ventilation variables, resulting in limitations in the applicability of models under different composting conditions. In addition, most existing studies are based on a single dataset, lacking verification of broad applicability under various composting conditions, and lacking in-depth analysis of the interaction and influence mechanism between variables.
[0005] This method is based on real composting experimental data and uses machine learning methods to construct a compost maturity prediction model to improve prediction accuracy, optimize compost management, and reduce the experimental time cost. In the study, multiple key factors affecting compost maturity are selected, and ventilation variables are particularly introduced to construct and evaluate the prediction capabilities of different machine learning models. The effectiveness of the model is verified through experiments to support the optimization of the composting process and provide a scientific basis for the resource utilization of rural organic waste. Summary of the Invention
[0006] The present invention proposes an aerobic composting optimization method for rural organic waste based on machine learning algorithms to improve the accuracy of compost maturity prediction, optimize key influencing factors, and enhance the applicability of the model, providing a scientific basis for the management and control of the composting process. The present invention improves and optimizes on the basis of existing research, mainly including the following aspects:
[0007] (1) Construct a data-driven compost prediction model to improve prediction accuracy: This method is based on the collected composting experimental data and uses machine learning algorithms (random forest, artificial neural network, support vector machine) to establish a prediction model to achieve efficient prediction of key indicators of compost maturity (C / N ratio, seed germination index GI, ratio of initial to final C / N ratio). Compared with traditional empirical analysis methods, it can more accurately describe the complex relationships between multivariables and improve the reliability of prediction.
[0008] (2) Optimize the data collection strategy to improve the generalization ability of the model: Extract key features before and after composting from multiple independent experiments and literature data, rather than relying solely on a single data source, so as to ensure that the model can cover a wider range of working conditions and make the prediction results more generally applicable.
[0009] (3) Optimize the input variables to enhance the applicability of the model: On the basis of existing machine learning research, the data features are further expanded, especially the introduction of ventilation-related variables, including ventilation frequency and average ventilation volume, to more comprehensively consider the impact of ventilation on the composting process.
[0010] (4) Feature importance analysis to enhance model interpretability: Analyze the factors affecting compost maturity using the SHAP method, quantify the contribution of each variable to the prediction result, and help identify key influencing factors.
[0011] By optimizing data input, increasing data coverage, and enhancing model interpretability, a more accurate compost prediction model applicable to various working conditions is constructed, providing technical support for the efficient resource utilization of rural organic waste and promoting the intelligent and refined management of composting technology.
[0012] To achieve the above object, the present invention provides the following technical solution: A method for predicting the maturity of rural organic waste compost and analyzing key factors based on machine learning, the method comprising the following steps:
[0013] Step 1, Data source and collection
[0014] The data of this method is sourced from published compost experiment literature, and the data is extracted and sorted through the tool WebPlotDigitizer, namely Version 4.8, Ankit Rohatgi, https: / / apps.automeris.io / wpd4 / ; The main parameters collected include the initial physical and chemical properties of the compost substrate, namely total organic carbon TOC, total Kjeldahl nitrogen TKN, pH, moisture content, electrical conductivity EC, initial carbon-nitrogen ratio C / N_Initial, external process control parameters, namely turning frequency, aeration continuity, average aeration volume, composting period, namely period and initial temperature, namely temperature_D0, and use the measurement indicators of the final compost maturity, namely C / N, that is, the carbon-nitrogen ratio of the final product, GI, that is, the seed germination index, IF-C / N, that is, the ratio of the initial value to the final value of C / N, as the target variables;
[0015] Among them, ventilation is a key factor affecting microbial activity, oxygen supply, and degradation rate; Different studies report ventilation parameters in different ways, so this method uses a standard formula to calculate aeration continuity and average ventilation rate to ensure data consistency and comparability;
[0016] 1.1 Aeration Continuity
[0017] Aeration continuity reflects the proportion of ventilation time in the total composting time and is calculated by the following formula 1:
[0018]
[0019] Where: N: number of ventilation cycles; T aeration : ventilation time of each cycle, i.e. min; T pause : The pause time of each cycle, i.e. the duration of no ventilation, i.e. min. The parameter value range is 0-1, and the larger the value, the higher the ventilation ratio;
[0020] 1.2 Average ventilation rate, i.e. AverageAerationRate
[0021] To quantify the oxygen supply level during composting, we calculated the average ventilation rate per unit dry matter, kgDM, according to Equation 2:
[0022]
[0023] Where: Q aeration : Ventilation rate during active ventilation phase, i.e. L / min; T aeration : Ventilation duration, i.e. min; T pause : Ventilation interval time, i.e. min; m DM : dry matter mass in the composting system, i.e. kg·DM;
[0024] This parameter standardizes the ventilation supply, making ventilation intensity comparable under different composting conditions;
[0025] 1.3 Conversion of different ventilation rate units
[0026] In different literatures, ventilation rate can be expressed in different ways, including:
[0027] Calculated by unit dry matter mass, i.e. L / min·kgDM;
[0028] By unit volume, i.e. m 3 / min·m 3 calculate;
[0029] Since ventilation rates based on material mass are more common, it is necessary to convert volume units to mass units according to Formula 3:
[0030]
[0031] Where: ρ: bulk density of compost material, i.e. kg / m 3 ;
[0032] Step 2: Data preprocessing
[0033] Data preprocessing is an important step in building machine learning models. This method mainly handles missing values: a deletion strategy is used to eliminate literature data with incomplete experimental methods or lack of key variables to ensure the quality and representativeness of model training data;
[0034] Step 3, Descriptive Statistical Analysis of Data
[0035] Before modeling, perform descriptive statistical analysis on the data, including:
[0036] Pearson correlation analysis: Calculate the Pearson correlation coefficients between input variables and draw a correlation heatmap, i.e., heatmap, to visually display the correlations between variables; these correlation analyses help understand the interactions between variables and provide a basis for model feature selection;
[0037] Step 4, Construction of Machine Learning Model
[0038] This method selects three common machine learning models:
[0039] Random Forest, i.e., RF, Random Forest: An ensemble learning method based on decision trees, with strong non-linear modeling capabilities and the ability to evaluate the importance of variables;
[0040] Artificial Neural Network, i.e., ANN, Artificial Neural Network: Simulates complex non-linear relationships through multi-layer neuron connections and is suitable for processing high-dimensional data;
[0041] Support Vector Machine, i.e., SVM, Support Vector Machine: Utilizes kernel functions to map high-dimensional feature spaces and performs regression prediction by maximizing the margin, suitable for small sample data;
[0042] Step 5, Model Training and Performance Evaluation
[0043] The performance of the model is mainly evaluated through the following two metrics:
[0044] Coefficient of determination: Measures the explanatory power of the model. The closer the value is to 1, the better the fitting effect of the model, as shown in Equation 4:
[0045]
[0046] where: y i represents the actual observed value, represents the model predicted value, is the mean of the actual observed values;
[0047] Root Mean Square Error: Used to measure the prediction error of the model. The smaller the value, the higher the prediction accuracy of the model, as shown in Equation 5:
[0048]
[0049] where: y irepresents the actual observed value, represents the model predicted value, and n is the number of samples;
[0050] Step 6, variable importance analysis
[0051] To further explore the key factors affecting compost maturity, this method uses SHAP, i.e., SHapley Additive exPlanations value analysis, to quantify the contribution of each input variable to the prediction result; the main steps of SHAP analysis are as follows:
[0052] Calculate the SHAP value and draw a SHAP summary plot, i.e., summaryplot, to show the influence degree of all variables on the prediction result;
[0053] Step 7, experimental verification
[0054] To evaluate the applicability of the model in the actual composting process, this method designs and implements an aerobic composting experiment on cow dung - straw mixture, inputs the experimental data into the optimal model, i.e., RF, calculates the predicted value and compares it with the actual experimental value to verify the feasibility of the machine - learning method in compost management.
[0055] Preferably, in step 1, assuming the bulk density is 500 kg / m 3 for calculation, 500 kg / m 3 is between the density ranges of common organic wastes, including livestock manure, food waste, and garden waste, and is a reasonable estimated value for conversion of different data sources.
[0056] Preferably, in step 5, for the prediction performance evaluation, a scatter plot is used to compare the experimental value and the model predicted value to visually analyze the fitting situation of the model, and finally the optimal model is selected for subsequent analysis.
[0057] Preferably, in step 6, combined with the SHAP importance ranking, the variables with the greatest influence on C / N, GI, and IF - C / N are screened out to provide an optimization direction for the actual composting experiment.
[0058] Compared with the prior art, the beneficial effects of the present invention are:
[0059] 1. By integrating data under different experimental conditions and introducing the ventilation variable, a machine - learning compost maturity prediction system based on multi - variable input and data - driven is constructed;
[0060] 2. By optimizing the key variable selection through SHAP analysis, the key influencing factors are identified, providing a reference basis for composting process optimization;
[0061] 3. By experimentally verifying the model prediction results, its reliability and practical application value under different composting conditions are ensured. Brief Description of the Drawings
[0062] Figure 1 Prediction accuracy R of different models of the present invention 2 and RMSE graph;
[0063] Figure 2 Graph of predicted data and actual experimental data of C / N, GI, and IF-C / N based on the optimized machine learning models SVM, ANN, and RF of the present invention. The diagonal line (y = x) represents the reference line, and the points closer to this line indicate higher prediction accuracy;
[0064] Figure 3 SHAP analysis graph of the prediction of C / N (a), GI (b), and IF-C / N (c) based on the random forest model of the present invention;
[0065] Figure 4 Graph of the verification results of the RF model of the present invention for C / N (a), GI (b), and IF-C / N (c) in a real composting experiment. Detailed Embodiment
[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0067] A method for predicting the maturity of rural organic waste composting and analyzing key factors based on machine learning, the method comprising the following steps:
[0068] Step 1, Data Source and Collection
[0069] The data for this method were obtained from published composting experimental literature and were extracted and organized using the tool WebPlotDigitizer (Version 4.8, Ankit Rohatgi, https: / / apps.automeris.io / wpd4 / ). The main parameters collected included the initial physical and chemical properties of the compost matrix, namely total organic carbon (TOC), total Kjeldahl nitrogen (TKN), pH, moisture, electrical conductivity (EC), initial carbon-nitrogen ratio (C / N_Initial), external process control parameters, namely turning frequency, aeration continuity, aeration in average, composting period (period), and initial temperature (temperature_D0). The target variables were the final maturity indicators C / N (the carbon-nitrogen ratio of the final product), GI (the germination index), and IF-C / N (the ratio of the initial to final C / N values).
[0070] Ventilation is a key factor affecting microbial activity, oxygen supply, and degradation rate. Different studies report ventilation parameters in different ways, so this method uses a standard formula to calculate ventilation continuity and average ventilation rate to ensure data consistency and comparability.
[0071] 1.1 Aeration Continuity
[0072] Ventilation continuity reflects the proportion of ventilation time in the total composting process and is calculated by the following formula 1:
[0073]
[0074] Where: N: number of ventilation cycles; T aeration : ventilation time of each cycle, i.e. min; T pause : The pause time of each cycle, i.e. the duration of no ventilation, i.e. min. The parameter value range is 0-1, and the larger the value, the higher the ventilation ratio;
[0075] 1.2 Average ventilation rate, i.e. AverageAerationRate
[0076] To quantify the oxygen supply level during composting, we calculated the average ventilation rate per unit dry matter, kgDM, according to Equation 2:
[0077]
[0078] Where: Q aeration : Ventilation rate during active ventilation phase, i.e. L / min; T aeration: Ventilation duration, i.e., min; T pause : Ventilation interval time, i.e., min; m DM : Dry matter mass in the composting system, i.e., kg·DM;
[0079] This parameter standardizes the ventilation supply, making the ventilation intensity comparable under different composting conditions;
[0080] 1.3 Conversion of different ventilation rate units
[0081] In different literatures, the ventilation rate can be expressed in different ways, including:
[0082] Calculated by unit dry matter mass, i.e., L / min·kgDM;
[0083] Calculated by unit volume, i.e., m 3 / min·m 3 ;
[0084] Since the ventilation rate is more general based on the material mass, it is necessary to convert the volume unit to the mass unit according to formula 3:
[0085]
[0086] Where: ρ: Bulk density of the composting material, i.e., kg / m 3 ;
[0087] Step 2, Data preprocessing
[0088] Data preprocessing is an important step in building a machine learning model. This method mainly performs missing value processing: adopting a deletion strategy to eliminate the literature data with incomplete experimental methods or lack of key variables to ensure the quality and representativeness of the model training data;
[0089] Step 3, Descriptive statistical analysis of data
[0090] Before modeling, perform descriptive statistical analysis on the data, including:
[0091] Pearson correlation analysis: Calculate the Pearson correlation coefficients between each input variable and draw a correlation heatmap, i.e., heatmap, to visually display the correlation between variables; These correlation analyses help to understand the interaction between variables and provide a basis for model feature selection;
[0092] Step 4, Construction of machine learning model
[0093] This method selects three common machine learning models:
[0094] Random Forest, i.e., RF: An ensemble learning method based on decision trees, with strong non-linear modeling capabilities and the ability to evaluate the importance of variables;
[0095] Artificial Neural Network, i.e., ANN: Simulates complex non-linear relationships through multi-layer neuron connections and is suitable for processing high-dimensional data;
[0096] Support Vector Machine, i.e., SVM: Utilizes kernel functions to map high-dimensional feature spaces and performs regression prediction by maximizing the margin, suitable for small-sample data;
[0097] Step 5, Model Training and Performance Evaluation
[0098] The performance of the model is mainly evaluated through the following two metrics:
[0099] Coefficient of determination: Measures the explanatory power of the model. The closer the value is to 1, the better the fitting effect of the model, as shown in Equation 4:
[0100]
[0101] where: y i represents the actual observed value, represents the model predicted value, is the mean of the actual observed values;
[0102] Root Mean Square Error: Used to measure the prediction error of the model. The smaller the value, the higher the prediction accuracy of the model, as shown in Equation 5:
[0103]
[0104] where: y i represents the actual observed value, represents the model predicted value, and n is the number of samples;
[0105] Step 6, Variable Importance Analysis
[0106] To further explore the key factors affecting compost maturity, this method uses SHAP, i.e., SHapley Additive exPlanations value analysis, to quantify the contribution of each input variable to the prediction result; The main steps of SHAP analysis are as follows: [[ID=4,8]]
[0107] Calculate SHAP values and draw a SHAP summary plot, i.e., summaryplot, to show the influence degree of all variables on the prediction result;
[0108] Step 7, Experimental Verification
[0109] To evaluate the applicability of the model in the actual composting process, this method designed and implemented an aerobic composting experiment on a cow manure-straw mixture, input the experimental data into the optimal model, i.e., RF, calculated the predicted values and compared them with the actual experimental values to verify the feasibility of the machine learning method in compost management.
[0110] In this embodiment, in step 1, assuming the bulk density is 500 kg / m 3 for calculation, 500 kg / m 3 is between the density ranges of common organic wastes, including livestock and poultry manure, food waste, and garden waste, and is a reasonable estimated value for the conversion of different data sources.
[0111] In this embodiment, in step 5, the prediction performance was evaluated by using a scatter plot to compare the experimental values with the model predicted values, visually analyzed the fitting situation of the model, and finally selected the optimal model for subsequent analysis.
[0112] In this embodiment, in step 6, combined with the SHAP importance ranking, the variables with the greatest influence on C / N, GI, and IF-C / N were screened out to provide an optimization direction for the actual composting experiment.
[0113] In summary:
[0114] 1. Select key time sections to enhance the applicability of the model
[0115] During the data collection process, based on a large number of published literatures, this method extracted the key data points at the initial stage and the end stage of composting to ensure the representativeness of the data and the wide applicability of the model. Through Pearson correlation analysis, as Figure 1 shown, the correlations between independent variables were evaluated to ensure that there would be no high collinearity among the input features during the modeling process, thereby improving the robustness of the model.
[0116] The analysis results show that the turning frequency is positively correlated with variables such as the composting cycle and the total organic carbon and the initial C / N ratio, while the moisture content is negatively correlated with variables such as the electrical conductivity and the moisture content and the initial C / N ratio, indicating the interactive effects of moisture, oxygen, and organic matter content during the composting process. In addition, the turning frequency has a weak correlation with the ventilation continuity, and the influence of the initial temperature on other variables is small, further verifying the complexity of the composting process.
[0117] 2. Improve the prediction accuracy and shorten the experimental cycle
[0118] Based on the collected experimental data, this method established an efficient compost maturity prediction model through data preprocessing, feature screening, model training, and verification. From the result Table 1 and Figure 2As can be seen, the RF model performs best in the prediction of the three target variables C / N, GI, and IF-C / N, with R 2 values reaching 0.93, 0.91, and 0.97 respectively, indicating that the model can fit the data well and capture the changing trends of key variables during the composting process. At the same time, its RMSE values are the lowest (C / N = 0.16, GI = 1.20, IF-C / N = 0.01), suggesting that the prediction error is small and the model stability is high. In contrast, ANN and SVM also have good prediction capabilities, but they are slightly inferior to RF overall. Especially in the prediction of GI, RF has the lowest RMSE (1.20), indicating that its prediction of the compost maturity index is more accurate. The prediction effects of the three models are far better than those of the traditional statistical regression method, significantly improving the prediction accuracy and reducing the time cost of composting experiments.
[0119] Table 1. Prediction accuracies (R 2 and RMSE) of different models
[0120]
[0121] 3. Variable importance analysis for optimizing compost management
[0122] This method uses the SHAP (Shapley Additive Explanations) method to interpret the model and identify the factors that have the most significant impact on compost maturity, such as Figure 3As shown. The results indicate that in the C / N prediction, the initial C / N ratio has the highest importance, suggesting that the initial carbon-nitrogen ratio is a key factor influencing the change in the final carbon-nitrogen ratio. In addition, the average ventilation rate and ventilation continuity also play important roles, reflecting the impact of ventilation conditions on organic matter degradation and nitrogen loss, while the effects of pH and TKN can be related to the nitrogen transformation process. For the GI prediction, the ventilation variables have the most significant influence, indicating that ventilation management directly affects microbial activity and the germination index of seeds. In addition, EC and TKN are also key variables, reflecting the changes in soluble salts and the impact of nitrogen supply on the change in GI, while the initial carbon-nitrogen ratio and pH also play a certain role in the GI prediction. In contrast, humidity and ventilation continuity have less influence on GI, which may be related to the control of experimental conditions. In the IF-C / N prediction, TKN is the most important factor affecting this index, and the ventilation variables also play an important role in the prediction, indicating that ventilation management has a direct impact on the stability and transformation of nitrogen. In addition, pH and the initial carbon-nitrogen ratio have a moderate influence, while temperature and humidity contribute less, suggesting that the direct impact of the initial temperature and moisture content of the compost on the final nitrogen stability is relatively limited. Overall, this method uses SHAP analysis to quantify the contributions of different variables to the prediction of compost maturity, clarifies the importance of ventilation management in optimizing the composting process, and also verifies the scientificity of the initial carbon-nitrogen ratio and TKN as key influencing factors. The results of this analysis provide data support for optimizing compost management and also provide a basis for subsequent experimental design.
[0123] 4. Verification of Model Reliability by Real Composting Experiments
[0124] To verify the reliability of the proposed prediction model, this method conducted a composting experiment on a cow manure-straw mixture and input the experimental data into the RF model to calculate the deviation between the predicted value and the experimental value. As Figure 4 shown, the experimental value of C / N is 13.08 and the predicted value is 12.56, with a small error, indicating that the model can accurately capture the change trend of C / N. The experimental value of GI is 0.89 and the predicted value is 0.95. The model slightly overestimates GI, which may be related to the complex interactions of some variables, including moisture, EC, and temperature, but the overall error is controlled within a reasonable range. The experimental value of IF-C / N is 1.60 and the predicted value is 1.54, with an extremely small error, indicating that the model has high accuracy in predicting this index. Overall, the random forest (RF) model performs well in the prediction of C / N, GI, and IF-C / N, with small prediction errors, and can effectively capture the change trends of key variables during the composting process for the assessment and management of compost maturity. In the future, the generalization ability and prediction accuracy of the model can be further improved by increasing the data volume or optimizing the input variables.
[0125] This method is based on real composting experimental data and uses machine learning methods to construct a compost maturity prediction model to improve prediction accuracy, optimize compost management, and shorten the experimental time cost.
[0126] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0127] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for predicting the maturity of rural organic waste composting and analyzing key factors based on machine learning, characterized in that: The method comprises the following steps: Step 1: Data Source and Collection The data for this method were obtained from published composting experimental literature and were extracted and organized using the tool WebPlotDigitizer (Version 4.8, Ankit Rohatgi, https: / / apps.automeris.io / wpd4 / ). The main parameters collected included the initial physical and chemical properties of the compost matrix, namely total organic carbon (TOC), total Kjeldahl nitrogen (TKN), pH, moisture, electrical conductivity (EC), initial carbon-nitrogen ratio (C / N_Initial), external process control parameters, namely turning frequency, aeration continuity, aeration in average, composting period (period), and initial temperature (temperature_D0). The target variables were the final maturity indicators C / N (the carbon-nitrogen ratio of the final product), GI (the germination index), and IF-C / N (the ratio of the initial to final C / N values). Ventilation is a key factor affecting microbial activity, oxygen supply, and degradation rate. Different studies report ventilation parameters in different ways, so this method uses a standard formula to calculate ventilation continuity and average ventilation rate to ensure data consistency and comparability. 1.1 Aeration Continuity The ventilation continuity reflects the proportion of ventilation time in the total time during the composting process and is calculated by the following formula 1: Where: N: number of ventilation cycles; T aeration : ventilation time of each cycle, i.e. min; T pause : The pause time of each cycle, i.e. the duration of no ventilation, i.e. min. The parameter value range is 0-1, and the larger the value, the higher the ventilation ratio; 1.2 Average ventilation rate, i.e. AverageAerationRate To quantify the oxygen supply level during composting, we calculated the average ventilation rate per unit dry matter, kgDM, according to Equation 2: Where: Q aeration : Ventilation rate during the active ventilation stage, i.e., L / min; T aeration : Ventilation duration, i.e., min; T pause : Ventilation interval time, i.e., min; m DM : Dry matter mass in the composting system, i.e., kg·DM; This parameter standardizes the ventilation supply, making ventilation intensity comparable under different composting conditions; 1.3 Conversion of different ventilation rate units In different literatures, ventilation rate can be expressed in different ways, including: Calculated by unit dry matter mass, i.e. L / min·kgDM; Calculated per unit volume, i.e., m 3 / min·m 3 ; Since ventilation rates based on material mass are more common, it is necessary to convert volume units to mass units according to Formula 3: Where: ρ: bulk density of the compost material, i.e., kg / m 3 ; Step 2: Data preprocessing Data preprocessing is an important step in building machine learning models. This method mainly handles missing values: a deletion strategy is used to eliminate literature data with incomplete experimental methods or lack of key variables to ensure the quality and representativeness of model training data; Step 3: Descriptive statistical analysis of data Before modeling, descriptive statistical analysis of the data was performed, including: Pearson correlation analysis: Calculate the Pearson correlation coefficient between each input variable and draw a heatmap to visually display the correlation between variables. This correlation analysis helps understand the interaction between variables and provides a basis for model feature selection. Step 4: Machine Learning Model Construction This method selects three common machine learning models: Random Forest, i.e., RF: An ensemble learning method based on decision trees, with strong non-linear modeling capabilities and the ability to evaluate the importance of variables; Artificial Neural Network, i.e., ANN: Simulates complex non-linear relationships through multi-layer neuron connections and is suitable for processing high-dimensional data; Support Vector Machine, i.e., SVM: Utilizes kernel functions to map high-dimensional feature spaces and performs regression prediction by maximizing the margin, suitable for small sample data; Step 5, Model Training and Performance Evaluation The performance of the model is mainly evaluated through the following two metrics: Coefficient of determination: Measures the explanatory power of the model. The closer the value is to 1, the better the fitting effect of the model, as shown in Equation 4: Where: y i represents the actual observed value, represents the model predicted value, is the mean of the actual observed values; Root Mean Square Error: Used to measure the prediction error of the model. The smaller the value, the higher the prediction accuracy of the model, as shown in Equation 5: where: y i represents the actual observed value, represents the model predicted value, and n is the number of samples; Step 6, Variable Importance Analysis To further explore the key factors affecting compost maturity, this method uses SHAP, i.e., SHapley Additive exPlanations value analysis, to quantify the contribution of each input variable to the prediction result; The main steps of SHAP analysis are as follows: Calculate SHAP values and draw a SHAP summary plot, i.e., summaryplot, to show the influence degree of all variables on the prediction result; Step 7, Experimental Verification To evaluate the applicability of the model in the actual composting process, this method designs and implements an aerobic composting experiment on cow manure-straw mixtures, inputs the experimental data into the optimal model, i.e., RF, calculates the predicted values and compares them with the actual experimental values to verify the feasibility of machine learning methods in compost management.
2. A method for predicting the maturity of rural organic waste compost and analyzing key factors based on machine learning according to claim 1, characterized in that: In Step 1, assume the bulk density is 500 kg / m 3 for calculation. 500 kg / m 3 is between the density ranges of common organic wastes, including livestock manure, food waste, and garden waste, and is a reasonable estimated value for conversion of different data sources.
3. A method for predicting the maturity of rural organic waste compost and analyzing key factors based on machine learning according to claim 1, characterized in that: In Step 5, the prediction performance is evaluated by comparing the experimental values with the model predicted values using a scatter plot to visually analyze the fitting situation of the model, and finally the optimal model is selected for subsequent analysis.
4. A method for predicting the maturity of rural organic waste composting and analyzing key factors based on machine learning according to claim 1, characterized in that: In Step 6, combined with the SHAP importance ranking, the variables with the greatest impact on C / N, GI, and IF-C / N are screened out to provide an optimization direction for the actual composting experiment.