Anaerobic fermentation process soft measurement modeling method based on TSSA-CNN-LSTM
Through the improved TSSA-CNN-LSTM model, the problem of insufficient model optimization in the existing technology is solved, efficient and accurate monitoring and early warning of the anaerobic fermentation process is achieved, the stability and wide application of the model are improved, and the intelligent management of the AD process is supported.
Patent Information
- Application Number
- CN202510457642.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-29
AI Technical Summary
The existing machine learning models have limited optimization degree in the anaerobic fermentation process, which is difficult to apply stably on different data sets, and lacks comprehensive system operation warnings and monitoring, which limits their comprehensive application potential in anaerobic fermentation process management.
Using the improved Chaos Sparrow Optimization Algorithm (TSSA) combined with the model of convolutional neural network (CNN) and long and short-term memory network (LSTM), a comprehensive soft measurement model of the anaerobic fermentation process is constructed through cross-validation and parameter optimization, and the key parameters are monitored in real time and abnormal warnings are performed.
It improves the stability and reliability of the model on different data sets, realizes efficient and accurate monitoring and early warning of the anaerobic fermentation process, and supports intelligent management of the AD process.
Smart Images

Figure CN120388639A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of soft measurement of anaerobic fermentation processes, and more specifically, to a soft measurement modeling method for anaerobic fermentation processes based on TSSA-CNN-LSTM. Background Art
[0002] Human production and life require energy. The current energy mix is dominated by non-renewable energy sources such as fossil fuels. Energy security and global energy consumption have become pressing issues for the international community. Growing awareness of climate change, population and economic growth, and the increasing demand for energy are making renewable energy economically attractive and environmentally acceptable. Food waste, containing organic matter such as fat, protein, and carbohydrates, can be treated using anaerobic digestion (AD), a method for treating various organic wastes, preventing pollution, and recovering energy. This process reduces environmental pollution and allows the organic fraction to be recovered as biogas.
[0003] The AD process consists of four interconnected phases: hydrolysis, acidification, acetogenesis, and methanogenesis. Each phase requires the complex and coordinated actions of various microorganisms. Because the biological process involves a series of reactions, AD performance is influenced by operating conditions, microbial type, and abundance. These factors are extremely complex and nonlinear, making it difficult to monitor AD progress.
[0004] With the development of computational algorithms and the availability of computing power, machine learning (ML) has emerged as a new data mining technique and modeling tool, and has been applied to the prediction of AD processes. It can achieve output prediction based on the potential interactions between input and output variables. Predicting AD performance through ML does not require an understanding of the process mechanism and is therefore an effective method for predicting biogas production. Currently, various ML algorithms, including artificial neural networks (ANNs), adaptive neuro-fuzzy interference systems (ANFISs), and random forests (RFs), have been used to model the complex nonlinear relationships of AD progression.
[0005] Previous studies have demonstrated the positive role of machine learning (ML) in predicting and monitoring anaerobic digestion (AD) processes, but there are still deficiencies. First, existing research mainly focuses on the performance evaluation of ML models themselves or the performance comparison between different models, while less attention is paid to how to further improve the prediction ability of ML models through optimization algorithms. This results in limited optimization of the models in practical applications and makes it difficult to fully exploit their prediction potential. Second, most current studies only validate the performance of ML models based on single or specific datasets, and this approach may limit the generalization ability of the models under different scenarios and conditions. Finally, most studies only focus on the prediction of biogas production in the AD process and ignore the importance of building a comprehensive early warning and monitoring system for the operation of the AD system, which to a certain extent limits the comprehensive application potential of ML technology in the management of the AD process.
[0006] To address these deficiencies, the present invention proposes an innovative solution, namely a CNN-LSTM model combined with an improved chaotic sparrow search algorithm (TSSA). The present invention adopts a cross-validation strategy to enhance the stability and reliability of the model on different datasets, thus overcoming the problem of model one-sidedness that may be caused by a single dataset. In addition, the present invention not only focuses on the prediction of biogas production, but is more committed to building a comprehensive early warning and monitoring system for the operation of the AD system. This system can real-time monitor the key parameters in the AD process, give early warnings of abnormal states, and comprehensively evaluate the performance of the system, providing more comprehensive and accurate technical support for the efficient, safe, and intelligent management of the AD process. Summary of the Invention
[0007] The object of the present invention is to overcome the deficiencies of the prior art and provide a soft-sensing modeling method for anaerobic digestion processes based on TSSA-CNN-LSTM.
[0008] The soft-sensing modeling method for anaerobic digestion processes based on TSSA-CNN-LSTM described in the present invention includes the following steps:
[0009] 1) Obtain sampling data of the anaerobic digestion process with a time-tagged sequence through on-site operation or experiments, and extract key feature variables from the collected anaerobic digestion data through a convolutional neural network (CNN) to reduce the dimension of each piece of data;
[0010] 2) Divide the sample data of the extracted basic feature variables into a training set and a test set, and perform normalization processing to eliminate the influence of the dimensions of different feature variables;
[0011] 3) Set the initial parameters of the Sparrow Search Algorithm (SSA), namely the number of iterations and the population size. Initialize the population using the chaotic mapping (Tent), and enhance the local search ability of SSA in the later stage using the non - linear inertia weight; increase the diversity of the population using the horizontal crossover strategy;
[0012] 4) Construct a long - short - term memory network (LSTM) soft - sensing model CNN - LSTM for the anaerobic fermentation process with the normalized data as the input. Use the chaotic sparrow search algorithm (TSSA) to search for the optimal parameters of the CNN - LSTM model, including the number of LSTM units, the learning rate, and the number of attention mechanism (Attention) heads;
[0013] 5) Substitute the obtained optimal parameter set into CNN - LSTM to form a high - precision soft - sensing model TSSA - CNN - LSTM for the anaerobic fermentation process, and predict the methane concentration and biogas production.
[0014] The convolutional layer is the basis of the entire network architecture. During forward propagation, the convolutional kernel establishes connections and shares weights among channels. When the data matrix passes through, the convolutional layer performs a dot - product operation on the data matrix within the defined region through the following formula:
[0015] y l =f(∑x l *W ij +b ij ) (1)
[0016] where l is the current channel; y l is the result of the convolutional operation in the specified channel; * is the convolutional operation; W ij is the weight matrix in the set region; b ij is the bias matrix in the same region.
[0017] After feature extraction in the convolutional layer, the output feature map is passed to the pooling layer for feature selection and information filtering through the following formula:
[0018] Z i,j =pool i,j (y l ) (2)
[0019] where Z i,j is the mapped result of the features after the pooling operation; pool i,j () represents the pooling operation.
[0020] In the soft - sensing modeling method for the anaerobic fermentation process based on TSSA - CNN - LSTM, the basic framework of LSTM mainly includes a forget gate, an input gate, and an output gate. Based on its gating unit and memory unit, it can effectively process long - term sequences, and its calculation formula is as follows:
[0021] f t = σ(W i ·[h t-1 , x t + b f )(3)
[0022] i t = σ(W i ·[h t-1 , x t + b i )(4)
[0023]
[0024] o t = σ(W0·[h t-1 , x t + b0)(7)
[0025] h t = o t * tanh(c t , x t )(8)
[0026] where f t , i t , o t are the input gate, forget gate, and output gate at time 0 respectively; C t is the output of the memory cell at time 0; h t is the hidden layer state; is the updated state of the memory cell; b i , b f , b o , b c are the bias terms; σ represents the sigmoid function, with a value range of 0 to 1; * represents element-wise multiplication, which is the point-wise multiplication between two vectors; tanh is the hyperbolic tangent function.
[0027] The TSSA of the described soft sensor modeling method for anaerobic fermentation process based on TSSA-CNN-LSTM enhances the local search ability through the following formula using a non-linear inertia weight strategy, adjusts the inertia weight in an exponentially decaying manner, making the exploration in the early stage of the search stronger and the local search ability stronger in the later stage, and improving the convergence accuracy:
[0028]
[0029] where ω is the inertia weight; MaxIt is the maximum number of iterations; t is the current number of iterations:
[0030] Using the horizontal cross strategy, linear weighting and difference terms are carried out through the following formula, so that the solutions after crossing maintain diversity and improve the global search ability:
[0031] X new = rX i +(1 - r)X j + c(X i - X j )(10)
[0032] Where X new is the new solution generated after crossing; r is a random factor that controls the information exchange ratio between individuals, and its value range is [0, 1]; X i , X j are the solution vectors of two individuals to be crossed; c is a random factor that controls the degree of mutation, and its value range is [0, 1].
[0033] The parameter settings of the Bayesian optimization algorithm for the soft sensor modeling method of the anaerobic fermentation process based on TSSA-CNN-LSTM include the number of iterations and the number of points. The parameter settings of LSTM include the number of units and the learning rate. The steps for the Bayesian optimization algorithm to optimize the CNN-LSTM parameters are as follows:
[0034] 1) Initialize CNN and LSTM, and use chaotic mapping to initialize the population of SSA;
[0035] 2) Use the variables selected by CNN as input and output to train LSTM;
[0036] 3) According to the population size, randomly select a part of the individuals from the population to be given the role of predators. According to the proportion of followers, randomly select a part from the remaining individuals as followers. Finally, the scouts will be selected from the remaining individuals;
[0037] 4) Use the non-linear inertia weight strategy to optimize the position update formula of the predator for global search; the predator will actively explore new solutions in the search space according to the global search strategy and update its own position; the expression in this stage is:
[0038]
[0039] Where is the value of the j-th dimension of the i-th sparrow at the t-th iteration; is the individual's historical best position; MaxIt is the maximum number of iterations, t is the current iteration number; η(0, 1) is a normal distribution with a mean of 0 and a standard deviation of 1; r2 is a generated random number, and its range is [0, 1];
[0040] Followers will perform local search according to the following formula, that is, search around the high-quality solutions found by the predators and adjust their positions accordingly:
[0041]
[0042] Where X tp is the optimal position occupied by the current discoverer; represents the worst position currently stored; A represents a 1×d matrix in which each element is randomly assigned 1 or -1, A + = A T (AA T ) -1 ;
[0043] When aware of danger, the scouts in the sparrow population will perform anti-predation behaviors, using the horizontal crossover strategy to enhance population diversity and global search ability. The expression in this stage is:
[0044]
[0045] Where is the historical optimal position of the i-th sparrow; is the historical optimal position of another sparrow individual in the population; r is a random number in [0, 1]; is a random number in [-1, 1];
[0046] 5) Accurately calculate the fitness of each sparrow individual as the standard for evaluating its performance; if the fitness function is satisfied, go to step 6); otherwise, repeat step 4);
[0047] 6) If the maximum number of iterations is reached, assign the obtained parameter optimal solution to the SVM; otherwise, repeat step 4). Description of the Drawings
[0048] Figure 1 is the flowchart of the CNN-LSTM soft sensor model;
[0049] Figure 2 is the overall process schematic diagram of the soft sensor system;
[0050] Figure 3 is the biogas production prediction result of the TSSA-CNN-LSTM soft sensor model. Detailed Implementation Modes
[0051] Combined with the embodiments, the technical solutions of this modeling are clearly and completely described. Using the soft sensor modeling method for anaerobic fermentation process based on TSSA-CNN-LSTM to perform soft sensor modeling for actual identical or similar processes, including the following steps:
[0052] 1) Collect sampling data of the anaerobic fermentation reaction process in a food waste treatment plant through on-site operations, including the percentage of solid content (TS), pH value, percentage of volatile suspended solid content (VS), chemical oxygen demand (COD), average flow rate, alkalinity (Alk), percentage of CO2 content, daily gas production, percentage of CH4 content, and VFA laboratory test values, a total of 240 groups;
[0053] 2) Use a convolutional neural network (CNN) to extract basic features from the collected anaerobic fermentation data as input variables;
[0054] 3) Take out 192 groups and 48 groups of data from the selected input variable sample data as the training set and the test set respectively, and then perform normalization to eliminate the influence of the units and dimensions of different variables. The mapping space is selected as (0, 1), and is used as the normalization criterion, where X0 is the historical data, X i is the minimum value in the historical data, X a is the maximum value, and X is the data sample after normalization;
[0055] 4) Set the parameters of the BOASSA algorithm and SVM: The number of iterations of BOASSA is 10, the number of individuals in the sparrow population is 6, and the dimension of the optimized parameters is 3;
[0056] 5) Use TSSA to optimize the parameters of CNN and LSTM, and assign the optimal parameter solutions to CNN and LSTM.
[0057] 6) Run the proposed TSSA-CNN-LSTM model to extract the time series features in the sample and predict the concentration of biogas accordingly.
[0058] The soft sensor model established by this method has good prediction accuracy for the biogas concentration in the embodiment, and the prediction results are as shown in the appendix Figure 3 . The results show that the established soft sensor model of biogas concentration can accurately estimate the biogas concentration in practical applications and has broad application prospects in the field of monitoring and controlling the anaerobic digestion process of food waste.
Claims
1. A soft sensor modeling method for anaerobic fermentation process based on TSSA-CNN-LSTM, characterized by including The following steps: 1) Obtain sampling data of the anaerobic fermentation process with a time-tagged sequence through on-site operation or experiment, and extract key feature variables from the collected anaerobic fermentation data through a convolutional neural network (CNN) to reduce the dimension of each data item; 2) Divide the sample data of the extracted basic feature variables into a training set and a test set, and perform normalization processing to eliminate the influence of the dimensions of different feature variables; 3) Set the initial parameters of the sparrow search algorithm (SSA), namely the number of iterations and the population size, initialize the population using a chaotic map (Tent), and enhance the local search ability of SSA in the later stage using a non-linear inertia weight; use a horizontal crossover strategy to increase the diversity of the population; 4) Construct a long short-term memory network (LSTM) soft-sensing model CNN-LSTM for the anaerobic fermentation process with the normalized data as input, and use the chaotic sparrow search algorithm (TSSA) to search for the optimal parameters of the CNN-LSTM model. These parameters include the number of LSTM units, the learning rate, and the number of attention mechanism (Attention) heads; 5) Substitute the obtained optimal parameter set into CNN-LSTM to form a high-precision soft-sensing model TSSA-CNN-LSTM for the anaerobic fermentation process, and predict the methane concentration and biogas production.
2. The soft sensor modeling method for anaerobic fermentation process based on TSSA-CNN-LSTM according to claim 1, wherein Use CNN to extract features from the anaerobic digestion process; CNN is a deep neural network designed specifically for feature extraction in the spatial dimension; the structure of CNN consists of an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The convolutional layer is the basis of the entire network architecture, and performs a dot product operation on the data matrix within the defined area through the following formula: y t = f(∑x i * W ij + b ij ) (1) where l is the current channel; y l is the result of the convolution operation in the specified channel; * is the convolution operation; W ij is the weight matrix in the set area; b ij is the bias matrix in the same area.
3. In the soft-sensing modeling method for the anaerobic fermentation process based on BO-CNN-LSTM described in claim 1, the basic framework of LSTM mainly includes a forget gate, an input gate, and an output gate. Based on its gating unit and memory unit, it can effectively process long sequences, and its calculation formula is as follows: f t = σ(W i · [h t-1 , x t + b f ) (2) i t = σ(W i · [h t-1 , x t + b i ) (3) o t = σ(W0 · [h t-1 , x t + b0) (6) h t = o t *tanh(c t , x t ) (7) where f t , i t , o t are the input gate, forget gate, and output gate at time 0, respectively; C t is the output of the memory cell at time 0; h t is the hidden layer state; is the updated state of the memory cell; b i , b f , b o , b c are bias terms; σ represents the sigmoid function, with a value range of 0 to 1; * represents element-wise multiplication, which is the point-wise multiplication between two vectors; tanh is the hyperbolic tangent function.
4. A soft sensor modeling method for anaerobic fermentation process based on TSSA-CNN-LSTM according to claim 1, characterized in that The non-linear inertia weight adjusts the inertia weight in an exponential decay manner through the following formula, making the exploration in the early stage of the search stronger and the local search ability in the later stage enhanced, thereby improving the convergence accuracy: Where ω is the inertia weight; MaxIt is the maximum number of iterations; t is the current number of iterations.
5. The soft sensor modeling method for anaerobic fermentation process based on TSSA-CNN-LSTM according to claim 1, characterized in that The horizontal crossover strategy performs linear weighting and difference terms through the following formula, making the solutions after crossover maintain diversity and improving the global search ability: X new = rX i +(1 - r)X j + c(X i- X j ) (9) Where X new is the new solution generated after crossover; r is a random factor that controls the information exchange ratio between individuals, and its value range is [0, 1]; X i , X j are the solution vectors of two individuals to be crossed; c is a random factor that controls the degree of mutation, and its value range is [0, 1].
6. A soft sensor modeling method for anaerobic fermentation process based on TSSA-CNN-LSTM according to claim 1, characterized in that The chaotic sparrow search algorithm (TSSA) optimizes the parameters of CNN-LSTM, and the specific steps are as follows: 1) Initialize CNN and LSTM, and initialize the population of SSA using a chaotic map; 2) Use the variables selected by CNN as input and output to train LSTM; 3) According to the population size, randomly select a part of the individuals from the population to be given the role of predators, and randomly select a part of the remaining individuals as followers according to the proportion of followers. Finally, the scouts will be selected from the remaining individuals; 4) Optimize the position update formula of the predator using a non - linear inertia weight strategy for global search; the predator will actively explore new solutions in the search space according to the global search strategy and update its own position; the expression at this stage is: where is the value of the j-th dimension of the i-th sparrow at the t-th iteration; is the individual's historical best position; MaxIt is the maximum number of iterations, and t is the current iteration number; η(0, 1) is a normal distribution with a mean of 0 and a standard deviation of 1; r2 is a generated random number in the range [0, 1]; The followers will perform local search according to the following formula, that is, search around the high - quality solutions found by the predator and adjust their positions accordingly: where X ip is the optimal position occupied by the current discoverer; represents the worst position currently stored; A represents a 1×d matrix in which each element is randomly assigned 1 or -1, A + = A T (AA T ) -I ; When aware of danger, the scouts in the sparrow population will perform anti - predation behavior, using a horizontal crossover strategy to enhance population diversity and global search ability. The expression at this stage is: where is the historical optimal position of the i-th sparrow; is the historical optimal position of another sparrow individual in the population; r is a random number in [0, 1]; c is a random number in [-1, 1]; 5) Accurately calculate the fitness of each sparrow individual as the criterion for evaluating its performance; if the fitness function is satisfied, go to step 6), otherwise repeat step 4); 6) If the maximum number of iterations is reached, assign the obtained optimal parameter solution to the SVM, otherwise repeat step 4).
Citation Information
Cited By
Artificial intelligence method for adaptive dynamic precise fermentation control
CN121325786A
An artificial intelligence method for adaptive dynamic precision fermentation control
CN121325786B