Big data-based lactobacillus fermentation process data processing method

The data processing method for lactic acid bacteria fermentation, built using big data and machine learning, solves the controllability and stability problems of traditional lactic acid bacteria fermentation processes, achieving efficient and low-consumption fermentation process control, and improving yield and product quality.

CN121963891APending Publication Date: 2026-05-01LIAONING ACAD OF AGRI SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LIAONING ACAD OF AGRI SCI
Filing Date
2026-01-14
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional lactic acid bacteria fermentation processes rely on empirical formulas and static models, which are difficult to adapt to complex processes with multiple variables coupled, resulting in poor controllability of the fermentation process, unstable yield, high energy consumption, and large fluctuations in product quality.

Method used

A data processing method based on big data for lactic acid bacteria fermentation is adopted. Data is collected collaboratively by multiple sensors to build a closed-loop control system for the entire process. Machine learning models are used for real-time prediction and optimization. Intelligent optimization algorithms are combined to generate process parameter adjustment schemes and dynamically update the model.

Benefits of technology

It achieves efficient, stable, and low-consumption operation of the fermentation process, improves yield and product quality consistency, and reduces energy consumption and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963891A_ABST
    Figure CN121963891A_ABST
Patent Text Reader

Abstract

The invention discloses a lactobacillus fermentation process data processing method based on big data, and belongs to the technical field of lactobacillus fermentation. Environmental parameters, biological parameters and material parameters are collected in real time through multiple sensors deployed in a fermentation tank, the parameters are uploaded to a cloud platform through an Internet of Things gateway, cleaning, alignment and normalization processing is carried out on collected time sequence data, a structured fermentation data set is constructed, time domain, frequency domain and statistical characteristics are extracted from the data set, and the time domain, the frequency domain and the statistical characteristics are analyzed. And constructing a high-dimensional feature vector. Through organic combination of the steps, a whole-process closed-loop control system from data acquisition, processing and analysis to decision execution is constructed. The system can sense dynamic changes of the fermentation process in real time, accurate prediction of key indexes and intelligent optimization of technological parameters are achieved in a data driving mode, and the limitation of a traditional control method is effectively overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of lactic acid bacteria fermentation technology, and more specifically, to a data processing method for lactic acid bacteria fermentation process based on big data. Background Technology

[0002] Lactic acid bacteria fermentation is a widely used bioprocess in the food, pharmaceutical, and feed industries. Traditional fermentation process control mainly relies on empirical formulas, manual adjustments, and static models, which are difficult to adapt to the complex process involving multiple variables such as batch differences in raw materials, environmental changes, and the metabolic dynamics of microorganisms.

[0003] Currently, real-time data generated during fermentation is often recorded in isolation or used only for post-process analysis, lacking systematic data integration and real-time intelligent decision support. This leads to problems such as poor controllability of the fermentation process, unstable yield, high energy consumption, and large fluctuations in product quality. Summary of the Invention

[0004] The purpose of this invention is to provide a data processing method for lactic acid bacteria fermentation process based on big data, in order to solve the problem that real-time data generated during the fermentation process is often recorded in isolation or only used for post-event analysis, lacking systematic data integration and real-time intelligent decision support. This leads to problems such as poor controllability of the fermentation process, unstable yield, high energy consumption, and large fluctuations in product quality.

[0005] A data processing method for lactic acid bacteria fermentation process based on big data includes the following steps: S1. Data Acquisition and Upload: Environmental parameters, biological parameters, and material parameters are collected in real time by various sensors deployed in the fermenter, and the parameters are uploaded to the cloud platform via an IoT gateway; S2. Data Preprocessing and Construction: The collected time-series data is cleaned, aligned, and normalized to construct a structured fermentation dataset; S3. Feature Extraction and Construction: Extract time-domain, frequency-domain, and statistical features from the dataset and construct high-dimensional feature vectors; S4. Indicator Prediction: Based on feature vectors, use a trained machine learning model to predict key fermentation indicators for future periods; S5. Optimization scheme generation: Based on the prediction results and combined with the preset optimization objectives, an intelligent optimization algorithm is used to generate a process parameter adjustment scheme; S6. Closed-loop regulation execution: The adjustment plan is sent to the fermentation control system to realize the closed-loop regulation of process parameters; S7. Model Update and Archiving: Based on the deviation between the actual fermentation results and the predicted values, the prediction model is dynamically updated, and the process data is archived to the knowledge base.

[0006] Through the above technical solutions and the organic combination of the above steps, a closed-loop control system covering the entire process from data acquisition, processing, analysis to decision execution has been constructed. This system can perceive the dynamic changes in the fermentation process in real time, and achieve accurate prediction of key indicators and intelligent optimization of process parameters through data-driven methods, effectively overcoming the limitations of traditional control methods. Among them, the collaborative acquisition of multiple sensors ensures the comprehensiveness and timeliness of data, the application of the cloud platform provides support for large-scale data storage and processing, the introduction of machine learning models improves the accuracy and robustness of predictions, intelligent optimization algorithms provide a scientific basis for the dynamic adjustment of process parameters, and the dynamic update mechanism of the model ensures that the system can continuously adapt to changes in the fermentation process, continuously optimize the control strategy, and ultimately achieve efficient, stable, and low-consumption operation of the fermentation process.

[0007] Preferably, the sensors in S1 include a temperature sensor, a pH sensor, a dissolved oxygen sensor, a turbidity sensor, and an exhaust gas analyzer, with a data acquisition frequency adjustable from 1 second to 10 minutes.

[0008] Through the above technical solution and the coordinated deployment of multiple types of sensors, it is possible to simultaneously capture physical parameters, biomass-related parameters and metabolite information in the fermentation environment, forming a multi-dimensional data acquisition network.

[0009] Preferably, the features extracted in S3 include mean, variance, slope, Fourier transform dominant frequency, and wavelet energy entropy.

[0010] Through the above technical solutions, and by deeply mining and integrating these multi-dimensional features, key information reflecting the inherent laws of the fermentation process can be extracted from the original monitoring data. Among them, the mean and variance can characterize the central tendency and fluctuation of parameters within a specific period, the slope can intuitively reflect the rate of change and trend direction of parameters, the Fourier transform main frequency helps to capture the periodic features in the data, and the wavelet energy entropy can effectively reveal the complexity of the energy distribution of the signal in different frequency bands. This provides comprehensive and representative input variables for the training of subsequent machine learning models, thereby improving the prediction accuracy of key indicators of the fermentation process.

[0011] Preferably, the machine learning model in S4 is an LSTM neural network, and the model training employs cross-validation and hyperparameter optimization. The above technical solutions leverage the unique advantage of LSTM neural networks in processing long-sequence data, effectively capturing the dynamic dependencies of various parameters over time during fermentation. This makes them particularly suitable for modeling complex biological processes with strong temporal characteristics, such as fermentation. Cross-validation, by dividing the dataset into multiple subsets and alternating between training and validation, avoids overfitting or underfitting caused by data partitioning bias, ensuring the model's stability and generalization ability under different data distributions.

[0012] Preferably, the optimization objectives in S5 include maximizing product yield, minimizing fermentation time, and stabilizing product quality, and the optimization algorithm is a genetic algorithm, particle swarm optimization, or Bayesian optimization.

[0013] Through the above technical solutions, the fermentation process can be comprehensively controlled by setting multi-dimensional optimization goals. Among them, the improvement of product yield is directly related to production efficiency, the reduction of fermentation time can significantly reduce energy and labor costs, and the stability of product quality is the core element to ensure the product's market competitiveness.

[0014] Preferably, the control system in S6 supports Modbus, OPC UA or MQTT protocol communication to achieve linkage control with the actuator.

[0015] Preferably, the data processing in S2 further includes outlier detection and interpolation, using a sliding window mean method or a linear interpolation method to process missing data.

[0016] Preferably, the control system in S6 includes a PLC, DCS, or edge computing controller, and supports Modbus, OPCUA, or MQTT protocol communication.

[0017] Compared with the prior art, the advantages of this invention are: 1. By organically combining the above steps, a closed-loop control system covering the entire process from data acquisition, processing, analysis to decision execution is constructed. This system can perceive the dynamic changes in the fermentation process in real time, and achieve accurate prediction of key indicators and intelligent optimization of process parameters through data-driven methods, effectively overcoming the limitations of traditional control methods. Among these features, the collaborative acquisition of multiple sensors ensures the comprehensiveness and timeliness of the data, the application of the cloud platform provides support for large-scale data storage and processing, the introduction of machine learning models improves the accuracy and robustness of predictions, intelligent optimization algorithms provide a scientific basis for the dynamic adjustment of process parameters, and the dynamic update mechanism of the model ensures that the system can continuously adapt to changes in the fermentation process, continuously optimize the control strategy, and ultimately achieve efficient, stable, and low-consumption operation of the fermentation process.

[0018] 2. Through the coordinated deployment of multiple types of sensors, physical parameters, biomass-related parameters, and metabolite information in the fermentation environment can be captured simultaneously, forming a multi-dimensional data acquisition network.

[0019] 3. By deeply mining and integrating these multi-dimensional features, key information reflecting the inherent laws of the fermentation process can be extracted from the original monitoring data. Among them, the mean and variance can characterize the central tendency and fluctuation of the parameters in a specific period, the slope can intuitively reflect the rate of change and trend direction of the parameters, and the Fourier transform can represent the complexity of the energy distribution of the main sign in different frequency bands. This provides comprehensive and representative input variables for the training of subsequent machine learning models, thereby improving the prediction accuracy of key indicators of the fermentation process. Attached Figure Description

[0020] Figure 1 This is a system diagram of the present invention. Detailed Implementation

[0021] Example 1: Optimization of Laboratory-Level Lactic Acid Bacteria Fermentation Process Step 1. System Configuration and Data Acquisition Uses a 5L fully automatic fermenter equipped with online pH, dissolved oxygen, temperature, foam and turbidity sensors; The above parameters are collected once per minute through the edge computing gateway. At the same time, the bacterial concentration and lactic acid concentration are manually sampled offline every 30 minutes and the measurement results are manually entered into the system as label data for model training and validation. The controlled variables include: fermentation temperature setting range of 30-40℃, stirring speed of 100-300 rpm, aeration rate of 0.5-1.5 vvm, and the rate of addition of alkali solution to adjust pH and nutrient supplementation.

[0022] Step 2. Data Processing and Feature Engineering For time-series data acquired online, a sliding window is used for smoothing to remove obvious abnormal pulses; For each sliding window, the following features are calculated: mean, variance, maximum value, minimum value, trend, and dominant frequency amplitude extracted via Fast Fourier Transform. Finally, a 35-dimensional feature vector is generated for each time point. The offline measured OD600 and lactic acid concentration data are aligned with the feature vector of the most recent time window to form a complete training sample set.

[0023] Step 3. Model Training and Prediction A Long Short-Term Memory (LSTM) network is used as the prediction model. The network structure consists of an input layer, two LSTM layers, and a fully connected output layer. The system was trained using data from the past 30 successful fermentation batches, with 80% of the data used as the training set and 20% as the validation set. The optimizer used was Adam, and the loss function was mean squared error. After the model is deployed, it receives the feature vector at the current moment in real time and predicts the trend of cell growth and product accumulation in the next 2 hours.

[0024] Step 4. Real-time optimization and control The optimization objective is set as follows: to maximize the final lactic acid concentration while ensuring that the fermentation cycle does not exceed 24 hours; Using a Bayesian optimization algorithm, based on the current predicted trend, the optimal parameter combination is calculated every 15 minutes, with the temperature setpoint and alkali addition rate for the next 2 hours as decision variables. The optimized temperature setpoint is automatically adjusted by the fermenter temperature control system, and the alkali addition rate is automatically replenished by controlling the frequency of the peristaltic pump.

[0025] Step 5. Implementation Results After optimization using this method, the average fermentation cycle of this strain was shortened from 26 hours to 22 hours, and the average final lactic acid concentration was increased by about 12%. The LSTM model's mean absolute error (MAE) for predicting lactic acid concentration over the next two hours remained stable within 5%, providing a reliable basis for optimized control.

[0026] Example 2: Multi-objective optimization of industrial-grade lactic acid fermentation Step 1. System Configuration and Data Acquisition Deployed on industrial production lines, the data comes from the factory's DCS system, integrating more than 50 process variables such as temperature, pH, DO, tank pressure, stirring current, inlet and outlet flow rate and CO2 / O2 concentration, and feed flow rate. The data is collected once per second and synchronized in real time to the factory's private cloud platform via an industrial data acquisition gateway.

[0027] Step 2. Data Processing and Integrated Modeling For high-dimensional, high-frequency data, correlation analysis and principal component analysis (PCA) are first performed to reduce the dimensionality of the original variables to 20 principal feature components. Using the gradient boosting tree model, we predicted the lactic acid yield, cell viability, and energy consumption index per unit product for the next 4 hours. The GBR model can better handle nonlinear relationships and feature importance ranking. The model is incrementally learned and updated every 8 hours using recent data to adapt to fluctuations in raw material batches.

[0028] Step 3. Multi-objective optimization and coordinated control Three optimization objectives were set: ① Maximize lactic acid yield; ② Minimize steam consumption per unit product; ③ Ensure that the final product quality indicators remain stable within the set range. A multi-objective particle swarm optimization algorithm was used to solve the problem. Optimization calculations were run every 30 minutes, outputting recommended adjustments to temperature, aeration rate, stirring speed, and sugar addition rate. After the optimization plan is confirmed by the operator, it is sent to the DCS for execution through the APC layer.

[0029] Step 4. Knowledge Base Application Establish a process knowledge base to store raw material information for each batch, full-process operation parameters, model decision records, and final production indicators; When changing raw material suppliers or starting a new strain, the system recommends initial process parameters and model-predicted baseline curves by searching the knowledge base for the most similar historical batches, thus accelerating process adaptation.

[0030] Step 5. Implementation Results Statistics after six months of implementation show that the average lactic acid yield increased by 8.5%, the steam consumption per unit product decreased by 12%, and the rate of premium-grade products increased from 95.2% to 98.7%. The system enables early warning of abnormalities in the fermentation process. By analyzing the abnormal upward trend of exhaust CO2, it successfully warned of multiple risks of contamination and avoided significant losses.

[0031] Comparative Example: Traditional Manual Experience Control Control method: On the same production line, using the same strains and raw materials, experienced engineers make manual judgments and adjustments based on fixed operating procedures and intermittent offline test results; Operating principles: Primarily relies on pH and DO setpoint control, with feeding based on a fixed schedule or a rough estimate of substrate consumption.

[0032] Testing and Results Analysis: Model prediction accuracy test Step 1: Divide the test set into independent test sets Key principle: Use data that the model has never seen before during training and validation.

[0033] Operation method: Divide into time sequence: For example, use the most recent 10% of batch data as the test set to avoid data leakage and ensure that the test set covers different production conditions, such as different raw material batches, different seasons, and different operating shifts, in order to test the model's generalization ability.

[0034] Step 2: Perform predictions and collect results Simulate a real-time operating environment on the time series of the test set: Rolling prediction: For each point in time, the model is input with current and historical data to predict the target value for one or more future time periods, and the model's predictions for all test samples are recorded.

[0035] Step 3: Obtain the actual value For each predicted time point, collect actual values ​​measured through reliable means to ensure that the actual values ​​and predicted values ​​are precisely aligned in time.

[0036] Step 4: Calculate accuracy evaluation indicators Choose a set of complementary indicators to comprehensively evaluate from different perspectives: Mean Absolute Error Calculation formula: MAE = (1 / n)Σ|y_i - ŷ_i| Where n is the number of samples, y_i is the true value, and ŷ_i is the predicted value; Root mean square error Calculation formula: RMSE = √[(1 / n)Σ(y_i - ŷ_i) 2 ] Used to measure the overall degree of deviation between predicted and actual values; Coefficient of determination Calculation formula: R 2 = 1 - Σ(y_i - ŷ_i) 2 / Σ(y_i - ȳ) 2 R reflects the model's ability to explain data variability. 2 The closer to 1, the better the model fit.

[0037] Step 5: Analysis and Visualization Draw a comparison chart: Plotting the predicted and actual curves on the same time axis clearly demonstrates the model's performance in trend following, peak prediction, and lag.

[0038] Draw the residual plot: Plot the distribution of prediction error as a function of the true value or time to check if the residuals are randomly distributed. If a clear pattern exists, it indicates that the model has a systematic bias.

[0039] Error distribution histogram: Check whether the distribution of errors approximates a normal distribution to determine whether there is a systematic overestimation or underestimation.

[0040] Step 6: Scenario-based stress testing Critical operating condition testing: This section specifically tests the model's prediction accuracy under critical or abnormal operating conditions such as process inflection points, material replenishment times, and the initial stage of contamination. This is a key aspect of verifying the model's robustness. Long-term forecast test: Test the accuracy decay of the model in predicting longer time periods such as 4 hours and 8 hours, and determine the effective prediction horizon of the model.

[0041] Performance Analysis: On the time series of the test set, a real-time operating environment is simulated. For each time point, current and historical data are input into the model to predict the target value for one or more future time steps. The model's prediction values ​​for all test samples are recorded. At the same time, metadata such as the input feature vector, model version number, and computation time during prediction should be recorded synchronously for subsequent retrospective analysis.

[0042] System response and stability testing Step 1: Preparation Phase Select a typical stable fermentation period as the starting point for testing, ensure that all sensors, actuators, and data links are working properly, and record the basic process data before testing.

[0043] Step 2: Perform the test At a predetermined time point, a disturbance of a certain level is implemented, and data of all relevant variables are recorded at a high frequency throughout the process, including the actual values ​​of all controlled variables, the set values ​​or operating quantities calculated by the algorithm, the actual output of the actuator, and the real-time predicted values ​​of the model for key indicators.

[0044] Step 3: Data Analysis and Evaluation Create a trend comparison chart: Put the trends of key variables before and after the disturbance together to clearly show the disturbance point, system response, and recovery process.

[0045] Ideal response: smooth curve transition, small overshoot, and fast recovery.

[0046] Adverse responses: violent oscillations, slow recovery, or even loss of control.

[0047] Calculate the quantitative indicators: Based on the indicators in Part 2, calculate the specific values.

[0048] Assess and control quality: Compare the response to process requirements to determine if it is acceptable.

[0049] Effect Analysis: Excellent dynamic performance: The system responds quickly and smoothly to changes in setpoints and external disturbances, and all dynamic indicators are superior to process control standards, proving the high efficiency of closed-loop optimization control; Stable and reliable operation: Long-term operation data shows that the system can continuously maintain process parameters within an extremely narrow optimization range, significantly reducing batch-to-batch differences and ensuring high product quality and consistency; Strong robustness and safety: When faced with abnormal operating conditions such as simulated sensor failures and actuator constraints, the system exhibits good fault tolerance and graceful degradation mechanism, which can maintain basic control or safe shutdown, thus ensuring production safety. Significant economic benefits: Statistical data during the testing period show that, compared with the traditional control mode, this system brings quantifiable economic benefits in terms of increasing productivity, reducing energy consumption per unit product, and reducing defective products, verifying its great application value.

[0050] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and modifications can be made to the present invention without departing from the spirit and scope thereof, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A data processing method for lactic acid bacteria fermentation process based on big data, characterized in that, The method specifically includes the following detailed steps: S1. Data Acquisition and Upload: Environmental parameters, biological parameters, and material parameters are collected in real time by various sensors deployed in the fermenter, and the parameters are uploaded to the cloud platform via an IoT gateway; S2. Data Preprocessing and Construction: The collected time-series data is cleaned, aligned, and normalized to construct a structured fermentation dataset; S3. Feature Extraction and Construction: Extract time-domain, frequency-domain, and statistical features from the dataset and construct high-dimensional feature vectors; S4. Indicator Prediction: Based on feature vectors, use a trained machine learning model to predict key fermentation indicators for future periods; S5. Optimization scheme generation: Based on the prediction results and combined with the preset optimization objectives, an intelligent optimization algorithm is used to generate a process parameter adjustment scheme; S6. Closed-loop regulation execution: The adjustment plan is sent to the fermentation control system to realize the closed-loop regulation of process parameters; S7. Model Update and Archiving: Based on the deviation between the actual fermentation results and the predicted values, the prediction model is dynamically updated, and the process data is archived to the knowledge base.

2. The method according to claim 1, characterized in that, The sensors in S1 include a temperature sensor, a pH sensor, a dissolved oxygen sensor, a turbidity sensor, and an exhaust gas analyzer, with a data acquisition frequency adjustable from 1 second to 10 minutes.

3. The method according to claim 1, characterized in that, The features extracted in S3 include mean, variance, slope, Fourier transform dominant frequency, and wavelet energy entropy.

4. The method according to claim 1, characterized in that, The machine learning model in S4 is an LSTM neural network, and the model training uses cross-validation and hyperparameter optimization.

5. The method according to claim 1, characterized in that, The optimization objectives in S5 include maximizing product yield, minimizing fermentation time, and stabilizing product quality. The optimization algorithm is a genetic algorithm, a particle swarm optimization algorithm, or a Bayesian optimization algorithm.

6. The method according to claim 1, characterized in that, The control system in S6 supports Modbus, OPC UA, or MQTT protocol communication to achieve linkage control with the actuator.

7. The method according to claim 1, characterized in that, The data processing in S2 also includes outlier detection and interpolation, using the sliding window mean method or linear interpolation method to process missing data.

8. The method according to claim 1, characterized in that, The control system in S6 includes a PLC, DCS, or edge computing controller, and supports Modbus, OPC UA, or MQTT protocol communication.