A microalgal growth curve prediction method based on transfer learning under small sample conditions
By combining transfer learning and deep learning methods, a microalgae growth curve prediction model was constructed, which solved the problems of data scarcity and complex nonlinear relationships in microalgae growth curve prediction, and achieved efficient and economical microalgae growth curve prediction.
Patent Information
- Application Number
- CN202411341853.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-09-25
AI Technical Summary
Existing technologies for predicting microalgae growth curves suffer from problems such as scarce data, complex nonlinear relationships, and numerous influencing factors, resulting in large prediction biases, high experimental costs, and insufficient model generalization ability.
By combining transfer learning and deep learning, and integrating a small amount of historical and simulated data, a microalgae growth dynamics model based on the logistic model is constructed. The model is then iteratively trained using an LSTM model and the Two-Stage TrAdaBoost.R2 algorithm to generate a microalgae growth curve prediction model.
It significantly improves the accuracy of prediction results under small sample conditions, reduces experimental costs, optimizes data utilization efficiency, enhances the generalization ability of the model, and overcomes the limitations of traditional methods.
Smart Images

Figure CN119294237B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the intersection of biotechnology and deep learning, specifically a method for predicting microalgae growth curves using transfer learning and deep learning techniques under conditions of small sample data. Background Technology
[0002] Microalgae, with their astonishing growth rate, highly efficient photosynthetic capacity (nearly ten times that of terrestrial plants), and ability to synthesize important biocompounds, are considered a highly promising renewable resource. Based on these unique advantages, microalgae have been widely used in the production of high-value-added products such as fuels, cosmetics, and nutrients, as well as in wastewater treatment and carbon emission mitigation. The growth curve of microalgae is a comprehensive indicator of its population growth, reflecting its ability to adapt to and survive in specific environments. Accurately predicting the growth curve of microalgae is crucial for developing cultivation programs, optimizing the cultivation environment, and controlling microalgae processing conditions.
[0003] Traditional microalgae industry predictions for microalgae growth status rely primarily on empirical summaries, simple kinetic models, or single machine learning models. However, these methods have significant limitations, including poor generalization performance and high requirements for prior historical data. Therefore, accurately simulating microalgae growth remains a challenge. Algal growth is influenced by multiple factors (such as light intensity, culture medium concentration, pH, temperature, and carbon dioxide concentration), which affect microbial growth through complex metabolic mechanisms. Therefore, capturing this multi-stage, nonlinear growth pattern is extremely challenging.
[0004] Existing kinetic models, such as the Monod, Droop, and Han models, suffer from difficulties in parameter estimation and require large amounts of experimental data. With the development of deep learning technology, data-driven models are increasingly being applied in the field of biochemical engineering. Among them, the LSTM model, due to its excellent time-series data processing capabilities and long-term dependency capture ability, shows great potential for widespread application in predicting biological growth states. However, as a deep learning model, the accuracy of the LSTM model is highly dependent on data quality. The scarcity of experimental data, measurement errors, and the high cost and long cycle of biological experiments and detection collectively limit the application of the LSTM model in microalgal growth prediction. Summary of the Invention
[0005] The core objective of this invention is to overcome the limitations of traditional growth kinetic models and LSTM models, addressing numerous challenges in microalgae growth prediction tasks, such as scarce data, complex nonlinear relationships, and numerous influencing factors. This invention aims to solve problems such as large prediction biases in microalgae growth curves, high experimental costs, and insufficient model generalization ability. To this end, this invention proposes an innovative solution: a kinetic-assisted microalgae growth curve prediction method combining transfer learning. This method significantly improves the accuracy of prediction results under small sample conditions by integrating limited historical and simulated data. It combines the advantages of kinetic models and LSTM models for microalgae growth curve prediction. This method is expected to overcome the bottlenecks of existing technologies, effectively reduce prediction bias, and decrease experimental costs.
[0006] The technical solution adopted by the present invention to achieve the above objectives is as follows:
[0007] A method for predicting microalgal growth curves based on transfer learning under small sample conditions, the method comprising the following steps:
[0008] Step S1: Prepare the cultures of the training group and the fitting group using the Box-Behnken method, and obtain the experimental data of the training group and the fitting group {key environmental influencing factors, biomass concentration values of the long-term liquid microalgae system}. The biomass concentration values of the long-term liquid microalgae system characterize the microalgae growth curve.
[0009] Step S2: Considering the light and dark cycle, construct a microalgae growth kinetic model based on the logistic model; use this model to iteratively learn and train with the fitted group data in the experimental data as input to obtain an ideal microalgae growth kinetic model;
[0010] Step S3: Using the key environmental influencing factors of the training group in the experimental data as input, the dense dynamic simulation data X″ of the microalgae system biomass concentration value is generated by using the ideal model of microalgae growth dynamics to characterize the dynamic growth curve and output it, thereby achieving generalization enhancement of data volume.
[0011] Step S4: Using an LSTM model as the prediction model, the {key environmental influencing factors and microalgae biomass concentration values} of the experimental data X′ and simulation data X″ are fused. The fitting learning relationship of the prediction model is iteratively trained using the Two-Stage TrAdaboost.R2 method to obtain the ideal prediction model. The ideal prediction model is then used to predict the corresponding microalgae biomass concentration values for the input key environmental influencing factors.
[0012] Step S5: Collect key environmental influencing factors of the water sample to be tested over a long period of time, input them into the trained ideal prediction model, and automatically predict and output the biomass concentration value of the microalgae system over a long period of time to characterize the microalgae growth trend over a long period of time.
[0013] The acquisition of experimental data in step S1 includes:
[0014] Step S11: Collect the long-term OD values of the liquid microalgae system in the training group and the fitting group;
[0015] Step S12: Fit the calibration curve and regression relationship F between the microalgal OD value and dry weight using the fitted group data;
[0016] Step S13: Use F to calibrate the microalgae OD value of the training group and obtain its corresponding calibrated biomass concentration value X′. Establish the {key environmental influencing factors and biomass concentration of long-term liquid microalgae system} for the training group and the fitting group.
[0017] Step S11 includes:
[0018] Step S11a: Prepare multiple training group microalgae culture solutions using culture medium and in combination with the variable range of key environmental factors, and collect the OD values of multiple long-term liquid microalgae systems under different variable environments by absorbance method;
[0019] Step S11b: Prepare the microalgae culture medium for the fitting group using a culture medium and in combination with standard key environmental factors, and collect the OD value of the long-term liquid microalgae system under standard variable environment by absorbance method.
[0020] The key environmental factors include nitrogen concentration, phosphorus concentration, light intensity, and light-dark ratio; the nitrogen concentration range is 335 mg·L⁻¹. -1 -385mg·L -1 The phosphorus concentration range is 6.4 mg·L⁻¹. -1 -25mg·L -1 The light intensity range is 26 μmol·m -2 ·s -1 -78 μmol·m -2 ·s -1 The light-to-dark ratio ranges from 12:12 to 24:0;
[0021] The preparation of multiple fitted microalgae culture media using culture medium and in conjunction with standard key environmental factors includes: setting the nitrogen element concentration range to 335 mg·L⁻¹. -1 -385mg·L -1 The phosphorus concentration range is 6.4 mg·L⁻¹. -1 -25mg·L -1 The light-to-dark ratio is 14:10, and the light intensity is 26 μmol·m⁻². -2·s -1 The culture was carried out for 120 hours until the microalgae culture medium entered the quiescent phase.
[0022] The microalgae growth curve data were collected using a UV-Vis spectrophotometer, measuring the absorbance (OD) of the microalgae at a wavelength of 750 nm according to sampling intervals. 750 The microalgae growth curve data under standard variable environment were obtained by measuring OD values every 4 hours.
[0023] The calibration curve and regression relationship between the fitted microalgal OD value and dry weight include: preparing microalgal dry powder from the microalgal culture medium of the fitted group, calculating the microalgal biomass concentration X based on the mass and volume of the dry powder, and characterizing the long-term growth curve data of the fitted group solid microalgal; obtaining the calibration curve and regression relationship F between the microalgal OD value and dry weight based on the OD value and the corresponding microalgal biomass concentration X at the OD.
[0024] Step S2 includes: Step S21, establishing a microalgae growth kinetic model, which includes sequentially interleaving multiple light cycle models and dark cycle models;
[0025] a) Using key environmental factors affecting microalgae cultivation as input variables and microalgae biomass concentration as the output variable, a growth kinetics loop model is constructed, as shown in the following equation: Factors related to light, nitrogen, and phosphorus are introduced and applied in a multiplicative model to the maximum specific growth rate μ in the logistic model. max As shown in equations (1)-(2); the light intensity is modeled using the Aiba equation, as shown in equation (3); the nitrogen and phosphorus elements are modeled using the Andrews equation, as shown in equations (4)-(5); the attenuation law of light propagating along the medium is described according to the Lamb-Beer law, as shown in equation (6).
[0026]
[0027]
[0028] Where X(t) is the microalgal biomass concentration at time t, in g·L⁻¹. -1 μ max Represents the maximum ratio growth rate, in h. -1 ;X max The maximum biomass concentration that the environment can sustain, expressed in g / L. -1 μ d The decay coefficient represents the rate at which microalgal cell populations decay in the environment, expressed in hours (h). -1 ;
[0029] Where f(I), f(N), and f(P) represent changes in light intensity, nitrogen element, and phosphorus element, respectively; where μ mThe ratio of growth rate constant, in units of h. -1 I represents light intensity, in μmol·m⁻¹ -2 ·s -1 ;K I,i is the light inhibition constant, used to represent the phenomenon that excessive light intensity inhibits photosynthesis, and its unit is μmol·m. -2 ·s -1 ;K I,s is the light saturation constant, used to indicate that higher light intensity does not significantly improve photosynthesis; the unit is μmol·m. -2 ·s -1 N and P are the concentrations of nitrogen and phosphorus, respectively, in mg·L. -1 ;K N,i ,K P,i These are the inhibition constants for nitrogen and phosphorus, respectively, in mg·L. -1 ;K N,s ,K P,s These are the saturation constants for nitrogen and phosphorus, respectively, in mg·L. -1 ;
[0030] Where I0 represents the incident light intensity, with units of μmol·m -2 ·s -1 Ka is the optical attenuation coefficient, with units of L·m. -1 ·g -1 Z represents the distance behind the irradiated area, in meters.
[0031] To reduce the impact of spatial dimension on the solution of the equation, the ten-step trapezoidal rule is used to transform equation (6) into equation (7):
[0032]
[0033] The matrix consumption rate can be calculated using the mass conservation equation:
[0034]
[0035] Among them, Y X / N ,Y XP It is the yield coefficient relative to the biomass production of the substrate; m N ,m P It is the specific consumption rate for cell maintenance.
[0036] b) Consider the light cycle process and the dark cycle process respectively, and use formulas (1)-(5) and (7)-(10) as the light cycle model;
[0037] Formulas (1), (4), (5), (9)-(11) are used as the dark cycle model; and the maximum specific growth rate is expressed by the following formula:
[0038] μ max =μ m ·f(N)·f(P) (12)
[0039] Step S2 further includes: Step S22, using the model to perform iterative learning with key environmental influencing factors as input, including using a two-stage parameter fitting method combined with a differential evolution algorithm to fit the parameters; including the following steps:
[0040] Step S221: Set the light-dark ratio and divide the fitted group data into light cycle data and dark cycle data; each part of the data is {key environmental influencing factors, long-term liquid microalgae system concentration};
[0041] Step S222: Initialize the parameters of the optical cycle model, fix the parameters of the dark cycle model, and use the optical cycle data from step S221 to perform parameter fitting only on the optical cycle model to obtain the dynamic parameters of the optical cycle model.
[0042] Step S223: Fix the dynamic parameters of the light cycle model obtained in step S222, and use the dark cycle data in step S221 to perform parameter fitting on the dark cycle model to obtain the dynamic parameters of the dark cycle model.
[0043] Step S4 includes the following steps:
[0044] Step S41: Combine the experimental data and dynamic simulation data under the training set cultivation conditions, where the experimental data serves as the target domain data T in transfer learning. target =(X exp ,Y exp Simulation data serves as the source domain data T. source =(X model ,Y model ); splicing the two together results in T = {(X model ,X exp ),(Y model ,Y exp )};
[0045] Step S42: Construct an LSTM model as a base learner and train it using the Two-Stage TrAdaboost.R2 method.
[0046] The Two-Stage TrAdaBoost.R2 algorithm in step S42 includes the following steps:
[0047] Step S421: Set the number of iterations S, t = 1, ..., S, and set the initial sample weights:
[0048] Step S422: Update the target domain instances according to the weight update strategy of the Adaboost.R2 algorithm; keep the source domain weights unchanged, and gradually reduce the weights of all target datasets T. source The weights of the middle samples and their proportion in the total weights ω1 are determined, and the optimal weight ratio is obtained through cross-validation; the auxiliary training set T is used. source The weights of the samples are updated; the weight update method is to call the base regressor (LSTM) to obtain a learner on the merged training set T, and then calculate the weight adjustment error for each sample. Reduce weights based on error values;
[0049] Step S423: Based on the determined weight ratio, update the target domain instances according to the weight update strategy of the Adaboost.R2 algorithm; perform weighted processing on data points with large errors; during this stage, the auxiliary training set T... source The weights of the samples remain unchanged; through cross-validation, the sample in the target training set T is found. target The model with the smallest regression error is taken as the final training result;
[0050] The weight update rules are as follows:
[0051]
[0052] Z t β is the normalization constant. t The goal is to balance the weight of the source domain data with the overall weight ω1 using a binary search. The search objective is to make the overall weight of the target domain equal to... This objective is not set for individual sample weights.
[0053] Step S424: Output the model with the smallest error. t , t = argmin(error) t ).
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] (1) Based on the logistic model, the factors affecting microalgae growth, such as light intensity, nitrogen concentration, and phosphorus concentration, are incorporated into the mathematical model in the form of a multiplicative model.
[0056] (2) Using a multi-stage modeling method, the growth process of microalgae under the influence of light-dark cycle is effectively simulated, and the prediction accuracy of the dynamic model under natural light-dark cycle environment is improved.
[0057] (3) By fitting the mathematical models of light cycle and dark cycle respectively through the two-stage parameter fitting method, the problem that the model is completely non-differentiable due to multi-stage modeling is solved, which makes it impossible to fit the dynamic parameters by conventional methods.
[0058] (4) Since the state at any time during the growth of microalgae is related to the previous state, the LSTM model is used as the prediction model to effectively capture the temporal relationship of microalgae growth and improve the prediction accuracy of the model.
[0059] (5) This invention cleverly utilizes the Two-Stage TrAdaBoost.R2 algorithm, a transfer learning strategy, to extract valuable information from the rich simulation data generated by the dynamic model, supplementing and amplifying the value of limited actual microalgae cultivation experimental data. This significantly improves the prediction accuracy of microalgae growth dynamics without heavily relying on new experimental data. This method not only optimizes data utilization efficiency but also enhances the model's generalization ability to unseen data through transferred knowledge, effectively alleviating the difficulty of data acquisition in industrial deep learning model applications and expanding the model's prediction boundary under large data volumes. Therefore, this invention not only improves prediction performance but also greatly alleviates the time and economic cost pressures of biological experiments, providing an efficient and economical solution for optimizing microalgae cultivation conditions and predicting biomass production. Attached Figure Description
[0060] Figure 1 This is a flowchart of the process of the present invention;
[0061] Figure 2 This is a multi-stage model operation process;
[0062] Figure 3 The process of training a model using the Two-stage TrAdaboost.R2 algorithm;
[0063] Figure 4 The prediction effect of microalgae growth curves under different scenarios of the present invention; Detailed Implementation
[0064] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.
[0066] (1) Based on experience, obtain the microalgae growth curves for the fitting group: design the nitrogen concentration as 335 mg·L. -1 The phosphorus concentration is 6.4 mg / L. -1 The light intensity range is 26 μmol·m -2 ·s -1 The light-to-dark ratio was 14:10. The OD values of the fitted group were measured every 4 hours using absorbance analysis. The samples were then processed into algal powder, and the corresponding dry weight was used as the biomass concentration.
[0067] Take two 2ml centrifuge tubes and dry them in an oven at 80℃ for 24 hours. Weigh the centrifuge tubes (m1) on a high-precision balance. At the same time, take 4ml of algal solution and put it into a centrifuge tube. Centrifuge at 3000r / min for 13min. Wash with ultrapure water and repeat three times. Collect the supernatant. Dry the centrifuge tubes and algal powder in an oven at 80℃ for 24 hours. Weigh the centrifuge tubes and algal powder (m2) on a high-precision balance. Calculate the dry weight of microalgae at the corresponding OD according to formula (1).
[0068]
[0069] In equation (1), X is the microalgal biomass concentration, in mg·L. -1 .
[0070] The regression relationship between OD values and biomass concentration values was determined. This relationship will then be used to uniformly convert OD values into biomass concentrations for the microalgal system.
[0071] (2) Based on the Box-Behnken design method, initial microalgal growth curve data were obtained as the source domain data for the training group and the fitting group data. Nitrogen concentration, phosphorus concentration, light intensity, and light-dark ratio were selected as key environmental factors, and variable ranges were designed based on literature and experience. The nitrogen concentration range was designed to be 335 mg·L⁻¹. -1 -385mg·L -1 The phosphorus concentration range is 6.4 mg / L. -1 -25mg·L -1 The light intensity range is 26 μmol·m -2 ·s -1 -78μmol m -2 ·s -1 The light-to-dark ratio ranges from 12:12 to 24:0;
[0072] (3) Considering the light-dark cycle, construct a microalgae growth dynamics model based on the logistic model;
[0073] The microalgae cultivation method involved in this invention is applied in industrial environments where light-dark ratios are controlled. Therefore, it is necessary to incorporate the light-dark ratio into the modeling process so that the subsequent transfer learning model can learn the dynamic process of the model under light-dark cycles in a small sample size. The light cycle model is as follows:
[0074]
[0075] Compared to the light cycle model, the dark cycle model eliminates the influence of light intensity, which is only reflected in the formula for the maximum specific growth rate.
[0076] μ max =μ m ·f(N)·f(P) (10)
[0077] The model operates by alternating between the light loop model and the dark loop model, with the output of the previous model serving as the input for the next. For detailed execution instructions, please refer to [link to relevant documentation]. Figure 2 .
[0078] Given that the dynamic model contains numerous empirical parameters that are difficult to determine directly, the experimental data-driven parameter calibration process is particularly important. Considering that model construction involves both light and dark cycles, this invention adopts the differential evolution algorithm as the parameter fitting optimization method. The differential evolution algorithm is an efficient non-gradient optimization strategy that aims to accurately optimize these key parameters by minimizing the mean square error between experimental observations and model simulation outputs. To further improve the accuracy and efficiency of parameter estimation, this invention implements a two-stage parameter optimization strategy to ensure that the model more accurately reflects the actual situation.
[0079] (4) Based on the designed experiment, use the kinetic model to generate complete kinetic curve data under the corresponding conditions;
[0080] (5) Use the Two-Stage TrAdaboost.R2 method to fuse experimental data and simulation data to train the LSTM model;
[0081] (5.1) The experimental data and dynamic simulation data under the training set cultivation conditions are spliced together. The experimental data serves as the target domain data T in transfer learning. target =(X exp ,Y exp Simulation data, as the source data T source =(X model ,Y model ); splicing the two together results in T = {(X model ,X exp ),(Y model ,Y exp )}.
[0082] (5.2) Use an LSTM model as the base learner. Train it using the Two-Stage TrAdaboost.R2 method;
[0083] To compensate for the impact of data quality on model training in basic LSTM model training, this invention employs the concept of transfer learning and uses the Two-Stage TrAdaBoost.R2 algorithm to fuse a large amount of simulation data with experimental data that is similar to the actual data distribution. Step (5.2) of the Two-Stage TrAdaBoost.R2 algorithm includes the following steps, the specific process of which can be found in [link to flowchart]. Figure 3 .
[0084] (5.2.1) Set the number of iterations S, t = 1, ..., S, and set the initial sample weights:
[0085] (5.2.2) Update the target domain instance according to the weight update strategy of the Adaboost.R2 algorithm. Keep the first n weights, i.e., the source domain weights, unchanged, and gradually reduce the weights of all source domain datasets T. source The weights of the samples in the middle are calculated, and their proportion in the total weights ω1 is determined. The optimal weight ratio is obtained through cross-validation. This stage aims to determine the target training set T. target Appropriate weights for instances in the training set T are used to avoid bias in the target training set T. target The amount of data is much smaller than the auxiliary training set T. source The resulting target training set T target The problem of excessively low initial weights, and the application of auxiliary training set T. source The weights of the samples are updated; the weight update method is to call the base regressor to obtain a learner on the merged training set T, and then calculate the weight adjustment error for each sample. Reduce weights based on error values;
[0086] (5.2.3) Based on the determined weight ratio, update the target domain instances according to the weight update strategy of the Adaboost.R2 algorithm. Weight distribution, i.e., weighting data points with larger errors, is performed during this stage. The auxiliary training set T... source The weights of the samples remain unchanged; only the hypotheses generated in this stage are stored and used to determine the output of the resulting model; through cross-validation, the hypothesis on the target training set T is found. target The model with the smallest regression error is taken as the final training result. The weight update rule is as follows:
[0087]
[0088] Z t β is the normalization constant. tThe goal is to balance the weight of the source domain data with the overall weight ω1 using a binary search. The search objective is to make the overall weight of the target domain equal to... This objective is not set for individual sample weights.
[0089] (5.2.4) The model with the smallest output error, model t , t = argmin(error) t ).
[0090] (6) Model prediction: Using the trained LSTM model, the new input data (including light intensity, light-dark ratio, N element concentration, and P element concentration) are predicted to obtain the predicted results of the microalgae growth curve.
[0091] The training data was split using both 5-fold and 2-fold validation methods, while maintaining the same test set, allowing the model to be trained on two different sized training sets. The former represents the case with a more abundant dataset, while the latter represents the case with a smaller sample size. The results were compared between an unmodified LSTM and a Trans-LSTM improved by the method described in this invention, and a model improvement rate (IMR) metric was defined to characterize the performance improvement of our method on the same training data. Key environmental factors from the test set were used as inputs to the four trained models to obtain the predicted growth curves, such as... Figure 4 As shown in the figure, (a), (b), (c), (d), and (e) are the prediction results of the trained model under five typical growth conditions. Figure (a) represents the case of low light intensity and low nutrient concentration; Figure (b) represents the case of high light intensity and low light-to-dark ratio; Figure (c) represents the case of medium light intensity; Figure (d) represents the case of high nutrient concentration and no dark cycle; and Figure (e) represents the case where all key factors are moderate. It can be seen that in most cases, both the original LSTM model and the Trans-LSTM model can predict the microalgae growth curve well, which stems from the temporal prediction characteristics of LSTM. However, in the case of Figure (a), although the LSTM model under small sample conditions effectively predicts the microalgae growth trend, its accuracy is insufficient. In the case of Figure (b), the LSTM model obviously failed to learn the characteristics of microalgae biomass concentration fluctuations in the later stages of growth under low light-to-dark ratio conditions, while the Trans-LSTM model, due to its large amount of relevant simulation data as an enhancement method, can capture these features. The overall performance is shown in Table 1. The results show that this invention can effectively improve the prediction accuracy of microalgae growth curves under small sample conditions.
[0092]
[0093] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for predicting microalgal growth curves based on transfer learning under small sample conditions, characterized in that: The method includes the following steps: Step S1: Prepare the cultures of the training group and the fitting group using the Box-Behnken method, and obtain the experimental data of the training group and the fitting group {key environmental influencing factors, biomass concentration values of the long-term liquid microalgae system}. The biomass concentration values of the long-term liquid microalgae system characterize the microalgae growth curve. Step S2: Considering the light and dark cycle, construct a microalgae growth kinetic model based on the logistic model; use this model to iteratively learn and train with the fitted group data in the experimental data as input to obtain an ideal microalgae growth kinetic model; Step S3: Using the key environmental influencing factors of the training group in the experimental data as input, the dense dynamic simulation data X″ of the microalgae system biomass concentration value is generated by using the ideal model of microalgae growth dynamics to characterize the dynamic growth curve and output it, thereby achieving generalization enhancement of data volume. Step S4: Using an LSTM model as the prediction model, the {key environmental influencing factors and microalgae biomass concentration values} of the experimental data X′ and simulation data X″ are fused. The fitting learning relationship of the prediction model is iteratively trained using the Two-Stage TrAdaboost.R2 method to obtain the ideal prediction model. The ideal prediction model is then used to predict the corresponding microalgae biomass concentration values for the input key environmental influencing factors. Step S5: Collect key environmental influencing factors of the water sample to be tested over a long period of time, input them into the trained ideal prediction model, and automatically predict and output the biomass concentration value of the microalgae system over a long period of time to characterize the microalgae growth trend over a long period of time.
2. The method for predicting microalgal growth curves based on transfer learning under small sample conditions according to claim 1, characterized in that, The acquisition of experimental data in step S1 includes: Step S11: Collect the long-term OD values of the liquid microalgae system in the training group and the fitting group; Step S12: Fit the calibration curve and regression relationship F between the microalgal OD value and dry weight using the fitted group data; Step S13: Use F to calibrate the microalgae OD value of the training group and obtain its corresponding calibrated biomass concentration value X′. Establish the {key environmental influencing factors and biomass concentration of long-term liquid microalgae system} for the training group and the fitting group.
3. The method for predicting microalgal growth curves based on transfer learning under small sample conditions according to claim 2, characterized in that, Step S11 includes: Step S11a: Prepare multiple training group microalgae culture solutions using culture medium and in combination with the variable range of key environmental factors, and collect the OD values of multiple long-term liquid microalgae systems under different variable environments by absorbance method; Step S11b: Prepare the microalgae culture medium for the fitting group using a culture medium and in combination with standard key environmental factors, and collect the OD value of the long-term liquid microalgae system under standard variable environment by absorbance method.
4. The method for predicting microalgal growth curves based on transfer learning under small sample conditions according to claim 3, characterized in that, The key environmental factors include nitrogen concentration, phosphorus concentration, light intensity, and light-dark ratio; the nitrogen concentration range is 335 mg·L⁻¹. -1 -385mg·L -1 The phosphorus concentration range is 6.4 mg·L⁻¹. -1 -25mg·L -1 The light intensity range is 26 μmol·m -2 ·s -1 -78 μmol·m -2 ·s -1 The light-to-dark ratio ranges from 12:12 to 24:0; The preparation of multiple fitted microalgae culture media using culture medium and in conjunction with standard key environmental factors includes: setting the nitrogen element concentration range to 335 mg·L⁻¹. -1 -385mg·L -1 The phosphorus concentration range is 6.4 mg·L⁻¹. -1 -25mg·L -1 The light-to-dark ratio is 14:10, and the light intensity is 26 μmol·m⁻². -2 ·s -1 The culture was carried out for 120 hours until the microalgae culture medium entered the quiescent phase.
5. The method for predicting microalgal growth curves based on transfer learning under small sample conditions according to claim 3, characterized in that, Microalgae growth curve data were collected using a UV-Vis spectrophotometer, measuring the absorbance (OD) of the microalgae at a wavelength of 750 nm at sampling intervals. 750 The microalgae growth curve data under standard variable environment were obtained by measuring OD values every 4 hours.
6. The method for predicting microalgal growth curves based on transfer learning under small sample conditions according to claim 3, characterized in that, The calibration curve and regression relationship between the fitted microalgal OD value and dry weight include: preparing microalgal dry powder from the microalgal culture medium of the fitted group, calculating the microalgal biomass concentration X based on the mass and volume of the dry powder, and characterizing the long-term growth curve data of the fitted group solid microalgal; obtaining the calibration curve and regression relationship F between the microalgal OD value and dry weight based on the OD value and the corresponding microalgal biomass concentration X at the OD.
7. The method for predicting microalgal growth curves based on transfer learning under small sample conditions according to claim 1, characterized in that, Step S2 includes: Step S21, establishing a microalgae growth kinetic model, which includes sequentially interleaving multiple light cycle models and dark cycle models; a) Using key environmental factors affecting microalgae cultivation as input variables and microalgae biomass concentration as the output variable, a growth kinetics loop model is constructed, as shown in the following equation: Factors related to light, nitrogen, and phosphorus are introduced and applied in a multiplicative model to the maximum specific growth rate μ in the logistic model. max As shown in equations (1)-(2); the light intensity is modeled using the Aiba equation, as shown in equation (3); the nitrogen and phosphorus elements are modeled using the Andrews equation, as shown in equations (4)-(5); the attenuation law of light propagating along the medium is described according to the Lamb-Beer law, as shown in equation (6). m max =μ m ·f(I)·f(N)·f(P) (2) I=I0exp(-K a XZ) (6) Where X(t) is the microalgal biomass concentration at time t, in g·L⁻¹. -1 μ max Represents the maximum ratio growth rate, in h. -1 ;X max The maximum biomass concentration that the environment can sustain, expressed in g / L. -1 μ d The decay coefficient represents the rate at which microalgal cell populations decay in the environment, expressed in hours (h). -1 ; Where f(I), f(N), and f(P) represent changes in light intensity, nitrogen element, and phosphorus element, respectively; where μ m The ratio of growth rate constant, in units of h. -1 I represents light intensity, in μmol·m⁻¹ -2 ·s -1 ;K I,i is the light inhibition constant, used to represent the phenomenon that excessive light intensity inhibits photosynthesis, and its unit is μmol·m. -2 ·s -1 ;K I,s is the light saturation constant, used to indicate that higher light intensity does not significantly improve photosynthesis; the unit is μmol·m. -2 ·s -1 N and P are the concentrations of nitrogen and phosphorus, respectively, in mg·L. -1 ;K N,i ,K P,i These are the inhibition constants for nitrogen and phosphorus, respectively, in mg·L. -1 ;K N,s ,K P,s These are the saturation constants for nitrogen and phosphorus, respectively, in mg·L. -1 ; Where I0 represents the incident light intensity, with units of μmol·m -2 ·s -1 Ka is the optical attenuation coefficient, with units of L·m. -1 ·g -1 Z represents the distance behind the irradiated area, in meters. To reduce the impact of spatial dimension on the solution of the equation, the ten-step trapezoidal rule is used to transform equation (6) into equation (7): The matrix consumption rate can be calculated using the mass conservation equation: Among them, Y X / N ,Y XP It is the yield coefficient relative to the biomass production of the substrate; m N ,m P It is the specific consumption rate for cell maintenance; b) Consider the light cycle process and the dark cycle process respectively, and use formulas (1)-(5) and (7)-(10) as the light cycle model; Formulas (1), (4), (5), (9)-(11) are used as the dark cycle model; and the maximum specific growth rate is expressed by the following formula: m max =μ m ·f(N)·f(P) (12).
8. The method for predicting microalgal growth curves based on transfer learning under small sample conditions according to claim 1, characterized in that, Step S2 further includes: Step S22, using the model to perform iterative learning with key environmental influencing factors as input, including using a two-stage parameter fitting method combined with a differential evolution algorithm to fit the parameters; including the following steps: Step S221: Set the light-dark ratio and divide the fitted group data into light cycle data and dark cycle data; each part of the data is {key environmental influencing factors, long-term liquid microalgae system concentration}; Step S222: Initialize the parameters of the optical cycle model, fix the parameters of the dark cycle model, and use the optical cycle data from step S221 to perform parameter fitting only on the optical cycle model to obtain the dynamic parameters of the optical cycle model. Step S223: Fix the dynamic parameters of the light cycle model obtained in step S222, and use the dark cycle data in step S221 to perform parameter fitting on the dark cycle model to obtain the dynamic parameters of the dark cycle model.
9. The method for predicting microalgal growth curves based on transfer learning under small sample conditions according to claim 1, characterized in that, Step S4 includes the following steps: Step S41: Combine the experimental data and dynamic simulation data under the training set cultivation conditions, where the experimental data serves as the target domain data T in transfer learning. target =(X exp ,Y exp Simulation data serves as the source domain data T. source =(X model ,Y model ); splicing the two together results in T = {(X model ,X exp ),(Y model ,Y exp )}; Step S42: Construct an LSTM model as a base learner and train it using the Two-Stage TrAdaboost.R2 method.
10. The method for predicting microalgal growth curves based on transfer learning under small sample conditions according to claim 9, characterized in that, The Two-Stage TrAdaBoost.R2 algorithm in step S42 includes the following steps: Step S421: Set the number of iterations S, t = 1, ..., S, and set the initial sample weights: Step S422: Update the target domain instances according to the weight update strategy of the Adaboost.R2 algorithm; keep the source domain weights unchanged, and gradually reduce the weights of all target datasets T. source The weights of the middle samples and their proportion in the total weights ω1 are determined, and the optimal weight ratio is obtained through cross-validation; the auxiliary training set T is used. source The weights of the samples are updated; the weight update method is to call the base regressor (LSTM) to obtain a learner on the merged training set T, and then calculate the weight adjustment error for each sample. Reduce weights based on error values; Step S423: Based on the determined weight ratio, update the target domain instances according to the weight update strategy of the Adaboost.R2 algorithm; perform weighted processing on data points with large errors; during this stage, the auxiliary training set T... source The weights of the samples remain unchanged; through cross-validation, the sample in the target training set T is found. target The model with the smallest regression error is taken as the final training result; The weight update rules are as follows: Z t β is the normalization constant. t The objective is to balance the weight of the source domain data with the overall weight ω1 using a binary search method; the search objective is to make the overall weight of the target domain equal to... This objective is not set for individual sample weights; Step S424: Output the model with the smallest error. t , t = argmin(error) t ).
Citation Information
Patent Citations
Photovoltaic output interval prediction method in small sample scene based on transfer learning
CN114897264A
Microalgae component detection method based on deep learning
CN117710816A