Optimization method for expression of algae cell secreting type recombinant protein
By combining deep neural network models and multi-level evaluation systems, the problem of yield fluctuations in the expression of recombinant proteins in algal cells was solved, enabling precise control and stability improvement of the fermentation process, thus ensuring product supply stability and production efficiency.
Patent Information
- Application Number
- CN202511325025.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-12-30
AI Technical Summary
In existing technologies, the yield of recombinant proteins expressed in algal cells fluctuates significantly, leading to unstable product supply and making it difficult to standardize process parameters and achieve multi-parameter synergistic optimization.
By employing a deep neural network model combined with a multi-level evaluation system, and constructing quality feature matrices and efficiency feature matrices, precise regulation of the recombinant protein expression process in algal cells is achieved. This includes the calculation of a pre-trained deep neural network model, process gain index, compensation index, and optimization index, combined with fuzzy logic control methods for real-time optimization.
It enables precise control of the expression process of recombinant proteins in algal cells, significantly reduces the yield difference between batches and in single-batch fermentation processes, and improves the stability and efficiency of production.
Smart Images

Figure CN121237208A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of bioinformatics and synthetic biology, and specifically relates to an optimized method for the expression of secretory recombinant proteins in algal cells. Background Technology
[0002] Algal cell recombinant protein expression systems are widely used in biopharmaceuticals, clinical diagnostics, industrial enzyme preparations, and functional proteins due to their advantages such as rapid growth, low cost, and high safety. Traditional algal cell recombinant protein expression processes mainly employ fixed-parameter control schemes, including preset key parameters such as light intensity, temperature, stirring speed, and aeration rate. This method is simple to operate and easy to standardize in industrial production, and remains the mainstream production mode. Process parameters are typically set based on laboratory optimization results and historical experience data, and fermentation conditions are maintained through simple process parameter monitoring and manual intervention.
[0003] However, the yield of recombinant proteins from algal cells often fluctuates significantly between different batches. This fluctuation manifests itself in several ways: First, even with the same process parameters and operating procedures, protein expression levels can differ by several times between batches, severely impacting product supply stability; second, protein expression levels often fluctuate irregularly during a single batch of fermentation, making it difficult to predict and control the optimal harvest time; third, even minor changes in external conditions such as environmental factors and batch-specific raw material variations can lead to significant yield fluctuations, making it difficult to achieve the desired results from standardizing process parameters.
[0004] The root cause of this yield fluctuation problem lies in the lack of precise control over the dynamic processes of algal cell growth and protein expression in current technologies. Fixed-parameter schemes cannot adapt to the dynamic changes in cellular metabolic states, while simple manual interventions often lag behind actual needs, resulting in unsatisfactory optimization effects. Furthermore, traditional methods struggle to establish the correlation between various process parameters, failing to achieve synergistic optimization of multiple parameters, which further exacerbates yield instability. Therefore, resolving the yield fluctuation problem has become a key bottleneck in improving the production efficiency of recombinant proteins from algal cells. Summary of the Invention
[0005] In view of this, the present invention provides an optimized method for the expression of secretory recombinant proteins in algal cells, which can solve the technical problem of significant yield fluctuations during the expression of recombinant proteins in algal cells in the prior art.
[0006] This invention is implemented as follows: An optimized method for the expression of secretory recombinant proteins in algal cells includes the following steps: collecting algal cell samples and screening the culture medium to obtain initial culture conditions with a cell density of 1,000,000 to 2,000,000 cells per milliliter, a protein expression level greater than 500 mg / L, and cell viability greater than 90%; inoculating the screened algal cells into the culture medium, recording and setting the initial culture parameters; collecting culture process parameters; constructing a feature matrix, which consists of a quality feature matrix and an efficiency feature matrix; importing a pre-trained deep neural network model, processing it through 5 hidden layers, each containing 64 to 128 neurons, and outputting an optimized culture parameter set and the optimal induction time point; adjusting the culture parameters according to the optimized culture parameter set; calculating the process gain index, compensation index, and optimization index; constructing a process judgment function; adding an inducer at the optimal induction time point; monitoring the real-time protein expression level, and collecting the fermentation supernatant when the process judgment function outputs a qualified judgment result; processing the fermentation supernatant by centrifugation; and filtering the centrifuged fermentation supernatant through a membrane to obtain the optimized recombinant protein product.
[0007] The culture process parameters include real-time light intensity, real-time temperature, real-time stirring speed, real-time aeration rate, real-time cell density, real-time protein expression level, and real-time cell viability; the initial light intensity is 5000 to 8000 lux, the initial temperature is 22 to 26°C, the initial stirring speed is 100 to 200 rpm, and the initial aeration rate is 0.5 to 1.0 L / min.
[0008] The quality feature matrix records real-time protein content, real-time purity, real-time bioactivity, real-time aggregation rate, and real-time degradation rate, while the efficiency feature matrix records real-time yield, real-time conversion rate, real-time expression efficiency, real-time cell density, and real-time glucose consumption.
[0009] The training dataset for the pre-trained deep neural network model includes historical fermentation data from 10,000 batches of recombinant protein. The training dataset construction steps include data collection, data cleaning, data standardization, data segmentation, and data augmentation. The training steps include model initialization, parameter optimization, model validation, model selection, and model ensemble.
[0010] Among them, the neuron weight parameters of the last four hidden layers out of the five hidden layers are determined by the backpropagation function. The input of the backpropagation function includes the current hidden layer input data, the target output data, the learning rate parameter, the momentum factor, the error threshold, and the neuron weight parameters of the previous hidden layer. The neuron weight parameters of the first hidden layer are determined by the autoencoder function.
[0011] The process gain index is calculated using a comprehensive evaluation equation. The inputs of the comprehensive evaluation equation include a quality characteristic matrix, an efficiency characteristic matrix, a quality weight coefficient, an efficiency weight coefficient, a time decay coefficient, and a process fluctuation coefficient. The output is a quantitative index characterizing fermentation performance.
[0012] The process compensation index is calculated using a deviation correction equation. The inputs to the deviation correction equation include the current process gain index, the historical average gain index, the standard deviation coefficient, the time series coefficient, the seasonality factor, and the process drift coefficient. The output is the correction value for batch-to-batch differences.
[0013] The process optimization index is calculated using an effect evaluation equation. The inputs of the effect evaluation equation include the process gain index before parameter adjustment, the process gain index after parameter adjustment, the time response coefficient, the parameter sensitivity coefficient, the system inertia coefficient, and the environmental impact coefficient. The output is the effect score of the process optimization.
[0014] The process judgment function takes the compensation index and the optimization index as inputs, uses fuzzy logic control to evaluate and optimize the fermentation process in real time, converts it into a three-level fuzzy set (low, medium, and high) through fuzzification, and uses a trapezoidal membership function to output process optimization suggestions and qualification judgment results.
[0015] The concentration of the inducer is 0.5 to 2 mmol / L; the centrifugation speed is 3000 to 5000 rpm; the centrifugation time is 15 to 20 minutes; the membrane filtration adopts a multi-stage series filtration system, including a 5-micron pre-filtration membrane, a 0.45-micron intermediate filtration membrane, and a 0.22-micron terminal filtration membrane.
[0016] Compared with existing technologies, this invention provides an optimization method for the expression of secretory recombinant proteins in algal cells. The intelligent optimization method proposed in this invention achieves precise control over the expression process of recombinant proteins in algal cells by establishing a deep neural network model and a multi-level evaluation system, effectively solving the problem of yield fluctuations. This method is the first to integrate quality feature matrices and efficiency feature matrices into a unified optimization framework, enabling comprehensive monitoring and predictive adjustment of the fermentation process.
[0017] The method of this invention has significant technical advantages. By using a pre-trained deep neural network model, it systematically learns the patterns and characteristics in historical fermentation data, establishes a mapping relationship between process parameters and yield, and realizes intelligent parameter adjustment. Through real-time calculation of the process gain index, it comprehensively evaluates the fermentation state and promptly identifies factors that may cause yield fluctuations. The introduction of a compensation index effectively balances batch-to-batch differences and improves process stability. The calculation of the optimization index quantifies the effect of parameter adjustment, providing a basis for continuous optimization.
[0018] This invention successfully solves the technical problem of significant yield fluctuations. Its core lies in establishing a data-driven dynamic optimization system. This system can accurately capture the changing trends during fermentation, enabling predictive adjustments to process parameters, while ensuring the reliability of optimization through multi-dimensional evaluation indicators. This method not only significantly reduces yield differences between batches but also improves the stability of the single-batch fermentation process. Attached Figure Description
[0019] Figure 1 This is a flowchart of the method of the present invention.
[0020] Figure 2 This is a graph showing the dynamic changes in PD-1 protein expression, cell density, and glucose concentration in Example 2.
[0021] Figure 3 This is a graph showing the trends in the bioactivity, purity, and aggregation rate of the PD-1 protein in Example 2. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0023] like Figure 1 The diagram shown is a flowchart of an optimized method for expressing secretory recombinant proteins in algal cells, provided by this invention. This method includes the following steps:
[0024] S01. Collect algal cell samples and screen the culture medium to obtain initial culture conditions with a cell density of 1,000,000 to 2,000,000 per milliliter, a protein expression level of more than 500 mg per liter, and a cell viability of more than 90%.
[0025] S02. The selected algal cells were inoculated into the culture medium, and the initial light intensity was recorded as 5000 to 8000 lux, the initial temperature as 22 to 26°C, the initial stirring speed as 100 to 200 rpm, and the initial aeration rate as 0.5 to 1.0 L / min.
[0026] S03. Collect culture process parameters, including real-time light intensity, real-time temperature, real-time stirring speed, real-time aeration rate, real-time cell density, real-time protein expression level, and real-time cell activity.
[0027] S04. Construct a feature matrix, which consists of a quality feature matrix and an efficiency feature matrix. The quality feature matrix records real-time protein content, real-time purity, real-time bioactivity, real-time aggregation rate, and real-time degradation rate. The efficiency feature matrix records real-time yield, real-time conversion rate, real-time expression efficiency, real-time cell density, and real-time glucose consumption. The feature matrix is used to characterize the quality and efficiency parameters of the culture process.
[0028] S05. Import the pre-trained deep neural network model. The input layer of the pre-trained deep neural network model receives the feature matrix and processes it through 5 hidden layers. Each hidden layer contains 64 to 128 neurons. The output layer generates the optimized culture parameter set and the best induction time point.
[0029] S06. Based on the optimized culture parameter set output by the pre-trained deep neural network model, adjust the real-time light intensity, the real-time temperature, the real-time stirring speed, and the real-time aeration rate.
[0030] S07. Calculate the process gain index, which is obtained based on the quality feature matrix and the efficiency feature matrix. The process gain index is used to characterize the overall performance of the recombinant protein expression process.
[0031] S08. Calculate the compensation index, which is obtained based on the deviation between the process gain index and the preset benchmark value. The compensation index is used to quantify the differences between batches.
[0032] S09. Calculate the optimization index, which is obtained based on the change in the process gain index before and after parameter adjustment. The optimization index is used to evaluate the effect of process parameter adjustment.
[0033] S10. Construct a process judgment function, wherein the process judgment function takes the compensation index and the optimization index as input and outputs process optimization suggestions and qualification judgment results; wherein, the process judgment function is used to determine whether the cultivation process has achieved the optimization target;
[0034] S11. Add an inducer at the optimal induction time point, wherein the concentration of the inducer is 0.5 to 2 mmol per liter;
[0035] S12. Monitor the real-time protein expression level. When the pass / fail result output by the process judgment function is qualified, collect the fermentation supernatant.
[0036] S13. The fermentation supernatant is centrifuged at a speed of 3000 to 5000 revolutions per minute for 15 to 20 minutes.
[0037] S14. The fermentation supernatant after centrifugation is filtered through a membrane with a pore size of 0.22 micrometers to obtain the optimized recombinant protein product.
[0038] The training dataset for the pre-trained deep neural network model includes 10,000 historical batches of recombinant protein fermentation data. Each batch of data records the light intensity, temperature, stirring speed, aeration rate, cell density, protein expression level, cell activity, protein content, purity, bioactivity, aggregation rate, degradation rate, yield, conversion rate, expression efficiency, and glucose consumption throughout the entire culture process. The training dataset construction steps include data acquisition, data cleaning, data standardization, data segmentation, and data augmentation. The training steps include model initialization, parameter optimization, model validation, model selection, and model ensemble.
[0039] In this design, the neuron weights of the last four hidden layers out of the five hidden layers are determined using a backpropagation function. The inputs to this function include the current hidden layer input data, the target output data, the learning rate parameter, the momentum factor, the error threshold, and the neuron weights of the previous hidden layer. The standard rank of the process control features is used as a regularization constraint for weight updates. The neuron weights of the first hidden layer are determined using an autoencoder function. The inputs to this autoencoder function include the feature matrix data, dimensionality reduction coefficients, the reconstruction error threshold, the learning rate parameter, the number of iterations, and the activation function type. The standard rank of the process control features is determined using a standard rank feature extraction function. The inputs to this function include algal cell growth curves, protein accumulation curves, glucose consumption curves, dissolved oxygen change curves, pH change curves, and fermentation cycle length. The standard rank reflects the correlation strength between process parameters and the complexity of the data structure, maintaining the stability of weight updates and ensuring the model captures key process features. This allows the deep neural network model to minimize errors while maintaining sensitivity to process features during training, thus improving the model's generalization ability and practicality.
[0040] The process judgment function is used to evaluate the current fermentation status and determine the process parameter adjustment plan. The inputs include the compensation index, optimization index, quality characteristic matrix, efficiency characteristic matrix, historical best batch parameters, and target yield parameters. The outputs include process parameter adjustment suggestions and qualification judgment results.
[0041] The process gain index is calculated using a comprehensive evaluation equation, which is used to quantify the overall performance of the fermentation process. The inputs include a quality characteristic matrix, an efficiency characteristic matrix, a quality weight coefficient, an efficiency weight coefficient, a time decay coefficient, and a process fluctuation coefficient. The output is a quantitative index characterizing the fermentation performance.
[0042] The process compensation index is calculated using a deviation correction equation, which is used to balance batch-to-batch differences. The inputs include the current process gain index, historical average gain index, standard deviation coefficient, time series coefficient, seasonality factor, and process drift coefficient. The output is the correction value for batch-to-batch differences.
[0043] The process optimization index is calculated using an effect evaluation equation, which measures the effectiveness of process parameter adjustments. The inputs include the process gain index before parameter adjustment, the process gain index after parameter adjustment, the time response coefficient, the parameter sensitivity coefficient, the system inertia coefficient, and the environmental impact coefficient. The output is the effect score of process optimization.
[0044] The specific implementation methods of the above steps are described in detail below. Step S01 is implemented using high-throughput culture medium screening technology. First, 96 different combinations of culture media are prepared, each containing different concentrations of nitrogen source, carbon source, inorganic salts, and trace elements. Small-scale culture experiments are conducted in 96-well plates. By comparing the growth status of algal cells in different culture media, the optimal culture medium composition is selected. Cell density is detected using flow cytometry, protein expression levels are detected using protein content assays, and cell viability is detected using trypan blue staining. The selected culture media have carbon source concentrations ranging from 10 to 30 g / L, nitrogen source concentrations ranging from 2 to 6 g / L, phosphorus source concentrations ranging from 0.5 to 2 g / L, and trace element concentrations ranging from 0.1 to 0.5 g / L. During the screening process, principal component analysis is used to optimize the culture medium formulation, establishing a correlation model between culture medium components and cell growth indicators to determine the optimal culture medium formulation.
[0045] The specific implementation of step S02 involves inoculating the screened algal cells into a 5-liter fermenter for scale-up culture, with an inoculation density of 50,000 to 100,000 cells per milliliter. An intelligent fermentation control system is used to precisely control the culture parameters, wherein the light intensity is provided by an adjustable LED light source, the temperature is controlled by the jacketed circulating water, the stirring speed is controlled by a variable frequency motor, and the aeration rate is controlled by a mass flow meter. The initial set values of the culture parameters are determined by response surface methodology, and a mathematical model between the culture parameters and the target product yield is established to obtain the optimal combination of initial parameters.
[0046] The specific implementation of step S03 involves using an online monitoring system to collect culture process parameters in real time. Light intensity is collected every 5 minutes using a light intensity sensor; temperature is collected every 1 minute using a temperature sensor; stirring speed is collected every 1 minute using a speed sensor; aeration rate is collected every 1 minute using a gas flow meter; cell density is collected every 30 minutes using an online turbidimeter; protein expression level is collected every 30 minutes using an online fluorescence detector; and cell viability is collected every 60 minutes using an online viability detector. The collected data is transmitted to the central control system via a data acquisition card, and outliers and noise are removed after data preprocessing.
[0047] The specific implementation of step S04 involves constructing a feature matrix. The quality feature matrix includes five indicators: real-time protein content determined by high-performance liquid chromatography, real-time purity determined by gel filtration chromatography, real-time bioactivity determined by enzyme activity assay, real-time aggregation rate determined by dynamic light scattering, and real-time degradation rate determined by protein quantification. The efficiency feature matrix includes five indicators: real-time yield determined by gravimetric method, real-time conversion rate determined by material balance algorithm, real-time expression efficiency determined by quantitative fluorescence method, real-time cell density determined by cell counting method, and real-time glucose consumption determined by glucose oxidase method. The feature matrix is preprocessed using a data standardization method to eliminate the influence of different dimensions.
[0048] The specific implementation of step S05 involves importing a pre-trained deep neural network model. The model is constructed using a deep learning framework. The input layer receives 10 feature parameters, and feature extraction and pattern recognition are performed through 5 hidden layers. Each hidden layer uses a modified linear unit activation function to prevent gradient vanishing. The output layer uses a linear activation function to generate 4 optimized training parameters and 1 optimal induction time point. The model training uses the stochastic gradient descent algorithm with a learning rate of 0.001, a batch size of 64, and 1000 training epochs. The model validation uses cross-validation with a validation set ratio of 20% and a test set ratio of 10%. The model evaluation uses root mean square error and coefficient of determination as evaluation metrics.
[0049] The specific implementation of step S06 involves real-time control of the fermentation system based on the optimized culture parameter set output by the deep neural network model. The optimized parameters are then distributed to each execution unit via a programmable controller. Light intensity is adjusted using pulse width modulation (PWM) technology, achieving continuous adjustment by changing the duty cycle of the LED light source, with an adjustment range of 1000 to 20000 lux and an adjustment accuracy of 100 lux. Temperature is adjusted using a proportional-integral-derivative (PID) control algorithm, achieving precise control by adjusting the opening of the circulating water valve, with an adjustment range of 15 to 35℃ and a control accuracy of 0.1℃. Stirring speed is adjusted using variable frequency drive (VFD) technology, achieving dynamic adjustment by changing the motor speed, with an adjustment range of 50 to 500 revolutions per minute and a control accuracy of 1 revolution per minute. Aeration is adjusted using mass flow control technology, achieving precise control by adjusting the gas flow valve opening, with an adjustment range of 0.1 to 5.0 liters per minute and a control accuracy of 0.1 liters per minute.
[0050] The specific implementation of step S07 involves calculating the process gain index, using a comprehensive evaluation equation to quantitatively assess the fermentation process, and transforming the quality characteristic matrix and efficiency characteristic matrix into a single evaluation index through a weighted summation method. The quality weight coefficients are determined using the analytic hierarchy process (AHP), including a protein content weight of 0.3, purity weight of 0.2, bioactivity weight of 0.2, aggregation rate weight of 0.15, and degradation rate weight of 0.15. The efficiency weight coefficients are determined using an expert scoring method, including a yield weight of 0.3, conversion rate weight of 0.2, expression efficiency weight of 0.2, cell density weight of 0.15, and glucose consumption weight of 0.15. The time decay coefficient is calculated using an exponential decay function to reflect the stability of process parameters over time. The process fluctuation coefficient is calculated using the coefficient of variation to reflect the degree of fluctuation of process parameters.
[0051] The specific implementation of step S08 involves calculating a compensation index and using a deviation correction equation to quantify and compensate for batch-to-batch differences. The deviation between the current process gain index and the historical average gain index is converted into a compensation index through standardization. The standard deviation coefficient is used to measure the range of process fluctuations, with a value ranging from 0.1 to 0.3. The time series coefficient is calculated using the exponential smoothing method to reflect the time correlation of process parameters, with a smoothing coefficient ranging from 0.1 to 0.3. The seasonality factor is calculated using the seasonal decomposition method to reflect the periodic changes of process parameters, with a decomposition order of 12. The process drift coefficient is calculated using the linear regression method to reflect the long-term trend of process parameters, with a confidence interval of 95%.
[0052] The specific implementation of step S09 involves calculating the optimization index and using an effect evaluation equation to quantitatively evaluate the effect of process parameter adjustment. The change in the process gain index before and after parameter adjustment is obtained through differential calculation. The time response coefficient is calculated using a first-order lag model to reflect the system response speed, with a time constant ranging from 0.1 to 1.0 hours. The parameter sensitivity coefficient is calculated using sensitivity analysis to reflect the degree of influence of process parameters on the target index, with a sensitivity threshold of 0.1. The system inertia coefficient is calculated using the step response method to reflect the system stability, with an inertia time constant ranging from 0.5 to 2.0 hours. The environmental impact coefficient is calculated using multiple regression analysis to reflect the influence of environmental factors on process parameters, with a significance level of 0.05.
[0053] The specific implementation of step S10 involves constructing a process judgment function and using fuzzy logic control to evaluate and optimize the fermentation process in real time. The compensation index and optimization index are used as input variables and converted into fuzzy sets through fuzzification. A trapezoidal membership function is used, and the fuzziness level is divided into three levels: low, medium, and high. Process optimization suggestions are generated through a fuzzy inference system, including the direction and magnitude of parameter adjustment. The inference rule base contains 27 rules. The qualified judgment result is obtained through defuzzification and the clear output is calculated using the centroid method. The judgment threshold is 0.8. When the optimization index is greater than 0.9 and the compensation index is less than 0.2, the process is judged to have reached the optimization target.
[0054] The specific implementation of step S11 involves adding the inducer at the optimal induction time point, which is predicted by a deep neural network model and is generally in the mid-to-late logarithmic growth phase of cells. The selection of the inducer considers the physiological state of the cells and the characteristics of the target protein. Commonly used inducers include isopropyl thiogalactoside, methanol, ethanol, etc. The inducer is added using a stepwise addition strategy. The initial addition concentration is 0.5 mmol / L, and the protein expression level is detected every 2 hours. The inducer is replenished according to the changes in expression level, with the replenishment amount being 0.2 to 0.5 mmol / L, and the final concentration not exceeding 2 mmol / L. During the induction process, the cell growth status is closely monitored, including indicators such as cell density, cell viability, and dissolved oxygen level, to ensure that the induction process proceeds smoothly.
[0055] The specific implementation of step S12 involves monitoring real-time protein expression levels using an online fluorescence detection system with a detection interval of 30 minutes; evaluating the fermentation status using a process judgment function, and stopping the fermentation process when the judgment result is qualified (i.e., the optimization index is greater than 0.9 and the compensation index is less than 0.2); collecting the fermentation supernatant using a sterile sampling system, with each sample volume being 50 ml, maintaining system airtightness during sampling to prevent exogenous contamination; immediately pre-cooling the collected supernatant at 4°C to reduce the risk of protein degradation; and cleaning and disinfecting the fermenter after each batch of fermentation to ensure sterile conditions for the next batch of fermentation.
[0056] The specific implementation of step S13 is to process the fermentation supernatant by centrifugation, using a large-capacity refrigerated centrifuge for solid-liquid separation; the centrifugation process adopts a step-by-step speed-up strategy, starting at 1000 rpm for 2 minutes, then increasing to 3000 rpm for 5 minutes, and finally increasing to the target speed of 4000 rpm for 15 minutes; the centrifugation temperature is controlled at 4℃ to reduce the risk of protein degradation and inactivation; the centrifuge tubes are made of biocompatible materials, with a capacity of 500 ml, and the number of centrifuge tubes is determined according to the volume of the fermentation supernatant; during the centrifugation process, attention is paid to rotor balance, and the liquid volume error of each centrifuge tube does not exceed 1 gram.
[0057] The specific implementation of step S14 involves membrane filtration of the centrifuged fermentation supernatant using a multi-stage cascade filtration system. The first stage uses a 5-micron pre-filtration membrane to remove residual cell debris and large particulate impurities. The second stage uses a 0.45-micron intermediate filtration membrane to remove fine particulate matter and some large molecular impurities. The third stage uses a 0.22-micron terminal filtration membrane to ensure the sterility of the product. The filtration process employs a constant pressure filtration mode, with the feed pressure controlled between 0.1 and 0.2 MPa and the filtration temperature controlled at 4°C. The filtration system is equipped with a differential pressure monitoring device, and the filter membrane is replaced when the differential pressure exceeds 0.05 MPa. The filtered product is collected in sterile containers and stored at 0 to 4°C. Product quality testing includes indicators such as protein content, purity, bioactivity, and sterility. Only products that pass the tests can proceed to the next purification step.
[0058] It should be noted that the protein expression level detection in step S03 employs an online detection technology based on fluorescent protein labeling. Specifically, the target recombinant protein is fused with green fluorescent protein (GFP) or other fluorescently labeled proteins to construct a fluorescently labeled recombinant protein expression vector. During fermentation, changes in fluorescence intensity in the culture medium are monitored in real time using an immersion fiber optic fluorescence probe. The excitation wavelength is set to 485 nm, and the emission wavelength is detected at 528 nm. A standard curve is established between fluorescence intensity and recombinant protein expression level, and the dynamic changes in protein expression level are extrapolated by monitoring changes in the fluorescence signal in real time. The detection system is equipped with an automatic calibration device, performing background fluorescence correction every 2 hours to eliminate interference from culture medium components and cell autofluorescence.
[0059] It should be noted that the real-time bioactivity detection in step S04 employs a specific enzyme activity assay, selecting the appropriate substrate reaction system based on the functional characteristics of the target recombinant protein. For recombinant proteins with enzyme activity, spectrophotometry is used to detect the absorbance change of the enzyme-catalyzed reaction product, with the reaction temperature controlled at 37℃ and the pH adjusted to the optimal range. During the detection process, 50 μL of fermentation supernatant is mixed with 200 μL of substrate solution, and kinetic analysis is performed using a microplate reader, recording the rate of absorbance change at wavelengths of 340 nm or 405 nm. By comparing with a standard enzyme activity control, the enzyme activity units per unit time are calculated, establishing a linear relationship between bioactivity and enzyme activity. For non-enzymatic recombinant proteins, appropriate biofunctional detection methods are used, such as binding activity assays or cell activity assays.
[0060] It should be noted that the real-time quantitative fluorescence detection of expression efficiency in step S04 is based on the principle of fluorescence resonance energy transfer (FRET). A dual-fluorescence labeling system is constructed, using cyan fluorescent protein (CFP) to label total cellular protein as an internal control and yellow fluorescent protein (YFP) to label the target recombinant protein as the detection target. The ratio of CFP to YFP fluorescence intensity in a single cell is detected using flow cytometry or fluorescence microscopy, and the expression ratio of recombinant protein relative to total protein is calculated. The expression efficiency is defined as the ratio of the fluorescence intensity of the target protein to the fluorescence intensity of the total cellular protein, multiplied by a correction factor. During the detection process, 1 ml of culture medium is automatically sampled every 30 minutes, and single-cell fluorescence detection is performed using a microfluidic chip. Fluorescence data from at least 1000 cells are statistically analyzed to calculate the average expression efficiency and its standard deviation.
[0061] It should be noted that the online fluorescence detection system in step S12 uses a multi-channel fiber optic fluorescence analyzer for continuous monitoring. The system is equipped with four independent excitation light sources, including 365 nm ultraviolet light, 485 nm blue light, 540 nm green light, and 640 nm red light, corresponding to the detection of different fluorescent protein markers. The fiber optic probe is directly inserted into the fermenter, and the probe tip is equipped with an anti-fouling coating and a self-cleaning device to prevent cell and culture medium components from adhering and affecting detection accuracy. The fluorescence signal is converted into an electrical signal by a photomultiplier tube, and then transmitted to the data acquisition system after being amplified and converted from analog to digital. The system uses time-gating technology to eliminate interference from scattered light, and multi-point calibration ensures that the detection linear range covers the protein concentration changes throughout the entire fermentation process. Optionally, the data acquisition frequency is set to once every 30 seconds to calculate the trend and rate of protein expression changes in real time.
[0062] The specific implementation of the deep neural network model training dataset includes: first, data collection, gathering data from the entire fermentation process of 10,000 batches of recombinant protein from history, with a collection frequency of once every 5 minutes; each batch of data includes 16 parameters such as light intensity and temperature; the data cleaning process uses an outlier detection algorithm to identify and delete outlier data points, a moving average method to handle data noise, and an interpolation algorithm to repair missing data; data standardization uses the min-max standardization method to map all feature values to the interval between 0 and 1, ensuring the comparability of features with different dimensions; data segmentation uses a stratified sampling method to divide the dataset into training, validation, and test sets in an 8:1:1 ratio to ensure the consistency of data distribution in each subset; data augmentation uses the sliding window method to generate time series samples, with a window size of 12 hours and a sliding step size of 1 hour, expanding the sample size to 3 times the original data.
[0063] The specific implementation of deep neural network model training includes the following stages: In the model initialization stage, an autoencoder is used to pre-train the parameters of the first hidden layer. The input is a 16-dimensional feature vector, the dimensionality reduction coefficient is 0.5, the reconstruction error threshold is 0.01, the learning rate is 0.001, the maximum number of iterations is 1000, and the activation function is the hyperbolic tangent function. In the parameter optimization stage, the last four hidden layers are trained using the backpropagation algorithm. The input includes the current layer's feature data and the target output data. The learning rate adopts an adaptive adjustment strategy, with an initial value of 0.01, decaying by 20% every 50 cycles. The momentum factor is set to 0.9, and the error threshold is 0.001. The standard rank of the process control features is used as a regularization constraint for weight updates. The standard rank is calculated through a feature extraction function. The input includes six process curve data, and the output is the matrix rank representing the correlation of process parameters.
[0064] In the model validation phase, cross-validation is used to evaluate model performance, employing 10-fold cross-validation. After each fold, the root mean square error and coefficient of determination are calculated. An early stopping strategy is used to prevent overfitting; training stops when the validation set error does not decrease for 10 consecutive epochs. In the model selection phase, a grid search method is used to optimize hyperparameters. The search space includes 64 to 128 hidden layer neurons, a learning rate of 0.0001 to 0.01, a batch size of 32 to 256, and a regularization coefficient of 0.0001 to 0.01. In the model ensemble phase, a voting method is used to integrate multiple trained models. The prediction results of different models are fused using a weighted average method, with the weight coefficients determined by the validation set performance.
[0065] During training, batch normalization is used to accelerate model convergence, and the output of each hidden layer is standardized with a batch size of 64. Random deactivation is used to prevent overfitting with a deactivation probability of 0.3. Residual connection structures are used to alleviate the gradient vanishing problem, with skip connections added between every two layers. Exponential moving average is used to smooth the parameter update process with a decay rate of 0.999. Learning rate annealing is used to optimize the training process with an initial learning rate of 0.01, which decays by 50% every 1000 epochs. Gradient clipping is used to prevent gradient explosion with a clipping threshold of 5.0.
[0066] After training, the model's final performance is evaluated using a test set, with test metrics including root mean square error, mean absolute error, coefficient of determination, and prediction accuracy. Sensitivity analysis is used to assess the importance of different input features, and the feature importance score is calculated using the ranking importance method. Interpretability analysis is used to understand the model's decision-making mechanism, and the local interpretability method is used to analyze the model's prediction basis under different operating conditions. Stability analysis is used to evaluate the model's robustness, and the Monte Carlo simulation method is used to test the model's sensitivity to input disturbances.
[0067] The mathematical expression of the characteristic matrix is as follows:
[0068]
[0069] In the formula, Q ij E represents the j-th quality characteristic at the i-th time point. ij Let j represent the j-th efficiency feature at the i-th time point, and n be the number of sampling points; where j = 1, 2, 3, 4, 5 correspond to protein content, purity, bioactivity, aggregation rate, and degradation rate in the quality feature matrix, and yield, conversion rate, expression efficiency, cell density, and glucose consumption in the efficiency feature matrix, respectively.
[0070] Optional, feature matrix construction process: First, collect quality feature data, including protein content measurements Q at n time points. i1 (Obtained by high performance liquid chromatography), purity value Q i2 (Obtained by gel filtration chromatography), bioactivity assay value Q i3 (Obtained by enzyme activity assay), aggregation rate value Q i4 (Obtained via dynamic light scattering method), degradation rate measurement value Q i5 (Obtained via protein quantification); then efficiency characteristic data, including yield measurements E, were collected. i1 (Obtained by gravimetric method), Conversion rate determination value E i2 (Obtained through material balance algorithm), expression efficiency measurement value E i3(Obtained via quantitative fluorescence method), cell density measurement value E i4 (Obtained by cell counting method), glucose consumption measurement value E i5 (Obtained via glucose oxidase method); these data are arranged in chronological order to form an n×10 feature matrix M.
[0071] The specific equation for calculating the process gain index is as follows:
[0072]
[0073] In the formula, G is the process gain index; Q i For quality characteristics; E i For efficiency characteristics; w qi For quality weighting coefficients; w ei α is the efficiency weighting coefficient; β are the feature type weights, and α + β = 1; γ is the time decay coefficient; λ is the time constant; t is the fermentation time; δ is the fluctuation coefficient; x i These are real-time values of the process parameters; is the average value of the process parameters; n is the number of sampling points.
[0074] Derivation of the process gain exponential equation: First, consider the weighted sum of quality characteristics. Weighting coefficient w qi The analytic hierarchy process (AHP) is used to determine this; then, a weighted sum of efficiency features is considered. Weighting coefficient w ei Determined through expert scoring; a time decay term e is introduced. -λt The stability of the parameters over time is reflected, where λ is obtained by fitting historical data; finally, a volatility term is introduced. It reflects the degree of parameter fluctuation. The process gain exponential equation comprehensively considers the characteristics of quality efficiency, time decay effect, and process variability. An exponential function is used to describe time decay because the influence of process parameters decreases exponentially over time.
[0075] The specific equation for calculating the process compensation index is as follows:
[0076]
[0077] In the formula, C is the compensation index; G is the current process gain index; σ is the historical average gain exponent; θ is the standard deviation coefficient; φ is the time series coefficient; i G is the autoregressive coefficient; p is the autoregressive order; t-i Historical process gain index; ω j Seasonal weighting; S j η is the seasonality factor; s is the seasonal cycle; η is the process drift coefficient; t is the time variable.
[0078] The process compensation index equation combines statistical standardization, time series analysis, and trend analysis methods. An autoregressive term is introduced to capture the time correlation of process parameters. The derivation process of the process compensation index equation is as follows: First, a standardization process is used. Eliminate the influence of dimensions; then introduce time series terms. Capturing the time correlation of parameters, autoregressive coefficient φ i Obtained through maximum likelihood estimation; seasonality term introduced. To handle periodic changes, seasonal weight ω j It is obtained through time series decomposition; finally, the process drift term ηt is added to reflect the long-term trend.
[0079] The specific equation for calculating the process optimization index is as follows:
[0080]
[0081] In the formula, O is the optimization index; G1 and G2 are the process gain indices before and after parameter adjustment, respectively; μ is the rate of change weight; v is the time response coefficient; τ is the time constant; t is the response time; and ρ is the sensitivity weight. For parameter x i Partial derivatives; Δx i ξ is the parameter adjustment amount; m is the number of parameters; ξ is the system inertia coefficient; k is the attenuation constant; ψ is the environmental impact coefficient; and F is the environmental factor.
[0082] Derivation of the process optimization exponential equation: First, introduce the rate of change term. Evaluate the effect of parameter adjustment; then add the time response term v(1-e) -t / τ Describe the dynamic characteristics of the system; introduce parameter sensitivity terms. Assess the impact of parameter adjustments; finally, consider the system inertia term ξe. -kt And the environmental impact item ψF.
[0083] The process optimization exponential equation integrates rate of change assessment, dynamic response characteristics, and parameter sensitivity analysis, and uses partial derivative terms to evaluate the sensitivity of parameter adjustments; the specific equation for calculating the process decision function is as follows:
[0084] D=κ1f(C)+κ2g(O)+κ3h(C,O);
[0085] In the formula, D is the output value of the decision function; f(C) is the membership function of the compensation index; g(O) is the membership function of the optimization index; h(C,O) is the interaction function; κ1, κ2, and κ3 are weight coefficients, and κ1+κ2+κ3=1; the membership function adopts the trapezoidal form:
[0086]
[0087] In the formula, a, b, c, and d are the inflection point parameters of the membership function.
[0088] The autoencoder function of the first hidden layer of a deep neural network is specifically represented as follows:
[0089] h = σ(W1x + b1);
[0090]
[0091] In the formula, h represents the output of the hidden layer; For reconstructing the output; x is the input data; W1 and W2 are weight matrices; b1 and b2 are bias vectors; σ is the activation function; L is the loss function; n is the number of samples; γ is the regularization coefficient; ||W|| F Let f be the Frobenius norm of the weight matrix.
[0092] The backpropagation function for subsequent hidden layers is expressed as follows:
[0093]
[0094] In the formula, ΔW l η is the weight update value of layer l; η is the learning rate; E is the error function; α is the momentum factor; ΔW l (t-1) y represents the weight update amount from the previous iteration; i This is the actual output; is the predicted output; β is the regularization coefficient; R(W) is the regularization term; rank(X) is the standard rank of the process feature matrix; L is the number of network layers.
[0095] The standard rank feature extraction function is expressed as follows:
[0096]
[0097] In the formula, r is the standard rank; X is the process feature matrix; m and n are the matrix dimensions; λ is the singular value weight; σ i is the singular value; k is the number of singular values selected; μ is the time entropy weight; H(t) is the time series entropy; the time series entropy is calculated using the following equation:
[0098]
[0099] In the formula, p i This represents the probability distribution of the time series. Introducing the standard rank of process features as a regularization term into the loss function of deep neural networks can improve the model's ability to recognize process features.
[0100] Specifically, the principle of this invention is based on an innovative combination of deep learning and process control theory. First, by constructing a feature matrix, various factors affecting yield during fermentation are systematically characterized. The quality feature matrix includes indicators such as protein content and purity, while the efficiency feature matrix includes parameters such as yield and conversion rate. This comprehensive data representation lays the foundation for precise control.
[0101] Deep neural network models are the core of achieving stable production control. This model, through a deep architecture with five hidden layers, is able to learn the complex nonlinear relationship between process parameters and production output. The first layer uses an autoencoder for feature extraction, improving the effectiveness of data representation; the latter four layers optimize weights through backpropagation and introduce the standard rank of process control features as a constraint, enhancing the model's ability to predict production fluctuations.
[0102] The process judgment system achieves precise yield control through three levels of indicators. The process gain index comprehensively evaluates fermentation performance and promptly identifies key factors affecting yield; the compensation index dynamically compensates for batch variations by considering historical data; and the optimization index assesses the impact of parameter adjustments on yield and guides optimization directions. This multi-level evaluation system ensures the accuracy and reliability of yield control.
[0103] The following provides a specific embodiment 1 of the present invention, and the specific implementation of each step in this embodiment 1 is described in detail below.
[0104] The specific implementation of step S01 involves using high-throughput culture medium screening technology. Small-scale culture experiments were conducted using 96-well plates to screen culture media. First, 96 different culture media compositions were prepared, including different concentrations of carbon source glucose (10-30 g / L), nitrogen source yeast extract (2-6 g / L), phosphorus source KH2PO4 (0.5-2 g / L), and trace element mixture (0.1-0.5 g / L). Orthogonal experimental design was used to construct the culture medium formulation combinations. Cell density was detected by flow cytometry, protein expression was detected using the Bradford protein assay, and cell viability was detected using trypan blue staining. The screened culture medium composition was optimized using principal component analysis. An optimal culture medium formulation was determined by establishing a response surface model between culture medium components and cell growth indicators. After three rounds of iterative optimization, initial culture conditions were finally obtained with a cell density of 1,000,000-2,000,000 cells / mL, a protein expression level greater than 500 mg / L, and cell viability greater than 90%.
[0105] The specific implementation of step S02 involves inoculating the screened algal cells into a 5-liter fermenter for scale-up culture, with the inoculation density controlled at 50,000 to 100,000 cells per milliliter; culturing is carried out in a bioreactor equipped with an intelligent control system, providing continuous initial light intensity of 5,000 to 8,000 lux through an adjustable LED light source, controlling the initial temperature within the range of 22 to 26°C through a jacketed circulating water system, controlling the initial stirring speed within the range of 100 to 200 rpm through a variable frequency motor, and controlling the initial aeration rate within the range of 0.5 to 1.0 liters per minute through a mass flow meter; a mathematical model between the culture parameters and the target product yield is constructed using response surface methodology, and the optimal combination of initial parameters is obtained through central composite design.
[0106] The specific implementation of step S03 involves using an online monitoring system to collect real-time parameters of the culture process, including light intensity, temperature, stirring speed, aeration rate, cell density, protein expression level, and cell viability. Specifically, light intensity is collected every 5 minutes using a light intensity sensor; temperature is collected every 1 minute using a PT100 temperature sensor; stirring speed is collected every 1 minute using a Hall effect speed sensor; aeration rate is collected every 1 minute using a thermal gas flow meter; cell density is collected every 30 minutes using an online turbidimeter; protein expression level is collected every 30 minutes using an online fluorescence detector; and cell viability is collected every 60 minutes using an online activity detector. The collected data is transmitted to the central control system via a 14-bit analog-to-digital converter data acquisition card. Outliers are removed using a sliding median filtering algorithm, and signal noise is processed using a wavelet transform denoising algorithm.
[0107] The specific implementation of step S04 is to construct a feature matrix, using the following mathematical expression:
[0108]
[0109] The quality characteristic matrix includes five indicators, and the real-time protein content Q is determined by high-performance liquid chromatography. i1 A C18 reversed-phase column was used, with a mobile phase of acetonitrile and water at a flow rate of 1 mL / min; the real-time purity Q was determined by gel filtration chromatography. i2 The chromatographic column was TSK G3000SWXL, the mobile phase was phosphate buffer, and the flow rate was 0.5 mL / min. Real-time biological activity Q was determined by enzyme activity assay. i3 The absorbance was measured at the maximum absorption wavelength using a spectrophotometer; the real-time aggregation rate Q was determined using the dynamic light scattering method. i4 Particle size distribution was determined using a light scattering instrument at 25℃; real-time degradation rate Q was determined using a protein quantification method. i5Total protein content was determined using the BCA method, and the degradation percentage was calculated. The efficiency characteristic matrix included five indicators, and real-time yield E was determined by gravimetric method. i1 A precision balance is used for weighing; the real-time conversion rate E is determined using a material balance algorithm. i2 The efficiency of substrate-to-product conversion was calculated; the real-time expression efficiency E was determined by quantitative fluorescence method. i3 The fluorescence intensity of the target protein was measured using a fluorescence spectrophotometer; the real-time cell density E was determined by cell counting. i4 The blood cell count was performed using a hemocytometer; the real-time glucose consumption E was determined by the glucose oxidase method. i5 The glucose was measured using a glucose assay kit.
[0110] The specific implementation of step S05 involves importing a pre-trained deep neural network model, constructing a 5-layer hidden neural network using a deep learning framework, receiving 10 feature parameters in the input layer, and performing feature extraction and pattern recognition through 64 to 128 neurons in each layer. Each hidden layer uses a modified linear unit activation function to prevent gradient vanishing; the first hidden layer is pre-trained using an autoencoder function.
[0111] h = σ(W1x + b1);
[0112]
[0113] The last four hidden layers are trained using the backpropagation function:
[0114]
[0115] The model parameters were optimized using stochastic gradient descent with a learning rate of 0.001, a batch size of 64, and 1000 training epochs. The model performance was evaluated using 10-fold cross-validation with a validation set ratio of 20% and a test set ratio of 10%. The root mean square error and coefficient of determination were used as evaluation metrics.
[0116] The specific implementation of step S06 involves using a programmable controller to control the fermentation system in real time based on the optimized culture parameter set output by the deep neural network model. Light intensity is adjusted using pulse width modulation (PWM) technology with an output frequency of 200 Hz. The adjustment is achieved by changing the duty cycle of the LED light source, resulting in continuous adjustment within the range of 1000 to 20000 lux with an adjustment accuracy of 100 lux. Temperature is adjusted using a proportional-integral-derivative (PID) control algorithm with a proportional coefficient of 2.0, an integral time of 120 seconds, and a derivative time of 30 seconds. Precise control within the range of 15 to 35°C is achieved by adjusting the opening of the circulating water valve, with a control accuracy of 0.1°C. The stirring speed is adjusted using variable frequency control technology, achieving dynamic adjustment within the range of 50 to 500 revolutions per minute by changing the motor speed, with a control accuracy of 1 revolution per minute. Aeration is adjusted using mass flow control technology, achieving precise control within the range of 0.1 to 5.0 liters per minute by adjusting the gas flow valve opening, with a control accuracy of 0.1 liters per minute.
[0117] The specific implementation of step S07 involves calculating the process gain index using the following comprehensive evaluation equation:
[0118]
[0119] The quality weighting coefficients were determined using the analytic hierarchy process (AHP), including a weight of 0.3 for protein content, 0.2 for purity, 0.2 for bioactivity, 0.15 for aggregation rate, and 0.15 for degradation rate. The efficiency weighting coefficients were determined using an expert scoring method, including a weight of 0.3 for yield, 0.2 for conversion rate, 0.2 for expression efficiency, 0.15 for cell density, and 0.15 for glucose consumption. The time decay coefficient was calculated using an exponential decay function with an initial value of 1.0 and a decay constant of 0.01 per hour. The process fluctuation coefficient was calculated using the coefficient of variation with a threshold set at 0.2.
[0120] The specific implementation of step S08 is to calculate the compensation index using the following deviation correction equation:
[0121]
[0122] The deviation between the current process gain index and the historical average gain index is converted into a compensation index through standardization, with the standard deviation coefficient ranging from 0.1 to 0.3; the time series coefficient is calculated using the exponential smoothing method, with the smoothing coefficient ranging from 0.1 to 0.3; the seasonality factor is calculated using the seasonal decomposition method, with a decomposition order of 12; and the process drift coefficient is calculated using the linear regression method, with a confidence interval of 95%.
[0123] The specific implementation of step S09 involves calculating the optimization index using the following effect evaluation equation:
[0124]
[0125] The change in the process gain index before and after parameter adjustment was obtained by differential calculation. The time response coefficient was calculated using a first-order lag model, with a time constant ranging from 0.1 to 1.0 hours. The parameter sensitivity coefficient was calculated using sensitivity analysis, with a sensitivity threshold of 0.1. The system inertia coefficient was calculated using the step response method, with an inertia time constant ranging from 0.5 to 2.0 hours. The environmental impact coefficient was calculated using multiple regression analysis, with a significance level of 0.05.
[0126] The specific implementation of step S10 is to construct a process decision function, using the following fuzzy logic control equation:
[0127] D=κ1f(C)+κ2g(O)+κ3h(C,O);
[0128] The compensation index and optimization index are fuzzified using a trapezoidal membership function:
[0129]
[0130] The weight coefficients κ1, κ2, and κ3 are determined by the fuzzy rule base and satisfy κ1+κ2+κ3=1; the process optimization suggestions are generated by the fuzzy inference system, including the direction and magnitude of parameter adjustment, and the inference rule base contains 27 rules; the qualified judgment result is defuzzified by the centroid method, and the judgment threshold is 0.8.
[0131] The specific implementation of step S11 involves adding the inducer at the optimal induction time point. The induction time point is determined by prediction using a deep neural network model, generally selected during the mid-to-late logarithmic growth phase of cells, when the cell density reaches 800,000 to 1,000,000 cells per milliliter. The choice of inducer is determined based on the cell physiological state and the characteristics of the target protein. For algal cells, isopropyl galactothioglycoside is commonly used as an inducer. The inducer is added using a stepwise addition strategy, with an initial concentration of 0.5 mmol / L. Protein expression levels are sampled and detected every 2 hours. When the rate of increase in expression decreases, 0.2 to 0.5 mmol / L of inducer is added, with the final concentration not exceeding 2 mmol / L. During the induction process, cell density, cell viability, dissolved oxygen levels, and other indicators are monitored every 30 minutes to ensure a stable induction process.
[0132] The specific implementation of step S12 involves monitoring real-time protein expression levels using an online fluorescence detection system every 30 minutes. The fermentation status is assessed using a process judgment function. When the judgment result is satisfactory, the fermentation supernatant is collected; specifically, the fermentation process is stopped when the optimization index is greater than 0.9 and the compensation index is less than 0.2. The fermentation supernatant is collected using a sterile sampling system, with each sample volume being 50 ml. The sampling tubing is made of 316L stainless steel. The system is kept sealed during sampling to prevent external contamination. The collected supernatant is immediately pre-cooled at 4°C, and subsequent separation and purification operations are completed within 2 hours. After each batch of fermentation, the supernatant is washed with 85°C warm water for 20 minutes, then disinfected with 0.5% peracetic acid solution for 15 minutes, and finally rinsed three times with sterile water.
[0133] The specific implementation of step S13 involves processing the fermentation supernatant using centrifugation. A large-capacity refrigerated centrifuge with a rotor radius of 25 cm is used for solid-liquid separation. The centrifugation process employs a step-by-step acceleration strategy: the initial speed is 1000 rpm, maintained for 2 minutes to allow the cells to gradually settle; then, the speed is increased to 3000 rpm and maintained for 5 minutes for preliminary separation; finally, the speed is increased to 4000 rpm and maintained for 15 minutes to complete the separation. The centrifugation temperature is strictly controlled at 4°C, and the centrifuge chamber is pre-cooled for 30 minutes to ensure temperature stability. The centrifuge tubes are made of polypropylene, with a single tube capacity of 500 ml and a maximum load capacity of 1000 g. The number of centrifuge tubes is determined based on the volume of the fermentation supernatant. During centrifugation, the liquid volume error of each centrifuge tube is controlled within 1 gram to ensure rotor balance.
[0134] The specific implementation of step S14 involves membrane filtration of the centrifuged fermentation supernatant using a multi-stage cascade filtration system. The first stage uses a pre-filtration membrane made of polypropylene with a 5-micron pore size and a filtration area of 0.1 square meters to remove residual cell debris and large particulate impurities. The second stage uses an intermediate filtration membrane made of polyethersulfone with a 0.45-micron pore size and a filtration area of 0.2 square meters to remove small particulate matter and some large molecular impurities. The third stage uses a terminal filtration membrane made of polyvinylidene fluoride with a 0.22-micron pore size and a filtration area of 0.5 square meters to ensure the sterility of the product. The filtration process uses a constant pressure filtration mode, with the feed pressure controlled between 0.1 and 0.2 MPa and the filtration temperature controlled at 4°C. The filtration system is equipped with a differential pressure monitoring device, and the appropriate level of filtration membrane is replaced when the differential pressure exceeds 0.05 MPa. The filtered product is collected in sterilized glass bottles, protected with nitrogen, sealed, and refrigerated at 0 to 4°C. Product quality testing includes indicators such as protein content, purity, bioactivity, and sterility. After passing the tests, the product proceeds to the next purification step.
[0135] The establishment of the deep neural network model training dataset involved collecting data from 10,000 historical batches of recombinant protein fermentation processes, with a collection frequency of once every 5 minutes. Each batch of data contained 16 process parameters. The data cleaning process employed the 3σ criterion for outlier detection and deletion, the 5-point moving average method to handle data noise, and the cubic spline interpolation algorithm to repair missing data. Data standardization used the min-max standardization method to map all feature values to the interval between 0 and 1. Data segmentation employed a stratified sampling method, dividing the training set, validation set, and test set in an 8:1:1 ratio. Data augmentation used the sliding window method, with a window size of 12 hours and a sliding step size of 1 hour.
[0136] The deep neural network model training process first employs an autoencoder to pre-train the first hidden layer, with a 16-dimensional feature vector as input. The dimensionality reduction coefficient is set to 0.5, the reconstruction error threshold to 0.01, the learning rate to 0.001, and the maximum number of iterations to 1000. The subsequent four hidden layers are trained using the backpropagation algorithm, with an initial learning rate of 0.01, decaying by 20% every 50 epochs, a momentum factor of 0.9, and an error threshold of 0.001. Ten-fold cross-validation is used to evaluate model performance. A grid search method is employed to optimize hyperparameters. Multiple models are integrated using a voting method. During training, strategies such as batch normalization, random deactivation, residual connection structures, and exponential moving averages are introduced to improve model performance.
[0137] The algal cell recombinant protein expression optimization method of this invention is particularly suitable for the expression and production of immune system-related proteins, including T cell surface markers (such as CD3, CD4, CD8, CD28), NK cell-related molecules (such as CD16, CD56), B cell markers (such as CD19), immunomodulatory molecules (such as PD-1, CD25, CD127), complement regulatory proteins (such as CD55, CD59, CD35), phagocytic cell-related molecules (such as CD64, CD11b), and cytotoxicity-related proteins (such as perforin). These immune-related proteins have important applications in clinical diagnostics, immunotherapy, and biopharmaceuticals. Many of them are targets of monoclonal antibodies or immune checkpoint molecules, requiring high stability in their expression levels. The optimization method of this invention can significantly improve the expression stability and yield consistency of these functional proteins, providing reliable process assurance for the large-scale production of immunotherapy products. Applicable algae include one or more cell lines from the following species: *Chlamydomonas reinhardtii*, *Chlorella* sp., *Dunaliella salina*, *Prorocentrum minimum*, *Alexandrium* sp., *Platymonass* sp., *Scenedesmus* sp., *Euglena* sp., *Porphyridium* sp., and *Nephroselmis* sp., with *Chlamydomonas reinhardtii* cell lines being preferred. These algal cell lines have undergone long-term laboratory domestication and genetic modification, establishing stable ribosome entry site sequences, chloroplast targeting sequences, and endoplasmic reticulum localization sequences, enabling efficient expression of the aforementioned human cell surface marker antibodies.
[0138] To better understand and implement this invention, Example 2 of a specific application scenario is provided below: In a PD-1 recombinant protein expression development project, researchers used the algal cell secretory recombinant protein expression optimization method of this invention for process optimization. First, algal cell samples were screened for culture media. Ninety-six different culture media formulations were designed for orthogonal experiments. By detecting indicators such as cell density, protein expression level, and cell viability, the optimal culture media formulation was obtained, as shown in Table 1.
[0139] Table 1. Optimal Culture Medium Formulation
[0140] Element Concentration (grams per liter) glucose 25.0 Yeast paste 4.5 <![CDATA[KH2PO4]]> 1.5 <![CDATA[MgSO4·7H2O]]> 0.5 <![CDATA[CaCl2]]> 0.2 Trace element mixture 0.3
[0141] Algal cells were cultured using this culture medium formulation, achieving a cell density of 1,568,000 cells / mL, a protein expression level of 682 mg / L, and a cell viability of 93.5%. The selected algal cells were then inoculated into a 5-liter fermenter for scale-up culture at a density of 75,000 cells / mL. The initial culture parameters are shown in Table 2.
[0142] Table 2 Initial Culture Parameter Settings
[0143] parameter numerical values Light intensity (lux) 6500 Temperature (°C) 24 Stirring speed (revolutions per minute) 150 Ventilation rate (liters per minute) 0.8 pH 7.2 Dissolved oxygen (%) 40
[0144] During the cultivation process, an online monitoring system was used to collect cultivation parameters in real time and construct a feature matrix. Partial data of the quality feature matrix and efficiency feature matrix are shown in Table 3.
[0145] Table 3. Partial Data Table of the Characteristic Matrix of the Cultivation Process
[0146]
[0147] Figure 2 The dynamic changes in PD-1 protein expression, cell density, and glucose concentration during culture are shown in the figure. As can be seen from the figure, PD-1 protein expression exhibits a typical S-shaped growth curve, reaching a rapid expression phase around 85.5 hours of culture; cell density also shows a similar growth trend, approaching 1 million cells per milliliter at the induction time point; glucose concentration shows an exponential decay trend, exhibiting a significant negative correlation with cell growth and protein expression. A pre-trained deep neural network model was imported, trained on 10,000 batches of historical data. The model performance evaluation results are shown in Table 4.
[0148] Table 4. Model Performance Evaluation Results
[0149] Evaluation indicators training set Validation set test set Root mean square error 0.052 0.068 0.072 Coefficient of determination 0.946 0.925 0.918 Mean Absolute Error 0.043 0.056 0.062 Accuracy (%) 94.5 92.8 91.6
[0150] The fermentation system is controlled in real time based on the optimized culture parameter set output by the model. The process gain index is calculated using the following equation: The weighting coefficients are shown in Table 5:
[0151] Table 5 Feature Weight Coefficients
[0152]
[0153]
[0154] At 85.5 hours of culture, the cell density reached 986,000 cells / mL, which was the optimal induction time predicted by the model. Isopropyl galactothioglycoside was added as an inducer at an initial concentration of 0.5 mmol / L. Protein expression was monitored every 2 hours thereafter, and the inducer was replenished based on changes in expression levels, with a final concentration of 1.8 mmol / L. At 168 hours of culture, the process judgment function output a pass / fail result, with an optimization index of 0.925 and a compensation index of 0.156. The fermentation process was then stopped, and the supernatant was collected. Figure 3 The changes in the bioactivity, purity, and aggregation rate of PD-1 protein during culture are shown in the figure. As can be seen from the figure, the bioactivity and purity of the protein both increased with the extension of culture time, eventually reaching 92.8% and 95.6%, respectively; while the aggregation rate increased slightly in the early stage of culture, but stabilized in the later stage, eventually remaining at a low level of 6.5%. The fermentation supernatant was treated by centrifugation and membrane filtration, and the quality indicators of the final PD-1 recombinant protein product are shown in Table 6.
[0155] Table 6 Product Quality Indicators
[0156] index numerical values Protein content (mg / L) 856 purity(%) 95.6 Bioactivity (%) 92.8 Aggregation rate (%) 6.5 Degradation rate (%) 4.2 Endotoxin (EU per milligram) 0.08 Aseptic testing qualified
[0157] Traditional PD-1 recombinant protein expression methods primarily utilize mammalian cell expression systems, which suffer from high culture costs, long culture cycles, and low yields. Furthermore, process parameter optimization relies heavily on empirical judgment and single-factor experiments, resulting in low optimization efficiency. This invention employs an algal cell expression system combined with deep learning algorithms to achieve intelligent optimization of process parameters, achieving significant improvements in the following aspects: 1. Expression level increased from 300-400 mg / L in traditional methods to 856 mg / L, an increase of approximately 100%; 2. Culture cycle shortened from 12-14 days in traditional methods to 7 days, an efficiency improvement of approximately 45%; 3. Culture cost reduced by approximately 60%; 4. Product quality became more stable, with the batch-to-batch variation coefficient decreasing from 0.25 to 0.08; 5. Process parameter optimization efficiency improved by approximately 300%, and the optimization results are more reliable.
[0158] It should be noted that the variables involved in this invention are explained in detail in Table 7 below.
[0159] Table 7 Variable Explanation Table
[0160]
[0161]
[0162] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for optimizing expression of a secreted recombinant protein in algal cells, comprising, The method comprises the following steps: Collecting algal cell samples and performing medium screening to obtain initial culture conditions with a cell density of 1,000,000 to 2,000,000 cells per milliliter, a protein expression amount of greater than 500 milligrams per liter, and a cell activity of greater than 90%; inoculating the screened algal cells in a culture medium, recording and setting initial culture parameters; collecting culture process parameters; Constructing a feature matrix composed of a quality feature matrix and an efficiency feature matrix; importing a pre-trained deep neural network model, processing through 5 hidden layers, each containing 64 to 128 neurons, and outputting an optimized culture parameter set and an optimal induction time point; adjusting the culture parameters according to the optimized culture parameter set; calculating a process gain index, a compensation index, and an optimization index; constructing a process determination function; adding an inducer at the optimal induction time point; Monitoring real-time protein expression amount, and collecting fermentation supernatant when the process determination function outputs a qualified determination result; Processing the fermentation supernatant by centrifugation; performing membrane filtration on the centrifuged fermentation supernatant to obtain an optimized recombinant protein product.
2. The algal cell-expressing recombinant protein secretion optimization method according to claim 1, wherein, The culture process parameters include real-time light intensity, real-time temperature, real-time stirring speed, real-time aeration rate, real-time cell density, real-time protein expression amount, and real-time cell activity; the initial light intensity is 5,000 to 8,000 lux, the initial temperature is 22 to 26°C, the initial stirring speed is 100 to 200 revolutions per minute, and the initial aeration rate is 0.5 to 1.0 liters per minute.
3. The algal cell-expressing recombinant protein secretion optimization method according to claim 2, wherein, The quality feature matrix records real-time protein content, real-time purity, real-time biological activity, real-time aggregation rate, and real-time degradation rate, and the efficiency feature matrix records real-time yield, real-time conversion rate, real-time expression efficiency, real-time cell density, and real-time glucose consumption amount.
4. The method for optimizing expression of a recombinant protein secreted by algal cells according to claim 3, wherein, The training data set of the pre-trained deep neural network model contains historical 10,000 batches of recombinant protein fermentation data, and the training data set construction steps include data acquisition, data cleaning, data standardization, data segmentation, and data enhancement; The training steps include model initialization, parameter optimization, model validation, model selection, and model integration.
5. The algal cell-expressing recombinant protein secretion optimization method of claim 4, wherein, The neuron weight parameters of the last 4 hidden layers in the 5 hidden layers are determined by a back propagation function, and the input of the back propagation function includes current hidden layer input data, target output data, learning rate parameter, momentum factor, error threshold, and neuron weight parameters of the previous hidden layer; the neuron weight parameters of the first hidden layer are determined by a self-encoder function.
6. The algal cell-expressing recombinant protein secretion optimization method of claim 5, wherein, The process gain index is calculated using a comprehensive evaluation equation, and the input of the comprehensive evaluation equation includes the quality feature matrix, the efficiency feature matrix, the quality weight coefficient, the efficiency weight coefficient, the time decay coefficient, and the process fluctuation coefficient, and the output is a quantitative index representing fermentation performance.
7. The algal cell-expressing recombinant protein secretion optimization method of claim 6, wherein, The process compensation index is calculated using a bias correction equation, and the input of the bias correction equation includes the current process gain index, the historical average gain index, the standard deviation coefficient, the time series coefficient, the seasonal factor, and the process drift coefficient, and the output is a correction value of batch-to-batch difference.
8. The algal cell-expressing recombinant protein secretion optimization method of claim 7, wherein, The process optimization index is calculated by an effect evaluation equation, wherein the input includes the process gain index before parameter adjustment, the process gain index after parameter adjustment, the time response coefficient, the parameter sensitivity coefficient, the system inertia coefficient and the environmental influence coefficient, and the output is the effect score of the process optimization.
9. The algal cell-expressing recombinant protein secretion optimization method of claim 8, wherein, The process determination function inputs the compensation index and the optimization index, uses the fuzzy logic control method to perform real-time evaluation and optimization on the fermentation process, converts into a low, medium and high three-level fuzzy set through fuzzy processing, uses a trapezoidal membership function, and outputs the process optimization suggestion and the qualified determination result.
10. The algal cell secretion optimized recombinant protein expression method of claim 9, wherein, The concentration of the inducer is 0.5 to 2 mmol / L; the centrifugal speed is 3000 to 5000 rpm, and the centrifugal time is 15 to 20 minutes; the membrane filtration uses a multi-stage series filtration system, including a 5-micron pre-filtration membrane, a 0.45-micron intermediate filtration membrane and a 0.22-micron terminal filtration membrane.