Microbial medium optimization method and system based on correlation coefficient significance analysis
Patent Information
- Application Number
- CN202310125434.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-17
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-02-17
AI Technical Summary
[0004]本发明的目的是提供一种基于相关系数显著性分析的微生物培养基优化方法及系统,可以解决常规PLSR回归方程异常“庞大和臃肿”以及简洁性、可解释性和预测泛化能力差的问题,同时计算过程简单,不需要过高的编程技术,适用性强
[0034]根据本发明提供的具体实施例,本发明公开了以下技术效果:本发明采用t检验法根据自变量与响应变量之间的皮尔逊相关系数矩阵分析自变量与响应变量之间的显著性水平,利用相关系数显著性分析可筛选显著性自变量,而不用繁琐的变量投影重要性技术,可以解决常规PLSR回归方程异常“庞大和臃肿”以及简洁性、可解释性和预测泛化能力差的问题,同时计算过程简单,不需要过高的编程计算技能即可便捷使用。
Smart Images

Figure CN116312817B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multivariate statistical technology, and in particular to a method and system for optimizing microbial culture media based on correlation coefficient significance analysis. Background Technology
[0002] In modern fermentation industry, culture medium optimization has always been one of the most important research topics. Optimization can not only effectively reduce fermentation costs but also improve the utilization rate of raw materials, the yield and quality of fermentation products, and reduce the difficulty of downstream processing in fermentation engineering. Existing culture medium optimization methods mainly rely on experimental design, with uniform design being a prominent example. This design only considers the "uniform distribution" of experimental points within the experimental range, without considering their "uniform comparability." Therefore, the number of experiments required for optimizing the same problem is significantly less than that of orthogonal experiments and central composite designs. Thus, using uniform design for culture medium optimization research can greatly save optimization costs and reduce optimization time. However, research shows that due to the limited number of experiments, uniform design suffers from typical "small sample size problems" and "multicollinearity" problems. Therefore, regression models based on the least squares method suffer from poor accuracy and low reliability.
[0003] Currently, the main methods for optimizing multiple regression modeling in "small sample" culture media include Partial Least Squares Regression (PLSR) and Support Vector Machines (SVM). Meanwhile, methods for addressing multicollinearity in multiple regression include PLSR, Ridge Regression, and LASSO Regression. Therefore, PLSR can simultaneously address both the "small sample" and "multicollinearity" problems faced by uniform design multiple regression modeling. Furthermore, studies have found that PLSR performs better than Ridge Regression in handling multicollinearity, primarily due to a smaller mean absolute percentage error and a larger multiple determination coefficient. PLSR also exhibits superior independent variable selection capabilities, significantly outperforming Ridge Regression and LASSO Regression. However, traditional PLSR models often tend to retain all modeling independent variables, resulting in an abnormally large and bloated model with poor simplicity and interpretability, and a degree of overfitting. Currently, Variable Importance in Projection (VIP) is commonly used to address this issue by removing redundant independent variables while retaining key ones. However, literature review found that the PLSR-VIP method has a relatively complicated calculation process and requires strong mathematical and computer programming skills, which to some extent limits its promotion and use. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for optimizing microbial culture media based on correlation coefficient significance analysis. This method can solve the problems of the abnormally large and cumbersome nature of conventional PLSR regression equations, as well as their poor simplicity, interpretability, and predictive generalization ability. At the same time, the calculation process is simple, does not require advanced programming skills, and has strong applicability.
[0005] To achieve the above objectives, the present invention provides the following solution:
[0006] A method for optimizing microbial culture media based on correlation coefficient significance analysis includes:
[0007] Multiple experimental samples were obtained by using a uniform design experiment with culture medium composition as the independent variable and experimental products as the response variable.
[0008] Based on each experimental sample, a partial least squares regression (PLSR) full model for each response variable is constructed using the PLSR method; the PLSR full model is a mathematical expression in quadratic polynomial form representing the relationship between the response variable and all the independent variables.
[0009] For any response variable, the Pearson correlation method is used to process the full PLSR model of the response variable to obtain the Pearson correlation coefficient matrix between the independent variables and the response variable;
[0010] The t-test method was used to obtain the p-value of the t-test between the independent variables and the response variable based on the Pearson correlation coefficient matrix between the independent variables and the response variable;
[0011] The independent variables are screened based on the t-test p-values of their respective independent variables and the response variable to obtain the significant independent variables corresponding to the response variable;
[0012] The PLSR selection model for the response variable is constructed using partial least squares method based on each experimental sample and each significant independent variable corresponding to the response variable; the PLSR selection model is a mathematical expression in quadratic polynomial form representing the relationship between the response variable and all the significant independent variables corresponding to the response variable.
[0013] Microbial culture media were constructed based on the PLSR selection model for all the aforementioned response variables.
[0014] Optionally, the t-test method is used to obtain the t-test p-value between each independent variable and the response variable based on the Pearson correlation coefficient matrix between the independent variables and the response variable, specifically as follows:
[0015] According to the formula Calculate the t-test t-value, where r represents the Pearson correlation coefficient matrix between the independent variable and the response variable, n represents the total number of experimental samples, and t represents the t-test t-value.
[0016] The tcdf function in the Matlab software package is used to calculate the t-test p-value based on the t-test t-value.
[0017] Optionally, the step of filtering the independent variables based on the t-test p-values of their respective variables and the response variable to obtain the significant independent variables corresponding to the response variable specifically includes:
[0018] For any independent variable, if the t-test p-value between the independent variable and the response variable is less than a set value, then the independent variable is determined to be a significant independent variable corresponding to the response variable.
[0019] Optionally, the culture medium components include: sodium glycerophosphate, (NH4)2SO4, CaSO4·2H2O and PTM1; the experimental products include culture medium biomass, xylanase activity and specific enzyme activity.
[0020] A microbial culture medium optimization system based on correlation coefficient significance analysis includes:
[0021] The experimental module is used to design a uniform experiment with culture medium composition as the independent variable and experimental product as the response variable to obtain multiple experimental samples.
[0022] The PLSR full model construction module is used to construct a PLSR full model for each of the experimental samples using partial least squares regression; the PLSR full model is a mathematical expression in quadratic polynomial form representing the relationship between the response variable and all the independent variables;
[0023] The Pearson module is used to process the full PLSR model of any response variable using the Pearson correlation method to obtain the Pearson correlation coefficient matrix between each independent variable and the response variable.
[0024] The t-test module is used to obtain the t-test p-value of each independent variable and the response variable based on the Pearson correlation coefficient matrix between the independent variables and the response variable using the t-test method.
[0025] The significant independent variable screening module is used to screen the independent variables based on the t-test p-values of the independent variables and the response variable to obtain the significant independent variables corresponding to the response variable;
[0026] The model selection module is used to construct a PLSR model for the response variable using partial least squares method based on each experimental sample and each significant independent variable corresponding to the response variable; the PLSR model is a mathematical expression in quadratic polynomial form representing the relationship between the response variable and all the significant independent variables corresponding to the response variable.
[0027] A microbial culture medium construction module is used to construct microbial culture media based on the PLSR selection model for all the aforementioned response variables.
[0028] Optionally, the t-test module specifically includes:
[0029] The t-test t-value calculation unit is used to calculate the t-value according to the formula. Calculate the t-test t-value, where r represents the Pearson correlation coefficient matrix between the independent variable and the response variable, n represents the total number of experimental samples, and t represents the t-test t-value.
[0030] The t-test p-value calculation unit is used to calculate the t-test p-value based on the t-test t-value using the tcdf function in the Matlab software package.
[0031] Optionally, the significant independent variable screening module specifically includes:
[0032] The significant independent variable screening unit is used to determine that, for any given independent variable, if the t-test p-value between the independent variable and the response variable is less than a set value, the independent variable is a significant independent variable corresponding to the response variable.
[0033] Optionally, the culture medium components include: sodium glycerophosphate, (NH4)2SO4, CaSO4·2H2O and PTM1; the experimental products include culture medium biomass, xylanase activity and specific enzyme activity.
[0034] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects: The present invention uses the t-test method to analyze the significance level between the independent variable and the response variable based on the Pearson correlation coefficient matrix between the independent variable and the response variable. The significance analysis of the correlation coefficient can be used to screen significant independent variables without the need for cumbersome variable projection importance techniques. This can solve the problems of the abnormal "large and bloated" nature of conventional PLSR regression equations, as well as their poor simplicity, interpretability, and predictive generalization ability. At the same time, the calculation process is simple and can be used conveniently without requiring excessive programming skills. Attached Figure Description
[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 A flowchart of a microbial culture medium optimization method based on correlation coefficient significance analysis provided in an embodiment of the present invention;
[0037] Figure 2 A flowchart illustrating the treatment of recombinant Pichia pastoris fermentation for xylanase production using the method provided in this embodiment of the invention and existing methods;
[0038] Figure 3 This is a comparison chart of predicted and experimental values of the full and selected PLSR models based on a single-objective genetic algorithm, provided in an embodiment of the present invention.
[0039] Figure 4 The following are evolutionary graphs of fitness functions for three optimization objectives in a full model and a single-objective genetic algorithm for a selected model, provided in embodiments of the present invention.
[0040] Figure 5 A comparison of Pareto front curves for full-model and model-selective dual-objective genetic algorithm optimization provided in this embodiment of the invention;
[0041] Figure 6 A comparison of the Pareto fronts for the full model and the selected model optimized by the three-objective genetic algorithm in the embodiments of the present invention. Detailed Implementation
[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0043] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0044] like Figure 1 As shown, this embodiment of the invention provides a method for optimizing microbial culture media based on correlation coefficient significance analysis, including:
[0045] Step 101: Using the culture medium composition as the independent variable and the experimental product as the response variable, a uniform design experiment was conducted to obtain multiple experimental samples.
[0046] Step 102: Construct a full PLSR model for each response variable using partial least squares regression based on each experimental sample; the full PLSR model is a mathematical expression in quadratic polynomial form representing the relationship between the response variable and all the independent variables.
[0047] Step 103: For any response variable, the Pearson correlation method is used to process the PLSR full model of the response variable to obtain the Pearson correlation coefficient matrix between the independent variables and the response variable.
[0048] Step 104: Using the t-test method, obtain the t-test p-values of the independent variables and the response variable based on the Pearson correlation coefficient matrix between the independent variables and the response variable.
[0049] Step 105: Based on the t-test p-values of the independent variables and the response variable, the independent variables are screened to obtain the significant independent variables corresponding to the response variable.
[0050] Step 106: Construct the PLSR selection model of the response variable using partial least squares method based on each experimental sample and each significant independent variable corresponding to the response variable; the PLSR selection model is a mathematical expression in quadratic polynomial form representing the relationship between the response variable and all the significant independent variables corresponding to the response variable.
[0051] Step 107: Construct a microbial culture medium based on the PLSR selection model for all the aforementioned response variables.
[0052] In practical applications, the t-test method is used to obtain the t-test p-value between the independent variables and the response variable based on the Pearson correlation coefficient matrix between the independent variables and the response variable. Specifically:
[0053] According to the formula Calculate the t-test t-value, where r represents the Pearson correlation coefficient matrix between the independent variable and the response variable, n represents the total number of experimental samples, and t represents the t-test t-value.
[0054] The tcdf function in the Matlab software package is used to calculate the t-test p-value based on the t-test t-value, where n-2 represents the degrees of freedom.
[0055] In practical applications, the step of filtering the independent variables based on the t-test p-values of their respective independent variables and the response variable to obtain the significant independent variables corresponding to the response variable specifically includes:
[0056] For any independent variable, if the t-test p-value between the independent variable and the response variable is less than a set value, then the independent variable is determined to be a significant independent variable corresponding to the response variable.
[0057] In practical applications, the culture medium components include: sodium glycerophosphate, (NH4)2SO4, CaSO4·2H2O and PTM1; the experimental products include culture medium biomass, xylanase activity and specific enzyme activity.
[0058] In practical applications, the value is set to 0.1.
[0059] This invention provides examples of improving xylanase production during recombinant Pichia pastoris fermentation using the above-described culture medium optimization method, and details the optimization effects of the above examples, such as... Figure 2 As shown, this embodiment integrates a series of strategies, namely, using leave-one-out cross-validation to extract latent variables and utilizing the coefficient of determination R. 2 To evaluate the goodness of fit of the model, the significance of the independent variables was assessed using a t-test based on the correlation coefficient r. Insignificant independent variables were removed, and a PLSR selection model was established. This successfully constructed a simple, robust, and highly accurate mathematical model. The specific technical solution is as follows:
[0060] 1. Experimental Materials and Measurement Methods
[0061] Strains: Recombinant Pichia pastoris GS115 / pPIC9K-xynZF-2, FMG medium as the initial medium, culture method: 30℃ shake flask culture (250mL shake flask, 25mL liquid volume, 250rpm). After 5 days of culture, the cell biomass was determined by turbidimetric assay, the xylanase activity (U / mL) in the culture medium was determined by DNS turbidimetric assay, and the total protein content (mg / mL) in the culture medium was determined by Bradford assay. Then, the xylanase specific enzyme activity (U / mg) was obtained by formula: xylanase enzyme activity (U / mL) / total protein content (mg / mL).
[0062] 2. Uniform Design Experiment
[0063] Multiple experimental samples were obtained using a uniform design with culture medium composition as the independent variable and experimental products as the response variable. Based on the previous Plackett–Burman experiment, the contents of some non-critical culture medium components were first specified: K₂SO₄ 14.0 g / L, MgSO₄·7H₂O 11.0 g / L, methanol 10.0 mL / L, and Tween 80 2.0 g / L. The initial pH was set at 5.5 (adjusted using 25% (v / v) NH₃·H₂O). For the four critical components CaSO₄·2H₂O, (NH₄)₂SO₄, PTM₁ (trace salt solution), and sodium glycerophosphate, the uniform design was used to determine the appropriate concentrations. The experiment was designed using a uniform design with four culture medium components (independent variables): sodium glycerophosphate, (NH4)2SO4, CaSO4·2H2O, and PTM1 (trace salt solution). A Pichia pastoris culture experiment was required. The experimental levels of the independent variables are shown in Table 1, and the experimental results are shown in Table 2. The biomass (OD) of the culture medium was also measured. 600 The xylanase activity and specific enzyme activity were used as response variables. A quadratic polynomial (Formula 1) was then used to characterize the mathematical relationship between the homogeneous design ratio of the culture medium and the response variables.
[0064]
[0065] Where β0 is the intercept term, β i The coefficient of the linear term, β ii It is the coefficient of the squared term, β ij X is the cross term coefficient (i≠j), ε is the model error term, i.e., the experimental error, and y is the symbol for the response variable. i and X j These are all independent variables, that is, the components of the culture medium.
[0066] Table 1 Uniform Design Independent variable experimental level table
[0067]
[0068] a X1, X2, X3, and X4 represent CaSO4·2H2O, (NH4)2SO4, PTM1 (trace salt solution), and sodium glycerophosphate, respectively.
[0069] Table 2 Uniform Design Experimental Results Table
[0070]
[0071]
[0072] a The culture medium components represented by X1, X2, X3 and X4 are the same as those in Table 1.
[0073] Table 1 shows that the uniform design of the culture medium ratio reconstructed the cell metabolic network, resulting in changes in yeast biomass, xylanase activity, and specific enzyme activity. Experiment 2 showed the highest xylanase activity, reaching 4144.658±118.647 U / mL, which is 139.09% of the lowest value (Experiment 11, 2979.793±387.439 U / mL). This indicates that the culture medium optimization was initially successful and can be further optimized through modeling. However, if formula (1) is used for least squares-based quadratic polynomial regression modeling, at least 1+4×(4+4) / 2=17 experiments are required, while the current number of experiments is only 12, which is a small sample problem. Therefore, only partial least squares can be used, and the least squares-based method cannot be used to establish the regression model.
[0074] 3. Correlation analysis of variables and its significance test
[0075] Pearson correlation coefficient analysis was performed on a uniform design dataset, and its significance was analyzed using a t-test. The results are shown in Table 3.
[0076] Table 3 shows that there are multiple multicollinearities among the independent variables, some of which are even highly significant, such as X1 and... Between, X2 and The correlation coefficients between the terms were all 0.973, and all were significant at the 0.01 level (p<0.01). Similarly, the correlations between X1 and X1X2, X1X3, and X1X4 were also significant (p<0.1). Even interaction terms without identical original independent variables, such as X1X4 and X2X3, were significantly correlated (p<0.1). This demonstrates that the uniform design regression modeling involved in this invention requires the use of partial least squares methods, rather than least squares-based methods, otherwise the model's robustness and predictive ability will be very poor.
[0077] Table 3. Pearson correlation coefficients and significance t-test results for the uniform design dataset.
[0078]
[0079] a The meanings of X1, X2, X3 and X4 are the same as in Table 1; b The meanings of Y1, Y2, and Y3 are the same as in Table 2.
[0080] "*” The correlation coefficient is significant at the 0.1 level.
[0081] "**” The correlation coefficient is significant at the 0.05 level.
[0082] "***” The correlation coefficient is significant at the 0.01 level.
[0083] 4. Establishment of the complete PLSR model
[0084] Based on the uniform design experiment, partial least squares regression (PLSR) was used to construct a complete PLSR model for each response variable (the target of culture medium optimization), and the coefficient of determination R (R²) of the complete PLSR model for each response variable, which measures the goodness of fit of the regression, was calculated. 2 The results are shown in Table 4. Specifically, latent variables were extracted using formula (2) based on leave-one-out cross-validation, and a full PLSR model for each response variable was established. The relevant calculation formulas are shown below:
[0085]
[0086] y i This is the i-th observation of the response variable y; an observation means an experimental value. Here, n is the number of samples in the model. PRESS is the fitted value of the response variable y after excluding the i-th sample and using the remaining samples to extract h latent variables for regression modeling, i.e., the theoretical value obtained using PLSR regression analysis (according to formula 1). h It is the sum of squared prediction errors of the "leave-one-out" method for the response variable y. It is the i-th fitted value of the response variable y in PLSR modeling based on h-1 latent variables extracted from all samples, SS h-1 It is the corresponding sum of squared errors. According to the basic criteria of cross-validation, when Adding new latent variables helps reduce model prediction error, and vice versa. After determining the optimal number of latent variables to extract, a full PLSR model (quadratic polynomial regression equation) including all independent variables is constructed, and the coefficient of determination R is calculated. 2 .
[0087] Table 4. Regression modeling results of the full PLSR model and the selected model based on leave-one-out cross-validation.
[0088]
[0089]
[0090] a The full model refers to a PLSR model that includes all types of independent variables (linear, quadratic, and interaction terms), while the selected model refers to a PLSR model obtained by regressing the response variable against key independent variables that are significantly correlated with the response variable. b LV S Indicate latent variables / principal components; c PRESS indicates the predicted sum of squared residuals; d SS represents the sum of squared errors; R represents the cross-validation determination coefficient. 2 .
[0091] As can be seen from Table 4, for all three response variables, only one principal component needs to be extracted for full PLSR model modeling. If all values are greater than 0.0975, the basic criteria for cross-validation are met. If one more latent variable is extracted... If the value becomes negative, it will lead to overfitting and impair the model's predictive ability. After reasonably extracting one latent variable, the PLSR full model will be successfully established, as shown in Table 5. It is easy to see that all four response variables produce fairly good regression models, all with fairly good coefficients of determination R. 2 The values were 0.877, 0.9685, and 0.9822, respectively, meaning that only 12.3%, 3.15%, and 1.78% of the total variation could not be explained by the full model.
[0092] Table 5. Regression Equations for the Full PLSR Model and Selected Model
[0093]
[0094]
[0095] a The meanings of full model and selected model are the same as in Table 3.
[0096] Furthermore, a scatter plot is created using the predicted values (x-axis) of the optimization objective determined by the full model against the corresponding experimental values (y-axis), as shown below. Figure 3 As shown, the scatter plots ("△") are closely spaced near the diagonal (i.e., the zero error bar, where predicted values equal experimental values), and the relative errors of all predicted values are within ±5%. These results indicate that the full model has excellent fitting accuracy and can be used to describe the quantitative relationship between culture medium components and optimization objectives.
[0097] 5. Establishment of PLSR selection model based on correlation coefficient significance analysis
[0098] Pearson correlation analysis was used to process each PLSR full model, calculating the Pearson correlation coefficient matrices between all independent variables, between the response variable, and between the independent and response variables in each PLSR full model. Then, a t-test was used to test the significance of the Pearson correlation coefficient matrices. If the significance level p between the independent and response variables was less than or equal to a set value, the independent variable was selected as a significant independent variable. Using the significant independent variable as the regression variable, PLSR regression modeling was performed again on the response variable to obtain the selected PLSR model for the response variable, and the R-squared value of the selected PLSR model for the response variable was calculated. 2 Since all PLSR full models contain independent variables that are not significantly correlated with the response variable (p>0.1), these insignificantly correlated independent variables should be deleted while retaining the significantly correlated independent variables, and a simplified model should be established accordingly, i.e., model selection. The specific method is to use the significant independent variables (p<0.1) selected from the full models of the three response variables as new input independent variables, and to perform a new round of PLSR modeling on the response variable, i.e., to extract latent variables by using formula (2) in sequence, and then establish the corresponding quadratic polynomial regression model. The corresponding results are shown in Tables 4 and 5 respectively. Figure 3 ,in, Figure 3 The middle (A) section represents biomass (OD). 600 ) result, Figure 3 Part (B) represents the results of xylanase activity. Figure 3 Section (C) represents the xylanase specific enzyme activity results. Clearly, for each of these three response variables, extracting one, rather than more, latent variables yields the best predictive power. (From Tables 4 and 5...) Figure 3 Some desirable properties of the PLSR model can be observed. First, not all deleted non-critical independent variables have absolute regression coefficients smaller than those of retained critical independent variables. For example, in the full model for the response variable Y2, the regression coefficients for the independent variables X3 and X4 to be deleted are -7.576 and 1.721, respectively, while the absolute values of the retained critical independent variables are... and The regression coefficients were -1.108 and 1.1789, respectively. Traditional stepwise regression models often tend to retain independent variables with larger regression coefficients while removing those with smaller coefficients. Secondly, for the response variables Y1, Y2, and Y3, although 7, 8, and 6 insignificant independent variables were removed from the full model, respectively, to establish selected models retaining 7, 6, and 8 significant independent variables, the differences in intercepts and regression coefficients between the selected and full models were not significant, and the signs of the regression coefficients were the same. Taking Y2 as an example, the intercepts of the full model and the selected model, as well as the independent variables X1 and X2,... X1X3 and X2X4 are 3661.446 and 3565.763, -15.403 and -18.216, 16.806 and 19.875, -1.108 and -1.311, 1.1789 and 1.394, -1.677 and -1.983, and 2.094 and 2.476, respectively. Third, comparing the full selection model, the absolute values of the regression coefficients of the corresponding significant independent variables retained in the selected model all show a slight increasing trend. This indicates a shift in the explanatory power for the response variable; that is, after the removal of insignificant independent variables from the selected model, the explanatory power previously attributed to insignificant independent variables in the full model has shifted to the significant independent variables in the selected model. Fourth, although the selected model removes a considerable number of insignificant independent variables, compared to the full model, the coefficient of determination R0 of the selected model is higher. 2 The coefficient of determination R only decreased slightly for the response variables Y1, Y2, and Y3. 2 The values decreased from 0.877, 0.9685, and 0.9822 in the full model to 0.8738, 0.9624, and 0.9551 in the selected model, with relative decreases of less than 3% for each. This demonstrates that establishing the selected model not only significantly simplifies the regression model, thereby improving its robustness and interpretability, but also largely maintains the model's fitting accuracy. Finally, Figure 3 This indicates that, similar to the full model, the scatter plot ("*") of the predicted values of the selected model (determined by the selected model) is also very close to the diagonal, and the relative errors of all predicted values are within ±5%.
[0099] 6. Solving the model using a single-objective genetic algorithm and experimental verification
[0100] The optimal solutions and optimal values of the full PLSR model and the selected model were obtained using a single-objective genetic algorithm (the optimal solution is the culture medium component value corresponding to the model's predicted best conditions, and the optimal value is the predicted value of the optimization objective corresponding to the model's predicted best conditions). The global optimization function `ga` in the Matlab software package was used to solve the full PLSR model and the finally optimized selected model for biomass, xylanase activity, and specific enzyme activity. The corresponding results are shown in […]. Figure 4 And Table 6, in which, Figure 4 The middle (A) section represents biomass (OD). 600 The solution result of ga) Figure 4 Part (B) represents the result of calculating the xylanase activity (ga). Figure 4 Part (C) represents the solution for the xylanase specific activity ga. QP represents the average fitness value of the entire model; QZ represents the optimal fitness value of the entire model; XP represents the average fitness value of the selected model, and XZ represents the optimal fitness value of the selected model. Clearly, after a finite number of evolutionary iterations, both models for the three response variables achieved good convergence, thus obtaining the optimal solution and optimal value. As shown in Table 5, the selected model converges faster due to its simplicity. Furthermore, the optimal values predicted by the selected model for all three response variables are greater than those predicted by the full model. This may be because the full model suffers from overfitting, resulting in high fitting accuracy but poor prediction accuracy. The selected model overcomes this drawback, as evidenced in Table 6. Validation experiments using the optimal solutions predicted by both models show that the experimental values for all three response variables are greater than the optimal values predicted by both models. Except for the full biomass model, the relative errors of the full and selected models for the three response variables are all less than 5%, indicating that the prediction accuracy of both models meets professional requirements. However, the relative errors of the selected model are all smaller than those of the full model, suggesting that the selected model's prediction accuracy is superior to the full model, indicating stronger generalization ability. In addition, both the full and selected models predicted the same optimal culture medium ratio for the three response variables, demonstrating the necessity of simplifying the full model and the robustness of the selected model.
[0101] Table 6 shows the optimization results of the full PLSR model and the selected model based on the single-objective genetic algorithm, and the results of their verification experiments.
[0102]
[0103] a The levels of the independent variables are consistent with the contents of the corresponding culture medium components in Table 1.
[0104] 7. Comparative Study of Multi-Objective Optimization Results of Two Models
[0105] Scatter plots were drawn to compare the fitting results of the full PLSR model and the selected model. Verification experiments were conducted on the multi-objective optimization genetic algorithm results of the full PLSR model and the selected model to compare the differences and advantages / disadvantages of the two models' predictive abilities. The multi-objective genetic algorithm function `gamultiobj` in the Matlab package was used to simultaneously optimize the full PLSR model and the selected model for biomass, xylanase activity, and specific enzyme activity using pairwise pairwise optimization and simultaneous optimization of the three response variables. The final Pareto optimal solution set distribution is shown below. Figure 5 and Figure 6 As shown, where Figure 5 Part (A) represents the dual-objective optimization of biomass and xylanase activity. Figure 5 Part (B) represents the dual-objective optimization of biomass and specific enzyme activity. Figure 5 Part (C) represents the dual-objective optimization of xylanase activity and specific enzyme activity. First, let's explain the Pareto optimal solution, also known as a non-dominated solution. It's mainly used when there are multiple optimization objectives. Due to conflicts and incomparability between these objectives, a solution that is best for one objective may be worst for others; it's difficult to find a solution that simultaneously optimizes all objectives. Therefore, for multi-objective optimization problems, there is usually a solution set rather than a single-objective optimization with a globally optimal solution. The characteristic of a non-dominated solution is that it's impossible to improve any objective function without weakening at least one other objective function. In other words, there is no solution in the feasible set where all objectives are not inferior to solution A, and at least one objective is superior to A. This is called a Pareto optimal solution. A set of optimal solutions for a set of objective functions is called the Pareto optimal set, or the Pareto front. All solutions in the Pareto front are non-dominated compared to solutions outside the Pareto front, and these solutions are also non-dominated among themselves. Therefore, these non-dominated solutions have the fewest objective conflicts compared to other solutions, providing decision-makers with a better choice space.
[0106] Secondly, the results of the bi-objective optimization of the response variables are analyzed, from... Figure 5It is easy to see that the distribution trends of Pareto optimal solutions in the full model and the selected model bi-objective optimization are basically the same, namely, they are Pareto front curves with a relatively uniform distribution and occasional discontinuities, each consisting of 18 Pareto optimal solutions. However, there are contradictory relationships between the pairwise response variables; that is, increasing one response variable will necessarily decrease the other. Therefore, if the two response variables are considered equally important, a compromise can be made by selecting the middle part of the Pareto optimal solution set; if the focus is on a certain optimization objective, the Pareto optimal solution set at the corresponding end can be selected. Most importantly, on the one hand, the Pareto front curves obtained from the three sets of bi-objective optimizations for the selected model are all located outside the Pareto front curves obtained from the corresponding full model bi-objective optimizations. Figure 5 As shown, the optimization of the two dual objectives, biomass vs. specific enzyme activity and xylanase activity vs. specific enzyme activity, is particularly effective. Figure 5 (B) and 5(C)), while the biomass vs. xylanase activity group was slightly lower, but except for a few scattered points, the Pareto front curve of the selected model was generally located outside the whole model. Figure 5 (A) Therefore, it can be said that the Pareto optimal solution set of the selected model constitutes the first-order Pareto front curve, while the Pareto optimal solution set of the full model constitutes the second-order Pareto front curve. In other words, the optimization objective predicted by the selected model is generally better than that of the full model, and it can obtain a larger optimization objective at the same time. The verification experimental results also show that the selected model is accurate in prediction and has good performance, as shown in Table 7. On the other hand, in the three sets of bi-objective optimizations, the coverage of the Pareto front curve of the selected model is significantly larger than that of the full model, and the distribution is more dispersed, that is, less crowded. This shows that the selected model not only provides decision-makers with a wider range of choices, but also that the Pareto optimal solution of the selected model is more representative and less substitutable. Although the Pareto front of biomass vs. xylanase activity of the full model has slightly better distribution continuity.
[0107] Final observation Figure 6 It can be seen that for the Pareto optimal solution set of the three-objective optimization, the Pareto front of the selected model is still located outside the full model, with a larger coverage and less crowding. This not only provides decision-makers with a larger selection space, but also the predicted Pareto optimal solution is better than that of the full model three-objective optimization, indicating that the Pareto optimal solution is more representative and less substitutable. The validation experiment selected at point 1 shows that the model prediction is accurate and the performance is good.
[0108] Table 7. Experimental Results of PLSR Full Model and Selected Model Validation Based on Multi-Objective Genetic Algorithm
[0109]
[0110] In addition to the above methods, this invention also provides a microbial culture medium optimization system based on correlation coefficient significance analysis, comprising:
[0111] The experimental module is used to design experiments uniformly with culture medium composition as the independent variable and experimental products as the response variable to obtain multiple experimental samples.
[0112] The PLSR full model construction module is used to construct a PLSR full model for each of the experimental samples using partial least squares regression. The PLSR full model is a mathematical expression in quadratic polynomial form representing the relationship between the response variable and all the independent variables.
[0113] The Pearson module is used to process the full PLSR model of any response variable using the Pearson correlation method to obtain the Pearson correlation coefficient matrix between each independent variable and the response variable.
[0114] The t-test module is used to obtain the t-test p-value of each independent variable and the response variable based on the Pearson correlation coefficient matrix between the independent variables and the response variable using the t-test method.
[0115] The significant independent variable screening module is used to screen the independent variables based on the t-test p-values of the independent variables and the response variable to obtain the significant independent variables corresponding to the response variable.
[0116] The model selection module is used to construct a PLSR model for the response variable using partial least squares method based on each experimental sample and each significant independent variable corresponding to the response variable; the PLSR model is a mathematical expression in quadratic polynomial form representing the relationship between the response variable and all the significant independent variables corresponding to the response variable.
[0117] A microbial culture medium construction module is used to construct microbial culture media based on the PLSR selection model for all the aforementioned response variables.
[0118] In practical applications, the t-test module specifically includes:
[0119] The t-test t-value calculation unit is used to calculate the t-value according to the formula. Calculate the t-test t-value, where r represents the Pearson correlation coefficient matrix between the independent variable and the response variable, n represents the total number of experimental samples, and t represents the t-test t-value.
[0120] The t-test p-value calculation unit is used to calculate the t-test p-value based on the t-test t-value using the tcdf function in the Matlab software package.
[0121] In practical applications, the significant independent variable screening module specifically includes:
[0122] The significant independent variable screening unit is used to determine that, for any given independent variable, if the t-test p-value between the independent variable and the response variable is less than a set value, the independent variable is a significant independent variable corresponding to the response variable.
[0123] In practical applications, the culture medium components include: sodium glycerophosphate, (NH4)2SO4, CaSO4·2H2O and PTM1; the experimental products include culture medium biomass, xylanase activity and specific enzyme activity.
[0124] This invention addresses the problems of conventional PLSR regression equations being excessively large and cumbersome, as well as their poor simplicity, interpretability, and predictive generalization ability. Furthermore, the calculation process is simple and requires minimal programming skills for convenient use.
[0125] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.
[0126] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for optimizing microbial culture media based on correlation coefficient significance analysis, characterized in that, include: Multiple experimental samples were obtained by using a uniform design experiment with culture medium composition as the independent variable and experimental products as the response variable. Based on each experimental sample, a partial least squares regression (PLSR) full model for each response variable is constructed using the partial least squares regression method; the PLSR full model is a mathematical expression in quadratic polynomial form representing the relationship between the response variable and all the independent variables; For any response variable, the Pearson correlation method is used to process the full PLSR model of the response variable to obtain the Pearson correlation coefficient matrix between the independent variables and the response variable; The t-test method is used to obtain the p-value of the t-test between the independent variables and the response variable based on the Pearson correlation coefficient matrix between the independent variables and the response variable; specifically: According to the formula Calculate the t-value for the t-test, where, r The Pearson correlation coefficient represents the relationship between the independent variable and the response variable. n The total number of experimental samples is represented by t, and t represents the t-value of the t-test. Using the Matlab software package tcdf The function calculates the t-test p-value based on the t-test t-value; The independent variables are screened based on the t-test p-values of their respective independent variables and the response variable to obtain the significant independent variables corresponding to the response variable; The PLSR selection model for the response variable is constructed using partial least squares method based on each experimental sample and each significant independent variable corresponding to the response variable; the PLSR selection model is a mathematical expression in quadratic polynomial form representing the relationship between the response variable and all the significant independent variables corresponding to the response variable. Microbial culture media were constructed based on the PLSR selection model for all the aforementioned response variables.
2. The method for optimizing microbial culture media based on correlation coefficient significance analysis according to claim 1, characterized in that, The step of filtering the independent variables based on the t-test p-values of their respective variables and the response variable to obtain the significant independent variables corresponding to the response variable specifically includes: For any independent variable, if the t-test p-value between the independent variable and the response variable is less than a set value, then the independent variable is determined to be a significant independent variable corresponding to the response variable.
3. The method for optimizing microbial culture media based on correlation coefficient significance analysis according to claim 1, characterized in that, The culture medium components include: sodium glycerophosphate, (NH4)2SO4, CaSO4·2H2O and PTM1; the experimental products include culture medium biomass, xylanase activity and specific enzyme activity.
4. A microbial culture medium optimization system based on correlation coefficient significance analysis, characterized in that, include: The experimental module is used to design a uniform experiment with culture medium composition as the independent variable and experimental product as the response variable to obtain multiple experimental samples. The PLSR full model construction module is used to construct a PLSR full model for each of the experimental samples using partial least squares regression; the PLSR full model is a mathematical expression in quadratic polynomial form representing the relationship between the response variable and all the independent variables; The Pearson module is used to process the full PLSR model of any response variable using the Pearson correlation method to obtain the Pearson correlation coefficient matrix between each independent variable and the response variable. The t-test module is used to obtain the t-test p-value of each independent variable and the response variable based on the Pearson correlation coefficient matrix between the independent variables and the response variable using the t-test method. The significant independent variable screening module is used to screen the independent variables based on the t-test p-values of the independent variables and the response variable to obtain the significant independent variables corresponding to the response variable; The model selection module is used to construct a PLSR model for the response variable using partial least squares method based on each experimental sample and each significant independent variable corresponding to the response variable; the PLSR model is a mathematical expression in quadratic polynomial form representing the relationship between the response variable and all the significant independent variables corresponding to the response variable. A microbial culture medium construction module is used to construct a microbial culture medium based on the PLSR selection model for all the aforementioned response variables; The t-test module specifically includes: The t-test t-value calculation unit is used to calculate the t-value according to the formula. Calculate the t-value for the t-test, where, r The Pearson correlation coefficient represents the relationship between the independent variable and the response variable. n The total number of experimental samples is represented by t, and t represents the t-value of the t-test. The t-test p-value calculation unit is used to calculate the p-value using the Matlab software package. tcdf The function calculates the t-test p-value based on the t-test t-value.
5. The microbial culture medium optimization system based on correlation coefficient significance analysis according to claim 4, characterized in that, The significant independent variable screening module specifically includes: The significant independent variable screening unit is used to determine that, for any given independent variable, if the t-test p-value between the independent variable and the response variable is less than a set value, the independent variable is a significant independent variable corresponding to the response variable.
6. The microbial culture medium optimization system based on correlation coefficient significance analysis according to claim 4, characterized in that, The culture medium components include: sodium glycerophosphate, (NH4)2SO4, CaSO4·2H2O and PTM1; the experimental products include culture medium biomass, xylanase activity and specific enzyme activity.
Citation Information
Patent Citations
Inversion method for copper elements in soil in vegetation-covered areas on basis of measured spectra of leaves
CN108663330A
Application of improved partial least squares regression method to microorganism culture medium optimization
CN108664719A
Petroleum concentration prediction method based on data rejection and local partial least squares
CN113094892A
Microbial culture medium optimization method and system
CN114373503A