Method for predicting fermentation data through multi-algorithm hybrid calculation based on random forest
By using a multi-algorithm hybrid calculation method based on random forests in the prediction of fermentation data, the problems of insufficient accuracy and poor robustness of traditional methods when processing complex fermentation data are solved, and high-precision fermentation process prediction and control are achieved.
Patent Information
- Application Number
- CN202510081355.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional fermentation data prediction methods have problems of insufficient accuracy and poor robustness when processing complex and changeable fermentation data, which are difficult to meet actual needs.
Using a multi-algorithm hybrid calculation method based on random forests, high-precision prediction of fermented data is achieved through data preprocessing, construction of random forests and KNN models and hyperparameter tuning, as well as model integration and uncertainty processing.
It improves the predictability and controllability of the fermentation process, enhances the accuracy and stability of the prediction results, and provides adaptability and uncertainty processing capabilities to new data.
Smart Images

Figure CN119940642A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of intelligent robots, and in particular to a method for predicting fermentation data based on multi-algorithm hybrid calculation based on random forest. Background Art
[0002] In the field of modern biotechnology, the optimization and control of the fermentation process is crucial. The fermentation process involves a variety of biochemical reactions, which are affected by a variety of environmental factors, such as ammonia quality, sugar addition, OD growth, pH value, temperature, tank pressure, air volume and dissolved oxygen. In order to achieve precise control of the fermentation process, these key parameters need to be monitored and predicted in real time. In recent years, with the development of big data and machine learning technologies, using these technologies to predict and optimize the fermentation process has become a research hotspot.
[0003] Traditional fermentation data prediction methods mostly rely on a single mathematical model or empirical formula. These methods often have limitations when dealing with complex and changeable fermentation data. First, a single model may not be able to fully capture the nonlinearity, dynamics and interaction between multiple variables of fermentation data. Second, when faced with large amounts of high-dimensional fermentation data, the computational efficiency and prediction accuracy of traditional methods are often difficult to meet actual needs. In addition, traditional methods lack the ability to adapt to new data and handle uncertainty, which limits their application in complex fermentation environments.
[0004] Therefore, a prediction method for fermentation data based on multi-algorithm hybrid computing of random forest was developed to overcome the limitations of traditional methods and achieve accurate prediction and control of the fermentation process. Summary of the invention
[0005] The purpose of the present invention is to make up for the shortcomings of the prior art and provide a prediction method for fermentation data based on multi-algorithm hybrid computing of random forests. This method achieves high-precision prediction of the fermentation process through sophisticated data preprocessing, construction and hyperparameter tuning of random forest and KNN models, as well as model integration and uncertainty processing. It effectively solves the problems of insufficient accuracy and poor robustness faced by traditional prediction technologies when processing complex fermentation data, and improves the predictability and controllability of the fermentation process.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a prediction method for fermentation data based on multi-algorithm hybrid calculation of random forest, the specific steps of the prediction method are:
[0007] S100, data preprocessing: obtain relevant data on ammonia quality, sugar addition, OD growth, pH, temperature, tank pressure, air volume, and dissolved oxygen from the data sensor in the fermentation tank, and decompose the data into trend terms, seasonal terms, and residual terms through the time series decomposition formula. Suppose the fermentation data vector collected at time t is x t =(A t , B t , O t , P t , T t , Pr t , F t , D t ) T , where A t Indicates the quality of ammonia water, B t Indicates the amount of sugar supplement, O t Indicates OD growth, P t Indicates pH value, T t Indicates temperature, Pr t Indicates tank pressure, F t Indicates air volume, D t represents the dissolved oxygen content, and the entire time series data set is represented by X = {x t |t=1,2,…,N}, the time series decomposition formula is: t =T t +S t +R t +∈ t , where T t is the trend term, S t is the seasonal term, R t is the residual term, ∈ t is the random error term, and the trend term calculation formula is: Where m is the moving average window size, weight w i The linear decreasing weighting method is used for calculation, and the formula is: The seasonal term calculation formula is: Where p is the seasonal period, K is the number of historical periods used to calculate the seasonal term, and the residual term R t The formula is obtained by subtracting the trend term and seasonal term from the original data: R t =x t -T t -S t -∈ t , smooth the trend term as a new feature, use the periodic model to fit and predict the seasonal term, merge the residual term with the prediction results of the seasonal term, perform outlier detection on the merged data, delete the detected outliers, and standardize the data:
[0008] S200, random forest model construction: from the preprocessed original training data set, a subset is generated by sampling with replacement, a subset of the same size as the original data set is generated, and the optimal feature split is selected based on the information gain criterion, and the above feature selection and splitting process is repeated, and the decision tree is continuously expanded until the above feature selection and splitting process is repeated until all samples in the node belong to the same category, and the above decision tree construction process is repeated T times multiple times, each time using a different sampling subset and a randomly selected feature subset to form a random forest. In the prediction stage, the output of the random forest is the average of the prediction results of all decision trees;
[0009] S300, KNN model construction: prepare a training data set containing multiple feature columns and target value columns, perform missing value processing, outlier detection and processing, feature scaling operations, perform KNN training, determine the value of parameter K through cross-validation, initialize the KNN model according to the K value and the selected distance measurement method, and when faced with a new sample that needs to be predicted, the model will calculate the distance between the new sample and all samples in the training data set according to the pre-set distance measurement method, and select the average of the K neighbor target values as the prediction result;
[0010] S400, hyperparameter tuning: For random forests, determine the hyperparameter ranges for the number of trees, maximum depth, minimum number of samples per node, and number of randomly selected features, construct a grid of all possible combinations through grid search, randomly search and randomly generate combinations, define an objective function, use hyperparameter combinations as input, build a Bayesian optimization model based on the prior distribution of hyperparameters and the objective function, use cross-validation to evaluate the performance of the model under different hyperparameter combinations, and select the hyperparameter combination that allows the model to generalize best; For KNN models, set the K value and the range of parameters related to the distance measurement method, construct a grid of all possible combinations through grid search, randomly search and randomly generate combinations, define an objective function, use hyperparameter combinations as input, build a Bayesian optimization model based on the prior distribution of hyperparameters and the objective function, use cross-validation to evaluate the performance of the model under different hyperparameter combinations, and select the hyperparameter combination that allows the model to generalize best;
[0011] S500, model integration and uncertainty processing: Combine the training model of KNN with the training model of random forest, use the random forest model to make preliminary predictions on new samples, and obtain the prediction results. By calculating the prediction variance between decision trees, the uncertainty judgment credibility of the random forest prediction is evaluated. The prediction confidence is lower than the threshold τ, and the sample is handed over to the KNN model for further screening and refinement of the prediction. The final prediction result integrates the prediction results of the two models through the weighted average fusion strategy, and evaluates the uncertainty of the prediction by calculating the prediction variance between decision trees. When the uncertainty of the model prediction exceeds the set threshold, the corresponding adjustment strategy is triggered. The uncertainty quantification of the KNN model estimates the uncertainty based on the distribution density and distance difference of neighbor samples. When the uncertainty of the model prediction exceeds the set threshold, the corresponding adjustment strategy is triggered.
[0012] Furthermore, in the S100, the trend item is smoothed by a weighted exponential smoothing formula in data preprocessing, and the formula is: in, is the new characteristic value of the trend term after smoothing at time t, x t is the original trend term value at time t, α is the smoothing coefficient, It is the trend item value after smoothing at the previous moment t-1.
[0013] Furthermore, in the S100, seasonal terms are fitted and predicted by Fourier series expansion formula in data preprocessing, and the formula is: Among them, S t is the seasonal term fitting value at time t, p is the seasonal period, a k and b k are the Fourier coefficients and Δt is the time offset.
[0014] Furthermore, in the S100, in the data preprocessing, the residual term and the seasonal term prediction result are combined by a weighted combination formula, and the formula is: Among them, M t is the result of merging at time t, is the predicted value of the seasonal term at time t, R t is the residual value at time t, and β is the weighting coefficient.
[0015] Furthermore, in the S100, outliers are detected and deleted by using an outlier detection formula in data preprocessing, and the formula is: Among them, Outlier t is the outlier judgment index at time t, M t is the merged data value at time t, is the average value of the combined data values at the past w time points, θ tis the dynamic threshold, calculated as: t = k × std t-w:t-1 , where k is the threshold multiple, std t-w:t-1 is the standard deviation of the combined data values at the past w time points.
[0016] Furthermore, in the S200, the optimal feature splitting is selected based on the information gain criterion in the random forest model construction, and the information entropy is calculated. Assume that the class label set of the samples in the data set D is C = {c 1 , c 2 , …, c m}, the number of samples is |D|, belonging to category c i The sample size is |D i |, then the information entropy Ent(D) of data set D is calculated as: Conditional entropy calculation, for feature A, the value set is {a 1 , a 2 , …, a v}, the dataset D is divided into v subsets D according to the value of feature A 1 , D 2 , …, D v , where D j Indicates that the feature A takes the value of a j The conditional entropy Ent(D|A) of the dataset D under the condition of feature A is calculated as follows: Information gain calculation and optimal feature selection. The information gain Gain(D, A) of feature A relative to data set D is calculated as: Gain(D, A) = Ent(D) - Ent(D|A). The information gain of each feature is calculated, and the feature with the largest information gain is selected as the optimal feature for splitting.
[0017] Furthermore, in the S300, the KNN model is constructed by determining the K value through a dynamic K value determination formula, and the calculation formula is: Where n is the number of samples in the training data set, x i is a characteristic value of the i-th sample in the training data set, is the mean of the feature in the training data set, and α and β are adjustment parameters.
[0018] Furthermore, in S300, the distance measurement is performed by using an adaptive distance measurement formula in the KNN model construction, and the formula is: Where x=(x 1 , x 2 , …, x n ) and y=(y 1 ,y 2 , …, y n) are two sample points, n is the number of features, and δ is an adaptive parameter.
[0019] Furthermore, in the S500, the uncertainty of the random forest prediction and the policy triggering are evaluated by calculating the prediction variance between the decision trees in the model integration and uncertainty processing. The calculation formula of the prediction variance is: when It is considered that the random forest has a low prediction confidence for the sample, triggering the corresponding strategy: data enhancement: appropriately transform the original training data and retrain the random forest model; model adjustment: adjust the hyperparameters of the random forest; integration strategy optimization: consider using other more complex integration methods or adjusting the weight calculation method of model fusion.
[0020] Furthermore, in the S500, uncertainty quantification of the KNN model in the model integration and uncertainty processing estimates uncertainty and strategy triggering according to the distribution density and distance difference of the neighbor samples, and calculates the distribution density standard deviation σρ of the neighbor samples, and the formula is: Among them, ρ ij is the local density of the jth neighbor sample, is the average value of the local density of K neighbor samples. At the same time, the standard deviation σ of the neighbor sample distance is calculated. d , the formula is: Among them, d(x, x ij ) is the distance between sample x and its jth neighbor sample, is the average distance of K neighbor samples, and the uncertainty index U of the KNN model KNN The calculation formula is: KNN =α 1 ·σ ρ +α 2 ·σ d , where α 1 and α 2 is the trade-off coefficient, when U KNN >θ, θ is the set uncertainty threshold of the KNN model, triggering the corresponding strategy: reselect neighbor samples: adjust the K value or distance measurement parameters, and recalculate the distance between the sample and other samples in the training data set; improve data preprocessing: perform further preprocessing operations on the data, adopt more complex feature engineering methods, perform more sophisticated processing on outliers, or try different data standardization methods; model fusion adjustment: according to the uncertainty of the KNN model, dynamically adjust its weight when fusing with the random forest model, or consider using other fusion methods.
[0021] Compared with the prior art, this method for predicting fermentation data based on multi-algorithm hybrid calculation based on random forest has the following beneficial effects:
[0022] 1. The present invention uses time series decomposition technology in the data preprocessing stage to decompose complex fermentation data into trend terms, seasonal terms and residual terms, so that subsequent processing and prediction are more accurate. By smoothing the trend terms and fitting and predicting the seasonal terms with a periodic model, the data processing efficiency and prediction accuracy are improved. At the same time, two machine learning algorithms, random forest and KNN, are introduced in the model construction stage. The performance of the model is optimized through hyperparameter tuning technology, and the prediction accuracy and stability are further improved. This multi-algorithm hybrid calculation method can give full play to the advantages of different algorithms, improve the generalization ability of the model, and provide strong support for the optimization and control of the fermentation process.
[0023] 2. The present invention evaluates the uncertainty of random forest prediction by calculating the prediction variance between decision trees, and triggers the corresponding adjustment strategy according to the uncertainty index. This uncertainty processing mechanism can timely discover and correct potential errors in the prediction results, improve the reliability and stability of the prediction results, and at the same time, a weighted average fusion strategy is used to integrate the prediction results of the random forest and KNN models, making full use of the advantages of the two models and further improving the accuracy of the prediction. A method for estimating the uncertainty of the KNN model based on the distribution density and distance difference of neighbor samples is also proposed, which provides a new idea for the quantification of model uncertainty and a more accurate basis for subsequent strategy adjustments.
[0024] Other advantages, objectives and features of the present invention will be set forth in part in the following description and, in part, will be apparent to those skilled in the art based on an examination of the following or may be taught from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0026] Figure 1 A flowchart of a method for predicting fermentation data based on multi-algorithm hybrid calculation based on random forest;
[0027] Figure 2 This is a schematic diagram of some test data of the high-density fermentation process of recombinant Escherichia coli;
[0028] Figure 3 This is a schematic diagram of the mean absolute error results when only KNN is used;
[0029] Figure 4 This is a schematic diagram of the mean absolute error results when only random forest is used;
[0030] Figure 5 A schematic diagram of the mean absolute error results when using a mixed model. DETAILED DESCRIPTION
[0031] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation mode, structure, characteristics and effects of the present invention are described in detail below in combination with the accompanying drawings and preferred embodiments.
[0032] Embodiment 1
[0033] Predicting the OD growth value in a fermentation process
[0034] The data sensors of the fermentation tanks are used to obtain the relevant data of ammonia quality, sugar addition, OD growth, pH, temperature, tank pressure, air volume, and dissolved oxygen. The data are decomposed into trend terms, seasonal terms, and residual terms through the time series decomposition formula. The fermentation data vector collected at time t is x t =(A t , B t , O t , P t , T t , Pr t , F t , D t ) T , where A t Indicates the quality of ammonia water, B t Indicates the amount of sugar supplement, O t Indicates OD growth, P t Indicates pH value, T t Indicates temperature, Pr t Indicates tank pressure, F t Indicates air volume, D t represents the dissolved oxygen content, and the entire time series data set is represented by X = {x t |t=1.2,…,N}, the time series decomposition formula is: t =T t +S t +R t +∈ t , where T t is the trend term, S t is the seasonal term, R t is the residual term, ∈ t is the random error term, and the trend term calculation formula is: Where m is the moving average window size, weight w iThe linear decreasing weighting method is used for calculation, and the formula is: The seasonal term calculation formula is: Where p is the seasonal period, K is the number of historical periods used to calculate the seasonal term, and the residual term R t The formula is obtained by subtracting the trend term and seasonal term from the original data: R t =x t -T t -S t -∈ t , the trend item is smoothed by the weighted exponential smoothing formula, the formula is: The trend term is smoothed by the weighted exponential smoothing formula, which is: The seasonal terms are fitted and predicted using the Fourier series expansion formula: The residual term and the seasonal term prediction results are combined through the weighted combination formula, the formula is: Outliers are detected and deleted through the outlier detection formula. The formula is: At the same time, the data is standardized.
[0035] Random forest model construction, from the preprocessed original training data set, generate subsets by sampling with replacement, select the optimal feature split based on the information gain criterion, and calculate the information entropy. Let the class label set of samples in the data set D be C = {c 1 , c 2 , …, c m}, the number of samples is |D|, belonging to category c i The sample size is |D i |, then the information entropy Ent(D) of data set D is calculated as: Conditional entropy calculation, for feature A, the value set is {a 1 , a 2 , …, a v}, the dataset D is divided into v subsets D according to the value of feature A 1 , D 2 , …, D v , where D j Indicates that the feature A takes the value of a j The conditional entropy Ent(D|A) of the dataset D under the condition of feature A is calculated as follows: Information gain calculation and optimal feature selection, the information gain Gain(D, A) calculation formula of feature A relative to data set D is: Gain(D, A) = Ent(D)-Ent(D|A), repeat the above feature selection and splitting process, and continuously expand the decision tree until all samples in the node belong to the same category, repeat the above decision tree construction process T times, each time using a different sampling subset and randomly selected feature subset to form a random forest. In the prediction stage, the output of the random forest is the average of the prediction results of all decision trees.
[0036] KNN model construction, prepare a training data set containing multiple feature columns and target value columns. Before KNN training, perform missing value processing, outlier detection and processing, and feature scaling operations. Determine the value of parameter K through cross-validation. The calculation formula is: The distance measurement is performed through the adaptive distance measurement formula, the formula is: The KNN model is initialized according to the K value and the selected distance measurement method. When a new sample needs to be predicted, the model will calculate the distance between the new sample and all samples in the training data set based on the preset distance measurement method, and select the average of the K neighbor target values as the prediction result.
[0037] Hyperparameter tuning: For random forests, determine the hyperparameter ranges for the number of trees, maximum depth, minimum number of samples per node, and number of randomly selected features; build combinations through grid search and random search, define the objective function, build a Bayesian optimization model based on the prior distribution of hyperparameters and the objective function, use cross-validation to evaluate performance, and select hyperparameter combinations that allow the model to generalize well. For the KNN model, set the value and range of parameters related to the distance metric method, build combinations through the grid search method, build a Bayesian optimization model and use cross-validation to evaluate, and select hyperparameter combinations that allow the model to generalize well.
[0038] Model integration and uncertainty processing, use the random forest model to make preliminary predictions for new samples, set the lower prediction confidence threshold as τ, and evaluate the uncertainty of random forest predictions by calculating the prediction variance between decision trees. The formula is: If the variance exceeds τ, the sample will be handed over to the KNN model for further screening and refinement of prediction. For random forest, when the calculated It is considered that the prediction confidence of the random forest for this sample is low, triggering the corresponding strategy, which is: data enhancement: appropriately transform the original training data and retrain the random forest model; model adjustment: adjust the hyperparameters of the random forest; integration strategy optimization: consider using other more complex integration methods or adjusting the weight calculation method of model fusion. For the KNN model, calculate the distribution density standard deviation σρ of the neighbor samples, the formula is: where ρij is the local density of the jth neighbor sample, is the average value of the local density of K neighbor samples. At the same time, the standard deviation σ of the neighbor sample distance is calculated. d , the formula is: Among them, d(x, x ij ) is the distance between sample x and its jth neighbor sample, is the average distance of K neighbor samples, and the uncertainty index U of the KNN model KNN The calculation formula is: KNN =α 1 ·σ ρ +α 2 ·σ d , where α 1 and α 2 is the trade-off coefficient, when U KNN >θ, θ is the set uncertainty threshold of the KNN model, triggering the corresponding strategy: reselect neighbor samples: adjust the K value or distance measurement parameters, and recalculate the distance between the sample and other samples in the training data set; improve data preprocessing: perform further preprocessing operations on the data, adopt more complex feature engineering methods, perform more sophisticated processing on outliers, or try different data standardization methods; model fusion adjustment: according to the uncertainty of the KNN model, dynamically adjust its weight when fusing with the random forest model, or consider using other fusion methods.
[0039] In summary, this embodiment demonstrates the prediction process of the OD growth value in a fermentation process. First, data preprocessing is performed, including data acquisition, decomposition, and outlier processing. Then, random forest and KNN models are constructed, and hyperparameters are tuned respectively. Finally, through model integration and uncertainty processing, the results of the two models are combined, and the corresponding strategy adjustment is triggered according to the uncertainty indicators and thresholds of each model to finally obtain the prediction result.
[0040] Embodiment 2
[0041] Data acquisition and preprocessing, the relevant data within a period of time are obtained from the data sensor in the fermentation tank, including ammonia quality, sugar addition, OD growth, PH, temperature, tank pressure, air volume, dissolved oxygen content, and the data is decomposed into trend terms, seasonal terms and residual terms through the time series decomposition formula. Let the fermentation data vector collected at time t be x t =(A t , B t , O t , P t , T t , Pr t , F t , D t ) T , where A t Indicates the quality of ammonia water, B t Indicates the amount of sugar supplement, O t Indicates OD growth, P t Indicates pH value, T t Indicates temperature, Pr t Indicates tank pressure, F t Indicates air volume, D t represents the dissolved oxygen content, and the entire time series data set is represented by X = {x t |t=1.2,…,N}, the time series decomposition formula is: t =T t +S t +R t +∈ t , where T t is the trend term, S t is the seasonal term, R t is the residual term, ∈ t is the random error term, and the trend term calculation formula is: Where m is the moving average window size, weight w i The linear decreasing weighting method is used for calculation, and the formula is: The seasonal term calculation formula is: Where p is the seasonal period, K is the number of historical periods used to calculate the seasonal term, and the residual term R t The formula is obtained by subtracting the trend term and seasonal term from the original data: R t =x t -T t -S t -∈ t , the trend item is smoothed, the formula is: The seasonal terms are fitted and predicted using the Fourier series expansion formula: The residual term is combined with the seasonal term forecast results through the weighted combination formula, the formula is: Outliers are detected and deleted through the outlier detection formula. The formula is: All data were normalized at the same time.
[0042] The random forest model is constructed by generating a subset with replacement sampling to generate a subset of the same size as the original data set, and selecting the optimal feature split based on the information gain criterion and calculating the information entropy. Suppose the class label set of the samples in the data set D is C = {c 1 , c 2 , …, c m}, the number of samples is |D|, belonging to category c i The sample size is |D i |, then the information entropy Ent(D) of data set D is calculated as: Conditional entropy calculation, for feature A, the value set is {a 1 , a 2 , …, a v}, the dataset D is divided into v subsets D according to the value of feature A 1 , D 2 , …, D v , where D j Indicates that the feature A takes the value of a j The conditional entropy Ent(D|A) of the dataset D under the condition of feature A is calculated as follows: Information gain calculation and optimal feature selection. The information gain Gain(D, A) of feature A relative to data set D is calculated as follows: Gain(D, A) = Ent(D) - Ent(D|A). The information gain of each feature is calculated, and the feature with the largest information gain is selected as the optimal feature for splitting. The above feature selection and splitting process is repeated, and the decision tree is continuously expanded until all samples in the node belong to the same category. The above decision tree construction process is repeated T times, each time using a different sampling subset and a randomly selected feature subset to form a random forest.
[0043] KNN model construction, prepare training data set, perform missing value processing, outlier detection and processing, feature scaling operations, and determine the value of parameter K through cross-validation. The calculation formula is: The distance measurement is performed through the adaptive distance measurement formula, the formula is: The KNN model is initialized according to the K value and the selected distance measurement method. When a new sample needs to be predicted, the model will calculate the distance between the new sample and all samples in the training data set based on the preset distance measurement method, and select the average of the K neighbor target values as the prediction result.
[0044] Hyperparameter tuning: Random forest hyperparameter tuning determines the hyperparameter ranges of the number of trees, maximum depth, minimum number of samples per node, and number of randomly selected features, and evaluates the performance of the model under different hyperparameter combinations through relevant methods. KNN model hyperparameter tuning: Set the K value and the range of parameters related to the distance measurement method, and evaluate the performance of the model under different hyperparameter combinations through relevant methods to select the appropriate hyperparameter combination.
[0045] Model integration and uncertainty handling, random forest prediction and uncertainty assessment evaluate the uncertainty and strategy triggering of random forest prediction by calculating the prediction variance between decision trees. The calculation formula of prediction variance is: when It is considered that the random forest has a low confidence in the prediction of this sample, and the corresponding strategy is triggered. The strategy is: Data enhancement: Appropriately transform the original training data, retrain the random forest model, further screen and evaluate the uncertainty of the KNN model, estimate the uncertainty and trigger the strategy based on the distribution density and distance difference of the neighbor samples, and calculate the standard deviation σ of the distribution density of the neighbor samples ρ , the formula is: Among them, ρ ij is the local density of the jth neighbor sample, is the average value of the local density of K neighbor samples. At the same time, the standard deviation σ of the neighbor sample distance is calculated. d , the formula is: Among them, d(x, x ij ) is the distance between sample x and its jth neighbor sample, is the average distance of K neighbor samples, and the uncertainty index U of the KNN model KNN The calculation formula is: KNN =α 1 ·σ ρ +α 2 ·σ d , where α 1 and α 2 is the trade-off coefficient, when U KNN >θ, θ is the uncertainty threshold of the KNN model, which triggers the corresponding strategy: reselect neighbor samples: adjust the K value or distance measurement parameter, recalculate the distance between the sample and other samples in the training data set, and integrate the prediction results of the two models through the weighted average fusion strategy.
[0046] In summary, the above embodiment demonstrates a fermentation data prediction method based on random forest and KNN, including data acquisition and preprocessing, involving time series decomposition, outlier processing and standardization, building a random forest model and a KNN model, performing feature selection, parameter determination and initialization, and hyperparameter tuning respectively; finally, performing model integration and uncertainty processing, combining the results of the two models and adjusting the strategy based on uncertainty.
[0047] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A method for predicting fermentation data based on multi-algorithm hybrid calculation of random forest, characterized in that: The specific steps of this prediction method are: S100, data preprocessing: obtain relevant data on ammonia quality, sugar addition, OD growth, pH, temperature, tank pressure, air volume, and dissolved oxygen from the data sensor in the fermentation tank, and decompose the data into trend terms, seasonal terms, and residual terms through the time series decomposition formula. Suppose the fermentation data vector collected at time t is x t =(A t , B t , O t , P t , T t , Pr t , F t , D t ) T , where A t Indicates the quality of ammonia water, B t Indicates the amount of sugar supplement, O t Indicates OD growth, P t Indicates pH value, T t Indicates temperature, Pr t Indicates tank pressure, F t Indicates air volume, D t represents the dissolved oxygen content, and the entire time series data set is represented by X = {x t |t=1,2,…,N}, the time series decomposition formula is: t =T t +S t +R t +∈ t , where T t is the trend term, S t is the seasonal term, R t is the residual term, ∈ t is the random error term, and the trend term calculation formula is: Where m is the moving average window size, weight w i The linear decreasing weighting method is used for calculation, and the formula is: The seasonal term is calculated as: Where p is the seasonal period, K is the number of historical periods used to calculate the seasonal term, and the residual term R t The formula is obtained by subtracting the trend term and seasonal term from the original data: R t =x t -T t -S t -∈ t , smooth the trend term as a new feature, use the periodic model to fit and predict the seasonal term, merge the residual term with the prediction result of the seasonal term, perform outlier detection on the merged data, delete the detected outliers, and standardize the data; S200, random forest model construction: from the preprocessed original training data set, a subset is generated by sampling with replacement, a subset of the same size as the original data set is generated, and the optimal feature split is selected based on the information gain criterion, and the above feature selection and splitting process is repeated, and the decision tree is continuously expanded until the above feature selection and splitting process is repeated until all samples in the node belong to the same category, and the above decision tree construction process is repeated T times multiple times, each time using a different sampling subset and a randomly selected feature subset to form a random forest. In the prediction stage, the output of the random forest is the average of the prediction results of all decision trees; S300, KNN model construction: prepare a training data set containing multiple feature columns and target value columns, perform missing value processing, outlier detection and processing, feature scaling operations, perform KNN training, determine the value of parameter K through cross-validation, initialize the KNN model according to the K value and the selected distance measurement method, and when faced with a new sample that needs to be predicted, the model will calculate the distance between the new sample and all samples in the training data set according to the pre-set distance measurement method, and select the average of the K neighbor target values as the prediction result; S400, hyperparameter tuning: For random forests, determine the hyperparameter ranges for the number of trees, maximum depth, minimum number of samples per node, and number of randomly selected features, construct a grid of all possible combinations through grid search, randomly search and randomly generate combinations, define an objective function, use hyperparameter combinations as input, build a Bayesian optimization model based on the prior distribution of hyperparameters and the objective function, use cross-validation to evaluate the performance of the model under different hyperparameter combinations, and select the hyperparameter combination that allows the model to generalize best; For KNN models, set the K value and the range of parameters related to the distance measurement method, construct a grid of all possible combinations through grid search, randomly search and randomly generate combinations, define an objective function, use hyperparameter combinations as input, build a Bayesian optimization model based on the prior distribution of hyperparameters and the objective function, use cross-validation to evaluate the performance of the model under different hyperparameter combinations, and select the hyperparameter combination that allows the model to generalize best; S500, model integration and uncertainty processing: Combine the training model of KNN with the training model of random forest, use the random forest model to make preliminary predictions on new samples, and obtain the prediction results. By calculating the prediction variance between decision trees, the uncertainty judgment credibility of the random forest prediction is evaluated. The prediction confidence is lower than the threshold τ, and the sample is handed over to the KNN model for further screening and refinement of the prediction. The final prediction result integrates the prediction results of the two models through the weighted average fusion strategy, and evaluates the uncertainty of the prediction by calculating the prediction variance between decision trees. When the uncertainty of the model prediction exceeds the set threshold, the corresponding adjustment strategy is triggered. The uncertainty quantification of the KNN model estimates the uncertainty based on the distribution density and distance difference of neighbor samples. When the uncertainty of the model prediction exceeds the set threshold, the corresponding adjustment strategy is triggered.
2. The method for predicting fermentation data based on multi-algorithm hybrid calculation based on random forest according to claim 1, characterized in that: In the S100, in the data preprocessing, the trend item is smoothed by a weighted exponential smoothing formula, and the formula is: in, is the new characteristic value of the trend term after smoothing at time t, x t is the original trend term value at time t, α is the smoothing coefficient, It is the trend item value after smoothing at the previous moment t-1.
3. The method for predicting fermentation data based on multi-algorithm hybrid calculation based on random forest according to claim 1, characterized in that: In the S100, the seasonal term is fitted and predicted by the Fourier series expansion formula in the data preprocessing, and the formula is: Among them, S t is the seasonal term fitting value at time t, p is the seasonal period, a k and b k are the Fourier coefficients and Δt is the time offset.
4. The method for predicting fermentation data based on multi-algorithm hybrid calculation based on random forest according to claim 1, characterized in that: In the S100, in the data preprocessing, the residual term and the seasonal term prediction result are combined by a weighted combination formula, and the formula is: Among them, M t is the result of merging at time t, is the predicted value of the seasonal term at time t, R t is the residual value at time t, and β is the weighting coefficient.
5. The method for predicting fermentation data based on multi-algorithm hybrid calculation based on random forest according to claim 1, characterized in that: In the S100, outliers are detected and deleted by using an outlier detection formula in data preprocessing. The formula is: Among them, Outlier t is the outlier judgment index at time t, M t is the merged data value at time t, is the average value of the combined data values at the past w time points, θ t is the dynamic threshold, calculated as: t =k×std t-w:t-1 , where k is the threshold multiple, std t-w:t-1 is the standard deviation of the combined data values at the past w time points.
6. The method for predicting fermentation data based on multi-algorithm hybrid calculation based on random forest according to claim 1, characterized in that: In the S200, the optimal feature splitting is selected based on the information gain criterion in the random forest model construction, and the information entropy is calculated. Assume that the category label set of the samples in the data set D is C = {c1, c2, ..., c m }, the number of samples is |D|, belonging to category c i The sample size is |D i |, then the information entropy Ent(D) of data set D is calculated as: Conditional entropy calculation, for feature A, the value set is {a1, a2, ..., a v }, the dataset D is divided into v subsets D according to the value of feature A 1 , D 2 , …, D v , where D j Indicates that the feature A takes the value of a j The conditional entropy Ent(D|A) of the dataset D under the condition of feature A is calculated as follows: Information gain calculation and optimal feature selection. The information gain Gain(D, A) of feature A relative to data set D is calculated as: Gain(D, A) = Ent(D) - Ent(D|A). The information gain of each feature is calculated, and the feature with the largest information gain is selected as the optimal feature for splitting.
7. The method for predicting fermentation data based on multi-algorithm hybrid calculation based on random forest according to claim 1, characterized in that: In the S300, the KNN model is constructed by using a dynamic K value determination formula to determine the K value, and the calculation formula is: Where n is the number of samples in the training data set, x i is a characteristic value of the i-th sample in the training data set, is the mean of the feature in the training data set, and α and β are adjustment parameters.
8. The method for predicting fermentation data based on multi-algorithm hybrid calculation based on random forest according to claim 1, characterized in that: In the S300, the distance measurement is performed by using an adaptive distance measurement formula in the KNN model construction, and the formula is: Where x = (x1, x2, ..., x n ) and y=(y1,y2,…,y n ) are two sample points, n is the number of features, and δ is an adaptive parameter.
9. The method for predicting fermentation data based on multi-algorithm hybrid calculation based on random forest according to claim 1, characterized in that: In the S500, the uncertainty of the random forest prediction and the policy triggering are evaluated by calculating the prediction variance between the decision trees in the model integration and uncertainty processing. The calculation formula of the prediction variance is: when It is considered that the random forest has a low prediction confidence for the sample, triggering the corresponding strategy: data enhancement: appropriately transform the original training data and retrain the random forest model; model adjustment: adjust the hyperparameters of the random forest; integration strategy optimization: consider using other more complex integration methods or adjusting the weight calculation method of model fusion.
10. The method for predicting fermentation data based on multi-algorithm hybrid calculation based on random forest according to claim 1, characterized in that: The uncertainty quantification of the KNN model in the model integration and uncertainty processing in S500 estimates uncertainty and strategy triggering according to the distribution density and distance difference of neighbor samples, and calculates the distribution density standard deviation σ of the neighbor samples. ρ , the formula is: Among them, ρ ij is the local density of the jth neighbor sample, is the average value of the local density of K neighbor samples. At the same time, the standard deviation σ of the neighbor sample distance is calculated. d , the formula is: Among them, d(x, x ij ) is the distance between sample x and its jth neighbor sample, is the average distance of K neighbor samples, and the uncertainty index U of the KNN model KNN The calculation formula is: KNN =α1·σ ρ +α2·σ d , where α1 and α2 are trade-off coefficients. KNN >θ, θ is the set uncertainty threshold of the KNN model, triggering the corresponding strategy: reselect neighbor samples: adjust the K value or distance measurement parameters, and recalculate the distance between the sample and other samples in the training data set; improve data preprocessing: perform further preprocessing operations on the data, adopt more complex feature engineering methods, perform more sophisticated processing on outliers, or try different data standardization methods; model fusion adjustment: according to the uncertainty of the KNN model, dynamically adjust its weight when fusing with the random forest model, or consider using other fusion methods.
Citation Information
Cited By
Partial pressure line loss rate prediction method and system based on random forest
CN121146126A
Service index anomaly attribution method, device, equipment and medium
CN121327375A