Papermaking enzyme two-way prediction method and system based on ensemble algorithm
By integrating algorithms and machine learning models, the complexity of multi-enzyme mixture prediction was solved, efficient two-way prediction of papermaking enzymes was achieved, more flexible tools were provided, and the application of bio-enzymes in the papermaking industry was promoted.
Patent Information
- Application Number
- CN202410129544.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-30
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-01-30
AI Technical Summary
It is difficult to effectively predict the enzymatic activity and reaction conditions of papermaking enzymes with existing technologies, especially in the case of multi-enzyme mixtures, which leads to complexity and difficulty in prediction.
An ensemble algorithm-based approach was adopted to establish a multi-enzyme dataset and perform bidirectional prediction by combining a machine learning model with a bagging algorithm and an optimization algorithm to predict the optimal application conditions and enzyme combinations of multiple enzymes.
It achieves efficient prediction of multi-enzyme application scenarios, reduces experimental time and cost, improves prediction performance and model robustness, and can reversely predict the optimal enzyme combination and conditions.
Smart Images

Figure CN118248236B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of artificial intelligence and pulp and papermaking, and particularly relates to a papermaking enzyme bidirectional prediction method and system based on an integrated algorithm, a terminal device, and a computer readable storage medium. BACKGROUND
[0002] In the pulp and papermaking industry, biological enzymes have become important biological catalysts and are widely used to improve paper quality, reduce production costs, and protect the environment. However, there are still some problems and shortcomings in the current research on papermaking enzymes. First, the activity of papermaking enzymes is affected by various factors such as temperature, pH value, and substrate concentration, so predicting the enzyme activity and reaction conditions of papermaking enzymes is a challenging task. Second, due to the variety of papermaking enzymes, there is interaction between different enzymes, so predicting the reaction conditions and enzyme activity of multiple enzyme mixtures is a more complex and difficult task. SUMMARY
[0003] To solve the above problems of the prior art, the application provides a papermaking enzyme bidirectional prediction method and system based on an integrated algorithm, a terminal device, and a computer readable storage medium. The machine learning model realizes bidirectional prediction of multiple enzyme application scenarios, which can not only predict the most suitable application conditions of multiple enzymes, but also inversely predict the most suitable enzyme combination under different conditions. Through bidirectional prediction, the performance and optimal combination of multiple enzymes under different conditions can be better understood, and more intelligent decisions can be made, providing more flexible and efficient tools. The application can promote the wider application of biological enzyme technology in the papermaking industry and play a guiding role in the green development of the light industry papermaking industry.
[0004] The first object of the application is to provide a papermaking enzyme bidirectional prediction method based on an integrated algorithm.
[0005] The second object of the application is to provide a papermaking enzyme bidirectional prediction system based on an integrated algorithm.
[0006] The third object of the application is to provide a terminal device.
[0007] The fourth object of the application is to provide a computer readable storage medium.
[0008] The first object of the application can be achieved by adopting the following technical solutions:
[0009] A papermaking enzyme bidirectional prediction method based on an integrated algorithm, the method comprising:
[0010] Obtaining a multi-enzyme dataset of papermaking enzymes and processing the multi-enzyme dataset; samples in the multi-enzyme dataset include multi-enzyme combination name and ratio, additive type and concentration, reaction condition and enzyme activity; the reaction condition includes incubation pH, incubation temperature and incubation time;
[0011] Training a machine learning model using the processed multi-enzyme dataset, and integrating the trained model using a Bagging algorithm to obtain an integrated model; input data of the machine learning model includes additive concentration and reaction condition, and output data includes enzyme activity;
[0012] According to the combination name and ratio of the multi-enzyme to be tested and the additive type, a corresponding integrated model is selected; according to the additive concentration and reaction condition of the multi-enzyme to be tested, the selected integrated model is used to predict the enzyme activity of the multi-enzyme; according to the predicted enzyme activity and the reaction condition, an optimization algorithm is used to obtain the optimal application condition and enzyme activity of the multi-enzyme;
[0013] A model pool is established using the trained model; the model pool takes additive type as key, and the corresponding value is a model list, and each model in the model list corresponds to a multi-enzyme combination name and ratio;
[0014] According to the additive type of the multi-enzyme to be tested, the enzyme activity of the multi-enzyme is predicted using the model pool; the model corresponding to the highest predicted enzyme activity value is selected from the model pool as a prediction model; the multi-enzyme combination name and ratio corresponding to the prediction model, and the predicted enzyme activity are taken as output results.
[0015] Further, the training of the machine learning model using the processed multi-enzyme dataset comprises:
[0016] First, the selected multi-enzyme combination name and ratio, and additive type are determined, and then a new dataset is obtained by screening the processed multi-enzyme dataset; samples in the new dataset include reaction condition, additive concentration and enzyme activity;
[0017] The machine learning model using different algorithms is trained using the new dataset; input data of the machine learning model includes additive concentration and reaction condition, and output is enzyme activity.
[0018] Further, the integration of the trained model using the Bagging algorithm to obtain an integrated model comprises:
[0019] The machine learning models trained on the same new dataset are integrated using the Bagging algorithm to obtain an integrated model; the integrated model corresponds to the determined multi-enzyme combination name and ratio, and additive type.
[0020] Further, the algorithm used by the machine learning model includes decision tree, random forest, support vector machine regression, Gaussian process regression and gradient boosting regression.
[0021] Further, the model pool is used to predict the enzyme activity of the multi-enzyme according to the type of the coenzyme of the multi-enzyme to be tested, comprising:
[0022] According to the type of the coenzyme of the multi-enzyme to be tested, a corresponding model set is found from the model pool;
[0023] According to the coenzyme concentration and reaction conditions of the multi-enzyme to be tested, the enzyme activity of the multi-enzyme is predicted by using each model in the found model set.
[0024] Further, the optimization algorithm is a genetic algorithm.
[0025] Further, the multi-enzyme data set is processed, comprising:
[0026] The samples in the multi-enzyme data set are sorted according to the multi-enzyme combination names;
[0027] The multi-enzyme combination names in the sorted samples are converted into strings;
[0028] The converted samples are sorted again according to the matching of the multi-enzyme combinations;
[0029] The samples after the two times of sorting are used as the processed multi-enzyme data set.
[0030] Further, the multi-enzyme data set of the papermaking enzyme is obtained, comprising:
[0031] The single enzyme activities of a plurality of papermaking enzymes under different papermaking conditions are obtained, and the single enzyme data is constituted by the papermaking conditions and the single enzyme activities; the papermaking conditions include the type of the papermaking enzyme, the reaction conditions, and the type and concentration of the coenzyme;
[0032] According to the single enzyme data of the papermaking enzyme, the multi-enzyme data set of the papermaking enzyme is obtained;
[0033] Wherein, the enzyme activity of the single enzyme is calculated according to the following formula to obtain the multi-enzyme enzyme activity:
[0034]
[0035] In the formula, a, b, …, n are the enzyme activities of a plurality of single enzymes, a max , b max , …, n max are the highest enzyme activities of the single enzymes corresponding to a, b, …, n, and A, B, …, N are the enzyme matching ratios of the single enzymes corresponding to a, b, …, n.
[0036] The second object of the application can be achieved by adopting the following technical scheme:
[0037] A papermaking enzyme bidirectional prediction system based on an integrated algorithm, the system comprising:
[0038] The data acquisition module is used for acquiring a multi-enzyme data set of papermaking enzymes and processing the multi-enzyme data set; samples in the multi-enzyme data set include multi-enzyme combination names and proportions, additive types and concentrations, reaction conditions and enzyme activities; the reaction conditions include incubation pH, incubation temperature and incubation time;
[0039] The training module is used for training a machine learning model by using the processed multi-enzyme data set;
[0040] The first prediction module is used for integrating the trained model by using a Bagging algorithm to obtain an integrated model; input data of the machine learning model include additive concentrations and reaction conditions, and output data include enzyme activities; according to combination names and proportions of a to-be-tested multi-enzyme and additive types, a corresponding integrated model is selected; according to additive concentrations and reaction conditions of the to-be-tested multi-enzyme, the enzyme activity of the multi-enzyme is predicted by using the selected integrated model; according to the predicted enzyme activity and the reaction conditions, an optimization algorithm is used to obtain the best application conditions and the enzyme activity of the multi-enzyme;
[0041] The second prediction module is used for establishing a model pool by using the trained model; the model pool takes additive types as keys, and corresponding values are model lists, each model in the model list corresponds to multi-enzyme combination names and proportions; according to additive types of a to-be-tested multi-enzyme, the enzyme activity of the multi-enzyme is predicted by using the model pool; a model corresponding to the highest predicted enzyme activity value is selected from the model pool as a prediction model; multi-enzyme combination names and proportions corresponding to the prediction model and the predicted enzyme activity are taken as output results.
[0042] The third object of the present application can be achieved by adopting the following technical solutions:
[0043] A terminal device includes a processor and a memory for storing a program executable by the processor, and when the processor executes the program stored in the memory, the above-mentioned bidirectional prediction method for papermaking enzymes based on an integrated algorithm is realized.
[0044] The fourth object of the present application can be achieved by adopting the following technical solutions:
[0045] A computer-readable storage medium stores a program, and when the program is executed by a processor, the above-mentioned bidirectional prediction method for papermaking enzymes based on an integrated algorithm is realized.
[0046] The present application has the following beneficial effects relative to the prior art:
[0047] 1、The method provided by the present application is simple to operate, time-consuming and cost-saving, and significantly reduces experimental time and cost. By using an integrated model, the advantages of multiple models can be combined to improve prediction performance, reduce model variance, improve model robustness and reduce the risk of overfitting;
[0048] 2. The method provided by the application realizes bidirectional prediction of multi-enzyme application scenarios, can not only predict the most suitable application conditions of multi-enzymes, but also can inversely predict the most suitable enzyme combination under different conditions. Through bidirectional prediction, the performance and optimal combination of multi-enzymes under different conditions can be predicted. The method can be used as a more flexible and efficient prediction tool. BRIEF DESCRIPTION OF DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only some embodiments of the application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of the drawings shown.
[0050] Figure 1 A simple flowchart of the bidirectional prediction method of papermaking enzymes based on the integrated algorithm in the embodiment 1 of the application;
[0051] Figure 2 A detailed flowchart of the bidirectional prediction method of papermaking enzymes based on the integrated algorithm in the embodiment 1 of the application;
[0052] Figure 3 An iteration diagram of the genetic algorithm in the embodiment 2 of the application;
[0053] Figure 4 A prediction result schematic diagram of the enzyme activity and application conditions of multi-enzymes in the embodiment 2 of the application;
[0054] Figure 5 A prediction result schematic diagram of the enzyme combination and enzyme activity of multi-enzymes in the embodiment 2 of the application;
[0055] Figure 6 A structure block diagram of the bidirectional prediction system of papermaking enzymes based on the integrated algorithm in the embodiment 3 of the application;
[0056] Figure 7 A structure block diagram of the computer device in the embodiment 4 of the application. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solutions and advantages of the embodiments of the application more clear, the following will combine the drawings in the embodiments of the application to clearly and completely describe the technical solutions in the embodiments of the application. Obviously, the described embodiments are part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application. It should be understood that the described specific embodiments are only used to explain the application, and are not used to limit the application.
[0058] Embodiment 1:
[0059] As shown in Figure 1 、 2 , the embodiment provides a two-way prediction method for papermaking enzymes based on an integrated algorithm, and specifically comprises the following steps:
[0060] S101, obtaining a multi-enzyme data set of papermaking enzymes.
[0061] Further, step S101 specifically comprises:
[0062] (1) obtaining single-enzyme data of papermaking enzymes.
[0063] The single-enzyme activity of a plurality of papermaking enzymes under different papermaking conditions is obtained through experimental simulation, and the single-enzyme data is composed of papermaking conditions and single-enzyme activity.
[0064] Among them, the papermaking conditions include but are not limited to any of the following: papermaking enzyme types, incubation pH, incubation temperature, incubation time, raw material types, metal ions, chemical additive types and concentrations, etc.
[0065] The chemical additive types include but are not limited to any of the following: cationic polyacrylamide CPAM, anionic polyacrylamide APAM, non-ionic polyacrylamide NPAM, polyethylene oxide PEO, cationic starch CS, oxidized starch OS, polyamide epichlorohydrin resin PAE, alkyl ketene dimer AKD, heavy calcium carbonate HCC, light calcium carbonate LCC, etc.
[0066] The papermaking conditions in the embodiment include papermaking enzyme types, incubation pH, incubation temperature, incubation time, chemical additive types and concentrations. The papermaking enzyme types are cellulase, lignin-degrading enzyme, hemicellulase, amylase, protease, esterolytic enzyme, lipase, pectinase, etc.
[0067] The enzyme activity test method in the embodiment can be selected from colorimetry, reduction method, etc., and can use a spectrophotometer, an enzyme marker, etc. to collect and analyze enzyme activity to obtain enzyme activity data when single enzymes act.
[0068] (2) obtaining a multi-enzyme data set of papermaking enzymes according to the single-enzyme data of papermaking enzymes.
[0069] The enzyme combination mode includes but is not limited to double enzymes and triple enzymes, and can also be four enzymes, five enzymes, etc.
[0070] The enzyme activity of single enzymes is calculated according to the following formula to obtain multi-enzyme enzyme activity:
[0071]
[0072] Wherein, a, b, …, n are the enzyme activities of a plurality of single enzymes, a max , b maxa, b, …, n max a, b, …, n
[0073] The sample in the multi-enzyme data set of papermaking enzymes is composed of reaction conditions, types and concentrations of additives, multi-enzyme combination names (papermaking enzyme types) and ratios, and multi-enzyme enzyme activity. Among them, the reaction conditions include incubation pH, incubation temperature and incubation time.
[0074] S102, processing the multi-enzyme data set.
[0075] The data in the multi-enzyme data set is sorted in a specific order to sort an enzyme list. Since the data in the excel table itself is not convenient for computer recognition and processing, the multi-enzyme combination name is sorted first here, and the sorted multi-enzyme combination name is combined into a string for subsequent computer processing. If the multi-enzyme ratio is provided, the sorted sample will also be sorted according to the multi-enzyme ratio, and the sorted sample will be used as the processed multi-enzyme data set.
[0076] S103, using the processed multi-enzyme data set to train the machine learning model for multi-enzyme prediction, and integrating the trained model to obtain an integrated model.
[0077] First, determine the enzyme combination name, ratio, and additive type to be selected, such as 'BA+XYN', '1:1', 'HCC', and then obtain a new data set through screening. The data set at this time is composed of data of incubation pH, additive concentration, incubation temperature, incubation time and enzyme activity; then use the samples in this data set to train the learning model corresponding to the 5 algorithms respectively, wherein the input of the learning model is the additive concentration, incubation pH, incubation temperature and incubation time, and the output is the enzyme activity; each data set obtains 5 trained models, and then integrates them into an integrated model by bagging method for subsequent optimization algorithm optimization.
[0078] After training, the 5 machine learning models corresponding to each trained data set are integrated into an integrated model using the Bagging integration method. The integrated model is evaluated using the leave-one-out method, and the evaluation indicators are RMSE and MAE.
[0079] The algorithms used by the multi-enzyme prediction machine learning model include decision tree (DT), random forest (RF), support vector regression (SVR), Gaussian process regression (GPR), and gradient boosting regressor (GBR).
[0080] The leave-one-out cross-validation method is a method for evaluating the performance of a machine learning model, which uses the leave-one-out (LOO) method to calculate the error of the model. In each iteration, one sample is used as the test set and the rest as the training set, and then the error of the model on the test set is calculated. This is a cross-validation method that can effectively evaluate the generalization ability of the model.
[0081] Bagging ensemble method is an ensemble learning technique that combines multiple models to improve prediction accuracy. It creates multiple sub-samples from the original training dataset, then trains a base model on each sub-sample, and finally combines the prediction results of these models. In the Bagging method, each sub-sample is obtained by randomly sampling the samples in the original dataset, and the sampling process allows repetition, that is, the same sample may be sampled into different sub-samples. For regression problems, simple averaging is usually used to combine the prediction results. Bagging method can effectively reduce the variance of the model. The calculation method of Bagging method can be expressed as follows:
[0082]
[0083] Where y represents the prediction result of the ensemble model, ai represents the weight of the ith base classifier, and Ci(x) represents the prediction result of the ith base classifier for sample x. The result of the ensemble model is the average of the prediction results of each sub-model.
[0084] In this embodiment, each new dataset corresponds to a specific enzyme combination name, ratio, and type of auxiliary agent, and also corresponds to an ensemble model, that is, a specific enzyme combination name, ratio, and type of auxiliary agent correspond to an ensemble model.
[0085] In this embodiment, the multi-enzyme data set after processing can be divided into multiple new data sets according to the enzyme combination name and ratio, and the type of auxiliary agent, so the ensemble model also has multiple.
[0086] S104, according to the enzyme combination name, ratio and additive type of the multi-enzyme to be tested, select the corresponding integrated model; input the additive concentration and reaction conditions of the multi-enzyme to be tested into the selected integrated model, and output the enzyme activity of the multi-enzyme; according to the enzyme activity and the reaction conditions, an optimization algorithm is used for optimization to obtain the best application conditions and enzyme activity of the multi-enzyme.
[0087] According to the given enzyme combination, ratio and additive type, select the corresponding integrated model; input the additive concentration, incubation pH, incubation temperature and incubation time in the selected integrated model, and predict the enzyme activity of the multi-enzyme; according to the predicted enzyme activity and the incubation pH, incubation temperature and incubation time, an optimization algorithm is used for optimization to obtain the best application conditions and enzyme activity of the multi-enzyme.
[0088] The integrated model obtained in step S103 is used for the optimization algorithm to solve the problem of input multi-enzyme + additive + ratio, and output the best application conditions. The optimization algorithm can be selected as genetic algorithm, Bayesian optimization, simulated annealing, and the preferred one is genetic algorithm. Genetic algorithm is an optimization search algorithm that simulates the principles of natural selection and genetic. The use of genetic algorithm can be different according to the specific application scene, but generally includes the following parts:
[0089] Encoding: convert the solution of the problem into a chromosome (or called gene), each chromosome has a certain number of genes;
[0090] Initial population: randomly generate a group of chromosomes to form an initial population;
[0091] Fitness function: evaluate the goodness of each chromosome according to its fitness. The fitness function is usually defined according to the objective function of the problem;
[0092] Selection operation: select excellent chromosomes according to the fitness function to enter the next generation. Common selection operators include roulette wheel selection, tournament selection, etc.;
[0093] Cross operation: cross the selected chromosomes to produce new offspring. Common cross operators include single-point cross, multi-point cross, etc.;
[0094] Mutation operation: make small random changes to the new generation to increase the diversity of the population;
[0095] Termination condition: when the preset number of iterations is reached or other termination conditions are met, stop the algorithm and output the optimal solution.
[0096] S105, use the trained model to establish a model pool, taking the additive type as the key.
[0097] According to the machine learning model trained in step S103, each machine learning model corresponds to a specific additive type and enzyme combination name and ratio.
[0098] Save all the trained models and establish a model pool with all the trained models, taking the aid species as the key and the corresponding value as the model list. Each model in the model list corresponds to an enzyme combination name and a ratio, that is, each model corresponds to a specific enzyme combination name and ratio.
[0099] The model pool refers to a collection of multiple independently trained models in machine learning or statistical modeling. Each model in the model pool is independent, and they can be models of the same type or different types. The purpose of establishing a model pool is to improve overall performance, robustness and generalization ability by combining the prediction results of multiple models. The model pool is used to store multiple enzyme models under different conditions so that the information of multiple models can be comprehensively utilized when predicting according to the input conditions.
[0100] S106, according to the aid species of the multi-enzyme to be tested, the enzyme activity of the multi-enzyme is predicted by using the model pool; the model corresponding to the highest predicted enzyme activity value is selected from the model pool as the prediction model; the multi-enzyme combination name and ratio corresponding to the prediction model, and the predicted enzyme activity are taken as the output result.
[0101] According to the aid species of the multi-enzyme to be tested, the corresponding model set is found from the model pool; according to the aid concentration and reaction conditions of the multi-enzyme to be tested, the enzyme activity of the multi-enzyme is predicted by using each model in the found model set, and the model corresponding to the highest predicted enzyme activity value is selected as the prediction model; the multi-enzyme combination name and ratio corresponding to the prediction model, and the predicted enzyme activity are taken as the output result.
[0102] Through the given aid species, the corresponding model set in the model pool is found, then each model found is traversed, the reaction conditions and the aid concentration are used as input for prediction, the model corresponding to the highest predicted enzyme activity value is selected as the prediction model; the multi-enzyme combination name and ratio corresponding to the prediction model, and the predicted enzyme activity are output.
[0103] Those skilled in the art can understand that all or part of the steps in the method of implementing the above embodiments can be instructed by a program to relevant hardware, and the corresponding program can be stored in a computer readable storage medium.
[0104] It should be noted that although the method operations of the above embodiments are described in a specific order in the accompanying drawings, this does not require or imply that the operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired result. On the contrary, the depicted steps can change the order of execution. Additionally or alternatively, certain steps can be omitted, combined into one step, and / or divided into multiple steps.
[0105] Embodiment 2:
[0106] This embodiment further illustrates the method provided in Embodiment 1, and is specifically as follows:
[0107] (1) Data set construction: Single-enzyme data of three enzymes, including cellulase EG, xylanase XYN and pectin lyase BA, were obtained through experimental simulation, and the selected papermaking conditions are shown in Table 1. The specific experimental method is as follows:
[0108] (I) Glucose / xylose / galacturonic acid-DNS standard curve drawing: 5 mg / mL (accurate to 0.0001 g, and the actual prepared concentration is calculated) glucose / xylose / galacturonic acid standard solution was prepared for standby, and 8 different concentration solutions of 0-5 mg / mL were diluted with pH buffer solution. After adding DNS solution and boiling, the standard curve was drawn with the content of the standard solution as the horizontal coordinate and the absorbance as the vertical coordinate.
[0109] (II) According to Table 1, the corresponding papermaking simulation solution was configured, different concentrations of additive solution / suspension were prepared with the corresponding pH buffer solution, and the corresponding substrate was prepared with deionized water. The substrate, additive and pH buffer were placed at the corresponding temperature and for the corresponding time for incubation.
[0110] (III) 0.1 mL of appropriately diluted enzyme solution was mixed with the papermaking simulation solution, and the reaction was accurately carried out for 10 min. Immediately after cooling, 0.3 ml of DNS was added, boiled for 5 min, and the absorbance at 540 nm was measured. According to the absorbance, the corresponding enzyme activity was obtained by entering the standard curve.
[0111] Under the optimum pH and temperature of the enzyme, the amount of enzyme required to hydrolyze the substrate to produce 1 μmoL of reducing sugar per minute was defined as one enzyme activity, and the unit was U / mL.
[0112] According to the enzyme activity of single enzyme, the multi-enzyme enzyme activity data was obtained by the following formula:
[0113]
[0114] Wherein, a, b, …, n are the enzyme activities of a plurality of single enzymes, a max , b max , …, n max are the highest enzyme activities of the corresponding single enzymes a, b, …, n, and A, B, …, N are the enzyme ratios of the corresponding single enzymes a, b, …, n.
[0115] The multi-enzyme combinations selected in this embodiment are EG+XYN, EG+BA, XYN+BA and EG+XYN+BA.
[0116] Table 1 Papermaking enzyme simulation conditions
[0117]
[0118] (2) Multi-enzyme data sorting and screening: The data set obtained in step (1) has a total of 642 data. According to a specific order, the multi-enzyme data list is sorted, and if the enzyme ratio is provided, the enzyme ratio is also sorted. Then, according to the selected multi-enzyme combination name, the type and ratio of the auxiliary agent, specific data is extracted from the data file to complete the induction and arrangement of the data set. For example, the multi-enzyme data corresponding to the multi-enzyme combination BA+XYN, the auxiliary agent HCC, and the ratio 1:1 are selected as a new data set.
[0119] (3) Model training and evaluation: The algorithms include decision tree (Decision Tree, DT), random forest (Random Forest, RF), support vector machine regression (Support vector regression, SVR), Gaussian process regression (Gaussian process regression, GPR), and gradient boosting regression (Gradient Boosting Regressor, GBR). The data set is used to train the machine learning model corresponding to different algorithms, and then the Bagging method is used to integrate the trained model, and the average method is used to predict the results of the integrated model. The leave-one-out cross-validation method is used to evaluate the model, and RMSE and MAE are used as evaluation indicators to calculate the model error. The calculation formulas of the above two evaluation indicators are as follows:
[0120]
[0121]
[0122] wherein y i is the true value of the integrated model, is the predicted value.
[0123] The evaluation results of each model are shown in Table 2.
[0124] Table 2 Performance of machine learning model in multi-enzyme prediction
[0125]
[0126]
[0127]
[0128]
[0129] (4) Optimal application conditions and enzyme activity prediction: the integrated model is optimized using a genetic algorithm. The fitness function is selected as the highest enzyme activity predicted by the model, and the inputs are multiple enzymes + additives + ratio, and the outputs are the optimal application conditions and predicted enzyme activity, i.e. pH, additive concentration, time, temperature and enzyme activity. In this example, the enzyme combination is BA + XYN, the additive is HCC, the ratio is 1:1, the genetic algorithm iteration chart is as shown in Figure 3 , and the prediction result is as shown in Figure 4 .
[0130] (5) Multiple enzyme combination and enzyme activity prediction: load the trained model pool and use this model pool to predict the enzyme combination and enzyme activity under given conditions. In this embodiment, the additive is CPAM, the pH is 7, the additive concentration is 0.001%, the incubation temperature is 60°C, and the incubation time is 20 min. The prediction result is as shown in Figure 5 .
[0131] Those skilled in the art can understand that all or part of the steps in the method of the above embodiments can be instructed by a program to relevant hardware, and the corresponding program can be stored in a computer readable storage medium.
[0132] It should be noted that although the method operations of the above embodiments are described in a specific order in the accompanying drawings, this does not require or imply that the operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired result. On the contrary, the steps depicted can change the order of execution. Additionally or alternatively, certain steps can be omitted, combined into one step, and / or divided into multiple steps.
[0133] Embodiment 3:
[0134] As shown in Figure 6 , the present embodiment provides a papermaking enzyme bidirectional prediction system based on an integrated algorithm, which comprises a data acquisition module 601, a training module 602, a first prediction module 603 and a second prediction module 604, wherein:
[0135] The data acquisition module 601 is used for acquiring a multiple enzyme data set of papermaking enzymes and processing the multiple enzyme data set. The samples in the multiple enzyme data set include multiple enzyme combination names and ratios, additive types and concentrations, reaction conditions and enzyme activities. The reaction conditions include incubation pH, incubation temperature and incubation time.
[0136] The training module 602 is used for training a machine learning model using the processed multiple enzyme data set.
[0137] The first prediction module 603 is used to integrate the trained models using the Bagging algorithm to obtain an integrated model; the input data of the machine learning model includes the concentration of auxiliary agents and reaction conditions, and the output data includes enzyme activity; the corresponding integrated model is selected based on the combination name and ratio of the multi-enzyme to be tested and the type of auxiliary agents; the enzyme activity of the multi-enzyme is predicted using the selected integrated model based on the auxiliary agent concentration and reaction conditions of the multi-enzyme to be tested; and the optimal application conditions and enzyme activity of the multi-enzyme are obtained using an optimization algorithm based on the predicted enzyme activity and reaction conditions;
[0138] The second prediction module 604 is used to establish a model pool using the trained model; the model pool uses the type of auxiliary agent as the key, and the corresponding value is a model list, and each model in the model list corresponds to the name and ratio of the multi-enzyme combination; according to the type of auxiliary agent of the multi-enzyme to be tested, the model pool is used to predict the enzyme activity of the multi-enzyme; the model corresponding to the highest predicted enzyme activity value is selected from the model pool as the prediction model; the multi-enzyme combination name and ratio corresponding to the prediction model, as well as the predicted enzyme activity are used as output results.
[0139] The specific implementation of each module in this embodiment can be found in the above-mentioned embodiment 1, and will not be described one by one here; it should be noted that the system provided in this embodiment is only illustrated by the division of the above-mentioned functional modules. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure can be divided into different functional modules to complete all or part of the functions described above.
[0140] Example 4:
[0141] This embodiment provides a terminal device, which can be a computer, such as Figure 7 As shown, it comprises a processor 702, a memory, an input device 703, a display 704, and a network interface 705 connected via a system bus 701. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium 706 and an internal memory 707. The non-volatile storage medium 706 stores an operating system, a computer program, and a database. The internal memory 707 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the processor 702 executes the computer program stored in the memory, the papermaking enzyme bidirectional prediction method based on the integrated algorithm of Example 1 is implemented as follows:
[0142] Obtaining a multi-enzyme data set of papermaking enzymes and processing the multi-enzyme data set; samples in the multi-enzyme data set include multi-enzyme combination names and ratios, adjuvant types and concentrations, reaction conditions, and enzyme activities; the reaction conditions include incubation pH, incubation temperature, and incubation time;
[0143] The machine learning model is trained by using the processed multi-enzyme data set, and the trained model is integrated by using a Bagging algorithm to obtain an integrated model; input data of the machine learning model includes an additive concentration and a reaction condition, and output data includes an enzyme activity;
[0144] According to the combination name and ratio of the to-be-tested multi-enzyme and the type of the additive, a corresponding integrated model is selected; according to the additive concentration and the reaction condition of the to-be-tested multi-enzyme, the selected integrated model is used to predict the enzyme activity of the multi-enzyme; and according to the predicted enzyme activity and the reaction condition, an optimization algorithm is used to obtain the optimal application condition and the enzyme activity of the multi-enzyme;
[0145] A model pool is established by using the trained model; the model pool takes the type of the additive as a key, and a corresponding value is a model list, and each model in the model list corresponds to a multi-enzyme combination name and ratio;
[0146] According to the type of the additive of the to-be-tested multi-enzyme, the model pool is used to predict the enzyme activity of the multi-enzyme; a model corresponding to the highest predicted enzyme activity value is selected from the model pool as a prediction model; and the multi-enzyme combination name and ratio corresponding to the prediction model and the predicted enzyme activity are taken as output results.
[0147] Embodiment 5:
[0148] The embodiment provides a computer readable storage medium, which stores a computer program, and when the computer program is executed by a processor, a papermaking enzyme bidirectional prediction method based on an integrated algorithm in the above embodiment 1 is realized, as follows:
[0149] A multi-enzyme data set of a papermaking enzyme is obtained and processed; samples in the multi-enzyme data set include a multi-enzyme combination name and ratio, a type and concentration of an additive, a reaction condition and an enzyme activity; the reaction condition includes an incubation pH, an incubation temperature and an incubation time;
[0150] The machine learning model is trained by using the processed multi-enzyme data set, and the trained model is integrated by using a Bagging algorithm to obtain an integrated model; input data of the machine learning model includes an additive concentration and a reaction condition, and output data includes an enzyme activity;
[0151] According to the combination name and ratio of the to-be-tested multi-enzyme and the type of the additive, a corresponding integrated model is selected; according to the additive concentration and the reaction condition of the to-be-tested multi-enzyme, the selected integrated model is used to predict the enzyme activity of the multi-enzyme; and according to the predicted enzyme activity and the reaction condition, an optimization algorithm is used to obtain the optimal application condition and the enzyme activity of the multi-enzyme;
[0152] A model pool is established by using the trained model; the model pool takes the type of the additive as a key, and a corresponding value is a model list, and each model in the model list corresponds to a multi-enzyme combination name and ratio;
[0153] According to the type of the co-agent of the multi-enzyme to be measured, the enzyme activity of the multi-enzyme is predicted by using the model pool; a model corresponding to the highest predicted enzyme activity value is selected from the model pool as a prediction model; and the multi-enzyme combination name and ratio corresponding to the prediction model and the predicted enzyme activity are taken as output results.
[0154] It should be noted that the computer readable storage medium of the embodiment can be a computer readable signal medium or a computer readable storage medium, or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection with one or more conductive wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0155] The above is only a preferred embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can make equivalent replacements or changes to the technical scheme and inventive concept of the present application within the scope disclosed by the present application, and such replacements or changes are also within the protection scope of the present application.
Claims
1. A bidirectional prediction method for papermaking enzymes based on an integrated algorithm, characterized in that: The method comprises: Obtaining a multi-enzyme data set of papermaking enzymes and processing the multi-enzyme data set; samples in the multi-enzyme data set include multi-enzyme combination names and ratios, adjuvant types and concentrations, reaction conditions, and enzyme activities; the reaction conditions include incubation pH, incubation temperature, and incubation time; The processed multi-enzyme dataset is used to train a machine learning model, and the trained model is integrated using a bagging algorithm to obtain an integrated model; the input data of the machine learning model includes the concentration of the auxiliary agent and the reaction conditions, and the output data includes the enzyme activity; Select the corresponding integrated model based on the combination name and ratio of the multi-enzyme to be tested and the type of auxiliary agent; use the selected integrated model to predict the enzyme activity of the multi-enzyme based on the auxiliary agent concentration and reaction conditions of the multi-enzyme to be tested; and use the optimization algorithm to obtain the optimal application conditions and enzyme activity of the multi-enzyme based on the predicted enzyme activity and reaction conditions; A model pool is established using the trained model; the model pool uses the type of additive as the key and the corresponding value as a model list, and each model in the model list corresponds to the name and ratio of the multi-enzyme combination; According to the type of additives of the multi-enzyme to be tested, a corresponding model set is found from the model pool; according to the additive concentration and reaction conditions of the multi-enzyme to be tested, each model in the found model set is used to predict the enzyme activity of the multi-enzyme; the model corresponding to the highest predicted enzyme activity value is selected from the model pool as the prediction model; the name and ratio of the multi-enzyme combination corresponding to the prediction model, as well as the predicted enzyme activity, are output as the result; Wherein, the obtaining of a multi-enzyme dataset of papermaking enzymes comprises: Obtaining the single enzyme activity of multiple papermaking enzymes under different papermaking conditions, where the single enzyme data is composed of the papermaking conditions and the single enzyme activity; the papermaking conditions include the type of papermaking enzyme, reaction conditions, and the type and concentration of additives; Based on the single enzyme data of papermaking enzymes, a multi-enzyme dataset of papermaking enzymes was obtained; The enzymatic activity of a single enzyme was calculated according to the following formula to obtain the multi-enzyme activity: In the formula, a, b, ..., n are the enzyme activities of multiple single enzymes, a max 、b max ,…,n max a, b, ..., n is the maximum enzyme activity of a single enzyme, and A, B, ..., N is the enzyme ratio of a, b, ..., n corresponding to a single enzyme.
2. The papermaking enzyme bidirectional prediction method according to claim 1, characterized in that: The method of training a machine learning model using the processed multi-enzyme dataset comprises: First, the name and ratio of the selected multi-enzyme combination and the type of auxiliary agent are determined. Then, a new dataset is generated by filtering the processed multi-enzyme dataset. The samples in the new dataset include reaction conditions, auxiliary agent concentrations, and enzyme activities. The new data set was used to train machine learning models using different algorithms separately; the input data of the machine learning model included additive concentration and reaction conditions, and the output was enzyme activity.
3. The papermaking enzyme bidirectional prediction method according to claim 2, characterized in that: The trained models are integrated using the Bagging algorithm to obtain an integrated model, including: The machine learning models trained on the same new data set are integrated using the Bagging algorithm to obtain an integrated model; the integrated model corresponds to the determined multi-enzyme combination name and ratio, as well as the type of auxiliary agent.
4. The method for bidirectional prediction of papermaking enzymes according to any one of claims 1 to 3, wherein The algorithms used in the machine learning model include decision tree, random forest, support vector machine regression, Gaussian process regression and gradient boosting regression.
5. The papermaking enzyme bidirectional prediction method according to claim 1, characterized in that: The optimization algorithm is a genetic algorithm.
6. The papermaking enzyme bidirectional prediction method according to claim 1, characterized in that: The multi-enzyme dataset is processed, including: Sort the samples in the multi-enzyme dataset according to the multi-enzyme combination name; Convert the multi-enzyme combination names in the sorted samples into strings; The transformed samples are sorted again according to the ratio of the multi-enzyme combination; The samples after two sequencing were used as the processed multi-enzyme dataset.
7. A papermaking enzyme bidirectional prediction system based on an integrated algorithm, characterized in that: The system comprises: A data acquisition module is used to acquire and process a multi-enzyme dataset of papermaking enzymes; the samples in the multi-enzyme dataset include the name and ratio of the multi-enzyme combination, the type and concentration of the adjuvant, the reaction conditions, and the enzyme activity; the reaction conditions include the incubation pH, incubation temperature, and incubation time; A training module for training a machine learning model using the processed multi-enzyme dataset; The first prediction module is used to integrate the trained models using the Bagging algorithm to obtain an integrated model; the input data of the machine learning model includes the concentration of auxiliary agents and reaction conditions, and the output data includes enzyme activity; the corresponding integrated model is selected based on the combination name and ratio of the multi-enzyme to be tested and the type of auxiliary agents; the enzyme activity of the multi-enzyme is predicted using the selected integrated model based on the auxiliary agent concentration and reaction conditions of the multi-enzyme to be tested; and the optimal application conditions and enzyme activity of the multi-enzyme are obtained using an optimization algorithm based on the predicted enzyme activity and reaction conditions; The second prediction module is used to establish a model pool using the trained model; the model pool uses the type of auxiliary agent as the key and the corresponding value is a model list, each model in the model list corresponds to the name and ratio of the multi-enzyme combination; according to the type of auxiliary agent of the multi-enzyme to be tested, the corresponding model set is found from the model pool; according to the auxiliary agent concentration and reaction conditions of the multi-enzyme to be tested, each model in the found model set is used to predict the enzyme activity of the multi-enzyme; the model corresponding to the highest predicted enzyme activity value is selected from the model pool as the prediction model; the name and ratio of the multi-enzyme combination corresponding to the prediction model, as well as the predicted enzyme activity, are output as the result; Wherein, the obtaining of a multi-enzyme dataset of papermaking enzymes comprises: Obtaining the single enzyme activity of multiple papermaking enzymes under different papermaking conditions, where the single enzyme data is composed of the papermaking conditions and the single enzyme activity; the papermaking conditions include the type of papermaking enzyme, reaction conditions, and the type and concentration of additives; Based on the single enzyme data of papermaking enzymes, a multi-enzyme dataset of papermaking enzymes was obtained; The enzymatic activity of a single enzyme was calculated according to the following formula to obtain the multi-enzyme activity: In the formula, a, b, ..., n are the enzyme activities of multiple single enzymes, a max 、b max ,…,n max a, b, ..., n is the maximum enzyme activity of a single enzyme, and A, B, ..., N is the enzyme ratio of a, b, ..., n corresponding to a single enzyme.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the bidirectional prediction method for papermaking enzymes according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Screening method and device, computer equipment, storage medium and program product
CN116994672A
Construction method of high-accuracy straw enzymolysis polysaccharide yield prediction model
CN117034774A