High-throughput fermentation strain screening method based on artificial intelligence

Through the high-throughput fermented bacterial strain screening method based on artificial intelligence, combined with a high-throughput experimental platform and a data-driven closed-loop optimization system, the problems of low efficiency, long cycle and poor targeting in traditional methods are solved, and rapid screening and optimization are achieved, the efficiency and performance of bacterial strain screening are improved, and the reliability and economicality of industrial applications are ensured.

CN120412752APending Publication Date: 2025-08-01ZHENGZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510501326.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Traditional fermented bacterial strain screening methods are inefficient, have long cycles and poor targeting, making them difficult to meet the needs of high-throughput screening in modern industries, and lack of data utilization and lack of systematic database support.

Method used

Using a high-throughput fermentation strain screening method based on artificial intelligence, combined with a high-throughput experimental platform, a data-driven closed-loop optimization system is built, and excellent strains are quickly screened through the AI initial screening model, and the fermentation conditions and genetic transformation are optimized using the AI performance prediction model to build a strain library database to achieve systematic storage and query of data.

Benefits of technology

It significantly improves the efficiency of strain screening and performance optimization accuracy, shortens the screening cycle, enhances the success rate and stability of complex bacterial flora construction, realizes systematic accumulation and reuse of data, and improves the reliability and economic benefits of industrial applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412752A_ABST
    Figure CN120412752A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of fermentation strain screening, and discloses a high-throughput fermentation strain screening method based on artificial intelligence. Comprising the following steps: acquiring standard strains from a public strain preservation center, separating potential functional strains from edible mushroom residues, utilizing a high-throughput fermentation platform, setting different fermentation conditions, monitoring physical parameters in real time through an online sensor, constructing an artificial intelligence model, and quickly screening out dominant strains through real-time data acquisition and model prediction. Designing an optimization experiment scheme by using an AI performance prediction model, and according to a high-throughput experiment result, carrying out multi-round iterative optimization to obtain a single strain or a composite flora with optimal performance; according to the method, actual edible fungus residues are used as fermentation substrates, multiple groups of repeated experiments are set, strains or composite flora with the optimal performance are screened out, and intelligent screening of high-throughput fermentation strains is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of fermentation strain screening. More specifically, the present invention relates to a high-throughput fermentation strain screening method based on artificial intelligence. Background Art

[0002] Fermentation strain screening is an important technology in the field of fermentation engineering, and is widely used in the resource utilization of edible mushroom residues, biofuel production, food fermentation, environmental governance and other fields. Traditional strain screening methods mainly rely on manual experiments and empirical judgments, and usually include steps such as isolating strains from natural environments, preliminarily screening through plate culture, and conducting small-scale fermentation experiments to verify performance. However, these methods have significant deficiencies, which limit their application efficiency and effect in modern industrial production.

[0003] First of all, the efficiency of traditional strain screening methods is low. Due to the diversity and complexity of strains, a large number of strains need to be tested one by one during the screening process, and the experimental workload is huge. For example, in the resource utilization of edible mushroom residues, strains that can efficiently degrade cellulose, lignin or produce specific metabolites need to be screened, and traditional methods usually require an experimental cycle of several months or even years. In addition, traditional methods rely on low-throughput experimental platforms, and only a few strains can be tested in each experiment, which is difficult to meet the high-throughput screening requirements of modern industries.

[0004] Secondly, the screening cycle of traditional methods is long. Strain screening not only needs to test the growth characteristics, metabolic characteristics and substrate utilization ability of strains, but also needs to determine the optimal fermentation conditions through multiple rounds of optimization experiments. Traditional optimization methods usually adopt the trial-and-error method or orthogonal experimental design, with limited variable combinations, which are difficult to comprehensively cover the complex fermentation condition space, resulting in a long optimization process and unsatisfactory results. For example, in the construction of complex microbial communities, traditional methods need to manually test the ratios and interactions of different strains, with strong blindness in experimental design and low success rate.

[0005] Thirdly, the targeting of traditional methods is poor. The performance of fermentation strains is comprehensively affected by substrate characteristics, fermentation conditions and strain genetic backgrounds. Traditional methods lack systematic data analysis and prediction means, and it is difficult to accurately screen for specific substrates (such as Lentinula edodes residues, Pleurotus ostreatus residues) or specific targets (such as high-yield reducing sugars, high enzyme activity). In addition, traditional methods have insufficient evaluation of the stability of microbial communities, and the complex microbial communities screened are prone to composition changes during long-term fermentation, affecting the reliability of industrial applications.

[0006] Finally, traditional methods have deficiencies in data utilization and knowledge accumulation. Experimental data is usually stored in a discrete form, lacking systematic database support, and it is difficult to achieve data sharing and reuse. This not only increases the workload of repeated experiments, but also limits the further optimization and promotion of strain screening technologies.

[0007] In view of the above problems, there is an urgent need for an efficient, rapid and highly targeted method for screening fermentation strains, so as to shorten the screening cycle, improve the screening efficiency, and achieve precise optimization of strain performance. By introducing artificial intelligence technology and combining with a high-throughput experimental platform, the present invention constructs a data-driven closed-loop optimization system, realizing the full-process intelligence from primary screening of strains to performance optimization and then to pilot-scale verification, significantly improving the screening efficiency and strain performance, and providing a new technical solution for the field of screening fermentation strains. Summary of the Invention

[0008] To overcome the above defects of the prior art and to achieve the above object, the present invention provides the following technical solution: A high-throughput fermentation strain screening method based on artificial intelligence, comprising: Step A1 is to construct a strain library: obtain standard strains from a public strain preservation center, isolate potential functional strains from edible mushroom residues, and screen strains with cellulose degradation, lignin degradation or metabolite generation functions from literature or patents, record the growth characteristics, metabolic characteristics and genomic information of the strains, and construct a strain library database that supports data import, query and update; Step A2 is to design small-scale fermentation experiments: use a high-throughput fermentation platform to set different fermentation conditions, including a pH range of 4.0 - 8.0, a temperature range of 25°C - 40°C, substrate concentrations including low, medium and high gradients, oxygen supply including aerobic, microaerobic and anaerobic conditions, and a stirring speed range of 100 - 300 rpm; Step A3 is data collection: Real-time monitor physical parameters through an online sensor, including pH, temperature, dissolved oxygen, cell concentration and stirring speed; Detect biochemical parameters through off-line sampling combined with high performance liquid chromatography, gas chromatography-mass spectrometry or enzyme-linked immunosorbent assay, including cell concentration, substrate consumption, metabolite generation and enzyme activity; Step A4 is to construct an AI primary screening model: Based on the data collected in Step A3, construct an artificial intelligence model. The artificial intelligence model takes the growth characteristics, metabolic characteristics and fermentation conditions of the strains as inputs and the preset screening criteria as outputs. The screening criteria include the synthesis potential of the target metabolite, specific enzyme activity or substrate degradation rate. Through real-time data collection and model prediction, quickly screen out strains with potential excellent performance and eliminate strains that do not meet the requirements; Step A5 is AI-assisted optimization experimental design: For the dominant strains obtained by primary screening in Step A4, use the AI performance prediction model to design an optimization experimental plan. The optimization experimental plan includes optimizing the culture medium formula, adjusting the fermentation conditions and performing genetic engineering transformation. The AI performance prediction model takes the fermentation conditions and strain characteristics as inputs and the strain performance parameters as outputs, and verifies the optimization plan through high-throughput experiments; Step A6 is model feedback and optimization: According to the results of high-throughput experiments, calculate the error between the experimental data and the prediction results of the AI model, feedback the experimental data to the AI model, update the training dataset, and optimize the model parameters through incremental learning or retraining. After multiple rounds of iterative optimization, obtain the single strain or composite microbial community with the optimal performance; Step A7 is pilot-scale fermentation experiment verification: For the strains or composite microbial communities screened in Step A6, conduct fermentation experiments in a 10-100L pilot-scale fermenter, use actual edible mushroom residue as the fermentation substrate, set multiple groups of repeated experiments, collect fermentation data, and verify their performance; Step A8 is performance evaluation and application: Compare the pilot-scale experimental data with the prediction results of the AI model, calculate the accuracy of the model prediction, screen out the strains or composite microbial communities with the optimal performance, and prepare them into fermentation inoculants for the resource utilization of mushroom residue or other fermentation fields.

[0009] Preferably, the strain library database in Step A1 includes: strain number, name, taxonomic status, source, growth characteristics, metabolic characteristics, and genomic information; Among them, the growth characteristics include the optimal pH, temperature range, and oxygen demand; the metabolic characteristics include the substrate utilization spectrum, types and yields of metabolites, and enzyme activities; the genomic information includes sequencing data, functional gene annotation, and metabolic pathway information.

[0010] Preferably, the high-throughput fermentation platform in Step A2 is a micro-bioreactor system, and the volume of each reactor is 1-10 mL, which can simultaneously run dozens to hundreds of fermentation experiments; The fermentation conditions include: the pH range is 4.0-8.0, the temperature range is 25°C-40°C, the substrate concentration includes three gradients of low, medium, and high, the oxygen supply includes aerobic, micro-aerobic, and anaerobic conditions, and the stirring speed range is 100-300 rpm.

[0011] Preferably, the data collection in Step A3 includes: Use an online sensor to record pH, temperature, dissolved oxygen, cell concentration, and stirring speed in real time; Regularly take samples, use high-performance liquid chromatography (HPLC) or gas chromatography-mass spectrometry (GC-MS) to analyze substrate consumption and metabolite production, and use an enzyme marker or spectrophotometer to measure enzyme activities; Among them, the substrates include cellulose, lignin, and protein, the metabolites include reducing sugars, organic acids, and volatile fatty acids, and the enzyme activities include cellulase, ligninase, and protease activities.

[0012] Preferably, the construction of the AI preliminary screening model in Step A4 includes the following steps: Clean the data collected in step 3, remove outliers, and perform standardization or normalization processing, and extract features. The features include the growth characteristics, metabolic characteristics, and environmental characteristics of the strain. Among them, the growth characteristics include the growth rate of the cells and the maximum cell concentration, the metabolic characteristics include the substrate consumption rate, the yield of metabolites, and the enzyme activity, and the environmental characteristics include pH, temperature, dissolved oxygen, and substrate concentration; According to the screening target, select a support vector machine, random forest, convolutional neural network, or gradient boosting tree as the model. Among them, the screening targets include the cellulose degradation rate, lignin degradation rate, metabolite yield, or enzyme activity; Divide the data into a training set, a validation set, and a test set. Among them, the training set accounts for 70%, the validation set accounts for 15%, and the test set accounts for 15%. Use the training set to train the model, optimize the model parameters through the validation set, including the learning rate, tree depth, or number of convolutional layers, and use the test set to evaluate the generalization ability of the model, and calculate the accuracy, mean square error, or coefficient of determination of the model; According to the preset screening criteria, predict the performance of the strain. The screening criteria include the synthesis potential of the target metabolite, specific enzyme activity, or substrate degradation rate. Among them, the target metabolites include reducing sugars, organic acids, or volatile fatty acids, the specific enzyme activities include cellulase activity, ligninase activity, or protease activity, and the substrate degradation rates include cellulose degradation rate or lignin degradation rate. Screen out potential dominant strains and eliminate strains that do not meet the requirements.

[0013] Preferably, the construction of the AI performance prediction model in step A5 includes: Set the input as fermentation conditions and strain characteristics. Among them, the fermentation conditions include pH, temperature, substrate concentration, oxygen supply, and stirring speed, and the strain characteristics include growth characteristics, metabolic characteristics, and genomic information. Among them, the growth characteristics include the optimal pH, temperature range, and oxygen demand, the metabolic characteristics include the substrate utilization spectrum, the types and yields of metabolites, and the enzyme activity, and the genomic information includes cellulose degradation-related genes, lignin degradation-related genes, or metabolic pathway information; Set the output as strain performance parameters, including substrate degradation rate, metabolite yield, and enzyme activity. Among them, the substrate degradation rate includes cellulose degradation rate or lignin degradation rate, the metabolite yield includes reducing sugar yield, organic acid yield, or volatile fatty acid yield, and the enzyme activity includes cellulase activity, ligninase activity, or protease activity; Use Bayesian optimization or genetic algorithms to search for the optimal combination of fermentation conditions to maximize the performance of the strain. Among them, Bayesian optimization constructs a surrogate model through Gaussian process regression, predicts the performance, and selects the next set of experimental conditions. Genetic algorithms iteratively optimize the fermentation conditions through crossover, mutation, and selection operations.

[0014] Preferably, the optimization of the experimental plan in step A5 includes: Predict the effects of different culture medium formulations, including carbon source, nitrogen source, and trace element concentrations, on strain performance through an AI model, and design optimization experiments; Predict the effects of different fermentation conditions, including pH, temperature, and dissolved oxygen, on strain performance through an AI model, and design optimization experiments; Combined with genomic information, predict the effects of key gene modifications on strain performance, and design gene editing experiments, including using CRISPR-Cas9 technology to modify cellulase genes or ligninase genes.

[0015] Preferably, the model feedback and optimization in step A6 include: Compare the high-throughput experimental results with the AI model prediction results, and calculate error metrics, including mean square error, root mean square error, or coefficient of determination, to evaluate the accuracy of model prediction; Feed the high-throughput experimental data back to the AI model, update the training dataset, and optimize the model parameters through incremental learning or retraining. Among them, incremental learning updates the model weights through an online learning algorithm, and retraining optimizes the model hyperparameters, including learning rate, tree depth, or number of convolutional layers, by re-partitioning the training set, validation set, and test set; According to the high-throughput experimental results, adjust the optimization objectives, including increasing substrate degradation rate, metabolite production, or enzyme activity, re-run the AI model, and design the next round of optimization experimental plans until the strain performance reaches the preset target or the performance improvement amplitude is lower than the preset threshold. Among them, the preset target includes a cellulose degradation rate greater than 90%, a metabolite production greater than 15 g / L, or an enzyme activity greater than 100 U / mL, and the preset threshold includes a performance improvement less than 5% in two consecutive rounds of optimization; Select strains with complementary functions from the optimized dominant strains to construct a complex microbial community. The complementary functions include cellulose degradation, lignin degradation, and metabolite production. Use the AI model to predict the effects of different strain ratios on fermentation performance, and design high-throughput experiments to verify the optimal ratio. Among them, the strain ratio is designed through orthogonal experiments or response surface methods, including proportional combinations such as 1:1, 2:1, and 1:2; Evaluate the stability of the complex microbial community during long-term fermentation through continuous passage experiments. The stability is analyzed by high-throughput sequencing technology to analyze the change rate of the microbial community composition, and eliminate the microbial community combinations with a change rate greater than 20%.

[0016] Preferably, the pilot fermentation experiment in step A7 includes: Use a 10 - 100 L pilot-scale fermenter to simulate industrial fermentation conditions; According to the optimization results of the high-throughput experiment, set the optimal pH, temperature, dissolved oxygen, stirring speed, and substrate concentration; Use actual edible mushroom residues as the fermentation substrate, including Lentinula edodes residues or Pleurotus ostreatus residues; Real-time monitor pH, temperature, dissolved oxygen, cell concentration, and stirring speed, and perform off-line detection of substrate consumption, metabolite production, enzyme activity, and microbial community composition; Calculate the substrate degradation rate, metabolite yield, fermentation efficiency, and microbial community stability, and compare the experimental results with the predicted results of the AI model.

[0017] Preferably, the performance evaluation and application in step A8 include: Compare the performance of the optimized strain or composite microbial community, and draw performance charts, including substrate degradation rate, metabolite yield, fermentation efficiency, and stability; Evaluate the culture cost, substrate utilization efficiency, and product value of the strain or microbial community; Upload the pilot experiment data, performance data of the preferred strain, and fermentation process parameters to the strain library database; Prepare the preferred strain or composite microbial community into a fermentation inoculant, including freeze-dried powder or liquid inoculant, conduct industrial tests, and transfer the technology to relevant enterprises or research institutions.

[0018] The technical effects and advantages of the high-throughput fermentation strain screening method based on artificial intelligence of the present invention: 1. Significantly improve the strain screening efficiency: By introducing AI technology and combining with a high-throughput experimental platform, the technical solution realizes the rapid primary screening and performance optimization of strains, significantly shortening the screening cycle. Through the AI primary screening model, potential dominant strains can be screened out from hundreds to thousands of strains in a short time. Compared with the traditional manual screening method, the efficiency is increased several times. The optimal fermentation condition combination can be quickly searched through the AI performance prediction model, reducing the number of experiments required by the traditional trial-and-error method.

[0019] 2. Improve the optimization accuracy of strain performance: Through the performance prediction of the AI model and data-driven closed-loop optimization, the technical solution can accurately optimize the fermentation conditions, culture medium formula, and gene modification plan of the strain, significantly improving the strain performance. The AI performance prediction model predicts the substrate degradation rate, metabolite yield, and enzyme activity, and the optimization results are more accurate. By predicting the role of key genes through AI and combining with the CRISPR-Cas9 technology, the cellulose degradation rate or metabolite yield of the strain is increased.

[0020] 3. Enhance the success rate and stability of composite microbial community construction: By predicting the strain ratio and microbial community stability through the AI model, the technical solution significantly improves the success rate of composite microbial community construction and ensures its stability in long-term fermentation. Use AI to assist in the design of orthogonal experiments or response surface method for verifying the optimal ratio of composite microbial community, reducing the blindness of traditional manual experiments. Analyze the change rate of microbial community composition through high-throughput sequencing technology to ensure the stability of the screened microbial community in continuous subculture.

[0021] 4. Systematic accumulation and reuse of data are achieved: By constructing a strain library database, the technical solution realizes the systematic storage, query, and update of experimental data, providing data support for subsequent research and applications. Uploading pilot experiment data and optimized strain performance data to the database provides a basis for subsequent technology development and promotion. Through pilot verification and application promotion, the technical solution ensures the industrial applicability of the screening results, reduces production costs, and improves economic benefits. Description of the Drawings

[0022] Figure 1 It is a step schematic diagram of a high-throughput fermentation strain screening method based on artificial intelligence according to the present invention. Detailed Embodiments

[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] Embodiment 1 Please refer to Figure 1 As shown, a high-throughput fermentation strain screening method based on artificial intelligence in this embodiment includes: Step A1: Construct a strain library; Strain sources: Public strain preservation centers include but are not limited to ATCC (American Type Culture Collection), CICC (China Center of Industrial Culture Collection), or DSMZ (German Collection of Microorganisms and Cell Cultures). When isolating strains from edible mushroom residues, the gradient dilution method is used. The residue samples (such as Lentinula edodes residues, Pleurotus ostreatus residues) are diluted and then spread on selective media (such as media containing cellulose or lignin), and cultured for 48 - 72 hours at 25°C - 37°C under aerobic or anaerobic conditions to isolate and purify the strains.

[0025] Characteristic records: Growth characteristics include optimal pH (determined by pH gradient incubation), temperature range (determined by temperature gradient incubation), and oxygen requirement (determined by anaerobic flasks or aerobic incubation). Metabolic characteristics include substrate utilization profile (determined by substrate consumption experiments, such as cellulose, lignin, and protein consumption), metabolite types and production (determined by HPLC or GC-MS, such as the production of reducing sugars, organic acids, and volatile fatty acids), and enzyme activity (determined by microplate reader, such as cellulase, ligninase, and protease activity). Genomic information is obtained through whole-genome sequencing, using platforms including but not limited to Illumina or PacBio. After assembly and annotation of the sequencing data, genes related to cellulose degradation (such as the cel gene), genes related to lignin degradation (such as the lcc gene), and metabolic pathway information are analyzed.

[0026] Database Construction: Build a strain library database using a relational database management system (e.g., MySQL) or a non-relational database (e.g., MongoDB). Fields in the database include strain number, name, taxonomic status, origin, growth characteristics, metabolic characteristics, and genomic information. The database supports data import, query, and update via SQL query language or API interfaces, and includes data backup and permission management capabilities.

[0027] Step A2: Design a small-scale fermentation experiment; High-throughput fermentation platforms utilize microbioreactor systems (such as the BioLector or Ambr systems), each with a volume of 1–10 mL, capable of running 48–96 fermentation experiments simultaneously. The reactors are equipped with online sensors for real-time monitoring of pH, temperature, dissolved oxygen, bacterial concentration, and agitation speed.

[0028] Fermentation condition settings: pH is controlled by an automatic titration system, ranging from 4.0 to 8.0 with a gradient interval of 0.5; temperature is controlled by a constant temperature water bath or heating module, ranging from 25°C to 40°C with a gradient interval of 5°C; substrate concentration is divided into three gradients: low (1% w / v), medium (5% w / v), and high (10% w / v), and the substrates include cellulose, lignin, or protein; oxygen supply is controlled by a gas mixer, aerobic conditions are dissolved oxygen >20%, microaerobic conditions are dissolved oxygen 5%-20%, and anaerobic conditions are dissolved oxygen <1%; stirring speed is controlled by a magnetic stirrer or mechanical stirrer, ranging from 100-300 rpm, with a gradient interval of 50 rpm.

[0029] Experimental Design: Full factorial experimental design or response surface methodology (RSM) was used to cover the diversity of strains in the strain library and the complexity of fermentation conditions. Each group of experiments was repeated 3-5 times to ensure data reliability.

[0030] Step A3: Data collection; Online monitoring: Online sensors include a pH electrode (measuring range 4.0-9.0, accuracy ±0.01), a temperature probe (measuring range 0°C-100°C, accuracy ±0.1°C), a dissolved oxygen electrode (measuring range 0%-100%, accuracy ±0.1%), an optical density probe (measuring wavelength 600nm, for determining bacterial concentration, accuracy ±0.01 OD units), and a rotation speed sensor (measuring range 0-1000 rpm, accuracy ±1 rpm). Data is recorded using a data acquisition system (such as LabVIEW or dedicated software) with a sampling frequency of once per minute.

[0031] Offline analysis: Samples were collected every 6-12 hours, with a volume of 0.5-1 mL. Substrate consumption was determined by HPLC using a C18 column, a methanol-water mobile phase, and a detection wavelength of 210 nm. Residual cellulose, lignin, or protein was determined. Metabolite production was determined by GC-MS using a DB-5MS column, helium carrier gas, and a mass spectrometry scan range of 50-500 m / z. Reducing sugars (e.g., glucose), organic acids (e.g., acetic acid, lactic acid), and volatile fatty acids (e.g., butyric acid) were determined. Enzyme activity was determined using a microplate reader. Cellulase activity was determined using the DNS method (3,5-dinitrosalicylic acid method), ligninase activity was determined using the ABTS method (2,2'-azino-bis(3-ethylbenzothiazoline-6-sulfonic acid) method), and protease activity was determined using the Folin-phenol reagent method. Detection wavelengths were 540 nm, 420 nm, and 660 nm, respectively.

[0032] Step A4: Build an AI preliminary screening model; Data preprocessing: Clean the data collected in step A3 to remove outliers (e.g., erroneous data caused by sensor failure). Use the Z-score method to eliminate data that deviates by more than three standard deviations from the mean. Normalize the data to a range of 0–1. Extract features, including growth characteristics (e.g., bacterial growth rate μ, maximum bacterial concentration OD_max), metabolic characteristics (e.g., substrate consumption rate (calculated as S_rate = (S_t1 - S_t2) / (t2 - t1), metabolite production P, enzyme activity E), and environmental characteristics (e.g., pH, temperature, dissolved oxygen, and substrate concentration).

[0033] Model Selection: Support vector machines (SVM, with radial basis function (RBF) kernel), random forests (RF, with 100 trees), convolutional neural networks (CNN, with 3 convolutional layers and 2 fully connected layers), or gradient boosting trees (XGBoost, with a tree depth of 6 and a learning rate of 0.1) were selected as models based on the screening objectives. Screening objectives included cellulose degradation rate, lignin degradation rate, metabolite yield (e.g., reducing sugar yield, in g / L), or enzyme activity (e.g., cellulase activity, in U / mL).

[0034] Model training and validation: The data is divided into a training set (70%), a validation set (15%), and a test set (15%). The model is trained using the training set, and the model parameters are optimized through the validation set (such as the C parameter of SVM, the number of trees of RF, the learning rate of CNN, the tree depth of XGBoost). The generalization ability of the model is evaluated using the test set, and the accuracy of the model (for classification tasks, the formula is Accuracy = number of correctly predicted samples / total number of samples), mean squared error (MSE for regression tasks), or coefficient of determination (R²) is calculated to evaluate the model performance. Early stopping strategy (earlystopping) is adopted for model training, and training stops when the validation set error does not decrease for 10 consecutive rounds to avoid overfitting.

[0035] Preliminary screening application: According to the preset screening criteria, the performance of strains is predicted. The screening criteria include the synthesis potential of target metabolites (such as reducing sugar production > 10 g / L, organic acid production > 5 g / L, volatile fatty acid production > 3 g / L), specific enzyme activities (such as cellulase activity > 50 U / mL, ligninase activity > 30 U / mL, protease activity > 20 U / mL), or substrate degradation rates (such as cellulose degradation rate > 70%, lignin degradation rate > 50%). Through model prediction, potential dominant strains (such as the top 10% of strains in performance scores) are screened out, and strains that do not meet the requirements (such as strains with performance scores below 50%) are eliminated. The model prediction results are plotted as a performance distribution map using a visualization tool (such as the Matplotlib library in Python) to intuitively show the performance differences of strains.

[0036] Step A5: AI-assisted optimization of experimental design; Construction of the AI performance prediction model: Inputs: Fermentation conditions include pH (range 4.0 - 8.0, interval 0.5), temperature (range 25°C - 40°C, interval 5°C), substrate concentration (range 1% - 10% w / v, interval 1%), oxygen supply (aerobic, microaerobic, anaerobic), and stirring speed (range 100 - 300 rpm, interval 50 rpm). Strain characteristics include growth characteristics (such as the optimal pH, temperature range, oxygen demand, determined by pH gradient, temperature gradient, and anaerobic bottle experiments respectively), metabolic characteristics (such as substrate utilization profile, types and yields of metabolites, enzyme activities, determined by HPLC, GC-MS, and microplate reader respectively), and genomic information (such as cellulose degradation-related gene cel5A, lignin degradation-related gene lcc1, metabolic pathway information, obtained through genome sequencing and KEGG database analysis).

[0037] Output: Strain performance parameters, including substrate degradation rate (such as cellulose degradation rate), metabolite production (such as reducing sugar production, unit: g / L), and enzyme activity (such as cellulase activity, unit: U / mL).

[0038] Optimization algorithms: Use Bayesian optimization (construct a surrogate model through Gaussian process regression, predict performance, and select the next set of experimental conditions, with the optimization goal of maximizing performance parameters) or genetic algorithm (iterate to optimize fermentation conditions with a crossover rate of 0.8, a mutation rate of 0.1, and a roulette wheel selection strategy, and the number of iterations is 50).

[0039] Optimization of experimental scheme: Optimization of culture medium: Predict the impact of different culture medium formulations on strain performance through an AI model. The components of the culture medium include carbon sources (such as glucose, cellulose, concentration range: 1% - 5% w / v), nitrogen sources (such as yeast extract, ammonium nitrate, concentration range: 0.5% - 2% w / v), and trace elements (such as Fe²⁺, Mn²⁺, concentration range: 0.01% - 0.1% w / v). Design optimization experiments and use the response surface method (RSM) or orthogonal experiments to verify the optimal formulation.

[0040] Optimization of fermentation conditions: Predict the impact of different fermentation conditions on strain performance through an AI model. The fermentation conditions include pH, temperature, substrate concentration, oxygen supply, and stirring speed. Design optimization experiments and use full factorial experiments or Box - Behnken design to verify the optimal conditions.

[0041] Genetic engineering modification: Combine genomic information to predict the impact of key gene modifications on strain performance. The key genes include cellulase genes (such as cel5A, up - regulating expression to increase cellulose degradation rate) and ligninase genes (such as lcc1, up - regulating expression to increase lignin degradation rate). Design gene editing experiments, use CRISPR - Cas9 technology to modify genes. The specific operations include designing gRNA (targeting the gene promoter region), constructing a Cas9 expression vector, transferring it into the strain by electroporation, and screening positive transformants.

[0042] Step A6: Model feedback and optimization Model evaluation: Compare the results of high - throughput experiments with the predictions of the AI model, and calculate error metrics, including mean squared error (MSE), root mean squared error (RMSE), and coefficient of determination (R²). If the error metrics exceed the preset thresholds (such as MSE > 0.1 or R² < 0.8), trigger model correction.

[0043] Model correction: Feed high-throughput experimental data back to the AI model to update the training dataset. Incremental learning updates the model weights through online learning algorithms (such as online gradient descent), and re-training optimizes the model hyperparameters (such as the C parameter of SVM, the number of trees in RF, the learning rate of CNN, the tree depth of XGBoost) by re-partitioning the training set (70%), validation set (15%), and test set (15%).

[0044] Iterative optimization: According to the high-throughput experimental results, adjust the optimization objectives (such as increasing the cellulose degradation rate to >90%, metabolite production to >15 g / L, enzyme activity to >100 U / mL), re-run the AI model, and design the next round of optimization experiment plans. Iteratively optimize until the strain performance reaches the preset goal or the performance improvement amplitude is lower than the preset threshold (such as the performance improvement <5% in two consecutive rounds of optimization).

[0045] Construction of complex microbial communities: Select strains with complementary functions (such as cellulose-degrading bacteria, lignin-degrading bacteria, metabolite-producing bacteria) from the optimized dominant strains to construct complex microbial communities. Use the AI model to predict the effects of different strain ratios on fermentation performance, design high-throughput experiments to verify the optimal ratio, and the ratio scheme is designed by orthogonal experiments or response surface methodology (RSM), including ratio combinations such as 1:1, 2:1, 1:2, etc.

[0046] Assessment of microbial community stability: Evaluate the stability of complex microbial communities during long-term fermentation through continuous subculture experiments (the number of subcultures is 10 times, and each culture lasts for 48 hours). The stability analyzes the change rate of the microbial community composition through high-throughput sequencing technology (such as 16S rRNA or ITS sequencing, the sequencing platform is Illumina MiSeq, and the data analysis software is QIIME), and eliminate the microbial community combinations with a change rate >20%.

[0047] Step A7: Verification by pilot-scale fermentation experiments; Fermentation equipment: Use 10 - 100 L pilot-scale fermenters (such as Sartorius Biostat or New Brunswick BioFlo), equipped with industrial-grade online sensors (pH electrodes, temperature probes, dissolved oxygen electrodes, rotational speed sensors) and an automatic control system (such as a PID controller) to simulate industrial fermentation conditions.

[0048] Fermentation conditions: According to the optimization results of high-throughput experiments, set the optimal pH (such as 5.5 ± 0.1, controlled by automatic titration), temperature (such as 30°C ± 0.1, controlled by jacket heating), dissolved oxygen (such as dissolved oxygen >20% under aerobic conditions, controlled by a gas mixer), stirring speed (such as 200 rpm ± 1, controlled by a variable frequency motor), and substrate concentration (such as 5% w / v cellulose).

[0049] Substrate selection: Use actual edible mushroom residues as the fermentation substrate, including Lentinula edodes residues (cellulose content is about 40%, lignin content is about 20%) or Pleurotus ostreatus residues (cellulose content is about 35%, lignin content is about 15%). The pretreatment of the residues includes crushing (particle size < 1 mm), sterilization (autoclaving at 121 °C for 20 minutes), and drying (drying at 60 °C until the moisture content < 10%).

[0050] Data collection: Real-time monitoring: Use industrial-grade online sensors to record pH (measurement range 4.0 - 9.0, accuracy ±0.01), temperature (measurement range 0 °C - 100 °C, accuracy ±0.1 °C), dissolved oxygen (measurement range 0% - 100%, accuracy ±0.1%), cell concentration (determined by dry weight method, accuracy ±0.01 g / L), and stirring speed (measurement range 0 - 1000 rpm, accuracy ±1 rpm) in real time. The data is recorded through a data acquisition system (such as SCADA or dedicated software), and the sampling frequency is 1 time per minute.

[0051] Offline detection: Sample once every 12 hours, and the sample volume is 10 - 20 mL. The substrate consumption is determined by high-performance liquid chromatography (HPLC). The chromatographic column is a C18 column, the mobile phase is a methanol-water mixture (volume ratio 1:1), the detection wavelength is 210 nm, and the residual amounts of cellulose and lignin are determined. The generation of metabolites is determined by gas chromatography-mass spectrometry (GC-MS). The gas chromatography column is a DB-5MS column, the carrier gas is helium, the mass spectrometry scanning range is 50 - 500 m / z, and the contents of reducing sugars (such as glucose), organic acids (such as acetic acid, lactic acid), and volatile fatty acids (such as butyric acid) are determined. The enzyme activity is determined by an enzyme-labeling instrument. The cellulase activity is determined by the DNS method (detection wavelength 540 nm), the ligninase activity is determined by the ABTS method (detection wavelength 420 nm), and the protease activity is determined by the Folin-phenol reagent method (detection wavelength 660 nm). The microbial community composition is analyzed by high-throughput sequencing technology. The sequencing platform is Illumina MiSeq, the data analysis software is QIIME, and the microbial community diversity index (such as Shannon index) and composition change rate are calculated.

[0052] Experiment repetition: Set 3 - 5 parallel repetitions for each group of experiments, and the fermentation cycle is 72 - 120 hours to ensure the reliability of the data.

[0053] Performance evaluation: Calculate the substrate degradation rate, metabolite yield (unit: g / L), fermentation efficiency, and microbial community stability (evaluated by the microbial community composition change rate, and the change rate < 20% is considered stable). Compare the experimental results with the predicted results of the AI model, and calculate the accuracy of the model prediction (evaluated by MSE, RMSE, or R²).

[0054] Step A8: Performance evaluation and application; Performance comparison: Compare the performance of the optimized strains or composite microbial communities, and draw performance charts, including substrate degradation rate, metabolite production, fermentation efficiency, and stability. Chart types include bar charts (comparing the performance of different strains), line charts (showing the change of performance over time), and radar charts (comprehensively evaluating multiple performance indicators). Chart drawing tools include the Matplotlib library in Python or Excel.

[0055] Economic evaluation: Evaluate the cultivation cost of strains or microbial communities (including medium cost, fermentation time cost, equipment depreciation cost, in yuan / L), substrate utilization efficiency, and product value (calculated according to the market price of metabolites, such as the market price of reducing sugar is 5000 yuan / ton).

[0056] Database update: Upload the pilot experiment data, performance data of selected strains, and fermentation process parameters to the strain library database. The uploaded data includes strain numbers, fermentation conditions (pH, temperature, dissolved oxygen, etc.), performance parameters (substrate degradation rate, metabolite production, etc.), and economic indicators (cultivation cost, product value, etc.). The database supports data export functions (such as CSV or JSON formats) for subsequent analysis and sharing.

[0057] Application and promotion: Prepare the selected strains or composite microbial communities into fermentation inoculants. The preparation methods include freeze-drying method (after concentrating the bacterial solution, adding protective agents such as glycerol, pre-freezing at -80°C and then vacuum freeze-drying to make freeze-dried powder) or liquid inoculant (directly filling the bacterial solution and adding stabilizers such as Tween80, storing at 4°C). Conduct industrial tests to verify the performance of the inoculant in a fermentation tank with a scale of more than 1000L. The fermentation cycle is 120 - 168 hours, and the verification indicators include substrate degradation rate, metabolite production, and inoculant stability. Transfer the technology to relevant enterprises or research institutions. The transferred content includes selected strains, fermentation process parameters, AI models, and databases. The transfer methods include technology licensing or technology transfer.

[0058] Example 2 Please refer to Figure 1 As shown, for the parts not described in detail in this example, refer to the description in Example 1. Provide a high-throughput fermentation strain screening system based on artificial intelligence, including: Screen a single cellulose-degrading strain based on AI, aiming to screen out a single strain with high efficiency in degrading cellulose from Lentinula edodes residue for the resource utilization of the residue.

[0059] Operation steps: Step A1: Construct a strain library; Isolate strains from Lentinula edodes residue. Take 10g of the residue sample, add 90mL of sterile water, shake and mix well, and then dilute it step by step to 10 -6, coated on a cellulose-containing selective medium (cellulose 1% w / v, yeast extract 0.5% w / v, agar 2% w / v), cultured at 30 °C for 48 hours, and the strains were isolated and purified. The strain types were identified by 16S rRNA sequencing, and 10 potential functional strains were obtained, including Trichoderma reesei, Bacillus subtilis, etc. Record the growth characteristics of the strains (optimum pH 5.5, optimum temperature 30 °C, aerobic), metabolic characteristics (cellulose degradation rate 50% - 70%, reducing sugar yield 5 - 10 g / L, cellulase activity 20 - 50 U / mL), and genomic information (sequencing platform is Illumina, analyze the cellulase gene cel5A). Construct a strain library database, and the fields include strain number, name, taxonomic status, source, growth characteristics, metabolic characteristics, and genomic information.

[0060] Step A2: Design a small-scale fermentation experiment; Use the BioLector micro-bioreactor system (reactor volume 2 mL, running 48 experiments), and set the fermentation conditions: pH 4.0 - 8.0 (interval 0.5), temperature 25 °C - 40 °C (interval 5 °C), substrate concentration 1% - 10% w / v (interval 1%), oxygen supply aerobic (dissolved oxygen > 20%), stirring speed 100 - 300 rpm (interval 50 rpm). The substrate is cellulose, and the fermentation period is 48 hours.

[0061] Step A3: Data collection; Real-time monitor pH, temperature, dissolved oxygen, cell concentration, and stirring speed through an online sensor, and record the data once per minute. Take samples once every 12 hours, use HPLC to measure the cellulose residue (C18 column, mobile phase methanol - water 1:1, detection wavelength 210 nm), use GC-MS to measure the reducing sugar yield (DB-5MS column, carrier gas helium, scanning range 50 - 500 m / z), and use an enzyme-labeling instrument to measure the cellulase activity (DNS method, detection wavelength 540 nm).

[0062] Step A4: Construct an AI primary screening model; Clean the collected data (remove data deviating from the mean by 3 times the standard deviation), and perform standardization processing (scale to the range of 0 - 1). Extract features, including growth features (bacterial growth rate μ, maximum bacterial concentration OD_max), metabolic features (cellulose consumption rate, reducing sugar yield, cellulase activity), and environmental features (pH, temperature, dissolved oxygen, substrate concentration). Select a random forest model (number of trees 100), with 70% for the training set, 15% for the validation set, and 15% for the test set. Optimize the parameters (tree depth 6), and evaluate the model performance (MSE = 0.05, R² = 0.92). According to the screening criteria (cellulose degradation rate > 70%, reducing sugar yield > 10 g / L, cellulase activity > 50 U / mL), screen out 3 dominant strains (numbered T1, T2, T3).

[0063] Step A5: AI-assisted optimization of experimental design; For the 3 dominant strains (numbered T1, T2, T3) obtained from the primary screening, use the AI performance prediction model to design an optimized experimental plan. The model inputs include fermentation conditions (pH 4.0 - 8.0, temperature 25°C - 40°C, substrate concentration 1% - 10% w / v, aerobic oxygen supply, stirring speed 100 - 300 rpm) and strain characteristics (optimum pH 5.5, optimum temperature 30°C, expression level of cellulase gene cel5A). The model output is the strain performance parameters (cellulose degradation rate, reducing sugar yield, cellulase activity). Use Bayesian optimization to search for the optimal fermentation conditions, with the optimization goal of maximizing the cellulose degradation rate. The optimization results are: the optimal conditions for strain T1 are pH 5.5, temperature 30°C, substrate concentration 5% w / v, and stirring speed 200 rpm. Verify the optimized conditions in the BioLector system, with a fermentation cycle of 48 hours and set 5 parallel replicates. The experimental results are: the cellulose degradation rate of strain T1 is increased to 92.3% (70.5% before optimization), the reducing sugar yield is increased to 15.6 g / L (10.2 g / L before optimization), and the cellulase activity is increased to 78.5 U / mL (52.3 U / mL before optimization).

[0064] Step A6: Model feedback and optimization; Compare the experimental results with the AI model prediction results, calculate the error metrics (MSE = 0.03, R² = 0.95), and the model prediction accuracy is relatively high. Feed the experimental data back to the AI model, update the training data set, and optimize the model parameters through incremental learning (adjust the number of trees in the random forest to 120). Rerun the model to predict the further optimization potential of strain T1, and the results show that the performance improvement is less than 5%, so terminate the iterative optimization.

[0065] Step A7: Verification by pilot-scale fermentation experiment; The strain T1 was verified in a 50 L pilot-scale fermenter, which was a Sartorius Biostat equipped with industrial-grade sensors and a PID controller. Lentinula edodes residue was used as the substrate (cellulose content 40%), and the substrate concentration after pretreatment was 5% w / v. The fermentation conditions were pH 5.5, temperature 30 °C, dissolved oxygen > 20%, and stirring speed 200 rpm. The fermentation cycle was 72 hours, and 3 parallel replicates were set. Samples were taken every 12 hours to measure the cellulose degradation rate, reducing sugar yield, and cellulase activity. The experimental results were: cellulose degradation rate 91.8%, reducing sugar yield 15.2 g / L, and cellulase activity 77.9 U / mL, which were consistent with the high-throughput experimental results. Comparing the experimental results with the AI model prediction results, MSE = 0.02 and R² = 0.96, indicating high model prediction accuracy.

[0066] Step A8: Performance evaluation and application; The performance of strain T1 was compared with that of the control strain (unoptimized Trichoderma reesei standard strain), and a bar chart was plotted. The results showed that the cellulose degradation rate of strain T1 increased by 30.2%, the reducing sugar yield increased by 48.5%, and the cellulase activity increased by 49.6%. Economic evaluation showed that the cultivation cost was 12 yuan / L, the product value (reducing sugar) was 23 yuan / L, and BCR = 1.92, indicating economic feasibility. Strain T1 was prepared into a freeze-dried bacterial agent (10% glycerol was added after the bacterial solution was concentrated, pre-frozen at -80 °C, and then vacuum freeze-dried), and a 1000 L industrial test was carried out. The results showed a cellulose degradation rate of 90.5% and a reducing sugar yield of 14.8 g / L, with high stability of the bacterial agent. The technology was transferred to a biotechnology company, and the transferred content included strain T1, fermentation process parameters, and the AI model.

[0067] Example 3 Please refer to Figure 1 As shown, for the parts not described in detail in this example, please refer to the descriptions in Examples 1 and 2. A high-throughput fermentation strain screening system based on artificial intelligence is provided, including: Construct a composite microbial community based on AI with the aim of screening a composite microbial community that can efficiently degrade cellulose and lignin from Pleurotus ostreatus residue for the resource utilization of the residue.

[0068] Operation steps: Step A1: Construct a strain library; The species of strains were identified by 16S rRNA or ITS sequencing, and 15 potential functional strains were obtained, including cellulose-degrading bacteria (such as Trichoderma reesei, Bacillus subtilis), lignin-degrading bacteria (such as Phanerochaete chrysosporium, Pseudomonas putida), and metabolite-producing bacteria (such as Lactobacillus plantarum).

[0069] Record the growth characteristics of the strains (optimum pH 4.5 - 6.0, optimum temperature 28°C - 35°C, aerobic or anaerobic), metabolic characteristics (cellulose degradation rate 40% - 70%, lignin degradation rate 30% - 60%, reducing sugar yield 3 - 8 g / L, organic acid yield 2 - 5 g / L, cellulase activity 15 - 40 U / mL, ligninase activity 10 - 30 U / mL), and genomic information (sequencing platform is Illumina, analyze cellulase gene cel5A and ligninase gene lcc1).

[0070] Construct a strain library database with fields including strain number, name, taxonomic status, source, growth characteristics, metabolic characteristics, and genomic information.

[0071] Step A2: Design a small-scale fermentation experiment; Use a BioLector micro-bioreactor system (reactor volume 2 mL, running 96 experiments), and set the fermentation conditions: pH 4.0 - 8.0 (interval 0.5), temperature 25°C - 40°C (interval 5°C), substrate concentration 1% - 10% w / v (interval 1%), oxygen supply aerobic (dissolved oxygen > 20%) or micro-aerobic (dissolved oxygen 5% - 20%), stirring speed 100 - 300 rpm (interval 50 rpm).

[0072] The substrate is Pleurotus ostreatus residue (cellulose content 35%, lignin content 15%), and the fermentation period is 48 hours.

[0073] Step A3: Data collection; Real-time monitor pH, temperature, dissolved oxygen, cell concentration, and stirring speed through an online sensor, and record data once per minute.

[0074] Take samples once every 12 hours. Use HPLC to determine the residual amounts of cellulose and lignin (C18 column, mobile phase methanol-water 1:1, detection wavelength 210 nm), use GC-MS to determine the yields of reducing sugar and organic acid (DB-5MS column, carrier gas helium, scanning range 50 - 500 m / z), and use an enzyme-labeling instrument to determine the activities of cellulase and ligninase (DNS method, detection wavelength 540 nm, ABTS method, detection wavelength 420 nm).

[0075] Step A4: Build the initial AI screening model; Clean the collected data (remove data deviating from the mean by 3 times the standard deviation) and perform standardization processing (scale to the range of 0 - 1).

[0076] Extract features, including growth features (bacterial growth rate μ, maximum bacterial concentration OD_max), metabolic features (cellulose consumption rate, lignin consumption rate, reducing sugar yield, organic acid yield, cellulase activity, ligninase activity), and environmental features (pH, temperature, dissolved oxygen, substrate concentration).

[0077] Select the gradient boosting tree model (XGBoost, tree depth 6, learning rate 0.1), with 70% for the training set, 15% for the validation set, and 15% for the test set. Optimize the parameters (number of trees 100) and evaluate the model performance (MSE = 0.04, R² = 0.93).

[0078] According to the screening criteria (cellulose degradation rate > 60%, lignin degradation rate > 50%, reducing sugar yield > 8 g / L, organic acid yield > 4 g / L, cellulase activity > 40 U / mL, ligninase activity > 30 U / mL), 5 dominant strains are screened out (numbered F1, F2, F3, W1, W2, where F1, F2, F3 are cellulose - degrading bacteria, and W1, W2 are lignin - degrading bacteria).

[0079] Step A5: AI - assisted optimization of experimental design For the 5 dominant strains obtained from the initial screening, use the AI performance prediction model to design an optimized experimental plan. The model inputs include fermentation conditions (pH 4.0 - 8.0, temperature 25°C - 40°C, substrate concentration 1% - 10% w / v, oxygen supply aerobic or micro - aerobic, stirring speed 100 - 300 rpm) and strain characteristics (optimum pH 4.5 - 6.0, optimum temperature 28°C - 35°C, expression level of cellulase gene cel5A, expression level of ligninase gene lcc1).

[0080] The model output is the strain performance parameters (cellulose degradation rate, lignin degradation rate, reducing sugar yield, organic acid yield, cellulase activity, ligninase activity).

[0081] Use Bayesian optimization to search for the optimal fermentation conditions, with the optimization goal of maximizing the cellulose degradation rate and lignin degradation rate.

[0082] The optimization results are as follows: The optimal conditions for strain F1 are pH 5.0, temperature 30°C, substrate concentration 5% w / v, oxygen supply aerobic, and stirring speed 200 rpm; The optimal conditions for strain W1 are pH 6.0, temperature 35°C, substrate concentration 5% w / v, microaerobic oxygen supply, and stirring speed 250 rpm. The optimized conditions were verified in the BioLector system with a fermentation cycle of 48 hours and 5 parallel replicates set up.

[0083] The experimental results were as follows: the cellulose degradation rate of strain F1 increased to 88.5% (65.2% before optimization), the reducing sugar yield increased to 12.3 g / L (8.5 g / L before optimization), and the cellulase activity increased to 65.7 U / mL (42.3 U / mL before optimization); The lignin degradation rate of strain W1 increased to 75.6% (52.1% before optimization), the organic acid yield increased to 6.8 g / L (4.2 g / L before optimization), and the ligninase activity increased to 48.9 U / mL (31.5 U / mL before optimization).

[0084] Step A6: Model feedback and optimization Comparing the experimental results with the AI model prediction results, calculating the error metrics (MSE = 0.03, R² = 0.94), the model has a relatively high prediction accuracy.

[0085] Feed the experimental data back to the AI model, update the training dataset, and optimize the model parameters through incremental learning (adjust the number of trees in XGBoost to 120). Rerun the model to predict the further optimization potential of strains F1 and W1. The results show that the performance improvement is <5%, and the iterative optimization is terminated.

[0086] Select strains with complementary functions (F1 is a cellulose-degrading bacterium and W1 is a lignin-degrading bacterium) from the optimized dominant strains to construct a complex microbial community.

[0087] Use the AI model to predict the effects of different strain ratios on fermentation performance, design a high-throughput experiment to verify the optimal ratio, and the ratio schemes include 1:1, 2:1, and 1:2. The experimental results were analyzed by the response surface method (RSM), fitting a quadratic polynomial model, and determining the optimal ratio as F1:W1 = 2:1.

[0088] The experimental results were as follows: the optimal performance of the complex microbial community was a cellulose degradation rate of 90.2%, a lignin degradation rate of 78.5%, a reducing sugar yield of 13.5 g / L, an organic acid yield of 7.2 g / L, a cellulase activity of 68.3 U / mL, and a ligninase activity of 50.1 U / mL.

[0089] Evaluate the stability of the complex microbial community through continuous subculture experiments (subculturing 10 times, with each culture for 48 hours). The high-throughput sequencing results showed that the change rate of the microbial community composition was 15.6% (<20%), indicating high microbial community stability.

[0090] Step A7: Verification by Pilot-Scale Fermentation Experiment Verify the composite microbial community (F1:W1 = 2:1) in a 50 L pilot-scale fermenter. The fermenter is a Sartorius Biostat, equipped with industrial-grade sensors and a PID controller.

[0091] Use Pleurotus ostreatus residue as the substrate (cellulose content 35%, lignin content 15%). After pretreatment, the substrate concentration is 5% w / v. The fermentation conditions are pH 5.5, temperature 32 °C, dissolved oxygen > 20%, and stirring speed 225 rpm. The fermentation cycle is 72 hours, and 3 parallel replicates are set. Samples are taken every 12 hours to measure the cellulose degradation rate, lignin degradation rate, reducing sugar yield, organic acid yield, cellulase activity, ligninase activity, and microbial community composition.

[0092] The experimental results are as follows: cellulose degradation rate 89.8%, lignin degradation rate 77.9%, reducing sugar yield 13.2 g / L, organic acid yield 7.0 g / L, cellulase activity 67.5 U / mL, ligninase activity 49.8 U / mL, and the change rate of microbial community composition is 16.2%, which is consistent with the high-throughput experimental results. Comparing the experimental results with the AI model prediction results, MSE = 0.02, R² = 0.95, and the model prediction accuracy is high.

[0093] Step A8: Performance Evaluation and Application; Compare the performance of the composite microbial community (F1:W1 = 2:1) with the control microbial community (unoptimized standard microbial community of Trichoderma reesei and Phanerochaete chrysosporium, ratio 1:1), and draw a bar chart. The results show that: the cellulose degradation rate of the composite microbial community increases by 38.2% (the control microbial community is 65.0%), the lignin degradation rate increases by 46.7% (the control microbial community is 53.1%), the reducing sugar yield increases by 58.8% (the control microbial community is 8.3 g / L), the organic acid yield increases by 66.7% (the control microbial community is 4.2 g / L), the cellulase activity increases by 60.7% (the control microbial community is 42.0 U / mL), and the ligninase activity increases by 63.3% (the control microbial community is 30.5 U / mL).

[0094] The economic evaluation shows that the cultivation cost is 15 yuan / L, the product value (reducing sugar and organic acid) is 28 yuan / L, and BCR = 1.87, which is economically feasible.

[0095] The composite microbial community was prepared into a liquid microbial agent (0.1% Tween 80 was added after the bacterial liquid was concentrated and stored at 4°C), and a 1000L industrial test was carried out. The results showed that the cellulose degradation rate was 89.2%, the lignin degradation rate was 77.5%, the reducing sugar yield was 12.8 g / L, the organic acid yield was 6.8 g / L, the change rate of the microbial community composition was 17.5%, and the stability of the microbial agent was high. The technology was transferred to an environmental protection technology company, and the transferred content included the composite microbial community, fermentation process parameters, AI model, and database.

[0096] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A high-throughput fermentation strain screening method based on artificial intelligence, characterized in that, Including: Step A1 is to construct a strain library: Obtain standard strains from public strain preservation centers, isolate potential functional strains from edible mushroom residues, and screen strains with cellulose degradation, lignin degradation, or metabolite generation functions from literature or patents. Record the growth characteristics, metabolic characteristics, and genomic information of the strains, and construct a strain library database that supports data import, query, and update; Step A2 is to design small-scale fermentation experiments: Using a high-throughput fermentation platform, set different fermentation conditions, including a pH range of 4.0 - 8.0, a temperature range of 25°C - 40°C, three gradients of substrate concentration including low, medium, and high, oxygen supply including aerobic, microaerobic, and anaerobic conditions, and a stirring speed range of 100 - 300 rpm; Step A3 is data collection: Real-time monitor physical parameters through online sensors, including pH, temperature, dissolved oxygen, cell concentration, and stirring speed; Detect biochemical parameters through off-line sampling combined with high-performance liquid chromatography, gas chromatography-mass spectrometry, or enzyme-linked immunosorbent assay, including cell concentration, substrate consumption, metabolite generation, and enzyme activity; Step A4 is to construct an AI preliminary screening model: Based on the data collected in Step A3, construct an artificial intelligence model. The artificial intelligence model takes the growth characteristics, metabolic characteristics, and fermentation conditions of the strains as inputs and the preset screening criteria as outputs. The screening criteria include the synthesis potential of the target metabolite, specific enzyme activity, or substrate degradation rate. Through real-time data collection and model prediction, quickly screen out strains with potential excellent performance and eliminate strains that do not meet the requirements; Step A5 is AI-assisted optimization experimental design: For the dominant strains preliminarily screened in Step A4, use the AI performance prediction model to design an optimized experimental plan. The optimized experimental plan includes optimizing the medium formula, adjusting the fermentation conditions, and performing genetic engineering transformation. The AI performance prediction model takes the fermentation conditions and strain characteristics as inputs and the strain performance parameters as outputs, and verify the optimization plan through high-throughput experiments; Step A6 is model feedback and optimization: According to the results of high-throughput experiments, calculate the error between the experimental data and the AI model prediction results, feedback the experimental data to the AI model, update the training data set, and optimize the model parameters through incremental learning or retraining. After multiple rounds of iterative optimization, obtain the single strain or composite microbial community with the best performance; Step A7 is pilot-scale fermentation experiment verification: For the strains or composite microbial community screened in Step A6, conduct fermentation experiments in a 10 - 100L pilot-scale fermentation tank, use actual edible mushroom residues as the fermentation substrate, set multiple repeated experiments, collect fermentation data, and verify their performance; Step A8 is performance evaluation and application: Compare the pilot-scale experimental data with the AI model prediction results, calculate the accuracy of the model prediction, screen out the strains or composite microbial community with the best performance, and prepare them into fermentation inoculants for resource utilization of mushroom residues or other fermentation fields.

2. The high-throughput fermentation strain screening method based on artificial intelligence according to claim 1, wherein The strain library database in Step A1 includes: strain number, name, taxonomic status, source, growth characteristics, metabolic characteristics, genomic information; Among them, the growth characteristics include the optimal pH, temperature range, and oxygen demand; the metabolic characteristics include the substrate utilization spectrum, types and yields of metabolites, and enzyme activities; the genomic information includes sequencing data, functional gene annotation, and metabolic pathway information.

3. The method for screening high-throughput fermentation strains based on artificial intelligence according to claim 2, wherein In step A2, the high-throughput fermentation platform is a micro-bioreactor system, and the volume of each reactor is 1-10 mL, and dozens to hundreds of fermentation experiments can be run simultaneously; The fermentation conditions include: the pH range is 4.0-8.0, the temperature range is 25°C-40°C, the substrate concentration includes three gradients of low, medium, and high, the oxygen supply includes aerobic, micro-aerobic, and anaerobic conditions, and the stirring speed range is 100-300 rpm.

4. The method for screening high-throughput fermentation strains based on artificial intelligence according to claim 3, wherein, In step A3, the data collection includes: Using on-line sensors to record pH, temperature, dissolved oxygen, cell concentration, and stirring speed in real time; Sampling regularly, using high-performance liquid chromatography or gas chromatography-mass spectrometry to analyze substrate consumption and metabolite production, and using an enzyme-labeling instrument or spectrophotometer to measure enzyme activities; Among them, the substrates include cellulose, lignin, and protein, the metabolites include reducing sugars, organic acids, and volatile fatty acids, and the enzyme activities include cellulase, ligninase, and protease activities.

5. The method for screening high-throughput fermentation strains based on artificial intelligence according to claim 4, characterized in that In step A4, the construction of the AI preliminary screening model includes the following steps: Clean the data collected in step 3, remove outliers, and perform standardization or normalization processing, and extract features. The features include the growth characteristics, metabolic characteristics, and environmental characteristics of the strain. Among them, the growth characteristics include the cell growth rate and the maximum cell concentration, the metabolic characteristics include the substrate consumption rate, metabolite production, and enzyme activities, and the environmental characteristics include pH, temperature, dissolved oxygen, and substrate concentration; According to the screening target, select support vector machine, random forest, convolutional neural network, or gradient boosting tree as the model. Among them, the screening targets include cellulose degradation rate, lignin degradation rate, metabolite production, or enzyme activity; Divide the data into a training set, a validation set, and a test set. Among them, the training set accounts for 70%, the validation set accounts for 15%, and the test set accounts for 15%. Use the training set to train the model, optimize the model parameters through the validation set, including the learning rate, tree depth, or number of convolutional layers, and use the test set to evaluate the generalization ability of the model, and calculate the accuracy, mean square error, or coefficient of determination of the model; According to the preset screening criteria, predict the performance of the strain. The screening criteria include the synthesis potential of the target metabolite, specific enzyme activity, or substrate degradation rate. Among them, the target metabolites include reducing sugars, organic acids, or volatile fatty acids, the specific enzyme activities include cellulase activity, ligninase activity, or protease activity, and the substrate degradation rates include cellulose degradation rate or lignin degradation rate. Screen out potential dominant strains and eliminate strains that do not meet the requirements.

6. The high-throughput fermentation strain screening method based on artificial intelligence according to claim 5, characterized in that In step A5, the construction of the AI performance prediction model includes: Set the input as fermentation conditions and strain characteristics. Among them, the fermentation conditions include pH, temperature, substrate concentration, oxygen supply, and agitation speed. The strain characteristics include growth characteristics, metabolic characteristics, and genomic information. Among them, the growth characteristics include the optimal pH, temperature range, and oxygen demand. The metabolic characteristics include the substrate utilization spectrum, types and yields of metabolites, and enzyme activities. The genomic information includes cellulose degradation-related genes, lignin degradation-related genes, or metabolic pathway information; Set the output as strain performance parameters, including substrate degradation rate, metabolite yield, and enzyme activity. Among them, the substrate degradation rate includes cellulose degradation rate or lignin degradation rate. The metabolite yield includes reducing sugar yield, organic acid yield, or volatile fatty acid yield. The enzyme activity includes cellulase activity, ligninase activity, or protease activity; Use Bayesian optimization or genetic algorithm to search for the optimal combination of fermentation conditions to maximize the strain performance. Among them, Bayesian optimization constructs a surrogate model through Gaussian process regression, predicts the performance, and selects the next set of experimental conditions. The genetic algorithm iteratively optimizes the fermentation conditions through crossover, mutation, and selection operations.

7. The high-throughput fermentation strain screening method based on artificial intelligence according to claim 6, wherein, The optimization experimental plan in step A5 includes: Predict the effects of different medium formulations, including carbon source, nitrogen source, and trace element concentrations, on the strain performance through an AI model, and design optimization experiments; Predict the effects of different fermentation conditions, including pH, temperature, and dissolved oxygen, on the strain performance through an AI model, and design optimization experiments; Combined with genomic information, predict the effects of key gene modifications on the strain performance, and design gene editing experiments, including using the CRISPR-Cas9 technology to modify cellulase genes or ligninase genes.

8. A high-throughput fermentation strain screening method based on artificial intelligence according to claim 7, characterized in that The model feedback and optimization in step A6 include: Compare the high-throughput experimental results with the AI model prediction results, and calculate error metrics, including mean square error, root mean square error, or coefficient of determination, to evaluate the accuracy of the model prediction; Feed the high-throughput experimental data back to the AI model, update the training dataset, and optimize the model parameters through incremental learning or retraining. Among them, incremental learning updates the model weights through an online learning algorithm, and retraining optimizes the model hyperparameters, including learning rate, tree depth, or number of convolutional layers, by re-partitioning the training set, validation set, and test set; According to the high-throughput experimental results, adjust the optimization objectives, including increasing the substrate degradation rate, metabolite yield, or enzyme activity, re-run the AI model, and design the next round of optimization experimental plan until the strain performance reaches the preset target or the performance improvement amplitude is lower than the preset threshold. Among them, the preset target includes a cellulose degradation rate greater than 90%, a metabolite yield greater than 15 g / L, or an enzyme activity greater than 100 U / mL, and the preset threshold includes a performance improvement less than 5% in two consecutive rounds of optimization; Select strains with complementary functions from the optimized dominant strains to construct a composite microbial community. The complementary functions include cellulose degradation, lignin degradation, and metabolite production. Use the AI model to predict the effects of different strain ratios on the fermentation performance, and design high-throughput experiments to verify the optimal ratio. Among them, the strain ratio is designed through orthogonal experiments or response surface methods, including proportional combinations such as 1:1, 2:1, and 1:2; The stability of the complex microbial community during long-term fermentation was evaluated through continuous subculture experiments. The stability was analyzed by the change rate of the microbial community composition using high-throughput sequencing technology, and the microbial community combinations with a change rate greater than 20% were eliminated.

9. A high-throughput fermentation strain screening method based on artificial intelligence according to claim 8, characterized in that The pilot-scale fermentation experiment in step A7 includes: Using a 10-100 L pilot-scale fermenter to simulate industrial fermentation conditions; Setting the optimal pH, temperature, dissolved oxygen, stirring speed, and substrate concentration according to the optimization results of the high-throughput experiment; Using actual edible mushroom residues as the fermentation substrate, including Lentinula edodes residues or Pleurotus ostreatus residues; Real-time monitoring of pH, temperature, dissolved oxygen, cell concentration, and stirring speed, and offline detection of substrate consumption, metabolite production, enzyme activity, and microbial community composition; Calculating the substrate degradation rate, metabolite yield, fermentation efficiency, and microbial community stability, and comparing the experimental results with the predicted results of the AI model.

10. A method for screening high-throughput fermentation strains based on artificial intelligence according to claim 9, characterized in that, The performance evaluation and application in step A8 include: Comparing the performance of the optimized strain or complex microbial community, and drawing performance charts, including substrate degradation rate, metabolite yield, fermentation efficiency, and stability; Evaluating the cultivation cost, substrate utilization efficiency, and product value of the strain or microbial community; Uploading the pilot-scale experiment data, preferred strain performance data, and fermentation process parameters to the strain library database; Preparing the preferred strain or complex microbial community into a fermentation inoculant, including freeze-dried powder or liquid inoculant, conducting industrial tests, and transferring the technology to relevant enterprises or research institutions.

Citation Information

Cited By

  • Probiotic product stability prediction model and construction method and application thereof

    CN121306280A