Method for determining high-ethanol-tolerance yeast strain based on multiple omics markers
By integrating gene expression levels, proteomics, and metabolomics data, a machine learning model was constructed to screen for yeast strains with high ethanol tolerance. This solved the problem of time-consuming and labor-intensive traditional methods, and enabled efficient and low-cost yeast strain screening and fermentation process optimization.
Patent Information
- Application Number
- CN202510692376.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-05-27
AI Technical Summary
Traditional yeast strain screening methods are time-consuming, labor-intensive, and costly, making it difficult to efficiently screen for yeast strains with high ethanol tolerance.
Using a multi-omics biomarker approach, gene expression levels, proteomics, and metabolomics data were integrated, and a strain screening model was constructed using machine learning algorithms. The characteristics of biomarkers such as GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine, and trehalose-6-phosphate were combined to screen for yeast strains with high ethanol tolerance.
It reduces the cost of yeast strain screening, improves screening efficiency and accuracy, and can more accurately identify yeast strains with high ethanol tolerance, guiding the optimization of industrial fermentation processes.
Smart Images

Figure CN120998302A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a method for determining high-ethanol-tolerant yeast strains based on multi-omics markers. BACKGROUND
[0002] Saccharomyces cerevisiae is the core strain of industrial ethanol fermentation. During ethanol fermentation, as the ethanol concentration increases, it will produce a stress effect on yeast cell growth. In order to survive and grow, yeast cells will produce corresponding stress response to cope with this stress. This coping mechanism is that yeast cells have developed ethanol tolerance. Ethanol tolerance directly determines the fermentation efficiency and product concentration of Saccharomyces cerevisiae. However, high-concentration ethanol can damage cell membrane integrity, inhibit enzyme activity and induce oxidative stress, leading to cell growth arrest and even death. It is generally believed that the effect of ethanol on yeast cells mainly manifests in three aspects, namely, inhibition of cell growth, cell survival and fermentation. Therefore, ethanol tolerance is often defined by the effect of ethanol on cell growth. The highest ethanol concentration that allows growth in the medium containing 1% to 14% ethanol at room temperature for 48 hours represents the ethanol tolerance level of yeast. Among them, the strain that can grow in 3% to 6% ethanol has poor ethanol tolerance, the strain that can grow in 6% to 10% ethanol has moderate ethanol tolerance, and the strain that can grow in 10% to 13% ethanol has high ethanol tolerance. This definition method is simple and is commonly used to screen yeast strains with ethanol tolerance. Improving the ethanol tolerance of yeast is of great significance to improve fermentation efficiency, yield and quality of the final product.
[0003] The traditional yeast strain screening method relies on culturing the strain in different ethanol concentration medium, and judging the ethanol tolerance of the yeast strain by measuring parameters such as growth curve, maximum specific growth rate and final cell density. However, this method requires a large number of experimental operations, which is time-consuming and labor-intensive, resulting in high screening cost.
[0004] Therefore, there is an urgent need for a method for reducing the screening cost of yeast strains. SUMMARY
[0005] Therefore, it is necessary to provide a method for determining high-ethanol-tolerant yeast strains based on multi-omics markers in view of the above technical problems. The method can reduce the screening cost of yeast strains.
[0006] The present application adopts the following technical solutions: The present application provides a method for determining high-ethanol-tolerant yeast strains based on multi-omics markers, comprising: obtaining the gene expression, protein expression profile and metabolite abundance of a plurality of unlabeled wild yeast strains; respectively performing Z-score standardization processing on the gene expression, protein expression profile and metabolite abundance, and aligning the three groups of data after Z-score standardization processing according to features; aligning the three groups of data after extracting the features, and extracting the standardized values of the GRE1 gene, the TPS1 gene, the HSP104 protein, the ADH2 protein, the CTA1 protein, the sphingosine, and the trehalose-6-phosphate, to form a plurality of biomarker feature combinations; inputting the plurality of biomarker feature combinations into a strain screening model to screen the industrial strains, to obtain a high ethanol tolerance probability score of each of the plurality of unlabeled wild yeast strains; the unlabeled wild yeast strain corresponding to the high ethanol tolerance probability score greater than or equal to the preset threshold is taken as a high ethanol tolerance yeast strain.
[0007] Preferably, the plurality of biomarker feature combinations are input into a strain screening model to screen the industrial strains, to obtain a high ethanol tolerance probability score of each of the plurality of unlabeled wild yeast strains, and specifically include: arranging the Z-score values of the 7 markers in the plurality of biomarker feature combinations in order as an input vector; In the strain screening model, each tree is independently traversed, each tree starts from a root node, splits according to a feature threshold, and distributes the unlabeled wild yeast strains to a specific leaf node, and the leaf node corresponds to a classification label. The proportion of high ethanol tolerance yeast strains in the specific leaf node during training is recorded as a local probability; The local probabilities of all trees are weighted and averaged to obtain a weighted average result, and the weighted average result is determined as the high ethanol tolerance probability score of the unlabeled wild yeast strain; the weight of the weighted average is the accuracy of the tree.
[0008] Preferably, the training process of the strain screening model specifically includes: obtaining a sample feature matrix, and dividing the sample feature matrix into a training set and a validation set; training the random forest model, the support vector machine model, the gradient boosting decision tree model, the naive Bayes model, the neural network model, and the generalized linear model through the training set, and adjusting the parameters of all models until all models are optimal; For each of the models, the validation set is input into the model to obtain a screening result of the model, the screening result of the model is compared with the true class of the test set to obtain an index of the model; the index includes prediction accuracy, sensitivity, and specificity; comparing and analyzing the indexes of all models, and determining the optimal model according to the comparison and analysis result; determining the optimal model as the strain screening model.
[0009] Preferably, the sample feature matrix is obtained, and specifically includes: obtain a plurality of high ethanol-tolerant yeast strains and a plurality of low ethanol-tolerant yeast strains as a sample training set; respectively obtain sample gene expression, sample protein expression profile and sample metabolite abundance of each yeast strain in the sample training set under different ethanol concentrations; respectively perform Z-score standardization processing on the sample gene expression, sample protein expression profile and sample metabolite abundance to obtain three groups of data; perform feature selection on the three groups of data after Z-score standardization processing to obtain a sample feature matrix; the sample feature matrix includes GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine and trehalose-6-phosphate.
[0010] Preferably, feature selection is performed on the three groups of data after Z-score standardization processing to obtain a sample feature matrix, specifically including: align the data after Z-score standardization processing across omics data by UniProt ID / KEGG ID; perform LASSO regression operation after aligning the data across omics data to screen a plurality of features; According to the screened plurality of features, a n×7 matrix is generated according to a row in the feature matrix corresponding to each yeast strain and a column corresponding to each biomarker, and the n×7 matrix is determined as the sample feature matrix; n the number of yeast strains in the sample training set.
[0011] Preferably, the screening process of the plurality of high ethanol-tolerant yeast strains and the plurality of low ethanol-tolerant yeast strains specifically includes: select candidate strains from a wild-type Saccharomyces cerevisiae strain library; Prepare an ethanol-containing medium; culture the candidate strains in the ethanol-containing medium; After culture, the candidate strains are subjected to primary screening and secondary screening, and the strains obtained by secondary screening are screened according to the high ethanol-tolerant strain screening standard and the low ethanol-tolerant strain screening standard to obtain a plurality of high ethanol-tolerant yeast strains and a plurality of low ethanol-tolerant yeast strains.
[0012] The present application provides a high ethanol-tolerant yeast strain determination device based on multi-omics markers, comprising: The acquisition module is used for acquiring the gene expression, protein expression profile and metabolite abundance of a plurality of unlabeled wild yeast strains; An alignment module is configured to perform Z-score standardization on the gene expression, protein expression profile and metabolite abundance respectively, and align the three groups of data after Z-score standardization according to features; A first determination module is configured to extract the standardized values of the GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine and trehalose-6-phosphate from the three groups of data after feature alignment, and form a plurality of biomarker feature combinations according to the extracted standardized values; A screening module is configured to input the plurality of biomarker feature combinations into a strain screening model to perform industrial strain screening, and obtain a high ethanol tolerance probability score of each of the plurality of unlabeled wild yeast strains; A second determination module is configured to take the unlabeled wild yeast strain corresponding to the high ethanol tolerance probability score greater than or equal to a preset threshold as a high ethanol tolerance yeast strain.
[0013] The application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the above-mentioned determination method of a high ethanol tolerance yeast strain based on a plurality of omics markers.
[0014] The application provides a computer device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the above-mentioned determination method of a high ethanol tolerance yeast strain based on a plurality of omics markers when executing the program.
[0015] The above-mentioned at least one technical solution adopted by the application can achieve the following beneficial effects: The gene expression, protein expression profile and metabolite abundance of a plurality of unlabeled wild yeast strains are obtained, the gene expression, proteomics and metabolomics data are integrated, more comprehensive biological information is provided, and the prediction accuracy and reliability are improved; the gene expression, protein expression profile and metabolite abundance are subjected to Z-score standardization respectively, and the three groups of data after Z-score standardization are aligned according to features; the standardized values of the GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine and trehalose-6-phosphate of the three groups of data after feature alignment are extracted, and a plurality of biomarker feature combinations are formed according to the extracted standardized values; the plurality of biomarker feature combinations are input into a strain screening model to perform industrial strain screening, and a high ethanol tolerance probability score of each of the plurality of unlabeled wild yeast strains is obtained, an efficient strain screening model is constructed, and the limitations of traditional methods in data processing and prediction ability are overcome; the unlabeled wild yeast strain corresponding to the high ethanol tolerance probability score greater than or equal to a preset threshold is taken as a high ethanol tolerance yeast strain. The method can reduce the screening cost of yeast strains. Attached Figure Description
[0016] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:
[0017] Figure 1 A schematic flowchart of a method for determining highly ethanol-tolerant yeast strains based on multi-omics biomarkers provided by the present invention; Figure 2 A performance comparison chart of six machine learning classifiers provided by this invention; Figure 3 A schematic diagram illustrating the global feature importance provided by this invention; Figure 4 This is a schematic diagram comparing the ethanol yield of the HT strain and the random strain provided by the present invention. Figure 5 A schematic diagram of the ROC curve for predicting tolerance of the independent test set (50 strains) provided by this invention; Figure 6 A schematic diagram of the technical route for constructing the multi-omics prediction model for ethanol tolerance in Saccharomyces cerevisiae provided by the present invention; Figure 7 A schematic diagram of a device for identifying highly ethanol-tolerant yeast strains based on multi-omics biomarkers provided by the present invention; Figure 8 This is a schematic diagram of a computer device for implementing a method for determining highly ethanol-tolerant yeast strains based on multi-omics biomarkers, as provided by the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0019] Traditional laboratory screening of yeast strains involves culturing yeast strains in ethanol media of varying concentrations and determining their ethanol tolerance using parameters such as growth curves, maximum specific growth rate, and final cell density. This method requires extensive experimental work, resulting in long screening cycles, low efficiency, and high costs.
[0020] Genetic engineering can improve the ethanol tolerance of yeast through gene knockout, gene overexpression, or gene editing techniques such as CRISPR-Cas9. However, this method requires a deep understanding of tolerance-related genes and raises concerns about genetic stability and transgenic safety. Therefore, analyzing the molecular mechanisms of ethanol tolerance and developing accurate prediction methods are of great significance for industrial strain selection, fermentation process optimization, and synthetic biology modification.
[0021] Researchers identified over 250 genes related to ethanol tolerance through whole-genome screening, focusing on vacuole, peroxisome, cell wall, and membrane-related pathways, and verified the role of the FPS1 gene (glycerol channel protein) in ethanol efflux. Through transcriptome and proteome analysis, researchers found significant changes in yeast mitochondria, endoplasmic reticulum function, signal transduction (such as GPCR pathway), and aromatic amino acid metabolism under ethanol stress. Through multi-omics (transcriptome, proteome, and metabolome) integration and network analysis, researchers proposed a differentiation mechanism for high-tolerance (HT) and low-tolerance (LT) phenotypes, emphasizing the roles of longevity pathways, peroxisome function, energy metabolism, and gene regulation, and established a biological mechanism called "ethanol stress buffering model." Existing research has only analyzed ethanol tolerance mechanisms through single or multiple biological omics (genomics, transcriptomics, proteomics, and metabolomics), without involving predictive model construction.
[0022] Single-omics studies, such as using genomics, transcriptomics, proteomics, or metabolomics alone to study yeast response mechanisms under high ethanol concentrations. Single-omics studies mainly focus on understanding how single-level biological information affects ethanol tolerance. Yeast cell response mechanisms involve genes, transcription, translation, protein function, metabolism, and other levels, and single-omics data cannot cover all these levels of information. Single-omics research data can only provide partial biological information and cannot fully reflect the yeast response mechanism under high ethanol concentrations, limiting the accuracy and comprehensiveness of the prediction results. Multi-omics integration studies, such as genomics, transcriptomics, proteomics, and metabolomics, provide a more comprehensive understanding of yeast response mechanisms under high ethanol concentrations. Although multi-omics integration provides a more comprehensive perspective, these studies mainly focus on mechanism exploration and do not further utilize these data for predictive evaluation.
[0023] In addition, traditional methods are limited by experimental conditions, with limited screening efficiency and result accuracy. In recent years, single-omics studies have attempted to analyze ethanol tolerance mechanisms, but due to the single dimension of data, it is difficult to fully reflect complex biological processes, limiting the application of high-efficiency screening.
[0024] The present application aims to solve the above technical problems, and provides a Saccharomyces cerevisiae ethanol tolerance prediction method integrating gene expression, proteomics, metabolomics data and machine learning algorithms, so as to improve the efficiency and accuracy of strain screening, reduce experimental workload, guide synthetic biology modification and fermentation process optimization, and meet the needs of industrial production.
[0025] The technical solutions provided by the embodiments of the present application will be described in detail below with reference to the drawings.
[0026] Devices such as desktop computers, servers, notebook computers, etc. can execute the solutions of the present application. For the convenience of description, the following will only take the server as the execution subject for description.
[0027] Figure 1 The present application is a method for determining a high-ethanol-tolerant yeast strain based on multi-omics markers. The process is as follows: S101: Obtain the gene expression, protein expression profile and metabolite abundance of a plurality of unlabeled wild yeast strains.
[0028] S102: Perform Z-score standardization processing on the gene expression, protein expression profile and metabolite abundance respectively, and align the three sets of data after Z-score standardization processing according to the characteristics.
[0029] Cross-omics data alignment: use the Weighted Gene Co-expression Network Analysis (WGCNA) package of R language to construct gene-protein-metabolite modules. Gene-protein-metabolite modules refer to the feature set composed of genes, proteins and metabolites highly associated with ethanol concentration identified through co-expression network analysis, reflecting the potential regulatory network or metabolic pathway. Network type: distinguish signed hybrid, module definition: minimum gene number = 30, cutting height = 0.25, screen modules significantly related to ethanol concentration (|r|>0.7, p<0.01).
[0030] S103: Extract the standardized values of GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine and trehalose-6-phosphate from the three sets of data after feature alignment, and form a multi-omics biomarker feature combination according to the extracted standardized values.
[0031] S104: Input the multi-omics biomarker feature combination into the strain screening model for industrial strain screening, and obtain the high-ethanol-tolerance probability score of each of the plurality of unlabeled wild yeast strains.
[0032] In an exemplary embodiment, the screening process of the plurality of high ethanol-tolerant yeast strains and the plurality of low ethanol-tolerant yeast strains specifically comprises: selecting candidate strains from a wild-type S. cerevisiae strain library; preparing an ethanol-containing medium; culturing the candidate strains in the ethanol-containing medium; after the culturing, performing primary screening and secondary screening on the candidate strains, and screening the strains obtained through the secondary screening according to the screening standard of the high ethanol-tolerant strains and the screening standard of the low ethanol-tolerant strains, to obtain the plurality of high ethanol-tolerant yeast strains and the plurality of low ethanol-tolerant yeast strains.
[0033] Specifically, the screening process of the plurality of high ethanol-tolerant yeast strains and the plurality of low ethanol-tolerant yeast strains is as follows: I. Strain source and pre-culture Strain library: candidate strains are selected from a wild-type S. cerevisiae strain library preserved in the laboratory.
[0034] Pre-culture conditions: each strain is recovered from a glycerol storage tube and inoculated into YPD liquid medium (20 g / L glucose, 20 g / L tryptone, 10 g / L yeast extract powder), and cultured at 28°C with 160 rpm shaking for 12 h to the logarithmic growth phase (OD600≈0.7).
[0035] II. Preparation of ethanol-containing medium Medium formula: basic YPD liquid medium (same as pre-culture), after sterilization and cooling, sterile anhydrous ethanol is added to a final concentration of 13% (v / v).
[0036] Control medium: YPD medium without ethanol, used to evaluate the basic growth performance of the strains.
[0037] III. Primary screening: solid plate tolerance test Gradient dilution: the pre-culture bacterial solution is gradient diluted to 10^-3~10^-5, and 100 μL is taken and spread on YPD solid plates containing 13% ethanol and control plates without ethanol.
[0038] Culture and observation: incubate at 28°C for 48~72 h, and the screening criteria are as follows: Primary screening HT candidate strains: form obvious colonies (colony diameter ≥1 mm) on 13% ethanol plates, and the number of colonies differs by ≤50% from the control group.
[0039] Primary screening LT candidate strains: no visible colonies or the number of colonies is reduced by ≥90% on 13% ethanol plates compared with the control group.
[0040] IV. Secondary screening: liquid culture growth curve analysis Inoculation and culture: The HT and LT candidate strains obtained by preliminary screening were inoculated into YPD liquid medium containing 13% ethanol at an inoculation amount of 3% (OD600≈0.7), and an ethanol-free control group was set up.
[0041] Dynamic monitoring of OD600: 28°C, 160 rpm shaking culture, OD600 value was measured every 2 h for 24 h, and the growth curve was drawn.
[0042] V. Screening criteria: HT strain: lag phase in 13% ethanol ≤6 h, log phase OD600 increase rate ≥0.1 / h, final OD600 value ≥60% of the control group.
[0043] LT strain: lag phase ≥12 h, OD600 increase rate ≤0.05 / h, final OD600 value ≤20% of the control group.
[0044] VI. Tolerance verification 1. Cell survival rate determination: The HT and LT candidate strains were cultured in 13% ethanol medium for 24 h, and then gradient diluted and plated on ethanol-free YPD plates.
[0045] Calculate survival rate: Survival rate (%) = (CFU of ethanol-treated group / CFU of control group) x 100.
[0046] HT criteria: survival rate ≥50%; LT criteria: survival rate ≤10%.
[0047] 2. Cell morphology observation: Take the 13% ethanol-treated 24 h bacterial solution, centrifuge to collect the bacterial cells, fix with 2.5% glutaraldehyde, and observe under scanning electron microscope.
[0048] HT strain: cell morphology is complete (elliptical, smooth surface); LT strain: cell rupture or severe shrinkage.
[0049] Data collection and training set construction: Select 10 strains (biological repeats) from each of the HT / LT groups and 5 ethanol concentrations as training set samples.
[0050] a. Gene sequencing: Use RNA-Seq technology to perform high-throughput sequencing of mRNA gene expression in Saccharomyces cerevisiae cells under different ethanol concentrations, and screen out genes that may be related to ethanol tolerance.
[0051] (1) Experimental design: For the training set strains, collect mid-log phase cells (OD600≈0.6) and fix them quickly with an RNA protection agent (such as RNAlater). Use strand-specific library preparation (retaining strand orientation information). Use Illumina NovaSeq6000 and adopt PE150 mode for sequencing, requiring data volume ≥20M clean reads / sample (to ensure detection of low-abundance genes).
[0052] (2) Bioinformatics analysis: Use fastp software to perform quality control on the raw sequencing data and filter out low-quality sequences and sequencing adapter sequences. Use hisat2 software to perform sequence alignment with Saccharomyces cerevisiae S288C (NCBI RefSeq ID: GCF_000146045.2) as the reference genome, and use stringtie software to assemble and quantify transcripts, outputting the TPM quantification of each gene.
[0053] (1) Experimental design: In the quantitative proteomics analysis, the S. cerevisiae cells were lysed and reduced and alkylated, then digested with Trypsin Gold, labeled with TMT 16plex peptides, fractionated by high-pH reverse-phase chromatography, and detected by Q Exactive HF-X mass spectrometer in DDA mode (MS1 120,000 resolution, MS2 45,000 resolution).
[0054] (2) Data analysis: The raw data was processed by MaxQuant (v2.1.0) software with Andromeda search engine, with the following parameter settings: parent ion mass tolerance 4.5 ppm, fragment ion tolerance 20 ppm, enzyme cleavage mode Trypsin / P (allowing up to 2 missed cleavage sites), fixed modification TMT 16plex (N-terminal and lysine) and cysteine alkylation (+57.021 Da), variable modifications including methionine oxidation (+15.995 Da) and protein N-terminal acetylation (+42.011 Da), database selection UniProt S. cerevisiae reference protein library (May 2024 version, containing common contaminants), and FDR <1% calculated by reverse bait database as the identification threshold.
[0055] c. Quantitative metabolomics analysis: Through mass spectrometry analysis, the abundance changes of metabolites in the fermentation broth under different ethanol concentrations were obtained.
[0056] (1) Experimental design The sample was pretreated by pre-cooled methanol-acetonitrile-water (4:4:2, v / v / v) to quench metabolic activity, ultrasonic-assisted extraction (300W, 2s / 3s pulse, 5min), and centrifugation (15,000g, 15min, 4℃) to take the supernatant and concentrate by nitrogen blowing. An ultra-high performance liquid chromatography (UPLC) system was used, coupled with two chromatographic columns:
[0057] Polar metabolites: ACQUITY UPLC BEH Amide column (2.1x100mm, 1.7μm), mobile phase: 25mM ammonium acetate + ammonia water (pH9.0) / acetonitrile.
[0058] Non-polar metabolites: ACQUITY UPLC HSS T3 column (2.1x100mm, 1.7μm), mobile phase: 0.1% formic acid water / acetonitrile. Mass spectrometry detection used Q-TOF 6600+ system, ESI positive and negative ion mode switching scan, parameter settings: ion source temperature 500℃, spray voltage ±5.5kV, collision energy 20-40V, mass range m / z 50-1500, resolution >30,000. An internal standard (L-2-chlorophenylalanine) was added to each batch of experiments to monitor data stability.
[0059] Raw mass spectrometry data was subjected to peak extraction, alignment and normalization (internal standard correction) by Progenesis QI. Metabolite identification was performed by accurate mass number (error <5ppm) and secondary spectrum matching (METLIN / HMDB database, similarity >80%).
[0060] Data preprocessing: Standardization: Due to the significant difference in the dimension of different omics data (such as gene TPM value range 0~10^5, metabolite peak area 10^3~10^8), the present application performs Z-score standardization on gene expression, protein and metabolite abundance by R language, eliminating the bias of different omics quantitative methods and experimental batches.
[0061] Dimension reduction: Principal component analysis (PCA) dimension reduction technique is used to reduce data dimension and retain key information.
[0062] In an exemplary embodiment, the training process of the strain screening model specifically comprises: obtaining a sample feature matrix, and dividing the sample feature matrix into a training set and a validation set; training a random forest model, a support vector machine model, a gradient boosting decision tree model, a naive Bayes model, a neural network model and a generalized linear model through the training set, and adjusting the parameters of all models until all models are optimal; for each of all models, inputting the validation set into the model to obtain a screening result of the model, comparing the screening result of the model with the true class of the test set, and obtaining an index of the model; the index includes prediction accuracy, sensitivity and specificity; comparing and analyzing the indexes of all models, and determining an optimal model according to the comparison and analysis result; and determining the optimal model as the strain screening model.
[0063] In an exemplary embodiment, the sample feature matrix is obtained, specifically comprising: obtaining a plurality of high-ethanol-tolerant yeast strains and a plurality of low-ethanol-tolerant yeast strains as a sample training set; obtaining sample gene expression, sample protein expression profile and sample metabolite abundance of each yeast strain in the sample training set under different ethanol concentrations respectively; performing Z-score standardization processing on the sample gene expression, the sample protein expression profile and the sample metabolite abundance respectively to obtain three groups of data; performing feature selection on the three groups of data after Z-score standardization processing to obtain a sample feature matrix; the sample feature matrix includes GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine and trehalose-6-phosphate.
[0064] Specifically, the present application adopts six machine learning algorithms in parallel training based on the caret package of R language to comprehensively evaluate the model performance, and the six machine learning algorithms are as follows: Random Forest (RF): based on Bootstrap sampling to construct multiple decision trees, and output classification results through voting mechanism, the default number of trees is set to 500, which is suitable for high-dimensional features and anti-overfitting.
[0065] Support Vector Machine (SVM): selecting a radial basis kernel function to realize nonlinear classification by maximizing the classification interval, and adjusting the penalty parameter C and the kernel parameter γ.
[0066] Gradient Boosting Decision Tree (GBDT): iteratively optimizing the residual error of the decision tree, the default learning rate is 0.1, and the tree depth is 3, which is suitable for scenarios where there is complex interaction between features.
[0067] Naive Bayes Model (NBM): Based on Bayes theorem, assuming feature independence, suitable for small sample data, probability threshold set to 0.5.
[0068] Neural Network (NN): Build a single hidden layer fully connected network, hidden layer node number is 1.5 times the number of features, activation function uses ReLU, and optimizer selects Adam.
[0069] Generalized Linear Model (GLM): Use Logistic regression (binary classification), L2 regularization to control coefficient weight, and logit as the link function.
[0070] The input data of the six models needs to be standardized (Z-score), and the category label (HT / LT) is coded as 1 / 0. The training set and test set are divided into 7:3 proportion by stratified splitting.
[0071] Optimize the six models.
[0072] Based on the pROC package of R language, optimize the hyperparameters of the model through ten-fold cross-validation, reduce overfitting, and improve the generalization ability of the model.
[0073] 1. The hyperparameter optimization range is: RF: Adjust mtry (the number of features randomly selected by each tree, range 1-7) and ntree (100-500).
[0074] SVM: Grid search C (0.1-10) and γ (0.01-0.1).
[0075] GBDT: Optimize nrounds (iteration number, 50-200), max_depth (3-6), and eta (learning rate, 0.01-0.3).
[0076] NN: Adjust the number of hidden layer nodes (5-15) and dropout rate (0.2-0.5).
[0077] GLM: Regularization parameter lambda (0.001-1).
[0078] 2. The cross-validation process is: Divide the training set into 10 equal parts, and train with 9 parts and validate with 1 part, repeat 10 times.
[0079] Record the AUC, sensitivity (recall rate), and specificity of each fold, and take the average as the basis for hyperparameter selection.
[0080] Early stopping strategy prevents overfitting, for example, terminate when GBDT iteration loss no longer decreases.
[0081] 3. Performance comparison and model selection: Figure 2 The performance comparison chart of the six machine learning classifiers provided by the present application is shown in FIG. 1, which comprehensively compares the six models from the three dimensions of prediction accuracy, sensitivity and specificity. The RF model is significantly better than other models in ROC, sensitivity and specificity, p<0.05, paired t test, and has high training efficiency, less than 5 minutes. Figure 2
[0082] The core hyperparameters of the random forest are shown in Table 1: Table 1 In an exemplary embodiment, feature selection is performed on the three groups of data after Z-score standardization, and a sample feature matrix is obtained, specifically including: aligning the data after Z-score standardization across omics data by UniProt ID / KEGG ID; after aligning the omics data, performing LASSO regression operation to screen a plurality of features; according to the screened plurality of features, generating an n×7 matrix according to a row in the feature matrix corresponding to each yeast strain and a column corresponding to each biomarker, and determining the n×7 matrix as the sample feature matrix. n The number of yeast strains in the sample training set.
[0083] Specifically, the feature selection process is: Cross-omics data alignment: using the Weighted Gene Co-expression Network Analysis (WGCNA) package of R language to construct gene-protein-metabolite modules. Gene-protein-metabolite module refers to a feature set composed of genes, proteins and metabolites highly associated with ethanol concentration identified by co-expression network analysis, reflecting the potential regulatory network or metabolic pathway. Network type: distinguish positive and negative regulation (signed hybrid), module definition: minimum gene number = 30, cutting height = 0.25, screen modules significantly related to ethanol concentration (|r|>0.7, p<0.01).
[0084] LASSO regression: further screen the features with the greatest contribution to the prediction model by LASSO regression in the glmnet package of R language to reduce the risk of overfitting.
[0085] The present application integrates the following 7 multi-omics biomarker feature combinations: Genes: GRE1 gene, TPS1 gene; Proteins: Heat shock protein 104 (HSP104), alcohol dehydrogenase 2 (ADH2), catalase 1 (CTA1); Metabolites: Sphinganine, Trehalose-6-phosphate.
[0086] Label: Binary classification label with HT / LT classification.
[0087] In an exemplary embodiment, a plurality of sets of student biomarker features are combined, input into a strain screening model for industrial strain screening, and a high ethanol tolerance probability score of each of a plurality of unlabeled wild yeast strains is obtained, specifically comprising: arranging the Z-score values of 7 markers in the plurality of sets of student biomarker features in order as an input vector; in the strain screening model, each tree is independently traversed, each tree starts from a root node, splits according to a feature threshold value, and assigns the unlabeled wild yeast strain to a specific leaf node, and the leaf node corresponds to a classification label; record the proportion of high ethanol tolerance yeast strains in a specific leaf node during training as a local probability; the local probabilities of all trees are weighted and averaged to obtain a weighted average result, and the weighted average result is determined as the high ethanol tolerance probability score of the unlabeled wild yeast strain; the weight of the weighted average is the accuracy of the tree.
[0088] Specifically, the SHAP values of the standardized multi-omics feature matrix (7 key biomarkers: 2 genes, 3 proteins, and 2 metabolites) using the trained optimal random forest model (RF, AUC = 0.95) are calculated using the R language shapr package to evaluate the importance of the features to the model, and a SHAP importance plot is drawn, Figure 3 A global feature importance diagram provided by the present application is shown in Figure 3 As shown, the bar chart shows the average |SHAP| value of the features, and the results show that the GRE1 gene, CTA1, and HSP104 are the top 3 contributing features.
[0089] Specifically, the probability score is used to guide industrial strain screening, and candidate strains suitable for high ethanol fermentation can be screened out.
[0090] S105: The unlabeled wild yeast strain corresponding to the high ethanol tolerance probability score greater than or equal to the preset threshold value is taken as the high ethanol tolerance yeast strain.
[0091] Specifically, the preset threshold is set according to specific engineering practice. In an industrial application example, the trained random forest model is applied to the multi-omics data (including gene sequencing, proteomics and metabolomics data) of 200 unannotated wild yeast strains, and the HT probability score (threshold >= 0.8) of each strain is output by the `predict()` function. Finally, 19 high-probability HT candidate strains are screened out; through small-scale fermentation verification (10% ethanol stress, 30°C culture for 48 hours), Figure 4 The schematic diagram of the comparison of the ethanol production rate of the HT strain provided by the present application and the random strain is shown in Figure 4 The ethanol production rate of these strains is 11.8±0.8 g / L on average, which is significantly improved by 18% (independent sample T test, p<0.01) compared with the random screening control group (10.0±1.2 g / L), and the expression of key tolerance genes (such as HSP104) is up-regulated by 2.1-3.5 times, which confirms that the model-guided directional screening can effectively improve the performance of the industrial strain.
[0092] In an exemplary embodiment, in the independent test set verification, the multi-omics data (including 7 marker characteristics: GRE1 gene, HSP104, etc.) of 50 known HT / LT phenotype S. cerevisiae strains pre-reserved as input are loaded into the trained random forest model for prediction, the ROC curve of the predicted probability and the true label is calculated by the roc() function of the pROC package, the result shows that the AUC reaches 0.88 (95% CI: 0.82-0.93), the sensitivity is 0.85, and the specificity is 0.83, and the ROC curve is drawn using ggplot2 Figure 5 , Figure 5 The schematic diagram of the tolerance prediction ROC curve of the independent test set (50 strains) provided by the present application is shown in Figure 5 The ROC curve is shown in the figure, the horizontal axis is the false positive rate, the vertical axis is the true positive rate, the diagonal line is used as a random reference, and the area under the curve is significantly higher than that of random guessing (DeLong test p=2.3x10^-6), which confirms that the model has stable prediction ability for unknown strains. Ethanol tolerance, and subsequent decision curve analysis (DCA) proves its clinical application value in industrial strain screening.
[0093] In an exemplary embodiment, the present application provides a technical route schematic diagram of the construction of the multi-omics prediction model of S. cerevisiae ethanol tolerance as shown in Figure 6 The schematic diagram of the construction of the multi-omics prediction model of S. cerevisiae ethanol tolerance is shown in Figure 6As shown, the Saccharomyces cerevisiae ethanol tolerance multi-omics prediction model includes a data collection stage, a feature engineering stage, and a modeling application stage. In the data collection stage, strain screening and culture and multi-omics data collection are performed to obtain sample label data and multi-omics quantitative data; in the feature engineering stage, multi-omics data standardization and feature selection are performed to obtain a standardized feature matrix and key biomarkers; in the modeling application stage, a feature matrix is constructed and a machine learning model is constructed to determine the optimal model as RF and determine the SHAP value representing feature importance; wild strain screening and ethanol yield detection are performed according to the optimal model and the SHAP value.
[0094] In the application of the determination method of a high-ethanol-tolerant yeast strain based on multi-omics markers provided by the present application, the determination method can be performed without Figure 1 The order of execution of each step shown can be determined as needed, and the present application does not limit the order of execution of each step.
[0095] In view of the above technical status, the advantages of the present application are as follows: Multi-level data integration: The present application integrates genomics, proteomics, and metabolomics data to provide more comprehensive biological information and improve the accuracy and reliability of prediction.
[0096] Application of machine learning: By applying machine learning algorithms, the present application can extract key features from multi-omics data and construct an efficient prediction model, overcoming the limitations of traditional methods in data processing and prediction ability.
[0097] Industrial application orientation: The present application focuses on providing a prediction tool that can be directly applied to industrial strain screening, fermentation process optimization, and synthetic biology modification, with output tolerance scores that can be directly used for production, improving the practicality and operability of the technology.
[0098] Improved screening efficiency and reduced cost: Through the prediction of the machine learning model, the experimental workload and time cost are reduced, improving the efficiency of strain screening and modification.
[0099] The above is a determination method of a high-ethanol-tolerant yeast strain based on multi-omics markers provided by one or more embodiments of the present application. Based on the same idea, the present application also provides a corresponding determination device of a high-ethanol-tolerant yeast strain based on multi-omics markers, as shown in Figure 7 .
[0100] Figure 7 A determination device of a high-ethanol-tolerant yeast strain based on multi-omics markers provided by the present application is shown in the schematic diagram, which comprises: The acquisition module 701 is configured to acquire the gene expression, protein expression profile, and metabolite abundance of a plurality of unlabeled wild yeast strains.
[0101] The alignment module 702 is used for Z-score standardization of the gene expression, protein expression profile and metabolite abundance respectively, and aligning the three groups of data after Z-score standardization according to features.
[0102] The first determination module 703 is used for extracting the standardized values of the GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine and trehalose-6-phosphate in the three groups of data after feature alignment, and composing a plurality of biomarker feature combinations according to the extracted standardized values.
[0103] The screening module 704 is used for inputting the plurality of biomarker feature combinations into a strain screening model for industrial strain screening to obtain a high ethanol tolerance probability score of each of the plurality of unlabeled wild yeast strains.
[0104] The second determination module 705 is used for taking the unlabeled wild yeast strain corresponding to the high ethanol tolerance probability score greater than or equal to a preset threshold as a high ethanol tolerance yeast strain.
[0105] The specific limitations of the device for determining a high ethanol tolerance yeast strain based on a plurality of omics markers can refer to the limitations of the method for determining a high ethanol tolerance yeast strain based on a plurality of omics markers, which will not be repeated here. Each module in the device for determining a high ethanol tolerance yeast strain based on a plurality of omics markers can be realized by software, hardware and combinations thereof. Each module can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to each module.
[0106] The application further provides a computer readable storage medium, which stores a computer program, and the computer program can be used to execute the above-mentioned Figure 3 The application provides a method for determining a high ethanol tolerance yeast strain based on a plurality of omics markers.
[0107] The application further provides a computer device, which can be used to execute the above-mentioned Figure 8 The structure of the computer device is shown in FIG. 1. Figure 8 As shown in FIG. 1, at the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory and a non-volatile memory, and can further include other hardware required by other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs to realize the above-mentioned Figure 1 The application provides a method for determining a high ethanol tolerance yeast strain based on a plurality of omics markers.
[0108] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiments of the methods. In the embodiments of the present application, any reference to memory, storage, database or other medium can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0109] The technical features of the above embodiments can be combined in any way. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, but as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.
Claims
1. A method for identifying highly ethanol-tolerant yeast strains based on multi-omics biomarkers, characterized in that, include: Gene expression levels, protein expression profiles, and metabolite abundances of multiple unlabeled wild-type yeast strains were obtained; The gene expression level, the protein expression profile, and the metabolite abundance were Z-score normalized respectively, and the three sets of data after Z-score normalization were aligned according to features; The standardized values of GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine and trehalose-6-phosphate were extracted from the three sets of data after feature alignment. Based on the extracted standardized values, multiple sets of biomarker feature combinations were formed. The multiple sets of biomarker features are combined and input into the strain screening model to screen industrial strains, and a high ethanol tolerance probability score is obtained for each of the multiple unlabeled wild yeast strains. Unlabeled wild yeast strains with a high ethanol tolerance probability score greater than or equal to a preset threshold are considered as high ethanol-tolerant yeast strains.
2. The method as described in claim 1, characterized in that, The process of combining the features of the multiple sets of biomarkers and inputting them into a strain screening model for industrial strain screening, to obtain a high ethanol tolerance probability score for each of the multiple unlabeled wild yeast strains, specifically includes: The Z-score values of the seven biomarkers in the multiple sets of biomarker feature combinations are arranged in order as the input vector; In the strain screening model, each tree is traversed independently. Each tree starts from the root node and splits according to the feature threshold. Unlabeled wild yeast strains are assigned to specific leaf nodes, and each leaf node corresponds to a classification label. The proportion of high ethanol-tolerant yeast strains in a specific leaf node during training is recorded as a local probability. The local probabilities of all trees are weighted and averaged to obtain a weighted average result. The weighted average result is determined as the high ethanol tolerance probability score of the unlabeled wild yeast strain; the weight of the weighted average is the precision of the tree.
3. The method as described in claim 1, characterized in that, The training process of the strain screening model specifically includes: Obtain the sample feature matrix and divide it into a training set and a validation set; The training set is used to train random forest model, support vector machine model, gradient boosting decision tree model, naive Bayes model, neural network model and generalized linear model, and the parameters of all models are adjusted until all models are optimal. For each of all models, the validation set is input into the model to obtain the model selection results. The model selection results are compared with the true categories of the test set to obtain the model's metrics. The metrics include prediction accuracy, sensitivity, and specificity. Compare and analyze the metrics of all models, and determine the optimal model based on the results of the comparative analysis; The optimal model is determined as the strain screening model.
4. The method as described in claim 3, characterized in that, The acquisition of the sample feature matrix specifically includes: Multiple high-ethanol-tolerant yeast strains and multiple low-ethanol-tolerant yeast strains were obtained as a sample training set; Gene expression levels, protein expression profiles, and metabolite abundances of each yeast strain in the training set were obtained at different ethanol concentrations. Z-score normalization was performed on the gene expression level, protein expression profile, and metabolite abundance of the sample to obtain three sets of data. Feature selection was performed on the three sets of data after Z-score standardization to obtain a sample feature matrix; the sample feature matrix includes GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine and trehalose-6-phosphate.
5. The method as described in claim 4, characterized in that, The step of performing feature selection on the three sets of data after Z-score standardization to obtain a sample feature matrix specifically includes: The data after Z-score normalization were aligned across omics data using UniProt ID / KEGG ID; After performing the cross-omics data alignment, LASSO regression was performed to filter out multiple features; Based on the selected features, an n×7 dimensional matrix is generated, with each yeast strain corresponding to one row of the feature matrix and each biomarker corresponding to one column. This n×7 dimensional matrix is then determined as the sample feature matrix. n This represents the number of yeast strains in the sample training set.
6. The method as described in claim 4, characterized in that, The screening process for multiple high-ethanol-tolerant yeast strains and multiple low-ethanol-tolerant yeast strains specifically includes: Candidate strains were selected from the wild-type Saccharomyces cerevisiae strain library; Prepare a culture medium containing ethanol; The candidate strains were cultured in an ethanol-containing medium; After the culture, the candidate strains are initially screened and then screened again. Based on the screening criteria for high ethanol-tolerant strains and low ethanol-tolerant strains, the strains obtained from the second screening are further screened to obtain multiple high ethanol-tolerant yeast strains and multiple low ethanol-tolerant yeast strains.
Citation Information
Patent Citations
Method for creating distiller's yeast ethanol high yield bacterial strain
CN101323837A
Acetic acid resistant ethanol producing wine making yeast strains and strain screening method
CN102146345A
Saccharomyces cerevisiae 6-phosphoric acid trehalose synthase genes as well as expression vectors and application thereof
CN103146721A
Saccharomyces cerevisiae engineering strain as well as preparation method, application and fermentation culture method thereof
CN105087407A
Recombinant saccharomyces cerevisiae engineering strain for high yield of squalene by improving ethanol tolerance and application
CN115786154A