A method for determining a high-ethanol-tolerant yeast strain based on multi-omics markers
By integrating gene expression levels, proteomics, and metabolomics data, a multi-omics biomarker screening model was constructed, which solved the problem of time-consuming and labor-intensive traditional yeast strain screening, and achieved efficient and accurate yeast strain screening to meet the needs of industrial production.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIYANG COLLEGE OF ZHEJIANG A & F UNIV
- Filing Date
- 2025-05-27
- Publication Date
- 2026-05-12
AI Technical Summary
Traditional yeast strain screening methods are time-consuming and labor-intensive, resulting in high screening costs. Furthermore, a single genomic study cannot fully reflect the yeast response mechanism at high ethanol concentrations, affecting screening efficiency and accuracy.
Using a multi-omics biomarker approach, gene expression, proteomics, and metabolomics data were integrated, and a strain screening model was constructed using Z-score normalization and machine learning algorithms to screen out yeast strains with high ethanol tolerance.
It reduces the cost of yeast strain screening, improves screening efficiency and accuracy, provides more comprehensive biological information, and guides the selection of industrial strains and the optimization of fermentation processes.
Smart Images

Figure CN120998302B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a method for identifying highly ethanol-tolerant yeast strains based on multi-omics biomarkers. Background Technology
[0002] Saccharomyces cerevisiae is the core microorganism for industrial ethanol fermentation. During ethanol fermentation, increasing ethanol concentration stresses yeast cell growth. To survive and grow, yeast cells develop corresponding stress responses to cope with this stress; this mechanism is known as ethanol tolerance. Ethanol tolerance directly determines the fermentation efficiency and product concentration of Saccharomyces cerevisiae. However, high concentrations of ethanol can damage cell membrane integrity, inhibit enzyme activity, and trigger oxidative stress, leading to cell growth arrest or even death. It is generally believed that ethanol's effects on yeast cells are mainly reflected in three aspects: inhibiting cell growth, cell survival, and fermentation. Therefore, ethanol tolerance is often defined by its effect on cell growth. At room temperature, the highest ethanol concentration allowed for growth after 48 hours of cultivation in a medium containing 1%–14% ethanol represents the yeast's ethanol tolerance level. Strains that can grow in 3%–6% ethanol have poor ethanol tolerance, 6%–10% has moderate tolerance, and 10%–13% has high tolerance. This definition method is relatively simple and is often used to screen for ethanol-tolerant strains. Improving yeast's ethanol tolerance is of great significance for increasing fermentation efficiency, yield, and the quality of the final product.
[0003] Traditional yeast strain screening methods rely on culturing strains in media with different ethanol concentrations and determining the ethanol tolerance of yeast strains by measuring parameters such as growth curves, maximum specific growth rate, and final cell density. However, this method requires extensive experimental procedures, is time-consuming and labor-intensive, resulting in high screening costs.
[0004] Therefore, there is an urgent need for a method to reduce the cost of yeast strain screening. Summary of the Invention
[0005] Therefore, it is necessary to provide a method for identifying highly ethanol-tolerant yeast strains based on multi-omics biomarkers to address the aforementioned technical problems. This method can reduce the screening cost of yeast strains.
[0006] The present invention adopts the following technical solution:
[0007] This invention provides a method for identifying highly ethanol-tolerant yeast strains based on multi-omics biomarkers, comprising:
[0008] Gene expression levels, protein expression profiles, and metabolite abundances of multiple unlabeled wild-type yeast strains were obtained;
[0009] Gene expression levels, protein expression profiles, and metabolite abundance were Z-score normalized, and the three sets of data after Z-score normalization were aligned according to features.
[0010] The standardized values of GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine and trehalose-6-phosphate were extracted from the three sets of data after feature alignment. Based on the extracted standardized values, multiple sets of biomarker feature combinations were formed.
[0011] Multiple sets of biomarker features were combined and input into the strain screening model to screen industrial strains, resulting in a high ethanol tolerance probability score for each of the multiple unlabeled wild yeast strains.
[0012] Unlabeled wild yeast strains with a high ethanol tolerance probability score greater than or equal to a preset threshold are considered as high ethanol-tolerant yeast strains.
[0013] Preferably, multiple sets of biomarker features are combined and input into a strain screening model for industrial strain screening to obtain a high ethanol tolerance probability score for each of the multiple unlabeled wild yeast strains, specifically including:
[0014] The Z-score values of the seven biomarkers in the multiple sets of biomarker feature combinations are arranged in order as the input vector;
[0015] In the strain screening model, each tree is traversed independently. Each tree starts from the root node and splits according to the feature threshold. Unlabeled wild yeast strains are assigned to specific leaf nodes, and each leaf node corresponds to a classification label.
[0016] The proportion of high ethanol-tolerant yeast strains in a specific leaf node during training is recorded as a local probability.
[0017] The local probabilities of all trees are weighted and averaged to obtain the weighted average result. The weighted average result is determined as the high ethanol tolerance probability score of the unlabeled wild yeast strains; the weight of the weighted average is the precision of the tree.
[0018] Preferably, the training process of the strain screening model specifically includes:
[0019] Obtain the sample feature matrix and divide it into a training set and a validation set;
[0020] The random forest model, support vector machine model, gradient boosting decision tree model, naive Bayes model, neural network model, and generalized linear model are trained using the training set, and the parameters of all models are adjusted until all models are optimal.
[0021] For each of all models, the validation set is input into the model to obtain the model selection results. The model selection results are compared with the true class in the test set to obtain the model's metrics. The metrics include prediction accuracy, sensitivity, and specificity.
[0022] Compare and analyze the metrics of all models, and determine the optimal model based on the results of the comparative analysis;
[0023] The optimal model was determined to be the strain screening model.
[0024] Preferably, obtaining the sample feature matrix specifically includes:
[0025] Multiple high-ethanol-tolerant yeast strains and multiple low-ethanol-tolerant yeast strains were obtained as a sample training set;
[0026] Gene expression levels, protein expression profiles, and metabolite abundances of each yeast strain in the training set were obtained at different ethanol concentrations.
[0027] Z-score normalization was performed on the gene expression level, protein expression profile, and metabolite abundance of the samples to obtain three sets of data.
[0028] Feature selection was performed on the three sets of data after Z-score standardization to obtain the sample feature matrix. The sample feature matrix includes GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine, and trehalose-6-phosphate.
[0029] Preferably, feature selection is performed on the three sets of data after Z-score standardization to obtain a sample feature matrix, specifically including:
[0030] The data after Z-score normalization were aligned across omics data using UniProt ID / KEGG ID;
[0031] After cross-omics data alignment, LASSO regression was performed to filter out multiple features;
[0032] Based on the selected features, an n×7 dimensional matrix is generated, with each yeast strain corresponding to one row of the feature matrix and each biomarker corresponding to one column. This n×7 dimensional matrix is then used as the sample feature matrix. n This represents the number of yeast strains in the sample training set.
[0033] Preferably, the screening process for multiple high-ethanol-tolerant yeast strains and multiple low-ethanol-tolerant yeast strains specifically includes:
[0034] Candidate strains were selected from the wild-type Saccharomyces cerevisiae strain library;
[0035] Prepare a culture medium containing ethanol;
[0036] The candidate strains were cultured in an ethanol-containing medium;
[0037] After cultivation, the candidate strains were initially screened and then screened again. Based on the screening criteria for high ethanol-tolerant strains and low ethanol-tolerant strains, the strains obtained from the second screening were further screened to obtain multiple high ethanol-tolerant yeast strains and multiple low ethanol-tolerant yeast strains.
[0038] This invention provides a device for identifying highly ethanol-tolerant yeast strains based on multi-omics biomarkers, comprising:
[0039] The acquisition module is used to acquire gene expression levels, protein expression profiles, and metabolite abundance of multiple unlabeled wild yeast strains;
[0040] The alignment module is used to perform Z-score normalization on gene expression levels, protein expression profiles, and metabolite abundance, and then align the three sets of data after Z-score normalization according to features.
[0041] The first determination module is used to extract the standardized values of GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine and trehalose-6-phosphate from the three sets of data after feature alignment, and to form multiple sets of biomarker feature combinations based on the extracted standardized values.
[0042] The screening module is used to combine multiple sets of biomarker features and input them into the strain screening model to screen industrial strains, and obtain a high ethanol tolerance probability score for each of the multiple unlabeled wild yeast strains.
[0043] The second determination module is used to identify unlabeled wild yeast strains with a high ethanol tolerance probability score greater than or equal to a preset threshold as high ethanol tolerance yeast strains.
[0044] The present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for determining highly ethanol-tolerant yeast strains based on multi-omics biomarkers.
[0045] The present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the above-mentioned method for determining a highly ethanol-tolerant yeast strain based on multi-omics biomarkers.
[0046] The above-mentioned at least one technical solution adopted in this invention can achieve the following beneficial effects:
[0047] This study obtained gene expression levels, protein expression profiles, and metabolite abundance data from multiple unlabeled wild-type yeast strains, integrating gene expression, proteomics, and metabolomics data to provide more comprehensive biological information and improve prediction accuracy and reliability. Z-score normalization was applied to gene expression levels, protein expression profiles, and metabolite abundance, and the three sets of data after Z-score normalization were aligned by features. Normalized values of GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine, and trehalose-6-phosphate were extracted from the feature-aligned data, and multiple sets of biomarker feature combinations were formed based on the extracted normalized values. These multiple sets of biomarker feature combinations were input into a strain screening model for industrial strain screening, obtaining a high ethanol tolerance probability score for each of the multiple unlabeled wild-type yeast strains. This constructed an efficient strain screening model, overcoming the limitations of traditional methods in data processing and prediction capabilities. Unlabeled wild-type yeast strains with a high ethanol tolerance probability score greater than or equal to a preset threshold were identified as high ethanol-tolerant yeast strains. This method can reduce the screening cost of yeast strains. Attached Figure Description
[0048] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:
[0049] Figure 1 A schematic flowchart of a method for determining highly ethanol-tolerant yeast strains based on multi-omics biomarkers provided by the present invention;
[0050] Figure 2 A performance comparison chart of six machine learning classifiers provided by this invention;
[0051] Figure 3 A schematic diagram illustrating the global feature importance provided by this invention;
[0052] Figure 4 This is a schematic diagram comparing the ethanol yield of the HT strain and the random strain provided by the present invention.
[0053] Figure 5 A schematic diagram of the ROC curve for predicting tolerance of the independent test set (50 strains) provided by this invention;
[0054] Figure 6 A schematic diagram of the technical route for constructing the multi-omics prediction model for ethanol tolerance in Saccharomyces cerevisiae provided by the present invention;
[0055] Figure 7 A schematic diagram of a device for identifying highly ethanol-tolerant yeast strains based on multi-omics biomarkers provided by the present invention;
[0056] Figure 8 This is a schematic diagram of a computer device for implementing a method for determining highly ethanol-tolerant yeast strains based on multi-omics biomarkers, as provided by the present invention. Detailed Implementation
[0057] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0058] Traditional laboratory screening of yeast strains involves culturing yeast strains in ethanol media of varying concentrations and determining their ethanol tolerance using parameters such as growth curves, maximum specific growth rate, and final cell density. This method requires extensive experimental work, resulting in long screening cycles, low efficiency, and high costs.
[0059] Genetic engineering can specifically enhance ethanol tolerance in yeast through gene knockout, gene overexpression, or gene editing technologies such as CRISPR-Cas9. However, this approach requires a deep understanding of tolerance-related genes and raises concerns about gene stability and transgenic safety. Therefore, elucidating the molecular mechanisms of ethanol tolerance and developing precise prediction methods are of great significance for the selection of industrial strains, optimization of fermentation processes, and synthetic biology modification.
[0060] Researchers identified over 250 genes associated with ethanol tolerance through genome-wide screening, focusing on vacuolar, peroxisome, cell wall, and membrane-related pathways, and verified the role of the FPS1 gene (glycerol channel protein) in ethanol efflux. Combining transcriptomic and proteomic analyses, they discovered significant changes in mitochondrial and endoplasmic reticulum function, signal transduction (such as the GPCR pathway), and aromatic amino acid metabolism in yeast under ethanol stress. Through multi-omics integration (transcriptomics, proteomics, and metabolomics) and network analysis, they proposed a differentiation mechanism for high tolerance (HT) and low tolerance (LT) phenotypes, emphasizing the regulatory roles of longevity pathways, peroxisome function, energy metabolism, and genes, and established a biological mechanism called the "ethanol stress buffer model." Existing studies have only elucidated the mechanisms of ethanol tolerance through single or multiple omics (genomics, transcriptomics, proteomics, and metabolomics) without addressing the construction of predictive models.
[0061] Single-omics studies, such as those using genomics, transcriptomics, proteomics, or metabolomics to investigate yeast's response to high ethanol concentrations, primarily focus on understanding how a single level of biological information affects ethanol tolerance. Yeast cell responses involve multiple levels, including genes, transcription, translation, protein function, and metabolism; single-omics data cannot encompass all these levels. Single-omics studies only provide partial biological information and cannot comprehensively reflect yeast's response mechanisms at high ethanol concentrations, thus limiting the accuracy and comprehensiveness of predictive results. Multi-omics integrated studies, such as those combining genomics, transcriptomics, proteomics, and metabolomics, offer a more comprehensive understanding of yeast's response mechanisms at high ethanol concentrations. While multi-omics integration provides a more comprehensive perspective, these studies primarily focus on mechanistic exploration and do not further utilize this data for predictive assessment.
[0062] Furthermore, traditional methods are limited by experimental conditions, resulting in limited screening efficiency and accuracy. In recent years, single-mathematical studies have attempted to elucidate the mechanisms of ethanol tolerance, but due to the limited data dimensions, they cannot fully reflect the complex biological processes, thus restricting the application of efficient screening.
[0063] The present invention aims to solve the above-mentioned technical problems and provide a method for predicting ethanol tolerance of Saccharomyces cerevisiae by integrating gene expression levels, proteomics, metabolomics data and machine learning algorithms, so as to improve the efficiency and accuracy of strain screening, reduce experimental workload, and guide synthetic biology modification and fermentation process optimization, thereby meeting the needs of industrial production.
[0064] The technical solutions provided by the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0065] Devices such as desktop computers, servers, and laptops are capable of executing the solutions of this invention. For ease of explanation, the following description will focus on servers as the executing entity.
[0066] Figure 1 This is a schematic flowchart of a method for determining highly ethanol-tolerant yeast strains based on multi-omics biomarkers according to the present invention, which specifically includes the following steps:
[0067] S101: Obtain gene expression levels, protein expression profiles, and metabolite abundances from multiple unlabeled wild yeast strains.
[0068] S102: Z-score normalization was performed on gene expression levels, protein expression profiles, and metabolite abundance, and the three sets of data after Z-score normalization were aligned according to features.
[0069] Cross-omics data alignment: Gene-protein-metabolite modules were constructed using the Weighted Gene Co-expression Network Analysis (WGCNA) package in R. These modules, identified through co-expression network analysis, consist of a set of genes, proteins, and metabolites highly correlated with ethanol concentration, reflecting potential regulatory networks or metabolic pathways. Network type: distinguishing between positive and negative regulation (signed hybrid); module definition: minimum number of genes = 30, cut height = 0.25; modules significantly correlated with ethanol concentration were selected (|r|>0.7, p<0.01).
[0070] S103: The standardized values of GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine and trehalose-6-phosphate of the three sets of data after feature alignment were extracted, and multiple sets of biomarker feature combinations were formed based on the extracted standardized values.
[0071] S104: Combine multiple sets of biomarker features and input them into the strain screening model to screen industrial strains, and obtain the high ethanol tolerance probability score for each of the multiple unlabeled wild yeast strains.
[0072] In an exemplary embodiment, the screening process for multiple high-ethanol-tolerant yeast strains and multiple low-ethanol-tolerant yeast strains specifically includes: selecting candidate strains from a wild-type Saccharomyces cerevisiae strain library; preparing an ethanol-containing culture medium; culturing the candidate strains in the ethanol-containing culture medium; after culturing, performing primary and secondary screening on the candidate strains, and screening the strains obtained from the secondary screening according to the screening criteria for high-ethanol-tolerant strains and low-ethanol-tolerant strains to obtain multiple high-ethanol-tolerant yeast strains and multiple low-ethanol-tolerant yeast strains.
[0073] Specifically, the screening process for multiple high-ethanol-tolerant yeast strains and multiple low-ethanol-tolerant yeast strains is as follows:
[0074] I. Strains and Pre-culture
[0075] Strain library: Candidate strains were selected from a library of wild-type Saccharomyces cerevisiae preserved in the laboratory.
[0076] Pre-culture conditions: Each strain was thawed from the glycerol cryopreservation tube and inoculated into YPD liquid medium (20 g / L glucose, 20 g / L tryptone, 10 g / L yeast extract) and cultured at 28℃ with shaking at 160 rpm for 12 h until the logarithmic growth phase (OD600≈ 0.7).
[0077] II. Preparation of Ethanol-Containing Culture Medium
[0078] Culture medium formulation: Basic YPD liquid medium (same as pre-culture), sterilized and cooled, then added with sterile anhydrous ethanol to a final concentration of 13% (v / v).
[0079] Control medium: YPD medium without ethanol, used to evaluate the basic growth performance of the strain.
[0080] III. Initial Screening: Solid Plate Tolerance Test
[0081] Serial dilution: The pre-cultured bacterial solution was serially diluted to 10^-3~10^-5, and 100 μL was spread on YPD solid plates containing 13% ethanol and control plates without ethanol.
[0082] Culture and observation: Incubate at 28℃ for 48-72 h. Screening criteria:
[0083] Initial screening of HT candidate strains: forming distinct colonies (colon diameter ≥ 1 mm) on 13% ethanol plates, with a colony count ≤ 50% different from the control group.
[0084] Initial screening of LT candidate strains: no visible colonies or a colony count reduced by ≥90% on 13% ethanol plates.
[0085] IV. Secondary Screening: Analysis of Liquid Culture Growth Curves
[0086] Inoculation and culture: The HT and LT candidate strains obtained from the initial screening were inoculated into YPD liquid medium containing 13% ethanol at a 3% inoculation rate (OD600 ≈ 0.7), while an ethanol-free control group was set up.
[0087] OD600 dynamic monitoring: cultured at 28℃ and 160 rpm with shaking, OD600 value was measured every 2 h until 24 h, and growth curve was plotted.
[0088] V. Screening Criteria:
[0089] HT strain: In 13% ethanol, the lag phase is ≤6 h, the OD600 growth rate during the logarithmic growth phase is ≥0.1 / h, and the final OD600 value is ≥60% of the control group.
[0090] LT strain: lag phase ≥12 h, OD600 growth rate ≤0.05 / h, final OD600 value ≤20% of control group.
[0091] VI. Tolerance Testing
[0092] 1. Cell viability assay:
[0093] The HT and LT candidates were cultured in 13% ethanol medium for 24 h, and then serially diluted and plated on ethanol-free YPD plates.
[0094] Calculate the survival rate: Survival rate (%) = (CFU in the ethanol treatment group / CFU in the control group) × 100.
[0095] HT standard: survival rate ≥ 50%; LT standard: survival rate ≤ 10%.
[0096] 2. Cell morphology observation:
[0097] The bacterial culture treated with 13% ethanol for 24 h was collected by centrifugation, fixed with 2.5% glutaraldehyde, and observed by scanning electron microscopy.
[0098] HT strains: intact cell morphology (oval, smooth surface); LT strains: ruptured or severely shrunken cells.
[0099] Data collection and training set construction:
[0100] Ten strains (biological replicates) were selected from each of the HT / LT group and five ethanol concentrations as training set samples.
[0101] a. Gene sequencing: RNA-Seq technology was used to perform high-throughput sequencing of gene expression mRNA in Saccharomyces cerevisiae cells at different ethanol concentrations to screen for genes that may be related to ethanol tolerance.
[0102] (1) Experimental design:
[0103] Cells in the mid-logarithmic growth phase (OD600≈0.6) of the training set strains were collected and rapidly fixed using an RNA protectant (such as RNAlater). Strand-specific libraries were prepared (preserving strand orientation information). Sequencing was performed using an Illumina NovaSeq 6000 in PE150 mode, requiring ≥20M clean reads / sample (to ensure detection of low-abundance genes).
[0104] (2) Bioinformatics analysis:
[0105] The raw sequencing data were quality controlled using FASTP software, and low-quality sequences and sequencing adapter sequences were filtered out. Hisat2 software was used to perform sequence alignment with the *Saccharomyces cerevisiae* S288C genome (NCBI RefSeq ID: GCF_000146045.2) as a reference genome. StringTie software was used to assemble and quantify transcripts, outputting the TPM quantification for each gene. b. Quantitative proteomics analysis: Protein expression profiles of *Saccharomyces cerevisiae* under different ethanol concentrations were analyzed using liquid chromatography-mass spectrometry (LC-MS).
[0106] (1) Experimental design:
[0107] In quantitative proteomics analysis, Saccharomyces cerevisiae cells were lysed, reduced and alkylated, digested with Trypsin Gold, and peptides were labeled with TMT 16plex. The peptides were fractionated by high-pH reversed-phase chromatography and detected by DDA mode using a Q Exactive HF-X mass spectrometer (MS1 120,000 resolution, MS2 45,000 resolution).
[0108] (2) Data Analysis:
[0109] Raw data were processed using MaxQuant (v2.1.0) software with the Andromeda search engine. Parameters were set as follows: precursor ion mass tolerance 4.5 ppm, fragment ion tolerance 20 ppm, enzyme digestion mode set to Trypsin / P (allowing a maximum of 2 missed cleavage sites), fixed modifications were TMT 16plex (N-terminus and lysine) and cysteine alkylation (+57.021 Da), and variable modifications included methionine oxidation (+15.995 Da) and protein N-terminal acetylation (+42.011 Da). The UniProt Saccharomyces cerevisiae reference protein library (May 2024 version, including common contaminants) was used as the database. The FDR <1% calculated using the reverse bait library was used as the identification threshold.
[0110] c. Quantitative metabolomics analysis: Mass spectrometry analysis was used to obtain the abundance changes of metabolites in the fermentation broth under different ethanol concentrations.
[0111] (1) Experimental Design
[0112] Sample pretreatment involved quenching metabolic activity with pre-cooled methanol-acetonitrile-water (4:4:2, v / v / v), followed by ultrasonic-assisted extraction (300W, 2s / 3s pulses, 5 min), centrifugation (15,000g, 15 min, 4℃), and concentration of the supernatant by nitrogen blowing. An ultra-high performance liquid chromatography (UPLC) system was used, equipped with two types of columns:
[0113] Polar metabolites: ACQUITY UPLC BEH Amide column (2.1×100mm, 1.7μm), mobile phase 25mM ammonium acetate + ammonia (pH9.0) / acetonitrile.
[0114] Nonpolar metabolites: An ACQUITY UPLC HSS T3 column (2.1 × 100 mm, 1.7 μm) was used with a mobile phase of 0.1% formic acid / acetonitrile. Mass spectrometry was performed using a Q-TOF 6600+ system with ESI switching between positive and negative ion modes. Parameter settings: ion source temperature 500℃, spray voltage ±5.5 kV, collision energy 20-40 V, mass range m / z 50-1500, resolution >30,000. An internal standard (L-2-chlorophenylalanine) was added to each batch of experiments to monitor data stability.
[0115] Raw mass spectrometry data underwent peak extraction, alignment, and normalization (internal standard correction) using Progenesis QI. Metabolite identification was performed using precise mass numbers (error <5 ppm) and secondary spectrum matching (METLIN / HMDB database, similarity >80%).
[0116] Data preprocessing:
[0117] Standardization: Due to the significant differences in the dimensions of different omics data (e.g., gene TPM values range from 0 to 10^5, and metabolite peak areas range from 10^3 to 10^8), this invention uses R language to perform Z-score standardization on gene expression levels, protein and metabolite abundance to eliminate biases from different omics quantification methods and experimental batches.
[0118] Dimensionality reduction: Principal component analysis (PCA) is applied to reduce the dimensionality of the data while retaining key information.
[0119] In an exemplary embodiment, the training process of the strain screening model specifically includes: obtaining a sample feature matrix and dividing the sample feature matrix into a training set and a validation set; training a random forest model, a support vector machine model, a gradient boosting decision tree model, a naive Bayes model, a neural network model, and a generalized linear model using the training set, and adjusting the parameters of all models until all models are optimal; for each of all models, inputting the validation set into the model to obtain the model's screening results, comparing the model's screening results with the true categories of the test set to obtain the model's metrics; the metrics include prediction accuracy, sensitivity, and specificity; performing comparative analysis on the metrics of all models, and determining the optimal model based on the comparative analysis results; and determining the optimal model as the strain screening model.
[0120] In an exemplary embodiment, obtaining the sample feature matrix specifically includes: obtaining multiple high-ethanol-tolerant yeast strains and multiple low-ethanol-tolerant yeast strains as a sample training set; obtaining the sample gene expression level, sample protein expression profile, and sample metabolite abundance for each yeast strain in the sample training set at different ethanol concentrations; performing Z-score normalization on the sample gene expression level, sample protein expression profile, and sample metabolite abundance to obtain three sets of data; performing feature selection on the three sets of data after Z-score normalization to obtain the sample feature matrix; the sample feature matrix includes the GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine, and trehalose-6-phosphate.
[0121] Specifically, this invention utilizes the caret package in R language to train six machine learning algorithms in parallel to comprehensively evaluate model performance. The six machine learning algorithms are as follows:
[0122] Random Forest (RF): Constructs multiple decision trees based on Bootstrap sampling and outputs classification results through a voting mechanism. The default number of trees is set to 500. It is suitable for high-dimensional features and is resistant to overfitting.
[0123] Support Vector Machine (SVM): It selects a radial basis kernel function and achieves non-linear classification by maximizing the classification margin. The penalty parameter C and the kernel parameter γ need to be adjusted.
[0124] Gradient Boosting Decision Tree (GBDT): Iteratively optimizes the residuals of the decision tree, with a default learning rate of 0.1 and a tree depth of 3. It is suitable for scenarios where there are complex interactions between features.
[0125] Naive Bayes Model (NBM): Based on Bayes' theorem, it assumes that features are independent and is suitable for small sample data. The probability threshold is set to 0.5.
[0126] Neural Network (NN): Construct a fully connected network with a single hidden layer. The number of nodes in the hidden layer is 1.5 times the number of features. Use ReLU as the activation function and Adam as the optimizer.
[0127] Generalized Linear Model (GLM): Uses Logistic regression (binary classification), L2 regularization to control coefficient weights, and logit as the link function.
[0128] The input data for the six models needs to be standardized (Z-score), the class labels (HT / LT) are encoded as 1 / 0, and the training set and test set are stratified in a 7:3 ratio.
[0129] The six models were optimized.
[0130] The pROC package, based on the R language, optimizes the hyperparameters of the model through 10-fold cross-validation, reducing overfitting and improving the model's generalization ability.
[0131] 1. The hyperparameter tuning range is:
[0132] RF: Adjust mtry (the number of features randomly selected from each tree, ranging from 1 to 7) and ntree (100 to 500).
[0133] SVM: Grid search C (0.1-10) and γ (0.01-0.1).
[0134] GBDT: Optimizes nrounds (number of iterations, 50-200), max_depth (3-6), and eta (learning rate, 0.01-0.3).
[0135] NN: Adjust the number of hidden layer nodes (5-15) and dropout rate (0.2-0.5).
[0136] GLM: Regularization parameter lambda (0.001-1).
[0137] 2. The cross-validation process is as follows:
[0138] The training set was divided into 10 equal parts, and 9 parts were used for training and 1 part for validation, which was repeated 10 times.
[0139] Record the AUC, sensitivity (recall), and specificity for each fold, and take the average value as the basis for hyperparameter selection.
[0140] Early stopping strategies prevent overfitting; for example, GBDT iterations terminate when the loss no longer decreases.
[0141] 3. Performance Comparison and Model Selection:
[0142] Figure 2 The performance comparison chart of the six machine learning classifiers provided in this invention is as follows: Figure 2 As shown, when comparing the six models from three dimensions of prediction accuracy, sensitivity, and specificity, the RF model significantly outperformed the other models in ROC, sensitivity, and specificity (p<0.05, paired t-test), and had high training efficiency, taking less than 5 minutes.
[0143] The core hyperparameters of random forest are shown in Table 1:
[0144] Table 1
[0145]
[0146] In an exemplary embodiment, feature selection is performed on the three sets of data after Z-score normalization to obtain a sample feature matrix. Specifically, this includes: aligning the Z-score normalized data across omics using UniProt ID / KEGG ID; performing LASSO regression after cross-omics data alignment to select multiple features; and generating an n×7 dimensional matrix based on the selected features, with each yeast strain corresponding to one row and each biomarker corresponding to one column. This n×7 dimensional matrix is then used as the sample feature matrix. n This represents the number of yeast strains in the sample training set.
[0147] Specifically, the feature selection process is as follows:
[0148] Cross-omics data alignment: Gene-protein-metabolite modules were constructed using the Weighted Gene Co-expression Network Analysis (WGCNA) package in R. These modules, identified through co-expression network analysis, consist of a set of genes, proteins, and metabolites highly correlated with ethanol concentration, reflecting potential regulatory networks or metabolic pathways. Network type: distinguishing between positive and negative regulation (signed hybrid); module definition: minimum number of genes = 30, cut height = 0.25; modules significantly correlated with ethanol concentration were selected (|r|>0.7, p<0.01).
[0149] LASSO Regression: The LASSO regression function in the glmnet package of R language is used to further filter out the features that contribute the most to the prediction model, thereby reducing the risk of overfitting.
[0150] This invention integrates the following 7 sets of multiple sets of biological marker features:
[0151] Genes: GRE1 gene, TPS1 gene;
[0152] Proteins: Heat shock protein 104 (HSP104), alcohol dehydrogenase 2 (ADH2), catalase 1 (CTA1);
[0153] Metabolites: Sphinganine, Trehalose-6-phosphate.
[0154] Tags: Classified by HT / LT as a binary tag.
[0155] In an exemplary embodiment, multiple sets of biomarker feature combinations are input into a strain screening model for industrial strain screening to obtain a high ethanol tolerance probability score for each of multiple unlabeled wild yeast strains. Specifically, this includes: arranging the Z-score values of the seven biomarkers in the multiple sets of biomarker feature combinations into an input vector; in the strain screening model, each tree is traversed independently, starting from the root node and splitting according to feature thresholds, assigning unlabeled wild yeast strains to specific leaf nodes, with each leaf node corresponding to a classification label; recording the proportion of high ethanol-tolerant yeast strains in specific leaf nodes during training as local probabilities; performing a weighted average on the local probabilities of all trees to obtain a weighted average result, which is then determined as the high ethanol tolerance probability score for unlabeled wild yeast strains; the weight of the weighted average is the accuracy of the tree.
[0156] Specifically, the importance of features to the model was assessed by calculating the SHAP values of a standardized multi-omics feature matrix (7 key biomarkers: 2 genes, 3 proteins, and 2 metabolites) using a pre-trained optimal random forest model (RF, AUC=0.95) via the R language's `shapr` package, and a SHAP importance plot was then generated. Figure 3 A schematic diagram illustrating the global feature importance provided by this invention, such as... Figure 3 As shown, the bar chart displays the average |SHAP| value of the features. The results show that GRE1 gene, CTA1, and HSP104 are the top three contributing features.
[0157] Specifically, probability scores are used to guide the screening of industrial strains, enabling the selection of candidate strains suitable for high-ethanol fermentation.
[0158] S105: Unlabeled wild yeast strains with a high ethanol tolerance probability score greater than or equal to a preset threshold are classified as high ethanol-tolerant yeast strains.
[0159] Specifically, the preset threshold is set according to specific engineering practices. In an industrial application example, the trained random forest model was applied to multi-omics data (including gene sequencing, proteomics, and metabolomics data) of 200 unlabeled wild yeast strains. The `predict()` function was used to output the HT probability score (threshold ≥ 0.8) for each strain, and 19 high-probability HT candidate strains were finally selected. The model was then validated through small-scale fermentation (10% ethanol stress, cultured at 30°C for 48 hours). Figure 4 This is a schematic diagram comparing the ethanol yield of the HT strain and the random strain provided by the present invention, as shown in the figure. Figure 4 As shown, the average ethanol yield of these strains reached 11.8±0.8 g / L, which was significantly higher than that of the randomized control group (10.0±1.2 g / L) by 18% (independent samples T test, p<0.01). Furthermore, the expression levels of key tolerance genes (such as HSP104) were upregulated by 2.1-3.5 times, confirming that model-guided targeted screening can effectively improve the performance of industrial strains.
[0160] In an exemplary embodiment, during independent test set validation, multi-omics data (containing 7 biomarker features: GRE1 gene, HSP104, etc.) of 50 pre-reserved Saccharomyces cerevisiae strains with known HT / LT phenotypes were used as input. A trained random forest model was loaded for prediction, and the ROC curve of the predicted probability versus the true label was calculated using the roc() function of the pROC package. The results showed an AUC of 0.88 (95% CI: 0.82-0.93), a sensitivity of 0.85, and a specificity of 0.83. The curve was plotted using ggplot2. Figure 5 , Figure 5This is a schematic diagram of the ROC curve for predicting tolerance of the independent test set (50 strains) provided by the present invention, as shown in the figure. Figure 5 The ROC curve shown has the false positive rate on the horizontal axis and the true positive rate on the vertical axis, with the diagonal used as a random reference. The area under the curve is significantly higher than that of random guessing (DeLong test p=2.3×10^-6), confirming that the model has a stable predictive ability for the ethanol tolerance of unknown strains. Subsequent decision curve analysis (DCA) demonstrated its clinical application value in industrial strain screening.
[0161] In one exemplary embodiment, the present invention provides as follows Figure 6 The diagram shown illustrates the technical route for constructing a multi-omics prediction model for ethanol tolerance in Saccharomyces cerevisiae. Figure 6 As shown, the multi-omics prediction model for ethanol tolerance in *Saccharomyces cerevisiae* includes a data acquisition phase, a feature engineering phase, and a modeling and application phase. In the data acquisition phase, strain screening and culture, as well as multi-omics data collection, are performed to obtain sample label data and multi-omics quantitative data. In the feature engineering phase, multi-omics data standardization and feature selection are performed to obtain a standardized feature matrix and key biomarkers. In the modeling and application phase, the feature matrix and machine learning model are constructed, the optimal model is determined as RF, and the SHAP value representing feature importance is determined. Based on the optimal model and SHAP value, wild-type strains are screened, and ethanol yield is detected.
[0162] When applying the method for determining highly ethanol-tolerant yeast strains based on multi-omics biomarkers provided by this invention, it is not necessary to rely on... Figure 1 The steps shown are executed in sequence. The specific execution order of each step can be determined as needed, and this invention does not impose any restrictions on it.
[0163] In view of the above-mentioned current state of the technology, the advantages of the present invention are as follows:
[0164] Multi-level data integration: This invention integrates genomics, proteomics, and metabolomics data, providing more comprehensive biological information and improving the accuracy and reliability of predictions.
[0165] Application of machine learning: By applying machine learning algorithms, this invention can extract key features from multi-omics data, build efficient prediction models, and overcome the limitations of traditional methods in data processing and prediction capabilities.
[0166] Industrial application-oriented: This invention focuses on providing a predictive tool that can be directly applied to industrial strain screening, fermentation process optimization, and synthetic biology modification, outputting a tolerance score that can be directly used in production. This improves the practicality and operability of the technology.
[0167] Improve screening efficiency and reduce costs: By using machine learning models for prediction, the workload and time cost of experiments are reduced, thus improving the efficiency of strain screening and modification.
[0168] The above describes a method for identifying highly ethanol-tolerant yeast strains based on multi-omics biomarkers, provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding apparatus for identifying highly ethanol-tolerant yeast strains based on multi-omics biomarkers, such as... Figure 7 As shown.
[0169] Figure 7 A schematic diagram of a device for identifying highly ethanol-tolerant yeast strains based on multi-omics biomarkers provided by the present invention includes:
[0170] The acquisition module 701 is used to acquire gene expression levels, protein expression profiles, and metabolite abundance of multiple unlabeled wild yeast strains.
[0171] Alignment module 702 is used to perform Z-score normalization on gene expression levels, protein expression profiles and metabolite abundance respectively, and align the three sets of data after Z-score normalization according to features.
[0172] The first determining module 703 is used to extract the standardized values of GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine and trehalose-6-phosphate from the three sets of data after feature alignment, and to form multiple sets of biomarker feature combinations based on the extracted standardized values.
[0173] The screening module 704 is used to combine multiple sets of biomarker features and input them into the strain screening model to screen industrial strains, thereby obtaining a high ethanol tolerance probability score for each of the multiple unlabeled wild yeast strains.
[0174] The second determining module 705 is used to identify unlabeled wild yeast strains with a high ethanol tolerance probability score greater than or equal to a preset threshold as high ethanol tolerance yeast strains.
[0175] Specific limitations regarding the apparatus for determining highly ethanol-tolerant yeast strains based on multi-omics biomarkers can be found in the above-described limitations of the method for determining highly ethanol-tolerant yeast strains based on multi-omics biomarkers, and will not be repeated here. Each module in the aforementioned apparatus for determining highly ethanol-tolerant yeast strains based on multi-omics biomarkers can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0176] The present invention also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 3 A method for identifying highly ethanol-tolerant yeast strains based on multi-omics biomarkers is provided.
[0177] The present invention also provides Figure 8 The schematic diagram of the computer device shown is as follows: Figure 8 As shown, at the hardware level, this computer device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then executes it to achieve the above. Figure 1 A method for identifying highly ethanol-tolerant yeast strains based on multi-omics biomarkers is provided.
[0178] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0179] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this invention.
Claims
1. A method for identifying highly ethanol-tolerant yeast strains based on multi-omics biomarkers, characterized in that, include: Gene expression levels, protein expression profiles, and metabolite abundances of multiple unlabeled wild-type yeast strains were obtained; The gene expression level, the protein expression profile, and the metabolite abundance were Z-score normalized respectively, and the three sets of data after Z-score normalization were aligned according to features; The standardized values of GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine and trehalose-6-phosphate were extracted from the three sets of data after feature alignment. Based on the extracted standardized values, multiple sets of biomarker feature combinations were formed. The features of the multiple sets of biomarkers are combined and input into the strain screening model for industrial strain screening to obtain a high ethanol tolerance probability score for each of the multiple unlabeled wild yeast strains; the strain screening model includes random forest model, support vector machine model, gradient boosting decision tree model, naive Bayes model, neural network model and generalized linear model. Unlabeled wild yeast strains with a high ethanol tolerance probability score greater than or equal to a preset threshold are considered as high ethanol-tolerant yeast strains.
2. The method as described in claim 1, characterized in that, The process of combining the features of the multiple sets of biomarkers and inputting them into a strain screening model for industrial strain screening, to obtain a high ethanol tolerance probability score for each of the multiple unlabeled wild yeast strains, specifically includes: The Z-score values of the seven biomarkers in the multiple sets of biomarker feature combinations are arranged in order as the input vector; In the strain screening model, each tree is traversed independently. Each tree starts from the root node and splits according to the feature threshold. Unlabeled wild yeast strains are assigned to specific leaf nodes, and each leaf node corresponds to a classification label. The proportion of high ethanol-tolerant yeast strains in a specific leaf node during training is recorded as a local probability. The local probabilities of all trees are weighted and averaged to obtain a weighted average result. The weighted average result is determined as the high ethanol tolerance probability score of the unlabeled wild yeast strain; the weight of the weighted average is the precision of the tree.
3. The method as described in claim 1, characterized in that, The training process of the strain screening model specifically includes: Obtain the sample feature matrix and divide it into a training set and a validation set; The training set is used to train random forest model, support vector machine model, gradient boosting decision tree model, naive Bayes model, neural network model and generalized linear model, and the parameters of all models are adjusted until all models are optimal. For each of all models, the validation set is input into the model to obtain the model selection results. The model selection results are compared with the true categories of the test set to obtain the model's metrics. The metrics include prediction accuracy, sensitivity, and specificity. Compare and analyze the metrics of all models, and determine the optimal model based on the results of the comparative analysis; The optimal model is determined as the strain screening model.
4. The method as described in claim 3, characterized in that, The acquisition of the sample feature matrix specifically includes: Multiple high-ethanol-tolerant yeast strains and multiple low-ethanol-tolerant yeast strains were obtained as a sample training set; Gene expression levels, protein expression profiles, and metabolite abundances of each yeast strain in the training set were obtained at different ethanol concentrations. Z-score normalization was performed on the gene expression level, protein expression profile, and metabolite abundance of the sample to obtain three sets of data. Feature selection was performed on the three sets of data after Z-score standardization to obtain a sample feature matrix; the sample feature matrix includes GRE1 gene, TPS1 gene, HSP104 protein, ADH2 protein, CTA1 protein, sphingosine and trehalose-6-phosphate.
5. The method as described in claim 4, characterized in that, The step of performing feature selection on the three sets of data after Z-score standardization to obtain a sample feature matrix specifically includes: The data after Z-score normalization were aligned across omics data using UniProt ID / KEGG ID; After performing the cross-omics data alignment, LASSO regression was performed to filter out multiple features; Based on the selected features, an n×7 dimensional matrix is generated, with each yeast strain corresponding to one row of the feature matrix and each biomarker corresponding to one column. This n×7 dimensional matrix is then determined as the sample feature matrix. n This represents the number of yeast strains in the sample training set.
6. The method as described in claim 4, characterized in that, The screening process for multiple high-ethanol-tolerant yeast strains and multiple low-ethanol-tolerant yeast strains specifically includes: Candidate strains were selected from the wild-type Saccharomyces cerevisiae strain library; Prepare a culture medium containing ethanol; The candidate strains were cultured in an ethanol-containing medium; After the culture, the candidate strains are initially screened and then screened again. Based on the screening criteria for high ethanol-tolerant strains and low ethanol-tolerant strains, the strains obtained from the second screening are further screened to obtain multiple high ethanol-tolerant yeast strains and multiple low ethanol-tolerant yeast strains.