Fatty acid metabolism regulation gene mining method based on transcriptome data driving

By integrating RNA-Seq data through the FindGene platform and using WGCNA analysis, key regulatory genes were screened out. Through gene editing, fatty acid metabolic pathways were optimized, solving the problem of limited fatty acid production capacity in existing technologies. This resulted in a significant increase in fatty acid yield and composition ratio, and has broad prospects for industrial application.

CN121237212APending Publication Date: 2025-12-30HANGZHOU LONGXING BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511110764.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing technologies lack systematic methods for mining fatty acid metabolism regulatory elements, traditional metabolic engineering strategies have limited improvement, public transcriptome resources are not fully utilized, there is a lack of efficient platform tools for integrating and utilizing data, and a lack of closed-loop validation systems, making it difficult to further break through fatty acid production capacity.

Method used

By developing the FindGene platform, integrating and standardizing large-scale RNA-Seq data, constructing gene co-expression networks using weighted gene co-expression network analysis (WGCNA), screening out key regulatory genes, and optimizing fatty acid metabolism pathways through gene editing methods (such as site-directed mutagenesis, overexpression, and knockout) to construct engineered strains.

Benefits of technology

It significantly increases fatty acid yield, improves fatty acid biosynthesis efficiency, and optimizes the proportion of fatty acid components, possessing broad potential for industrial applications, especially in the field of bio-based fuel and biomaterial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237212A_ABST
    Figure CN121237212A_ABST
Patent Text Reader

Abstract

The invention discloses a fatty acid metabolism regulation gene mining method based on transcriptome data driving, a FindGene platform developed by the invention can efficiently integrate and standardize large-scale RNA-Seq data, systematic mining function related genes are analyzed through a weighted gene co-expression network, and the gene mining efficiency is improved. The limitation of dependence on small-scale experiments and manual screening in the past is broken through, and the efficiency and accuracy of key regulation gene screening are greatly improved. According to the schizosaccharomyces cerevisiae engineering strain QH-7 constructed by carrying out site-directed mutagenesis, overexpression and combined transformation on a key gene obtained by screening, the yield of fatty acid reaches 2.12 g / L and is increased by about 118.6% compared with that of a wild strain, and the biosynthesis efficiency of the fatty acid is greatly improved. According to the modified strain, the total fatty acid content is increased, meanwhile, the proportion of unsaturated fatty acids (C16: 1, C18: 1 and C18: 2) is remarkably increased, and the requirements for high-quality lipid raw materials in the fields of downstream biofuels, nutritional supplements and the like are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology, specifically relating to a method for mining fatty acid metabolism regulatory genes based on transcriptome data. Background Technology

[0002] Long-chain fatty acids (LCFAs) and their derivatives are widely used in renewable fuels, biodegradable plastics, nutritional supplements (such as omega-3 fatty acids), and lipid-based drug delivery systems. Currently, the demand for efficient and sustainable bio-based production systems is constantly growing. Replacing traditional petrochemical sources with microbial cell factories has become an important direction for the development of the fatty acid industry.

[0003] Among the many microbial chassis, Yersinia lipophila ( Yarrowia lipolytica or Y. lipolytica Due to its excellent lipid accumulation ability, it has become an important model for research in synthetic biology and metabolic engineering. (Natural) Y. lipolytica Under nutritional stress, its lipid content can reach more than 40% of the cell dry weight, and it can utilize inexpensive substrates such as crude glycerol, lignocellulosic acid hydrolysate, and industrial byproducts. Through systemic metabolic engineering, the engineered lines reportedly achieve lipid accumulation of about 90% of the dry weight (Athanasios B, Thierry C, Jean MN). Yarrowia lipolytica A model and a tool to understand the mechanisms implicated in lipid accumulation [J]. Biochimie. 2009, 91(6):692-696., demonstrating great potential for industrial applications. Furthermore, Y. lipolytica With advantages such as simple genetic manipulation, strong tolerance, and complete genome annotation, it is an ideal host for the large-scale synthesis of fatty acids and their derivatives.

[0004] However, currently targeting Y. lipolytica Metabolic engineering research primarily focuses on the regulation of metabolic precursors (such as acetyl-CoA and NADPH supply) or lipid synthesis pathways (such as structural genes like ACC1, DGA1, and DGA2). While these strategies can initially increase yields, they suffer from diminishing returns, mainly because they neglect the fact that lipid metabolism is controlled by a complex transcriptional regulatory network. Existing research suggests that lipid synthesis is finely regulated by various transcription factors, including nutrient sensing, redox balance, and stress response, involving multiple levels such as transcription factors, non-coding RNA, and epigenetic modifications. However, current research on... Y. lipolyticaThe work on identifying systemic regulatory elements of fatty acid metabolism remains fragmented, relying heavily on model yeasts (Saccharomyces cerevisiae). Saccharomyces cerevisiae The speculative homology comparisons lack empirical verification from baseline species.

[0005] On the other hand, public databases have accumulated a large amount of data. Y. lipolytica RNA-Seq data theoretically holds information that reveals key regulatory factors. However, due to inconsistent technical standards, varying experimental conditions, and non-standard data processing procedures, direct integration and analysis between different studies is difficult. Therefore, there is currently a lack of a universal platform for systematically integrating transcriptome data and discovering functionally regulatory genes at the global network level.

[0006] Existing research has shown that systems biology methods based on weighted co-expression network analysis (WGCNA) have demonstrated significant effectiveness in model microbial systems such as *Aspergillus niger* and *Escherichia coli*. By integrating multi-omics transcriptional data, researchers have successfully identified novel metabolic regulatory hubs, leading to a breakthrough improvement in the biosynthetic efficiency of target products. This suggests that employing similar systems biology strategies in *Yarrowia lipolytica*, through in-depth analysis of the transcriptome's dynamic network, could potentially reveal previously unknown key regulatory elements in fatty acid metabolism pathways, thus providing theoretical support for precise metabolic engineering.

[0007] The existing technology still has the following main problems and shortcomings: (1) There is a lack of systematic methods for exploring fatty acid metabolism regulatory elements.

[0008] Currently, the screening of fatty acid synthesis-related genes is mostly based on speculative homology comparisons, which rely on human experience and lack methods for systematic association screening using global transcriptome data, resulting in the easy omission of potentially efficient regulatory nodes.

[0009] (2) Traditional metabolic engineering strategies have limited improvement.

[0010] Existing fatty acid synthesis engineering based on single-gene manipulation (overexpression or knockout) often encounters bottleneck problems, such as metabolic network feedback inhibition, byproduct accumulation, and increased energy burden, resulting in diminishing marginal returns to fatty acid production.

[0011] (3) Public transcriptome resources are not fully utilized.

[0012] Despite a large number Y. lipolytica While RNA-Seq data is publicly available, there are currently no efficient platforms or tools to integrate and utilize this data for functional gene mining due to the large heterogeneity of its sources and the lack of a unified standardized processing procedure.

[0013] (4) There is a lack of a closed-loop verification system that combines transcriptome data with genetic engineering.

[0014] In existing studies, even when new regulatory candidate genes are discovered, there is often a lack of systematic functional verification, combinatorial engineering optimization, and mechanism analysis, which affects further breakthroughs in fatty acid production capacity. Summary of the Invention

[0015] This invention aims to solve the existing Yarrowia lipolytica The following problems exist in the fatty acid production system of (Yarrowia lipolyticis): (1) There is a lack of systematic analytical methods for the fatty acid metabolism regulatory network; (2) Traditional metabolic engineering methods (such as simple structural gene overexpression or knockout) are easily subject to feedback regulation and metabolic bottlenecks, resulting in limited yield improvement; (3) The utilization rate of public transcriptome resources is low, and there is a lack of efficient and standardized bioinformatics mining tools, making it difficult to quickly screen functional regulatory genes; (4) At present, the field of metabolic engineering has not yet formed a rational design framework and implementation path based on the combination optimization of multidimensional regulatory mechanisms of key regulatory genes, making it difficult to achieve a systematic improvement in the biosynthetic efficiency of fatty acids in Yersinia lipolyticis.

[0016] A transcriptome-driven method for mining fatty acid metabolism regulatory genes includes the following steps: S1. Acquire transcriptome data and perform standardized preprocessing to establish a unified expression matrix; S2, construct a gene co-expression network from the standardized transcriptome data to obtain several modules, and select the modules related to fatty acid metabolism pathways; S3 uses known key structural genes in the three major metabolic pathways of fatty acid synthesis, fatty acid accumulation and fatty acid β-oxidation as anchors to extract genes that are significantly co-expressed with the key structural genes used as anchors in the selected modules. S4 selects common genes identified from different genes in the same metabolic pathway as candidate regulatory genes.

[0017] Preferably, in step S1, the transcriptome data is raw sequencing data in the format of .fastq, .fq, .fastq.gz or .fq.gz, which serves as the raw FASTQ file.

[0018] More preferably, in step S1, the original FASTQ file is subjected to sequencing quality assessment, and low-quality bases and adapter contamination are removed.

[0019] More preferably, sequencing quality assessment of the original FASTQ file includes: The quality of sequencing base by base is shown by box plots based on the Phred score, illustrating the distribution of base quality in each sequencing cycle. Sequence-wise average quality distribution to identify the proportion of low-quality reads; GC content distribution to detect library construction or contamination preferences; Sequence repeat levels are assessed to evaluate the presence of PCR over-amplification or high-abundance biological sequences. The contamination status of the joint and the enrichment of high-frequency sequences help to identify joint residues or exogenous contamination. And k-mer enrichment analysis, used to discover over-enrichment of specific short sequences at specific locations.

[0020] More preferably, the removal of low-quality bases is as follows: at the 3′ end, using Phred quality fraction Q20 as the threshold, all bases with a quality lower than 20 are pruned sequentially until the quality of all bases in the remaining sequence is ≥20. If the read length after pruning is less than 50bp, it is discarded. Save the trimmed read segment with the suffix _trimmed.fq.gz and store it in the trimmed_results directory.

[0021] More preferably, the _trimmed.fq.gz file is input into HISAT2 for genome alignment, generating .sam format output, and the results are saved in the hisat2_results directory; The samtools view is called sequentially to convert all .sam files to .bam format, and samtools sort is used to sort each .bam file, generating _sorted.bam files, which are stored in the bam_results and sorted_bam_results directories respectively; Based on the selected GTF annotation file, all _sorted.bam files are passed to FeatureCounts in a multi-threaded manner for read count statistics, and the original gene × sample count expression matrix featurecounts_results.txt is output. Based on the expression matrix of counts and gene length information parsed from the GTF file, the RPK value is calculated, and then TPM normalization is performed to obtain the TPM matrix; The standardized TPM matrix is ​​transformed element by element using log2(x+1) to obtain a uniform expression matrix.

[0022] Preferably, in step S2, the WGCNA algorithm is used to construct a gene co-expression network; In step S3, the key structural genes of the fatty acid synthesis pathway are ACC1, FAS1, and FAS2; the key structural genes of the accumulation pathway are DGA2, GPD1, and SCT1; and the key structural genes of the β-oxidation pathway are FAA1 and PEX10.

[0023] Preferably, in step S4, Wayne analysis is applied to obtain the final six high-confidence candidate regulatory genes, namely YALI0B12342g, YALI0A07733g, YALI0C03003g, YALI0C16797g, YALI0A20207g, and YALI0D01001g.

[0024] This invention further provides the application of fatty acid metabolism regulatory genes in regulating yeast fatty acid synthesis, wherein the gene editing method is at least one of the following: The YALI0B12342g gene was subjected to a G643R site-directed mutation. Overexpress at least one of the YALI0A07733g and YALI0A20207g genes; Knock out at least one of the genes YALI0C03003g, YALI0C16797g, and YALI0D01001g.

[0025] Beneficial effects of this invention: (1) Efficient screening of key regulatory genes: The FindGene platform developed in this invention can efficiently integrate and standardize the processing of large-scale RNA-Seq data. Through weighted gene co-expression network analysis, it systematically mines functionally related genes, breaking through the limitations of previous reliance on small-scale experiments and manual screening, and greatly improving the efficiency and accuracy of screening key regulatory genes.

[0026] (2) Significantly increases fatty acid production: By implementing site-directed mutagenesis, overexpression, and combined modification of key genes selected through screening, the engineered fission yeast strain QH-7 was constructed, which achieved a fatty acid yield of 2.12 g / L, an increase of approximately 118.6% compared to the wild-type strain, thus realizing a significant improvement in fatty acid biosynthesis efficiency.

[0027] (3) Optimize the proportion of fatty acid components: The modified strain of this invention increases the total fatty acid content while significantly increasing the proportion of unsaturated fatty acids (C16:1, C18:1, C18:2), which helps meet the demand for high-quality lipid raw materials in downstream fields such as biofuels and nutritional supplements.

[0028] (4) Good scalability and application prospects: The method of this invention is not only applicable to fission yeast, but can also be extended to fatty acid metabolism engineering of other lipid-accumulating microorganisms (such as Cutaneotrichosporon, Rhodotorula, etc.), and has broad industrial application potential, especially in the fields of bio-based fuels, biomaterials and high-value lipid products production. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of the FindGene platform's workflow.

[0030] Figure 2 It showcases the three major metabolic pathways of fatty acid synthesis, accumulation, and β-oxidation, along with their key structural genes.

[0031] Figure 3 This demonstrates the process of screening candidate regulatory genes based on co-expression networks in different metabolic pathways. Figure 3 In the diagram, a, b, and c represent the fatty acid synthesis pathway, fatty acid accumulation pathway, and fatty acid β-oxidation pathway, respectively.

[0032] Figure 4 This demonstrates the FindGene platform's validation of the screening results.

[0033] Figure 5 Construction and fatty acid yield analysis of the G643R mutant strain of the YALI0B12342g gene. Figure 5 The 'a' in the diagram represents the construction of the G643R mutant homologous arm plasmid. Figure 5 In the figure, b represents a comparison of fatty acid production between the QH-1 mutant and the Po1f wild-type strain. Figure 5 In the figure, 'c' represents the variation in the content of different fatty acid components (C16:1, C18:1, C18:2, etc.).

[0034] Figure 6 Analysis of fatty acid production in candidate gene overexpression strains. Figure 6 In the diagram, 'a' represents the structure of the candidate gene overexpression plasmid. Figure 6 In the figure, b represents the effect of overexpression of different candidate genes on the total amount of fatty acids. Figure 6 In the figure, 'c' represents the effect of overexpression on the proportion of fatty acid components.

[0035] Figure 7 Analysis of fatty acid production in candidate gene knockout strains. Figure 7 In the diagram, 'a' represents a gene knockout strategy and a PCR screening method. Figure 7 In the figure, b represents the change in total fatty acid content after knocking out different candidate genes. Figure 7 In the figure, 'c' represents the effect of knockout on the proportion of fatty acid components.

[0036] Figure 8This study analyzes the growth curves and fatty acid yields of multi-site combined modified plants. Figure 8 In the figure, 'a' represents a comparison of the growth curves of engineered plants QH-1, QH-5, QH-6, and QH-7. Figure 8 In the figure, b represents a comparison of fatty acid yields among different strains. Figure 8 In the figure, 'c' represents the analysis of changes in fatty acid composition of the engineered strain. Detailed Implementation

[0037] Materials and reagents: Yeast strain: Yarrowia lipolytica Po1f (MatA, leu2-270, ura3-302, xpr2-322, axp1-2).

[0038] Plasmids: pCRISPRyl (for gene knockout and site-directed mutagenesis), pINA1269 (for target gene overexpression), pINA1312 (for constructing homologous arm plasmids).

[0039] Culture medium: YPD liquid medium (10 g / L yeast extract, 20 g / L peptone, 20 g / L glucose). Selective media: SD-Leu, SD-Leu / Ura.

[0040] Example 1

[0041] (1) Develop a standardized transcriptome data analysis platform, such as FindGene. Figure 1 The diagram shows the workflow of the FindGene platform, illustrating its core modules, including the RNA-Seq data normalization module (Trim-Galore, FastQC, Hisat2, FeatureCounts) and the weighted gene co-expression network analysis module (WGCNA). It also demonstrates the process of predicting gene interactions and screening potential regulatory genes based on expression data.

[0042] This application constructs a tool named FindGene, developed based on Python (versions 3.12.9 / 3.7.7 / 3.10.16) and R (4.4.0).

[0043] FindGene integrates modules for FastQC (data quality control), Trim-Galore (read pruning), Hisat2 (genome alignment), FeatureCounts (gene quantification), WGCNA (weighted gene co-expression network analysis), TPM normalization, and log(TPM+1) transformation.

[0044] Through batch processing, standardized preprocessing (TPM normalization and log2(x+1) transformation) of 447 Y. lipolytica RNA-Seq data was achieved, and a unified expression matrix was established.

[0045] The data processing procedure is as follows: The first step was to select a directory of raw sequencing data in Tkinter GUI (a standard GUI toolkit in Python) containing .fastq / .fq / .fastq.gz / .fq.gz formats (where " / " indicates an "OR" relationship between the formats). A total of 447 publicly available [databases] were collected. Y. lipolytica RNA-Seq data. The system automatically identifies and loads a list of all FASTQ files that meet the suffix requirements.

[0046] The second step involves using FastQC to assess the sequencing quality of the original FASTQ file, including: Per-base sequence quality is represented by box plots showing the distribution of base quality across each sequencing cycle based on Phred scores; per-sequence quality scores identify the proportion of low-quality reads; per-sequence GC content is used to detect library construction or contamination bias; sequence duplication levels assess the presence of PCR overamplification or high-abundance biological sequences; adapter content and overrepresented sequences help identify adapter residues or exogenous contamination; and k-mer content is used to identify over-enrichment of specific short sequences at specific locations.

[0047] FastQC generates Pass / Warning / Fail graded reports for each module, helping researchers quickly identify and optimize downstream trimming, filtering, or library construction processes.

[0048] The third step involves calling Trim Galore to remove low-quality bases and adapter contaminants from all original FASTQ files. Low-quality base removal is performed at the 3′ end, using a Phred quality score (Q20) as the threshold, sequentially trimming all bases with a quality score below 20 until all bases in the remaining sequence have a quality score ≥ 20. Reads shorter than 50 bp after trimming are discarded. The default Illumina aptamer sequence (AGATCGGAAGAGC) is used for adapter removal. The trimmed reads are saved with the _trimmed.fq.gz extension in the trimmed_results directory.

[0049] The fourth step involves FindGene automatically extracting the index base name based on the user-specified HISAT2 index file and inputting all _trimmed.fq.gz files into HISAT2 for genome alignment, generating .sam format output, and saving the results in the hisat2_results directory.

[0050] The fifth step involves sequentially calling samtools view to convert all .sam files to .bam format, and then using samtools sort to sort each .bam file, generating _sorted.bam files, which are stored in the bam_results and sorted_bam_results directories respectively.

[0051] Step 6: Based on the GTF annotation file selected by the user, the software passes all _sorted.bam files to FeatureCounts in a multi-threaded manner (-T 4) for read count statistics, and outputs the original gene × sample count expression matrix featurecounts_results.txt.

[0052] In the seventh step, FindGene calculates RPK (Reads Per Kilobase) in the background based on the counts matrix and gene length information parsed from the GTF file, and further performs TPM (Transcripts Per Million) standardization.

[0053] In the eighth step, the software applies log2(x+1) transformation to each element of the standardized TPM matrix to obtain a unified standardized expression matrix, which can be used for subsequent co-expression network analysis (WGCNA) and differential expression analysis.

[0054] (2) Screening of key regulatory genes related to fatty acid metabolism based on WGCNA method.

[0055] For the standardized transcriptome data, the WGCNA algorithm was used to construct a gene co-expression network, which yielded several modules. Modules highly associated with key fatty acid metabolism pathways (synthesis, accumulation, and β-oxidation) were identified.

[0056] Using known key structural genes in the three major fatty acid metabolic pathways—fatty acid synthesis, fatty acid accumulation, and fatty acid β-oxidation—as anchors, genes significantly co-expressed with these anchor genes were extracted from selected modules. Specifically, candidate genes strongly co-expressed with known key genes (ACC1, FAS1 / FAS2, DGA2, GPD1, SCT1, FAA1, PEX10) were screened. The key structural genes in the fatty acid synthesis pathway module were ACC1, FAS1, and FAS2; those in the accumulation pathway module were DGA2, GPD1, and SCT1; and those in the β-oxidation pathway were FAA1 and PEX10.

[0057] The stability and reliability of the screening results are ensured by using three-level data grouping validation (low heterogeneity group, medium heterogeneity group, and full data group).

[0058] The dataset consists of three sets of data: the low-heterogeneity set comprises 86 samples from the same laboratory, with highly consistent experimental conditions and sequencing procedures, resulting in minimal batch effects. This set is used to assess network stability under conditions of near-zero heterogeneity. The medium-heterogeneity set consists of approximately 50% (197 samples) randomly selected from 447 samples, with varying sample sources and experimental conditions. This set is used to simulate moderate inter-laboratory and inter-batch variability. The high-heterogeneity set includes all 447 publicly available RNA-Seq samples, covering the maximum heterogeneity across multiple laboratories, culture conditions, sequencing platforms, and other sources. This set is used for the final global co-expression network analysis. After constructing the WGCNA co-expression network, we first selected known anchor genes (ACC1, FAS1, FAS2; DGA2, GPD1, SCT1; FAA1, PEX10) from the three major metabolic pathways of fatty acid biosynthesis, accumulation, and β-oxidation. We then extracted lists of genes significantly co-expressed with each anchor gene from the corresponding modules of subset 1 (n=86), subset 2 (n=197), and the entire dataset (n=447). Subsequently, we applied Venn analysis (Venn diagram) to find the intersection of genes from each subset that were independently screened multiple times for the same metabolic pathway, eliminating the heterogeneity between subsets and thus retaining high-confidence candidate regulatory genes co-expressed in the corresponding anchor gene groups across all subsets. The screening was performed separately for each of the three pathways, and the intermediate processes and final results are as follows: Figures 2-4As shown in the figure, six key candidate regulatory genes were finally identified: YALI0B12342g (NCBI Gene ID: 2907123), YALI0A07733g (NCBI Gene ID: 2906089), YALI0C03003g (NCBI Gene ID: 2909057), YALI0C16797g (NCBI Gene ID: 2909298), YALI0A20207g (NCBI Gene ID: 2906000), and YALI0D01001g (NCBI Gene ID: 2910413).

[0059] Example 2

[0060] Functional validation and genetic modification.

[0061] (1) Site-directed gene mutation: The mutant strain QH-1 was constructed by performing site-directed mutagenesis at the G643R site on YALI0B12342g.

[0062] The site-directed mutagenesis method for YALI0B12342g G643R is as follows: A gRNA sequence targeting YALI0B12342g was designed and inserted into the pCRISPRyl vector to construct the pCRISPRyl_B12342g plasmid.

[0063] Using the Po1f genome as a template, a 1.2 kb upstream homologous arm and a 1.2 kb downstream homologous arm were amplified by PCR, and the G643R mutation was introduced in the middle to construct the homologous recombination plasmid pHR_B-G643R.

[0064] Centrifuge 500 µL of overnight culture medium to remove the supernatant, and resuspend the cell pellet in 0.1 M LiAc solution. Then, add 50% PEG 4000, 2 M LiAc, 10 mM DTT, 100 µg single-stranded vector DNA (ssDNA), and pCRISPRyl_B12342g and pHR_B-G643R plasmids (1 µg each) at a 1:1 molar ratio to the suspension and mix well. Incubate the mixture at 37°C for 30 min, and then heat shock at 42°C for 30 min to promote plasmid entry. After heat shock, add 1 mL of sterile ddH2O to dilute the mixture, centrifuge to remove the supernatant, resuspend the cells, and plate them on SD plates containing 5-FOA. Incubate at 30°C for 2-3 days. Further colony PCR and sequencing verification were performed on surviving clones to obtain the mutant strain QH-1 carrying YALI0B12342g G643R.

[0065] Fatty acid yield determination method: Yeast cells were disrupted by acid hydrolysis to extract lipids. After extraction, fatty acid methyl esters (FAMEs) were esterified using the methanol-hydrochlorination method and quantitatively analyzed by gas chromatography-flame ionization detector (GC-FID). Sample analysis was performed using a DB-WAX capillary column with nitrogen as the carrier gas in split injection mode. C19:0 fatty acid methyl ester was used as an internal standard for standardized quantification to ensure accurate and reliable data.

[0066] The construction process of the mutant strain and the results of fatty acid yield analysis are as follows: Figure 5 As shown in the figure. The results showed that the total fatty acid production of strain QH-1 was increased by 75% compared with that of wild type, and it was significantly enriched with unsaturated fatty acids (C16:1, C18:1, C18:2).

[0067] (2) Gene overexpression: The YALI0A07733g and YALI0A20207g genes were constructed into the pINA1269 vector (driven by the hp4d promoter and terminated by the XPR2t terminator).

[0068] First, the coding regions of YALI0A07733g and YALI0A20207g were amplified from the wild-type Po1f genome using primers, and then placed into the pINA1269 plasmid using the Gibson assembly method to form pINA1269_A07733g and pINA1269_A20207g, respectively. These plasmids carried the hp4d strong promoter, the XPR2 terminator, and the LEU2 selection marker. Subsequently, 1 μg of each recombinant plasmid was taken and transformed together with strain QH-1 (which already contained YALI0B12342g-G643R) using the lithium acetic acid / PEG heat shock method. After transformation, the plasmids were plated on SD-Leu plates and incubated at 30°C. The colonies were cultured at ℃ for 2-3 days to obtain white monoclonal antibodies. Colony PCR (using plasmid-specific primers) was then used to verify correct integration of the expression cassette, resulting in overexpression strains named QH-5 (YALI0A07733g overexpression) and QH-6 (YALI0A20207g overexpression), respectively. The plasmid information for the gene overexpression strains and the fatty acid yield analysis results are shown below. Figure 6 As shown in the figure. The results indicate that overexpression of YALI0A07733g can increase fatty acid production by 117%, and overexpression of YALI0A20207g can increase it by 307%.

[0069] (3) Gene knockout: YALI0C03003g, YALI0C16797g, and YALI0D01001g were knocked out respectively.

[0070] First, CRISPR-Cas9 vectors (pCRISPRyl_C03003g, pCRISPRyl_C16797g, pCRISPRyl_D01001g) and homologous recombination donor vectors (pHR_C03003g, pHR_C16797g, pHR_D01001g) targeting each gene were constructed. The latter inserted a 30 bp nonsense sequence into the coding region of each gene and flanked by approximately 1.2 kb homologous arms. 1 µg of each vector was co-transformed with 500 µL of overnight cultured Po1f bacterial suspension (OD600≈2–4). After washing with ddH2O and 0.1 M LiAc, the cells were resuspended in 200 µL of 50% PEG 4000 solution, and 10 µL of 2 M LiAc, 20 µL of 1 M DTT, 15 µL of denatured ssDNA and vector mixture were added. The mixture was incubated at 30°C for 30 min, followed by heat shock at 42°C for 30 min. After transformation, the cells were recovered and plated on selective plates containing LEU / antibiotics, and cultured at 30°C for 2–3 days. After picking single clones, colony PCR was performed using P3 / P2_C0, P3 / P2_C1, or P3 / P2_D primers. The PCR products were sequenced to confirm that the target gene had been replaced by a 30 bp nonsense fragment, resulting in knockout strains QH-2, QH-3, and QH-4.

[0071] Gene knockout strain construction information and fatty acid production analysis results are as follows: Figure 7 As shown, the total amount of fatty acids in the knockout strain increased slightly, but the changes in fatty acid composition were unfavorable (the proportion of saturated fatty acids increased, while the proportion of unsaturated fatty acids decreased).

[0072] Example 3

[0073] Multi-site combined engineering transformation.

[0074] In the mutant QH-1 background, the overexpression cassettes YALI0A07733g and YALI0A20207g were further integrated to construct the dual-site integrated strain QH-7.

[0075] First, an sgRNA targeting the POX3 site (sequence GTACTGAATCTGGGACTGGT) was designed and cloned into the pCRISPRyl vector (pCRISPRyl_POX3). Simultaneously, a recombinant plasmid pHR_POX3_A0 containing a 1.2 kb upstream homologous arm of POX3, an hp4d promoter-YALI0A07733g-XPR2t expression cassette, and a 1.2 kb downstream homologous arm was constructed. pCRISPRyl_POX3 and pHR_POX3_A0 were co-transformed into QH-1 cells. After transformation, the cells were plated on SD-Leu plates and cultured at 30 ℃ for 2-3 days to obtain single colonies. Colony PCR was performed using primers targeting the integration site to confirm the acquisition of the POX3 site-integrated strain QH-5. Subsequently, using QH-5 as the receptor, the same strategy was employed to design sgRNA (pCRISPRyl_POX2, sequence GCATAGACATAGGCCAGAAG) and recombinant plasmid (pHR_POX2_A2, containing hp4d-YALI0A20207g-XPR2t expression cassette and homologous arms flanking POX2) targeting the POX2 site. A second round of co-transformation and SD-Leu screening were performed, and dual-site integration was identified by PCR. Finally, an engineered strain QH-7 carrying both hp4d-YALI0A07733g-XPR2t and hp4d-YALI0A20207g-XPR2t was obtained.

[0076] Growth curves and fatty acid yield analysis results of multi-site combined modification engineering plants are as follows: Figure 8 As shown, the total fatty acid yield of strain QH-7 reached 2.12 g / L, an increase of 118.6% compared with the initial Po1f wild type, and the proportion of unsaturated fatty acids increased significantly.

Claims

1. A method for mining genes regulating fatty acid metabolism based on transcriptome data, characterized in that, The method comprises the following steps: S1, obtaining transcriptome data and performing standardization preprocessing to establish a unified expression matrix; S2, constructing a gene co-expression network for the standardized transcriptome data to obtain a plurality of modules, and selecting a module related to a fatty acid metabolic pathway; S3, taking a known key structural gene in a fatty acid synthesis, fatty acid accumulation and fatty acid beta-oxidation metabolic pathway as an anchor point, and extracting a gene significantly co-expressed with the key structural gene as the anchor point in the selected module; S4, taking a common gene selected from different genes in the same metabolic pathway as a candidate regulatory gene.

2. The method for mining fatty acid metabolism regulatory genes based on transcriptome data as described in claim 1, characterized in that, In step S1, the transcriptome data is raw sequencing data in.fastq,.fq,.fastq.gz or.fq.gz format, as a raw FASTQ file. 3.The method of claim 2, wherein the method comprises: determining a gene set of a sample of a first group and a gene set of a sample of a second group; and identifying a gene set of a sample of a third group based on the gene set of the sample of the first group and the gene set of the sample of the second group. In step S1, the raw FASTQ file is subjected to sequencing quality evaluation, and low-quality bases and adapter contamination are removed. 4.The method of claim 3, wherein the method is characterized by, The sequencing quality evaluation of the raw FASTQ file comprises: Base-by-base sequencing quality, box plot display of the distribution of base quality on each sequencing cycle based on Phred score; Average quality distribution per sequence to identify the proportion of low-quality reads; GC content distribution for detecting library construction or contamination bias; Sequence repeat level to assess the presence of PCR over-amplification or high-abundance biological sequences; Adapter contamination and high-frequency sequence enrichment to help confirm adapter residues or exogenous contamination; 5. The method for mining fatty acid metabolism regulatory genes based on transcriptome data as described in claim 3, characterized in that, And k-mer enrichment analysis for discovering excessive enrichment of specific short sequences at specific positions. Removing low-quality bases: trimming all bases with quality less than 20 in turn at the 3' end with Phred quality score Q20 as the threshold, until all bases in the remaining sequence have a quality of ≥20, and if the read length is less than 50 bp after trimming, discarding it; 6.The method of claim 5, wherein the method is characterized by, Save the trimmed reads as a _trimmed.fq.gz suffix in the trimmed_results directory. Input the _trimmed.fq.gz file into HISAT2 for genome alignment to generate a.sam format output, and save the result in the hisat2_results directory; Call samtools view in turn to convert all.sam files to.bam format, and use samtools sort to sort each.bam file to generate a _sorted.bam file, which is stored in the bam_results and sorted_bam_results directories, respectively; According to the selected GTF annotation file, pass all _sorted.bam files to FeatureCounts in a multi-threaded manner to perform read count statistics, and output a gene-sample raw counts expression matrix featurecounts_results.txt; Based on the counts expression matrix and the gene length information parsed from the GTF file, calculate the RPK value, and further perform TPM standardization to obtain a TPM matrix; The log2(x+1) conversion is applied to the normalized TPM matrix element by element to obtain a unified expression matrix.

7. The method for mining fatty acid metabolism regulatory genes based on transcriptome data as described in claim 1, characterized in that, In step S2, the WGCNA algorithm is used to construct a gene co-expression network. In step S3, the key structural genes of the fatty acid synthesis pathway are ACC1, FAS1 and FAS2; the key structural genes of the accumulation pathway are DGA2, GPD1 and SCT1; and the key structural genes of the beta-oxidation pathway are FAA1 and PEX10. 8.The method of claim 1, wherein the method is characterized by, In step S4, the final 6 high-confidence candidate regulatory genes are obtained by applying the Venn analysis, and are YALI0B12342g, YALI0A07733g, YALI0C03003g, YALI0C16797g, YALI0A20207g and YALI0D01001g.

9. Use of a gene regulating fatty acid metabolism for regulating the synthesis of fatty acids in yeast, characterized in that, The gene editing mode is at least one of the following: YALI0B12342g gene is subjected to G643R site-directed mutation; At least one of YALI0A07733g and YALI0A20207g genes is overexpressed; At least one of YALI0C03003g, YALI0C16797g and YALI0D01001g genes is knocked out.