Big Data Association Mining and Data Processing Methods for Fungal Secondary Metabolism Pathways
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]为了弥补以上不足,本发明提供了真菌次生代谢通路组学大数据关联挖掘数据处理方法,旨在改善现有的真菌次生代谢通路关联分析大都采用统一时间点组学数据直接关联分析,容易造成跨组学关联关系失真的问题
[0047]1、本发明中,通过获取真菌在多个培养时间点的基因组层数据、转录组层数据、蛋白组层数据、代谢组层数据及培养环境层数据,并对基因组层数据、转录组层数据、蛋白组层数据和代谢组层数据执行跨组学动态时滞校正,进一步基于时滞校正后组学矩阵计算通路可信度、识别通路演化状态、计算候选因果强度、构建动态关联网络以及计算通路驱动指数,从而改善了现有的真菌次生代谢通路关联分析大都采用统一时间点组学数据直接关联分析,由于转录表达、蛋白表达与代谢产物积累之间存在跨组学时间滞后,从而造成跨组学关联关系失真的问题。
Smart Images

Figure CN122575497A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent processing and analysis technology of multi-omics data on fungal secondary metabolism, and in particular to a data processing method for big data correlation mining of fungal secondary metabolic pathway omics. Background Technology
[0002] Fungi produce a wealth of secondary metabolites, including polyketides, nonribosomal peptides, terpenes, and alkaloids, which have wide applications in pharmaceutical development, agricultural production, food processing, and biomanufacturing. With the advancements in high-throughput sequencing, proteomics, metabolomics, and bioinformatics, researchers can simultaneously acquire multi-level omics data, including fungal genomes, transcriptomes, proteomes, and metabolomes, and combine this with culture environment parameters to study fungal secondary metabolic pathways. Currently, research on fungal secondary metabolic pathways typically employs multi-omics joint analysis, identifying biosynthetic gene clusters, key enzyme expression characteristics, and metabolite variation patterns to establish relationships between different omics layers. This allows for the identification of potential secondary metabolic pathways and their regulatory mechanisms, providing data support for the discovery of novel natural products and the targeted mining of high-value metabolites.
[0003] Existing association analyses of fungal secondary metabolic pathways mostly use direct association analysis of omics data at a unified time point. However, due to the cross-omics time lag between transcriptional expression, protein expression, and metabolite accumulation, cross-omics associations are often distorted. Summary of the Invention
[0004] To address the above shortcomings, this invention provides a data processing method for big data association mining of fungal secondary metabolic pathways, aiming to improve the problem that existing fungal secondary metabolic pathway association analyses mostly use direct association analysis of omics data at a unified time point, which easily leads to distortion of cross-omics association relationships.
[0005] This invention provides the following technical solution: a data processing method for big data association mining of fungal secondary metabolite pathways, comprising the following steps:
[0006] S1. Obtain genomic, transcriptomic, proteomic, metabolomic, and culture environment data of fungi at multiple culture time points;
[0007] S2. Perform cross-omics dynamic time lag correction on genomic layer data, transcriptomic layer data, proteomic layer data and metabolomic layer data to generate time lag corrected omics matrix;
[0008] S3. Based on the temporal change characteristics of the data at each level in the time-lag corrected omics matrix, calculate the pathway confidence of each secondary metabolic pathway, and identify the pathway evolution status of each secondary metabolic pathway at multiple culture time points based on the pathway confidence.
[0009] S4. Within the time window defined by the pathway evolution state of each secondary metabolic pathway, calculate the candidate causal strength of the genomic layer data, transcriptomic layer data, proteomic layer data and metabolomic layer data in the time-lag corrected omics matrix, and generate the candidate causal strength corresponding to each association.
[0010] S5. Construct a dynamic association network based on the pathway evolution status of each secondary metabolic pathway and the candidate causal strength corresponding to each association, and calculate the pathway driving index of each secondary metabolic pathway based on the dynamic association network.
[0011] By employing the above technical solution, genomic, transcriptomic, proteomic, metabolomic, and culture environment data of fungi are acquired at multiple culture time points. Cross-omics dynamic time lag correction is performed on the genomic, transcriptomic, proteomic, and metabolomic data. Furthermore, based on the time lag-corrected omics matrix, pathway credibility is calculated, pathway evolution status is identified, candidate causal strength is calculated, dynamic association networks are constructed, and pathway driving indices are calculated. This improves upon the problem that most existing fungal secondary metabolic pathway association analyses use direct association analysis of omics data at a unified time point, which leads to distortion of cross-omics association relationships due to cross-omics time lags between transcriptional expression, protein expression, and metabolite accumulation.
[0012] Furthermore, in S2, the step of performing cross-omics dynamic time delay correction includes:
[0013] Based on cross-omics time series, the time-lag correlations between transcriptome layer data and proteome layer data, between proteome layer data and metabolome layer data, and between transcriptome layer data and metabolome layer data at different time scales were calculated.
[0014] Generate dynamic time delay functions based on the time delay correlations at different time scales;
[0015] Generate a time delay weight matrix based on the correlation strength corresponding to the dynamic time delay function.
[0016] Further, in S2, the step of generating the time-delay-corrected omics matrix includes:
[0017] A unified reference timeline was established based on multiple culture time points;
[0018] Based on the dynamic time delay function, time position mapping is performed on transcriptome layer data, proteome layer data and metabolome layer data to generate time-corrected data;
[0019] Weight correction is performed on the time-delay weight matrix to generate weighted correction data;
[0020] The genomic layer data and weighted correction data are recombined based on a unified reference time axis to generate a time-lag corrected omics matrix.
[0021] Furthermore, in S3, the step of calculating the pathway confidence of each secondary metabolic pathway includes:
[0022] Gene cluster expression sequences were extracted based on the temporal correspondence between genomic layer data and transcriptomic layer data in the time-lag corrected omics matrix.
[0023] Enzyme expression sequences were extracted based on the correspondence between proteomic layer data and gene cluster expression sequences in the time-lag corrected omics matrix.
[0024] Metabolite response sequences were extracted from metabolomics layer data in the time-lag corrected omics matrix, and the metabolite response sequences were time-aligned with gene cluster expression sequences and enzyme expression sequences.
[0025] Pathway reliability is calculated based on gene cluster expression sequences, enzyme expression sequences, and metabolite response sequences.
[0026] Furthermore, in S3, the step of identifying the pathway evolution state of each secondary metabolic pathway at multiple culture time points includes:
[0027] Based on pathway reliability and the key enzyme change sequences and metabolic flow change sequences in the time-lag corrected omics matrix, a pathway status observation sequence was constructed.
[0028] A pathway evolution state set is constructed based on the pathway state observation sequence. The pathway evolution state set includes activation state, enhancement state, stable state, decay state and reconstructed state.
[0029] Constraints are applied to the set of evolution states of the pathway based on the partial time point state annotations corresponding to the gold standard pathway.
[0030] Based on the pathway state observation sequence and the pathway evolution state set, the pathway evolution state sequence of each secondary metabolic pathway at multiple culture time points is generated.
[0031] Further, in S4, the step of calculating the candidate causal strength includes:
[0032] Based on the genomic, transcriptomic, proteomic, and metabolomic data in the time-lag corrected omics matrix, cross-omics time-corresponding variable pairs were constructed.
[0033] Calculate dynamic time-delay causal relationships based on time-corresponding variable pairs across omics and time windows defined by pathway evolution states;
[0034] Calculation of independence constraints for execution conditions is based on dynamic time-delay causal relationships combined with a pre-set set of proxy control variables;
[0035] Candidate causal strengths are generated by fusing the calculation results based on conditional independence constraints with pathway structure consistency information.
[0036] Further, in S4, the step of generating candidate causal strengths corresponding to each association includes:
[0037] A set of cross-omics associations was constructed based on genomic, transcriptomic, proteomic, and metabolomic data.
[0038] Establish a corresponding mapping relationship based on the candidate causal strength and the set of cross-omics associations;
[0039] The mapping relationship is divided into time windows based on the path evolution state;
[0040] Based on the time window segmentation results, candidate causal strengths are generated for each association.
[0041] Furthermore, in S5, the step of constructing a dynamic correlation network and calculating the pathway driving index of each secondary metabolic pathway based on the dynamic correlation network includes:
[0042] A set of cross-omics associations was constructed based on the genomic layer data, transcriptomic layer data, proteomic layer data, metabolomic layer data, and pathway evolution status of each secondary metabolic pathway in the time-lag corrected omics matrix.
[0043] Based on the candidate causal strength, the set of cross-omics associations is assigned an association strength value and the associations are filtered to generate a set of filtered cross-omics associations.
[0044] Based on the filtered set of cross-omics associations, a network structure is mapped according to a time window to generate a dynamic association network;
[0045] The pathway driving index is calculated based on the subnetwork structure corresponding to each secondary metabolic pathway in the dynamic correlation network.
[0046] The present invention has the following beneficial effects:
[0047] 1. In this invention, by acquiring genomic, transcriptomic, proteomic, metabolomic, and culture environment data of fungi at multiple culture time points, and performing cross-omics dynamic time lag correction on the genomic, transcriptomic, proteomic, and metabolomic data, the invention further calculates pathway credibility, identifies pathway evolution states, calculates candidate causal strength, constructs dynamic association networks, and calculates pathway driving indices based on the time lag-corrected omics matrix. This improves upon the existing fungal secondary metabolic pathway association analysis, which mostly uses direct association analysis of omics data at a unified time point. Due to the cross-omics time lag between transcriptional expression, protein expression, and metabolite accumulation, the cross-omics association relationships are distorted.
[0048] 2. In this invention, the reliability of the pathway is calculated by extracting gene cluster expression sequences, enzyme expression sequences and metabolite response sequences, and the pathway evolution state sequences at multiple culture time points are generated by combining key enzyme change sequences and metabolic flow change sequences. This establishes a temporal evolution process description mechanism for secondary metabolic pathways, thereby improving the problem that most existing fungal secondary metabolic pathway analyses use static expression levels to evaluate pathway activity. Due to the lack of continuous characterization of pathway activity changes, it is difficult to identify pathway state changes.
[0049] 3. In this invention, candidate causal strength is calculated by constructing cross-omics time-corresponding variable pairs, and a dynamic association network is constructed based on the candidate causal strength and pathway evolution status, as well as the pathway driving index is calculated. This forms a temporal network expression structure of cross-omics associations, thereby improving the problem that most existing fungal secondary metabolic pathway mining methods use correlation network analysis, which makes it difficult to distinguish the differences in the contribution of different associations to pathway changes, thus causing difficulties in identifying key driving pathways and key regulatory relationships. Attached Figure Description
[0050] Figure 1 This is a flowchart illustrating the data processing method for big data association mining of fungal secondary metabolite pathways proposed in this invention.
[0051] Figure 2 This is a schematic diagram of the process of cross-omics dynamic time-delay correction and generation of time-delay corrected omics matrix in this invention;
[0052] Figure 3 This is a schematic diagram of the process for calculating the reliability of pathways and identifying pathway evolution states in this invention.
[0053] Figure 4 This is a schematic diagram of the process for calculating candidate causality strength and generating candidate causality strength of association relationships in this invention.
[0054] Figure 5This is a schematic diagram illustrating the process of constructing the dynamic correlation network and calculating the path-driven index in this invention. Detailed Implementation
[0055] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0056] In embodiments of the present invention, a data processing method for big data association mining of fungal secondary metabolite pathways is provided, such as... Figure 1 As shown, the process includes the following steps: S1, acquiring genomic layer data, transcriptomic layer data, proteomic layer data, metabolomic layer data and culture environment layer data of fungi at multiple culture time points;
[0057] like Figure 2 As shown, S2 performs cross-omics dynamic time lag correction on genomic layer data, transcriptomic layer data, proteomic layer data and metabolomic layer data to generate time lag corrected omics matrix;
[0058] Furthermore, in S2, the steps for performing cross-omics dynamic time lag correction include:
[0059] Based on cross-omics time series, the time-lag correlations between transcriptome layer data and proteome layer data, between proteome layer data and metabolome layer data, and between transcriptome layer data and metabolome layer data at different time scales were calculated.
[0060] Generate dynamic time delay functions based on the time delay correlations at different time scales;
[0061] Generate a time delay weight matrix based on the correlation strength corresponding to the dynamic time delay function.
[0062] Specifically, after acquiring genomic, transcriptomic, proteomic, metabolomic, and culture environment data at multiple culture time points, cross-omics dynamic time lag correction was performed on the transcriptomic, proteomic, and metabolomic data to eliminate time shifts caused by biological process transmission between different omics layers. First, a cross-omics time series was constructed according to the culture time sequence, where the transcriptomic data was denoted as... Proteomic layer data are denoted as Metabolomics data are denoted as The time variable is denoted as The unit is hours; in this scheme, the incubation time points can be determined according to the experimental design, for example, by selecting the optimal collection time. , , , , and Time series data were generated from multiple time points; subsequently, time-lag relationships at different time scales were calculated for transcriptome-proteome, proteome-metabolome, and transcriptome-metabolome data, respectively; for any two omics layer time series... and The correlation strength under different time delay conditions is calculated using the time-delay cross-correlation function:
[0063] ;
[0064] in, Indicates the time delay, in hours; Indicates the number of time points; and These represent the average values of the corresponding time series; Indicates the time delay as The correlation strength over time is determined; the location with the strongest correlation strength is searched within short-term, medium-term, and long-term time scales, with the short-term scale being preferred. The preferred time scale is the medium time scale. Long-term scale is preferred After obtaining the optimal time delay at the corresponding time scale, a set of time delay correlations is formed. Since the transcriptional regulation rate, protein expression rate, and metabolite accumulation rate of fungal secondary metabolism vary at different culture stages, a dynamic time delay function is further constructed based on the time delay correlations obtained at multiple time scales. This is used to describe the time-delay change trajectory during the cultivation process; in this scheme, the dynamic time-delay function is generated using time series interpolation, so that each time point corresponds to a dynamic time-delay value; for example, obtaining the value in the early stage of cultivation. Acquired in the later stages of cultivation The dynamic time delay function reflects the gradual change of the time delay over time; subsequently, a time delay weight matrix is constructed based on the correlation strength corresponding to the dynamic time delay function. The matrix elements are defined as follows: ;in, Indicates the first The omics layer and the first The time-delay correlation strength corresponding to each omics layer; Indicates the number of omics layers involved in the calculation; The normalized time-delay weights are represented; the time-delay weight matrix is used to characterize the relative contribution of time-delay correlations between different omics layers; after the above processing, a dynamic time-delay function and a time-delay weight matrix are obtained, where the dynamic time-delay function is used for the subsequent time position mapping process, and the time-delay weight matrix is used for the subsequent weight correction process of time-corrected data, thereby generating a time-delay corrected omics matrix, which provides a unified time reference basis for subsequent pathway credibility calculation, pathway evolution state identification, and candidate causal strength calculation.
[0065] Furthermore, in S2, the steps for generating the time-delay-corrected omics matrix include:
[0066] A unified reference timeline was established based on multiple culture time points;
[0067] Based on the dynamic time delay function, time position mapping is performed on transcriptome layer data, proteome layer data and metabolome layer data to generate time-corrected data;
[0068] Weight correction is performed on the time-delay weight matrix to generate weighted correction data;
[0069] The genomic layer data and weighted correction data are recombined based on a unified reference time axis to generate a time-lag corrected omics matrix.
[0070] Specifically, in obtaining the dynamic time delay function and time delay weight matrix Then, a time-lag corrected omics matrix is generated; firstly, a unified reference time axis is established based on all culture time points. The unified reference timeline is directly constructed from the experimental sampling time series, and this method is preferred in this scheme. 12h, 24h, 36h, 48h and As a unified reference time node; then, a dynamic time delay function is used to perform time position mapping on the transcriptomic, proteomic, and metabolomic layer data; using the original time series of any one of the genome layers. For example, among which This represents the omics measurement values at the corresponding time points. The sampling time is represented in hours, and the time delay output by the dynamic time delay function is denoted as... If the unit is hours, then the time coordinate after time mapping is defined as follows: ;in, This represents the corrected time coordinates; since the corrected time coordinates usually do not fall exactly on the unified reference time axis nodes, time interpolation is used to obtain the omics data values on the unified reference time axis; for the reference time node... The corresponding time correction data is denoted as After the time-corrected data is generated, a time-delay weight matrix is introduced for weight correction; let the first... Each omics layer at the reference time point Time correction data is The corresponding weight is The weighted correction data is then defined as: ;in, This represents the weighted omics data; the time-delay weight matrix is derived from the normalized calculation results of the time-delay correlation between different omics layers, therefore the weight values can reflect the contribution of the corresponding omics layer's time mapping results in the cross-omics time alignment process; in this scheme, the elements in the time-delay weight matrix are preferably distributed as follows: The interval is defined, and the sum of all weights is 1. After weight correction, the genomic layer data and the corresponding weighted corrected data of each omics layer are recombined according to a unified reference time axis. The genomic layer data usually remains stable during culture, so it is directly mapped to the unified reference time axis. The transcriptome layer data, proteome layer data, and metabolome layer data are recombined using weighted corrected data. Finally, a time-lag corrected omics matrix is formed. Where OM represents the time-delay corrected omics matrix; Represents genomic layer data; This represents the transcriptome layer data after time correction and weight adjustment. This represents the proteomic layer data after time correction and weight adjustment. This represents the metabolomics layer data after time correction and weight adjustment. The time-lag corrected omics matrix realizes the corresponding expression of data from different omics layers under a unified reference time axis. Its output is used as the data input for subsequent pathway credibility calculation. By extracting gene cluster expression sequences, enzyme expression sequences and metabolite response sequences, we can further carry out secondary metabolic pathway credibility analysis and pathway evolution status identification.
[0071] like Figure 3 As shown in Figure S3, based on the temporal change characteristics of the data at each level in the time-lag corrected omics matrix, the pathway confidence of each secondary metabolic pathway is calculated, and the pathway evolution status of each secondary metabolic pathway at multiple culture time points is identified based on the pathway confidence.
[0072] Furthermore, in S3, the steps for calculating the pathway confidence of each secondary metabolic pathway include:
[0073] Gene cluster expression sequences were extracted based on the temporal correspondence between genomic layer data and transcriptomic layer data in the time-lag corrected omics matrix.
[0074] Enzyme expression sequences were extracted based on the correspondence between proteomic layer data and gene cluster expression sequences in the time-lag corrected omics matrix.
[0075] Metabolite response sequences were extracted from metabolomics layer data in the time-lag corrected omics matrix, and the metabolite response sequences were time-aligned with gene cluster expression sequences and enzyme expression sequences.
[0076] Pathway reliability is calculated based on gene cluster expression sequences, enzyme expression sequences, and metabolite response sequences.
[0077] Specifically, after obtaining the time-lag corrected omics matrix, the pathway confidence of each secondary metabolic pathway is further calculated. First, a pathway gene set is established based on the annotation results of biosynthetic gene clusters in the genome-level data. Then, the transcriptome-level data corresponding to the target biosynthetic gene clusters under a unified reference time axis are extracted in chronological order to form the gene cluster expression sequences, denoted as... ,in Indicates the first Gene cluster expression values corresponding to each reference time point This represents the total number of reference time points; subsequently, based on the correspondences of core synthases, modifying enzymes, and transporters recorded in the genome layer data, corresponding protein abundance information is extracted from the proteome layer data to form enzyme expression sequences, denoted as... ,in Indicates the first Enzyme expression values corresponding to reference time points; further, based on the end products, marker metabolic peaks, and detectable intermediate metabolites corresponding to the target secondary metabolic pathway, corresponding metabolic response information is extracted from the metabolomics layer data to form a metabolite response sequence, denoted as... ,in Indicates the first Metabolite response values corresponding to each reference time point; since cross-omics dynamic time lag correction has been completed in the previous steps, gene cluster expression sequences, enzyme expression sequences, and metabolite response sequences all correspond to a unified reference time axis, allowing for direct time alignment analysis; to calculate pathway reliability, the average correlation coefficient between gene expression sequences within a gene cluster is first calculated, based on the gene cluster expression sequences Obtain gene cluster expression consistency Subsequently, the correlation matching degree between the gene cluster expression sequence and the enzyme expression sequence was calculated to obtain the enzyme expression matching degree. And introduce a time delay weight matrix Time-delay weighted correction is performed; based on the correspondence between metabolite response sequences and enzyme expression sequences, the metabolic response matching degree corresponding to the final product, marker metabolic peak, and detectable intermediate is calculated. The probability of path activity within each time window is calculated based on a sliding window to obtain the sustained activity level. The above four indicators are respectively denoted as , , , The entropy weighting method is used to calculate the information entropy corresponding to each indicator, and the weight coefficient of each indicator is automatically determined based on the information entropy. , , , Credibility of the fusion computing pathway: ;in, Indicates the reliability of the pathway. , , , Automatically generated by the entropy weight method and satisfying + + + =1; The calculated pathway confidence value ranges from 0 to 1. Together with the key enzyme change sequence and the metabolic flow change sequence, a pathway state observation sequence is constructed, which further generates pathway evolution state sequences corresponding to multiple culture time points.
[0078] Furthermore, in S3, the steps for identifying the pathway evolution status of each secondary metabolic pathway at multiple culture time points include:
[0079] Based on pathway reliability and the key enzyme change sequences and metabolic flow change sequences in the time-lag corrected omics matrix, a pathway status observation sequence was constructed.
[0080] A pathway evolution state set is constructed based on the pathway state observation sequence. The pathway evolution state set includes activation state, enhancement state, stable state, decay state and reconstruction state.
[0081] Constraints are applied to the set of evolution states of the pathway based on the partial time point state annotations corresponding to the gold standard pathway.
[0082] Based on the pathway state observation sequence and the pathway evolution state set, the pathway evolution state sequence of each secondary metabolic pathway at multiple culture time points is generated.
[0083] Specifically, after obtaining the pathway confidence levels for each secondary metabolic pathway, the dynamic evolution of these pathways during the culture process is further identified. First, key enzyme expression sequences corresponding to the target secondary metabolic pathway are extracted from the time-lag corrected omics matrix, and the changes between adjacent time points are calculated to form key enzyme change sequences, denoted as... ,in Indicates the first Key enzyme changes at each time point; simultaneously, metabolite response sequences corresponding to the target secondary metabolic pathways are extracted, and metabolite flow change sequences are calculated based on metabolite abundance changes between adjacent time points, denoted as... ,in Indicates the first The metabolic flow change values corresponding to each time point; then, the pathway confidence sequence, key enzyme change sequence, and metabolic flow change sequence are combined according to a unified reference time axis to form a pathway status observation sequence: ;in, Indicates the first The observation state vector corresponding to each time node; Indicates the first The reliability of the pathway corresponding to each time point; Indicates the first Key enzyme changes at each time point; Indicates the first Metabolic logistics changes at each time point; after completing the observation sequence construction, a set of pathway evolution states is established: ;in, Indicates the active state. Indicates an enhanced state. Indicates a stable state. Indicates the decay state. The reconstructed state is represented; to describe the temporal transition relationship between states, a semi-supervised Hidden Markov Model is used to establish the state transition structure, and the state transition probability is defined as: ;in, Representing state Transition to state The probability of each time point belonging to a different pathway is determined. In this scheme, partial time-point state labels corresponding to the gold standard pathway are introduced as state constraint information. The gold standard pathway library is preferably derived from experimentally validated secondary metabolic pathways such as the aflatoxin pathway, penicillin pathway, lovastatin pathway, or gibberellin pathway. For example, if an enhanced state is known at a certain time point in the gold standard pathway, the corresponding state label is written into the training samples, and the state probability of that time point is constrained, thus forming a state learning process guided by partially known states. Subsequently, the probability distribution of each time point belonging to different pathway evolution states is calculated based on the observed state vector and the state transition probability. ;in, Indicates the first The posterior probability of each time point belonging to a certain path evolution state is obtained; further, the Viterbi path search method is used to obtain the state path with the highest posterior probability, forming a path evolution state sequence: Where PS represents the pathway evolution state sequence corresponding to the target secondary metabolic pathway; Indicates the first The path evolution state corresponding to each time node; The value of belongs to the state set The final obtained pathway evolution state sequence is used to characterize the dynamic changes of the target secondary metabolic pathway at different culture stages, and serves as the data input for subsequent candidate causal strength calculations. By constructing cross-omics time-corresponding variable pairs within the time window corresponding to the same pathway evolution state, further cross-omics causal relationship analysis is carried out.
[0084] like Figure 4 As shown in S4, within the time window defined by the pathway evolution state of each secondary metabolic pathway, candidate causal strength is calculated for the genomic layer data, transcriptomic layer data, proteomic layer data and metabolomic layer data in the time-lag corrected omics matrix, and candidate causal strength is generated for each association.
[0085] Furthermore, in S4, the steps for calculating candidate causal strength include:
[0086] Based on the genomic, transcriptomic, proteomic, and metabolomic data in the time-lag corrected omics matrix, cross-omics time-corresponding variable pairs were constructed.
[0087] Calculate dynamic time-delay causal relationships based on time-corresponding variable pairs across omics and time windows defined by pathway evolution states;
[0088] Calculation of independence constraints for execution conditions is based on dynamic time-delay causal relationships combined with a pre-set set of proxy control variables;
[0089] Candidate causal strengths are generated by fusing the calculation results based on conditional independence constraints with pathway structure consistency information.
[0090] Specifically, after obtaining the pathway evolution state sequence, candidate causal strength calculations are further performed within the time window corresponding to the same pathway evolution state. First, cross-omics time-corresponding variable pairs are constructed based on the genomic, transcriptomic, proteomic, and metabolomic data in the time-lag corrected omics matrix. Genomic data is used to determine the association between biosynthetic gene clusters and their corresponding transcripts, enzymes, and metabolites. Based on this, time-corresponding variable pairs are formed between gene expression variables, protein expression variables, and metabolite response variables, denoted as... ,in Represents the front-end variable of the causal chain. This represents the variables at the back end of the causal chain. In this scheme, the preferred time-corresponding variable pairs include those between transcriptome-level data and proteome-level data, between proteome-level data and metabolome-level data, and between transcriptome-level data and metabolome-level data. Subsequently, the time axis is divided into windows based on the pathway evolution state sequence, and time nodes in the same pathway evolution state are assigned to the same analysis window. Within each time window, the dynamic time lag relationship between variables is determined using the dynamic time lag function obtained in the previous steps, and dynamic Granger analysis is used to calculate the dynamic time lag causal relationship. ;in, Representing variables For variables Dynamic time-delay causality score; This indicates the value of the backend variable at the current time point; to Indicates the preceding Historical values of backend variables corresponding to each time point; The dynamic time delay is represented as The corresponding value of the front-end variable at that time; Var represents the residual variance; The historical order is represented, and in this scheme, a value of 2 to 5 is preferred. After the dynamic time-delay causality calculation is completed, a set of proxy control variables is introduced to perform conditional independence constraint calculations. The set of proxy control variables is denoted as... The set of proxy control variables is derived from culture environment layer data and annotated global regulatory factor data, preferably including... Relevant expression data of the regulation module Expression data related to regulatory modules, Velvet complex, environmental stress response factors, and culture environment parameters were collected; subsequently, variables were calculated. With variables Partial correlation coefficient under the constraints of the set of proxy control variables: ;in, The conditional independence constraint score is represented by _Corr_, which represents the conditional correlation calculation result. When the association between variables mainly originates from common regulatory factors, the corresponding score decreases after conditional independence constraints. When the association between variables remains stable, the corresponding score remains at a high level after conditional independence constraints. Further, pathway structure consistency information is constructed based on the biosynthetic gene cluster structure information, enzyme catalytic sequence information, and metabolite synthesis sequence information in the genomic layer data, denoted as ___. Higher consistency scores are assigned when the variable relationships conform to the biosynthetic pathway sequence, and lower consistency scores are assigned when the variable relationships do not conform to the known pathway structure. Finally, dynamic time-delay causal relationships, conditional independence constraints, and pathway structure consistency information are integrated to generate candidate causal strengths. Where CCD represents the candidate causal strength; and Indicates the fusion weights; in this scheme, , and The preferred method is to use a normalized weight allocation method and satisfy the following conditions: After the above calculations, the candidate causal strength results for each cross-omics time-corresponding variable pair are obtained. The candidate causal strength results serve as the data basis for subsequent cross-omics association assignments, and further establish the mapping relationship between candidate causal strength and cross-omics associations to provide edge weight data for the construction of dynamic association networks.
[0091] Furthermore, in S4, the steps for generating candidate causal strengths corresponding to each association include:
[0092] A set of cross-omics associations was constructed based on genomic, transcriptomic, proteomic, and metabolomic data.
[0093] Establish a corresponding mapping relationship based on the candidate causal strength and the set of cross-omics associations;
[0094] The mapping relationship is divided into time windows based on the path evolution state;
[0095] Based on the time window segmentation results, candidate causal strengths are generated for each association.
[0096] Specifically, after obtaining the candidate causal strengths for each cross-omics time-corresponding variable pair, candidate causal strength results for each association are further generated. First, a set of cross-omics associations is constructed based on the annotation relationships of biosynthetic gene clusters in the genomic layer data, and the correspondences between the transcriptomic, proteomic, and metabolomic layer data. This set of cross-omics associations is denoted as... ,in Indicates the first Cross-omics associations, This represents the total number of associations; in this scheme, cross-omics associations preferably include associations between gene expression and protein expression, associations between protein expression and metabolite response, and associations between gene expression and metabolite response; subsequently, a mapping relationship is established between the candidate causal strengths obtained in the preceding steps and the set of cross-omics associations, forming an association causal mapping set: ;in, Represents a set of causal mappings indicating relationships; Indicates the first The candidate causal strengths corresponding to each cross-omics association are calculated. These candidate causal strengths are derived from the fusion calculation of dynamic time-delay causal relationships, conditional independence constraints, and pathway structure consistency information; therefore, each association corresponds to a unique candidate causal strength value. After constructing the mapping relationships, the time windows of the mapping relationships are divided according to the pathway evolution state sequence obtained in the preceding steps. Let the pathway evolution state sequence be denoted as: ;in, Indicates the first The path evolution state corresponding to each time node; This represents the total number of time nodes; time nodes that are continuously in the same evolutionary state along the same path are divided into the same time window, forming a set of time windows: ;in, Indicates the first One time window; This represents the total number of time windows; subsequently, each relationship in the causal mapping set is mapped to its corresponding time window according to its time node, forming a window relationship set: ;in, Indicates the first The set of associations corresponding to each time window; since the same association may correspond to different candidate causal strengths in different pathway evolution states, the candidate causal strength results of associations under time window constraints are further generated: Wherein, CCDR represents the candidate causal strength result of the association relationship; Indicates the first The relationship in the first Candidate causal strengths within a time window; after the above processing, each cross-omics association obtains candidate causal strength results under the corresponding pathway evolution state and time window constraints, thus forming a set of correspondences between associations, time windows, and candidate causal strengths; this result serves as the data input for subsequent dynamic association network construction. During the construction of the dynamic association network, candidate causal strengths are written into the network structure as association edge weights, and combined with the pathway evolution state to complete the generation of the dynamic association network and the calculation of the pathway driving index.
[0097] like Figure 5 As shown in Figure S5, a dynamic association network is constructed based on the pathway evolution status of each secondary metabolic pathway and the candidate causal strength corresponding to each association relationship, and the pathway driving index of each secondary metabolic pathway is calculated based on the dynamic association network.
[0098] Furthermore, in S5, the steps of constructing a dynamic correlation network and calculating the pathway driving index of each secondary metabolic pathway based on the dynamic correlation network include:
[0099] A set of cross-omics associations was constructed based on the genomic layer data, transcriptomic layer data, proteomic layer data, metabolomic layer data, and pathway evolution status of each secondary metabolic pathway in the time-lag corrected omics matrix.
[0100] Based on the candidate causal strength, the cross-omics association set is assigned an association strength value and the association is filtered to generate a filtered cross-omics association set.
[0101] Based on the filtered set of cross-omics associations, a network structure is mapped according to a time window to generate a dynamic association network;
[0102] The pathway driving index is calculated based on the subnetwork structure corresponding to each secondary metabolic pathway in the dynamic correlation network.
[0103] Specifically, after obtaining the causal strength results of candidate associations, a dynamic association network is further constructed and the pathway driving index is calculated. First, a cross-omics association set is established based on the genomic, transcriptomic, proteomic, and metabolomic data, as well as pathway evolution state sequences, in the time-lag corrected omics matrix. Each gene cluster, transcript, key enzyme, and metabolite serves as a network node, and the node set is denoted as: ;in, Indicates the first Each node; based on the candidate causal strength results of the association relationships obtained in the previous steps, a set of association edges is established: ;in, Indicates the first There are 10 associated edges; then the candidate causal strengths are written into the corresponding associated edges to form the associated strength assignment results: ;in, Indicates the weight of the associated edge; This represents the candidate causal strength of the corresponding association. To ensure the consistency of the network structure, the association edges are filtered. In this scheme, the candidate causal strength threshold is preferably the value corresponding to the upper quartile of the distribution of all candidate causal strengths. The association edges that meet the filtering conditions are retained to form the filtered association set. ;in, This represents the set of relationships after filtering. This represents the threshold for selecting candidate causal strength; subsequently, a set of time windows is constructed based on the time windows corresponding to the pathway evolution state sequence. ;in, Indicates the first A time window is defined; the filtered set of relationships is mapped to the network structure according to the corresponding time window, forming a dynamic relationship network. ;in, Indicates the first The dynamic correlation network corresponding to each time window; Indicates the first A set of nodes within a time window; Indicates the first The set of associated edges within each time window; since different time windows correspond to different pathway evolution states, the dynamic association network can characterize the cross-omics association structure of each secondary metabolic pathway at different evolutionary stages; further, pathway subnetworks are extracted based on the nodes and associated edges corresponding to each secondary metabolic pathway, denoted as: ;in, Indicates the first Subnetworks corresponding to each secondary metabolic pathway; Represents the set of sub-network nodes; Represent the set of associated edges in the subnetwork; then calculate the mean causal strength of all candidate associated edges in the subnetwork: ;in, This represents the average causal strength of the subnetwork; This represents the number of edges associated with the subnetwork; further, the number of path evolution state transitions is counted and denoted as... The number of state transitions is derived from the cumulative result of state changes at adjacent time points in the evolutionary state sequence of the corresponding pathway; the contribution rate of metabolites is statistically calculated and denoted as... , ;in, Let Q represent the metabolite response value of the p-th pathway at time t, where Q represents the total number of pathways and T represents the set of all time points. The metabolite contribution rate is derived from the proportion of the corresponding pathway's metabolite response sequence to the total metabolite response. The statistical environmental response intensity is denoted as... , ;in, This represents the sequence of parameters in the culture environment layer, and Corr represents the correlation coefficient calculation. This represents the metabolite response sequence of the p-th pathway; the environmental response intensity is derived from the correlation between the culture environment layer data and the pathway metabolite response sequence; finally, the pathway driving index is calculated: ;in, Indicates the first Pathway driving index corresponding to each secondary metabolic pathway; , , as well as The normalized weight coefficients satisfy: In this scheme, the entropy weight method is preferred to determine the weight coefficients. After the calculation is completed, the pathway driving index results corresponding to all secondary metabolic pathways are obtained. The pathway driving index results are used to sort and analyze each secondary metabolic pathway, identify biosynthetic gene clusters with high driving capacity, key regulatory nodes and high-potential secondary metabolic pathways, and serve as data input for subsequent prediction of secondary metabolic potential.
[0104] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A data processing method for big data association mining of fungal secondary metabolite pathways, characterized in that, Includes the following steps: S1. Obtain genomic, transcriptomic, proteomic, metabolomic, and culture environment data of fungi at multiple culture time points; S2. Perform cross-omics dynamic time lag correction on genomic layer data, transcriptomic layer data, proteomic layer data and metabolomic layer data to generate time lag corrected omics matrix; S3. Based on the temporal change characteristics of the data at each level in the time-lag corrected omics matrix, calculate the pathway confidence of each secondary metabolic pathway, and identify the pathway evolution status of each secondary metabolic pathway at multiple culture time points based on the pathway confidence. S4. Within the time window defined by the pathway evolution state of each secondary metabolic pathway, calculate the candidate causal strength of the genomic layer data, transcriptomic layer data, proteomic layer data and metabolomic layer data in the time-lag corrected omics matrix, and generate the candidate causal strength corresponding to each association. S5. Construct a dynamic association network based on the pathway evolution status of each secondary metabolic pathway and the candidate causal strength corresponding to each association, and calculate the pathway driving index of each secondary metabolic pathway based on the dynamic association network.
2. The data processing method for fungal secondary metabolomics big data association mining according to claim 1, characterized in that, In S2, the step of performing cross-omics dynamic time lag correction includes: Based on cross-omics time series, the time-lag correlations between transcriptome layer data and proteome layer data, between proteome layer data and metabolome layer data, and between transcriptome layer data and metabolome layer data at different time scales were calculated. Generate dynamic time delay functions based on the time delay correlations at different time scales; Generate a time delay weight matrix based on the correlation strength corresponding to the dynamic time delay function.
3. The data processing method for fungal secondary metabolomics big data association mining according to claim 1, characterized in that, In S2, the step of generating the time-delay-corrected omics matrix includes: A unified reference timeline was established based on multiple culture time points; Based on the dynamic time delay function, time position mapping is performed on transcriptome layer data, proteome layer data and metabolome layer data to generate time-corrected data; Weight correction is performed on the time-delay weight matrix to generate weighted correction data; The genomic layer data and weighted correction data are recombined based on a unified reference time axis to generate a time-lag corrected omics matrix.
4. The data processing method for fungal secondary metabolomics big data association mining according to claim 1, characterized in that, In S3, the step of calculating the pathway confidence of each secondary metabolic pathway includes: Gene cluster expression sequences were extracted based on the temporal correspondence between genomic layer data and transcriptomic layer data in the time-lag corrected omics matrix. Enzyme expression sequences were extracted based on the correspondence between proteomic layer data and gene cluster expression sequences in the time-lag corrected omics matrix. Metabolite response sequences were extracted from metabolomics layer data in the time-lag corrected omics matrix, and the metabolite response sequences were time-aligned with gene cluster expression sequences and enzyme expression sequences. Pathway reliability is calculated based on gene cluster expression sequences, enzyme expression sequences, and metabolite response sequences.
5. The data processing method for fungal secondary metabolomics big data association mining according to claim 1, characterized in that, In S3, the step of identifying the pathway evolution status of each secondary metabolic pathway at multiple culture time points includes: Based on pathway reliability and the key enzyme change sequences and metabolic flow change sequences in the time-lag corrected omics matrix, a pathway status observation sequence was constructed. A pathway evolution state set is constructed based on the pathway state observation sequence. The pathway evolution state set includes activation state, enhancement state, stable state, decay state and reconstructed state. Constraints are applied to the set of evolution states of the pathway based on the partial time point state annotations corresponding to the gold standard pathway. Based on the pathway state observation sequence and the pathway evolution state set, the pathway evolution state sequence of each secondary metabolic pathway at multiple culture time points is generated.
6. The data processing method for fungal secondary metabolomics big data association mining according to claim 1, characterized in that, In S4, the step of calculating the candidate causal strength includes: Based on the genomic, transcriptomic, proteomic, and metabolomic data in the time-lag corrected omics matrix, cross-omics time-corresponding variable pairs were constructed. Calculate dynamic time-delay causal relationships based on time-corresponding variable pairs across omics and time windows defined by pathway evolution states; Calculation of independence constraints for execution conditions is based on dynamic time-delay causal relationships combined with a pre-set set of proxy control variables; Candidate causal strengths are generated by fusing the calculation results based on conditional independence constraints with pathway structure consistency information.
7. The data processing method for fungal secondary metabolomics big data association mining according to claim 1, characterized in that, In S4, the step of generating candidate causal strengths corresponding to each association includes: A set of cross-omics associations was constructed based on genomic, transcriptomic, proteomic, and metabolomic data. Establish a corresponding mapping relationship based on the candidate causal strength and the set of cross-omics associations; The mapping relationship is divided into time windows based on the path evolution state; Based on the time window segmentation results, candidate causal strengths are generated for each association.
8. The data processing method for fungal secondary metabolomics big data association mining according to claim 1, characterized in that, In S5, the step of constructing a dynamic correlation network and calculating the pathway driving index of each secondary metabolic pathway based on the dynamic correlation network includes: A set of cross-omics associations was constructed based on the genomic layer data, transcriptomic layer data, proteomic layer data, metabolomic layer data, and pathway evolution status of each secondary metabolic pathway in the time-lag corrected omics matrix. Based on the candidate causal strength, the set of cross-omics associations is assigned an association strength value and the associations are filtered to generate a set of filtered cross-omics associations. Based on the filtered set of cross-omics associations, a network structure is mapped according to a time window to generate a dynamic association network; The pathway driving index is calculated based on the subnetwork structure corresponding to each secondary metabolic pathway in the dynamic correlation network.