A lung adenocarcinoma recurrence prediction method and system based on multi-omics data analysis

By constructing a metabolism-epigenetic association topology map and an epigenetic metabolism two-way distillation framework, multi-omics data are integrated to generate a lung adenocarcinoma recurrence risk prediction model. This solves the problem that existing technologies cannot effectively integrate metabolomics and transcriptomics data, and achieves efficient prediction and personalized intervention for lung adenocarcinoma recurrence.

CN120853876BActive Publication Date: 2026-05-29NANCHANG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANCHANG UNIV
Filing Date
2025-07-02
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively integrate metabolomics and transcriptomics data, making it difficult to accurately identify early high-risk signals of lung adenocarcinoma recurrence and predict spatial invasion pathways, thus limiting the development and implementation of personalized intervention strategies.

Method used

By constructing a metabolism-epigenetic association topology, we screened significant causal relationships between methylation changes and metabolic fluctuations, generated significant causal edge sets of methylation metabolism, constructed a two-way distillation framework for epigenetic metabolism, fused multi-omics data and generated metabolic pathway activity feature vectors, simulated metabolic pathway characteristics, constructed a lung adenocarcinoma recurrence risk prediction model, triggered high-risk early warning signals through time series analysis and solved the reaction-diffusion equation to simulate the perturbation propagation path.

Benefits of technology

This improved the accuracy and robustness of predicting lung adenocarcinoma recurrence, revealed hidden molecular interaction mechanisms, and enhanced the synergistic performance and interpretability of the prediction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853876B_ABST
    Figure CN120853876B_ABST
Patent Text Reader

Abstract

The application discloses a lung adenocarcinoma recurrence prediction method and system based on multi-omics data analysis, relates to the technical field of precision medicine, and comprises the following steps: collecting folate metabolism group data, a transcription group expression spectrum and methylation level data of a target gene promoter region of tumor tissue of a lung adenocarcinoma patient; constructing a metabolism-epigenetic correlation topology graph based on a spatial adjacent relationship, screening a significant causal relationship between methylation variation and metabolism fluctuation through causal analysis, and generating a methylation metabolism significant causal edge set; fusing simulation metabolism channel characteristics and the methylation metabolism significant causal edge set, constructing a lung adenocarcinoma recurrence risk prediction model, triggering a high-risk early warning signal through time sequence analysis of risk factor fluctuation; through the construction of the metabolism-epigenetic correlation topology graph and the screening of the significant causal edge set, the causal correlation modeling of multi-omics data in the spatial dimension is realized, the robustness of the recurrence prediction model is improved, the prediction performance is synergistically enhanced, and the accuracy of lung adenocarcinoma recurrence prediction is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of precision medicine technology, and in particular to a method and system for predicting recurrence of lung adenocarcinoma based on multi-omics data analysis. Background Technology

[0002] Predicting recurrence in lung adenocarcinoma is a key challenge in precision medicine, with the core challenge being the precise analysis of the molecular dynamics of the tumor microenvironment. In recent years, integrated analysis of multi-omics data (such as metabolomics and epigenetics) has provided new perspectives for studying recurrence mechanisms, particularly the association between folate metabolism networks and DNA methylation in tumor progression. Spatial gradient changes in folate metabolite concentrations can reflect the state of tumor metabolic reprogramming, while the regulation of promoter region methylation levels directly affects metabolic pathway activity; the synergistic effect of these two factors has potential predictive value for recurrence risk.

[0003] Current technologies typically employ correlation analysis or independent pathway enrichment strategies for multi-omics fusion, which struggles to capture the nonlinear responses and time-dependent characteristics of biological processes. In particular, in the scenario of predicting lung adenocarcinoma recurrence, the co-evolutionary mechanism between abnormal folate metabolism and DNA methylation dysregulation has not been adequately modeled, preventing existing models from effectively identifying early high-risk signals and predicting spatial invasion pathways. This directly limits the development and implementation of personalized intervention strategies. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a method for predicting recurrence of lung adenocarcinoma based on multi-omics data analysis to solve the problem of not being able to effectively integrate metabolomics and transcriptomics data for recurrence prediction.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, this invention provides a method for predicting recurrence of lung adenocarcinoma based on multi-omics data analysis. This method includes: collecting folate metabolome data, transcriptome expression profiles, and methylation level data of target gene promoter regions from tumor tissues of lung adenocarcinoma patients; constructing a metabolism-epigenetic association topology based on spatial adjacency relationships; screening for significant causal relationships between methylation changes and metabolic fluctuations through causal analysis to generate a significant causal edge set of methylation metabolism; constructing a two-way distillation framework for epigenetic metabolism based on the significant causal edge set of methylation-metagenesis; fusing folate metabolome data and transcriptome data in a teacher network to generate a metabolic pathway activity feature vector; optimizing methylation level data through a student network to obtain simulated metabolic pathway features; fusing the simulated metabolic pathway features with the significant causal edge set of methylation metabolism to construct a lung adenocarcinoma recurrence risk prediction model; triggering a high-risk warning signal through time-series analysis of risk factor fluctuations; and simulating the propagation path of folate metabolism network perturbations by solving reaction-diffusion equations to output a recurrence prediction report with a spatial invasion heatmap and a time evolution curve.

[0008] As a preferred embodiment of the lung adenocarcinoma recurrence prediction method based on multi-omics data analysis described in this invention, the specific steps for collecting folate metabolome data, transcriptome expression profiles, and methylation level data of target gene promoter regions from lung adenocarcinoma patient tumor tissues are as follows.

[0009] Lung adenocarcinoma tumor tissue was frozen and fixed, and serially sectioned into folic acid metabolome sections, transcriptome sections, and methylation detection region sections;

[0010] A micro-region marker coordinate system was preset on a glass slide, and dynamic adaptive laser desorption ionization imaging was performed on folic acid metabolome slices to obtain folic acid metabolome data.

[0011] In situ single-cell methylation barcode probe hybridization was performed on methylation detection region slices to detect methylation level data of target gene promoter regions. Transcriptome slices were spatially barcoded and whole transcriptome expression data were obtained through in situ RNA capture probe hybridization to obtain transcriptome expression profiles.

[0012] As a preferred embodiment of the lung adenocarcinoma recurrence prediction method based on multi-omics data analysis described in this invention, the specific steps for generating a significant causal edge set of methylation metabolism are as follows:

[0013] Based on the micro-region labeling coordinate system, the spatial coordinates of folate metabolomics data, transcriptomics data and methylation level data are finely aligned to generate spatial adjacency relationships; a metabolism-epigenetic association topology map is constructed by weighting the spatial adjacency relationships and the similarity of folate metabolomics data.

[0014] The initial causal edges in the metabolism-epigenesis association topology graph are iteratively optimized using a meta-reinforcement learning algorithm. With local causal direction consistency and statistical significance as constraints, a set of significant causal edges of methylation metabolism across spatial locations is selected and generated.

[0015] As a preferred embodiment of the lung adenocarcinoma recurrence prediction method based on multi-omics data analysis described in this invention, the steps of constructing an epigenetic-metabolic two-way distillation framework based on a significant causal edge set of methylation metabolism, folic acid metabolomics data and transcriptomics data being fused in a teacher network to generate metabolic pathway activity feature vectors are as follows.

[0016] Based on the significant causal edge set of methylation metabolism, a biological interaction logic network describing folate metabolomics data, transcriptome expression profiles and methylation level data is constructed. The teacher network forms a cross-modal fusion pathway, and the student network forms a causal-driven simulation pathway and a joint training mechanism, which constitutes an epigenetic-metabolic dual-pathway distillation framework.

[0017] Based on the teacher network, a metapath-aware attention mechanism is used to fuse folic acid metabolomics data with transcriptome expression profiles to generate metabolic pathway activity feature vectors.

[0018] As a preferred embodiment of the lung adenocarcinoma recurrence prediction method based on multi-omics data analysis described in this invention, the following steps are taken: methylation level data and significant causal edge sets of methylation metabolism are input into the student network; the first distillation path dynamically filters key regulatory sites in the methylation level data through causal masking; and the second distillation path aligns metabolic pathway activity feature vectors through an adversarial loss function to generate simulated metabolic pathway features.

[0019] As a preferred embodiment of the lung adenocarcinoma recurrence prediction method based on multi-omics data analysis described in this invention, the specific steps of triggering a high-risk early warning signal through time-series analysis of risk factor fluctuations are as follows:

[0020] Initial weights are generated based on multi-omics data inference, and dynamic correction and time-series sliding window smoothing are combined with student network simulation residual feedback to form time-varying weights of significant causal boundary set of methylation metabolism. Through dynamic simulation of methylation level data by student network, activity changes are calculated by difference and weighted by combining the time-varying weights of significant causal boundary set of methylation metabolism to generate metabolic pathway activity fluctuations.

[0021] By spatiotemporally fusing simulated metabolic pathway features with time-varying weights of significant causal edge sets of methylation metabolism, a lung adenocarcinoma recurrence risk prediction model is constructed based on a deep hierarchical survival model.

[0022] By using the survival analysis calculation logic of the lung adenocarcinoma recurrence risk prediction model, the combined effect of metabolic pathway activity fluctuations and methylation metabolism significant causal boundary set time-varying weights is quantified to generate time-series analysis risk factor fluctuations, detect residual fluctuation anomalies in time-series analysis risk factor fluctuations, and trigger high-risk warning signals.

[0023] As a preferred embodiment of the lung adenocarcinoma recurrence prediction method based on multi-omics data analysis described in this invention, the specific steps are as follows: The recurrence prediction report is generated by solving the reaction-diffusion equation to simulate the propagation path of folic acid metabolism network perturbations and outputting a spatial invasion heatmap and a temporal evolution curve.

[0024] Based on the simulated metabolic pathway characteristics and the time-varying weights of the significant causal edge set of methylation metabolism, the perturbation source, diffusion coefficient and reaction rate of the folate metabolism network are defined by the reaction-diffusion equation, and a dynamic model of the propagation of metabolic perturbation in the tumor space is constructed.

[0025] The reaction-diffusion equation is discretized and numerically solved to simulate the spatiotemporal propagation of metabolic perturbations in the spatial omics coordinate system. A standardized invasion intensity distribution is generated and mapped to the tumor spatial omics coordinate system to generate a spatial invasion heat map aligned with anatomical images and mark high invasion risk areas.

[0026] Extract the time evolution curve of high-risk areas and combine it with high-risk early warning signals to generate a recurrence prediction report.

[0027] Secondly, this invention provides a lung adenocarcinoma recurrence prediction system based on multi-omics data analysis, including a data acquisition module, a margin set module, a teacher network module, a student network optimization module, a construction module, and a reporting module. The data acquisition module is used to collect folate metabolome data, transcriptome expression profiles, and methylation level data of target gene promoter regions from tumor tissues of lung adenocarcinoma patients. The margin set module is used to construct a metabolism-epigenetic association topology based on spatial adjacency relationships, and to screen for significant causal relationships between methylation changes and metabolic fluctuations through causal analysis, generating significant causal margin sets for methylation metabolism. The teacher network module is used to construct a significant causal margin set for methylation-metagenesis relationships. The system employs a two-way distillation framework for epigenetic metabolism, using a fruit edge set to integrate folate metabolomics and transcriptomics data within a teacher network and generate metabolic pathway activity feature vectors. A student network optimization module optimizes methylation level data to obtain simulated metabolic pathway features. A construction module integrates these simulated metabolic pathway features with significant causal edge sets related to methylation metabolism to construct a lung adenocarcinoma recurrence risk prediction model, triggering high-risk warning signals through time-series analysis of risk factor fluctuations. A reporting module simulates the propagation path of folate metabolic network perturbations by solving reaction-diffusion equations, outputting a recurrence prediction report with spatial invasion heatmaps and temporal evolution curves.

[0028] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the lung adenocarcinoma recurrence prediction method based on multi-omics data analysis as described in the first aspect of the present invention.

[0029] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the lung adenocarcinoma recurrence prediction method based on multi-omics data analysis as described in the first aspect of the present invention.

[0030] The beneficial effects of this invention are as follows: By constructing a metabolic epigenetic association topology map and screening significant causal edge sets, causal association modeling of multi-omics data in the spatial dimension is achieved, preserving the biological significance of spatial locations and revealing hidden molecular interaction mechanisms, thereby improving the robustness and specificity of the recurrence prediction model; furthermore, an epigenetic dual-pathway distillation framework is constructed to generate simulated metabolic pathway features, which are then integrated by a teacher network to fuse multi-omics data and optimized by a student network to simulate features, achieving knowledge transfer dimensionality reduction, dynamic simulation to enhance interpretability, and synergistic enhancement of prediction performance; breaking through the limitations of single-omics data, uncovering hidden molecular interaction networks, and significantly improving the accuracy of lung adenocarcinoma recurrence prediction. Attached Figure Description

[0031] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is a flowchart of a lung adenocarcinoma recurrence prediction method based on multi-omics data analysis.

[0033] Figure 2 This is a schematic diagram of a lung adenocarcinoma recurrence prediction system based on multi-omics data analysis.

[0034] Figure 3 This is a flowchart of the dual-path distillation framework for epigenetic metabolism.

[0035] Figure 4 This is a flowchart of the process of solving the reaction-diffusion equation. Detailed Implementation

[0036] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0037] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0038] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0039] Reference Figures 1-4 This is one embodiment of the present invention, which provides a method for predicting recurrence of lung adenocarcinoma based on multi-omics data analysis, including the following steps:

[0040] S1. Collect folate metabolome data, transcriptome expression profiles, and methylation level data of target gene promoter regions from tumor tissues of lung adenocarcinoma patients.

[0041] Lung adenocarcinoma tumor tissue was cryofixed and serially sectioned into folic acid metabolome sections, transcriptome sections, and methylation detection region sections.

[0042] Specifically, lung adenocarcinoma tumor tissue samples were rapidly frozen and fixed in liquid nitrogen to maintain molecular integrity. The frozen tissue was then continuously cut into 10 μm thick slices using a cryostat (example temperature range: -20°C to -25°C). Each group of three slices was labeled as follows: folic acid metabolome slice (example temperature range: immediately transferred to -80°C for storage to avoid metabolite degradation), transcriptome slice (immersed in RNAlater stabilizing solution overnight at 4°C to stabilize RNA), and methylation detection region slice (fixed in 4% paraformaldehyde and embedded in paraffin blocks).

[0043] A micro-region marker coordinate system was preset on a glass slide, and dynamic adaptive laser desorption / ionization imaging was performed on folic acid metabolome slices to obtain folic acid metabolome data.

[0044] Specifically, after importing the spatial coordinate data of the preset micro-region marker coordinate system on the glass slide into laser desorption / ionization imaging, a dynamic adaptive algorithm is used to calibrate the laser focusing parameters in real time according to the reflected light intensity of the slice surface. The tissue slice is scanned along the preset coordinate path with a step accuracy of 5μm. The mass spectrometry signal intensity values ​​of folic acid metabolites (such as 5-MTHF, 10-CHO-THF, etc.) in each micro-region are captured simultaneously by linear ion trap mass spectrometry. After peak area normalization processing combined with the NIST metabolite database, a two-dimensional spatial concentration gradient matrix is ​​generated. The detection blind zone caused by tissue folds is eliminated by spatial interpolation algorithm, and metabolomics data that are strictly aligned with the micro-region coordinate system are output.

[0045] It should be noted that fluorescent microspheres (1 μm in diameter and 488 nm / 520 nm excitation / emission wavelengths in the example) are embedded as spatial anchors during the frozen sectioning of tumor tissue. Combined with anatomical landmarks (such as vascular bifurcation points) from preoperative CT / MRI images and whole-slice scan images of pathological sections (0.5 μm / pixel resolution in the example), the affine transformation matrix is ​​calculated using the Elastix rigid registration tool to map the multi-omics detection regions (folate metabolism, transcriptome, methylation) to a unified three-dimensional spatial grid, generating a micro-region coordinate system.

[0046] In situ single-cell methylation barcode probe hybridization was performed on methylation detection region slices to detect methylation level data of target gene promoter regions. Transcriptome slices were spatially barcoded and whole transcriptome expression data were obtained through in situ RNA capture probe hybridization to obtain transcriptome expression profiles.

[0047] Specifically, after fixing the methylation detection region slices onto a glass slide with a preset micro-region coordinate system, in situ hybridization is performed using methylation-specific probes carrying fluorescent labels. Single-cell fluorescence intensity is quantitatively detected using high-resolution confocal microscopy to generate a methylation value matrix. The coordinates of the fluorescent probes are scanned using confocal microscopy. A transcriptome capture probe chip containing spatial barcodes is then covered on the same glass slide surface, and the spatial barcode coordinates are recorded. After in situ reverse transcription amplification, a cDNA library is extracted and high-throughput sequencing is performed. After UMI correction and spatial barcode deconvolution, a standardized gene expression matrix is ​​output. The methylation matrix and transcriptome expression matrix are aligned using spatial encoding in the micro-region coordinate system to form spatially corresponding methylation level data and transcriptome expression profiles.

[0048] S2. Construct a metabolism-epiota association topology based on spatial adjacency relationships, screen significant causal relationships between methylation changes and metabolic fluctuations through causal analysis, and generate significant causal edge sets of methylation metabolism.

[0049] Based on the micro-region marker coordinate system, the spatial coordinates of folic acid metabolomics data, transcriptomics data and methylation level data are finely adjusted and aligned to generate spatial adjacency relationships.

[0050] Specifically, after inputting the spatial coordinates of metabolomics data, the spatial barcode coordinates of transcriptomics data, and the fluorescent probe coordinates of methylation data into multi-omics registration, based on the reference points in the preset micro-region marker coordinate system, the spatial offset of the three sets of data is calculated at the sub-pixel level using a feature point matching algorithm. Using the metabolomics coordinates as the reference, the coordinate offset of transcriptomics and methylation data is corrected by affine transformation matrix (including translation, rotation, and scaling parameters). The spatial overlap of key markers (such as high folic acid metabolite concentration regions and corresponding high gene expression regions) is cross-validated. Abnormal data points with offsets exceeding the tissue deformation region are removed. The metabolomics concentration matrix, methylation value matrix, and transcriptomics FPKM matrix with strictly aligned spatial coordinates are output to obtain the spatial adjacency relationship.

[0051] A metabolic epigenetic association topology was constructed by weighting spatial adjacency and the similarity of folic acid metabolome data.

[0052] Specifically, based on the dual weights of spatial adjacency and folic acid metabolome data similarity, a weighted average fusion strategy (with biological weight accounting for 60% and statistical weight accounting for 40% for example values) is used to construct the edge weight matrix of the metabolic epigenetic association topology graph. High-weight sub-networks are identified through a community detection algorithm (Louvain method), and the node embedding vector of the topology graph is generated by aggregating neighborhood features in combination with a graph attention network, thus outputting the metabolic epigenetic association topology graph.

[0053] It should be noted that the dual weighting of folate metabolome data similarity is generated by fusing biological association weights and statistical similarity weights. Based on the methylation value matrix and metabolome concentration matrix, the dynamic causal effect of methylation sites on metabolites is inferred through Markov chain Monte Carlo (MCMC) sampling, the posterior probability distribution of the causal effect size is calculated, and significant causal edges are screened. Subsequently, the results are verified by CRISPR interference experiments (e.g., knocking down the CpG island (cg07865057) in the KRAS promoter region resulted in a 32% decrease in 5-MTHF concentration (p=0.008)). The causal direction verified by the experiments is preserved, and the results of multi-omics inference and experimental verification are integrated to generate biological association weights. The statistical similarity weights are Gaussian kernel similarity constructed by the Euclidean distance between samples of the metabolite concentration matrix, and weighted average (fusing the two types of weights to obtain the dual weighting of folate metabolome data similarity).

[0054] The meta-reinforcement learning algorithm is used to iteratively optimize the initial causal edges in the metabolic epigenetic association topology graph. With local causal direction consistency and statistical significance as constraints, a set of significant causal edges of methylation metabolism across spatial locations is screened and generated.

[0055] Specifically, based on the initial causal edge set (node ​​= spatial micro-region, edge = association strength fused by dual weights) in the metabolic epigenetic association topology graph, a meta-reinforcement learning algorithm (MAML framework, inner loop policy network is PPO, outer loop meta-update step size 0.01) is adopted with local causal direction consistency (Jacobi matrix gradient direction matching rate ≥95%) and statistical significance as constraints. The policy network is trained by iteratively sampling spatial micro-region subgraphs to predict causal edge addition and deletion actions (action space: retain / delete / enhance weights). The policy network parameters are optimized using the reward function (causal direction matching score + statistical significance score - sparsity penalty term). After each iteration, the causal edge set is updated and the global causal network modularity is recalculated to screen out methylation metabolism significant causal edge sets that cross spatial locations and satisfy the constraints.

[0056] It should be noted that a Bayesian dynamic causal model was used to analyze the methylation and metabolite concentration sequences across time points, calculate the Granger causal effect strength of methylation sites on metabolites, and screen preliminary causal associations. The results were verified by CRISPR interference experiments. After knocking out specific methylation sites (such as the CpG island of the KRAS promoter), changes in metabolite concentration were detected, and the experimentally verified causal edges were preserved. The results of multi-omics causal inference and biological verification were integrated to generate an initial set of causal edges based on the metabolic epigenetic association topology graph.

[0057] It should also be noted that the training process of the Yess dynamic causal model aligns the methylation value matrix and metabolomics data (such as folate metabolite 5-MTHF) at multiple time points, constructs a lag dynamic causal equation, uses MCMC sampling (such as the NUTS algorithm) to estimate the posterior distribution of causal effect coefficients, screens significant causal edges based on Bayesian factors, and verifies the model through CRISPR knockout experiments (such as a ≥30% decrease in metabolite concentration after targeting methylation sites). The model is then optimized by combining cross-validation (time window segmentation) and sensitivity analysis (prior robustness test) to generate a biologically validated dynamic causal network, resulting in the trained Yess dynamic causal model.

[0058] S3. An epigenetic-metabolic dual-pathway distillation framework was constructed based on the significant causal edge set of methylation metabolism. Folic acid metabolomics data and transcriptomics data were fused in the teacher network to generate metabolic pathway activity feature vectors.

[0059] Based on the significant causal edge set of methylation metabolism, a biological interaction logic network describing folate metabolome data, transcriptome expression profiles and methylation level data is constructed. The teacher network forms a cross-modal fusion pathway, and the student network forms a causal-driven simulation pathway and a joint training mechanism, which constitutes an epigenetic-metabolic dual-pathway distillation framework.

[0060] Specifically, based on significant causal edge sets of methylation metabolism, a biological interaction logic network describing the relationship between metabolomics data, transcriptome expression profiles, and methylation level data is identified and constructed. This interaction logic network, as a core component of the teacher network, captures the complex relationships between different biological modalities. Specific metabolomics data, transcriptome expression profiles, and methylation level data are input into the teacher network. Deep nonlinear transformations and feature extraction are performed through the interaction logic network to learn the intrinsic connections between multi-source data and generate corresponding high-level feature representations. The feature representations generated by the teacher network guide the learning process of the student network, enabling it to simulate the causal driving relationship from epigenetic information (such as methylation) to metabolic change pathways, constructing a preliminary causal pathway mapping model. Under the joint training mechanism, a two-path distillation method is used to simultaneously optimize both the teacher and student networks. This allows the student network to not only effectively inherit the knowledge and patterns learned by the teacher network but also achieve higher generalization ability under resource constraints, constructing a complete epigenetic-metabolic two-path distillation framework.

[0061] It should be noted that, through a joint training mechanism, high-level features extracted from the teacher network are used as knowledge-guiding signals input into the student network. The student network simulates causal driving pathways based on these features and generates corresponding predictive outputs. By designing a unified objective function, the parameters of the teacher and student networks are optimized so that the output of the student network approximates the output of the teacher network as closely as possible. Joint gradient backpropagation is then performed in conjunction with the supervision signals from real data. The teacher network continuously refines its expression of multimodal data interaction logic, while the student network gradually improves its causal inference ability and generalization performance. This achieves synergistic improvement in knowledge transfer and causal modeling between the student and teacher networks, completing the training of the epigenetic-metabolic dual-path distillation framework.

[0062] Based on the teacher network, a metapath-aware attention mechanism is used to fuse folic acid metabolomics data with transcriptome expression profiles to generate metabolic pathway activity feature vectors.

[0063] Specifically, based on the teacher network, a heterogeneous graph structure is constructed through predefined metapaths (such as "metabolite → methylation site → gene") to encode folic acid metabolome data and transcriptome expression profiles into node feature vectors. Metapath-aware attention mechanism is used to calculate metabolite-gene cross-modal interaction weights (such as attention scores for 5-MTHF and MTHFR) along the metapath. Feature embeddings under multiple paths are aggregated (concatenation + dimensionality reduction by fully connected layers). After optimization by a multi-layer graph convolutional network, a cross-modal fused metabolic pathway activity feature vector is generated. The feature vector is then aligned with experimentally validated metabolic pathway activity labels through contrastive learning loss (InfoNCE) to output the metabolic pathway activity feature vector.

[0064] S4. Methylation level data were optimized using a student network to obtain simulated metabolic pathway characteristics.

[0065] Methylation level data and significant causal edge sets of methylation metabolism are input into the student network. The first distillation path dynamically filters key regulatory sites in the methylation level data through causal masks, while the second distillation path aligns metabolic pathway activity feature vectors through an adversarial loss function to generate simulated metabolic pathway features.

[0066] Specifically, methylation level data and a significant causal edge set of methylation metabolism are input into the student network. The first distillation path uses a dynamic causal mask (example values ​​are generated as a binary mask based on the causal edge weights, with a threshold of 0.7) to filter key regulatory sites (such as CpG island methylation data of the KRAS promoter). The filtered methylation data is then input into a lightweight graph convolutional network (2 layers, hidden layer dimensions 128→256, activation function GELU) to generate preliminary simulated metabolic features. The second distillation path inputs the preliminary simulated features into an adversarial discriminator (such as a multilayer perceptron, hidden layers 512→256→1, activation function LeakyReLU) and aligns the distribution with the metabolic pathway activity feature vector generated by the teacher network. The generator loss is optimized by minimizing the Wasserstein distance, and the distillation loss (mean squared error) and causal direction regularization term are jointly optimized to generate simulated metabolic pathway features that are consistent with the feature distribution of the teacher network and have reasonable biological logic.

[0067] S5. By integrating simulated metabolic pathway characteristics with significant causal boundary sets of methylation metabolism, a lung adenocarcinoma recurrence risk prediction model is constructed, and high-risk warning signals are triggered by time-series analysis of risk factor fluctuations.

[0068] Initial weights are generated based on multi-omics data inference. The student network simulates residual feedback, performs dynamic correction, and smooths the time-series sliding window to form time-varying weights for significant causal boundary sets of methylation metabolism. Through dynamic simulation of methylation level data by the student network, the activity changes are calculated by difference and weighted by the time-varying weights of significant causal boundary sets of methylation metabolism to generate metabolic pathway activity fluctuations.

[0069] Specifically, the multi-omics data includes folate metabolomics, transcriptomics, and methylation level data. Based on the integrated analysis of multi-omics data, statistical models are used to infer the potential correlation strength between molecular levels, generating initial weights as the basic input for subsequent dynamic correction. A student network simulation error and residual feedback mechanism is constructed, and iterative optimization is performed using the initial weights as driving parameters to achieve dynamic correction of the weights. A time-series sliding window smoothing algorithm is added to suppress noise in the time dimension of the corrected weight sequence, forming time-varying weights of methylation metabolism significant causal edge sets with time resolution. On this basis, the student network is further used to perform time-series dynamic simulation of methylation level data to obtain the trend of methylation state changes at different time points. The difference method is used to calculate the activity change rate of each node, and the change signals are weighted and aggregated in combination with the aforementioned time-varying weights to characterize the activity fluctuation trajectory of each metabolic pathway in different time periods, thus obtaining the activity fluctuation of metabolic pathways.

[0070] By spatiotemporally fusing simulated metabolic pathway features with time-varying weights of significant causal boundary sets of methylation metabolism, a lung adenocarcinoma recurrence risk prediction model is constructed based on a deep hierarchical survival model.

[0071] Specifically, based on a deep hierarchical survival model, this study integrates temporal methylation data, metabolomic concentrations, and transcriptomic expression profiles. A teacher network (based on the attention mechanism of the KEGG pathway metapath) generates high-fidelity metabolic pathway activity features, while a student network uses dynamic causal masks to screen key methylation regulatory sites and generate simulated metabolic features. Combined with the initial causal edge set inferred from the Bayesian dynamic causal model (BDCM) and temporal residual feedback correction, dynamic time-varying weights are generated. A spatiotemporal heterogeneous graph is constructed (nodes = spatial micro-regions, edges = spatiotemporal adjacency + causal weights), which is mapped to a risk score through a fully connected layer. A high-risk spatiotemporal heatmap generation module is also integrated to form a lung adenocarcinoma recurrence risk prediction model.

[0072] By using the survival analysis calculation logic of the lung adenocarcinoma recurrence risk prediction model, the combined effect of metabolic pathway activity fluctuations and methylation metabolism significant causal boundary set time-varying weights is quantified to generate time-series analysis risk factor fluctuations, detect residual fluctuation anomalies in time-series analysis risk factor fluctuations, and trigger high-risk warning signals.

[0073] Specifically, the survival analysis calculation logic of the lung adenocarcinoma recurrence risk prediction model is input into the joint effect quantification module along with the time-varying weights of metabolic pathway activity fluctuations and methylation metabolism significant causal boundary sets. A time-series risk score is calculated using a linear weighted formula, generating a risk factor fluctuation curve (e.g., risk score increasing from 0.6 to 0.9). This yields the time-series analysis of risk factor fluctuations, calculates the residual of the current risk score, collects and organizes the risk scores of known patients at different time points and their corresponding clinical outcomes, and defines the residual by calculating the difference between the predicted and actual observed risk scores. Distribution fitting is performed on all collected residuals, and the standard deviation of the residuals is estimated using a normal distribution. Based on the obtained standard deviation, a residual threshold is set; if the residual continuously exceeds the residual threshold, a high-risk warning signal is triggered.

[0074] S6. By solving the reaction-diffusion equation, the perturbation propagation path of the folic acid metabolism network is simulated, and a relapse prediction report with a spatial invasion heatmap and temporal evolution curve is output.

[0075] Based on the simulated metabolic pathway characteristics and the time-varying weights of the significant causal edge set of methylation metabolism, a dynamic model of the propagation of metabolic perturbations in the tumor space is constructed by defining the perturbation source, diffusion coefficient and reaction rate of the folate metabolism network through the reaction-diffusion equation.

[0076] Specifically, based on the time-varying weights of the simulated metabolic pathway feature vectors generated by the student network and the significant causal boundary set of methylation metabolism, metabolic perturbation sources, diffusion coefficients, and reaction rates are defined through the reaction-diffusion equation. Spatiotemporal partial differential equations are constructed and solved discretized in the tumor spatial coordinate system (micro-area label registration accuracy ±50μm) using the finite difference method. Boundary conditions are initialized by combining clinical imaging data, and the parameters of the metabolic perturbation propagation dynamics model are dynamically corrected by fusing real-time metabolomics data through a Kalman filter to generate the metabolic perturbation propagation dynamics model.

[0077] The reaction-diffusion equation is discretized and numerically solved to simulate the spatiotemporal propagation of metabolic perturbations in the spatial omics coordinate system. A standardized invasion intensity distribution is generated and mapped to the tumor spatial omics coordinate system to generate a spatial invasion heatmap aligned with anatomical images and mark high-invasion-risk areas.

[0078] Specifically, based on the discretized numerical solution of the reaction-diffusion equation, the metabolic disturbance source (simulating metabolic characteristic fluctuations), diffusion coefficient, and reaction rate (weighted sum of causal effects) are input into the metabolic disturbance propagation dynamics model. Through iterative calculation (example values ​​are 1000 iterations, residual convergence threshold), the spatiotemporal propagation process of metabolic disturbance in the spatial omics coordinate system is simulated, and the standardized invasion intensity distribution is output. The invasion intensity is mapped to the tumor spatial omics coordinate system through an affine transformation matrix to generate a spatial invasion heatmap and obtain high invasion risk areas.

[0079] It should be noted that time-series data of metabolite concentrations (e.g., 5-MTHF) were collected from healthy individuals (N=100 in the example) and postoperative stable patients (N=200 in the example). The residuals between the values ​​and the predicted values ​​were calculated, and the 95th percentile of the residuals in the healthy group was used as the initial threshold. The residual convergence threshold was adjusted to match the clinical recurrence risk sensitivity through ROC curve optimization, and verified by CRISPR perturbation experiments. Combining computational efficiency and model stability, the residual convergence threshold was obtained.

[0080] Extract the time evolution curve of high-risk areas and combine it with high-risk early warning signals to generate a recurrence prediction report.

[0081] Specifically, based on the high-risk areas marked in the spatial invasion heatmap, the standardized invasion intensity value change sequence over time is extracted. After smoothing the noise with a Savitzky-Golay filter, a time evolution curve is generated (the horizontal axis is postoperative time, and the vertical axis is invasion intensity). Combined with high-risk warning signals, a logistic regression model (input = slope of time curve + warning signal intensity, output = recurrence probability) is used to calculate the dynamic risk score. Key trend features are extracted through time window sliding analysis to generate a structured recurrence prediction report.

[0082] This embodiment also provides a lung adenocarcinoma recurrence prediction system based on multi-omics data analysis, including: a data acquisition module, a margin set module, a teacher network module, a student network optimization module, a construction module, and a reporting module; the data acquisition module is used to collect folate metabolome data, transcriptome expression profiles, and methylation level data of target gene promoter regions from tumor tissues of lung adenocarcinoma patients; the margin set module is used to construct a metabolism-epigenetic association topology graph based on spatial adjacency relationships, screen significant causal relationships between methylation changes and metabolic fluctuations through causal analysis, and generate significant causal margin sets for methylation metabolism; the teacher network module is used to construct a significant causal margin set for methylation-metagenesis... The system constructs a dual-pathway distillation framework for epigenetic metabolism, folic acid metabolomics data and transcriptomics data are integrated in the teacher network to generate metabolic pathway activity feature vectors; a student network optimization module is used to optimize methylation level data through the student network to obtain simulated metabolic pathway features; a construction module is used to integrate simulated metabolic pathway features with significant causal edge sets of methylation metabolism to construct a lung adenocarcinoma recurrence risk prediction model, and trigger high-risk warning signals through time-series analysis of risk factor fluctuations; a reporting module is used to simulate the perturbation propagation path of the folic acid metabolism network by solving the reaction-diffusion equation, and output a recurrence prediction report with spatial invasion heatmap and temporal evolution curve.

[0083] This embodiment also provides a computer device applicable to the lung adenocarcinoma recurrence prediction method based on multi-omics data analysis, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the lung adenocarcinoma recurrence prediction method based on multi-omics data analysis as proposed in the above embodiment.

[0084] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0085] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the lung adenocarcinoma recurrence prediction method based on multi-omics data analysis as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0086] In summary, this invention improves the robustness and specificity of recurrence prediction models by: constructing a metabolic epigenetic association topology map and screening significant causal edge sets; preserving the biological significance of spatial locations and revealing hidden molecular interaction mechanisms; constructing an epigenetic dual-pathway distillation framework to generate simulated metabolic pathway features; integrating multi-omics data through a teacher network and optimizing methylation data and simulating features through a student network; achieving knowledge transfer dimensionality reduction, dynamic simulation to enhance interpretability, and synergistic enhancement of prediction performance; and overcoming the limitations of single-omics data by mining hidden molecular interaction networks, significantly improving the accuracy of lung adenocarcinoma recurrence prediction.

[0087] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for predicting recurrence of lung adenocarcinoma based on multi-omics data analysis, characterized in that: include, Data on folate metabolomics, transcriptomic expression profiles, and methylation levels of target gene promoter regions were collected from tumor tissues of lung adenocarcinoma patients. A metabolic-epiota association topology graph is constructed based on spatial adjacency. Significant causal relationships between methylation changes and metabolic fluctuations are screened through causal analysis, and a significant causal edge set of methylation metabolism is generated. The specific steps for generating a significant causal edge set for methylation metabolism are as follows. Based on the micro-region marker coordinate system, the spatial coordinates of folic acid metabolomics data, transcriptomics data and methylation level data are finely aligned to generate spatial adjacency relationships. A metabolism-epigenetic association topology was constructed by weighting both spatial adjacency and the similarity of folic acid metabolome data. The initial causal edges in the metabolism-epigenesis association topology graph are iteratively optimized using a meta-reinforcement learning algorithm. With local causal direction consistency and statistical significance as constraints, a set of significant causal edges of methylation metabolism across spatial locations is selected and generated. An epigenetic dual-pathway distillation framework was constructed based on the methylation-metabolism significant causal edge set, folic acid metabolomics data and transcriptomics data were fused in the teacher network and metabolic pathway activity feature vectors were generated. The proposed epigenetic-metabolic two-pathway distillation framework, based on significant causal boundary sets of methylation metabolism, integrates folic acid metabolomics and transcriptomics data within a teacher network to generate metabolic pathway activity feature vectors. The specific steps are as follows: Based on the significant causal edge set of methylation metabolism, a biological interaction logic network describing folate metabolomics data, transcriptome expression profiles and methylation level data is constructed. The teacher network forms a cross-modal fusion pathway, and the student network forms a causal-driven simulation pathway and a joint training mechanism, which constitutes an epigenetic-metabolic dual-pathway distillation framework. Based on the teacher network, a metapath-aware attention mechanism is used to fuse folic acid metabolomics data and transcriptome expression profiles to generate metabolic pathway activity feature vectors. Simulated metabolic pathway characteristics were obtained by optimizing methylation level data through a student network. By fusing simulated metabolic pathway characteristics with significant causal boundary sets of methylation metabolism, a lung adenocarcinoma recurrence risk prediction model is constructed, and high-risk warning signals are triggered by time-series analysis of risk factor fluctuations. By solving the reaction-diffusion equation to simulate the perturbation propagation path of the folic acid metabolism network, a relapse prediction report with a spatial invasion heatmap and temporal evolution curve is output.

2. The method for predicting lung adenocarcinoma recurrence based on multi-omics data analysis as described in claim 1, characterized in that: The specific steps for collecting folate metabolomics data, transcriptomic expression profiles, and methylation level data of target gene promoter regions from lung adenocarcinoma patient tumor tissues are as follows. Lung adenocarcinoma tumor tissue was frozen and fixed, and serially sectioned into folic acid metabolome sections, transcriptome sections, and methylation detection region sections; A micro-region marker coordinate system was preset on a glass slide, and dynamic adaptive laser desorption ionization imaging was performed on folic acid metabolome slices to obtain folic acid metabolome data. In situ single-cell methylation barcode probe hybridization was performed on methylation detection region slices to detect methylation level data of target gene promoter regions. Transcriptome slices were spatially barcoded and whole transcriptome expression data were obtained through in situ RNA capture probe hybridization to obtain transcriptome expression profiles.

3. The method for predicting lung adenocarcinoma recurrence based on multi-omics data analysis as described in claim 1, characterized in that: Methylation level data and significant causal edge sets of methylation metabolism are input into the student network. The first distillation path dynamically filters key regulatory sites in the methylation level data through causal masks, while the second distillation path aligns metabolic pathway activity feature vectors through an adversarial loss function to generate simulated metabolic pathway features.

4. The method for predicting lung adenocarcinoma recurrence based on multi-omics data analysis as described in claim 1, characterized in that: The specific steps for triggering high-risk early warning signals through time-series analysis of risk factor fluctuations are as follows: Initial weights are generated based on multi-omics data inference, and dynamic correction and time-series sliding window smoothing are combined with student network simulation residual feedback to form time-varying weights of significant causal boundary set of methylation metabolism. The methylation level data is dynamically simulated through student network, and the activity change is calculated by difference and weighted by the time-varying weights of significant causal boundary set of methylation metabolism to generate metabolic pathway activity fluctuations. By spatiotemporally fusing simulated metabolic pathway features with time-varying weights of significant causal edge sets of methylation metabolism, a lung adenocarcinoma recurrence risk prediction model is constructed based on a deep hierarchical survival model. By using the survival analysis calculation logic of the lung adenocarcinoma recurrence risk prediction model, the combined effect of metabolic pathway activity fluctuations and methylation metabolism significant causal boundary set time-varying weights is quantified to generate time-series analysis risk factor fluctuations, detect residual fluctuation anomalies in time-series analysis risk factor fluctuations, and trigger high-risk warning signals.

5. The method for predicting lung adenocarcinoma recurrence based on multi-omics data analysis as described in claim 4, characterized in that: The process involves simulating the propagation path of folic acid metabolic network perturbations by solving the reaction-diffusion equation, and outputting a relapse prediction report with spatial invasion heatmaps and temporal evolution curves. The specific steps are as follows: Based on the simulated metabolic pathway characteristics and the time-varying weights of the significant causal edge set of methylation metabolism, the perturbation source, diffusion coefficient and reaction rate of the folate metabolism network are defined by the reaction-diffusion equation, and a dynamic model of the propagation of metabolic perturbation in the tumor space is constructed. The reaction-diffusion equation is discretized and numerically solved to simulate the spatiotemporal propagation of metabolic perturbations in the spatial omics coordinate system. A standardized invasion intensity distribution is generated and mapped to the tumor spatial omics coordinate system to generate a spatial invasion heat map aligned with anatomical images and mark high invasion risk areas. Extract the time evolution curve of high-risk areas and combine it with high-risk early warning signals to generate a recurrence prediction report.

6. A lung adenocarcinoma recurrence prediction system based on multi-omics data analysis, based on the lung adenocarcinoma recurrence prediction method based on multi-omics data analysis according to any one of claims 1 to 5, characterized in that: It includes a data acquisition module, a side set module, a teacher network module, a student network optimization module, a construction module, and a reporting module; The acquisition module is used to collect folate metabolome data, transcriptome expression profiles, and methylation level data of target gene promoter regions from tumor tissues of lung adenocarcinoma patients. The edge set module is used to construct a metabolism-epigenetic association topology graph based on spatial adjacency relationships. It uses causal analysis to screen for significant causal relationships between methylation changes and metabolic fluctuations, and generates significant causal edge sets for methylation metabolism. The teacher network module is used to construct an epigenetic two-way distillation framework based on the methylation-metabolism significant causal edge set. It integrates folate metabolomics data and transcriptomics data in the teacher network and generates metabolic pathway activity feature vectors. The student network optimization module is used to optimize methylation level data through the student network to obtain simulated metabolic pathway characteristics. A module is built to fuse simulated metabolic pathway features with significant causal edge sets of methylation metabolism to construct a lung adenocarcinoma recurrence risk prediction model, and trigger high-risk warning signals by analyzing the fluctuations of risk factors over time. The reporting module is used to simulate the propagation path of folic acid metabolic network perturbations by solving the reaction-diffusion equation, and output a relapse prediction report with a spatial invasion heatmap and a time evolution curve.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the lung adenocarcinoma recurrence prediction method based on multi-omics data analysis as described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the lung adenocarcinoma recurrence prediction method based on multi-omics data analysis as described in any one of claims 1 to 5.