Crohn disease early diagnosis marker screening method fusing multiple omics data

By constructing a deep multi-omics time-series fusion network model, we screened and validated early diagnostic biomarkers for Crohn's disease, which solved the problem of insufficient shallow integration and validation of multi-omics data in existing technologies, and achieved early risk warning and improved biomarker reliability.

CN121838867APending Publication Date: 2026-04-10WENZHOU MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies lack screening methods for early diagnostic biomarkers for Crohn's disease that can deeply integrate multi-omics data, focus on the temporal dynamics of disease evolution, and lack systematic validation, resulting in diagnostic delays and low likelihood of biomarker clinical translation.

Method used

A deep multi-omics temporal fusion network model was constructed, including a multi-omics feature encoding layer, a cross-omics attention fusion layer, a temporal dynamic modeling layer, and an early state classification layer. The model was trained with longitudinal multi-time point data, candidate biomarkers were screened, and the model was validated and in vitro functional experiments were conducted in an independent cohort to build an early diagnostic model.

Benefits of technology

It significantly improves the biological interpretability and predictive efficacy of biomarker sets, enables early risk warning, enhances the reliability and clinical translation feasibility of biomarkers, and ensures the possibility of shifting the diagnostic window and early intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838867A_ABST
    Figure CN121838867A_ABST
Patent Text Reader

Abstract

The invention discloses a Crohn disease early diagnosis marker screening method fused with multi-omics data. The method comprises the following steps: constructing a Crohn disease evolution time sequence specificity multi-omics sample set, and synchronously obtaining genome, transcriptome, proteome and metabolome data through longitudinal sampling; a deep multi-omics time sequence fusion network model is constructed and applied to screen a candidate marker set, the network realizes multi-dimensional data deep integration through a multi-omics feature coding layer and a cross-omics cross attention fusion layer, and an evolution law of markers is captured through a time sequence dynamic modeling layer; the technical problems that in an existing marker screening method, multi-omics data integration is shallow, disease dynamic time sequence information is neglected, and a verification system is incomplete are solved, and the Crohn disease early diagnosis marker with high sensitivity and specificity can be efficiently and accurately screened out from complex multi-omics data; and the accuracy and clinical transformation potential of early diagnosis are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of medical diagnostics and bioinformatics, and in particular to a method for screening early diagnostic biomarkers for Crohn's disease that integrates multi-omics data. Background Technology

[0002] Crohn's disease is a chronic and relapsing inflammatory bowel disease with a complex pathogenesis involving the interaction of genetic, immune, microbial, and environmental factors. Currently, the diagnosis of Crohn's disease relies on a combination of clinical symptoms, endoscopic findings, imaging, and histopathological examination. These methods often only provide definitive conclusions when the disease has progressed to a stage of significant inflammation or tissue damage, leading to diagnostic delays and missing the optimal intervention window. Therefore, there is an urgent clinical need for reliable biomarkers capable of identifying Crohn's disease at an early stage, or even in the preclinical phase.

[0003] With the development of high-throughput detection technologies, single-mathematical techniques such as genomics, transcriptomics, proteomics, and metabolomics have been widely used to identify Crohn's disease-related biomarkers. For example, studies have reported gene loci (such as NOD2) associated with Crohn's disease risk, differentially expressed gene transcripts, specific serum cytokines, or fecal metabolites. However, biomarkers screened based solely on single-mathematical data often suffer from low sensitivity, insufficient specificity, or poor stability, making them difficult to apply to complex early diagnostic scenarios. This is because the development of Crohn's disease is a dynamic process driven by multiple factors and pathways, and information at the single molecular level cannot fully capture its complexity.

[0004] The application of existing multi-omics data integration and analysis strategies in Crohn's disease biomarker screening is still in its early stages and has significant limitations, specifically in the following three aspects: First, existing methods are mostly "shallow integrations," such as simply listing differentially identified molecules from different omics or compiling lists after correlation analysis. This approach fails to achieve organic fusion of multi-omics data at the deep feature level at the algorithmic level, and cannot reveal the regulatory networks and synergistic relationships between different levels of biomolecules (such as how gene variations affect gene expression, and consequently protein function and metabolites), resulting in weak interpretability and limited improvement in predictive efficacy of the screened biomarker set. Second, existing studies generally ignore the temporal dimension of disease development. Crohn's disease evolves dynamically from a subclinical state to clinical onset, while most current studies are based on cross-sectional sample designs, resulting in static "snapshot" data that cannot distinguish which molecular changes are early driving events of the disease and which are late secondary phenomena. This makes the screened biomarkers potentially only related to the disease state, rather than true early warning signals. Third, there is a lack of a systematic closed-loop validation process from computational prediction to biological function and clinical validation. Many studies stop at a list of biomolecules that show good classification performance in training cohorts, without validation in rigorously designed independent time-series cohorts, and lack subsequent in vitro functional experiments to support their biological rationale, resulting in low likelihood of clinical translation of biomarkers.

[0005] In summary, existing technologies lack a method for systematically screening and validating early diagnostic biomarkers for Crohn's disease that can deeply integrate multi-omics data and pay particular attention to the temporal dynamics of disease evolution. This invention aims to address this technical problem. Summary of the Invention

[0006] To achieve the above objectives, this invention provides a method for screening early diagnostic biomarkers for Crohn's disease by integrating multi-omics data, the method comprising the following steps: Step 1: Construct a time-specific multi-omics sample set for Crohn's disease evolution: Obtain longitudinal biological samples from multiple time points, including Crohn's disease patients, high-risk individuals, and healthy controls. Simultaneously perform genomics, transcriptomics, proteomics, and metabolomics tests on each sample at each time point to generate a multi-omics dataset with time-series and status labels. Step 2: Construct and apply a deep ensemble analysis network to screen a set of candidate biomarkers: Construct a deep multi-omics temporal fusion network model, which sequentially includes a multi-omics feature encoding layer, a cross-omics attention fusion layer, a temporal dynamic modeling layer, and an early state classification layer; train the deep multi-omics temporal fusion network model using the multi-omics sample set generated in Step 1 to distinguish between early-transformed individuals and healthy individuals; based on the trained deep multi-omics temporal fusion network model, extract the features that contribute the most to the classification, backtrack to the original omics data, and screen a set of candidate biomarkers; Step 3: Systematically validate the candidate biomarker set and build an early diagnostic model: Validate the predictive efficacy of the candidate biomarker set in an independent time series queue, conduct in vitro functional experiments to validate key candidate biomarkers, and build an early diagnostic model based on the validated biomarker subset.

[0007] Preferably, in step 1, constructing a time-series-specific multi-omics sample set for Crohn's disease evolution specifically includes: Step 1.1: Determine the study subject cohort: The study subject cohort includes a cohort of patients diagnosed with Crohn's disease, a cohort with high-risk factors for Crohn's disease but without clinical symptoms, and a healthy control cohort; Step 1.2: Design the longitudinal sampling plan: For all the above-mentioned study subject cohorts, including the patient cohort, the high-risk cohort, and the healthy control cohort, conduct no less than three longitudinal follow-up samplings, with each sampling interval being a fixed period; for the Crohn's disease patient cohort, the sampling time points cover the disease activity period and the remission period; for the high-risk cohort and the healthy control cohort, sampling is conducted at the same time interval. Step 1.3: Collect and process biological samples: At each sampling time point, peripheral blood and fecal samples are collected simultaneously from each research subject; Step 1.4: Perform multi-omics data detection: Perform whole-genome sequencing on blood samples to generate genomic data, perform RNA sequencing on peripheral blood mononuclear cells in blood samples to generate transcriptomics data, perform non-targeted proteomics mass spectrometry analysis on blood serum to generate proteomics data, and perform non-targeted metabolomics mass spectrometry analysis on fecal samples to generate metabolomics data. Step 1.5: Perform data quality control and labeling: Perform standardized quality control processing on the generated raw omics data, including removing low-quality data, batch effect correction, and data normalization; assign a queue type label, sampling time point number label, and disease activity index label to each omics data of each sample to form a structured time-specific multi-omics sample set.

[0008] Preferably, generating multi-omics data in step 1.4 specifically includes: The genomic data includes single nucleotide polymorphism data and copy number variation data detected from whole genome sequencing results; The transcriptomics data includes quantitative data on gene expression levels obtained from RNA sequencing results; The proteomics data includes serum proteins and their relative abundance data identified from non-targeted proteomics mass spectrometry analysis; The metabolomics data includes fecal metabolites and their relative abundance data identified from untargeted metabolomics mass spectrometry analysis.

[0009] Preferably, in step 2, the specific construction and training process of the deep multi-omics temporal fusion network model includes: Step 2.1: Construct a multi-omics feature coding layer: The multi-omics feature coding layer contains four independent coding sub-networks, namely, a genome coding sub-network, a transcriptome coding sub-network, a proteome coding sub-network, and a metabolome coding sub-network; each coding sub-network receives the corresponding type of raw omics data as input, and maps the high-dimensional features to a unified low-dimensional feature space through a fully connected operation, outputting the feature coding vector of the corresponding omics; Step 2.2: Constructing a cross-omics attention fusion layer: The cross-omics attention fusion layer receives feature encoding vectors from four omics at the same time point. In this layer, the feature encoding vector of each omics is used as the query vector in sequence, and the feature encoding vectors of the other three omics are used together as the key vector and value vector. The relevance weight between the query omics and other omics is calculated through an attention mechanism, and the value vectors are weighted and summed to obtain the feature representation of the query omics after cross-omics information enhancement. This process is repeated to enhance the feature encoding vector of each omics. Finally, the four enhanced feature representations are concatenated to output the deep fusion feature vector at that time point. Step 2.3: Constructing the Temporal Dynamic Modeling Layer: The temporal dynamic modeling layer consists of a temporal convolutional network. The deep fusion feature vectors of all time series of the same individual are input into the temporal convolutional network in chronological order. The temporal convolutional network contains multiple convolutional layers with dilated causal convolutional structures to capture the long-term evolution pattern of the fusion features in the temporal dimension. The output of the last time step of the temporal convolutional network is used as a summary feature vector representing the multi-omics dynamic evolution pattern of the individual. Step 2.4: Construct an early state classification layer: The early state classification layer consists of a fully connected neural network that receives the summary feature vector and outputs a binary classification probability to predict whether an individual belongs to the early transformation type or the healthy type. Step 2.5, Model Training: Divide the multi-omics sample set generated in Step 1 into a training set and a validation set according to individuals. Using the training set data, with the goal of minimizing the cross-entropy loss of early state classification, train all parameters of the deep multi-omics temporal fusion network model through the backpropagation algorithm, and use the validation set for performance monitoring to prevent overfitting.

[0010] Preferably, in step 2, the process of extracting the features that contribute most to classification and backtracking to the original omics data to screen for a set of candidate biomarkers includes: On the trained deep multi-omics temporal fusion network model, the gradient-based ensemble gradient method is used to calculate the importance score of the early state classification layer output to each feature of the model input layer. Set an importance score threshold to filter out the original omics data features whose importance scores are higher than the threshold; The selected features are categorized according to their respective omics types, resulting in lists of biomolecules from genomics, transcriptomics, proteomics, and metabolomics. These biomolecule lists together constitute a set of candidate biomarkers.

[0011] Preferably, in step 3, the process of verifying the predictive performance of the candidate marker set in an independent time-series queue specifically includes: Step 3.1: Obtain an independent time series queue: Collect a brand new independent time series queue. The inclusion criteria, sampling scheme, sample type and multi-omics detection process of the independent time series queue are exactly the same as those in Step 1, and ensure that the individuals in the independent time series queue are not included in the sample set used in Step 1. Step 3.2: Evaluate the predictive performance: Input the multi-omics data from the independent time series cohorts into the deep multi-omics time series fusion network model trained in Step 2 to obtain the model's early risk prediction probability for each individual at each time point; calculate the model's prediction accuracy, sensitivity, specificity, positive predictive value, and negative predictive value based on the predefined transformation events and observation endpoints.

[0012] Preferably, in step 3, the in vitro functional experimental verification specifically includes: Step 3.3: Select key regulatory markers: From the candidate marker set, select a combination of markers that have known or predicted regulatory relationships and belong to different omics. The combination of markers constitutes a key regulatory node. The regulatory relationships include the relationship between transcription factors and target genes, enzymes and metabolites, and upstream and downstream molecules of signaling pathways. Step 3.4: Design cell model experiments: Use human intestinal epithelial cell lines or immune cell lines as in vitro models; use gene overexpression or gene silencing technology to change the expression level of genes or proteins in key node markers; Step 3.5, Simulated Inflammatory Stimulation and Phenotypic Detection: Cells are stimulated with tumor necrosis factor-α or lipopolysaccharide to simulate the intestinal inflammatory environment; phenotypic changes related to Crohn's disease pathology in the cell model are detected, including cell barrier permeability, inflammatory cytokine secretion levels, and cell apoptosis rate; simultaneously, mass spectrometry or enzyme-linked immunosorbent assay (ELISA) is used to detect changes in the levels of relevant proteins or metabolites in key node biomarkers to confirm the functional roles of these biomarkers in the inflammatory pathway.

[0013] Preferably, in step 3, the specific process of constructing the simplified early diagnosis index model includes: Step 3.6: Determine the final biomarker subset: Based on the stability analysis results of independent time-series cohort validation and the biological supporting evidence of in vitro functional experiments, select the final biomarker subset from the candidate biomarker set; Step 3.7: Construct a logistic regression model: Using the standardized expression levels or concentrations of all biomolecules in the final biomarker subset at the baseline time point or a specific time point as input features, and whether an individual eventually develops Crohn's disease as a binary output label, the logistic regression algorithm is used to fit the model parameters on the training set data to obtain the simplified early diagnosis index model. Step 3.8: Determine the diagnostic threshold: Calculate the risk probability output by the simplified early diagnosis index model on the validation set, and select the risk probability value that maximizes the Youden index as the positive judgment threshold for early diagnosis through receiver operating characteristic curve analysis.

[0014] Preferably, the specific calculation process of the cross-omics attention fusion layer in step 2.2 includes: For a given time point, let G be the genomics feature coding vector, T be the transcriptomics feature coding vector, P be the proteomics feature coding vector, and M be the metabolomics feature coding vector; First, the genomics feature encoding vector G is used as the query vector. The concatenation of the transcriptomics feature encoding vector T, the proteomics feature encoding vector P, and the metabolomics feature encoding vector M is used as the key vector K1 and the value vector V1. The attention weight matrix is ​​calculated, and the value vector V1 is weighted and summed to obtain the genomics enhancement feature vector G'. Secondly, the transcriptomics feature encoding vector T is used as the query vector, and the concatenation of the genomics feature encoding vector G, the proteomics feature encoding vector P, and the metabolomics feature encoding vector M is used as the key vector K2 and the value vector V2. The attention weight matrix is ​​calculated, and the value vector V2 is weighted and summed to obtain the transcriptomics enhancement feature vector T'. Next, the proteomics feature encoding vector P is used as the query vector, and the concatenation of the genomics feature encoding vector G, transcriptomics feature encoding vector T, and metabolomics feature encoding vector M is used as the key vector K3 and the value vector V3. The attention weight matrix is ​​calculated, and the value vector V3 is weighted and summed to obtain the proteomics enhancement feature vector P'. Finally, the metabolomics feature encoding vector M is used as the query vector, and the concatenation of the genomics feature encoding vector G, transcriptomics feature encoding vector T, and proteomics feature encoding vector P is used as the key vector K4 and the value vector V4. The attention weight matrix is ​​calculated, and the value vector V4 is weighted and summed to obtain the metabolomics enhancement feature vector M'. The genomics-enhanced feature vector G', transcriptomics-enhanced feature vector T', proteomics-enhanced feature vector P', and metabolomics-enhanced feature vector M' are concatenated to generate a deep fusion feature vector for that time point.

[0015] Preferably, the structure and processing flow of the temporal convolutional network in step 2.3 include: The temporal convolutional network consists of a one-dimensional dilated causal convolutional layer, a weight normalization layer, a rectified linear unit activation function, and a random deactivation layer stacked sequentially. The dilation coefficient of the one-dimensional dilated causal convolutional layer increases exponentially with the network depth to ensure that the receptive field can cover the entire input temporal sequence. The deep fusion feature vectors of the same individual at multiple time points are arranged in chronological order and input into a temporal convolutional network. The temporal convolutional network performs convolution operations in the time dimension, and each convolutional layer depends only on the information of the current and previous time points to ensure temporal causality. The output feature map of the last convolutional layer of the temporal convolutional network at the last time step, after a global pooling operation, forms a summary feature vector representing the multi-omics dynamic evolution pattern of the individual.

[0016] The beneficial effects of this invention are: 1. This invention fundamentally changes the integration model of multi-omics data by constructing a deep neural network containing a cross-omics attention fusion layer. This technical feature enables the model to actively and dynamically learn and model the nonlinear interactions and regulatory relationships between features at different omics levels, rather than simply performing data splicing or correlation calculations. It can extract biologically relevant biomarker combinations from complex data; these combinations better reflect the overall pathological mechanisms of diseases, thereby significantly improving the biological interpretability and synergistic predictive efficacy of the selected biomarker set. 2. This invention innovatively combines longitudinal follow-up data with deep learning temporal modeling by constructing a time-specific sample set and designing a time-series dynamic modeling layer. This enables the model to learn the continuous evolution and trajectory of biomarkers before the clinical onset of disease, effectively distinguishing between stable, trending early abnormal signals and transient physiological fluctuations. The selected biomarkers are not only related to disease status but also possess the potential to provide risk warnings before the appearance of clinical symptoms, thus truly shifting the diagnostic window forward and providing the possibility for early intervention. 3. This invention establishes a systematic validation process by integrating independent time-series cohort validation and in vitro functional experimental validation as essential steps with the preliminary computational screening stage. The screened biomarkers undergo rigorous multi-dimensional validation across three dimensions: statistical reliability, independent reproducibility, and biological functionality. This not only significantly improves the reliability and credibility of the final biomarker subset but also lays a solid foundation for the subsequent development into diagnostic tools or models with clear biological significance and clinical applicability, significantly enhancing the clinical translation feasibility and application value of the entire technical solution. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of the steps of the method of the present invention; Figure 2 The flowchart below shows the specific steps of constructing a time-specific multi-omics sample set for the evolution of Crohn's disease in the method of this invention. Figure 3 This is a flowchart illustrating the specific steps involved in constructing and training the deep multi-omics temporal fusion network model in the method of this invention. Detailed Implementation

[0019] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. It should also be noted that, to make the embodiments more comprehensive, the following embodiments are the best and preferred embodiments, and those skilled in the art can use other alternative methods to implement some well-known technologies; moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.

[0020] Please see Figures 1-3This invention provides a method for screening early diagnostic biomarkers for Crohn's disease by integrating multi-omics data. First, a time-series-specific multi-omics sample set for Crohn's disease evolution is constructed. This step is specifically implemented by establishing a research cohort comprising at least three groups: confirmed Crohn's disease patients, asymptomatic individuals with high-risk factors, and a healthy control group. All enrolled subjects are longitudinally tracked, and biological samples are collected at fixed time intervals (e.g., every 6 months), with a total sampling frequency of no less than three times.

[0021] At each sampling point, peripheral venous blood and stool samples were collected simultaneously from each subject. Blood samples were subsequently used for whole-genome sequencing, peripheral blood mononuclear cell RNA sequencing, and serum non-targeted proteomics analysis, respectively; stool samples were used for non-targeted metabolomics analysis. All generated raw data underwent rigorous quality control, batch correction, and normalization, and were assigned three-dimensional labels: cohort affiliation (patient / high-risk / healthy), sampling sequence number (time point 1, 2, 3...), and clinical disease activity index (if applicable). This resulted in a structured, time-labeled dataset integrating genomics, transcriptomics, proteomics, and metabolomics. Next, a deep ensemble analysis network was constructed and applied to screen a set of candidate biomarkers.

[0022] The network model is constructed using four sequentially connected layers: the first layer is a multi-omics feature encoding layer, which uses four independent neural network submodules to receive and process raw data from one omics, converting it into a low-dimensional feature vector of a uniform scale; the second layer is a cross-omics attention fusion layer, which receives feature vectors from four omics at the same time point and generates a deep fusion feature vector that integrates multi-dimensional information by having each omics feature take turns acting as an "interrogator" to retrieve and fuse information from other omics; the third layer is a temporal dynamic modeling layer, which inputs a series of deep fusion feature vectors of the same individual at different time points into a temporal convolutional network specifically for processing sequential data to capture the dynamic patterns of feature changes over time and outputs a summary feature vector representing the individual's full-cycle dynamics; the fourth layer is an early state classification layer, a simple classification network responsible for determining whether an individual belongs to the "high-risk to disease conversion" type based on the summary feature vector. The model is trained using the sample set constructed in the first step, with the training goal of enabling the model to accurately distinguish between early high-risk individuals who eventually become patients and individuals who remain healthy throughout.

[0023] After training, specific algorithms (such as gradient backtracking) are used to analyze the key features upon which the model makes its judgments. These features are then back-mapped back to the initial genomic, transcriptomic, proteomic, and metabolomic data to screen a set of candidate biomarkers. Finally, the candidate biomarker set is systematically validated, and an early diagnostic model is built. This requires repeating the above sampling and detection process in a completely new cohort of volunteers, entirely independent of the training set, to verify the predictive accuracy of the aforementioned network model and candidate biomarkers in this new cohort.

[0024] Simultaneously, for the screened biomarkers that may be located in key biological pathways (e.g., a gene and a metabolite it regulates), functional validation was performed in laboratory cell models. Changes in the levels of related metabolites and cellular inflammatory responses were observed by altering gene expression to confirm their biological validity. Based on a subset of biomarkers that passed independent and functional validation, a diagnostic index model was constructed using simplified models such as logistic regression to calculate the risk of early disease with only a small number of biomarkers, and its diagnostic threshold was determined.

[0025] By employing "longitudinal sampling-temporal modeling," the problem of existing cross-sectional studies being unable to capture early dynamic signals was solved; by using "cross-attention deep fusion," the deficiency of simple data splicing in revealing deep connections between omics was overcome; and by employing a triple validation system of "computational screening-independent validation-functional experiments," the reliability and clinical translation potential of the final biomarker were ensured, thus systematically achieving efficient and accurate screening of early diagnostic biomarkers for Crohn's disease.

[0026] In one possible implementation, step 1.1, determining the study subject cohort, requires adherence to clearly defined clinical inclusion and exclusion criteria. The inclusion criteria for the Crohn's disease patient cohort are: newly diagnosed patients in active or remission phases based on clinical, endoscopic, imaging, and pathological criteria.

[0027] The inclusion criteria for the high-risk factor cohort were: having at least one first-degree relative with Crohn's disease, or carrying a known high-risk genotype (such as a specific variant of the NOD2 gene) through genetic testing, but currently having no intestinal symptoms and normal colonoscopy results. The healthy control cohort consisted of volunteers with no history of gastrointestinal diseases, no family history of autoimmune diseases, and matched age and sex. All participants were required to sign informed consent forms. The longitudinal sampling protocol in step 1.2 requires a detailed follow-up plan to be developed in advance.

[0028] For example, baseline visits (time point 0), 6-month visits (time point 1), and 12-month visits (time point 2) should be set. For the patient group, a standard disease activity index (such as the Crohn's Disease Activity Index, CDAI) should be used for assessment and recording at each visit. For the high-risk and healthy groups, visits should be conducted at the same time points, and any new symptoms should be recorded. The biosample collection and processing procedure in step 1.3 must be standardized. Peripheral blood should be collected in EDTA anticoagulant tubes and serum separation tubes, and plasma, serum, and peripheral blood mononuclear cells should be separated, aliquoted, and cryopreserved (at -80 degrees Celsius) within a specific time after collection. Stool samples should be collected by participants at home using pre-prepared preservation solution collection boxes and delivered to the laboratory within 24 hours, homogenized, aliquoted, and cryopreserved.

[0029] Step 1.4, multi-omics data analysis, was performed using a professional platform. Whole-genome sequencing utilized a high-throughput sequencing platform to sequence blood DNA, with bioinformatics workflows used to detect single nucleotide polymorphisms and copy number variations. RNA sequencing involved extracting total RNA from isolated peripheral blood mononuclear cells, constructing libraries, and sequencing. Gene expression matrices were obtained through alignment and counting. Serum non-targeted proteomics employed liquid chromatography-mass spectrometry (LC-MS / MS), acquiring mass spectra using a data-dependent acquisition mode, identifying peptides and proteins using database searches, and calculating relative intensities. Fecal non-targeted metabolomics also used LC-MS / MS, identifying compounds and obtaining their peak areas by comparing them to a public metabolite database. Step 1.5, data quality control and labeling, is crucial for data preparation. Quality control included: removing samples with insufficient sequencing depth or excessively low mass spectrometry signal intensity; correcting for technical variations introduced by different testing batches using statistical methods (such as the ComBat algorithm); and performing logarithmic transformation and standardization (such as Z-score standardization) on each omics dataset to meet the requirements of subsequent analysis. The annotation process involves associating each data file of each sample with a metadata table, which clearly records the sample ID, the queue to which it belongs (encoded as a numeric label), the sampling time point number, and the corresponding clinical activity index score.

[0030] This invention provides a high-quality, multi-dimensional, structured data foundation with temporal and clinical phenotypic labels. This rigorous cohort design ensures sample representativeness; longitudinal sampling provides timeline information on disease evolution; simultaneous multi-omics testing guarantees data comparability and correlation; and strict quality control and systematic annotation provide reliable and uniformly formatted input for subsequent complex data analysis and machine learning modeling, which are prerequisites for the correct implementation and effective results of all subsequent analytical steps.

[0031] In one possible implementation, the generation of genomic data begins with quality filtering and removal of adapter sequences from the raw sequences produced by whole-genome sequencing. High-quality reads are then aligned with a human reference genome. Based on the alignment results, specialized software for detecting genetic variations is used to identify single-base differences (single nucleotide polymorphisms) by comparing the sample sequences with the reference sequences, and to record their chromosomal location, reference base, and variant base.

[0032] Simultaneously, the software analyzes the continuity of read depth coverage across the genome, using statistical methods to determine whether there are copy number increases or decreases (copy number variations) in specific genomic regions, and records their start and end positions and copy number status. For transcriptomics data generation, after similar quality control and alignment of RNA sequencing data, gene quantification software is used to count the reads uniquely aligned to each gene coding region, obtaining the raw expression count for each gene in each sample. To eliminate the influence of sequencing depth differences between samples, the raw counts are subsequently standardized, for example, by calculating the number of reads from a specific gene per million reads, thus obtaining quantitative data on gene expression levels.

[0033] The generation of proteomics data involves peak extraction, deconvolution, and calibration of the raw mass spectrometry (MS) spectra. The resulting peptide mass-to-charge ratio and retention time information are then compared with protein sequence databases to identify the peptides and their corresponding proteins. For each identified protein, the relative abundance data in the sample is obtained by normalizing the peak area or intensity of its corresponding peptide in the MS spectrum (e.g., dividing the intensity of all proteins by the total intensity of the sample). The generation process for metabolomics data is similar to that of proteomics. MS data is compared with a metabolite database to identify metabolites. Based on the peak area of ​​the characteristic ions of the identified metabolites in the MS spectrum, after normalization correction using internal standards or the total peak area, the relative abundance data of the metabolite in the fecal sample is obtained.

[0034] This invention clarifies the specific transformation process from raw instrument data to a standardized data matrix suitable for analysis. By defining the specific content and format of each omics data point (variable sites, expression levels, relative abundance), it ensures that different omics data have clear, calculable, and standardized numerical characteristics during subsequent fusion analysis. This clear data definition is the foundation for effective mathematical operations of technical features such as the "cross-omics cross-attention fusion layer," avoiding analytical errors or model failures caused by ambiguous data formats or unclear meanings.

[0035] In one possible implementation, step 2.1 involves constructing a multi-omics feature encoding layer. For genomic data, which is typically sparse information on variant sites (e.g., genotypes at millions of sites), the genomic encoding sub-network first encodes the genotype (e.g., AA, Aa, aa) of each site using one-hot encoding, and then maps the high-dimensional sparse vector to a dense vector of a predetermined dimension (e.g., 128-dimensional) through a fully connected layer. Transcriptome, proteome, and metabolome data are typically continuous expression levels or abundance vectors of genes / proteins / metabolites. The corresponding encoding sub-networks are each multi-layer fully connected neural networks, with the input dimension being the number of corresponding omics features. After several layers of nonlinear transformation, the final output is a low-dimensional vector with the same dimension as the genomic feature vector (e.g., also 128-dimensional). In this way, biological data from four different scales and dimensions are uniformly encoded into the same feature space, facilitating subsequent comparison and fusion.

[0036] In step 2.2, a cross-omics attention fusion layer is constructed. The input to this layer is four 128-dimensional feature vectors at the same time point. The method for implementing cross-attention is as follows: the four vectors are concatenated into a matrix, and the linear transformation matrices of the query, key, and value are calculated separately. For the genomic vector, its transformed vector is used as the query; the transformed vectors of the transcriptome, proteome, and metabolome are concatenated as the key and value. By calculating the dot product of the query and the key and normalizing it, a set of attention weights is obtained. These weights are then used to perform a weighted summation of the value vectors to obtain a new vector that "integrates information from other omics from a genomic perspective." This process is repeated sequentially with the transcriptome, proteome, and metabolome vectors as queries, ultimately resulting in four new 128-dimensional vectors. These four new vectors are then concatenated again to form a 512-dimensional vector, which is the "deeply fused feature vector" at that time point.

[0037] In step 2.3, a temporal dynamic modeling layer is constructed. This layer employs a temporal convolutional network. T 512-dimensional deep-fused feature vectors of the same individual at T time points are arranged chronologically into a T×512 matrix. The temporal convolutional network consists of multiple stacked one-dimensional convolutional blocks. Each block contains a one-dimensional causal convolution (ensuring the output at time t depends only on the input at time t and earlier), an activation function, and a normalization layer. The "dilation" coefficient of the convolutional kernels multiplies layer by layer, enabling higher-level convolutions to cover a long historical period. The network ultimately performs global pooling on the temporal dimension, or takes the output of the last time step, forming a fixed-length (e.g., 256-dimensional) "summary feature vector" that encapsulates the dynamic evolution pattern of the individual's multi-omics features during the observation period.

[0038] In step 2.4, an early state classification layer is constructed. This layer is a simple multilayer perceptron, which takes a 256-dimensional summary feature vector as input, passes through one or more fully connected layers and non-linear activation functions, and finally passes through an output layer to reduce the dimension to 2. The Softmax function is used to output the probability of belonging to the two categories of "early transformation" and "healthy".

[0039] In step 2.5, model training takes place. The sample set is randomly divided into a training set (70%) and a validation set (30%) based on individual IDs, ensuring that all time point data for the same individual exists in only one set. Using the training set data, the cross-entropy between the aforementioned classification probabilities and the true labels is used as the loss function, and all weight parameters of the model are iteratively updated through backpropagation and an optimizer (such as Adam). During training, the model's classification accuracy is evaluated on the validation set after a certain number of iterations. Training is stopped early when the validation set performance no longer improves to prevent the model from overfitting the training data.

[0040] This invention provides a complete and operable deep learning model architecture and training scheme. It achieves co-spatial mapping of multi-source heterogeneous data through a feature encoding layer; creates a mechanism for proactive information exchange between omics through a cross-attention fusion layer, which is the core of deep integration; effectively models temporal dependencies through temporal convolutional layers; and finally, through end-to-end training, the entire model can automatically learn the complex mapping relationship from raw multi-omics temporal data to early disease states. This design enables the model to automatically mine important patterns from massive, dynamic, and multi-dimensional data, far superior to traditional methods that rely on manual feature engineering.

[0041] In one possible implementation, on a well-trained and stable deep multi-omics temporal fusion network model, interpretability analysis is needed to identify the key input features driving the model's "early conversion" judgment. Ensemble gradient methods are an effective means of achieving this. The basic idea is to calculate the gradient of the model output (e.g., the probability of the "early conversion" category) with respect to the input features, and then integrate the input from the baseline value (e.g., an all-zero vector) along the path to the actual value. The resulting integral is the contribution or importance score of that feature to the final output.

[0042] In practice, for a "early transformation" positive sample in the validation set, its real, standardized multi-omics data is input into the model. The ensemble gradient of the model's output probability relative to each original feature of the input layer (i.e., each SNP site, each gene expression level, and each protein or metabolite abundance) is calculated. This process requires dense sampling and gradient calculation for each feature within its value range, ultimately assigning a scalar importance score to each feature. The importance scores of all features constitute a vector with the same dimension as the original input.

[0043] Next, a threshold needs to be set to filter important features. A practical approach is to calculate the average importance score of each feature across all validation set samples (or all “early conversion” samples) and then sort them from highest to lowest. Features in the top 1% (e.g., the top 1%) can be selected, or features whose absolute importance score is greater than a multiple of the average score of all features (e.g., twice the standard deviation) can be selected as significant features.

[0044] Finally, based on the index positions of these significant features in the original data matrix, their corresponding biological entities are traced back. For example, if the 305th input feature is selected, and the first 1000 columns of the data matrix are genomic features, while columns 1001 to 5000 are transcriptomic features, then the index indicates that the feature originates from the transcriptome and corresponds to a specific gene. In this way, significant single nucleotide polymorphisms or copy number variation regions from genomics data, significantly differentially expressed genes from transcriptomics, significantly differentially expressed proteins from proteomics, and significantly differentially expressed metabolites from metabolomics can be listed separately. This list constitutes the "candidate biomarker set."

[0045] This invention provides a reliable technical approach to transform the decision-making basis of a "black box" neural network model into a concrete and understandable list of biomarkers. Ensemble gradient method, as a well-known attribution method, has a clear calculation process and relatively stable results. By setting thresholds based on statistical distribution, the most relevant features can be objectively selected, avoiding subjective assumptions. This step is a crucial bridge in the entire method's transition from a "predictive model" to a "discovery tool," enabling the results of deep learning methods to connect with traditional biological knowledge and providing clear target molecules for subsequent verification and application.

[0046] In one possible implementation, step 3.1, obtaining an independent time-series cohort, is the cornerstone of the entire validation process. This cohort must be completely independent of any data used during the model development phase. Participants should be re-recruited at a different research center or geographic region, following the exact same clinical criteria, and using the exact same longitudinal study protocol (same sampling intervals, same sample types, and same sample pretreatment and preservation procedures).

[0047] All biological samples should be sent to the same testing laboratory or a laboratory whose performance has been rigorously verified to be consistent, using the exact same technical platform and bioinformatics analysis workflow to generate multi-omics data. This strict control is to ensure that the heterogeneity of the validation data comes only from individual biological differences and the natural progression of disease, rather than technical bias, thus making the validation results convincing. In step 3.2, the predictive efficacy evaluation is the key to quantifying the model's generalization ability. Standardized multi-omics data from all individuals and all time points in the independent time series cohorts are input into a pre-trained deep multi-omics time series fusion network model with fixed parameters.

[0048] For each individual, the model outputs their "early conversion risk probability" at each follow-up visit. For evaluation, the observation endpoint and conversion event need to be predefined. For example, the observation endpoint could be set at 24 months after enrollment, and the "conversion event" defined as a clinical diagnosis of Crohn's disease during this period. Two main evaluation strategies can then be employed: one is time-point specific evaluation, such as using the risk probability output by the model at baseline visit to assess its predictive power for conversion within the next 24 months; the other is time-series dynamic evaluation, such as using the highest risk probability output by the model throughout the follow-up period for each individual, or the trend of the risk probability increasing over time, as a predictive indicator. Based on the predicted probabilities and actual conversion outcomes, a series of standardized performance metrics can be calculated, including accuracy (the proportion of all correct predictions), sensitivity (the proportion of those who were correctly predicted to convert), specificity (the proportion of those who were correctly predicted to be negative among those who did not convert), positive predictive value (the proportion of those predicted to convert who actually converted), and negative predictive value (the proportion of those predicted to be negative who actually did not convert).

[0049] Plotting the receiver operating characteristic (ROC) curve and calculating the area under the curve can provide a comprehensive evaluation of the model's discriminative ability.

[0050] Testing the model's performance on a completely new and independent dataset provides the strongest evidence of its generalization ability (i.e., its effectiveness when applied to new patients). This is an indispensable step in evaluating the clinical value of any diagnostic model. A rigorously controlled validation process ensures the impartiality and reliability of the evaluation results. The calculated quantitative indicators provide objective and comparable data to determine whether the model and its selected biomarkers are truly effective, serving as a crucial basis for subsequent decisions regarding larger-scale clinical trials or the development of diagnostic products.

[0051] In one possible implementation, step 3.3, selecting key node biomarkers, requires combining bioinformatics analysis with biological knowledge. From the computationally screened set of candidate biomarkers, molecular pairs or groups that may have a direct relationship in known biological pathways or regulatory networks are sought. For example, a transcription factor gene highly expressed in the transcriptome may also have its encoded protein showing abundance variations in proteomic data, and downstream genes or metabolic enzymes known to be regulated by this transcription factor may also appear in the candidate list.

[0052] This combination constitutes a "critical regulatory node." In step 3.4, designing cell model experiments requires selecting a suitable in vitro system. Human colon adenocarcinoma cell lines (such as Caco-2) can differentiate into models with intestinal epithelial cell characteristics under specific culture conditions, often used for intestinal barrier function studies. Human monocyte cell lines (such as THP-1) can be induced to differentiate into macrophages, used to simulate immune inflammatory responses. Based on the functional hypothesis of the selected critical node, the most relevant cell line is chosen. Using molecular biology techniques, for candidate genes, forced high or low expression of the gene is achieved in cells by transfecting it with its overexpression plasmid or small interfering RNA. In step 3.5, simulating inflammatory stimulation and phenotypic detection aims to verify the function of biomarkers in the pathological process. After completing gene manipulation, cells are treated for a period of time using classic inflammatory stimuli, such as tumor necrosis factor-α or lipopolysaccharide.

[0053] Subsequently, a series of cellular phenotypes associated with Crohn's disease pathology were examined: the barrier integrity of intestinal epithelial cells was assessed using transmembrane resistance measurements or fluorescently labeled dextran permeability assays; the secretion levels of key inflammatory factors such as interleukin-1β, interleukin-6, and tumor necrosis factor-α in cell culture supernatants were detected using enzyme-linked immunosorbent assay (ELISA); and the proportion of apoptotic cells was detected using flow cytometry. Simultaneously, changes in the levels of proteins or metabolites corresponding to key nodes in cells or culture supernatants were detected using the same mass spectrometry techniques or specific immunological methods (such as Western blotting or ELISA) as in the discovery phase. If the effects of gene manipulation (overexpression or knockdown) directly caused the expected changes in related proteins / metabolites and were further significantly correlated with changes in cellular inflammatory phenotypes (such as increased barrier disruption and increased secretion of inflammatory factors), then at the functional level, it would support the fact that this candidate biomarker does indeed participate in the regulation of Crohn's disease-related inflammatory pathways.

[0054] This invention elevates computational discoveries based on data mining to the level of experimental biological validation. This validation not only confirms that the correlation between biomarkers and diseases is not statistically coincidental, but also reveals their potential biological mechanisms and causal directions. The results of in vitro functional experiments imbue candidate biomarkers with biological significance, greatly enhancing their persuasiveness as reliable diagnostic targets and providing preliminary clues for the development of potential future drug targets. This is a crucial step in realizing "translational medicine" from the laboratory to clinical application.

[0055] In one possible implementation, step 3.6, determining the final subset of biomarkers, is a comprehensive decision-making process. The final subset should meet two core criteria: first, it should demonstrate stable and excellent predictive performance in independent time-series cohort validation, for example, its expression levels show a sustained and significant difference between transformed and non-transformed individuals, and it should contribute significantly as model input; second, it should obtain positive supporting evidence from in vitro functional experiments, indicating its involvement in disease-related biological processes. Based on these criteria, a smaller (e.g., 5-10) molecular combination with more robust evidence is selected from a larger set of candidate biomarkers. In step 3.7, a logistic regression model is constructed.

[0056] Logistic regression models are simple, stable, and easily interpreted clinically. The input features of the model are measurements of each molecule in the final biomarker subset at a single, specific time point (typically the baseline time point for readily available clinical samples). These measurements (such as standardized readings of gene expression, protein or metabolite concentrations) undergo the same standardization process as the training set. The output label is binary: whether the individual has progressed to Crohn's disease within the predetermined follow-up period. The logistic regression model is fitted using the initial training set data (or a combination of training and validation set data) from model development. The fitting process involves using algorithms such as maximum likelihood estimation to find a set of coefficients that best matches the model's calculated conversion probability to the actual probability.

[0057] The final model is a linear formula: Risk Score = Coefficient 1 × Biomarker A Level + Coefficient 2 × Biomarker B Level + ... + Constant Terms. The risk score is then substituted into the Sigmoid function to obtain the conversion risk probability between 0 and 1. In step 3.8, the diagnostic threshold is determined. The fitted logistic regression model is applied to an independent validation set (or a reserved test set) to calculate the risk probability for each individual. Receiver Operating Characteristic (ROC) curves are plotted based on the actual outcomes (converted / unconverted) and the calculated risk probabilities for all individuals.

[0058] Each point on the curve corresponds to a specific risk probability threshold (i.e., the cutoff value for a positive result). By calculating the sum of sensitivity and specificity (i.e., the Youden index) for each threshold, the threshold that maximizes the Youden index is identified as the optimal diagnostic threshold. This threshold will be used in future clinical applications: when a new individual's biomarker test value is input into the model, if the calculated risk probability is higher than this threshold, it is classified as "high-risk positive," and enhanced monitoring or early intervention is recommended; if it is lower than the threshold, it is classified as "low-risk negative."

[0059] This invention represents a transformation from complex deep learning research models to simple, practical, and deployable clinical diagnostic tools. The logistic regression model requires only the detection of a few biomarkers, is computationally simple and fast, and is well-suited for clinical laboratory settings. Clearly defined diagnostic thresholds provide clear clinical decision points. This "simplified model," while maintaining core predictive performance, significantly reduces the cost and technical barriers to future clinical applications, making it a key step in moving this invention from a technical solution to a practical product and social benefit.

[0060] In one possible implementation, the core of this layer is to compute a projection of a "query," "key," and "value" for each omics feature. First, learnable weight matrices are defined for the four input feature vectors to project each omics feature vector onto the query space, key space, and value space. Assuming a unified feature dimension of D, the dimensions of the projected query, key, and value vectors are D_q, D_k, and D_v, respectively, typically D_k = D_v. Specifically, for a genomic feature vector G, its query vector Q_G = G × W_QG is first calculated, where W_QG is the genomic query projection matrix.

[0061] Next, the bond and value vectors of the transcriptome, proteome, and metabolome feature vectors T, P, and M are calculated: K_T = T × W_KT, V_T = T × W_VT; K_P = P × W_KP, V_P = P × W_VP; K_M = M × W_KM, V_M = M × W_VM. Then, these three bond vectors are concatenated along the feature dimension to form the joint bond matrix K_combined = Concat(K_T, K_P, K_M). Similarly, the three value vectors are concatenated to form the joint value matrix V_combined = Concat(V_T, V_P, V_M).

[0062] Subsequently, the dot product of the genomic query and the joint key matrix is ​​calculated, scaled by dividing by the square root of the key vector dimension, and then normalized using the Softmax function to obtain a set of attention weights. Formulaically, this calculates the similarity between each key vector in Q_G and K_combined. Then, this set of attention weights is used to perform a weighted summation on the joint value matrix V_combined, i.e., the weights are multiplied by each value vector and then summed to obtain the final genomic enhancement feature vector G'. This process can be described as follows: G' is the result of weighted combination of transcriptomic, proteomic, and metabolomic information based on their "relevance" to genomic features, and then summed or concatenated with the original genomic features (depending on implementation details). For the transcriptomic feature vector T, the exact same process is repeated, but this time T is used as the query source, projected to obtain Q_T, while the keys and values ​​of G, P, and M are concatenated as the retrieved object, attention weights are calculated, and weighted summation is performed to obtain T'. Symmetrical operations are also performed on the proteomic feature vector P and the metabolomic feature vector M.

[0063] Ultimately, four enhanced feature vectors G', T', P', and M' are obtained. To form the final output at this time point, these four vectors are usually concatenated directly along the feature dimension to generate a longer deep fusion feature vector.

[0064] By allowing each omics to take turns acting as an "active queryer," the model can dynamically and selectively draw supplementary information from other omics. For example, for an individual with a risk variant in their genome, the model might "notic" changes in transcriptome expression related to the gene pathway with that variant when calculating genomic features, thereby strengthening the expression of that feature. This dynamic, content-dependent information fusion approach is better able to characterize complex biological regulatory networks than simple static weighting or splicing. It is the core technical guarantee for achieving deep integration in this invention and also provides a certain foundation for the interpretability of the model (through analysis of attention weights).

[0065] In one possible implementation, the core component of a temporal convolutional network is a one-dimensional dilated causal convolutional layer. "One-dimensional" means the convolution operation is performed along the time series dimension; "causal" means the convolution output at time t depends only on elements in the input up to time t, achieved through zero-padding at the front end, ensuring the model doesn't use future information during prediction; "dilation" refers to the input stride skipped when the convolutional kernel processes the input, and the dilation coefficient d determines the size of the kernel's receptive field (which grows exponentially with the number of layers). The network is composed of multiple such dilated causal convolutional blocks stacked together.

[0066] The specific operation flow of each convolutional block is as follows: First, a one-dimensional dilated causal convolution operation is performed on the input sequence. Assume the input is a sequence of length T with C_in features, the convolution kernel width is k, and the dilation coefficient is d. At each time point, the convolution operation performs a weighted linear combination of the C_in features from the current time point and its preceding d×(k-1) time points (due to dilation) to generate a new feature. After convolution, a new sequence of length T with C_out features is output. Second, weight normalization is performed on each feature output by the convolution. This is an improved normalization technique that helps stabilize training. Then, nonlinearity is introduced through a rectified linear unit activation function.

[0067] Finally, a random deactivation layer is applied, randomly setting the output of a subset of neurons to zero during training. This is an effective regularization technique to prevent overfitting. Residual connections are typically added between convolutional blocks to alleviate the vanishing gradient problem in deep networks. At the end of the entire temporal convolutional network, a fixed-length summary vector needs to be extracted from the time-varying feature sequence. A common approach is to directly take the output feature map of the last convolutional block at the last time step (i.e., the end of the sequence). Due to the design of dilated causal convolution, this terminal feature has "seen" the historical information of the entire input sequence. This feature map is then flattened or compressed through a fully connected layer to form the final "summary feature vector." This vector encapsulates the evolution trajectory and final state of the individual's multi-omics fusion features throughout the entire observation time window.

[0068] This invention provides an efficient and powerful temporal modeling method. Compared to traditional recurrent neural networks, temporal convolutional networks achieve longer effective historical memory through dilated convolution, and the training process can be parallelized, resulting in higher computational efficiency. Causality ensures the model's suitability for predictive tasks. This structure is particularly well-suited for processing equally spaced longitudinal medical data like that presented in this invention, effectively capturing complex patterns of biomarker levels rising, falling, fluctuating, or remaining stable over time. Incorporating this temporal dynamic information into the model is a core innovation that distinguishes it from diagnostics using only single-time-point data. This allows the model to identify early high-risk individuals with normal static levels but abnormal dynamic trajectories, thereby significantly improving the sensitivity of early warning.

[0069] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0070] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for screening early diagnostic biomarkers for Crohn's disease by integrating multi-omics data, characterized in that, The method includes the following steps: Step 1: Construct a time-specific multi-omics sample set for Crohn's disease evolution: Obtain longitudinal biological samples from multiple time points, including Crohn's disease patients, high-risk individuals, and healthy controls. Simultaneously perform genomics, transcriptomics, proteomics, and metabolomics tests on each sample at each time point to generate a multi-omics dataset with time-series and status labels. Step 2: Construct and apply a deep ensemble analysis network to screen a set of candidate biomarkers: Construct a deep multi-omics temporal fusion network model, which sequentially includes a multi-omics feature encoding layer, a cross-omics attention fusion layer, a temporal dynamic modeling layer, and an early state classification layer; train the deep multi-omics temporal fusion network model using the multi-omics sample set generated in Step 1 to distinguish between early-transformed individuals and healthy individuals; based on the trained deep multi-omics temporal fusion network model, extract the features that contribute the most to the classification, backtrack to the original omics data, and screen a set of candidate biomarkers; Step 3: Systematically validate the candidate biomarker set and build an early diagnostic model: Validate the predictive efficacy of the candidate biomarker set in an independent time series queue, conduct in vitro functional experiments to validate key candidate biomarkers, and build an early diagnostic model based on the validated biomarker subset.

2. The method for screening early diagnostic biomarkers for Crohn's disease by fusing multi-omics data according to claim 1, characterized in that, In step 1, constructing a time-specific multi-omics sample set for Crohn's disease evolution specifically includes: Step 1.1: Determine the study subject cohort: The study subject cohort includes a cohort of patients diagnosed with Crohn's disease, a cohort with high-risk factors for Crohn's disease but without clinical symptoms, and a healthy control cohort; Step 1.2: Design the longitudinal sampling plan: For all the above-mentioned study subject cohorts, including the patient cohort, the high-risk cohort, and the healthy control cohort, conduct no less than three longitudinal follow-up samplings, with each sampling interval being a fixed period; for the Crohn's disease patient cohort, the sampling time points cover the disease activity period and the remission period; for the high-risk cohort and the healthy control cohort, sampling is conducted at the same time interval. Step 1.3: Collect and process biological samples: At each sampling time point, peripheral blood and fecal samples are collected simultaneously from each research subject; Step 1.4: Perform multi-omics data detection: Perform whole-genome sequencing on blood samples to generate genomic data, perform RNA sequencing on peripheral blood mononuclear cells in blood samples to generate transcriptomics data, perform non-targeted proteomics mass spectrometry analysis on blood serum to generate proteomics data, and perform non-targeted metabolomics mass spectrometry analysis on fecal samples to generate metabolomics data. Step 1.5: Perform data quality control and labeling: Perform standardized quality control processing on the generated raw omics data, including removing low-quality data, batch effect correction, and data normalization; assign a queue type label, sampling time point number label, and disease activity index label to each omics data of each sample to form a structured time-specific multi-omics sample set.

3. The method for screening early diagnostic biomarkers for Crohn's disease by fusing multi-omics data according to claim 2, characterized in that, Step 1.4, generating multi-omics data, specifically includes: The genomic data includes single nucleotide polymorphism data and copy number variation data detected from whole genome sequencing results; The transcriptomics data includes quantitative data on gene expression levels obtained from RNA sequencing results; The proteomics data includes serum proteins and their relative abundance data identified from non-targeted proteomics mass spectrometry analysis; The metabolomics data includes fecal metabolites and their relative abundance data identified from untargeted metabolomics mass spectrometry analysis.

4. The method for screening early diagnostic biomarkers for Crohn's disease by fusing multi-omics data according to claim 1, characterized in that, In step 2, the specific construction and training process of the deep multi-omics temporal fusion network model includes: Step 2.1: Construct a multi-omics feature coding layer: The multi-omics feature coding layer contains four independent coding sub-networks, namely, a genome coding sub-network, a transcriptome coding sub-network, a proteome coding sub-network, and a metabolome coding sub-network; each coding sub-network receives the corresponding type of raw omics data as input, and maps the high-dimensional features to a unified low-dimensional feature space through a fully connected operation, outputting the feature coding vector of the corresponding omics; Step 2.2: Constructing a cross-omics attention fusion layer: The cross-omics attention fusion layer receives feature encoding vectors from four omics at the same time point. In this layer, the feature encoding vector of each omics is used as the query vector in sequence, and the feature encoding vectors of the other three omics are used together as the key vector and value vector. The relevance weight between the query omics and other omics is calculated through an attention mechanism, and the value vectors are weighted and summed to obtain the feature representation of the query omics after cross-omics information enhancement. This process is repeated to enhance the feature encoding vector of each omics. Finally, the four enhanced feature representations are concatenated to output the deep fusion feature vector at that time point. Step 2.3: Constructing the Temporal Dynamic Modeling Layer: The temporal dynamic modeling layer consists of a temporal convolutional network. The deep fusion feature vectors of all time series of the same individual are input into the temporal convolutional network in chronological order. The temporal convolutional network contains multiple convolutional layers with dilated causal convolutional structures to capture the long-term evolution pattern of the fusion features in the temporal dimension. The output of the last time step of the temporal convolutional network is used as a summary feature vector representing the multi-omics dynamic evolution pattern of the individual. Step 2.4: Construct an early state classification layer: The early state classification layer consists of a fully connected neural network that receives the summary feature vector and outputs a binary classification probability to predict whether an individual belongs to the early transformation type or the healthy type. Step 2.5, Model Training: Divide the multi-omics sample set generated in Step 1 into a training set and a validation set according to individuals. Using the training set data, with the goal of minimizing the cross-entropy loss of early state classification, train all parameters of the deep multi-omics temporal fusion network model through the backpropagation algorithm, and use the validation set for performance monitoring to prevent overfitting.

5. The method for screening early diagnostic biomarkers for Crohn's disease by fusing multi-omics data according to claim 4, characterized in that, In step 2, the process of extracting the features that contribute most to classification and backtracking to the original omics data to screen for a set of candidate biomarkers includes: On the trained deep multi-omics temporal fusion network model, the gradient-based ensemble gradient method is used to calculate the importance score of the early state classification layer output to each feature of the model input layer. Set an importance score threshold to filter out the original omics data features whose importance scores are higher than the threshold; The selected features are categorized according to their respective omics types, resulting in lists of biomolecules from genomics, transcriptomics, proteomics, and metabolomics. These biomolecule lists together constitute a set of candidate biomarkers.

6. The method for screening early diagnostic biomarkers for Crohn's disease by fusing multi-omics data according to claim 1, characterized in that, In step 3, the process of verifying the predictive performance of the candidate marker set in an independent time-series queue specifically includes: Step 3.1: Obtain an independent time series queue: Collect a brand new independent time series queue. The inclusion criteria, sampling scheme, sample type and multi-omics detection process of the independent time series queue are exactly the same as those in Step 1, and ensure that the individuals in the independent time series queue are not included in the sample set used in Step 1. Step 3.2: Evaluate the predictive performance: Input the multi-omics data from the independent time series cohorts into the deep multi-omics time series fusion network model trained in Step 2 to obtain the model's early risk prediction probability for each individual at each time point; calculate the model's prediction accuracy, sensitivity, specificity, positive predictive value, and negative predictive value based on the predefined transformation events and observation endpoints.

7. The method for screening early diagnostic biomarkers for Crohn's disease by fusing multi-omics data according to claim 1, characterized in that, In step 3, the in vitro functional experimental verification specifically includes: Step 3.3: Select key regulatory markers: From the candidate marker set, select a combination of markers that have known or predicted regulatory relationships and belong to different omics. The combination of markers constitutes a key regulatory node. The regulatory relationships include the relationship between transcription factors and target genes, enzymes and metabolites, and upstream and downstream molecules of signaling pathways. Step 3.4: Design cell model experiments: Use human intestinal epithelial cell lines or immune cell lines as in vitro models; use gene overexpression or gene silencing technology to change the expression level of genes or proteins in key node markers; Step 3.5, Simulated Inflammatory Stimulation and Phenotypic Detection: Cells are stimulated with tumor necrosis factor-α or lipopolysaccharide to simulate the intestinal inflammatory environment; phenotypic changes related to Crohn's disease pathology in the cell model are detected, including cell barrier permeability, inflammatory cytokine secretion levels, and cell apoptosis rate; simultaneously, mass spectrometry or enzyme-linked immunosorbent assay (ELISA) is used to detect changes in the levels of relevant proteins or metabolites in key node biomarkers to confirm the functional roles of these biomarkers in the inflammatory pathway.

8. The method for screening early diagnostic biomarkers for Crohn's disease by fusing multi-omics data according to claim 1, characterized in that, In step 3, the specific process of constructing the simplified early diagnosis index model includes: Step 3.6: Determine the final biomarker subset: Based on the stability analysis results of independent time-series cohort validation and the biological supporting evidence of in vitro functional experiments, select the final biomarker subset from the candidate biomarker set; Step 3.7: Construct a logistic regression model: Use the standardized expression level or concentration of all biomolecules in the final biomarker subset at the baseline time point or a specific time point as the input feature, and use whether the individual eventually develops Crohn's disease as the binary output label. Use the logistic regression algorithm to fit the model parameters on the training set data to obtain a simplified early diagnosis index model. Step 3.8: Determine the diagnostic threshold: Calculate the risk probability output by the simplified early diagnosis index model on the validation set, and select the risk probability value that maximizes the Youden index as the positive judgment threshold for early diagnosis through receiver operating characteristic curve analysis.

9. The method for screening early diagnostic biomarkers for Crohn's disease by fusing multi-omics data according to claim 4, characterized in that, The specific calculation process of the cross-omics attention fusion layer described in step 2.2 includes: For a given time point, let G be the genomics feature coding vector, T be the transcriptomics feature coding vector, P be the proteomics feature coding vector, and M be the metabolomics feature coding vector; First, the genomics feature encoding vector G is used as the query vector. The concatenation of the transcriptomics feature encoding vector T, the proteomics feature encoding vector P, and the metabolomics feature encoding vector M is used as the key vector K1 and the value vector V1. The attention weight matrix is ​​calculated, and the value vector V1 is weighted and summed to obtain the genomics enhancement feature vector G'. Secondly, the transcriptomics feature encoding vector T is used as the query vector, and the concatenation of the genomics feature encoding vector G, the proteomics feature encoding vector P, and the metabolomics feature encoding vector M is used as the key vector K2 and the value vector V2. The attention weight matrix is ​​calculated, and the value vector V2 is weighted and summed to obtain the transcriptomics enhancement feature vector T'. Next, the proteomics feature encoding vector P is used as the query vector, and the concatenation of the genomics feature encoding vector G, transcriptomics feature encoding vector T, and metabolomics feature encoding vector M is used as the key vector K3 and the value vector V3. The attention weight matrix is ​​calculated, and the value vector V3 is weighted and summed to obtain the proteomics enhancement feature vector P'. Finally, the metabolomics feature encoding vector M is used as the query vector, and the concatenation of the genomics feature encoding vector G, transcriptomics feature encoding vector T, and proteomics feature encoding vector P is used as the key vector K4 and the value vector V4. The attention weight matrix is ​​calculated, and the value vector V4 is weighted and summed to obtain the metabolomics enhancement feature vector M'. The genomics-enhanced feature vector G', transcriptomics-enhanced feature vector T', proteomics-enhanced feature vector P', and metabolomics-enhanced feature vector M' are concatenated to generate a deep fusion feature vector for that time point.

10. The method for screening early diagnostic biomarkers for Crohn's disease by fusing multi-omics data according to claim 4, characterized in that, The structure and processing flow of the temporal convolutional network described in step 2.3 include: The temporal convolutional network consists of a one-dimensional dilated causal convolutional layer, a weight normalization layer, a rectified linear unit activation function, and a random deactivation layer stacked sequentially. The dilation coefficient of the one-dimensional dilated causal convolutional layer increases exponentially with the network depth to ensure that the receptive field can cover the entire input temporal sequence. The deep fusion feature vectors of the same individual at multiple time points are arranged in chronological order and input into a temporal convolutional network. The temporal convolutional network performs convolution operations in the time dimension, and each convolutional layer depends only on the information of the current and previous time points to ensure temporal causality. The output feature map of the last convolutional layer of the temporal convolutional network at the last time step, after a global pooling operation, forms a summary feature vector representing the multi-omics dynamic evolution pattern of the individual.