Multi-omics deep learning based colorectal cancer liver metastasis detection method and system

By combining transcriptome analysis and deep learning models, biomarkers with CRC metastasis commonality and CRLM specificity were screened out, and a cross-omics integrated framework was constructed to solve the problems of early prediction and biomarker identification in colorectal cancer liver metastasis detection, achieving high-precision early detection and personalized treatment support.

CN121281845BActive Publication Date: 2026-03-27ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies have poor early prediction performance in the detection of colorectal cancer liver metastases. Single-mathematical methods ignore metabolic characteristics and immune environment information, have poor cross-platform model adaptability, lack reliable biomarkers, and are difficult to achieve accurate detection and support for treatment strategies.

Method used

By combining large-scale transcriptome analysis and deep learning models, and through feature gene cluster screening and batch effect correction, a cross-omics integrated framework is constructed to screen out biomarkers with CRC metastasis commonality and CRLM specificity. Combined with metabolomics data, a multi-layer core metabolic network is constructed to realize early detection and personalized treatment strategies.

Benefits of technology

It achieves high-precision prediction of liver metastases before they are visible on imaging, identifies biomarkers with mechanistic significance, supports personalized treatment decisions, and improves the ability to screen and assess the risk of colorectal cancer liver metastases early.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121281845B_ABST
    Figure CN121281845B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-omics deep learning colorectal cancer liver metastasis detection method and system, comprising: extracting the feature gene group with CRC transfer-metabolic dual function, and construct CRC transfer risk prediction model foundation, based on analysis CRC transfer risk probability to screen the molecular marker with CRC transfer nature and CRLM specific characteristics as initial candidate biomarker;Meanwhile, from the metabolome to obtain the differential expression metabolite corresponding to CRLM patient sample and LCRC patient sample metabolome, and the enrichment pathway analysis of transcriptome and metabolome is combined to construct CRLM multi-layer core metabolic network, for the screening core biomarker of colorectal cancer liver metastasis accurate detection in accordance with, entire method can effectively screen the biomarker for detecting colorectal cancer liver metastasis.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of bioinformatics, deep learning and tumor molecular diagnosis, and particularly relates to a colorectal cancer liver metastasis detection method and system based on multi-omics deep learning. BACKGROUND

[0002] Colorectal cancer (CRC) is the third most common malignant tumor in the world, and liver metastasis (CRLM) is one of the main causes of death. Early detection and risk assessment of liver metastasis can significantly improve the resectability and prognosis.

[0003] Current clinical early CRLM screening mainly relies on imaging (CT, MRI, PET, ultrasound) combined with serum tumor markers (such as CEA, CA19-9) detection. This kind of method has high sensitivity in finding formed metastatic lesions, but has limited predictive ability in the preoperative and early stages of tumor. The spatial resolution and detection window of imaging often make it difficult to identify small metastatic lesions before they form larger lesions. Such early signs of lesions are often discovered through molecular detection, while the specificity of conventional serum markers in early liver metastasis patients is insufficient, resulting in a high rate of false positives or false negatives.

[0004] In recent years, some studies have used genomic or transcriptomic data for prognosis evaluation, such as KRAS / NRAS, BRAF, or microsatellite stability (MSS) and instability (MSI) typing. This kind of method can guide targeted therapy or immunotherapy for liver metastasis patients to some extent, but still has significant limitations in early screening of liver metastasis: limited information dimension: single omics cannot capture the complex changes of tumor metabolic reprogramming and immune microenvironment; insufficient prediction performance: due to the multi-factor regulation of the metastasis process, including metabolic state, immune escape mechanism and microenvironment heterogeneity, relying only on genotype or transcriptome local features is difficult to achieve high sensitivity and high specificity prediction, and is limited to the explanation of the biological mechanism of metastasis; poor cross-platform portability: there are significant batch effects between different sequencing platforms and cohorts, and the generalization performance of a single sequencing platform on different datasets is low.

[0005] Some studies have attempted to integrate transcriptome, proteome, metabolome and other data for tumor typing or mechanism research. For example, Sinkala et al. analyzed the correlation between metabolic gene variation and drug response in multiple cancers; however, most of these multi-omics methods remain at the level of statistical correlation or difference analysis, based on relatively direct linear relationships and relatively simple feature dimension analysis, lacking analysis of complex high-dimensional features and non-linear relationships in real medical scenarios, and lacking high-performance prediction models for early clinical screening. At the same time, in practical applications, multi-omics models often encounter the following difficulties: difficulty in integrating cross-platform data: data sources are diverse and batch differences are obvious, direct integration can introduce noise and weaken the generalization ability of the model; high-dimensional data modeling limitations: metabolome and transcriptome feature dimensions are large, and there are many non-linear relationships, traditional machine learning models (such as random forests and support vector machines) have insufficient ability to analyze non-linear relationships in high-dimensional data, and are prone to overfitting and poor generalization; lack of precise markers: the marker screening of existing multi-omics research often lacks mechanism verification and clinical application evaluation, and it is difficult to implement in specific scenarios such as preoperative prediction, risk stratification, and treatment decision support.

[0006] In summary, the existing liver metastasis risk prediction technology has the following common deficiencies: poor early prediction performance: imaging and traditional markers have limited prediction ability when patients have no obvious tumor load in the early stage; incomplete utilization of omics information: single omics ignores key information such as metabolic characteristics and immune environment, reducing the discriminant power and comprehensive biological function analysis ability; poor cross-cohort stability: the model lacks adaptability and portability, and lacks a robust batch correction framework; lack of clinical verification of markers: it is difficult to provide mechanism-based and verified molecular targets for risk assessment.

[0007] Therefore, there is an urgent need for a deep learning framework that can stably operate in cross-platform, integrated multi-omics data, both for early prediction and to identify markers with mechanism significance and clinical translation potential, thereby supporting individualized monitoring and treatment strategy development. SUMMARY

[0008] In view of the above, the purpose of the present application is to provide a multi-omics deep learning-based colorectal liver metastasis detection method and system that combines large-scale transcriptome analysis, deep learning modeling, and serum metabolomics cross-omics set methods for early detection, biomarker discovery, risk assessment, and individualized treatment strategy support of colorectal liver metastasis (CRLM). The method and system can be extended to metastasis risk prediction and molecular mechanism research of other malignant tumor types.

[0009] To achieve the above-mentioned purpose of the application, the multi-omics deep learning-based colorectal liver metastasis detection method provided by the embodiment comprises the following steps:

[0010] Screening a feature gene group with CRC metastasis-metabolic reprogramming bifunctionality based on transcriptome dataset, and sample batch effect correction for large-scale transcriptomic cohort for deep learning and prediction;

[0011] Using the feature gene group as an input item, using the feature gene group to conduct deep learning on CNN to construct a liver metastasis risk prediction model, calculating a transcriptome risk score based on the metastasis risk probability of the CRC metastasis risk prediction model, and performing correlation analysis of metastatic CRC (hereinafter referred to as mCRC) and CRLM occurrence results according to the transcriptome risk score, screening high correlation features with CRC metastasis and CRLM specificity as initial candidate biomarkers;

[0012] Obtaining differential expression metabolites corresponding to the metabolic group of CRLM patient samples and LCRC patient samples, performing enrichment pathway analysis on the initial candidate biomarkers corresponding to the transcriptome and the differential expression metabolites corresponding to the metabolic group, constructing a CRLM multi-layer core metabolic network with shared enrichment pathways, and marking the candidate biomarkers within the pathways, using the CRLM multi-layer core metabolic network to test the CRLM resolution performance of the candidate biomarkers, and screening core biomarkers, which are used for precise detection of colorectal cancer liver metastasis.

[0013] Preferably, sample batch effect correction is performed based on the transcriptome dataset, including:

[0014] All transcriptome datasets are combined according to shared genes, and the sequencing platform is used as a batch variable and the disease state is used as a biological grouping variable. ComBat is used for batch effect correction, and the biological information related to the disease state is preserved. The correction effect is visualized and verified by PCA, and the average silhouette coefficient of the sample is calculated ≤ 0.1 to confirm that the batch effect is reduced.

[0015] Preferably, the feature gene group of the patient sample is used to calculate the transcriptome risk score based on the liver metastasis risk probability of the liver metastasis risk prediction model, including:

[0016] The average value of the liver metastasis risk probability of the feature gene group of the patient sample in the five sub-models contained in the liver metastasis risk prediction model is used as the transcriptome risk score.

[0017] Preferably, the correlation analysis of metastatic CRC and CRLM occurrence results is performed according to the transcriptome risk score, and high correlation features with CRC metastasis and CRLM specificity are screened as initial candidate biomarkers, including:

[0018] Screening high metastasis risk features from the feature gene group based on transcriptome risk score, calculating the correlation coefficient of all features in the mCRC cohort and the CRLM cohort with the transcriptome risk score respectively, and screening the top N features in the correlation coefficient ranking from the mCRC cohort and the CRLM cohort respectively to obtain the intersection, and the feature is the core feature, which is the molecular marker with CRC metastasis property and CRLM specificity, and is used as the initial candidate biomarker.

[0019] Preferably, the consensus enrichment pathway refers to the intersection analysis of the metabolic pathway enriched with the initial candidate biomarker and the metabolic pathway enriched with the differentially expressed metabolite, and the key metabolic pathway enriched at the same time is obtained, including the retinol metabolic pathway and the tryptophan metabolic pathway.

[0020] Preferably, the CRLM multi-layer core metabolic network constructed based on the consensus enrichment pathway includes key metabolic pathways, biomarkers of regulatory genes, transcription factors and metabolites, wherein the transcription factors act on the biomarkers, the biomarkers act on the pathways, and the pathways produce metabolites.

[0021] According to the CRLM multi-layer core metabolic network, a CRLM core metabolic feature map is drawn, and the candidate biomarkers are marked in the CRLM core metabolic feature map.

[0022] Preferably, the CRLM multi-layer core metabolic network is used to test the CRLM resolution performance of the candidate biomarkers, and the core biomarkers are screened, including:

[0023] (1) Performance evaluation of candidate biomarkers:

[0024] All the candidate biomarkers marked in the CRLM multi-layer core metabolic network are used to distinguish colorectal cancer liver metastasis (CRLM) samples from colorectal cancer primary (CRC PR) samples; the area under the receiver operating characteristic curve of each candidate biomarker is calculated to quantify its discrimination performance, and the marker with the best discrimination performance and the highest sensitivity is selected as the candidate marker for CRLM early screening;

[0025] (2) Performance verification and determination of core biomarkers:

[0026] The optimal candidate biomarker screened in step (1) is compared with a plurality of reported CRLM prediction biomarkers in performance benchmarking to evaluate the discrimination sensitivity of the candidate marker for CRLM patients. Through ROC curve and AUC value comparison analysis, the sensitivity difference of the candidate marker compared with the known CRLM screening marker for CRLM patients is evaluated, and the core biomarker is determined according to the evaluation result.

[0027] Preferably, the method further comprises verifying the mechanism of action of the core biomarker in the key metabolic pathway in combination with the metabolic monitoring data, and implementing metabolic regulation based on the mechanism of action, i.e., metabolic reprogramming.

[0028] To achieve the above-mentioned purposes of the application, the embodiments of the present application further provide a colorectal cancer liver metastasis detection system based on multi-omics deep learning, comprising:

[0029] A dual-function feature screening module is used for screening a feature gene group with CRC metastasis-metabolic reprogramming dual-function based on a transcriptome dataset, and performing sample batch effect correction on a transcriptomic cohort for deep learning and prediction;

[0030] An initial candidate biomarker screening is used for using the feature gene group as an input item, using the feature gene group to perform deep learning on a CNN to construct a liver metastasis risk prediction model, calculating a transcriptome risk score based on a metastasis risk probability of the CRC metastasis risk prediction model, and performing correlation analysis of metastatic CRC and CRLM occurrence results according to the transcriptome risk score, screening high-correlation features with CRC metastasis properties and CRLM specificity as initial candidate biomarkers;

[0031] A core biomarker screening module is used for obtaining differential expression metabolites corresponding to the metabolome of the CRLM patient sample and the LCRC patient sample, performing enrichment pathway analysis on the initial candidate biomarkers corresponding to the transcriptome and the differential expression metabolites corresponding to the metabolome, constructing a CRLM multi-layer core metabolic network with shared enrichment pathways, and marking the candidate biomarkers in the pathways, using the CRLM multi-layer core metabolic network to test the CRLM resolution performance of the candidate biomarkers, and screening core biomarkers, which are used for precise detection of colorectal cancer liver metastasis.

[0032] Compared with the prior art, the present application has at least the following beneficial effects:

[0033] Based on the extraction of the feature gene group with CRC metastasis-metabolic dual-function and the construction of the CRC metastasis risk prediction model, the present application screens molecular markers with CRC metastasis properties and CRLM specificity features as initial candidate biomarkers based on the CRC metastasis risk probability; meanwhile, differential expression metabolites corresponding to the metabolome of the CRLM patient sample and the LCRC patient sample are obtained from the technical point of view of metabolomics, and a CRLM multi-layer core metabolic network is constructed in combination with the enrichment pathways of the transcriptome and the metabolome, which is used for screening core biomarkers for precise detection of colorectal cancer liver metastasis. The whole method can effectively screen biomarkers for detecting colorectal cancer liver metastasis. BRIEF DESCRIPTION OF DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0035] Figure 1 is a flowchart of the colorectal cancer liver metastasis detection method based on multi-omics deep learning provided by the embodiments;

[0036] Figure 2 is a construction flowchart of the feature candidate pool provided by the embodiments;

[0037] Figure 3 is a construction flowchart of the CRC metastasis risk prediction model and the transcriptome risk score calculation provided by the embodiments;

[0038] Figure 4 is a screening result of the feature gene group and a batch effect correction effect provided by the embodiments;

[0039] Figure 5 is a structural schematic diagram of the one-dimensional convolutional neural network provided by the embodiments;

[0040] Figure 6 is a training process, internal validation and external validation performance diagram of the CRC metastasis risk prediction model provided by the embodiments;

[0041] Figure 7 is a benchmark test diagram of the model of the present application and 10 classical machine learning classifiers for identifying the performance of patients with metastatic colorectal cancer provided by the embodiments;

[0042] Figure 8 and Figure 9 is a screening diagram of the features related to the transcriptome risk score (TRS) provided by the embodiments;

[0043] Figure 10 is a functional analysis diagram of the features related to the transcriptome risk score (TRS) provided by the embodiments;

[0044] Figure 11 is a construction flowchart of the CRLM multi-layer core metabolic network provided by the embodiments;

[0045] Figure 12 and Figure 13 is an intersection analysis of the initial candidate biomarkers (genes with metastasis-promoting function, i.e. TRS core features) and the metabolic pathways regulating the metastasis function to identify the key CRLM-related pathways provided by the embodiments;

[0046] Figure 14A CRLM multi-layer core metabolic network diagram provided by the embodiment;

[0047] Figure 15 A method for screening optimal markers from candidate markers provided by the embodiment;

[0048] Figure 16 A CRC cell line verification of the metastasis-promoting phenotype of ACMSD provided by the embodiment;

[0049] Figure 17 A clinical verification of the metastasis-promoting phenotype of the marker ACMSD provided by the embodiment;

[0050] Figure 18 A mechanism verification of the metastasis-promoting phenotype of ACMSD verified by the CRC cell line provided by the embodiment;

[0051] Figure 19 A structure diagram of a colorectal cancer liver metastasis detection system based on multi-omics deep learning provided by the embodiment. DETAILED DESCRIPTION

[0052] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the protection scope of the present application.

[0053] The inventive concept of the present application is that: in view of the following problems existing in the current colorectal cancer liver metastasis (CRLM) early detection and risk assessment technology: the current methods relying on imaging and conventional serum tumor markers are insufficient in sensitivity and specificity in predicting liver metastasis risk in the early stage of clinical, especially when the metastatic lesions have not formed significant imaging features; the molecular detection methods based on a single omics (such as genome, transcriptome) have limited information dimension, ignoring important mechanism factors such as metabolic reprogramming and immune microenvironment changes, and cannot accurately depict the multi-factor role in the liver metastasis process; the existing models have poor generalization ability across sequencing platforms or different patient cohorts, lack of batch effect correction and feature extraction framework that can be stably run in multiple source cohorts; the existing multi-omics research lacks mechanism verification and clinical application evaluation in marker screening, and it is difficult to directly translate the results into preoperative accurate prediction, risk stratification and drug guidance.

[0054] To address the above issues, the technical solution proposed in this invention combines large-scale transcriptomics analysis, deep learning models, and a cross-omics integration method of serum metabolomics to establish an analytical framework that can stably operate on multiple data sources, achieving the following technical tasks: When liver metastases are not yet visible on clinical imaging, the risk of liver metastasis in individual patients can be predicted with high accuracy through gene and molecular characteristics, enabling early screening of CRC patients for metastasis, early detection of patients with high metastasis risk for medical intervention, and saving personal and public medical costs; Based on cross-omics characteristics, metabolic pathways and gene molecules closely related to liver metastasis can be identified and pinpointed, discovering novel biomarkers with mechanistic significance; Through clinical samples and in vitro validation, the association between biomarkers and pathological stage, recurrence risk, immune microenvironment characteristics, and sensitivity to specific targeted drugs can be clarified, providing quantifiable evidence for personalized treatment strategies and improving the effectiveness of precision treatment for patients;

[0055] Therefore, this invention aims to effectively overcome the shortcomings of existing early prediction technologies for colorectal cancer liver metastasis in terms of information utilization, model stability, biomarker screening and validation, and to achieve the integrated technical goal of metastasis risk assessment, mechanism analysis and treatment decision support.

[0056] like Figure 1 As shown in the embodiment, the method for detecting colorectal cancer liver metastases based on multi-omics deep learning includes the following steps:

[0057] S1, based on the transcriptome dataset, screened characteristic gene groups with dual functions of CRC transfer-metabolic reprogramming, and performed sample batch effect correction on the transcriptomics cohort.

[0058] In this embodiment, transcriptome datasets containing clinical annotations for colorectal cancer were obtained from multiple public databases (such as GEO, cBioPortal, TCGA, GSE131418, etc.). The data types included microarray expression profiling (GPL570 platform) and RNA-seq expression data (Illumina NGS sequencing). The clinical labels of these data were distinguished as primary CRC (CRC-PR) and metastatic CRC (mCRC), and pathological M stage was marked in samples involved in CRLM.

[0059] In the embodiments, such as Figure 2As shown in FIG. 1A, in screening the characteristic gene group with CRC metastasis-metabolism dual function, cancer metabolism-related genes were extracted from the published tumor metabolic reprogramming gene set, intersected with the mCRC vs CRC-PR regulatory-related differentially expressed genes (DEGs) in the GSE131418 dataset (containing the largest scale of CRC samples), to obtain the characteristic gene group related to CRC metastasis regulation and metabolic function regulation, and to constitute the functional gene pool with CRC metastasis-metabolism dual function as the characteristic candidate pool for model input.

[0060] As shown in FIG. 1A, in screening the characteristic gene group with CRC metastasis-metabolism dual function, cancer metabolism-related genes were extracted from the published tumor metabolic reprogramming gene set, intersected with the mCRC vs CRC-PR regulatory-related differentially expressed genes (DEGs) in the GSE131418 dataset (containing the largest scale of CRC samples), to obtain the characteristic gene group related to CRC metastasis regulation and metabolic function regulation, and to constitute the functional gene pool with CRC metastasis-metabolism dual function as the characteristic candidate pool for model input. Figure 3 As shown in FIG. 1A, in screening the characteristic gene group with CRC metastasis-metabolism dual function, cancer metabolism-related genes were extracted from the published tumor metabolic reprogramming gene set, intersected with the mCRC vs CRC-PR regulatory-related differentially expressed genes (DEGs) in the GSE131418 dataset (containing the largest scale of CRC samples), to obtain the characteristic gene group related to CRC metastasis regulation and metabolic function regulation, and to constitute the functional gene pool with CRC metastasis-metabolism dual function as the characteristic candidate pool for model input.

[0061] As shown in FIG. 1A, in screening the characteristic gene group with CRC metastasis-metabolism dual function, cancer metabolism-related genes were extracted from the published tumor metabolic reprogramming gene set, intersected with the mCRC vs CRC-PR regulatory-related differentially expressed genes (DEGs) in the GSE131418 dataset (containing the largest scale of CRC samples), to obtain the characteristic gene group related to CRC metastasis regulation and metabolic function regulation, and to constitute the functional gene pool with CRC metastasis-metabolism dual function as the characteristic candidate pool for model input. Figure 4 As shown in FIG. 1A, in screening the characteristic gene group with CRC metastasis-metabolism dual function, cancer metabolism-related genes were extracted from the published tumor metabolic reprogramming gene set, intersected with the mCRC vs CRC-PR regulatory-related differentially expressed genes (DEGs) in the GSE131418 dataset (containing the largest scale of CRC samples), to obtain the characteristic gene group related to CRC metastasis regulation and metabolic function regulation, and to constitute the functional gene pool with CRC metastasis-metabolism dual function as the characteristic candidate pool for model input. Figure 4 FIG. 1A shows the results of principal component analysis (PCA) of metastatic colorectal cancer (mCRC) and primary colorectal cancer (CRC PR) groups in the GSE131418 dataset. Figure 4 FIG. 1A shows the results of principal component analysis (PCA) of metastatic colorectal cancer (mCRC) and primary colorectal cancer (CRC PR) groups in the GSE131418 dataset. Figure 4 FIG. 1A shows the results of principal component analysis (PCA) of metastatic colorectal cancer (mCRC) and primary colorectal cancer (CRC PR) groups in the GSE131418 dataset. Figure 4 FIG. 1A shows the results of principal component analysis (PCA) of metastatic colorectal cancer (mCRC) and primary colorectal cancer (CRC PR) groups in the GSE131418 dataset. Figure 4 FIG. 1A shows the results of principal component analysis (PCA) of metastatic colorectal cancer (mCRC) and primary colorectal cancer (CRC PR) groups in the GSE131418 dataset. Figure 4 FIG. 1A shows the results of principal component analysis (PCA) of metastatic colorectal cancer (mCRC) and primary colorectal cancer (CRC PR) groups in the GSE131418 dataset.

[0062] S2, using the feature gene group as an input item, using the feature gene group to perform deep learning on the CNN to construct a liver metastasis risk prediction model, calculating a transcriptome risk score based on a metastasis risk probability of the CRC metastasis risk prediction model, and performing correlation analysis of metastatic CRC and CRLM occurrence results according to the transcriptome risk score, screening high-correlation features with CRC metastasis specificity and CRLM specificity as initial candidate biomarkers;

[0063] In the embodiment, as shown in Figure 3 The batch-corrected patient samples are proportionally divided into a training set and a validation set; in each cross-validation fold, the Z-score standardization parameter is calculated and applied to the validation and test sets.

[0064] As shown in Figure 5 The model architecture of the CRC metastasis risk prediction model adopts a one-dimensional convolutional neural network (1D-CNN), which is used to extract transcriptome features from the feature gene group and predict the CRC metastasis risk probability. The specific structure is: the first convolution-pooling module: convolution layer (Conv1D) (number of convolution kernels = 32, convolution kernel size = 5, activation function ReLU) + maximum pooling layer (MaxPooling1D) (pooling window size = 2 + dropout rate = 0.3); ReLu activation function formula: f(x) = max(0, x); the second convolution-pooling module: convolution layer (Conv1D) (number of convolution kernels = 64, convolution kernel size = 5, activation function ReLU) + maximum pooling layer (MaxPooling1D) (pooling window size = 2 + dropout rate = 0.4); fully connected module: flattening layer (Flatten) + connection layer (Dense) (neurons = 64, activation function ReLU, dropout rate = 0.5); output layer: connection layer (Dense) (neurons = 1, activation function Sigmoid) outputs the metastasis risk probability 0,1.

[0065] In the embodiments, when the CNN is trained using the feature gene group in the feature candidate pool, the loss function is binary cross-entropy, the optimizer is Adam, and the initial learning rate is 0.001. Five-fold stratified cross-validation is also performed (4 folds are used for training, and 1 fold is used for verification), and each fold generates a sub-model, and there are 5 sub-models. Each fold sub-model uses the early stopping method strategy (Early Stopping, patience = 10) and the learning rate plateau decay strategy (Reduce LR On Plateau, decay factor factor = 0.5, patience = 5) to avoid overfitting. When predicting the CRC metastasis risk, the feature gene group of the patient sample is input into the five sub-models for integrated training and prediction, and the metastasis risk probability is calculated after forming the integrated model.

[0066] The embodiments also provide the training process, internal validation, and external validation performance chart of the CRC metastasis risk prediction model of the application, as shown in Figure 6 Figure 6 In the chart A, A is the loss curve of each fold sub-model on the training set and the validation set over time, which is used to evaluate the training convergence and potential overfitting. Figure 6 In the chart B, A is the accuracy curve of each fold sub-model on the training set and the validation set. Figure 6 In the chart C, A is the area under the receiver operating characteristic curve (AUC) of each fold sub-model on the validation set. Figure 6 In the chart D, A is the precision and recall index of each fold sub-model on the validation set. Figure 6 In the chart E, A is the confusion matrix of the internal validation set. Figure 6 In the chart F, A is the area under the receiver operating characteristic curve (AUC) of the internal validation set. Figure 6 In the chart G, A is the receiver operating characteristic curve (ROC) and the precision-recall (PR) curve of the internal validation set, and the corresponding AUC value. Figure 6 In the chart H, A is the confusion matrix of the external test queue (CRLM). Figure 6 In the chart I, A is the area under the receiver operating characteristic curve (AUC) of the external test set. Figure 6 In the chart J, A is the receiver operating characteristic curve (ROC) and the precision-recall (PR) curve of the external test set, and the corresponding AUC value. In the internal five-fold cross-validation, the CRC metastasis risk prediction model of the application has an AUC of 0.92, a precision of 86%, and an accuracy of 85% for the mCRC vs. CRC-PR classification; on the external independent test set (CRLM data set), the model performance remains stable, with an AUC of 0.97, an accuracy of 92%, and a precision of 94% for the CRLM vs. CRC-PR classification.

[0067] ​The embodiments also provide a comparison of the model of the application with 10 kinds of machine learning models with high classification performance such as LightGBM, Random Forest, XGBoost, etc. Figure 7 As shown in Figure 7 FIG. 3A is a comparison of classification performance on an internal validation set, where the x-axis represents AUC and the y-axis represents accuracy. The model located in the upper right quadrant represents the best performance. Figure 7 FIG. 3B is a comparison of classification performance on an external test set, where the x-axis represents AUC and the y-axis represents accuracy. The model located in the upper right quadrant represents the best performance. Tables 1 and 2 also illustrate the accuracy, AUC, precision, and recall of each model on the internal validation set and the external test set. Analysis Figure 7 , Tables 1 and 2, the performance of the method of the application on an independent test set has the smallest decline, the classification performance is the best, and the generalization ability is significantly better than that of existing methods.

[0068] Table 1: Baseline test of classification performance of the validation set

[0069]

[0070] Table 2: Baseline test of classification performance of the external test set

[0071]

[0072] As shown in Figure 3 , when calculating the transcriptome risk score based on the CRC metastasis risk probability, the average value of the liver metastasis risk probability of the characteristic gene group of the patient sample in the five sub-models included in the CRC metastasis risk prediction model is taken as the transcriptome risk score (TRS).

[0073] In the embodiments, the transcriptome risk score (TRS) is also used for correlation analysis of the occurrence of metastatic CRC and CRLM, and high correlation features with CRC metastasis specificity and CRLM specificity are screened as initial candidate biomarkers. The specific process is as follows: based on the transcriptome risk score TRS, high metastasis risk features are screened from the characteristic gene group, and the threshold is set to 0.5. When TRS ≥ 0.5, the characteristic gene group of the patient is determined to be high risk, i.e., the probability of metastasis is higher, and TRS < 0.5 is determined to be low risk, i.e., the probability of metastasis is lower.

[0074] In the feature correlation analysis, the correlation coefficients of all features in the mCRC cohort and the CRLM cohort with the transcriptome risk score TRS are calculated respectively, and the top N (for example, 50) features with the highest correlation coefficients in the mCRC cohort and the CRLM cohort are screened respectively to obtain the intersection of the features, which are the core features. The core features are molecular markers with CRC metastasis properties and CRLM specific features, and are used as initial candidate biomarkers for early screening of CRLM.

[0075] The embodiments specifically provide a diagram for screening and functional analysis of features related to the transcriptome risk score (TRS), Figure 8 The top 20 features with the strongest positive or negative correlation with TRS are shown in the combined mCRC and CRC PR cohort (training set and internal validation set). Figure 9 The top 20 features with the strongest positive or negative correlation with TRS are shown in the external test cohort (CRLM and CRC PR), where the correlation is calculated using. Figure 10 In the middle A, the correlation coefficient of each feature in the training set and the internal validation set with the transcriptome risk score TRS is calculated. Figure 8 And Figure 9 The intersection of the top 50 TRS-related features identified in the evaluation using Spearman correlation is obtained, and 22 core TRS features are obtained. Figure 10 In the middle B, pathway enrichment analysis is performed on the 22 core TRS features, and the significantly enriched metabolic and signaling pathways are highlighted. Figure 10 In the middle C, differential expression analysis of the core TRS features between CRLM and CRC PR samples in the GSE50760 dataset is performed, and the results are presented in the form of log2(FPKM+1). The red box represents the CRLM sample, and the gray box represents the CRC PR sample; the significance level is marked with asterisks. The statistical difference between groups is evaluated using t-test. *P < 0.05; **P < 0.01; ***P < 0.001; ****P < 0.0001; ns, not significant.

[0076] In S3, the differentially expressed metabolites corresponding to the CRLM patient samples and the LCRC patient samples are obtained, the initial candidate biomarkers corresponding to the transcriptome and the differentially expressed metabolites corresponding to the metabolome are subjected to enrichment pathway analysis, a CRLM multi-layer core metabolic network with common enriched pathways is constructed, and the candidate biomarkers within the pathways are labeled. The CRLM multi-layer core metabolic network is used to test the CRLM resolution performance of the candidate biomarkers, and the core biomarkers are screened. The core biomarkers are used for precise detection of colorectal cancer liver metastasis, as shown in Figure 11 .

[0077] In the embodiments, the differentially expressed metabolites of the metabolome data were also obtained, and the analysis process of the metabolome data is a conventional prior art. Specifically, first, the collected serum samples of the CRLM patients and the LCRC patients were subjected to standard LC-MS / MS non-targeted metabolome detection, using a Waters HSS T3 C18 column and gradient elution, and the ESI source and MRM detection were performed according to standard parameters. Then, differential metabolite analysis was performed, and specifically, the orthogonal partial least squares discriminant analysis (OPLS-DA) was used to distinguish the component differences between the sample groups of the metabolome, and the variable importance in projection (variable importance in projection) was used to screen the differential metabolites (DEMs) with VIP>1 and P<0.05.

[0078] In the embodiments, the initial candidate biomarkers corresponding to the transcriptome were also subjected to enrichment pathway analysis. Specifically, the initial candidate biomarkers were subjected to pathway enrichment analysis in the GO and KEGG biological pathway databases, and were significantly enriched in retinol metabolism, tryptophan metabolism and other liver metastasis-related metabolic pathways. The differentially expressed metabolites detected by metabolomics were also subjected to pathway enrichment analysis. The pathways enriched by the initial candidate biomarkers and the pathways enriched by the differentially expressed metabolites were cross-analyzed, that is, the intersection was taken, and the key pathways commonly enriched across the omics were obtained as the common enrichment pathways, including the two key pathways of retinol metabolism and tryptophan metabolism driving CRLM (GSEA NES>2.2, P.adjust<0.001).

[0079] Figure 12 and Figure 13 The intersection analysis of the enrichment pathways of the initial candidate biomarkers (genes with metastasis-promoting functions, that is, the TRS core features) and the metabolites regulating the metastasis function was given to identify the key CRLM-related pathways. Figure 12 The workflow for metabolomics analysis was as follows: serum samples of patients with LCRC (n=15) and CRLM (n=15) were collected, and metabolites were sequenced by LC-MS, followed by metabolite clustering, differential expression analysis and pathway enrichment analysis. Figure 13 In FIG. 1A, A is the OPLS-DA score plot showing clear separation between the CRLM and LCRC samples. Figure 13 In FIG. 1B, B is a volcano plot highlighting the differentially expressed metabolites (DEMs) associated with CRLM. The metabolites were identified as DEMs based on |log2 fold change|>1 or <−1 and P value<0.05. Figure 13The Venn diagram in the middle illustrates the overlap between pathways enriched from the initial candidate biomarkers and those enriched from DEMs. Only pathways satisfying p.adjust < 0.05 or FDR significance are retained for intersection analysis. Figure 13 The diagram in section D summarizes the cross-pathways and their significant enrichment in transcriptomics and metabolomics analyses. Retinol metabolism and tryptophan metabolism are among the most representative enriched pathways. Figure 13 China E and Figure 13 In the figure, F represents the gene set enrichment analysis (GSEA) curves for retinol metabolism pathway and tryptophan metabolism in the CRLM sample (GSE50760 human liver metastasis sequencing dataset).

[0080] In the embodiments, a CRLM multilayer core metabolic network with shared enrichment pathways was also constructed, such as Figure 14 As shown, candidate biomarkers within the pathway are labeled. Specifically, the CRLM multilayer core metabolic network, which enriches the pathway, includes key metabolic pathways, biomarkers of regulatory genes, transcription factors, and metabolites. Transcription factors act on biomarkers, biomarkers act on pathways, and pathways produce metabolites. Based on the CRLM multilayer core metabolic network, a CRLM core metabolic feature map is drawn, and candidate biomarkers are labeled in the CRLM core metabolic feature map.

[0081] In this embodiment, the CRLM multilayer core metabolic network was also used to test the CRLM resolution performance of candidate biomarkers, and core biomarkers (such as ACMSD) were screened. These core biomarkers are used for the precise detection of colorectal cancer liver metastases. The specific process is as follows:

[0082] (1) Performance evaluation of candidate biomarkers:

[0083] All candidate biomarkers labeled in the CRLM multilayer core metabolic network (including ACMSD, ADH1A, CYP2C8, CYP1A2, RDH16, ALDH8A1, CYP2A6, CYP2C9, HSD17B6, and DHTKD1) were used to distinguish between colorectal cancer liver metastases (CRLM) samples and primary colorectal cancer lesions (CRC PR) samples. The discriminative performance of each candidate biomarker was quantified by calculating the area under the receiver operating characteristic (ROC) curve (AUC). The biomarker with the best discriminative performance and highest sensitivity was selected as the candidate biomarker for early CRLM screening. Figure 15 A.

[0084] The AUC value of the ACMSD was 0.97 (95% confidence interval of 0.93-1), demonstrating the best discriminative performance among all candidate biomarkers.

[0085] (2) Performance verification and determination of core biomarkers:

[0086] The optimal candidate biomarker screened in step (1) is compared with a plurality of reported CRLM prediction biomarkers (including CXCL12, FGF19, CXCR4, EGFR, CDX2, ERBB3, ERBB2 and PTGS2) for performance benchmarking, and the candidate marker is evaluated for its sensitivity in distinguishing CRLM patients. The sensitivity of the candidate marker in distinguishing CRLM patients from known CRLM screening markers is evaluated by ROC curve and AUC value comparison analysis. The core biomarker is determined according to the evaluation results. Among them, the classification accuracy of the candidate marker ACMSD in distinguishing CRLM from CRC samples is better than all reported biomarkers; finally, ACMSD is determined as the core biomarker for precise detection of colorectal liver metastasis, as shown in Figure 15 B.

[0087] Most of the biomarkers in the metabolic network have strong performance in distinguishing CRLM patients (AUC: 0.89-0.97), and ACMSD is the best molecular marker for distinguishing CRLM patients, and is a new molecular marker for CRLM screening; compared with the currently reported molecular markers for diagnosis of CRLM patients (ERBB2, FGF19 and CDX2, etc.), ACMSD is still the best molecular marker for distinguishing CRLM patients.

[0088] In the embodiments, the mechanism of the core biomarker in the key metabolic pathway is also verified by combining metabolic monitoring data, and metabolic regulation is realized based on the mechanism, i.e. metabolic reprogramming. Specifically for ACMSD, the mechanism of ACMSD in the tryptophan-NAD + metabolic pathway (promoting the conversion of ACMS to picolinic acid, reducing the synthesis efficiency of quinic acid and NAD + ) and leading to metabolic reprogramming is verified.

[0089] In the embodiments, the core biomarker screened is also clinically verified and evaluated for molecular characteristics to verify the effectiveness of the core biomarker in detecting colorectal liver metastasis, and specific immunohistochemical verification (IHC), immune microenvironment analysis, and drug sensitivity prediction and drug treatment strategies are performed.

[0090] For immunohistochemical verification (IHC), ACMSD antibody is used to detect patient tumor tissue sections taken from cancer tissues of CRLM patients and LCRC patients, and the expression level is quantified; CRC cell lines SW480 and NCI-H716 are used, SW480 is knocked down using biotechnology (shRNA lentivirus infection), and CRC cell strains with target gene knockdown are constructed, as shown in Figure 16 A andFigure 16 WB experiments verified that the protein expression of ACMSD was knocked down after SW480 was infected with shRNA virus, as shown in FIG. 25B. Transwell experiments were used to verify the regulation of target gene ACMSD on the metastasis of CRC cells, and the data of metastatic cells were quantified, as shown in FIG. 25C and FIG. 25D. Figure 16 Figure 16 Figure 17 Figure 17 FIG. 25B shows the protein expression of ACMSD in the tissues of CRLM and LCRC patients, and the statistical correlation between the expression of ACMSD and the risk of postoperative recurrence of CRC patients was analyzed, as shown in FIG. 25C. The expression of ACMSD in the tissues of CRLM and LCRC patients was analyzed, as shown in FIG. 25D. Figure 17 Figure 17 Figure 18 Figure 18 Figure 18

[0091] For immune microenvironment analysis, ssGSEA algorithm was used to calculate the immune cell infiltration in CRLM patient tissues, and the differences in immune activation and immune suppression cells between the high expression group of ACMSD were evaluated. In combination with the sequencing data of CRLM patient tissues, the expression of immune checkpoint genes (PDCD1, LAG3, HAVCR2, etc.) was analyzed to evaluate the immunotherapy potential.

[0092] For drug sensitivity prediction and drug treatment strategy, the correlation between ACMSD expression and CRLM patient therapy (chemotherapy and targeted therapy) sensitivity was calculated based on real world data (data set of cohort with CRLM patient drug administration and drug administration results); the correlation between ACMSD expression and CRC targeted therapy (EGFR / VEGFR inhibitor) sensitivity was calculated based on the drug administration data of GDSC (Genomics of Drug Sensitivity in Cancer, cell drug sensitivity database); OncoPredict algorithm was used for modeling to predict the probability of benefit of patients with high ACMSD expression under specific targeted therapy.

[0093] Figure 19 ​​​​​​​​​As shown, the embodiment also provides a colorectal cancer liver metastasis detection system 30 based on multi-omics deep learning, comprising: a dual-function feature screening module 31, an initial candidate biomarker screening module 32, and a core biomarker screening module 33, wherein the dual-function feature screening module 31 is used for sample batch effect correction based on a transcriptome dataset and screening of a feature gene group with CRC metastasis-metabolic dual function; the initial candidate biomarker screening module 32 is used for deep learning of CNN to construct a liver metastasis risk prediction model using the feature gene group, calculating a transcriptome risk score based on the liver metastasis risk probability of the liver metastasis risk prediction model of the feature gene group of the patient sample, and performing feature association analysis of CRC features and CRLM features according to the transcriptome risk score, screening molecular markers with CRC metastasis characteristics and CRLM specific features as initial candidate biomarkers; the core biomarker screening module 33 is used for obtaining differential expression metabolites corresponding to the CRLM patient sample and the LCRC patient sample metabolic group, performing enrichment pathway analysis on the initial candidate biomarkers corresponding to the transcriptome and the differential expression metabolites corresponding to the metabolic group, constructing a CRLM multi-layer core metabolic network with common enrichment pathways, and marking candidate biomarkers within the pathway, testing the CRLM resolution performance of the candidate biomarkers using the CRLM multi-layer core metabolic network, and screening core biomarkers, the core biomarkers are used for precise detection of colorectal cancer liver metastasis.

[0094] It should be noted that the above embodiment provides a colorectal cancer liver metastasis detection system based on multi-omics deep learning, which is used for precise detection of colorectal cancer liver metastasis, and the above functions are divided into different functional modules for illustration, and the above functions can be completed by different functional modules according to needs, that is, the internal structure of the terminal or server is divided into different functional modules to complete all or part of the above described functions. In addition, the colorectal cancer liver metastasis detection system based on multi-omics deep learning provided by the above embodiment and the colorectal cancer liver metastasis detection method based on multi-omics deep learning belong to the same concept, and the specific implementation process is described in detail in the colorectal cancer liver metastasis detection method based on multi-omics deep learning, which will not be repeated here.

[0095] The specific embodiments described above have explained the technical solutions and advantages of the present application, and it should be understood that the above description is only the most preferred embodiment of the present application, and is not used to limit the present application. Any modification, supplement and equivalent replacement made within the principle range of the present application should be included in the protection scope of the present application.

Claims

1. A method for screening biomarkers for diagnosing liver metastases from colorectal cancer, characterized in that, Includes the following steps: Feature gene clusters with dual functions of CRC transfer-metabolic reprogramming were screened based on transcriptome datasets, and batch effect correction was performed on the transcriptomics cohort for deep learning and prediction. Using the characteristic gene cluster as input, a liver metastasis risk prediction model was constructed by deep learning of CNN using the characteristic gene cluster. The transcriptome risk score was calculated based on the metastasis risk probability of the CRC metastasis risk prediction model. The association between metastatic CRC and CRLM was analyzed based on transcriptome risk scores. Highly correlated features with CRC metastasis generality and CRLM specificity were screened as initial candidate biomarkers. This included: screening high metastasis risk features from characteristic gene groups based on transcriptome risk scores; calculating the correlation coefficients between all features and transcriptome risk scores in both the mCRC and CRLM cohorts; and selecting the intersection of the top N features with the highest correlation coefficients from both the mCRC and CRLM cohorts to obtain the core features. These core features are molecular biomarkers with CRC metastasis generality and CRLM specificity, and are used as initial candidate biomarkers. Differentially expressed metabolites corresponding to the metabolomes of CRLM and LCRC patient samples were obtained. Enrichment pathway analysis was performed on the initial candidate biomarkers corresponding to the transcriptome and the differentially expressed metabolites corresponding to the metabolome. A CRLM multilayer core metabolic network with shared enriched pathways was constructed, and candidate biomarkers within the pathways were labeled. The shared enriched pathways refer to the key metabolic pathways that are simultaneously enriched by cross-analysis of the metabolic pathways enriched by the initial candidate biomarkers and the metabolic pathways enriched by the differentially expressed metabolites. These include the retinol metabolic pathway and the tryptophan metabolic pathway. The CRLM multilayer core metabolic network includes key metabolic pathways, biomarkers of regulatory genes, transcription factors, and metabolites. Transcription factors act on biomarkers, biomarkers act on pathways, and pathways produce metabolites. A CRLM core metabolic feature map was drawn based on the CRLM multilayer core metabolic network, and candidate biomarkers were labeled in the CRLM core metabolic feature map. The CRLM multilayer core metabolic network was used to test the CRLM resolution performance of candidate biomarkers, and ADH1A was selected as the core biomarker. The core biomarker is used for the precise detection of colorectal cancer liver metastasis.

2. The method for screening biomarkers for diagnosing liver metastasis of colorectal cancer according to claim 1, characterized in that, Batch effect correction based on transcriptome datasets includes: All transcriptome datasets were merged according to common genes, and the sequencing platform was used as a batch variable and disease status as a biological grouping variable. ComBat was used for batch effect correction, while retaining the biological information related to disease status. The correction effect was visualized and verified using PCA, and the batch effect was confirmed to be reduced when the mean silhouette coefficient of the samples was ≤ 0.

1.

3. The method for screening biomarkers for diagnosing liver metastasis of colorectal cancer according to claim 1, characterized in that, Based on the characteristic gene clusters of patient samples, a transcriptomic risk score is calculated in the liver metastasis risk prediction model, including: The average probability of liver metastasis risk in the five sub-models included in the liver metastasis risk prediction model for the characteristic gene groups of the patient sample was calculated as the transcriptome risk score.

4. The method for screening biomarkers for diagnosing liver metastasis of colorectal cancer according to claim 1, characterized in that, Also includes: By combining metabolic monitoring data, we can verify the mechanism of action of core biomarkers in key metabolic pathways and achieve metabolic regulation based on the mechanism of action, i.e., metabolic reprogramming.

5. A biomarker screening system for diagnosing liver metastases from colorectal cancer, characterized in that, include: A dual-function feature screening module is used to screen feature gene groups with dual functions of CRC transfer-metabolic reprogramming based on transcriptome datasets, and to perform sample batch effect correction on transcriptomics cohorts for deep learning and prediction. The initial candidate biomarker screening module is used to use the characteristic gene cluster as input, and to use the characteristic gene cluster to perform deep learning on CNN to build a liver metastasis risk prediction model. The transcriptome risk score is calculated based on the metastasis risk probability of the CRC metastasis risk prediction model. The association between metastatic CRC and CRLM was analyzed based on transcriptome risk scores. Highly correlated features with CRC metastasis generality and CRLM specificity were screened as initial candidate biomarkers. This included: screening high metastasis risk features from characteristic gene groups based on transcriptome risk scores; calculating the correlation coefficients between all features and transcriptome risk scores in both the mCRC and CRLM cohorts; and selecting the intersection of the top N features with the highest correlation coefficients from both the mCRC and CRLM cohorts to obtain the core features. These core features are molecular biomarkers with CRC metastasis generality and CRLM specificity, and are used as initial candidate biomarkers. The core biomarker screening module is used to obtain differentially expressed metabolites corresponding to the metabolomes of CRLM and LCRC patient samples. Enrichment pathway analysis is performed on initial candidate biomarkers corresponding to the transcriptome and differentially expressed metabolites corresponding to the metabolome. A CRLM multilayer core metabolic network with shared enriched pathways is constructed, and candidate biomarkers within the pathways are labeled. Shared enriched pathways refer to the key metabolic pathways simultaneously enriched by cross-analysis of metabolic pathways enriched by the initial candidate biomarkers and those enriched by differentially expressed metabolites, including the retinol metabolic pathway and the tryptophan metabolic pathway. The CRLM multilayer core metabolic network includes key metabolic pathways, biomarkers of regulatory genes, transcription factors, and metabolites. Transcription factors act on biomarkers, biomarkers act on pathways, and pathways produce metabolites. A CRLM core metabolic feature map is drawn based on the CRLM multilayer core metabolic network, and candidate biomarkers are labeled in the CRLM core metabolic feature map. The CRLM resolution performance of the candidate biomarkers is tested using the CRLM multilayer core metabolic network, and ADH1A is selected as the core biomarker. This core biomarker is used for precise detection of colorectal cancer liver metastases.

Citation Information

Patent Citations

  • Acquisition method of biomarker for assisting CRLM early diagnosis

    CN120564842A

  • A method for determining biomarker(s) pertaining to mental health disorder(s) and biomarker(s) determined therefrom

    WO2025012793A1