Multi-omics pathological analysis system and method for colorectal cancer liver metastasis risk prediction
By developing a multiomic pathological analysis system, integrating multi-source data and using deep learning models, the problem of insufficient accuracy and reliability of predicting liver metastasis risk in colorectal cancer in the prior art is solved, and accurate prediction of liver metastasis risk in colorectal cancer patients and support for the support of personalized treatment plans.
Patent Information
- Application Number
- CN202510208038.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-25
AI Technical Summary
When predicting the risk of liver metastasis in colorectal cancer, the prior art relies on limited clinical indicators and pathological characteristics, ignores the molecular heterogeneity of cancer, and is difficult to process and analyze large-scale and multi-type data, resulting in insufficient prediction accuracy, stability and reliability.
A multiomic pathological analysis system was developed, and by integrating GEO, TCGA databases and hospital-local multi-source data, using data preprocessing, cell subtype identification, CRLM scoring construction, deep learning model construction and training processes, it can achieve accurate prediction of liver metastasis risk in colorectal cancer patients.
It improves the accuracy and reliability of predicting liver metastasis risk in colorectal cancer, provides a scientific basis for the formulation of individualized treatment plans, and through the application of deep learning models, it achieves rapid and accurate identification of large-scale image data, reducing interference from human factors.
Smart Images

Figure CN120048529A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of biomedicine and cancer diagnosis, and specifically to a multi-omics pathological analysis system and method for predicting the risk of liver metastasis in colorectal cancer. Background Art
[0002] As one of the common malignant tumors globally, colorectal cancer has a high incidence and mortality rate. During the progression of colorectal cancer, liver metastasis is a crucial link, seriously affecting the prognosis and quality of life of patients. With the rapid development of bioinformatics and omics technologies, more and more researchers have started to focus on using multi-omics data for cancer research and prediction. Multi-omics data covers information at multiple levels including the genome, transcriptome, and proteome, providing unprecedented opportunities for comprehensively analyzing the mechanisms of cancer occurrence, development, and metastasis. Against this technical background, it is particularly important to develop a multi-omics pathological analysis system that can accurately predict the risk of liver metastasis in colorectal cancer.
[0003] Although traditional methods have achieved certain results in the diagnosis and treatment of colorectal cancer, there are still many deficiencies in predicting the risk of liver metastasis. Traditional prediction models often rely on limited clinical indicators and pathological features, ignoring the molecular heterogeneity of cancer. In addition, traditional analysis methods are overwhelmed when dealing with large-scale and multi-type data and are difficult to fully exploit the potential information in the data. Therefore, traditional methods still need to be improved in terms of prediction accuracy, stability, and reliability. A more comprehensive, accurate, and efficient prediction system is urgently needed to meet clinical needs and patient expectations.
[0004] Therefore, developing a multi-omics pathological analysis system and method for predicting the risk of liver metastasis in colorectal cancer will help improve the early diagnosis rate of liver metastasis in colorectal cancer and provide more personalized treatment plans and prognostic evaluations for patients. Summary of the Invention
[0005] The purpose of the present invention is to make up for the deficiencies of the existing technology and provide a multi-omics pathological analysis system and method for predicting the risk of liver metastasis in colorectal cancer. By integrating multi-source data from GEO, TCGA databases, and the hospital's local data, and using data preprocessing, cell subtype identification, CRLM score construction, deep learning model construction and training processes, the present invention realizes the accurate prediction of the risk of liver metastasis in colorectal cancer patients, not only improving the accuracy and reliability of the prediction, but also providing strong support for the formulation of individualized treatment plans.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: On the one hand, a multi-omics pathological analysis system for predicting the risk of colorectal cancer liver metastasis, the system includes a data acquisition module, a data preprocessing module, a cell analysis module, a score construction module, a model construction and training module, and a prediction and evaluation module;
[0007] The data acquisition module collects single-cell data sets, spatial transcriptome data, bulk RNA-seq data, clinical data, and whole-slide images WSIs of colorectal cancer patients from the GEO database, TCGA database, and the hospital local database, and performs format conversion and standardization processing on the data;
[0008] The data preprocessing module performs corresponding preprocessing on the collected different types of data using the SeuratR software package respectively;
[0009] Based on single-cell gene expression and chromosome sorting data, the cell analysis module calculates the CNV score of cells using the inferCNV software package, distinguishes malignant epithelial cells from non-malignant epithelial cells, and identifies LMTMCs, where LMTMCs represents the liver metastasis-triggering malignant cell subtype. The CytoTRACE software package and Monocle are used to evaluate its stemness and pseudotemporal trajectory analysis respectively, and CellChat is used to analyze the interaction between LMTMCs and other cell types;
[0010] The score construction module uses the OCLR algorithm in BulkRNA-seq data to calculate the relative abundance of LMTMCs in patients to construct a CRLM score system, where OCLR represents orthogonal categorical logistic regression and CRLM represents colorectal cancer liver metastasis. Analyze its correlation with clinicopathological features, and use Kaplan-Meier survival analysis to verify its predictive value for patient survival prognosis;
[0011] The model construction and training module selects the ResNet18 architecture to construct a deep learning model, inputs the preprocessed tumor patches, adjusts the model parameters through the Adam optimizer, pre-trains using the data set, evaluates the model performance using the AUC index, and fine-tunes for the CRLM task;
[0012] The prediction and evaluation module inputs the preprocessed WSIs data of the patient to be predicted into the trained model for patch-level classification prediction, and uses the majority voting algorithm to aggregate the patch-level classification results into WSIs-level prediction results.
[0013] Furthermore, in the data preprocessing module, the corresponding processing is carried out for different types of collected data using the SeuratR software package: for scRNA-seq data, the SeuratR software package is used to remove low-quality cells with more than 6,000 or less than 200 expressed genes and a mitochondrial read ratio exceeding 30%, perform normalization and standardization operations on the sample gene expression matrix, and use the harmonyR package to integrate cell data; for spatial transcriptome data, the SeuratR software package is used for normalization operations and cell type label mapping transfer; for WSIs data, it is cut into appropriate patch levels, and data augmentation and normalization are performed.
[0014] Furthermore, in the cell data analysis module, the identification of liver metastasis-triggering malignant cells LMTMCs is carried out. The inferCNV software package is used to calculate the CNV score of cells. Hierarchical clustering of all endothelial cells ECs is performed based on the CNV score through the clustering method of similarity field and density field. Endothelial cells ECs with a low CNV score and close to the reference cells are identified as non-malignant cells, while ECs with a high CNV score are identified as malignant cells. The malignant cells are clustered again, and the ratio of the actual cell number to the expected cell number of each tumor cell subtype in samples with different metastasis states of CRC is calculated. The formula is: Where the expected cell number is calculated based on the overall cell ratio of the sample. Ro / e > 1 for a cell subtype in the tissue indicates that this subtype prefers to be distributed in this tissue. The tumor cell subpopulation mainly distributed in the LMCRC tissue and the CRLM colorectal cancer liver metastasis tissue is identified as LMTMCs.
[0015] Furthermore, the cell analysis module performs hierarchical clustering of all endothelial cells ECs and fibroblasts based on the CNV score through the clustering method of similarity field and density field:
[0016] (1) For any two cells i and j, their CNV scores are CNV i and CNV j respectively. Define the similarity S ij between them as: Where α is the similarity decay coefficient, which is used to adjust the influence of the CNV score difference on the similarity;
[0017] (2) For cell i, its density ρ i is defined as the sum of the similarities of cells within a certain range around it: ρ i = ∑ j∈N(i) S ij , where N(i) represents the neighborhood of cell i, which is the absolute value of the difference in CNV scores |CNV i -CNV j |;
[0018] (3) Select the cell with the highest density as the initial clustering center C 1 , and for subsequent clustering centers, consider the isolation of cells. Define the isolation index I k of cell i from the selected clustering center C ik as: where m is the number of selected clustering centers, and each time select the cell with the largest I ik as the new clustering center C m+1 , until the predetermined number of clusters K is satisfied;
[0019] (4) For each cell i, calculate its similarity to all clustering centers C k and assign cell i to the cluster with the highest similarity:
[0020] (5) Let the CNV score of the reference cell be CNV ref , and for the ECs in each cluster, calculate their average similarity to the reference cell The formula is: where n is the number of ECs in the cluster, S i,ref is the similarity between cell i and the reference cell, calculated according to the formula in step (1). If is greater than the similarity threshold τ and the average CNV score of the cluster is lower, τ is the similarity threshold used to determine whether the cluster is malignant, then the ECs in the cluster are identified as non-malignant cells; otherwise, they are identified as malignant cells.
[0021] Further, the scoring construction module uses the OCLR algorithm to calculate the relative abundance of LMTMCs in patients. Let the gene expression matrix of the patient sample be X, where x ij represents the expression value of the i-th gene in the j-th sample. For each sample j, the OCLR algorithm constructs a classification model by learning the characteristic gene expression pattern of LMTMCs. Let the parameters of the model be θ, then the probability P(y j = 1|x j ,θ) of sample j belonging to the LMTMCs category is expressed by the logistic regression function as: where y j is the category label of sample j, y j = 1 indicates belonging to the LMTMCs category, y j = 0 indicates not belonging, x j is the gene expression vector of sample j, and θ Tis the transpose of the parameter vector θ. By training the OCLR model, the relative abundance scores of LMTMCs for each sample are obtained and defined as the CRLM score. The score represents the relative enrichment degree of LMTMCs in the patient samples. The higher the score, the more obvious the malignant cell characteristics related to liver metastasis in the sample.
[0022] Furthermore, the cell analysis module uses the Kaplan-Meier survival analysis method to analyze the impact of the CRLM score on the patient's survival prognosis. Taking the median of the CRLM score as the cut-off value, the patients are divided into a high-score group and a low-score group. Let t be the survival time, d be the death event, d = 1 indicates death, d = 0 indicates survival, and n be the number of patients at risk. At time t, the estimated value of the survival probability S(t) is calculated by the formula: where t i is the time point of the death event, d i is the number of deaths occurring at time t i and n i is the number of patients at risk at time t i By plotting the Kaplan-Meier survival curve, the survival situations of the patients in the high-score group and the low-score group are compared.
[0023] Furthermore, the model construction and training module selects the ResNet18 architecture to construct a deep learning model, which consists of an input layer, a convolutional layer, a pooling layer, a residual block, and a fully connected layer:
[0024] (1) Input layer: Receives the preprocessed tumor patch image with the size set as H×W×C, where H is the height, W is the width, and C is the number of channels;
[0025] (2) Convolutional layer: The convolutional layer is used for feature extraction and there are multiple of them. Each convolutional layer consists of a convolutional kernel, a bias, and an activation function. Let the size of the convolutional kernel be k×k, the stride be s, and the padding be p. For the input feature map X, the formula for the output feature map Y is: where W is the convolutional kernel weight, b is the bias, and i, j are the position indices of the output feature map;
[0026] (3) Pooling layer: Mainly performs sampling to reduce the size of the feature map. Let the window be a 2×2 window and the stride be 2. The calculation formula is: Y ij = max(X (2i,2j) , X (2i,2j+1) , X (2i+1,2j) , X (2i+1,2j+1) ), where X is the input feature map, Y is the output feature map, and i, j are the position indices of the output feature map;
[0027] (4) Residual block: It consists of two convolutional layers and a skip connection. Let the input be x, and the output after passing through two convolutional layers be F(x). The output y of the residual block is: y = F(x) + x. When the number of input and output channels is inconsistent, a 1×1 convolution is used to transform the number of channels of the input x to match the output dimension. The formula is: y = F(x) + W 1×1 x, where W 1×1 is the weight of the 1×1 convolution, and there are multiple residual blocks. The residual blocks in different stages are different in terms of the convolutional kernel size, stride, and number of channels setting;
[0028] (5) Fully connected layer: Located at the end of the network, it maps the extracted features to the final class space. Let the dimension of the output feature vector be d, the number of classes be C, the weight of the fully connected layer be W fc , and the bias be b fc , and the output z is: z = W fc x + b fc , in the prediction of the risk of colorectal cancer liver metastasis, C is set to 2, representing metastasis and non - metastasis. The z is converted into a probability distribution through the softmax function. The formula is: where p i is the probability that the sample belongs to the i - th class, and z i is the i - th value output by the fully connected layer.
[0029] Furthermore, the model construction and training module adjusts the model parameters through the Adam optimizer. Let t represent the current training iteration number, and θ t represent the model parameters W or b at the t - th iteration, and g t is the gradient of the parameter θ at the t - th iteration. Calculate the partial derivative of the parameter θ with respect to the loss function L: Calculate the first - order moment estimate m t and the second - order moment estimate v t , and the formula is: m t = β 1 m t-1 +(1 - β 1 )g t , where β 1 and β 2 are hyperparameters. Calculate the corrected first - order moment estimate and the second - order moment estimate The formula is: The update formula of the model parameters is: where α is the learning rate, and ∈ is a constant used to prevent the denominator from being zero.
[0030] Furthermore, the prediction and evaluation module uses the majority voting algorithm to aggregate the patch-level classification results into the WSI-level prediction results, and counts the number of patches predicted to be positive for liver metastasis, denoted as n positive and the number of patches predicted to be negative for liver metastasis, denoted as n negative , n positive > n negative , then the WSI is determined to be positive for liver metastasis; n positive < n negative , the WSI is determined to be negative for liver metastasis. Let the final WSI-level prediction result be y WSI , whose value is 0 or 1, representing negative or positive liver metastasis, and the formula is: When n positive = n negative , further examinations are carried out.
[0031] On the other hand, a multi-omics pathological analysis method for predicting the risk of colorectal liver metastasis, the analysis method includes the following steps:
[0032] S100, data acquisition: Download single-cell datasets, spatial transcriptome data, bulk RNA-seq data, and clinical data from the GEO database and TCGA database, and obtain WSI data of colorectal cancer patients from the hospital;
[0033] S200, data preprocessing: Perform quality control on single-cell data, remove low-quality cells, perform normalization, standardization, and batch effect correction, integrate and preprocess spatial transcriptome data, cut WSIs into patches, and perform data augmentation and normalization;
[0034] S300, cell subtype identification and analysis: Use the inferCNV software package to calculate the cell CNV score, distinguish malignant and non-malignant epithelial cells, perform clustering analysis on malignant cells to determine LMTMCs, use the CytoTRACE software package and Monocle to evaluate their stemness and perform pseudotime trajectory analysis respectively, and analyze the interactions between LMTMCs and other cell types at the single-cell and spatial transcriptome levels through CellChat;
[0035] S400, CRLM score construction and analysis: Use the OCLR algorithm to calculate the CRLM score of patients, representing the relative abundance of LMTMCs, analyze its correlation with clinicopathological features, and perform Kaplan-Meier survival analysis;
[0036] S500, Model Construction and Training: Build a deep learning model based on the ResNet18 architecture. Input the preprocessed tumor patches into the model, adjust the model parameters through the Adam optimizer, perform pre-training using the dataset, evaluate the model performance using the AUC metric, and fine-tune for the colorectal cancer liver metastasis task.
[0037] S600, Risk Prediction and Evaluation: Input the preprocessed WSIs data of the patient to be predicted into the trained model for patch-level prediction, obtain the WSIs-level prediction result through majority voting, and output the liver metastasis risk prediction value.
[0038] Compared with the prior art, the multi-omics pathological analysis system and method for predicting the risk of colorectal cancer liver metastasis have the following beneficial effects:
[0039] First, by integrating multi-omics data, including single-cell datasets, spatial transcriptome data, bulk RNA-seq data, clinical data, and whole-slide images WSIs, the present invention realizes a comprehensive and accurate prediction of the risk of colorectal cancer liver metastasis. This cross-dimensional data fusion and analysis not only improve the accuracy and reliability of the prediction but also provide a powerful tool for in-depth exploration of the molecular mechanisms of tumor occurrence, development, and metastasis. Especially in the cell analysis module, by identifying and analyzing the malignant cell subset LMTMCs of liver metastasis, cell subtypes and molecular characteristics closely related to liver metastasis can be revealed, providing a scientific basis for formulating personalized treatment plans.
[0040] Second, by selecting the ResNet18 architecture to build a deep learning model and combining the Adam optimizer for parameter adjustment, the present invention can effectively process large-scale and high-dimensional image data, quickly and accurately identify tumor patches, and aggregate the patch-level classification results into WSIs-level prediction results through the majority voting algorithm. This intelligent prediction method not only greatly improves work efficiency but also reduces the interference of human factors, providing strong support for the early detection and precise treatment of the risk of colorectal cancer liver metastasis.
[0041] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0043] Figure 1 It is a framework diagram of a multi-omics pathological analysis system for predicting the risk of liver metastasis in colorectal cancer;
[0044] Figure 2 It is a flowchart of a multi-omics pathological analysis method for predicting the risk of liver metastasis in colorectal cancer. Detailed implementation manners
[0045] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, detail the specific implementation manners, structures, features and their effects of the present invention as follows.
[0046] Embodiment 1:
[0047] Hospital clinical research scenario
[0048] Researchers collected the clinical data of 100 colorectal cancer patients in the past 5 years from the local hospital database, including the basic information of the patients (age, gender, medical history), pathological diagnosis results, treatment process and follow-up data. At the same time, the whole slide image (WSIs) data of these patients was obtained. In addition, relevant single-cell datasets, spatial transcriptome data and bulk RNA-seq data were downloaded from the GEO database and TCGA database to supplement the information required for the research.
[0049] For single-cell data, the Seurat R package was used to remove low-quality cells with more than 6000 or less than 200 expressed genes and a mitochondrial read ratio exceeding 30%. The sample gene expression matrix was normalized and standardized, and the harmony R package was used to integrate the cell data to eliminate batch effects.
[0050] For spatial transcriptome data, the Seurat R package was used for normalization operations and cell type label mapping transfer to make its data format and cell type annotation meet the requirements of subsequent analysis.
[0051] The WSIs data was cut into appropriate patch levels, for example, each patch was 256×256 pixels in size, and data augmentation (such as rotation and flipping operations) and normalization processing were performed to increase the diversity and comparability of the data.
[0052] Calculate the CNV scores of cells using the inferCNV software package, distinguish malignant and non-malignant epithelial cells, perform hierarchical clustering on endothelial cells and fibroblasts based on the CNV scores through clustering methods of similarity field and density field. For any two cells i and j, their CNV scores are CNV i and CNV j , define the similarity S ij between them as: For cell i, its density ρ i is defined as the sum of similarities of cells within a certain range around it: ρ i = ∑ j∈N(i) S ij . Select the cell with the maximum density as the initial clustering center C 1 . For subsequent clustering centers, consider the isolation of cells. Define the isolation index I k of cell i from the selected clustering center C ik as: Each time, select the cell with the maximum I ik as the new clustering center C m+k , until the predetermined number of clusters K is satisfied. For each cell i, calculate its similarity k to all clustering centers C and assign cell i to the cluster with the maximum similarity: Let the CNV score of the reference cell be CNV ref . For the ECs in each cluster, calculate their average similarity with the reference cell. The formula is: If is less than the similarity threshold τ and the average CNV score of this cluster is relatively low, then the ECs in this cluster are identified as malignant cells, and further cluster analysis of malignant cells is performed to determine liver metastasis-triggering malignant cells LMTMCs. Cluster the malignant cells again, calculate the ratio of the actual cell number to the expected cell number of each tumor cell subtype in samples with different metastatic states of CRC. The formula is: where the expected cell number is calculated based on the overall cell proportion of the sample. Ro / e > 1 for a cell subtype in the tissue indicates that this subtype prefers to be distributed in this tissue. The tumor cell subsets mainly distributed in the LMCRC tissue and the CRLM colorectal cancer liver metastasis tissue are identified as LMTMCs.
[0053] The stemness of LMTMCs was evaluated using the CytoTRACE software package, and it was found that some LMTMCs had a high stemness score, indicating stronger self-renewal and differentiation abilities, which might play a key role in the process of liver metastasis. Meanwhile, pseudotime trajectory analysis was performed using Monocle to observe the differentiation path and key time nodes of LMTMCs during tumor development.
[0054] By analyzing the interactions between LMTMCs and other cell types (such as immune cells, fibroblasts) at the single-cell and spatial transcriptome levels using CellChat, it was found that there was close signal communication between LMTMCs and certain immunosuppressive cells, which might inhibit the immune response and promote tumor liver metastasis by secreting specific cytokines.
[0055] The OCLR algorithm was used to calculate the relative abundance of LMTMCs in patients. Let the gene expression matrix of patient samples be X, where x ij represents the expression value of the i-th gene in the j-th sample. For each sample j, the OCLR algorithm constructs a classification model by learning the characteristic gene expression pattern of LMTMCs. Let the parameters of the model be θ, then the probability P(y j =1|x j ,θ) is expressed by the logistic regression function as: After analysis, it was found that the CRLM score was significantly correlated with the clinicopathological features of patients' tumor stage and lymph node metastasis. For example, patients with a later tumor stage and positive lymph node metastasis generally had a higher CRLM score.
[0056] Kaplan-Meier survival analysis was performed. Using the median of the CRLM score as the cut-off value, patients were divided into a high-score group and a low-score group. Let t be the survival time, d be the death event, d = 1 indicates death, d = 0 indicates survival, and n be the number of patients at risk. At time t, the estimated value of the survival probability S(t) was calculated using the formula: The survival curve was plotted and the log-rank test was used to verify the reliability of the CRLM score as an independent risk factor for survival prognosis.
[0057] The ResNet18 architecture was selected to construct a deep learning model. Its input layer received preprocessed tumor patch images (size 256×256×3). In the convolutional layer, let the convolutional kernel size be k×k, the stride be s, and the padding be p. For the input feature map X, the calculation formula for the output feature map Y is: Image features were extracted through multiple convolutional layers, and the pooling layer was used for sampling to reduce the size of the feature map. Taking a 2×2 window with a stride of 2 as an example, the calculation formula is: Y ij =max(X(2i,2j) , X (2i,2j+1) , X (2i+1,2j) , X (2i+1,2j+1) ), The residual block consists of two convolutional layers and a skip connection. Let the input be x, and the output after passing through the two convolutional layers be F(x). The output y of the residual block is: y = F(x) + x. When the number of input and output channels is inconsistent, a 1×1 convolution is used to transform the number of channels of the input x to match the output dimension, and the formula becomes: y = F(x) + W 1×1 x. The residual blocks in different stages are adjusted according to the network depth in terms of the convolutional kernel size, stride, and number of channels. The fully connected layer maps the extracted features to the final class space (two categories: with metastasis and without metastasis). Let the dimension of the output feature vector be d, the number of classes be C, and the weight of the fully connected layer be W fc , and the bias be b fc , and the output z is: z = W fc x + b fc , In the prediction of colorectal cancer liver metastasis risk, C is set to 2, representing with metastasis and without metastasis. The z is converted into a probability distribution through the softmax function, and the formula is:
[0058] The model parameters are adjusted by the Adam optimizer. Let t represent the current training iteration number, and θ t represent the model parameters W or b at the t-th iteration, and g t is the gradient of the parameter θ at the t-th iteration, and its calculation is based on the partial derivative of the loss function L with respect to the parameter θ: Calculate the first-order moment estimate m t and the second-order moment estimate v t , and the formula is: m t = β 1 m t-1 +(1 - β 1 )g t , where β 1 and β 2 are hyperparameters, calculate the corrected first-order moment estimate and the second-order moment estimate , and the formula is: The update formula of the model parameters is: Use the collected dataset for pre-training, and adopt the AUC index to evaluate the model performance. After multiple iterative trainings, the AUC value of the model gradually increases and reaches above 0.85. Then, fine-tune for the colorectal cancer liver metastasis task to further improve the prediction accuracy of the model.
[0059] After preprocessing the WSIs data of the patient to be predicted, input it into the trained model for patch-level prediction. For example, for the WSIs of a certain patient, the model predicts 100 patches after cutting them. Among them, the number of patches predicted as positive for liver metastasis is 60, and the number of patches predicted as negative for liver metastasis is 40. Through the majority voting algorithm, since 60 > 40, this WSIs is determined to be positive for liver metastasis, and the liver metastasis risk prediction value is output, providing a decision-making reference for clinicians and assisting in formulating personalized treatment plans.
[0060] In summary, in the hospital clinical research example, first integrate multi-channel data, covering the clinical and WSIs data of 100 patients in the hospital and relevant information in public databases. Improve the data quality through strict preprocessing, accurately analyze cell subtypes with the help of the inferCNV tool, clarify the characteristics of LMTMCs and the cell-cell interaction mechanism, lay a foundation for the research, construct a CRLM scoring system and verify its association with clinicopathology and prognosis, provide a basis for patient stratification, build a model based on ResNet18, achieve high performance through Adam optimization training, and finally use the model to predict the liver metastasis risk at the WSIs level of patients, assist doctors in formulating personalized treatment strategies, play a key role in clinical practice, and provide strong support for the diagnosis and treatment of colorectal liver metastasis.
[0061] Example Two:
[0062] Drug R & D Scenario
[0063] A pharmaceutical company collected a large amount of single-cell data sets, spatial transcriptome data, bulk RNA-seq data, and clinical data of colorectal cancer patients from public databases (GEO database and TCGA database) for the research and development of a new drug for colorectal liver metastasis. The patient information covers different races, ages, and disease stages, with a total of 500 sample data. At the same time, it obtained the WSIs data of some patients from a cooperative hospital to enrich the research data sources.
[0064] According to the method in the invention, strict quality control is carried out on the single-cell data, removing low-quality cells that do not meet the requirements, and performing normalization, standardization, and batch effect correction to ensure the accuracy and comparability of the data. Integrate and preprocess the spatial transcriptome data to enable effective combined analysis with other data types. Cut the WSIs data into appropriate patches and perform data augmentation and normalization operations to provide sufficient and high-quality data for subsequent model training.
[0065] Calculate the cell CNV score using the inferCNV software package to accurately distinguish malignant and non-malignant epithelial cells. Hierarchically cluster all endothelial cells (ECs) based on the CNV score by the clustering methods of similarity field and density field. ECs with a low CNV score and close to the reference cells are identified as non-malignant cells, while ECs with a high CNV score are identified as malignant cells. Cluster the malignant cells again and calculate the ratio of the actual cell number to the expected cell number of each tumor cell subtype in samples with different metastatic states of CRC. The formula is: Ro / e > 1 of the cell subtype in its tissue indicates that this subtype prefers to be distributed in this tissue. The tumor cell subset mainly distributed in the LMCRC tissue and the CRLM (colorectal cancer liver metastasis) tissue is identified as LMTMCs. Use the CytoTRACE software package and Monocle to evaluate its stemness and perform pseudotemporal trajectory analysis respectively. The study finds that under certain specific gene expression patterns, the stemness of LMTMCs is significantly enhanced, and its differentiation trajectory during tumor development is closely related to liver metastasis.
[0066] Deeply analyze the interaction between LMTMCs and other cell types through CellChat and find that some new intercellular signaling pathways are activated during liver metastasis. For example, signal transduction occurs between LMTMCs and tumor-associated macrophages through specific ligand-receptor pairs, promoting the migration and invasion of tumor cells. These findings provide potential targets for drug development.
[0067] Use the OCLR algorithm to calculate the relative abundance of LMTMCs in patients. Let the gene expression matrix of the patient sample be X, where x ij represents the expression value of the i-th gene in the j-th sample. For each sample j, the OCLR algorithm constructs a classification model by learning the characteristic gene expression pattern of LMTMCs. Let the parameter of the model be θ, then the probability P(y j = 1|x j ,θ) of sample j belonging to the LMTMCs category is represented by the logistic regression function as: Analyze its correlation with clinicopathological features. The results show that the CRLM score is significantly associated with the size, grade of the tumor, and the level of serum tumor markers of the patients. Use the Kaplan-Meier survival analysis method to analyze the impact of the CRLM score on the survival prognosis of patients. According to the median of the CRLM score as the cut-off value, divide the patients into a high-score group and a low-score group. Let t be the survival time, d be the death event, d = 1 indicates death, d = 0 indicates survival, and n be the number of patients at risk. At time t, calculate the estimated value of the survival probability S(t). The formula is: By plotting the Kaplan-Meier survival curve, the survival situations of patients in the high-score group and the low-score group were compared, providing important reference indicators for evaluating drug efficacy and patient prognosis during the drug R & D process.
[0068] A deep learning model was built based on the ResNet18 architecture, and the model structure was optimized according to the data characteristics. At the input layer, preprocessed tumor patch images (the size was adjusted to 512×512×3 according to the actual situation) were received. The parameter settings of the convolutional layer, pooling layer, and residual blocks were adjusted through multiple experiments to improve the model's ability to extract tumor features. The fully connected layer mapped the features to the final class space and converted them into a probability distribution through the softmax function.
[0069] The model parameters were adjusted by the Adam optimizer. Let t represent the current training iteration number, and θ t represent the model parameters W or b at the t-th iteration, and g t be the gradient of the parameter θ at the t-th iteration, and its calculation was based on the partial derivative of the loss function L with respect to the parameter θ: Calculate the first-order moment estimate m t and the second-order moment estimate v t , and the formula was: m t =β 1 m t-1 +(1 - β 1 )g t , Calculate the corrected first-order moment estimate and the second-order moment estimate , and the formula was: The update formula of the model parameters was: After being trained and optimized with a large amount of data sets, the AUC index of the model reached above 0.9, indicating that the model had high prediction performance. During the training process, the performance indicators of the model were continuously monitored, and the model was fine-tuned according to the results of the validation set to ensure the generalization ability and stability of the model.
[0070] The WSIs data of the animal model (simulating colorectal cancer liver metastasis) in preclinical drug trials were preprocessed and then input into the trained model for patch-level prediction. For example, in the prediction of the WSIs data of the animal model in a certain drug intervention group, the model analyzed its patches, and the WSIs-level prediction results were obtained through the majority voting algorithm. According to the prediction results, the inhibitory effect of the drug on colorectal cancer liver metastasis was evaluated, providing data support for the early screening and optimization of drug R & D, accelerating the drug R & D process, and improving the R & D success rate.
[0071] In summary, in the pharmaceutical R & D example, the pharmaceutical company widely collected 500 sample data, including various types of data with rich sources. After ensuring its usability through data preprocessing, it deeply explored the information at the cellular level, mined potential targets related to LMTMCs, and the CRLM score effectively associated clinical features with prognosis, guiding the pharmaceutical R & D process. It carefully constructed and trained a deep learning model, achieved a high AUC value through parameter optimization, used the model to predict the WSI data of animal models, and evaluated the effect of drugs in inhibiting liver metastasis. It is of great significance in the early screening and optimization links of pharmaceutical R & D, accelerating the R & D process and improving the success rate, opening up a new way to overcome the problem of colorectal liver metastasis.
[0072] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modifications, equivalent changes and modifications made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A multi-omics pathology analysis system for predicting the risk of colorectal cancer liver metastasis, characterized by: The system includes a data acquisition module, a data preprocessing module, a cell analysis module, a scoring construction module, a model construction and training module, and a prediction and evaluation module; The data acquisition module collects single-cell data sets, spatial transcriptome data, batch RNA-seq data, clinical data, and whole-slice image WSIs of colorectal cancer patients from the GEO database, TCGA database, and local hospital database, and performs format conversion and standardization on the data; The data preprocessing module uses the SeuratR software package to perform corresponding preprocessing on different types of collected data; The cell analysis module uses the inferCNV software package to calculate the CNV score of cells based on single-cell gene expression and chromosome sorting data, distinguishes malignant epithelial cells from non-malignant epithelial cells, and identifies LMTMCs, where LMTMCs represent a subtype of malignant cells that trigger liver metastasis, and uses the CytoTRACE software package and Monocle to evaluate their stemness and pseudo-time trajectory analysis, respectively, and uses CellChat to analyze the interaction between LMTMCs and other cell types; The scoring construction module uses the OCLR algorithm in BulkRNA-seq data to calculate the relative abundance of LMTMCs of patients to construct a CRLM scoring system, where OCLR represents orthogonal categorical logistic regression and CRLM represents colorectal cancer liver metastasis, analyzes its correlation with clinical pathological characteristics, and uses Kaplan-Meier survival analysis to verify its predictive value for patient survival prognosis; The model building and training module uses the ResNet18 architecture to build a deep learning model, inputs the preprocessed tumor patch, adjusts the model parameters through the Adam optimizer, uses the data set for pre-training, uses the AUC indicator to evaluate the model performance, and performs fine-tuning for the CRLM task; The prediction and evaluation module inputs the preprocessed WSIs data of the patient to be predicted into the trained model, performs patch-level classification prediction, and uses the majority voting algorithm to aggregate the patch-level classification results into WSIs-level prediction results.
2. The multi-omics pathology analysis system for predicting the risk of colorectal cancer liver metastasis according to claim 1, characterized in that: In the data preprocessing module, the SeuratR software package is used to perform corresponding processing on different types of collected data: for scRNA-seq data, the SeuratR software package is used to remove low-quality cells with more than 6,000 or less than 200 genes expressed and a mitochondrial read ratio of more than 30%, the sample gene expression matrix is normalized and standardized, and the harmonyR package is used to integrate cell data; For spatial transcriptome data, the SeuratR software package was used for normalization and cell type label mapping transfer; for WSIs data, it was cut into appropriate patch levels, and data enhancement and normalization were performed.
3. The multi-omics pathology analysis system for predicting the risk of colorectal cancer liver metastasis according to claim 1, characterized in that: In the cell data analysis module, the identification of liver metastasis-triggered malignant cells LMTMCs is performed, and the CNV score of the cells is calculated using the inferCNV software package. All endothelial cells ECs are hierarchically clustered based on the CNV score by clustering methods of similarity field and density field. Endothelial cells ECs with low CNV scores close to reference cells are identified as non-malignant cells, while ECs with high CNV scores are identified as malignant cells. The malignant cells are clustered again, and the ratio of the actual cell number to the expected cell number of each tumor cell subtype in CRC samples with different metastatic states is calculated. The formula is: The expected number of cells was calculated based on the overall cell ratio of the sample. The Ro / e>1 of the cell subtype in the tissue indicated that the subtype preferred to be distributed in the tissue. The tumor cell subpopulation mainly distributed in LMCRC tissue and CRLM colorectal cancer liver metastasis tissue was identified as LMTMCs.
4. The multi-omics pathology analysis system for predicting the risk of colorectal cancer liver metastasis according to claim 3, characterized in that: The cell analysis module hierarchically clusters all endothelial cells ECs and fibroblasts based on CNV scores using clustering methods of similarity fields and density fields: (1) For any two cells i and j, their CNV scores are CNV i and CNV j , define the similarity S between them ij for: Where α is the similarity decay coefficient; (2) For cell i, its density ρ i It is defined as the sum of similarities of cells within a certain range around it: ρ i =∑ j∈N(i) S ij , where N(i) represents the neighborhood of cell i and is the absolute value of the difference in CNV scores |CNV i -CNV j |; (3) Select the cell with the largest density as the initial cluster center C1. Consider the isolation of cells for subsequent cluster centers and define the distance from cell i to the selected cluster center C1. k Isolation index I ik for: Where m is the number of selected cluster centers, and each time I is selected ik The largest cell is used as the new cluster center C m+1 , until the predetermined number of clusters K is met; (4) For each cell i, calculate its distance to all cluster centers C k The similarity S iCk , assign cell i to the cluster with the greatest similarity: (5) Let the CNV score of the reference cell be CNV ref For each EC in each cluster, the average similarity between the ECs and the reference cells is calculated. The formula is: Where n is the number of ECs in the cluster, S i,ref is the similarity between cell i and the reference cell, calculated according to the formula in step (1). If The average CNV score of the cluster is greater than the similarity threshold τ If the mRNA expression level is lower, the ECs in this cluster are identified as non-malignant cells; otherwise, they are identified as malignant cells.
5. The multi-omics pathology analysis system for predicting the risk of colorectal cancer liver metastasis according to claim 1, characterized in that: The scoring construction module uses the OCLR algorithm to calculate the relative abundance of LMTMCs of patients. Let the gene expression matrix of the patient sample be X, where x ij represents the expression value of the ith gene in the jth sample. For each sample j, the OCLR algorithm constructs a classification model by learning the characteristic gene expression pattern of LMTMCs. Assuming the parameter of the model is θ, the probability P(y j =1|x j ,θ) is expressed by the logistic regression function as: where y j is the category label of sample j, y j =1 means it belongs to the LMTMCs category, y j =0 means it does not belong to, x j is the gene expression vector of sample j, θ T It is the transpose of the parameter vector θ. By training the OCLR model, the relative abundance score of LMTMCs for each sample is obtained.
6. The multi-omics pathology analysis system for predicting the risk of colorectal cancer liver metastasis according to claim 1, characterized in that: The cell analysis module uses the Kaplan-Meier survival analysis method to analyze the effect of CRLM score on the patient's survival prognosis. According to the median of CRLM score as the cutoff value, the patients are divided into a high score group and a low score group. Let t be the survival time, d be the death event, d=1 means death, d=0 means survival, n is the number of patients at risk, and at time t, the estimated value of the survival probability S(t) is calculated, and the formula is: where t i is the time point when the death occurred, d i is at time t i Number of deaths that occurred, n i is at time t i The number of patients at risk was calculated, and the survival of patients in the high-score group and the low-score group was compared by drawing the Kaplan-Meier survival curve.
7. The multi-omics pathology analysis system for predicting the risk of colorectal cancer liver metastasis according to claim 1, characterized in that: The model building and training module uses the ResNet18 architecture to build a deep learning model, which consists of an input layer, a convolutional layer, a pooling layer, a residual block, and a fully connected layer: (1) Input layer: receives the preprocessed tumor patch image, with the size set to H×W×C, where H is the height, W is the width, and C is the number of channels; (2) Convolutional layer: There are multiple convolutional layers used for feature extraction. Each convolutional layer consists of a convolution kernel, a bias, and an activation function. Assume that the convolution kernel size is k×k, the step size is s, and the padding is p. For the input feature map X, the calculation formula for the output feature map Y is: Where W is the convolution kernel weight, b is the bias, and i,j is the position index of the output feature map; (3) Pooling layer: mainly performs sampling to reduce the size of the feature map. The window is set to 2×2 and the step size is 2. The calculation formula is: Y ij =max(X (2i,2j) ,X (2i,2j+1) ,X (2i+1,2j) ,X (2i+1,2j+1) ), where X is the input feature map, Y is the output feature map, and i, j are the position indices of the output feature map; (4) Residual block: It consists of two convolutional layers and a skip connection. Let the input be x, and the output after two convolutional layers be F(x). The output y of the residual block is: y = F(x) + x. When the number of input and output channels is inconsistent, the number of channels of the input x is transformed through 1×1 convolution to match the output dimension. The formula is: y = F(x) + W 1×1 x, where W 1×1 It is the weight of 1×1 convolution, which contains multiple residual blocks. The residual blocks at different stages have different settings in convolution kernel size, step size and number of channels. (5) Fully connected layer: Located at the end of the network, it maps the extracted features to the final category space. Let the dimension of the output feature vector be d, the number of categories be C, and the weight of the fully connected layer be W. fc , with a bias of b fc , the output z is: z = W fc x+b fc ,In the risk prediction of colorectal cancer liver metastasis, C is set to 2, representing metastasis and no metastasis, and z is converted to probability distribution by the softmax function, and the formula is: where p i is the probability that the sample belongs to the i-th class, z i is the i-th value output by the fully connected layer.
8. The multi-omics pathology analysis system for predicting the risk of colorectal cancer liver metastasis according to claim 1, characterized in that: The model building and training module adjusts the model parameters through the Adam optimizer. Let t represent the current number of training iterations, θ t represents the model parameters W or b, g at the tth generation t is the gradient of the parameter θ at the tth iteration, and calculates the partial derivative of the loss function L with respect to the parameter θ: Calculate the first moment estimate m of the parameter t and the second-order moment estimate v t , the formula is: m t =β1m t-1 +(1-β1)g t , Where β1 and β2 are hyperparameters, calculate the corrected first-order moment estimate and the second moment estimate The formula is: The update formula of the model parameters is: where α is the learning rate and ∈ is a constant used to prevent the denominator from being zero.
9. The multi-omics pathology analysis system for predicting the risk of colorectal cancer liver metastasis according to claim 1, characterized in that: The prediction and evaluation module uses a majority voting algorithm to aggregate the classification results at the patch level into prediction results at the WSI level, and counts the number of patches predicted to be positive for liver metastasis among all patches. positive and the number of patches predicted to be negative for liver metastasis n negative , n positive >n negative , then the WSIs are judged as positive for liver metastasis; n positive <n negative , the WSIs are judged to be negative for liver metastasis, and the final WSIs-level prediction result is y WSI , whose value is 0 or 1, indicating negative or positive liver metastasis, and the formula is: When n appears positive =n negative , for further inspection.
10. A multi-omics pathological analysis method for predicting the risk of colorectal cancer liver metastasis, characterized in that: The analytical method comprises the following steps: S100, data acquisition: download single-cell datasets, spatial transcriptome data, bulk RNA-seq data, and clinical data from the GEO database and TCGA database, and obtain WSIs data of colorectal cancer patients from hospitals; S200, data preprocessing: quality control of single-cell data, removal of low-quality cells, normalization, standardization and batch effect correction, integration and preprocessing of spatial transcriptome data, cutting WSIs into patches and performing data enhancement and normalization; S300, cell subtype identification and analysis: the inferCNV software package was used to calculate the cell CNV score to distinguish between malignant and non-malignant epithelial cells, and the malignant cells were clustered to identify LMTMCs. The CytoTRACE software package and Monocle were used to assess their stemness and perform pseudo-time trajectory analysis, respectively. CellChat was used to analyze the interaction between LMTMCs and other cell types at the single-cell and spatial transcriptome levels. S400, CRLM score construction and analysis: The OCLR algorithm was used to calculate the patient's CRLM score, which represents the relative abundance of LMTMCs, and its correlation with clinical pathological characteristics was analyzed, and Kaplan-Meier survival analysis was performed; S500, model construction and training: build a deep learning model based on the ResNet18 architecture, input the preprocessed tumor patch into the model, adjust the model parameters through the Adam optimizer, use the dataset for pre-training, use the AUC metric to evaluate the model performance, and fine-tune it for the colorectal cancer liver metastasis task; S600, risk prediction and assessment: The WSIs data of the patient to be predicted are preprocessed and input into the trained model for patch-level prediction. The WSIs-level prediction results are obtained through majority voting, and the liver metastasis risk prediction value is output.
Citation Information
Patent Citations
Marker combination for predicting prognosis of liver cancer and application thereof
CN114107511A
Hepatocellular carcinoma prognosis biomarker and application thereof
CN115807089A
Hepatocellular carcinoma prognosis layering construction method integrating single cell and batch transcriptome pyroptosis characteristics
CN117497184A
Skin melanin nevus benign and malignant prediction method based on omics sequencing
CN118581219A
Gene marker related to breast cancer prognosis and detection kit
CN118745463A
Cited By
System for realizing precise prognosis prediction of gastric cancer patient based on deep learning
CN120511037A
A system for realizing precise prognosis prediction of gastric cancer patients based on deep learning
CN120511037B
Spatial omics-based intestinal cancer metastasis prediction method and device, medium and equipment
CN120913863A
Metastasis prediction method, device, medium and equipment for colorectal cancer based on spatial omics
CN120913863B
Construction method of cell-level pathological feature and gene mutation data matching model, cell-level pathological feature and gene mutation data matching model and application thereof
CN121438941A