Multisomic pathological analysis system and method for predicting risk of colorectal cancer liver metastasis
By integrating multi-omics data and deep learning models, the shortcomings of traditional methods in predicting the risk of colorectal cancer liver metastasis have been addressed, enabling accurate risk assessment and personalized treatment support.
Patent Information
- Application Number
- CN202510208038.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-02-25
AI Technical Summary
Existing technologies rely on limited clinical indicators in predicting the risk of colorectal cancer liver metastasis, neglecting molecular-level heterogeneity and struggling to handle large-scale, multi-type data, resulting in insufficient prediction accuracy and reliability.
By integrating multi-source data from GEO, TCGA, and local hospitals, and utilizing data preprocessing, cell subtype identification, CRLM scoring construction, and deep learning models, a multi-omics pathological analysis system was constructed using software packages such as SeuratR, inferCNV, CytoTRACE, Monocle, CellChat, and ResNet18 to achieve accurate prediction of the risk of colorectal cancer liver metastasis.
It improves the accuracy and reliability of predicting the risk of liver metastasis in colorectal cancer, provides a scientific basis for personalized treatment plans, reduces interference from human factors, and improves work efficiency.
Smart Images

Figure CN120048529B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of biomedicine and cancer diagnostic technology, specifically to a multi-omics pathological analysis system and method for predicting the risk of colorectal cancer liver metastasis. Background Technology
[0002] Colorectal cancer, as one of the most common malignant tumors worldwide, has a persistently high incidence and mortality rate. Liver metastasis is a crucial step in the progression of colorectal cancer, severely impacting patients' prognosis and quality of life. With the rapid development of bioinformatics and omics technologies, more and more researchers are focusing on using multi-omics data for cancer research and prediction. Multi-omics data encompasses information at multiple levels, including the genome, transcriptome, and proteome, providing unprecedented opportunities for a comprehensive understanding of the mechanisms of cancer occurrence, development, and metastasis. Against this technological backdrop, developing a multi-omics pathological analysis system capable of accurately predicting the risk of colorectal cancer liver metastasis is particularly important.
[0003] While traditional methods have achieved some success in the diagnosis and treatment of colorectal cancer, they still have many shortcomings in predicting the risk of liver metastasis. Traditional predictive models often rely on limited clinical indicators and pathological features, ignoring the heterogeneity of cancer at the molecular level. In addition, traditional analytical methods are inadequate when dealing with large-scale, multi-type data, making it difficult to fully extract the potential information in the data. Therefore, traditional methods still need to be improved in terms of predictive accuracy, stability, and reliability. A more comprehensive, accurate, and efficient predictive system urgently needs to be developed to meet clinical needs and patient expectations.
[0004] Therefore, developing a multi-omics pathological analysis system and method for predicting the risk of colorectal cancer liver metastasis will help improve the early diagnosis rate of colorectal cancer liver metastasis and provide patients with more personalized treatment plans and prognostic assessments. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a multi-omics pathological analysis system and method for predicting the risk of liver metastasis in colorectal cancer. By integrating multi-source data from GEO, TCGA databases and local hospital data, and utilizing data preprocessing, cell subtype identification, CRLM score construction, and deep learning model construction and training processes, this invention achieves accurate prediction of the risk of liver metastasis in colorectal cancer patients. This not only improves the accuracy and reliability of the prediction but also provides strong support for the development of individualized treatment plans.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: On the one hand, a multi-omics pathological analysis system for predicting the risk of colorectal cancer liver metastasis, the system including a data acquisition module, a data preprocessing module, a cell analysis module, a scoring construction module, a model construction and training module, and a prediction and evaluation module;
[0007] The data acquisition module collects single-cell datasets, spatial transcriptome data, batch RNA-seq data, clinical data, and whole-slice images (WSIs) from colorectal cancer patients from the GEO database, TCGA database, and local hospital database, and performs format conversion and standardization processing on the data.
[0008] The data preprocessing module uses the SeuratR software package to perform corresponding preprocessing for different types of collected data.
[0009] The cell analysis module uses single-cell gene expression and chromosome sequencing data to calculate the CNV score of cells using the inferCNV software package, distinguishes between malignant and non-malignant epithelial cells, and identifies LMTMCs, which represent liver metastasis triggering malignant cell subtypes. The stemness and pseudo-temporal trajectory analysis of LMTMCs are evaluated using the CytoTRACE software package and Monocle, respectively, and the interaction between LMTMCs and other cell types is analyzed using CellChat.
[0010] The scoring construction module uses the OCLR algorithm to calculate the relative abundance of LMTMCs in patients in BulkRNA-seq data to construct the CRLM scoring system, where OCLR represents orthogonal categorical logistic regression and CRLM represents colorectal cancer liver metastasis. The correlation between CRLM and clinicopathological features is analyzed, and the predictive value of CRLM for patient survival prognosis is verified by Kaplan-Meier survival analysis.
[0011] The model building and training module uses the ResNet18 architecture to build a deep learning model. The preprocessed tumor patch is input, the model parameters are adjusted by the Adam optimizer, the model is pre-trained using the dataset, the AUC metric is used to evaluate the model performance, and fine-tuning is performed for the CRLM task.
[0012] The prediction and evaluation module inputs the preprocessed WSIs data of the patients to be predicted into the trained model, performs patch-level classification prediction, and uses a majority voting algorithm to aggregate the patch-level classification results into WSIs-level prediction results.
[0013] Furthermore, the data preprocessing module utilizes the SeuratR software package to process different types of collected data accordingly: for scRNA-seq data, the SeuratR software package removes low-quality cells with more than 6000 or fewer than 200 expressed genes and mitochondrial read ratios exceeding 30%, normalizes and standardizes the sample gene expression matrix, and integrates cell data using the harmonyR software package; for spatial transcriptome data, the SeuratR software package performs normalization and cell type label mapping transfer; for WSIs data, it is segmented into appropriate patch levels and subjected to data augmentation and normalization.
[0014] Furthermore, in the cell data analysis module, the identification of liver metastasis-triggered malignant cells (LMTMCs) is performed. The inferCNV software package is used to calculate the CNV score of the cells. Based on the CNV score, all endothelial cells (ECs) are hierarchically clustered using similarity field and density field clustering methods. Endothelial cell ECs with low CNV scores close to the reference cells are identified as non-malignant cells, while ECs with high CNV scores are identified as malignant cells. The malignant cells are then clustered again, and the ratio of the actual number of each tumor cell subtype to the expected number of cells in CRC samples at different metastatic states is calculated using the following formula: The expected cell count is calculated based on the overall cell proportion of the sample. A Ro / e > 1 for a cell subtype in the tissue indicates that the subtype is preferred to be distributed in the tissue. The tumor cell subpopulation mainly distributed in LMCRC tissue and CRLM colorectal cancer liver metastasis tissue was identified as LMTMCs.
[0015] Furthermore, the cell analysis module performs hierarchical clustering of all endothelial cells (ECs) and fibroblasts based on CNV scores using clustering methods based on similarity fields and density fields:
[0016] (1) For any two cells i and j, their CNV scores are respectively CNV i and CNV j Define the similarity S between them. ij for: Where α is the similarity decay coefficient, used to adjust the impact of CNV score differences on similarity;
[0017] (2) For cell i, its density ρ i Defined as the sum of the similarities of cells within a certain radius around it: ρ i =∑ j∈N(i) S ij Where N(i) represents the neighborhood of cell i, and |CNV| is the absolute value of the difference in CNV scores. i -CNV j |;
[0018] (3) Select the cell with the highest density as the initial cluster center C1. For subsequent cluster centers, consider the isolation of cells and define the distance from cell i to the selected cluster center C1. k Isolation index I ik for: Where m is the number of selected cluster centers, and each time I is selected... ik The largest cell serves as the new cluster center C m+1 This continues until the predetermined number of clusters K is met;
[0019] (4) For each cell i, calculate its distance to all cluster centers C. k similarity Assign cell i to the cluster with the highest similarity:
[0020] (5) Let the CNV score of the reference cell be CNV. ref For each cluster of ECs, calculate their average similarity to the reference cell. The formula is: Where n is the number of ECs in the cluster, S i,ref It is the similarity between cell i and the reference cell, calculated according to the formula in step (1). If The average CNV score of the cluster is greater than the similarity threshold τ. If the value is low, τ is the similarity threshold used to determine whether a cluster is malignant; if it is, ECs in that cluster are identified as non-malignant cells, otherwise they are identified as malignant cells.
[0021] Furthermore, the scoring construction module uses the OCLR algorithm to calculate the relative abundance of LMTMCs in patients. Let X be the gene expression matrix of the patient sample, where x ij Let represent the expression value of the i-th gene in the j-th sample. For each sample j, the OCLR algorithm constructs a classification model by learning the characteristic gene expression patterns of LMTMCs. Let the parameters of the model be θ, then the probability P(y) of sample j belonging to the LMTMCs category is... j =1|x j The logistic regression function (θ) is used to express the following: Where y j y is the class label of sample j. j =1 indicates that it belongs to the LMTMCs category, y j =0 indicates that it does not belong to x j Let θ be the gene expression vector of sample j. TIt is the transpose of the parameter vector θ. By training the OCLR model, the relative abundance score of LMTMCs for each sample is obtained and defined as the CRLM score. The score represents the relative enrichment of LMTMCs in the patient sample. The higher the score, the more obvious the malignant cell features related to liver metastasis in the sample.
[0022] Furthermore, the cell analysis module uses the Kaplan-Meier survival analysis method to analyze the impact of CRLM scores on patient survival prognosis. Based on the median CRLM score as the cutoff value, patients are divided into high-score and low-score groups. Let t be the survival time, d be the death event (d=1 represents death, d=0 represents survival), and n be the number of patients at risk. At time t, the estimated survival probability S(t) is calculated using the following formula: Where t i It is the time point in time when the death occurred, d i At time t i The number of deaths that occurred, n i At time t i The number of patients at risk was determined by plotting Kaplan-Meier survival curves to compare the survival of patients in the high-scoring and low-scoring groups.
[0023] Furthermore, the model building and training module uses the ResNet18 architecture to build a deep learning model, which consists of an input layer, convolutional layers, pooling layers, residual blocks, and fully connected layers.
[0024] (1) Input layer: Receives the preprocessed tumor patch image, with a size of H×W×C, where H is the height, W is the width, and C is the number of channels;
[0025] (2) Convolutional Layers: Convolutional layers are used for feature extraction. There are multiple convolutional layers, each consisting of a convolutional kernel, a bias, and an activation function. Let the kernel size be k×k, the stride be s, and the padding be p. For the input feature map X, the formula for calculating the output feature map Y is: Where W is the convolution kernel weight, b is the bias, and i,j are the position indices of the output feature map;
[0026] (3) Pooling layer: mainly for sampling, reducing the size of the feature map. Assuming a 2×2 window and a stride of 2, the calculation formula is: Y... ij =max(X (2i,2j) ,X (2i,2j+1) ,X (2i+1,2j) ,X (2i+1,2j+1) ), where X is the input feature map, Y is the output feature map, and i,j are the position indices of the output feature map;
[0027] (4) Residual Block: Consists of two convolutional layers and a skip connection. Let the input be x, and the output after the two convolutional layers be F(x). The output y of the residual block is: y = F(x) + x. When the number of input and output channels is inconsistent, the number of channels of the input x is transformed by a 1×1 convolution to match the output dimension. The formula is: y = F(x) + W 1×1 x, where W 1×1 These are the weights of a 1×1 convolution, which contain multiple residual blocks. The residual blocks at different stages have different settings in terms of kernel size, stride, and number of channels.
[0028] (5) Fully connected layer: Located at the end of the network, it maps the extracted features to the final class space. Let the dimension of the output feature vector be d, the number of classes be C, and the weight of the fully connected layer be W. fc The bias is b fc The output z is: z = W fc x+b fc In the prediction of colorectal cancer liver metastasis risk, C is set to 2, representing metastasis and no metastasis. The z-value is converted into a probability distribution using the softmax function, as shown in the formula: Where p i z is the probability that a sample belongs to the i-th class. i It is the i-th value output by the fully connected layer.
[0029] Furthermore, the model building and training module adjusts the model parameters using the Adam optimizer, where t represents the current training iteration number, and θ... t The model parameters W or b, g represent the parameters at the t-th generation. t It is the gradient of parameter θ at the t-th iteration, calculated using the partial derivative of the loss function L with respect to parameter θ: First-moment estimate of the parameter m t and second-order moment estimate v t The formula is: m t =β1m t-1 +(1-β1)g t , Where β1 and β2 are hyperparameters, the corrected first-order moment estimate is calculated. and second-order moment estimation The formula is: The formula for updating the model parameters is: Where α is the learning rate, and ∈ is used to prevent the denominator from being a zero constant.
[0030] Furthermore, the prediction and evaluation module uses a majority voting algorithm to aggregate the patch-level classification results into WSIs-level prediction results, and counts the number n of patches predicted as positive for liver metastasis among all patches. positiveThe number of patches n predicted to be negative for liver metastasis. negative n positive >n negative If so, the WSIs are considered positive for liver metastasis; n positive <n negative The WSIs were determined to be negative for liver metastasis. Let the final WSIs grade prediction result be y. WSI Its value is 0 or 1, indicating negative or positive liver metastasis, and the formula is: When n appears positive =n negative Further investigation will be conducted.
[0031] On the other hand, a multi-omics pathological analysis method for predicting the risk of colorectal cancer liver metastasis includes the following steps:
[0032] S100, Data Acquisition: Download single-cell datasets, spatial transcriptome data, batch RNA-seq data, and clinical data from the GEO and TCGA databases, and obtain WSIs data from colorectal cancer patients from hospitals;
[0033] S200, Data Preprocessing: Quality control of single-cell data, removal of low-quality cells, normalization, standardization and batch effect correction, integration and preprocessing of spatial transcriptome data, cutting WSIs into patches and performing data augmentation and normalization.
[0034] S300, Cell Subtype Identification and Analysis: The CNV score of cells was calculated using the inferCNV software package to distinguish between malignant and non-malignant epithelial cells. Cluster analysis of malignant cells was used to identify LMTMCs. CytoTRACE software package and Monocle were used to evaluate their stemness and perform pseudo-temporal trajectory analysis. CellChat was used to analyze the interaction between LMTMCs and other cell types at the single-cell and spatial transcriptome levels.
[0035] S400, CRLM Score Construction and Analysis: The OCLR algorithm was used to calculate the CRLM score of patients, which represents the relative abundance of LMTMCs. The correlation between the score and clinicopathological features was analyzed, and Kaplan-Meier survival analysis was performed.
[0036] S500, Model Building and Training: A deep learning model is built based on the ResNet18 architecture. The preprocessed tumor patch is input into the model, the model parameters are adjusted by the Adam optimizer, the model is pre-trained using the dataset, the AUC index is used to evaluate the model performance, and fine-tuning is performed for the colorectal cancer liver metastasis task.
[0037] S600, Risk Prediction and Assessment: The WSIs data of the patients to be predicted are preprocessed and input into the trained model for patch-level prediction. The WSIs-level prediction results are obtained through majority voting, and the predicted value of liver metastasis risk is output.
[0038] Compared with existing technologies, this multi-omics pathological analysis system and method for predicting the risk of colorectal cancer liver metastasis has the following advantages:
[0039] I. This invention integrates multi-omics data, including single-cell datasets, spatial transcriptome data, batch RNA-seq data, clinical data, and whole-slice images (WSIs), to achieve a comprehensive and accurate prediction of the risk of colorectal cancer liver metastasis. This cross-dimensional data fusion and analysis not only improves the accuracy and reliability of the prediction but also provides a powerful tool for in-depth exploration of the molecular mechanisms of tumor occurrence, development, and metastasis. In particular, in the cell analysis module, the identification and analysis of the malignant cell subset LMTMCs in liver metastasis can reveal cell subtypes and molecular characteristics closely related to liver metastasis, providing a scientific basis for developing personalized treatment plans.
[0040] Second, this invention utilizes the ResNet18 architecture to construct a deep learning model and combines it with the Adam optimizer for parameter tuning. This enables the effective processing of large-scale, high-dimensional image data, rapid and accurate identification of tumor patches, and aggregation of patch-level classification results into WSIs-level prediction results through a majority voting algorithm. This intelligent prediction method not only greatly improves work efficiency but also reduces the interference of human factors, providing strong support for the early detection and precision treatment of colorectal cancer liver metastasis risk.
[0041] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0043] Figure 1 A framework diagram of a multi-omics pathological analysis system for predicting the risk of liver metastasis in colorectal cancer;
[0044] Figure 2A flowchart of a multi-omics pathological analysis method for predicting the risk of liver metastasis in colorectal cancer. Detailed Implementation
[0045] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0046] Example 1:
[0047] Hospital clinical research scenario
[0048] Researchers collected clinical data from 100 colorectal cancer patients over the past five years from the hospital's local database, including basic patient information (age, gender, medical history), pathological diagnosis results, treatment process, and follow-up data. They also obtained whole-slice images (WSIs) of these patients. In addition, they downloaded relevant single-cell datasets, spatial transcriptome data, and batch RNA-seq data from the GEO and TCGA databases to supplement the information needed for the study.
[0049] For single-cell data, the SeuratR package was used to remove low-quality cells with more than 6,000 or fewer than 200 expressed genes and a mitochondrial readout rate of more than 30%. The gene expression matrix of the samples was normalized and standardized, and the cellular data was integrated using the harmonyR package to eliminate batch effects.
[0050] The spatial transcriptome data were normalized and cell type label mapping was performed using the SeuratR software package to ensure that the data format and cell type annotation met the requirements for subsequent analysis.
[0051] WSIs data are segmented into appropriate patch levels, such as each patch being 256×256 pixels in size, and data augmentation (such as rotation and flipping operations) and normalization are performed to increase data diversity and comparability.
[0052] The CNV score of cells was calculated using the inferCNV software package to distinguish between malignant and non-malignant epithelial cells. Hierarchical clustering of endothelial cells and fibroblasts was performed based on the CNV score using similarity field and density field clustering methods. For any two cells i and j, their CNV scores are respectively... i and CNV j Define the similarity S between them. ij for: For cell i, its density ρ i Defined as the sum of the similarities of cells within a certain radius around it: ρ i =∑ j∈N(i) Sij The cell with the highest density is selected as the initial cluster center C1. For subsequent cluster centers, considering the isolation of cells, the distance from cell i to the selected cluster center C is defined. k Isolation index I ik for: Each time I select I ik The largest cell serves as the new cluster center C m+k Until the predetermined number of clusters K is met, for each cell i, calculate its distance to all cluster centers C. k similarity Assign cell i to the cluster with the highest similarity: Let the CNV score of the reference cell be CNV. ref For each cluster of ECs, calculate their average similarity to the reference cell. The formula is: if The average CNV score of the cluster is less than the similarity threshold τ. If the value is low, then ECs in this cluster are identified as malignant cells. Further cluster analysis of malignant cells identifies liver metastasis triggering malignant cells (LMTMCs). The malignant cells are then clustered again, and the ratio of the actual number of each tumor cell subtype to the expected number of cells in CRC samples at different metastatic states is calculated using the following formula: The expected cell count is calculated based on the overall cell proportion of the sample. A Ro / e > 1 for a cell subtype in the tissue indicates that the subtype is preferred to be distributed in the tissue. The tumor cell subpopulation mainly distributed in LMCRC tissue and CRLM colorectal cancer liver metastasis tissue was identified as LMTMCs.
[0053] The stemness of LMTMCs was assessed using the CytoTRACE software package. Some LMTMCs had high stemness scores, indicating that they have stronger self-renewal and differentiation capabilities and may play a key role in liver metastasis. Meanwhile, pseudo-time trajectory analysis was performed using Monocle to observe the differentiation pathways and key time points of LMTMCs in the process of tumor development.
[0054] CellChat analysis of the interactions between LMTMCs and other cell types (such as immune cells and fibroblasts) at the single-cell and spatial transcriptome levels revealed close signaling communication between LMTMCs and certain immunosuppressive cells, which may suppress immune responses and promote tumor liver metastasis by secreting specific cytokines.
[0055] The OCLR algorithm was used to calculate the relative abundance of LMTMCs in patients. Let X be the gene expression matrix of the patient sample, where x ijLet represent the expression value of the i-th gene in the j-th sample. For each sample j, the OCLR algorithm constructs a classification model by learning the characteristic gene expression patterns of LMTMCs. Let the parameters of the model be θ, then the probability P(y) of sample j belonging to the LMTMCs category is... j =1|x j The logistic regression function (θ) is used to express the following: Analysis revealed a significant correlation between CRLM scores and patients' tumor stage, lymph node metastasis, and clinicopathological features. For example, patients with later-stage tumors and positive lymph node metastasis generally had higher CRLM scores.
[0056] Kaplan-Meier survival analysis was performed, using the median CRLM score as the cutoff value to divide patients into high-score and low-score groups. Let t be the survival time, d be the death event (d=1 for death, d=0 for survival), and n be the number of patients at risk. At time t, the estimated survival probability S(t) was calculated using the following formula: Survival curves were plotted and the log-rank test was used to verify the reliability of the CRLM score as an independent risk factor for survival prognosis.
[0057] A deep learning model is constructed using the ResNet18 architecture. Its input layer receives a preprocessed tumor patch image (256×256×3). In the convolutional layers, the kernel size is k×k, the stride is s, and the padding is p. For the input feature map X, the formula for calculating the output feature map Y is: Image features are extracted through multiple convolutional layers, and pooling layers are used for sampling to reduce the size of the feature map. Taking a 2×2 window and a stride of 2 as an example, the calculation formula is: Y ij =max(X (2i,2j) ,X (2i,2j+1) ,X (2i+1,2j) ,X (2i+1,2j+1) The residual block consists of two convolutional layers and a skip connection. Let the input be x, and the output after the two convolutional layers be F(x). The output y of the residual block is: y = F(x) + x. When the number of input and output channels is inconsistent, a 1×1 convolution is used to transform the number of channels of the input x to match the output dimension. The formula then becomes: y = F(x) + W. 1×1 x, the residual blocks at different stages have their kernel size, stride, and number of channels adjusted according to the network depth. The fully connected layer maps the extracted features to the final class space (with and without transfer). Let the output feature vector dimension be d, the number of classes be C, and the weights of the fully connected layer be W. fc The bias is b fc The output z is: z = W fc x+b fcIn the prediction of colorectal cancer liver metastasis risk, C is set to 2, representing metastasis and no metastasis. The z-value is converted into a probability distribution using the softmax function, as shown in the formula:
[0058] The model parameters are adjusted using the Adam optimizer. Let t represent the current number of training iterations, and θ... t The model parameters W or b, g represent the parameters at the t-th generation. t It is the gradient of parameter θ at the t-th iteration, which is calculated based on the partial derivative of the loss function L with respect to parameter θ: First-moment estimate of the parameter m t and second-order moment estimate v t The formula is: m t =β1m t-1 +(1-β1)g t , Where β1 and β2 are hyperparameters, the corrected first-order moment estimate is calculated. and second-order moment estimation The formula is: The formula for updating the model parameters is: The model was pre-trained using the collected dataset and its performance was evaluated using the AUC metric. After multiple iterations of training, the model's AUC value gradually increased to over 0.85. Then, it was fine-tuned for the task of colorectal cancer liver metastasis to further improve the model's prediction accuracy.
[0059] The WSIs data of the patients to be predicted are preprocessed and then input into the trained model for patch-level prediction. For example, for a certain patient's WSIs, the model predicts 100 patches after segmentation. Among them, 60 patches are predicted to be positive for liver metastasis, and 40 patches are predicted to be negative for liver metastasis. Through the majority voting algorithm, since 60 > 40, the WSIs are judged to be positive for liver metastasis, and the predicted value of liver metastasis risk is output to provide decision-making reference for clinicians and assist in the formulation of personalized treatment plans.
[0060] In summary, in the hospital clinical research implementation, data from multiple channels was first integrated, covering clinical and WSIs data of 100 patients within the hospital and relevant information from public databases. Data quality was improved through rigorous preprocessing. The inferCNV tool was used to accurately analyze cell subtypes, clarify the characteristics of LMTMCs and the intercellular interaction mechanism, laying the foundation for the research. A CRLM scoring system was constructed and its association with clinicopathology and prognosis was verified, providing a basis for patient stratification. A model based on ResNet18 was built and optimized with Adam to achieve high performance. Finally, the model was used to predict the risk of WSIs-level liver metastasis in patients, assisting physicians in developing personalized treatment strategies and playing a crucial role in clinical practice, providing strong support for the diagnosis and treatment of colorectal cancer liver metastasis.
[0061] Example 2:
[0062] Drug development scenario
[0063] To develop a novel drug targeting liver metastases from colorectal cancer, a pharmaceutical company collected a large amount of single-cell datasets, spatial transcriptome data, batch RNA-seq data, and clinical data from publicly available databases (GEO and TCGA databases). These data covered patients of different races, ages, and disease stages, totaling 500 samples. In addition, the company obtained some patients' WSIs data from partner hospitals to enrich its research data sources.
[0064] According to the method in the invention, strict quality control is performed on single-cell data to remove low-quality cells that do not meet the requirements, and normalization, standardization and batch effect correction are performed to ensure the accuracy and comparability of the data. Spatial transcriptome data is integrated and preprocessed so that it can be effectively combined with other data types for analysis. WSIs data are cut into appropriate patches and data augmentation and normalization operations are performed to provide sufficient and high-quality data for subsequent model training.
[0065] The inferCNV software package was used to calculate cell CNV scores to accurately distinguish between malignant and non-malignant epithelial cells. Based on CNV scores, hierarchical clustering of all endothelial cells (ECs) was performed using similarity field and density field clustering methods. ECs with low CNV scores close to the reference cells were identified as non-malignant cells, while ECs with high CNV scores were identified as malignant cells. Malignant cells were then clustered again, and the ratio of the actual to the expected number of each tumor cell subtype in CRC metastatic samples was calculated using the following formula: The Ro / e>1 ratio of the cell subtype in the tissue indicates that this subtype is preferred to be distributed in this tissue. The tumor cell subpopulation mainly distributed in LMCRC and CRLM colorectal cancer liver metastasis tissues was identified as LMTMCs. The stemness and pseudo-temporal trajectory analysis were evaluated using CytoTRACE software and Monocle, respectively. The study found that the stemness of LMTMCs was significantly enhanced under certain gene expression patterns, and their differentiation trajectory in the process of tumor development was closely related to liver metastasis.
[0066] In-depth analysis of the interactions between LMTMCs and other cell types using CellChat revealed the activation of novel intercellular signaling pathways during liver metastasis. For example, LMTMCs and tumor-associated macrophages communicate via specific ligand-receptor pairs, promoting tumor cell migration and invasion. These findings provide potential targets for drug development.
[0067] The OCLR algorithm was used to calculate the relative abundance of LMTMCs in patients. Let X be the gene expression matrix of the patient sample, where x ij Let represent the expression value of the i-th gene in the j-th sample. For each sample j, the OCLR algorithm constructs a classification model by learning the characteristic gene expression patterns of LMTMCs. Let the parameters of the model be θ, then the probability P(y) of sample j belonging to the LMTMCs category is... j =1|x j The logistic regression function (θ) is used to express the following: The correlation between CRLM score and clinicopathological features was analyzed. Results showed a significant association between CRLM score and tumor size, grade, and serum tumor marker levels. The Kaplan-Meier survival analysis method was used to analyze the impact of CRLM score on patient survival prognosis. Patients were divided into high-score and low-score groups based on the median CRLM score as the cutoff value. Let t be survival time, d be the death event (d=1 for death, d=0 for survival), and n be the number of patients at risk. At time t, the estimated survival probability S(t) was calculated using the following formula: By plotting Kaplan-Meier survival curves and comparing the survival of patients in the high-scoring and low-scoring groups, an important reference indicator is provided for evaluating drug efficacy and patient prognosis during the drug development process.
[0068] A deep learning model was built based on the ResNet18 architecture. The model structure was optimized according to the characteristics of the data. In the input layer, the preprocessed tumor patch image was received (the size was adjusted to 512×512×3 according to the actual situation). The parameter settings of the convolutional layer, pooling layer and residual block were tested and adjusted many times to improve the model's ability to extract tumor features. The fully connected layer maps the features to the final class space and converts them into a probability distribution through the softmax function.
[0069] The model parameters are adjusted using the Adam optimizer. Let t represent the current number of training iterations, and θ... t The model parameters W or b, g represent the parameters at the t-th generation. t It is the gradient of parameter θ at the t-th iteration, which is calculated based on the partial derivative of the loss function L with respect to parameter θ: First-moment estimate of the parameter m t and second-order moment estimate v t The formula is: m t =β1m t-1 +(1-β1)g t , Calculate the corrected first-order moment estimate and second-order moment estimation The formula is: The formula for updating the model parameters is: After training and optimization on a large dataset, the model's AUC reached over 0.9, indicating that the model has high predictive performance. During the training process, the model's performance metrics were continuously monitored, and the model was fine-tuned based on the results of the validation set to ensure the model's generalization ability and stability.
[0070] WSIs data from animal models (simulating colorectal cancer liver metastasis) in preclinical drug trials are preprocessed and then input into a trained model for patch-level prediction. For example, in the prediction of WSIs data from animal models in a certain drug intervention group, the model analyzes its patches and obtains WSIs-level prediction results through a majority voting algorithm. Based on the prediction results, the inhibitory effect of the drug on colorectal cancer liver metastasis is evaluated, providing data support for early screening and optimization of drug development, accelerating the drug development process, and improving the success rate of development.
[0071] In summary, in this drug development example, the pharmaceutical company extensively collected data from 500 samples, including diverse data types and abundant sources. After data preprocessing to ensure its usability, the company delved into cellular-level information, explored potential targets related to LMTMCs, and demonstrated that CRLM scores effectively correlated with clinical characteristics and prognosis, guiding the drug development process. A deep learning model was carefully constructed and trained, and high AUC values were achieved through parameter optimization. The model was used to predict WSIs data from animal models to evaluate the drug's inhibitory effect on liver metastasis. This approach is significant in the early screening and optimization stages of drug development, accelerating the development process, improving the success rate, and opening up new avenues for overcoming the challenge of colorectal cancer liver metastasis.
[0072] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A multi-omic pathological analysis system for predicting the risk of colorectal liver metastasis, characterized in that, The system comprises a data acquisition module, a data preprocessing module, a cell analysis module, a score construction module, a model construction and training module, and a prediction and evaluation module; The data acquisition module collects single-cell data sets, spatial transcriptome data, bulk RNA-seq data, clinical data and whole-slide images WSIs of colorectal cancer patients from GEO databases, TCGA databases and local hospital databases, and performs format conversion and standardization processing on the data; The data preprocessing module respectively utilizes Seurat R software package to perform corresponding preprocessing on the collected different types of data; The cell analysis module calculates the CNV score of cells based on single-cell gene expression and chromosome ordering data using the inferCNV software package, distinguishes between malignant epithelial cells and non-malignant epithelial cells, and identifies LMTMCs, wherein LMTMCs represent liver metastasis-triggered malignant cell subtypes, uses CytoTRACE software package and Monocle to evaluate the stemness and pseudo-temporal trajectory analysis of LMTMCs, respectively, and uses CellChat to analyze the interaction between LMTMCs and other cell types; The score construction module uses OCLR algorithm to calculate the relative abundance of patient LMTMCs in Bulk RNA-seq data to construct CRLM score system, wherein OCLR represents orthogonal category logistic regression, CRLM represents colorectal cancer liver metastasis, analyzes the correlation with clinical pathological characteristics, and verifies the prediction value of the score system for patient survival prognosis by using Kaplan-Meier survival period analysis; The model construction and training module selects ResNet18 architecture to construct a deep learning model, inputs the preprocessed tumor patch, adjusts the model parameters through Adam optimizer, pre-trains the model using the data set, evaluates the model performance using AUC index, and fine-tunes the model for CRLM task; The prediction and evaluation module inputs the preprocessed WSI data of the patient to be predicted into the trained model, performs patch-level classification prediction, and aggregates the patch-level classification results into WSI-level prediction results using the majority voting algorithm.
2. The multi-omic pathological analysis system for predicting the risk of colorectal liver metastasis according to claim 1, characterized in that, In the data preprocessing module, different types of data collected are respectively processed using Seurat R software package: for scRNA-seq data, Seurat R software package is used to remove low-quality cells with more than 6000 or less than 200 genes expressed and mitochondrial read proportion exceeding 30%, normalize and standardize the sample gene expression matrix, and integrate cell data using harmony R package; For spatial transcriptome data, Seurat R software package is used for normalization operation and cell type label mapping transfer; for WSIs data, the data is cut into appropriate patch level and subjected to data enhancement and normalization.
3. The multi-omic pathological analysis system for predicting the risk of colorectal liver metastasis according to claim 1, characterized in that, The identification of liver metastasis-triggered malignant cells (LMTMCs) in the cell data analysis module, the CNV score of the cells is calculated using the inferCNV software package, all endothelial cells (ECs) are hierarchically clustered based on the CNV score by the clustering method of similarity field and density field, the endothelial cells (ECs) with low CNV score and close to the reference cells are identified as non-malignant cells, and the ECs with high CNV score are identified as malignant cells, the malignant cells are clustered again, and the ratio of the actual number of cells of each tumor cell subtype in the CRC different metastasis state sample to the expected number of cells is calculated, the formula is: Wherein the expected number of cells is calculated based on the overall cell proportion of the sample, and Ro / e>1 of the cell subtype in the tissue indicates that the subtype is preferentially distributed in the tissue, and the tumor cell subpopulation mainly distributed in the LMCRC tissue and the CRLM colorectal liver metastasis tissue is identified as LMTMCs.
4. The multi-omic pathological analysis system for predicting the risk of colorectal liver metastasis according to claim 3, characterized in that, The cell analysis module performs hierarchical clustering of all endothelial cells ECs and fibroblasts based on CNV score through clustering method of similarity field and density field: (1) For any two cells i and j with CNV scores CNV i and CNV j , define their similarity S ij as: where a is a similarity decay coefficient; (2) For a cell i, its density p i is defined as the sum of similarity of cells within a certain range around it: p i =∑ j∈N(i) S ij where N(i) denotes the neighborhood of cell i, and is the absolute value of the difference |CNV i -CNV j | of CNV score. (3) Select the cell with the largest density as the initial cluster center C1, and consider the isolation of the cell for the subsequent cluster center. The isolation index I of the cell i to the selected cluster center C k is defined as: ik where m is the number of selected cluster centers, and the cell with the largest I ik is selected as the new cluster center C m+1 each time until the predetermined cluster number K is met. (4) For each cell i, compute its similarity S to all cluster centers C k iCk Assign cell i to the cluster with the largest similarity: (5) Set the CNV score of the reference cell as CNV ref For each EC in the cluster, calculate its average similarity to the reference cell The formula is: where n is the number of ECs in the cluster, S i,ref is the similarity of cell i to the reference cell, calculated according to the formula in step (1), if is greater than a similarity threshold τ and the average CNV score of the cluster is lower, then the ECs in the cluster are identified as non-malignant cells; otherwise, as malignant cells.
5. The multi-omic pathological analysis system for predicting the risk of colorectal liver metastasis according to claim 1, wherein, The scoring module adopts OCLR algorithm to calculate the relative abundance of LMTMCs of the patient, and the gene expression matrix of the patient sample is X, wherein x ij represents the expression value of the ith gene in the jth sample, and for each sample j, the OCLR algorithm constructs a classification model by learning the feature gene expression pattern of LMTMCs, and the parameters of the model are θ, then the probability P(y j =1|x j , θ) that the sample j belongs to the LMTMCs category is expressed by a logistic regression function as follows: where y j is the category label of the sample j, y j =1 indicates belonging to the LMTMCs category, y j =0 indicates not belonging, x j is the gene expression vector of the sample j, and θ T is the transpose of the parameter vector θ, and the relative abundance score of LMTMCs of each sample is obtained by training the OCLR model.
6. The multi-omic pathological analysis system for predicting the risk of colorectal liver metastasis according to claim 1, wherein, The cell analysis module analyzes the influence of CRLM score on the survival prognosis of patients by using Kaplan-Meier survival analysis method, according to the median of CRLM score as the cutoff value, the patients are divided into high score group and low score group, set t as survival time, d as death event, d=1 indicates death, d=0 indicates survival, n as the number of patients at risk, at time t, the estimated value of survival probability S(t) is calculated, the formula is: Where t i is the time point of death event, d i is the number of deaths occurring at time t i , n i is the number of patients at risk at time t i , by drawing Kaplan-Meier survival curve, the survival conditions of patients in high score group and low score group are compared.
7. The multi-omic pathological analysis system for predicting risk of colorectal liver metastasis according to claim 1, wherein, The model construction and training module selects a ResNet18 architecture to construct a deep learning model, and the deep learning model includes an input layer, a convolutional layer, a pooling layer, a residual block, and a fully connected layer: (1) Input layer: receiving pre-processed tumor patch images, with a size of HxWxC, H is the height, W is the width, and C is the number of channels; (2) Convolutional layer: the convolutional layer is used for feature extraction, has a plurality of, each convolutional layer is composed of a convolution kernel, a bias and an activation function, assuming that the convolution kernel size is k x k, the step is s, the padding is p, for the input feature map X, the calculation formula of the output feature map Y is: Wherein W is the convolution kernel weight, b is the bias, i, j is the position index of the output feature map; (3) Pooling layer: mainly for sampling, reducing the size of the feature map, set the window to 2*2 window, step to 2, the calculation formula is: Y ij = max(X (2i,2j) ,X (2i,2j+1) ,X (2i+1,2j) ,X (2i+1,2j+1) ), wherein X is the input feature map, Y is the output feature map, i, j is the position index of the output feature map; (4) Residual block: composed of two convolutional layers and a skip connection, let the input be x, the output after two convolutional layers be F(x), and the output y of the residual block be: y = F(x) + x, when the number of input and output channels is inconsistent, the channel number of the input x is changed through a 1x1 convolution to match the output dimension, and the formula is: y = F(x) + W 1×1 x, where W 1×1 is the weight of the 1x1 convolution, contains multiple residual blocks, and the residual blocks at different stages are different in convolution kernel size, step size, and channel number setting; (5) Full connection layer: located at the end of the network, mapping the extracted features to the final class space, assuming the dimension of the output feature vector is d, the number of classes is C, the weight of the full connection layer is W fc , and the bias is b fc , the output z is: z = W fc x + b fc In the prediction of colorectal cancer liver metastasis risk, C is set to 2, representing metastasis and non-metastasis, and z is converted to a probability distribution by a softmax function, the formula is: Where p i is the probability that the sample belongs to the ith class, and z i is the ith value of the full connection layer output.
8. The multi-omic pathological analysis system for predicting the risk of colorectal liver metastasis according to claim 1, wherein, The model building and training module adjusts the model parameters using the Adam optimizer. Let t represent the current training iteration number, and θ... t The model parameters W or b, g represent the parameters at the t-th generation. t It is the gradient of parameter θ at the t-th iteration, calculated using the partial derivative of the loss function L with respect to parameter θ: First-moment estimate of the parameter m t and second-order moment estimate v t The formula is: m t =β1m t-1 +(1-β1)g t , Where β1 and β2 are hyperparameters, the corrected first-order moment estimate is calculated. and second-order moment estimation The formula is: The formula for updating the model parameters is: Where α is the learning rate, and ∈ is used to prevent the denominator from being a zero constant.
9. The multi-omic pathological analysis system for predicting risk of colorectal liver metastasis according to claim 1, wherein, The prediction and evaluation module uses a majority voting algorithm to aggregate patch-level classification results into WSIs-level prediction results, and counts the number n of patches predicted as positive for liver metastasis among all patches. positive The number of patches n predicted to be negative for liver metastasis. negative n positive >n negative If so, the WSIs are considered positive for liver metastasis; n positive <n negative The WSIs were determined to be negative for liver metastasis. Let the final WSIs grade prediction result be y. WSI Its value is 0 or 1, indicating negative or positive liver metastasis, and the formula is: When n appears positive =n negative Further investigation will be conducted.
10. A multi-omic pathological analysis method for predicting the risk of colorectal liver metastasis, characterized in that, The analysis method comprises the following steps: S100, data acquisition: downloading single-cell data sets, spatial transcriptome data, batch RNA-seq data and clinical data from GEO database and TCGA database, and obtaining WSI data of colorectal cancer patients from hospitals; S200, data preprocessing: quality control of single-cell data, removal of low-quality cells, normalization, standardization and batch effect correction, integration and preprocessing of spatial transcriptome data, cutting of WSI into patches and data enhancement and normalization; S300, cell subtype identification and analysis: using the inferCNV software package to calculate cell CNV scores, distinguishing between malignant and non-malignant epithelial cells, clustering analysis of malignant cells to determine LMTMCs, using CytoTRACE software package and Monocle to evaluate their stemness and perform pseudo-temporal trajectory analysis, and analyzing the interaction between LMTMCs and other cell types at the single-cell and spatial transcriptome levels by CellChat; S400, CRLM score construction and analysis: using OCLR algorithm to calculate the CRLM score of the patient, representing the relative abundance of LMTMCs, analyzing its correlation with clinical pathological characteristics, and performing Kaplan-Meier survival analysis; S500, model construction and training: based on ResNet18 architecture to build a deep learning model, input the pre-processed tumor patch into the model, adjust the model parameters through Adam optimizer, use the data set for pre-training, use AUC index to evaluate the model performance, and fine-tune for colorectal cancer liver metastasis task; S600, risk prediction and evaluation: input the pre-processed WSI data of the patient to be predicted into the trained model, perform patch-level prediction, obtain WSI-level prediction results through majority voting, and output liver metastasis risk prediction value.
Citation Information
Patent Citations
Hepatocellular carcinoma prognosis biomarker and application thereof
CN115807089A
Hepatocellular carcinoma prognosis layering construction method integrating single cell and batch transcriptome pyroptosis characteristics
CN117497184A