Spatial omics-based intestinal cancer metastasis prediction method and device, medium and equipment
By collecting and processing multi-omics data, spatially enhanced expression profiles and heterogeneity features are generated, solving the problem of insufficient multi-scale data integration in colorectal cancer metastasis risk prediction, and realizing high-precision quantification of the tumor microenvironment and early metastasis identification.
Patent Information
- Application Number
- CN202511456553.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-13
AI Technical Summary
Existing technologies struggle to effectively integrate multi-scale data and accurately quantify the spatial heterogeneity of the tumor microenvironment, resulting in insufficient accuracy in predicting colorectal cancer metastasis risk.
By collecting multi-omics data, performing modality alignment and quality control processing, spatially enhanced expression spectra and spatial heterogeneity features are generated. Cross-modal semantic embedding and multi-scale graph construction are used to input the pre-trained metastasis risk prediction model, and output a spatial heatmap of liver metastasis probability and a list of key driving features.
It significantly improves the sensitivity and spatial positioning accuracy of colorectal cancer metastasis prediction, and achieves high-precision quantification of the tumor microenvironment and accurate identification of early metastasis risk.
Smart Images

Figure CN120913863A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical treatment, in particular to a colorectal cancer metastasis prediction method and device based on spatial omics, a medium and equipment. BACKGROUND
[0002] Early prediction of colorectal cancer metastasis is crucial for clinical treatment. Traditional pathological examination relies on morphological observation, which can identify some high-risk features, but it is difficult to quantify the spatial heterogeneity of tumor microenvironment. Molecular detection techniques such as gene expression profiling can provide tumor molecular characteristics, but they lose key spatial information and cannot reflect regional differences within the tumor. Single-cell sequencing improves resolution, but still cannot retain spatial distribution information of cells in situ tissues.
[0003] In recent years, spatial omics technology has made it possible to obtain gene expression and its spatial distribution simultaneously, but existing spatial transcriptome detection has limited resolution, and a single data point may contain mixed cell signals, affecting the identification of micro-metastases. In addition, how to effectively integrate multi-scale data such as molecular expression, cell interaction and tissue structure, and establish a high-precision prediction model is still a technical problem to be solved. Existing prediction methods mostly rely on single data or simple feature stacking, and fail to fully utilize the complementarity of multi-source data, resulting in insufficient prediction accuracy of early metastasis risk. SUMMARY
[0004] In view of the above problems, the present application provides a colorectal cancer metastasis prediction method and device based on spatial omics, medium and equipment, which realizes high-sensitivity metastasis risk detection through dynamic optimization of spatial resolution and multi-scale feature collaborative modeling, and solves the problem of high early metastasis missed detection rate.
[0005] To achieve the above purpose, in a first aspect, the present application provides a colorectal cancer metastasis prediction method based on spatial omics, comprising:
[0006] Collecting original multi-omics data of an original colorectal cancer tissue sample, the original multi-omics data comprising first spatial transcriptome data, first single-cell RNA sequencing data and first pathological image data;
[0007] Performing modal alignment and quality control processing on the original multi-omics data to obtain preprocessed multi-omics data, the preprocessed multi-omics data comprising second spatial transcriptome data, second single-cell RNA sequencing data and second pathological image data;
[0008] Performing cross-modal semantic embedding on the second spatial transcriptome data based on the second single-cell RNA sequencing data to generate a spatially enhanced expression profile, the cross-modal semantic embedding comprising single-cell-space feature alignment based on contrastive learning and probabilistic sparse signal completion, and the spatially enhanced expression profile containing cell type deconvolution results and confidence scores of metastasis-related genes;
[0009] performing multi-scale graph construction on the second spatial transcriptome data and the second pathology image data, extracting spatial heterogeneity features, the multi-scale graph construction including microscopic cell interaction graph, mesoscopic functional region graph and macroscopic tissue structure hierarchical modeling, and the spatial heterogeneity features including tumor-immune adjacency strength, blood vessel invasion front spatial gradient and high-risk sub-region topological relationship;
[0010] inputting the spatial enhanced expression profile and the spatial heterogeneity features into a pre-trained metastasis risk prediction model, and outputting a spatial heat map containing liver metastasis probability and a key driving feature list;
[0011] based on the spatial heat map and the key driving feature list, generating a clinical prediction report of colorectal cancer metastasis risk, the report including high-risk area positioning, metastasis probability score and treatment response prediction.
[0012] Further, the original multi-omics data is subjected to modal alignment and quality control processing to obtain preprocessed multi-omics data, including:
[0013] filtering out spatial sites satisfying a preset quality threshold from the first spatial transcriptome data as the second spatial transcriptome data, the preset quality threshold including a gene detection number threshold and an upper limit of mitochondrial gene proportion;
[0014] extracting cell sequencing data in the first single-cell RNA sequencing data matching the anatomical region of the second spatial transcriptome data as the second single-cell RNA sequencing data, the anatomical region matching being achieved through tissue section coordinate mapping;
[0015] selecting an image region corresponding to the collection position of the second spatial transcriptome data from the first pathology image data as the second pathology image data, the spatial deviation of the collection position in matching being not more than a preset alignment tolerance;
[0016] establishing a spatial correlation relationship among the second spatial transcriptome data, the second single-cell RNA sequencing data and the second pathology image data, so that the three have a unified coordinate reference system;
[0017] performing batch effect correction on the second spatial transcriptome data to eliminate technical variations among different samples;
[0018] performing low-quality cell filtering on the second single-cell RNA sequencing data to retain transcriptome data meeting cell integrity standards;
[0019] generating preprocessed multi-omics data, and storing a plurality of preprocessed multi-omics data in a quality control database.
[0020] Further, based on the second single-cell RNA sequencing data, the second spatial transcriptome data is subjected to cross-modal semantic embedding to generate a spatial enhanced expression profile, including:
[0021] extracting a cell type feature matrix from the second single-cell RNA sequencing data, the cell type feature matrix containing a marker gene expression profile of each cell type;
[0022] aligning each spatial site of the second spatial transcriptome data with the cell type feature matrix by contrastive learning, and calculating a similarity score of each spatial site with each cell type;
[0023] performing probabilistic cell type deconvolution on the spatial sites based on the similarity score to generate a cell type proportion matrix, the deconvolution being optimized by joint non-negative matrix factorization and Markov random field;
[0024] performing sparse signal completion on lowly expressed genes of the second spatial transcriptome data, the completed signal being derived from an expression distribution prior of the same type of cells in single-cell data;
[0025] fusing the cell type proportion matrix and the completed gene expression data to generate a spatially enhanced expression profile, the spatially enhanced expression profile containing cell composition information and gene expression confidence scores of each spatial site.
[0026] Further, the probabilistic cell type deconvolution on the spatial sites based on the similarity score to generate a cell type proportion matrix, the deconvolution being optimized by joint non-negative matrix factorization and Markov random field, includes:
[0027] constructing a non-negative matrix factorization objective function of the spatial sites-cell types by taking the similarity score as an initial weight, and establishing a neighborhood relation graph of the Markov random field based on spatial coordinate information of the second spatial transcriptome data, the neighborhood relation graph containing topological connection relations between the spatial sites;
[0028] introducing a spatial smoothing constraint term in the non-negative matrix factorization objective function, the spatial smoothing constraint term adjusting cell type proportion differences between adjacent sites through the Markov random field;
[0029] updating the cell type proportion matrix iteratively by using an alternating optimization algorithm, each iteration including:
[0030] fixing the spatial constraint term, and optimizing the cell type proportion of the non-negative matrix factorization by a multiplication update rule, and fixing the cell type proportion, and optimizing the spatial consistency by graph Laplacian regularization;
[0031] terminating the optimization when a Frobenius norm change of the cell type proportion matrix of adjacent iterations is less than a preset convergence threshold;
[0032] normalizing the optimized cell type proportion matrix to ensure that a sum of the cell type proportions of each spatial site is 1;
[0033] An output cell type proportion matrix is generated, with rows corresponding to spatial sites and columns corresponding to cell types, and element values being probabilistic proportion weights.
[0034] Further, multi-scale graph construction is performed on the second spatial transcriptome data and the second pathology image data to extract spatial heterogeneity features, including:
[0035] Based on the single-cell resolution expression profile of the second spatial transcriptome data, a microscopic cell interaction network graph is constructed, which is generated through ligand-receptor pair co-expression analysis and cell adjacency relationship calculation;
[0036] The second pathology image data is segmented into tissue regions to identify tumor core regions, invasion front regions, and interstitial regions, and combined with the expression features of the corresponding regions of the second spatial transcriptome data to construct a mesoscopic functional region atlas;
[0037] The tissue structure features of the second pathology image data and the global expression pattern of the spatial transcriptome data are integrated at the whole slice level to establish a macroscopic tissue metastasis trend prediction graph;
[0038] The tumor-immune cell interaction intensity features are extracted from the microscopic cell interaction network graph, including the spatial co-localization frequency of immune checkpoint molecule pairs;
[0039] The spatial gradient features of the vascular invasion front are calculated from the mesoscopic functional region atlas, including the decay coefficient of endothelial marker expression with distance;
[0040] The high-risk sub-region topological features are extracted from the macroscopic tissue metastasis trend prediction graph, including the spatial autocorrelation index of the pre-metastasis microenvironment;
[0041] The tumor-immune cell interaction intensity features, vascular invasion front spatial gradient features, and high-risk sub-region topological features are fused into spatial heterogeneity features.
[0042] Further, the tissue structure features of the second pathology image data and the global expression pattern of the spatial transcriptome data are integrated at the whole slice level to establish a macroscopic tissue metastasis trend prediction graph, including:
[0043] The second pathology image data is scanned at the whole slice level to extract tissue structure features, including tumor gland arrangement directionality, interstitial fibrosis distribution density, and necrotic region spatial proportion;
[0044] The global expression pattern of the second spatial transcriptome data is subjected to spatial autocorrelation analysis to calculate the spatial clustering index of each gene expression hotspot region, and the global expression pattern includes the expression amount spatial distribution of epithelial-mesenchymal transition markers, angiogenic factors, and immune suppression-related molecules;
[0045] A multi-modal fusion model is constructed based on a graph convolutional neural network, and a feature alignment is performed on the tissue structure features and the global expression pattern. The alignment process is realized by a cross-modal attention mechanism to allocate weights of the pathological image features and the transcriptome expression features, and a fused multi-modal feature matrix is output.
[0046] A macroscopic tissue metastasis trend prediction graph is established in the multi-modal feature matrix. The macroscopic tissue metastasis trend prediction graph simulates potential metastasis paths of tumor cells along the tissue structure features and the expression gradient features by a random walk algorithm.
[0047] The macroscopic tissue metastasis trend prediction graph is topologically optimized to identify high-risk metastasis sub-regions. The topological optimization includes calculating the spatial density of the metastasis path intersection points and the metastasis direction consistency coefficient.
[0048] A final macroscopic tissue metastasis trend prediction graph is output. The macroscopic tissue metastasis trend prediction graph includes a metastasis probability heat map, a high-risk sub-region boundary label, and a main metastasis path vector field.
[0049] Further, the spatial enhanced expression profile and the spatial heterogeneity feature are input into a pre-trained metastasis risk prediction model to output a spatial heat map containing liver metastasis probability and a key driving feature list, including:
[0050] The spatial enhanced expression profile is subjected to feature standardization processing to make the expression amounts of different genes comparable, including Z-score normalization and batch effect correction.
[0051] The spatial heterogeneity feature is subjected to dimension compression, and principal component analysis method is used to retain principal component features with a contribution rate exceeding a preset threshold.
[0052] The standardized spatial enhanced expression profile and the compressed spatial heterogeneity feature are subjected to feature splicing to generate a fusion feature matrix.
[0053] The fusion feature matrix is input into a pre-trained graph neural network model. The graph neural network model aggregates multi-scale spatial information based on an attention mechanism.
[0054] The liver metastasis probability of each spatial site is calculated by the graph neural network model, and a spatial heat map is generated based on the probability value. The spatial heat map maintains a spatial correspondence relationship with the original tissue section.
[0055] Gradient backpropagation method is used to extract key driving features from the graph neural network model, including gene expression features and spatial topological features with the highest contribution to the prediction results.
[0056] The key driving features are sorted according to the feature contribution, and a key driving feature list containing feature names, contribution scores, and biological explanations is generated.
[0057] The spatial heat map is stored in association with the list of key driving features, and a mapping relationship with the clinicopathological parameters is established.
[0058] In a second aspect, the present application also provides a device for predicting intestinal cancer metastasis based on spatial omics, which is suitable for the method of the first aspect. The device comprises a data acquisition module, a data processing module, a logical operation module, and a report generation module. The data acquisition module is used to acquire original multi-omics data of an original intestinal cancer tissue sample, and the original multi-omics data comprises first spatial transcriptome data, first single-cell RNA sequencing data, and first pathological image data. The data processing module is used to perform modality alignment and quality control processing on the original multi-omics data to obtain preprocessed multi-omics data, and the preprocessed multi-omics data comprises second spatial transcriptome data, second single-cell RNA sequencing data, and second pathological image data. The logical operation module is used to perform cross-modality semantic embedding on the second spatial transcriptome data according to the second single-cell RNA sequencing data to generate a spatially enhanced expression profile, and the cross-modality semantic embedding comprises single-cell-space feature alignment based on contrastive learning and probabilistic sparse signal completion. The spatially enhanced expression profile contains cell type deconvolution results and confidence scores of metastasis-related genes. Multi-scale graph construction is performed on the second spatial transcriptome data and the second pathological image data to extract spatial heterogeneity features. The multi-scale graph construction comprises hierarchical modeling of microscopic cell interaction graphs, mesoscopic functional region graphs, and macroscopic tissue structures. The spatial heterogeneity features comprise tumor-immune adjacency strength, blood vessel invasion front spatial gradient, and high-risk sub-region topological relationship. The spatially enhanced expression profile and the spatial heterogeneity features are input into a pre-trained metastasis risk prediction model to output a spatial heat map containing liver metastasis probability and a list of key driving features. The report generation module is used to generate a clinical prediction report of intestinal cancer metastasis risk according to the spatial heat map and the list of key driving features. The report comprises high-risk area positioning, metastasis probability scoring, and treatment response prediction.
[0059] In a third aspect, the present application also provides a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the method of the first aspect.
[0060] In a fourth aspect, the present application also provides an electronic device comprising a memory and a processor, wherein the memory is used to store one or more computer program instructions, and the one or more computer program instructions are executed by the processor to implement the method of the first aspect.
[0061] Distinguish from the prior art, the technical scheme provides a kind of intestinal cancer metastasis prediction method, device, medium and equipment based on space omics, method includes: original multi-omics data is collected, modal alignment and quality control processing are carried out, obtain the pre-processing multi-omics data including second spatial transcriptome data, second single cell RNA sequencing data and second pathological image data;Second single cell RNA sequencing data is based on second spatial transcriptome data and carries out cross-modal semantic embedding, generates spatial enhanced expression profile;Second spatial transcriptome data and second pathological image data are carried out multi-scale graph construction, and spatial heterogeneity feature is extracted;Spatial enhanced expression profile and spatial heterogeneity feature are input into pre-trained metastasis risk prediction model, and output liver metastasis probability spatial heat map and key driving feature list;Finally, the clinical prediction report containing high-risk area positioning is generated.The micro-metastasis signal is effectively identified by dynamic optimization of spatial resolution in the application;By multi-scale feature collaborative modeling, the spatial heterogeneity features of tumor microenvironment can be accurately quantified, which can significantly improve the sensitivity and spatial positioning accuracy of intestinal cancer metastasis prediction.
[0062] The above invention content related record is only the summary of the technical scheme of the present application, in order to let the ordinary skilled in the art can more clearly understand the technical scheme of the present application, then can be implemented according to the content of the description and the drawing record, and in order to let the above-mentioned purpose and other purposes, characteristics and advantages of the present application can be more easily understood, the following combining with the specific embodiment of the present application and the drawing are described. BRIEF DESCRIPTION OF DRAWINGS
[0063] The drawings are only used to show the principles, implementation modes, applications, characteristics and effects of the specific embodiments of the present application and other related contents, and cannot be considered as the limitation of the present application.
[0064] In the drawings of the specification:
[0065] Figure 1 The method step chart of steps S101 to S106 of the prediction method described in the specific embodiment is shown in the figure;
[0066] Figure 2 The method step chart of steps S201 to S207 of the prediction method described in the specific embodiment is shown in the figure;
[0067] Figure 3 The method step chart of steps S301 to S305 of the prediction method described in the specific embodiment is shown in the figure;
[0068] Figure 4 The method step chart of steps S401 to S403 of the prediction method described in the specific embodiment is shown in the figure;
[0069] Figure 5 The structure schematic diagram of the electronic device described in the specific embodiment is shown in the figure.
[0070] The reference signs related in the above-described drawings are explained as follows.
[0071] 1. An electronic device;
[0072] 11. A memory;
[0073] 12. A processor. DETAILED DESCRIPTION
[0074] To explain possible application scenarios, technical principles, specific schemes that can be implemented, purposes and effects that can be achieved, etc. of the present application in detail, the following will be described in detail in combination with specific embodiments listed and with the aid of the drawings. The embodiments described in the present document are only used to more clearly illustrate the technical schemes of the present application, and therefore only serve as examples, and cannot limit the protection scope of the present application.
[0075] In the present document, the term "embodiment" means that the specific features, structures or characteristics described in combination with the embodiment can be included in at least one embodiment of the present application. The term "embodiment" appearing at various positions in the specification does not necessarily refer to the same embodiment, and does not particularly limit the independence or association between other embodiments. In principle, in the present application, as long as there is no technical contradiction or conflict, each technical feature mentioned in each embodiment can be combined in any way to form a corresponding implementable technical scheme.
[0076] Unless otherwise defined, the meanings of the technical terms used in the present document are the same as those commonly understood by those skilled in the art to which the present application belongs; the use of related terms in the present document is only for the purpose of describing specific embodiments, and is not intended to limit the present application.
[0077] In the description of the present application, the phrase "and / or" is a description of the logical relationship between objects, which means that there can be three relationships, for example, A and / or B, which means that there are three cases: A exists, B exists, and A and B exist at the same time. In addition, the character " / " in the present document generally represents a "or" logical relationship between the associated objects before and after it.
[0078] In the present application, phrases such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual quantity, primary and secondary or order relationship between the entities or operations.
[0079] In the present application, the "comprise", "include", "have" or other similar open-ended expressions used in a statement are intended to cover the non-exclusive inclusion, and these expressions do not exclude the presence of additional elements in the process, method or product comprising the elements, so that the process, method or product comprising a series of elements can not only include those limited elements, but also include other elements not explicitly listed, or also include the elements inherent to such process, method or product.
[0080] As the same understanding as in the "Guidelines for Examination", in the present application, the expressions such as "greater than", "less than", "exceed" are understood as not including the number; the expressions such as "above", "below", "within" are understood as including the number. In addition, in the description of the embodiments of the present application, the meaning of "multiple" is more than two (including two), and similar expressions related to "multiple" are also understood in this way, for example, "multiple groups", "multiple times" and the like, unless otherwise explicitly specified.
[0081] In the description of the embodiments of the present application, the spatially related expressions used, such as "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "perpendicular", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. The indicated orientation or position relationship is based on the orientation or position relationship shown in the specific embodiment or the drawing, and is only for the convenience of describing the specific embodiments of the present application or for the reader to understand, and does not indicate or imply that the indicated device or component must have a specific position, a specific orientation, or be constructed or operated in a specific orientation, and therefore cannot be understood as a limitation on the embodiments of the present application.
[0082] Please refer to Figure 1 In the first aspect, the embodiments provide a method for predicting intestinal cancer metastasis based on spatial omics, comprising:
[0083] S101, collecting original multi-omics data of original intestinal cancer tissue samples, the original multi-omics data comprising first spatial transcriptome data, first single-cell RNA sequencing data and first pathological image data;
[0084] S102, modal alignment and quality control processing of the original multi-omics data are performed to obtain preprocessed multi-omics data, the preprocessed multi-omics data comprising second spatial transcriptome data, second single-cell RNA sequencing data and second pathological image data;
[0085] S103, performing cross-modal semantic embedding on the second spatial transcriptome data based on the second single-cell RNA sequencing data to generate a spatially enhanced expression profile, the cross-modal semantic embedding comprising contrastive learning based single-cell-spatial feature alignment and probabilistic sparse signal completion, and the spatially enhanced expression profile comprising cell type deconvolution results and confidence scores of metastasis related genes;
[0086] S104, performing multi-scale graph construction on the second spatial transcriptome data and the second pathology image data to extract spatial heterogeneity features, the multi-scale graph construction comprising microscopic cell interaction graph, mesoscopic functional region graph and macroscopic tissue structure hierarchical modeling, and the spatial heterogeneity features comprising tumor-immune adjacency strength, blood vessel invasion front spatial gradient and high-risk sub-region topological relationship;
[0087] S105, inputting the spatially enhanced expression profile and the spatial heterogeneity features into a pre-trained metastasis risk prediction model to output a spatial heat map containing liver metastasis probability and a list of key driving features;
[0088] S106, generating a clinical prediction report of colorectal cancer metastasis risk based on the spatial heat map and the list of key driving features, the report comprising high-risk area positioning, metastasis probability score and treatment response prediction.
[0089] In step S101, the first spatial transcriptome data is a gene expression profile preserving the spatial position of the tissue, which can be obtained through a commercial spatial transcriptome sequencing platform; the first single-cell RNA sequencing data is the transcriptome data of a single tumor microenvironment cell, which is prepared using a clinical-grade single-cell separation platform to ensure that a sufficient number of epithelial cells and immune cells are obtained for subsequent analysis; and the first pathology image data is a digitally scanned H&E stained section, which has a resolution that can clearly distinguish cell nuclei to meet the requirements of pathological diagnosis. It should be noted that the simultaneous collection of raw multi-omics data is achieved by a sample identification system to achieve spatiotemporal matching, and the pretreatment of the tissue sample needs to be completed within a specified time after excision to maintain RNA integrity.
[0090] In step S102, modal alignment maps the data obtained by different detection techniques to a unified coordinate system, and precise registration of the spatial transcriptome chip and the tissue section is achieved through tissue structure markers in the pathology image. In the quality control process, preferably, the second spatial transcriptome data is confirmed by detecting the RNA quality of the internal reference genes, the second single-cell RNA sequencing data filters low-quality cells and double-cell signals, and the second pathology image data verifies the scanning integrity and staining uniformity. Further, the preprocessed data needs to meet the minimum quality control standards for downstream analysis, for example, a specified number of valid genes need to be detected at each detection point of the spatial transcriptome.
[0091] In step S103, single-cell-space feature alignment maps gene expression patterns at single-cell resolution to spatial transcriptomic sites, achieved by similarity measure based on cell type signatures; probabilistic sparse signal completion infers undetected gene expression in spatial transcriptome using single-cell data, with prior distribution set as known molecular signatures of metastasis-related pathways in intestinal cancer. In the generated spatially enhanced transcriptome, cell type deconvolution results are proportional distribution of different cell types in each spatial site, and confidence scores of metastasis-related genes are correlation strength between genes and metastasis phenotype evaluated by cross-validation.
[0092] In step S104, micro-cell interaction map depicts spatial relationship between individual tumor cells and neighboring immune cells, achieved by calculating nuclear distance and contact area; meso-functional region map divides tumor subregions with similar gene expression patterns, which can be generated by graph-based clustering algorithm; macro-tissue structure map analyzes overall growth pattern and invasion characteristics of tumor, which is consistent with the evaluation standard of vascular invasion in pathological diagnosis. In the extracted spatial heterogeneity features, tumor-immune adjacency strength is used to quantify the surrounding degree of immune cells to tumor cells, and vascular invasion front spatial gradient is used to measure the variation pattern of metastasis-related gene expression along the blood vessels.
[0093] In step S105, the metastasis risk prediction model refers to a deep learning system trained by clinical data, whose input layer receives both gene expression features and spatial topology features. The output spatial heat map visualizes the metastasis probability to the two-dimensional projection of the original tissue location, which is generated by interpolation algorithm commonly used in medical image analysis; the key driver feature list is the molecular and spatial features with the highest contribution to the prediction results, preferably extracted by feature importance ranking algorithm and confirmed by pathological experts.
[0094] In step S106, high-risk area positioning refers to mapping the areas in the spatial heat map that exceed the clinical threshold to the coordinates of the pathological section; metastasis probability score is a composite risk assessment value integrating molecular and spatial features, whose grading standard is consistent with the metastasis risk stratification of clinical guidelines; treatment response prediction is evidence-based medicine evidence matching targeted treatment plan according to key driver features.
[0095] The embodiment utilizes single-cell resolution data to enhance the resolution capability of spatial transcriptome, captures the heterogeneity characteristics of tumor microenvironment through pathological image guided multi-scale spatial modeling, and finally fuses molecular expression patterns and spatial topological relationship to evaluate the metastasis risk. By integrating spatial transcriptome data, single-cell RNA sequencing data and pathological image data, multi-dimensional accurate prediction of colorectal cancer metastasis risk is realized. Spatially enhanced expression profile solves the problem of insufficient resolution of traditional spatial transcriptome data, multi-scale graph construction realizes cross-level feature extraction from cell interaction to tissue structure, and clinical prediction report combines molecular characteristics and pathological morphology organically, providing a comprehensive solution with spatial resolution and clinical interpretability for colorectal cancer metastasis risk evaluation, which helps to guide individualized treatment decisions.
[0096] Please refer to Figure 2 In some embodiments, the original multi-omics data is subjected to modal alignment and quality control processing to obtain preprocessed multi-omics data, including:
[0097] S201, filtering out spatial sites meeting a preset quality threshold from the first spatial transcriptome data as second spatial transcriptome data, the preset quality threshold including a gene detection number threshold and an upper limit of mitochondrial gene proportion;
[0098] S202, extracting cell sequencing data matching the anatomical region of the second spatial transcriptome data from the first single-cell RNA sequencing data as second single-cell RNA sequencing data, the anatomical region matching being achieved through tissue section coordinate mapping;
[0099] S203, selecting an image region corresponding to the collection position of the second spatial transcriptome data from the first pathological image data as second pathological image data, the spatial deviation of the collection position in matching being not more than a preset alignment tolerance;
[0100] S204, establishing a spatial correlation relationship between the second spatial transcriptome data, the second single-cell RNA sequencing data and the second pathological image data, so that the three have a unified coordinate reference system;
[0101] S205, batch effect correction is performed on the second spatial transcriptome data to eliminate technical variations between different samples;
[0102] S206, filtering low-quality cells is performed on the second single-cell RNA sequencing data to retain transcriptome data meeting cell integrity standards;
[0103] S207, generating preprocessed multi-omics data, and storing the plurality of preprocessed multi-omics data to a quality control database.
[0104] In step S201, a quality threshold is preset for evaluating spatial site data quality, wherein a gene detection number threshold is determined by a technical white paper of a sequencing platform, representing a minimum number of genes that should be detected for each site; and a mitochondrial gene proportion upper limit is dynamically adjusted according to a cell activity detection result, for identifying a cell damage degree. This step retains second spatial transcriptome data with biological significance through quality filtering, while removing noise signals caused by technical factors.
[0105] In step S202, anatomical region matching spatially corresponds a single cell sample source location with a spatial transcriptome chip coordinate, preferably, coordinate mapping is realized by an affine transformation algorithm in a digital pathology system, and this process needs to correct deformation errors caused by tissue section processing. The extracted second single cell RNA sequencing data ensures that a main cell subpopulation of a target region is included, providing an accurate single cell reference map for subsequent spatial analysis.
[0106] In step S203, a preset alignment tolerance refers to a maximum allowed positional deviation between a pathology image and spatial transcriptome data, which is calculated and determined according to a tissue section thickness and an image resolution, and the selected second pathology image data completely covers a detection region of the spatial transcriptome, preferably, sub-millimeter level alignment accuracy is realized through an image registration function of a digital pathology software.
[0107] In step S204, a spatial correlation is a unified spatial coordinate system of multi-modal data, which is constructed based on an H&E staining structure of a pathology image, and the second spatial transcriptome data and the second single cell RNA sequencing data are mapped to the same reference system through a feature point matching algorithm. This step retains spatial accuracy of original data, so that data of different modalities can be jointly analyzed at the same spatial scale.
[0108] In step S205, batch effect correction preferably adopts a negative binomial distribution model to identify technical variations, and a ComBat algorithm is used to eliminate influences of sample preparation and sequencing batches. The corrected second spatial transcriptome data retains true biological differences, ensuring reliability of cross-sample comparison.
[0109] In step S206, a cell integrity standard includes indicators such as gene detection number and ribosomal gene proportion, and high-quality cells are screened through automatic clustering analysis combined with manual review. The filtered second single cell RNA sequencing data improves accuracy of downstream analysis and avoids interference caused by low-quality cells.
[0110] In step S207, a quality control database can be constructed using a clinical research data management system, and multi-omics data is preprocessed and stored in HDF5 format with metadata index, supporting standardized management and traceability query of multi-center research data.
[0111] The embodiment realizes the accurate integration of spatial transcriptome, single-cell sequencing and pathological image data through a systematic multi-modal data quality control and alignment process. High-quality spatial sites are selected based on technical parameters and biological characteristics, affine transformation is used to achieve accurate matching of anatomical regions, image registration is used to ensure multi-modal spatial alignment, and batch correction and cell filtering are used to eliminate technical variations, finally a standardized dataset under a unified coordinate system is constructed. The embodiment combines histological positioning accuracy with molecular feature analysis, making spatial multi-omics data retain original biological characteristics and be comparable, providing high-fidelity input for subsequent cross-scale analysis and improving the reliability and depth of multi-modal data integration.
[0112] Referring to Figure 3 In some embodiments, the second spatial transcriptome data is subjected to cross-modal semantic embedding based on the second single-cell RNA sequencing data to generate spatially enhanced expression profiles, including:
[0113] S301. Extracting a cell type feature matrix from the second single-cell RNA sequencing data, the cell type feature matrix containing a marker gene expression profile of each cell type;
[0114] S302. Comparatively learning and aligning each spatial site of the second spatial transcriptome data with the cell type feature matrix to calculate a similarity score of the spatial site with each cell type;
[0115] S303. Probabilistic cell type deconvolution of the spatial site based on the similarity score to generate a cell type proportion matrix, the deconvolution using joint optimization of non-negative matrix factorization and Markov random field;
[0116] S304. Sparse signal completion for low-expression genes of the second spatial transcriptome data, the completed signal derived from the expression distribution prior of the same type of cells in the single-cell data;
[0117] S305. Fusing the cell type proportion matrix and the completed gene expression data to generate a spatially enhanced expression profile, the spatially enhanced expression profile containing cell composition information and gene expression confidence score of each spatial site.
[0118] In step S301, the cell type feature matrix is a cell subpopulation molecular feature map obtained by single-cell clustering analysis, and the construction process includes: using standard single-cell analysis processes such as Seurat to normalize the second single-cell RNA sequencing data, identifying main cell subpopulations through PCA dimension reduction and UMAP visualization, and finally extracting marker genes differentially expressed in each cell type to form a feature matrix. The row vector of the cell type feature matrix represents the gene, the column vector represents the cell type, and the matrix element is the average expression of the marker gene in the specific cell type.
[0119] In step S302, the alignment of the contrastive learning is to learn the shared feature space of the single cell and the spatial data by the neural network model. Specifically, the maximum mean difference is used as the distribution distance measurement, and the mixed expression profile of the spatial site is projected to the single cell feature space. The similarity score is calculated by the cosine similarity, reflecting the matching degree of the spatial site expression pattern and the reference profile of each cell type. The score calculation considers the dropout noise characteristics of the spatial transcriptome.
[0120] In step S303, the process of joint optimization of non-negative matrix factorization and Markov random field can be understood as: the cell type proportion is preliminarily estimated by NMF, and the MRF model is introduced to integrate the cell composition similarity constraint of adjacent spatial sites. The energy function includes a data fitting term and a spatial smoothing term. Preferably, the optimization process uses a variational expectation maximization algorithm, in which the spatial smoothing weight is adaptively adjusted according to the continuity of the tissue region.
[0121] In step S304, cells of the same type have a conservative gene expression program, which provides a basis for sparse signal completion for low-expression genes. For genes not detected in the spatial data, the completion value is sampled from the expression distribution of the cell type in the single cell data, and the distribution parameters are preferably estimated by Bayes, considering covariates such as gene length and sequencing depth. Further, the completed expression value is accompanied by a confidence score, reflecting the expression stability of the gene in the corresponding cell type.
[0122] In step S305, the fusion process preferably uses a weighted splicing method. The cell type proportion matrix provides spatial microenvironment composition information, and the completed gene expression data retains the original detection signal. The two are associated through the spatial coordinates. In the generated spatial enhanced expression profile, the gene expression confidence score comprehensively considers the original UMI count quality, the reliability of the completed signal, and the cell type specificity.
[0123] The embodiment transmits the fine classification ability at the single cell level to the spatial data by establishing a cross-modal semantic correspondence, while retaining the original spatial topological information. The probabilistic deconvolution handles the cell type colocalization, and the sparse completion restores the biological signals lost due to technical reasons. The final output of the spatial enhanced expression profile has both the classification accuracy of single cell resolution and the in-situ retention advantage of spatial transcriptome, and is particularly suitable for the analysis of highly heterogeneous tissues such as tumor microenvironment, and can accurately analyze the complex spatial interaction mode of malignant cells and stromal cells.
[0124] See Figure 4 In some embodiments, the spatial sites are probabilistically deconvoluted based on the similarity score to generate a cell type proportion matrix. The deconvolution uses joint optimization of non-negative matrix factorization and Markov random field, including:
[0125] S401, constructing a non-negative matrix factorization objective function of spatial sites-cell types with the similarity score as an initial weight, establishing a neighborhood relation graph of a Markov random field based on spatial coordinate information of the second spatial transcriptome data, the neighborhood relation graph comprising a topological connection relation between spatial sites;
[0126] S402, introducing a spatial smoothing constraint term into the non-negative matrix factorization objective function, the spatial smoothing constraint term adjusting a difference in proportions of cell types between adjacent sites through the Markov random field;
[0127] S403, iteratively updating a cell type proportion matrix by using an alternating optimization algorithm, each iteration comprising:
[0128] fixing the spatial constraint term, optimizing the cell type proportion of the non-negative matrix factorization through a multiplication update rule, and fixing the cell type proportion, optimizing spatial consistency through graph Laplacian regularization;
[0129] terminating the optimization when a Frobenius norm change of the cell type proportion matrix of adjacent iterations is less than a preset convergence threshold;
[0130] normalizing the optimized cell type proportion matrix to ensure that a sum of cell type proportions of each spatial site is 1;
[0131] outputting the cell type proportion matrix, a row of the cell type proportion matrix corresponding to a spatial site, a column of the cell type proportion matrix corresponding to a cell type, and an element value being a probabilistic proportion weight.
[0132] In step S401, the non-negative matrix factorization objective function is used to decompose mixed expression profiles of spatial sites into a linear combination of a cell type feature matrix and a proportion matrix, wherein the similarity score as an initial weight can effectively guide the decomposition direction, and the cross-modal correspondence relationship between single-cell data and spatial data is used, so that the decomposition result is more in line with biological reality. The neighborhood relation graph of the Markov random field is constructed through spatial coordinate information, preferably, the topological connection between sites is determined by using a Delaunay triangulation algorithm, which can adaptively reflect the spatial organization structure of different density regions, and provide accurate neighborhood definition for subsequent spatial constraints.
[0133] In step S402, the spatial smoothing constraint term is introduced into the objective function in the form of graph Laplacian regularization, which maintains spatial continuity by adjusting the difference degree of cell type proportions between adjacent sites, wherein the constraint strength parameter is dynamically adjusted according to the heterogeneity degree of the tissue region, which avoids the boundary blur caused by excessive smoothing and effectively suppresses the false difference caused by technical noise; the potential function design of the Markov random field considers the biological characteristics of cell type colocalization, so that the model can distinguish between real cell mixing regions and artificial artifacts.
[0134] In step S403, the alternating optimization algorithm gradually approaches the optimal solution through a decomposition iterative process, in which the multiplicative update rule ensures the non-negativity of the cell type proportions, and the graph Laplacian regularization efficiently realizes the spatial consistency constraint through matrix operations. The setting of the convergence threshold comprehensively considers the balance between analytical precision and computational efficiency, and the final result after normalization processing meets the basic requirements of probability distribution. Further, by introducing a temperature coefficient, the information of the minor cell components in the microenvironment is retained.
[0135] This embodiment organically combines the molecular precision of single-cell data analysis with the topological information of spatial transcriptome, in which the construction of adaptive neighborhood relationship and dynamic constraint strength enable the model to adapt to the clear boundary of tumor core area and the gradual change characteristics of infiltration area, providing a reliable computational framework for quantitative analysis of complex tissue microenvironment, especially suitable for application scenarios such as tumor immunotherapy response prediction that require accurate spatial composition information.
[0136] In some embodiments, multi-scale graph construction is performed on the second spatial transcriptome data and the second pathology image data to extract spatial heterogeneity features, including:
[0137] Based on the single-cell resolution expression profile of the second spatial transcriptome data, a microscopic cell interaction network graph is constructed, which is generated through ligand-receptor pair co-expression analysis and cell adjacency relationship calculation;
[0138] The second pathology image data is segmented into tissue regions to identify tumor core area, invasion front area and interstitial area, and combined with the expression features of the corresponding regions of the second spatial transcriptome data, a mesoscopic functional region atlas is constructed;
[0139] The tissue structure features of the second pathology image data at the whole slice level and the global expression pattern of the spatial transcriptome data are integrated to establish a macroscopic tissue metastasis trend prediction graph;
[0140] The tumor-immune cell interaction intensity features are extracted from the microscopic cell interaction network graph, including the spatial co-localization frequency of immune checkpoint molecule pairs;
[0141] The spatial gradient features of the vascular invasion front are calculated from the mesoscopic functional region atlas, including the decay coefficient of endothelial marker expression with distance change;
[0142] High-risk sub-region topological features are extracted from the macroscopic tissue metastasis trend prediction graph, including the spatial autocorrelation index of the microenvironment before metastasis;
[0143] The tumor-immune cell interaction intensity features, vascular invasion front spatial gradient features and high-risk sub-region topological features are fused as spatial heterogeneity features.
[0144] In this embodiment, the micro-cell interaction network graph is a cell-cell interaction relationship graph constructed based on single-cell resolution expression profiles, which is generated by analyzing the co-expression pattern of ligand-receptor pairs and combining cell spatial adjacency relationship calculation, representing the molecular communication characteristics between cells in the tumor microenvironment. The co-expression analysis of ligand-receptor pairs uses a verified intercellular interaction database as a reference, and the cell adjacency relationship is determined by Voronoi diagram division of spatial coordinates. This multi-modal integration method can accurately identify biologically meaningful cell interaction events.
[0145] The mesoscopic functional region map is obtained by segmenting the tissue regions of the pathological image, wherein the tumor core region is identified by cell density and nuclear atypia characteristics, the invasion front region is determined according to the morphological characteristics of the tumor-stromal interface, and the stromal region is defined by collagen fiber distribution and immune cell infiltration. The expression characteristics of the second spatial transcriptome data of each functional region are integrated by spatial coordinate alignment.
[0146] The macroscopic tissue metastasis trend prediction map integrates the histopathological features of the whole slice level and the global expression pattern of the spatial transcriptome, wherein the histological structure features include morphological indicators such as the degree of vascular invasion and the distribution of necrotic areas, and the global expression pattern is extracted after dimensionality reduction by principal component analysis. The macroscopic tissue metastasis trend prediction map models the spatial propagation rule of the histological structure by graph neural network, which is used to infer the potential path of tumor metastasis.
[0147] The tumor-immune cell interaction intensity feature is extracted from the micro-network graph, focusing on the spatial co-localization of immune checkpoint molecule pairs, and its calculation process considers the product of ligand-receptor expression and the weighted relationship of cell distance. The spatial gradient feature of the vascular invasion front is obtained by fitting the decay curve of endothelial marker expression with distance, wherein the calculation of the decay coefficient uses a nonlinear regression method. The spatial autocorrelation index in the high-risk sub-region topology feature can be quantified by Moran's I statistic, reflecting the regional homogeneity of the microenvironment before metastasis.
[0148] This embodiment constructs a complete spatial heterogeneity analysis framework by extracting and integrating features at three scales of micro, meso and macro. The micro-cell interaction network graph reveals the cell interaction mechanism, the mesoscopic functional region map locates the functional region characteristics, and the macroscopic tissue metastasis trend prediction map grasps the organizational evolution trend. This multi-scale integration method can comprehensively analyze the spatial dynamic rules of tumor development. In terms of immunotherapy response prediction, by integrating the immune checkpoint co-localization feature and the vascular invasion pattern, more accurate spatial basis can be provided for treatment plan selection.
[0149] In some embodiments, the integration of the tissue structure features of the second whole-slice level pathological image data and the global expression patterns of the spatial transcriptome data establishes a macroscopic tissue metastasis trend prediction graph, including:
[0150] Whole-slice scanning is performed on the second pathological image data to extract tissue structure features, including tumor gland arrangement directionality, interstitial fibrosis distribution density, and necrotic area spatial proportion.
[0151] Spatial autocorrelation analysis is performed on the global expression patterns of the second spatial transcriptome data to calculate the spatial clustering index of each gene expression hotspot region. The global expression patterns include the spatial distribution of expression amounts of epithelial-mesenchymal transition markers, angiogenic factors, and immunosuppression-related molecules.
[0152] A multi-modal fusion model is constructed based on a graph convolutional neural network to align the tissue structure features and the global expression patterns. The alignment process realizes the weight distribution of pathological image features and transcriptome expression features through a cross-modal attention mechanism, and outputs a fused multi-modal feature matrix.
[0153] A macroscopic tissue metastasis trend prediction graph is established in the multi-modal feature matrix. The macroscopic tissue metastasis trend prediction graph simulates the potential metastasis path of tumor cells along the tissue structure features and expression gradient features through a random walk algorithm.
[0154] Topological optimization is performed on the macroscopic tissue metastasis trend prediction graph to identify high-risk metastasis sub-regions. The topological optimization includes calculating the spatial density of metastasis path intersection points and the consistency coefficient of metastasis direction.
[0155] The final macroscopic tissue metastasis trend prediction graph is outputted. The macroscopic tissue metastasis trend prediction graph includes a metastasis probability heat map, a high-risk sub-region boundary label, and a main metastasis path vector field.
[0156] In this embodiment, the tumor gland arrangement directionality refers to the distribution characteristics of the main axis direction of the gland structure identified through digital pathology image analysis. Preferably, a gray level co-occurrence matrix combined with a direction filter is used for calculation to reflect the invasive growth pattern of tumor tissue. The interstitial fibrosis distribution density can be quantitatively analyzed by a deep learning segmentation model (such as U-Net) on the Masson staining area, and further, non-specific staining interference such as blood vessel wall is excluded. The necrotic area spatial proportion is determined by morphological processing combined with threshold segmentation. It should be noted that consistency verification with artificial interpretation results is performed to ensure clinical relevance.
[0157] The identification of hotspots of gene expression in spatial autocorrelation analysis can adopt spatial statistical methods, in which the spatial clustering index of epithelial-mesenchymal transition markers is calculated by the local Moran's I algorithm, the expression gradient characteristics of angiogenic factors suggest that thin plate spline interpolation is used for smoothing processing, and the distribution pattern analysis of immunosuppression-related molecules needs to consider the truncation effect of tissue boundaries. The analysis of global expression patterns maintains data consistency with spatially enhanced expression profiles.
[0158] Preferably, the implementation of the cross-modal attention mechanism is based on the Transformer architecture, in which the query vector comes from the pathological image features, the key-value pair comes from the transcriptome features, and the weight assignment process needs to add relative position encoding to maintain spatial correspondence. Further, the generation of multi-modal feature matrix can adopt graph convolution operation, and the construction of adjacency matrix should be clearly distinguished from the scale of the microcellular interaction network graph.
[0159] When simulating the potential metastasis path of tumor cells along the tissue structure characteristics and expression gradient characteristics by random walk algorithm, it needs to be noted that: the anisotropic migration weight provided by the arrangement direction of the gland, the resistance coefficient derived from the fibrosis density, and the absolute prohibition constraint of the necrotic area. Preferably, the walk step should be parameterized according to the histological scale, and the simulation results need to be ensured by multiple independent repetitions.
[0160] In the process of topological optimization of the macroscopic tissue metastasis trend prediction map, the calculation of the metastasis direction consistency coefficient excludes the interference of the local vortex area, and the determination of the high-risk sub-region boundary is combined with the tumor front area defined by pathology for calibration. The final vector field visualization can use streamline tracking algorithm, and keep accurate correspondence with the original spatial coordinate system.
[0161] The present embodiment constructs a multi-scale linked tumor metastasis prediction framework, which integrates pathological features such as tumor gland arrangement directionality and interstitial fibrosis distribution density with molecular expression patterns such as epithelial-mesenchymal transition markers and angiogenic factors through graph convolutional neural network for multi-modal fusion, uses cross-modal attention mechanism to realize feature alignment, and simulates the migration path of tumor cells along the tissue structure and expression gradient by random walk algorithm. The present embodiment establishes a prediction bridge from micro-molecular features to macro-tissue morphology, and the output macro-tissue metastasis trend prediction map contains metastasis probability heat map, high-risk sub-region boundary annotation and main metastasis path vector field, which provides a decision basis with spatial resolution and molecular specificity for clinical evaluation of tumor metastasis risk, helps to guide surgical boundary planning and targeted therapy strategy, and is especially suitable for evaluating the metastasis risk of highly heterogeneous tumors and guiding precise intervention.
[0162] In some embodiments, the spatially enhanced expression profile and the spatial heterogeneity feature are input into a pre-trained metastasis risk prediction model, and a spatial heat map containing liver metastasis probability and a key driving feature list are output, including:
[0163] The spatially enhanced expression profile is subjected to feature standardization processing to make the expression amounts of different genes comparable, including Z-score normalization and batch effect correction;
[0164] The spatial heterogeneity feature is subjected to dimension compression, and principal component analysis method is used to retain principal component features with a contribution rate exceeding a preset threshold;
[0165] The standardized spatially enhanced expression profile and the compressed spatial heterogeneity feature are subjected to feature splicing to generate a fusion feature matrix;
[0166] The fusion feature matrix is input into a pre-trained graph neural network model, and the graph neural network model aggregates multi-scale spatial information based on an attention mechanism;
[0167] The liver metastasis probability of each spatial site is calculated by the graph neural network model, and a spatial heat map is generated based on the probability value, and the spatial heat map maintains a spatial correspondence with the original tissue section;
[0168] Gradient backpropagation method is used to extract key driving features from the graph neural network model, including gene expression features and spatial topological features with the highest contribution to the prediction results;
[0169] The key driving features are sorted according to the feature contribution, and a key driving feature list containing feature names, contribution scores and biological explanations is generated;
[0170] The spatial heat map and the key driving feature list are stored in association, and a mapping relationship with clinical pathological parameters is established.
[0171] In this embodiment, the standardization processing of the spatially enhanced expression profile uses Z-score normalization method to eliminate the dimensional difference between gene expression amounts, and ComBat algorithm is used for batch effect correction to ensure the comparability between different samples. Further, specific correction factors for different tissue regions can be introduced in the standardization process, and different normalization strategies are used for different microenvironment regions such as tumor core region and infiltration front, so as to better retain the biologically relevant spatial expression variation.
[0172] The spatial heterogeneity feature represents the spatial variation pattern of gene expression in the tissue microenvironment, and its dimension compression process is realized by principal component analysis, in which the preset threshold is preferably set to principal component features with a cumulative contribution rate of 85%, which can be determined by elbow rule or parallel analysis. Preferably, when retaining principal components, spatial autocorrelation analysis results can be combined to preferentially select feature components that present a significant aggregation pattern in space.
[0173] The fusion feature matrix generated by feature splicing is used to integrate molecular expression profiles and spatial distribution features. The implementation is preferably transverse splicing along the feature dimension, and the correspondence of the original spatial coordinates is maintained. Optionally, mutual information analysis can be performed on the two types of features before splicing, and the splicing weight is dynamically adjusted according to the correlation between the features to avoid interference of redundant information.
[0174] The pre-trained graph neural network model realizes multi-scale spatial information aggregation based on the multi-head attention mechanism, wherein the construction of the graph structure considers both spatial proximity and feature similarity connections, and the node feature update process adopts the graph attention network (GAT) architecture. Further, the attention mechanism can introduce a spatial distance decay factor, so that the information transmission weight of adjacent nodes decreases exponentially with the increase of distance, which is more consistent with the biological diffusion law.
[0175] The generation of the spatial heat map is realized by mapping the predicted probability value of each spatial site to the corresponding tissue region, and the spatial correspondence is ensured by registration to the H&E stained section. Preferably, the heat map generation can be superimposed with histological structure boundary information, so that the prediction result is directly corresponding to the pathological morphological features.
[0176] When extracting key driving features by the gradient back propagation method, the integral gradient method is used to calculate the feature contribution degree, and the gene expression features and spatial topological features that have the greatest impact on liver metastasis prediction are identified. Optionally, the contribution degree calculation can be verified by combining Shapley value to ensure the stability of feature importance evaluation.
[0177] The biological interpretation of the key driving feature list can be realized by KEGG pathway enrichment analysis and spatial co-localization verification, and the mapping relationship with clinical pathological parameters is established by calculating the Spearman rank correlation coefficient. Further, a feature-phenotype association network can be established to visualize the complex interactions between key driving genes and clinical indicators. Preferably, the association storage of spatial heat map and feature list adopts a spatial database structure, supports joint query based on location and features, and can be extended to integrate electronic medical record data to realize multi-dimensional association analysis.
[0178] The embodiment realizes multi-scale information fusion from molecular expression features to tissue spatial heterogeneity by splicing the standardized spatial enhanced expression profile and the spatial heterogeneity features compressed by principal component analysis into a fusion feature matrix, and inputting the graph neural network model based on the attention mechanism. The spatial heat map generated by the method accurately corresponds to the original tissue section, and the key driving feature list extracted by gradient back propagation reveals the molecular mechanism of liver metastasis. The prediction model established in the embodiment not only retains the molecular specificity of spatial transcriptome data, but also integrates the spatial heterogeneity features of the tissue microenvironment. The output spatial heat map and driving feature list provide a decision basis for clinical treatment with spatial resolution and molecular interpretation, and are particularly suitable for evaluating liver metastasis risk and guiding individualized treatment.
[0179] In a second aspect, the present application also provides a spatial omics-based intestinal cancer metastasis prediction device, which is suitable for the method of the first aspect. The device comprises a data acquisition module, a data processing module, a logical operation module, and a report generation module. The data acquisition module is used to acquire original multi-omics data of an original intestinal cancer tissue sample, and the original multi-omics data comprises first spatial transcriptome data, first single-cell RNA sequencing data, and first pathological image data.
[0180] The data processing module is used to perform modal alignment and quality control processing on the original multi-omics data to obtain preprocessed multi-omics data, and the preprocessed multi-omics data comprises second spatial transcriptome data, second single-cell RNA sequencing data, and second pathological image data.
[0181] The logical operation module is used to perform cross-modal semantic embedding on the second spatial transcriptome data according to the second single-cell RNA sequencing data to generate a spatial enhanced expression profile, the cross-modal semantic embedding comprises single-cell-spatial feature alignment based on contrastive learning and probabilistic sparse signal completion, and the spatial enhanced expression profile comprises cell type deconvolution results and confidence score of metastasis-related genes; multi-scale graph construction is performed on the second spatial transcriptome data and the second pathological image data to extract spatial heterogeneity features, the multi-scale graph construction comprises hierarchical modeling of microscopic cell interaction graph, mesoscopic functional region graph, and macroscopic tissue structure, and the spatial heterogeneity features comprise tumor-immune adjacency strength, blood vessel invasion front spatial gradient, and high-risk sub-region topological relationship; the spatial enhanced expression profile and the spatial heterogeneity features are input into a pre-trained metastasis risk prediction model to output a spatial heat map containing liver metastasis probability and a key driving feature list;
[0182] The report generation module is used to generate a clinical prediction report of intestinal cancer metastasis risk according to the spatial heat map and the key driving feature list, and the report comprises high-risk area positioning, metastasis probability score, and treatment response prediction.
[0183] In the present embodiment, the data acquisition module can be a hardware system integrating a high-throughput sequencer and a digital pathology scanner, which is used to automatically collect original multi-omics data of intestinal cancer tissue samples, and ensures the compatibility of data input through a standardized interface.
[0184] The data processing module is deployed on a high-performance computing node, and is configured with a modal alignment algorithm and a quality control pipeline, and is used to generate preprocessed multi-omics data that can be directly used for analysis. Preferably, the data processing module accelerates large-scale data processing through a parallel computing architecture.
[0185] The logic operation module is preferably implemented by a GPU-accelerated server, and its core is a pre-trained graph neural network model. Through the built-in cross-modal embedding unit and multi-scale graph analysis engine, the input spatial enhanced expression profile and spatial heterogeneity features are converted into liver metastasis risk prediction results. The output interface of the logic operation module supports standardized data exchange formats.
[0186] The report generation module integrates visualization tools and clinical report templates, which are used to automatically combine the spatial heat map and the list of key driving features output by the model into an interactive HTML report, and support direct calling by the electronic medical record system of medical institutions.
[0187] The present embodiment realizes the full-process automation of intestinal cancer liver metastasis prediction through modular design. The data acquisition module ensures the standardized acquisition of multi-omics data, the data processing module completes the seamless integration of multi-modal information, the logic operation module realizes efficient analysis of complex spatial features based on hardware acceleration, and the report generation module provides clinically friendly decision support output. The device converts the advanced spatial omics analysis method into a standardized and deployable medical equipment solution, significantly improves the reliability and clinical applicability of intestinal cancer metastasis risk assessment, and provides end-to-end technical support from experimental data to diagnosis and treatment suggestions for the pathology department.
[0188] In a third aspect, the present embodiment also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement the method of the first aspect.
[0189] The computer program involved in the embodiment can be stored in a computer device readable storage medium, including but not limited to magnetic disk, magnetic tape, magnetic card, floppy disk, flash memory, optical disc, optical card, read-only memory (ROM), random access memory (RAM), erasable programmable ROM (EPROM) and electrically erasable programmable ROM (EEPROM) and the like, and also includes other biological, physical or chemical structures that can realize similar or equivalent functions as the above-mentioned storage media, such as DNA, RNA, protein and the like units with information storage ability. In specific embodiments, the storage medium involved can be one of the above-mentioned medium types, or a combination of the above-mentioned medium types. In different embodiments, the computer program involved in the embodiment can be centrally stored in a single medium, or can be distributedly stored in multiple media. The storage medium containing the computer device readable storage medium can be a non-volatile memory or a random access memory. These computer device readable storage media can be built-in in the device, or connected with the device as an external device or part of the external device. In some embodiments, the storage medium with the computer device readable storage medium is deployed locally; in other embodiments, the storage medium can also be deployed remotely from the processor, for example, network attached storage accessed via RF circuit or external port and communication network, wherein the communication network can be Internet, one or more intranets, local area network (LAN), wide area network (WAN), storage area network (SAN) and the like, or appropriate combination thereof, as long as the access of the computer device to the storage medium can be realized. In addition, the computer program involved in the embodiment can be stored in plaintext / ciphertext form, or can be designed as training data, and integrated and reorganized by model training to be implicitly saved in the parameter state of the deep neural network or other machine learning model.
[0190] Please refer to Figure 5 In the fourth aspect, the embodiment also provides an electronic device 1, including a memory 11 and a processor 12, the memory 11 is used for storing one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor 12 to realize the method in the first aspect.
[0191] The processor described in the embodiment can be implemented by hardware, firmware, software or a combination thereof, and can use at least one of a circuit, a single or multiple application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), central processing units (CPUs), controllers, microcontrollers, microprocessors, and other physical, biological or chemical structures that can realize similar or equivalent functions to the above-mentioned processors, such as biological neurons, quantum computing units, DNA computing units, etc., so that the processor can execute part or all of the steps or any combination of the steps mentioned in the computer programs or methods of various embodiments of the present application.
[0192] Compared with the prior art, the above technical solution has the following beneficial effects: by integrating spatial transcriptome data, single-cell RNA sequencing data and pathological image data, a multi-scale spatial omics analysis system is constructed, single-cell resolution features are fused with spatial transcriptome data by using cross-modal semantic embedding technology to generate spatially enhanced expression profiles containing cell type deconvolution results; by hierarchical modeling of microcell interaction graphs, mesoscopic functional region graphs and macroscopic tissue structure graphs, spatial heterogeneity features containing tumor-immune adjacency strength, vascular invasion front spatial gradient and high-risk subregion topological relationship are extracted; finally, based on the graph neural network, molecular expression and spatial topological features are fused to output liver metastasis probability spatial heat map and key driving feature list; finally, a clinical prediction report of high-risk area positioning, metastasis probability scoring and treatment response prediction is generated. The above technical solution realizes multi-dimensional correlation analysis from gene expression to tissue morphology, provides a single-cell precision and spatial resolution solution for colorectal liver metastasis risk assessment, has advantages in prediction accuracy and clinical applicability, and has application value for guiding individualized treatment decisions.
[0193] Finally, it should be noted that the above embodiments have been described in the specification and drawings of the application, but this does not limit the patent protection scope of the application. Any equivalent structure or equivalent process replacement or modification based on the essential concept of the application, using the content described in the specification and drawings of the application, and directly or indirectly implementing the technical solutions of the above embodiments in other related technical fields, etc., are all included in the patent protection scope of the application.
Claims
1. A method for predicting colorectal cancer metastasis based on spatial omics, characterized by, The method comprises the following steps: Collecting original multi-omics data of original intestinal cancer tissue samples, the original multi-omics data comprising first spatial transcriptome data, first single-cell RNA sequencing data and first pathological image data; Performing modal alignment and quality control processing on the original multi-omics data to obtain preprocessed multi-omics data, the preprocessed multi-omics data comprising second spatial transcriptome data, second single-cell RNA sequencing data and second pathological image data; Performing cross-modal semantic embedding on the second spatial transcriptome data based on the second single-cell RNA sequencing data to generate a spatially enhanced expression profile, the cross-modal semantic embedding comprising single-cell-space feature alignment based on contrastive learning and probabilistic sparse signal completion, and the spatially enhanced expression profile comprising cell type deconvolution results and confidence scores of metastasis-related genes; Performing multi-scale graph construction on the second spatial transcriptome data and the second pathological image data to extract spatial heterogeneity features, the multi-scale graph construction comprising microscopic cell interaction graph, mesoscopic functional region graph and hierarchical modeling of macroscopic tissue structure, and the spatial heterogeneity features comprising tumor-immune adjacency strength, blood vessel invasion front spatial gradient and high-risk sub-region topological relationship; Inputting the spatially enhanced expression profile and the spatial heterogeneity features into a pre-trained metastasis risk prediction model to output a spatial heat map comprising liver metastasis probability and a list of key driving features; Generating a clinical prediction report of intestinal cancer metastasis risk based on the spatial heat map and the list of key driving features, the report comprising high-risk area positioning, metastasis probability score and treatment response prediction.
2. The method of predicting intestinal cancer metastasis based on spatial omics according to claim 1, characterized by, Performing modal alignment and quality control processing on the original multi-omics data to obtain preprocessed multi-omics data, comprising: Filtering spatial sites satisfying a preset quality threshold from the first spatial transcriptome data as second spatial transcriptome data, the preset quality threshold comprising a gene detection number threshold and an upper limit of mitochondrial gene proportion; Extracting cell sequencing data in the first single-cell RNA sequencing data matching the anatomical region of the second spatial transcriptome data as second single-cell RNA sequencing data, the anatomical region matching being achieved through tissue section coordinate mapping; Selecting an image region corresponding to the collection position of the second spatial transcriptome data from the first pathological image data as second pathological image data, the spatial deviation of the collection position in matching being not more than a preset alignment tolerance; Establishing a spatial correlation relationship among the second spatial transcriptome data, the second single-cell RNA sequencing data and the second pathological image data, so that the three have a unified coordinate reference system; Performing batch effect correction on the second spatial transcriptome data to eliminate technical variations among different samples; Filtering low-quality cells on the second single-cell RNA sequencing data to retain transcriptome data meeting cell integrity standards; Generating preprocessed multi-omics data, and storing a plurality of the preprocessed multi-omics data in a quality control database. 3.The method of predicting intestinal cancer metastasis based on spatial omics according to claim 1, characterized in that, Performing cross-modal semantic embedding on the second spatial transcriptome data based on the second single-cell RNA sequencing data to generate a spatially enhanced expression profile, comprising: extracting a cell type feature matrix from the second single-cell RNA sequencing data, the cell type feature matrix comprising a marker gene expression profile of each cell type; aligning each spatial site of the second spatial transcriptome data with the cell type feature matrix by contrastive learning, and calculating a similarity score of each spatial site with each cell type; performing probabilistic cell type deconvolution on the spatial sites based on the similarity score, and generating a cell type proportion matrix, the deconvolution being optimized by joint non-negative matrix factorization and Markov random field; performing sparse signal completion on lowly expressed genes of the second spatial transcriptome data, the completed signal being derived from the expression distribution prior of the same type of cells in the single-cell data; fusing the cell type proportion matrix and the completed gene expression data to generate a spatially enhanced expression profile, the spatially enhanced expression profile comprising cell composition information and gene expression confidence score of each spatial site.
4. The method of predicting intestinal cancer metastasis based on spatial omics according to claim 3, characterized by, performing probabilistic cell type deconvolution on the spatial sites based on the similarity score, and generating a cell type proportion matrix, the deconvolution being optimized by joint non-negative matrix factorization and Markov random field, comprising: constructing a non-negative matrix factorization objective function of spatial sites-cell types by taking the similarity score as an initial weight, and establishing a neighborhood relationship graph of Markov random field based on spatial coordinate information of the second spatial transcriptome data, the neighborhood relationship graph comprising topological connection relationship between spatial sites; introducing a spatial smoothing constraint term in the non-negative matrix factorization objective function, the spatial smoothing constraint term adjusting the cell type proportion difference between adjacent sites through Markov random field; updating the cell type proportion matrix iteratively by an alternating optimization algorithm, each iteration comprising: fixing the spatial constraint term, and optimizing the cell type proportion of non-negative matrix factorization by multiplication update rule, and fixing the cell type proportion, and optimizing spatial consistency by graph Laplacian regularization; terminating the optimization when the Frobenius norm change of the cell type proportion matrix of adjacent iterations is less than a preset convergence threshold; normalizing the optimized cell type proportion matrix to ensure that the sum of cell type proportions of each spatial site is 1; outputting the cell type proportion matrix, the rows of the cell type proportion matrix corresponding to spatial sites, the columns corresponding to cell types, and the element values being probabilistic proportion weights. 5.The method of predicting intestinal cancer metastasis based on spatial omics according to claim 1, characterized in that, performing multi-scale graph construction on the second spatial transcriptome data and the second pathological image data, and extracting spatial heterogeneity features, comprising: constructing a microscopic cell interaction network graph based on the single-cell resolution expression profile of the second spatial transcriptome data, the microscopic cell interaction network graph being generated by ligand-receptor pair co-expression analysis and cell adjacency relationship calculation; segmenting the second pathological image data into tissue regions, identifying tumor core region, invasion front region and interstitial region, and constructing a mesoscopic functional region atlas in combination with the expression features of the corresponding regions of the second spatial transcriptome data; integrating the tissue structure features of the second pathological image data at the whole slice level and the global expression pattern of the spatial transcriptome data to establish a macroscopic tissue metastasis trend prediction graph; extracting tumor-immune cell interaction intensity features from the microcosmic cell interaction network graph, including the spatial co-localization frequency of immune checkpoint molecule pairs; calculating the spatial gradient features of the vascular invasion front from the mesoscopic functional region map, including the decay coefficient of endothelial marker expression with distance; extracting high-risk sub-region topology features from the macroscopic tissue metastasis trend prediction map, including the spatial autocorrelation index of the microenvironment before metastasis; fusing the tumor-immune cell interaction intensity features, vascular invasion front spatial gradient features, and high-risk sub-region topology features into spatial heterogeneity features. 6.The method of predicting intestinal cancer metastasis based on spatial omics according to claim 5, characterized in that, integrating the tissue structure features of the second pathological image data at the whole slice level and the global expression pattern of the spatial transcriptome data to establish a macroscopic tissue metastasis trend prediction map, including: performing whole slice scanning on the second pathological image data to extract tissue structure features, including tumor gland arrangement directionality, interstitial fibrosis distribution density, and necrotic area spatial proportion; performing spatial autocorrelation analysis on the global expression pattern of the second spatial transcriptome data to calculate the spatial clustering index of each gene expression hotspot region, including the expression amount spatial distribution of epithelial-mesenchymal transition markers, angiogenic factors, and immune suppression-related molecules; constructing a multi-modal fusion model based on a graph convolutional neural network to align the tissue structure features and global expression pattern, the alignment process achieving weight distribution of pathological image features and transcriptome expression features through cross-modal attention mechanism, outputting a fused multi-modal feature matrix; establishing a macroscopic tissue metastasis trend prediction map in the multi-modal feature matrix, the macroscopic tissue metastasis trend prediction map simulating potential metastasis paths of tumor cells along tissue structure features and expression gradient features through a random walk algorithm; performing topological optimization on the macroscopic tissue metastasis trend prediction map to identify high-risk metastasis sub-regions, including calculating the spatial density of metastasis path intersection points and the consistency coefficient of metastasis direction; outputting the final macroscopic tissue metastasis trend prediction map, which includes a metastasis probability heat map, high-risk sub-region boundary annotation, and main metastasis path vector field.
7. The spatial omics-based colorectal cancer metastasis prediction method of claim 1, characterized in that, inputting the spatially enhanced expression profile and spatial heterogeneity features into a pre-trained metastasis risk prediction model to output a spatial heat map containing liver metastasis probability and a list of key driving features, including: performing feature standardization processing on the spatially enhanced expression profile to make the expression amounts of different genes comparable, including Z-score normalization and batch effect correction; performing dimension compression on the spatial heterogeneity features, retaining principal component features with a contribution rate exceeding a pre-set threshold using principal component analysis method; concatenating the standardized spatially enhanced expression profile and compressed spatial heterogeneity features to generate a fusion feature matrix; inputting the fusion feature matrix into a pre-trained graph neural network model, which aggregates multi-scale spatial information based on attention mechanism; Calculate the liver metastasis probability of each spatial site through the graph neural network model, and generate a spatial heat map based on the probability value, which maintains a spatial correspondence with the original tissue section; Adopt the gradient back propagation method to extract key driving features from the graph neural network model, including gene expression features and spatial topology features with the highest contribution to the prediction results; Sort the key driving features according to the feature contribution, and generate a key driving feature list containing feature names, contribution scores and biological interpretations; Store the spatial heat map and the key driving feature list in association, and establish a mapping relationship with the clinical pathological parameters.
8. A device for predicting colorectal cancer metastasis based on spatial omics, characterized in that, The device is suitable for the method of any one of claims 1 to 7, and the device comprises: A data acquisition module is configured to acquire original multi-omics data of an original intestinal cancer tissue sample, wherein the original multi-omics data comprises first spatial transcriptome data, first single-cell RNA sequencing data, and first pathological image data; A data processing module is configured to perform modal alignment and quality control processing on the original multi-omics data to obtain preprocessed multi-omics data, wherein the preprocessed multi-omics data comprises second spatial transcriptome data, second single-cell RNA sequencing data, and second pathological image data; A logical operation module is configured to perform cross-modal semantic embedding on the second spatial transcriptome data based on the second single-cell RNA sequencing data to generate a spatially enhanced expression profile, wherein the cross-modal semantic embedding comprises single-cell-space feature alignment based on contrastive learning and probabilistic sparse signal completion, and the spatially enhanced expression profile comprises cell type deconvolution results and confidence scores of metastasis-related genes; the second spatial transcriptome data and the second pathological image data are subjected to multi-scale graph construction to extract spatial heterogeneity features, wherein the multi-scale graph construction comprises hierarchical modeling of microscopic cell interaction graphs, mesoscopic functional region graphs, and macroscopic tissue structures, and the spatial heterogeneity features comprise tumor-immune adjacency strength, blood vessel invasion front spatial gradient, and high-risk sub-region topological relationship; the spatially enhanced expression profile and the spatial heterogeneity features are input into a pre-trained metastasis risk prediction model to output a spatial heat map containing liver metastasis probability and a key driving feature list; A report generation module is configured to generate a clinical prediction report of intestinal cancer metastasis risk based on the spatial heat map and the key driving feature list, wherein the report comprises high-risk area positioning, metastasis probability scoring, and treatment response prediction.
9. A computer readable storage medium having stored thereon computer program instructions, wherein, The computer program instructions, when executed by a processor, implement the method of any one of claims 1 to 7.
10. An electronic device comprising a memory and a processor, characterized in that, The memory is configured to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Colorectal cancer space transcriptome prediction method and device based on deep learning and storage medium
CN119887761A
Multi-omics pathological analysis system and method for colorectal cancer liver metastasis risk prediction
CN120048529A
Cell type and cell abundance identification method and system based on cross-modal training
CN120148030A
Multimodal foundation model for patient risk stratification
US20250166827A1
Cited By
Operation area lesion image segmentation method based on multi-scale dynamic segmentation kernel
CN121437540A
Breast cancer detection method and system based on morphological image and space transcriptome cross-graph collaborative learning
CN121582228A
Breast cancer detection method and system based on morphological image and spatial transcriptome cross-map collaborative learning
CN121582228B
Anatomical atlas-based tumor recurrence and metastasis risk prediction method and system
CN121687519A
Tumor recurrence and metastasis risk prediction method and system based on anatomical atlas
CN121687519B