A disease target network construction method fusing pathological images and spatial transcriptome

By fusing pathological images with spatial transcriptome data, a disease target network was constructed, which solved the problem of inaccurate disease regional localization in existing technologies, achieved systematic analysis of multi-level molecular mechanisms, and improved the analytical capabilities of spatial omics data and the accuracy of pathological diagnosis.

CN122117003APending Publication Date: 2026-05-29INNOVATION CENTER OF YANGTZE RIVER DELTA ZHEJIANG UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INNOVATION CENTER OF YANGTZE RIVER DELTA ZHEJIANG UNIVERSITY
Filing Date
2026-04-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing spatial transcriptomics and single-cell transcriptomics analysis methods struggle to accurately locate disease-related spatial regions, and pathological or imaging information is not effectively integrated into the molecular analysis process, resulting in insufficient accuracy in disease diagnosis and pathological assessment.

Method used

By fusing pathological images and spatial transcriptome data, regions of interest are defined, and single-cell and spatial transcriptome data are integrated to construct disease-related spatial transcriptome molecular networks. Pathological or imaging data are used as the starting point for spatial analysis, and spatial deconvolution or cell mapping algorithms are combined for data registration and mapping. Cell type and gene enrichment scores are calculated to construct gene-gene or cell-gene spatial transcriptome molecular networks.

Benefits of technology

It significantly improves the biological accuracy and pathological relevance of spatial analysis results, achieves unified modeling of cell composition, gene expression and molecular interaction networks, enhances the comprehensive analytical capabilities of spatial omics data, has good multimodal adaptability and robustness, and can identify key cell types and molecular networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122117003A_ABST
    Figure CN122117003A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of biomedical data analysis and spatial omics data processing, and particularly relates to a disease target network construction method fusing pathological images and spatial transcriptome. Based on automatic or manual definition of a region of interest according to pathological characteristics, single-cell data and spatial transcriptome data are integrated; the spatial enrichment degree of different cell types and genes is quantitatively scored; and finally, a spatial co-localization network of genes and cells is constructed in the region of interest. The method can be compatible with various pathological imaging methods, and is suitable for irregularly shaped and significantly spatially heterogeneous disease tissues, overcoming the limitations of traditional spatial analysis methods guided by transcriptome characteristics in disease region positioning, and realizing systematic analysis from spatial pathology positioning to cell, gene and molecular correlation levels. The method is suitable for spatial mechanism research of complex diseases such as cardiovascular diseases, tumors and neurodegenerative diseases, and has high biological interpretation value and practical application significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedical data analysis and spatial omics data processing technology, and in particular relates to a method for constructing a disease target network that integrates pathological images and spatial transcriptomics. Background Technology

[0002] Single-cell RNA sequencing (scRNA-seq) technology enables quantitative analysis of RNA expression in tissues at single-cell resolution, revealing the heterogeneity of cell populations at the transcriptional level. However, scRNA-seq cannot preserve the original spatial structure information of tissues. Spatial transcriptomics, on the other hand, can detect gene expression distribution while preserving the spatial structure of tissues, thus enabling joint analysis of gene expression and spatial location information. This has shown significant application value in the study of spatial heterogeneity in disease tissues.

[0003] However, existing analytical methods typically rely primarily on gene expression characteristics or cellular composition similarity to delineate and interpret spatial regions. Disease-related regions are often indirectly inferred through transcriptomic features rather than being directly defined based on pathological or imaging characteristics. For disease tissues with irregular lesion morphology, involving multiple tissue structures, or exhibiting significant spatial heterogeneity, the aforementioned analytical strategies struggle to accurately correspond to actual pathological lesions.

[0004] Histopathological staining, immunofluorescence imaging, and medical imaging techniques (such as magnetic resonance imaging) can directly reflect disease-related structural abnormalities and microenvironmental characteristics at the spatial level, and are commonly used in disease diagnosis and pathological evaluation. However, in the current technological system, pathological and imaging information are mostly used as auxiliary verification methods and have not yet been systematically incorporated into the analysis process of spatial transcriptomics and single-cell data. This makes it difficult to achieve a direct correlation between pathological regions and their corresponding cellular composition, gene expression characteristics, and molecular interaction networks, which limits the application value of spatial transcriptomics data in disease mechanism research and clinical translation. Summary of the Invention

[0005] To address the challenges of accurately locating disease-related spatial regions and effectively integrating pathological or imaging information into molecular analysis processes in existing spatial and single-cell transcriptomic analysis methods, this invention provides a method for constructing a disease target network that integrates pathological images and spatial transcriptomics.

[0006] This method uses pathological or imaging data as the starting point for spatial analysis. By registering pathological images with spatial transcriptome data, it defines disease-related regions of interest and integrates single-cell transcriptome and spatial transcriptome data within these regions. It then performs joint analysis on cell type distribution, gene expression characteristics, and molecular interaction networks to construct a disease-related spatial transcriptome molecular network.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0008] In a first aspect, the present invention provides a method for constructing a disease target network that integrates pathological images and spatial transcriptomics, comprising the following steps:

[0009] Tissue slide data containing pathological or imaging information and corresponding spatial transcriptome sequencing data are acquired. The pathological images are then cropped, scaled, and coordinate-transformed to establish a correspondence between the pixel coordinates of the pathological images and the coordinates of the spatial transcriptome capture points. The pathological images include, but are not limited to, immunofluorescence images, histopathological staining images, or medical imaging data. The spatial transcriptome data is stored in the form of a spatial expression matrix, containing spatial location and gene expression information.

[0010] Cell type annotation is performed on the single-cell data of the tissue sample to obtain single-cell objects containing cell expression matrices and cell annotation information, wherein the cell annotation information includes cell type, state, or origin information. Using existing spatial deconvolution or cell mapping algorithms, the single-cell annotation information is mapped to the spatial expression sub-matrix, thereby obtaining the cell composition information corresponding to each spatial location within the Region of Interest (ROIs).

[0011] Based on the registered pathological images, regions of interest (ROIs) are defined in the images according to preset pathological feature thresholds or manual annotation methods, and a set of spatial transcriptome capture point barcodes corresponding to the ROIs is collected. The ROIs may correspond to disease-related lesion areas, marginal areas, or other spatial regions with specific pathological features. Within the ROIs, enrichment scores for different cell types and genes are calculated. Based on spatial co-localization relationships and a predefined molecular interaction database, a gene-gene or cell-gene spatial transcriptome molecular network is constructed to quantitatively characterize the association strength in specific pathological spatial microenvironments.

[0012] The generated spatial molecular network can be compared and analyzed in different pathological regions, providing quantitative analysis results for spatial mechanism research based on disease pathological regions.

[0013] Furthermore, the definition of the region of interest according to the preset pathological feature threshold or manual annotation method is specifically as follows:

[0014] (a) Perform intensity normalization processing on the pathological image, select spatial capture points with enhanced pathological signals as spatial center points according to a preset threshold, and construct the region of interest based on the spatial center points;

[0015] (b) Define the region of interest in the pathological image by manually drawing.

[0016] Furthermore, the spatial deconvolution or cell mapping algorithm includes at least one of cell2location, CellTrek, or Tangram.

[0017] Furthermore, within the region of interest, the specific method for calculating the enrichment scores of different cell types and genes within the region of interest is as follows:

[0018] (S51) Method for calculating the enrichment score of different cell types in the region of interest: For each cell type, count the number of cells mapped to the region of interest, and normalize it with the total number of cells of that cell type in the overall sample, and use the ratio as the enrichment score of that cell type.

[0019] (S52) The calculation of enrichment scores for different genes within the region of interest includes any of the following methods:

[0020] (a) When the region of interest is automatically identified, the spatial correlation score is calculated based on the Pearson correlation coefficient: the expression level of each gene in different regions of interest is correlated with the pathological signal intensity of the corresponding region, and the Pearson correlation coefficient between the expression value of the gene in each region of interest and the corresponding pathological signal intensity is used as the enrichment score of the gene.

[0021] (b) When the region of interest is determined manually or semi-automatically, the Mann-Whitney rank-sum test is used to divide the region into two groups based on the number of spatial capture points inside and outside the region of interest set. The U statistic or Z statistic of the gene between the two groups is calculated, and the U value, Z value or the corresponding p value is used as the enrichment score.

[0022] (S53) Sort the genes according to the spatial association score, and select the preset number of genes with the highest score and the preset number of genes with the lowest score for subsequent analysis.

[0023] Furthermore, the disease target network uses proteins as nodes, protein-protein interactions as edges, and the spatial co-expression coefficient as the edge weight. The spatial co-expression coefficient is calculated using the Pearson correlation coefficient.

[0024] Secondly, the present invention provides a system for implementing the above-mentioned method for constructing a disease target network that integrates pathological images and spatial transcriptomics, comprising:

[0025] The data acquisition and registration module is used to acquire spatial transcriptome data and corresponding pathological images, and to establish the spatial correspondence between the two.

[0026] The Region of Interest (ROI) definition module is used to define the ROI based on pathological images and obtain the corresponding spatial transcriptome capture points;

[0027] The single-cell processing module is used to process single-cell transcriptome data and obtain annotated single-cell objects.

[0028] The spatial mapping module is used to map single-cell objects to spatial transcriptome capture points to obtain cellular composition information;

[0029] The enrichment analysis module is used to calculate the enrichment scores of cell types and genes in regions of interest, and to screen cells and genes for analysis.

[0030] The network construction module is used to calculate the spatial co-expression coefficient based on the expression information and protein-protein interaction relationships of selected genes, and to construct a disease target network.

[0031] Thirdly, the present invention also provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores one or more programs, which can be executed by one or more processors to realize the above-described method for constructing a disease target network that integrates pathological images and spatial transcriptomes.

[0032] The present invention has the following beneficial effects:

[0033] (1) This invention uses pathological images or medical imaging information as the prior basis for spatial analysis, and changes the definition of disease-related regions from the traditional "inference based on transcriptome features" to "direct definition based on real pathological signals", which significantly improves the biological accuracy and pathological relevance of spatial analysis results and overcomes the problem of inaccurate correspondence between spatial regions and real lesions in existing methods.

[0034] (2) This invention constructs a multimodal joint analysis framework that integrates pathological images, spatial transcriptome and single-cell transcriptome data, realizing unified modeling of cell composition, gene expression and molecular interaction networks, thereby enabling systematic analysis of disease-related multi-level molecular mechanisms in the same spatial background and improving the comprehensive analysis capability of spatial omics data.

[0035] (3) The present invention has good multimodal adaptability and versatility, and can be applied to various types of pathological or imaging data such as immunofluorescence, histopathological staining and magnetic resonance imaging. It can stably identify key cell types and molecular networks under different imaging signal guidance.

[0036] (4) This invention does not rely on complete cell reference information. Even if some cell types are missing in single-cell data, the corresponding molecular functional modules can still be reconstructed based on spatial expression correlation, demonstrating strong robustness and the ability to infer unknown cell components.

[0037] (5) The method of the present invention has good scalability and can be combined with the continuously developing image analysis technology to dynamically analyze the changes in spatial molecular networks during the occurrence and development of diseases, and has important potential for scientific research and clinical translational applications. Attached Figure Description

[0038] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0039] Figure 1 The flowchart illustrates a method for constructing a disease target network that integrates pathological images and spatial transcriptomics, as provided by this invention.

[0040] Figure 2 This provides multimodal single-cell-space transcriptome data of the brains of Alzheimer's disease mice; among which, Figure 2 In the diagram, A represents the flowchart for constructing a multimodal single-cell-space transcriptome dataset of Alzheimer's disease (AD) mouse brains; Figure 2 B in the figure represents a visualization of multimodal single-cell and spatial transcriptome data of the AD mouse brain.

[0041] Figure 3 Region of interest identification for different modalities; among which, Figure 3 In this context, A represents the accumulation of Aβ in each 5×FAD sample and the identification of regions of interest based on Aβ. Figure 3 B in the model identifies the FA difference region and the region of interest based on the FA difference in each 5×FAD sample. Figure 3 In this context, C represents the cell type score based on Aβ. Figure 3 In this context, D represents the cell type score based on FA differences.

[0042] Figure 4 SpaceNet reveals the spatial network of Aβ-guided AD mice; Figure 4 In this context, A represents a spatial protein network associated with Aβ pathology. Figure 4 B in the diagram represents a spatially co-expressed network of immune-related proteins. Figure 4C in the image represents the immunofluorescence staining image of Aβ (green) and GFAP (yellow) in a 5×FAD sample.

[0043] Figure 5 SpaceNet reveals the spatial network of AD mice guided by magnetic resonance imaging (MRI); among which, Figure 5 In this context, A represents a spatial protein network associated with pathological conditions observed in magnetic resonance imaging. Figure 5 B in the diagram represents a spatially co-expressed network of immune-related proteins. Figure 5 C in the diagram represents a spatial protein network guided by Aβ pathology and magnetic resonance imaging pathology. Detailed Implementation

[0044] The following will describe the concept and technical effects of the present invention clearly and completely with reference to embodiments, so as to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention.

[0045] This invention provides a method for constructing a disease target network that integrates pathological images and spatial transcriptomics. The flowchart of the method is shown below. Figure 1 As shown, it includes the following steps:

[0046] 1. Data Acquisition, Preprocessing, and Spatial Registration: Acquire spatial transcriptome sequencing data of the target tissue, as well as pathological or imaging data corresponding to the spatial transcriptome sections. Pathological or imaging data includes, but is not limited to, immunofluorescence images, histopathological staining images, magnetic resonance imaging data, or other medical imaging data. Spatial transcriptome data is stored in the form of a spatial expression matrix, containing spatial location and gene expression information.

[0047] Preprocessing of pathological images and spatial transcriptome data based on prior knowledge includes, but is not limited to, image cropping, scale unification, and coordinate system transformation, to ensure a one-to-one correspondence between the pixel coordinates of the pathological images and the coordinates of the spatial transcriptome capture points.

[0048] For example, for processing pathological images stained with HE, the Space Ranger provided by 10x Genomics can be used to spatially integrate the images with spatial transcriptome data.

[0049] 2. Pathology-guided spatial region definition: Based on the registered pathological image, disease-related spatial regions are identified in the pathological image using automatic identification, manual or semi-manual annotation methods according to a preset pathological signal intensity threshold. These regions are then mapped to a spatial transcriptome coordinate system to obtain the corresponding set of spatial capture points. The specific steps are as follows:

[0050] (2.1) Based on the registered pathological images, the correspondence between the images and the coordinates of the spatial transcriptome capture points is included, providing a coordinate transformation basis for subsequent region mapping.

[0051] (2.2) For the definition of regions of interest with reliable quantitative biomarkers, an automated region identification method based on pathological signal intensity is used. This method first performs intensity normalization processing on the quantitative pathological image to improve the objectivity of spatial region definition and cross-sample comparability. The normalization formula is:

[0052]

[0053] in, This represents the normalized fluorescence intensity of the i-th spot. This represents the original fluorescence intensity of the i-th spot. This represents the maximum intensity value of all spots in the image.

[0054] Based on a preset threshold (ranging from 1 to 5), spatial capture points with significantly enhanced pathological signals are selected. Using these points as spatial centers, and combining the pathological signal intensity with a custom search radius parameter, disease-related regions of interest are automatically constructed. The calculation formula for each disease-related region of interest is as follows:

[0055]

[0056] in, Indicates each region of interest related to a disease. The normalized fluorescence intensity represents the center point of the space, and search_radius represents a custom search radius parameter.

[0057] The expression for the Region of Interest (ROI) is as follows:

[0058] .

[0059] For cases where the lesion area has an irregular shape or involves multiple pathological areas, the region of interest is defined by manual annotation. Multiple regions of interest can be defined in the pathological image by drawing polygons using online tools (such as SpaceNet), and then analyzed separately.

[0060] 3. Single-cell data integration and spatial mapping: Obtain single-cell or single-nucleus transcriptome sequencing data corresponding to the tissue sample, perform quality control, standardization and cell type annotation on the single-cell data, and obtain single-cell objects containing cell expression matrix and cell annotation information.

[0061] Seurat is preferred for processing. Quality control aims to screen out low-expression genes and low-quality cells, including but not limited to: removing cells with excessively low or high gene counts (e.g., below 200 or above 6000 genes), removing cells with abnormally high mitochondrial gene expression (e.g., exceeding 10%-20%), and filtering out low-abundance genes expressed in a small number of cells (e.g., genes expressed in fewer than 3 cells). Standardization is used to eliminate the influence of sequencing depth differences and technical noise on expression levels, ensuring comparability between different cell types. Cell type annotation is based on cluster analysis results combined with known marker gene expression characteristics. Reference databases or automated annotation tools (such as SingleR, CellTypist, etc.) can be used to assist in cell type analysis to obtain single-cell data objects containing cell expression matrices and cell type labels.

[0062] By employing spatial deconvolution or cell mapping algorithms such as cell2location, CellTrek, or Tangram included in SpaceNet, single-cell objects are mapped to spatial transcriptome capture points, thereby obtaining the cell composition ratio or cell type assignment results at each spatial location within the region of interest.

[0063] 4. Disease-related spatial cellular composition and gene analysis: After identifying the region of interest, the degree of association between different cell types and genes and disease-related spatial regions is quantitatively assessed based on the mapping results from single-cell data to spatial transcriptomes.

[0064] 4.1 For any cell type, for each cell type, count the number of cells mapped to the region of interest, and normalize this number with the total number of cells of that cell type in the overall sample. Use this proportion as the enrichment score for that cell type. The enrichment score is as follows: .

[0065] Defined as:

[0066]

[0067] in, This indicates the number of cells of a specific cell type within the set of regions of interest. This indicates the total number of cells of this cell type in the entire tissue sample.

[0068] 4.2 Quantitative scoring of gene enrichment levels

[0069] (a) For a region of interest determined manually or semi-automatically, the enrichment score for any gene j is: . Defined as:

[0070]

[0071] in, This represents the average expression level of gene j in the region of interest m. This represents the average expression level of gene j across all regions of interest. This represents the normalized fluorescence intensity of the region of interest m, while This represents the average fluorescence intensity across all regions of interest.

[0072] (b1) For regions of interest determined manually or semi-automatically, all capture points are divided into two groups of samples based on whether they are within the region of interest. The Mann-Whitney U test is used to calculate the statistic, and the smaller of the two values ​​is taken as the final test statistic.

[0073] The calculation formula is as follows: ;

[0074] Where n1 represents the number of spatial capture points within the region of interest, and n2 represents the number of spatial capture points outside the region of interest; for the gene j to be analyzed, all spatial capture points are sorted according to the expression level of gene j, R1 represents the sum of all indices within the region of interest, and R2 represents the sum of all indices outside the region of interest.

[0075] After obtaining the statistic U, consult the critical value table for the rank-sum test based on the sample sizes n1 and n2 to determine the corresponding probability value. .

[0076] When p < 0.05, the gene is determined to be significantly enriched or significantly missing in the region of interest; or a significance threshold can be preset, or genes can be directly sorted according to the U value to screen for spatially specific expression genes.

[0077] (b2) For any gene j, calculate the standardized Z-statistic of the difference in its expression between the region of interest and the region outside the region, using the following formula:

[0078]

[0079] Where n1 and n2 represent the number of spatial capture points within and outside the region of interest, respectively, and R1 represents the rank-sum statistic of gene j's expression level within the region of interest.

[0080] when Time indicates gene High expression levels were enriched in regions of interest. Time indicates gene The expression is lost in the region of interest. In one embodiment of the invention, the top 100 and bottom 100 genes are selected for subsequent analysis.

[0081] 5. Spatial protein interaction analysis

[0082] First, based on the spatial transcriptome data from step 1, the expression information of selected genes (the top 100 genes and the 100th gene in each spatial capture point) is obtained, and known or predicted protein-protein interaction relationships between the corresponding proteins of the genes are obtained from a protein interaction database (e.g., the STRING database). For gene pairs with protein-protein interaction relationships, the spatial co-expression coefficient is calculated based on their expression distribution at the spatial capture points, and a spatial protein-protein interaction network is constructed accordingly.

[0083] For each protein-protein interaction gene pair selected The co-expression coefficient in the spatial dimension was calculated using the Pearson correlation coefficient. The calculation formula is as follows:

[0084] ;

[0085] in, This represents the co-expression coefficient between gene i and gene j. and These represent their average expression levels at spatial point m, while and This indicates their average expression level across all points.

[0086] In the spatial protein interaction network, network nodes represent proteins, network edges represent protein interaction relationships, and edge weights are determined by spatial co-expression coefficients. The final spatial protein interaction network is constructed by incorporating spatial covariance information of gene expression on the basis of traditional protein interaction relationships, realizing quantitative characterization of protein interaction relationships in the spatial dimension, thereby more accurately identifying functional modules and core molecules related to disease regions.

[0087] The present invention also provides a system for implementing the above-mentioned method for constructing a disease target network that integrates pathological images and spatial transcriptomics, comprising:

[0088] The data acquisition and registration module is used to acquire spatial transcriptome data and corresponding pathological images, and to establish the spatial correspondence between the two.

[0089] The Region of Interest (ROI) definition module is used to define the ROI based on pathological images and obtain the corresponding spatial transcriptome capture points;

[0090] The single-cell processing module is used to process single-cell transcriptome data and obtain annotated single-cell objects.

[0091] The spatial mapping module is used to map single-cell objects to spatial transcriptome capture points to obtain cellular composition information;

[0092] The enrichment analysis module is used to calculate the enrichment scores of cell types and genes in regions of interest, and to screen cells and genes for analysis.

[0093] The network construction module is used to calculate the spatial co-expression coefficient based on the expression information and protein-protein interaction relationships of selected genes, and to construct a disease target network.

[0094] The present invention also provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the above-described method for constructing a disease target network that integrates pathological images and spatial transcriptomes.

[0095] Example 1

[0096] A disease target network was constructed based on a mouse brain dataset of Alzheimer's disease (AD).

[0097] This invention constructs a dataset of AD mouse brains, including magnetic resonance imaging (MRI), immunofluorescence images (containing Aβ staining results), spatially resolved transcriptome sequencing data, and single-cell transcriptome sequencing data of the mouse brains. Figure 2 The data was stored in the GEO database: GSE253570. Immunofluorescence images and spatially resolved transcriptome sequencing data were derived from the same tissue section, ensuring a precise spatial correspondence between image and transcriptome information; the MRI data was spatially aligned with the spatially resolved transcriptome sequencing data.

[0098] To ensure unbiased analysis, this study employed two independent biomarker-guided strategies for defining regions of interest: one based on Aβ pathological deposition, and the other based on microstructural changes reflected on MRI. For example... Figure 3As shown in AB, for Aβ-stained region of interest (ROI) selection, the Aβ fluorescence signal intensity is first normalized, and the normalized intensity is mapped to each spot in the spatial transcriptome. Then, spots with normalized intensities higher than a preset threshold (greater than 0 in this case) are defined as candidate center points, and local circular regions are constructed centered on their spatial coordinates. The radius of these circular regions is determined by the original signal intensity and the user-defined search parameter (set to 1.5 in this case). For MRI-guided ROI selection, fractional anisotropy (FA) maps are calculated based on diffusion-weighted imaging data from MRI to characterize the integrity of brain tissue microstructure. Regions showing significant differences in FA between the 5×FAD and WT groups are screened, and these regions are mapped to the empty transcriptome coordinates, defined as ROI regions. Irregular regions can be selected using relevant online tools (e.g., SpaceNet).

[0099] This method demonstrates good robustness under various image-guided conditions. Regardless of the image features used for guidance, the identified key cell types are highly similar, primarily enriched in cortical excitatory neuronal populations, including glutamatergic neurons in layers L2-L3, L4-L5, and L6 (Isocortex L2-L3, Isocortex L4-L5, Isocortex L6), while also covering neurons in the olfactory (OLF) and hippocampal (HIP) regions. Figure 3 (CD in the middle).

[0100] In this embodiment, the scores of all genes in the region of interest are first calculated: for automatically identified regions based on Aβ staining, the Pearson correlation coefficient is used for scoring; for manually selected regions guided by MRI, the Mann-Whitney U test is used, and the Z-statistic is used for ranking in the U test. Finally, the top 100 and bottom 100 genes with the highest scores are selected as the candidate disease-related gene set. This gene set is then input into a protein-protein interaction (PPI) database to construct a network, and the Pearson correlation coefficient is used to quantify the spatial expression pattern correlation between proteins (or genes). In the final constructed spatial network, nodes represent proteins (or genes), and edges represent the correlation strength between proteins (or genes) or between proteins (or genes) and cell types, with darker colors indicating stronger correlations. Figure 4 A and Figure 5 (A in the middle). Figure 4 B and Figure 5B in the diagram specifically illustrates an immune response-related network characterized by spatial co-expression, where nodes represent proteins (or genes) and are colored red (upregulated genes) or blue (downregulated genes). Edges are colored from red (spatial co-expression) to blue (spatial negative correlation).

[0101] like Figure 4 As shown in Figure A, the disease target network constructed under Aβ signaling guidance mainly focuses on the aforementioned cortical excitatory neuronal population. Notably, these modules are also closely related to astrocytes and immune response signals. The astrocyte-associated module, with GfAP as its core node, shows a significant positive correlation between its spatial expression level and Aβ deposition intensity, and is primarily enriched in the Astro1 subset, suggesting it is a representative marker molecule of Aβ-associated reactive astrocytes. Astro1 represents the Aβ-associated reactive astrocyte population. Immunofluorescence staining further confirmed the presence of GFAP. + Astrocytes exhibit spatial aggregation around Aβ-detected plaques. Figure 4 (C in the equation), thus verifying the spatial prediction results of SpaceNet.

[0102] The immune response module contains several classic immune and inflammation-related genes, including Ctss, C1qa, C1qc, Cst3, and Trem2. These genes exhibit significant co-expression patterns spatially and form a dense functional subnetwork connected by high-weighted edges. Trem2, as a key regulatory node, together with molecules such as Apoe, constitutes the known TREM2-APOE signaling axis, a pathway that plays a crucial role in microglia-mediated Aβ clearance and immune responses. Notably, although microglia were absent from the snRNA-seq reference data used for anatomical analysis, SpaceNet reconstructed the microglia-related gene module based solely on spatial correlation, demonstrating its ability to infer missing cellular components.

[0103] like Figure 5 As shown in Figure A, the disease target network constructed under MRI signal guidance also indicates that the neuronal population in the cerebral cortex is significantly affected, exhibiting different molecular and cellular characteristics compared to the Aβ-guided condition. SpaceNet identified a spatially highly colocalized myelin-related protein module ( Figure 5The module B contains key proteins such as MOG, MBP, MOBP, MAG, PLP1, and PLLP. The genes corresponding to these proteins are mainly involved in neuronal myelination and maintenance, and show an overall downregulation trend in the FA differential region, suggesting oligodendrocyte-mediated impairment of myelin structural integrity. At the cell type level, the aforementioned myelin-related genes are significantly enriched in the OLG1 subset, indicating that this subset is the main myelin-forming cell population and is damaged in the disease. Under MRI guidance, SpaceNet systematically analyzed cell type-specific responses related to brain microstructural changes at the spatial molecular network level, revealing the synergistic relationship between neuronal dysfunction and oligodendrocyte myelin damage. Compared with Aβ-guided conditions, the molecular network revealed by MRI guidance focuses more on changes in neurofibrillary structure and myelin integrity, demonstrating the complementarity of different imaging signals in spatial disease analysis. Figure 5 (C in the middle).

[0104] By introducing different types of pathological and imaging signals as guiding conditions within the same analytical framework, the method of this invention can construct corresponding spatial disease networks for different disease characteristics. This ensures the stability of the key cell type identification results and reflects the differential molecular and cellular changes corresponding to different signal sources, thereby verifying the universality and reliability of the method under multimodal imaging guidance conditions.

[0105] The above description of specific embodiments of the present invention does not limit the present invention. Those skilled in the art can make various changes or modifications based on the present invention, and as long as they do not depart from the spirit of the present invention, they should all fall within the scope of the appended claims.

Claims

1. A method for constructing a disease target network integrating pathological images and spatial transcriptomics, characterized in that, Includes the following steps: S1. Obtain spatial transcriptome sequencing data and corresponding pathological images of the target tissue, and preprocess the pathological images and spatial transcriptome sequencing data to establish the correspondence between the pixel coordinates of the pathological images and the coordinates of the spatial transcriptome capture points. S2. Define the region of interest in the pathological image by automatic identification, manual or semi-manual annotation according to the preset pathological signal intensity threshold, and collect the set of spatial transcriptome capture points corresponding to the region of interest; S3. Obtain single-cell or single-nucleus transcriptome sequencing data corresponding to the target tissue, and perform quality control, standardization and cell type annotation to obtain a single-cell object containing cell expression matrix and cell annotation information. S4. Using spatial deconvolution or cell mapping algorithms, single-cell objects are mapped to spatial transcriptome capture points to obtain the cell composition ratio and cell type results at each spatial location within the region of interest. S5. Calculate the enrichment scores of different cell types and genes in the region of interest, and select the genes for subsequent analysis. S6. Obtain the expression information of the selected gene at each spatial capture point and the interaction relationship of its corresponding protein. For gene pairs with protein interaction relationships, calculate the spatial co-expression coefficient based on their expression distribution at the spatial capture points in the region of interest, and construct a disease target network based on the spatial co-expression coefficient.

2. The method for constructing a disease target network according to claim 1, characterized in that, In step S1, the pathological image is an immunofluorescence image, a histopathological staining image, or magnetic resonance imaging data.

3. The method for constructing a disease target network according to claim 1, characterized in that, In step S1, the preprocessing specifically involves: cropping the pathological image, unifying its scale, and transforming its coordinate system to establish a correspondence between the pixel coordinates of the pathological image and the coordinates of the spatial transcriptome capture points.

4. The method for constructing a disease target network according to claim 1, characterized in that, In step S2, the method for defining the region of interest in the pathological image includes at least one of the following: (a) Perform intensity normalization processing on the pathological image, select spatial capture points with enhanced pathological signals as spatial center points according to a preset threshold, and construct the region of interest based on the spatial center points; (b) Define the region of interest in the pathological image by manually drawing.

5. The method for constructing a disease target network according to claim 1, characterized in that, In step S4, the spatial deconvolution or cell mapping algorithm includes at least one of cell2location, CellTrek, or Tangram.

6. The method for constructing a disease target network according to claim 1, characterized in that, Step S5 specifically includes: (S51) Method for calculating the enrichment score of different cell types in the region of interest: For each cell type, count the number of cells mapped to the region of interest, and compare it with the total number of cells of that cell type in the overall sample, using the ratio as the enrichment score of that cell type; (S52) Calculate the enrichment scores of different genes within the region of interest, including any of the following methods: (a) When the region of interest is automatically identified, the spatial correlation score is calculated based on the Pearson correlation coefficient: the expression level of each gene in different regions of interest is correlated with the pathological signal intensity of the corresponding region, and the Pearson correlation coefficient between the expression value of the gene in each region of interest and the corresponding pathological signal intensity is used as the enrichment score of the gene. (b) When the region of interest is determined manually or semi-automatically, the Mann-Whitney rank-sum test is used to divide the region into two groups according to the number of spatial capture points inside and outside the region of interest set. The U statistic or Z statistic of the gene between the two groups is calculated, and the U value, Z value or corresponding p value is used as the enrichment score. (S53) Sort the genes according to the enrichment score, and select the preset number of genes with the highest score and the preset number of genes with the lowest score for subsequent analysis.

7. The method for constructing a disease target network according to claim 1, characterized in that, In step S6, the spatial co-expression coefficient is calculated using the Pearson correlation coefficient.

8. The method for constructing a disease target network according to claim 1, characterized in that, In step S6, the disease target network uses proteins as nodes, protein-protein interactions as edges, and the spatial co-expression coefficient as edge weights.

9. A system for implementing the method for constructing a disease target network integrating pathological images and spatial transcriptomics as described in any one of claims 1-8, characterized in that, include: The data acquisition and registration module is used to acquire spatial transcriptome data and corresponding pathological images, and to establish the spatial correspondence between the two. The Region of Interest (ROI) definition module is used to define the ROI based on pathological images and obtain the corresponding spatial transcriptome capture points; The single-cell processing module is used to process single-cell transcriptome data and obtain annotated single-cell objects. The spatial mapping module is used to map single-cell objects to spatial transcriptome capture points to obtain cellular composition information; The enrichment analysis module is used to calculate the enrichment scores of cell types and genes in regions of interest, and to screen cells and genes for analysis. The network construction module is used to calculate the spatial co-expression coefficient based on the expression information and protein-protein interaction relationships of selected genes, and to construct a disease target network.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the method for constructing a disease target network that integrates pathological images and spatial transcriptomes as described in any one of claims 1-8.