Space transcriptome and space metabolome integration method based on deep learning

By integrating spatial transcriptomic and metabolomic data using deep learning methods, the technical differences and batch effects in data integration were resolved, achieving efficient integration across modalities and samples, and improving analytical accuracy and the precision of biological information.

CN120998315APending Publication Date: 2025-11-21ZHEJIANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511096999.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing methods for integrating spatial transcriptomics and spatial metabolomics data suffer from technical differences, batch effects, and challenges in quantifying modal contributions, resulting in insufficient accuracy of analytical results and a lack of effective cross-modal and cross-sample integration tools.

Method used

A deep learning-based spatial multi-omics data integration method is adopted. By aligning affine transformation matrices and unifying resolution, and combining variational autoencoders and conditional decoders, cross-modal and cross-sample integration of ST and SM data is achieved, generating joint embeddings and performing visualization analysis.

Benefits of technology

It improves the accuracy and reliability of data analysis, enhances the correlation between samples from different sources, provides more precise and comprehensive biological information, and supports tumor microenvironment research and the discovery of targets for immunometabolic therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998315A_ABST
    Figure CN120998315A_ABST
Patent Text Reader

Abstract

The invention discloses a space transcriptome and space metabolome integration method based on deep learning. The method comprises the following steps: firstly, acquiring original space transcriptomics data and original space metabonomics data of a biological tissue, and performing data preprocessing to obtain a preprocessed biological tissue data set; then, aligning space metabonomics data in the preprocessed biological tissue data set to space transcriptomics data, unifying data resolution, obtaining corrected space metabonomics data, and updating the biological tissue data set; and finally, generating to-be-integrated sample data from the newest biological tissue data set, inputting the to-be-integrated sample data into the spatial multi-omics data integration model, and outputting final joint embedding by the model, so that cross-modal and cross-sample effective integration of the data can be realized. According to the method, the problems of form and resolution inconsistency and batch effect caused by technical difference of ST and SM data are solved, and high-precision and interpretable spatial multi-omics integration analysis is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of spatial multi-omics data analysis, and specifically relates to a data integration method based on deep learning for spatial transcriptome and spatial metabolome. BACKGROUND

[0002] SM technology measures the intensity of metabolites and their spatial location in tissues through mass spectrometry imaging (MSI) technology. The combination of spatial transcriptomics (ST) data and spatial metabolomics (SM) data not only enables comprehensive characterization of the dynamic changes in the tissue microenvironment, but also reveals the mechanisms of metabolic reprogramming in the tumor microenvironment. In particular, in cancer research, this integration provides a unique perspective for in-depth understanding of tumor progression and treatment response.

[0003] However, the integration of ST and SM data faces many challenges. First, the technical differences are an important issue. Typically, ST and SM data come from adjacent tissue sections, which leads to differences in spatial morphology and resolution between the two types of data. In addition, ST technology is based on sequencing technology, while SM relies on mass spectrometry imaging, which results in significant differences in data morphology, resolution, and distribution. Second, batch effects are also a problem that must be overcome when integrating across samples. Technical batch and biological sample heterogeneity can cause data bias, affecting the accuracy of analysis results. Third, how to quantify the contribution of different modalities to the analysis results is still a major difficulty in current methods. Existing integration methods often cannot clearly assess the specific contribution of ST and SM to the analysis, thereby limiting the in-depth explanation of biological mechanisms.

[0004] Currently, existing multi-omics integration methods mainly focus on single-cell data or the integration of spatial proteomics (SP) and spatial epigenomics (SE) data, but there are still no specialized integration tools for ST and SM data. Therefore, it is particularly important to develop an ST and SM data analysis method that can not only integrate across modalities, but also integrate across samples. SUMMARY

[0005] In order to solve the problems and needs in the background art, the present application provides a spatial transcriptome and spatial metabolome integration method based on deep learning, named SpatialMETA. The present application realizes the cross-modal and cross-sample integration of ST and SM data through a spatial multi-omics data integration model, can accurately capture the spatially related ST-SM patterns, and supports interactive visualization analysis downstream, is suitable for tumor microenvironment, neuroscience and developmental biology research, and provides an innovative tool for tumor microenvironment research and immune metabolic therapy target discovery. This not only has important significance for accurately depicting the spatial transcriptomic and spatial metabolomic characteristics in the tissue microenvironment, but also provides a more in-depth analysis scheme for understanding complex biological processes.

[0006] The technical scheme of the present application comprises the following steps:

[0007] One, a spatial transcriptome and spatial metabolome integration method based on deep learning

[0008] Step 1: Obtain the original spatial transcriptomic data and the original spatial metabolomic data of the biological tissue and perform data preprocessing to obtain the preprocessed biological tissue data set;

[0009] Step 2: Align the spatial metabolomic data of the same tissue section or adjacent tissue section in the preprocessed biological tissue data set to the spatial transcriptomic data and unify the data resolution to obtain the corrected spatial metabolomic data and update the biological tissue data set;

[0010] Step 3: Select the tissue sample to be integrated and the data modality from the latest biological tissue data set and generate the to-be-integrated sample data, input the to-be-integrated sample data into the spatial multi-omics data integration model, and the model outputs the final joint embedding.

[0011] The integration method further comprises the following steps:

[0012] Step 4: Select and generate other to-be-integrated sample data from the biological tissue data set, and obtain the corresponding joint embedding after processing by the spatial multi-omics data integration model respectively.

[0013] In step 1, the data preprocessing includes data format normalization processing.

[0014] In the step 2, the spatial metabolomics data is aligned to the spatial transcriptomics data, including aligning the spatial point coordinate system in the spatial metabolomics data to the spatial point coordinate system in the spatial transcriptomics data by using an affine transformation matrix A to obtain a new spatial point coordinate system of the spatial metabolomics data; or based on the spatial point coordinates in the spatial metabolomics data and the histological images in the spatial transcriptomics data, aligning the spatial point coordinate system in the spatial metabolomics data to the spatial point coordinate system in the spatial transcriptomics data by using an affine transformation matrix A to obtain a new spatial point coordinate system of the spatial metabolomics data.

[0015] The aligning the spatial point coordinate system in the spatial metabolomics data to the spatial point coordinate system in the spatial transcriptomics data by using an affine transformation matrix A to obtain a new spatial point coordinate system of the spatial metabolomics data specifically includes:

[0016] First, an initial affine transformation matrix A is generated according to the spatial point coordinates in the spatial metabolomics data and the spatial point coordinates in the spatial transcriptomics data; then, the spatial point coordinate system in the spatial metabolomics data is aligned to the spatial point coordinate system in the spatial transcriptomics data by using the current affine transformation matrix A to obtain the aligned spatial point coordinates of the spatial metabolomics data;

[0017] Then, the Euclidean distance between the aligned spatial point coordinates of the spatial metabolomics data and the spatial point coordinates in the spatial transcriptomics data is calculated; the optimization objective is to minimize the Euclidean distance, and the large deformation differential isomorphic metric mapping model uses a stochastic gradient descent method to optimize and solve the affine transformation matrix A to minimize the optimization objective, and the optimal affine transformation matrix A is obtained; finally, the spatial point coordinate system in the spatial metabolomics data is aligned to the spatial point coordinate system in the spatial transcriptomics data by using the optimal affine transformation matrix A to obtain a new spatial point coordinate system of the spatial metabolomics data.

[0018] The aligning the spatial point coordinate system in the spatial metabolomics data to the spatial point coordinate system in the spatial transcriptomics data by using an affine transformation matrix A to obtain a new spatial point coordinate system of the spatial metabolomics data specifically includes:

[0019] Firstly, the histological image in the spatial transcriptomic data is converted into a two-dimensional point set corresponding to its spatial coordinates, and the two-dimensional point set of the ST image is obtained after the registration information provided by the ST sequencing platform is used to unify the two to the same coordinate system; then the spatial point coordinates in the spatial metabolomic data are extracted by using the concave hull algorithm to obtain the two-dimensional point set of the SM tissue shape; an initial affine transformation matrix A is generated according to the two-dimensional point set of the ST image and the two-dimensional point set of the SM tissue shape; the spatial point coordinate system in the spatial metabolomic data is aligned to the histological image coordinate system in the spatial transcriptomic data by using the current affine transformation matrix A, and the two-dimensional point set of the aligned SM tissue shape is obtained; then the Euclidean distance between the two-dimensional point set of the aligned SM tissue shape and the two-dimensional point set of the ST image is calculated; the random gradient descent method is used to optimize and solve the current affine transformation matrix A to minimize the optimization target, so that the optimal affine transformation matrix A is obtained; finally, the spatial point coordinate system in the spatial metabolomic data is aligned to the histological image coordinate system in the spatial transcriptomic data by using the optimal affine transformation matrix A, and a new spatial point coordinate system of the spatial metabolomic data is obtained.

[0020] In step 2, the spatial metabolomic data is unified with the spatial transcriptomic data to have the same data resolution, which specifically includes:

[0021] Firstly, the spatial point coordinates in the spatial transcriptomic data are extracted and used as the coordinates of the spatial metabolomic data in the new spatial point coordinate system, denoted as SM reference spatial points; and after the spatial point coordinates in the preprocessed spatial metabolomic data are aligned to the spatial transcriptomic data, the obtained spatial metabolomic data coordinates are denoted as SM aligned spatial points.

[0022] Then, the K SM aligned spatial points closest to each SM reference spatial point are found; the spatial points in the found K SM aligned spatial points that are too far away from the current SM reference spatial point are removed, and a plurality of SM aligned spatial points corresponding to the current SM reference spatial point are obtained; finally, the metabolic signal values of the plurality of SM aligned spatial points are averaged and used as the metabolic signal value of the current SM reference spatial point; the remaining SM reference spatial points are processed iteratively to obtain the metabolic signal values of the remaining SM reference spatial points.

[0023] In step 3, the data modalities of the sample data to be integrated include at least spatial metabolomic data; when the data modalities are only spatial metabolomic data, the sample data to be integrated comes from multiple tissue samples; when the data modalities are spatial metabolomic data and spatial transcriptomic data, the sample data to be integrated includes one spatial metabolomic data and one spatial transcriptomic data from the same tissue sample, or paired spatial metabolomic data and spatial transcriptomic data from multiple tissue samples.

[0024] In step 3, the spatial multi-omics data integration model includes two batch-invariant encoders and two batch-variant decoders, after the spatial metabolomics data in the sample data to be integrated is input into the first batch-invariant encoder, the first batch-invariant encoder outputs a low-dimensional representation of SM; after the spatial transcriptomics data in the sample data to be integrated is input into the second batch-invariant encoder, the second batch-invariant encoder outputs a low-dimensional representation of ST; the outputs of the two batch-invariant encoders are spliced to obtain a joint representation q joint , based on the joint representation q joint , a joint embedding z joint of a multivariate normal distribution is generated joint , the joint embedding z joint is taken as the input of the two batch-variant decoders, the first batch-variant decoder reconstructs the spatial metabolomics data according to the joint embedding z joint , the second batch-variant decoder reconstructs the spatial transcriptomics data according to the joint embedding z

[0025]

[0026] , wherein, are the reconstruction mean, variance and gating weight of the spatial transcriptomics data after reconstruction based on the joint embedding z joint , are the reconstruction mean and variance of the spatial metabolomics data after reconstruction based on the joint embedding z joint , is the first batch-variant decoder, is the second batch-variant decoder, and B represents batch information; when only a single tissue slice is integrated in the sample data to be integrated, B is set to null; when multiple tissue slices are integrated, the batch information B is used to eliminate batch effects;

[0027] The distribution mean of the low-dimensional representation of SM is calculated and obtained, and the distribution mean of the low-dimensional representation of ST is calculated and obtained, the distribution mean of the low-dimensional representation of SM is input into the first batch-variant decoder, and the distribution mean of the low-dimensional representation of ST is input into the second batch-variant decoder, the first batch-variant decoder further reconstructs the spatial metabolomics data according to the distribution mean μ sm of the low-dimensional representation of SM, and the second batch-variant decoder further reconstructs the spatial transcriptomics data according to the distribution mean μ st of the low-dimensional representation of ST; the two batch-variant decoders satisfy the following formula:

[0028]

[0029] wherein, are the distribution mean μ st reconstructed mean, variance and gating weight of the reconstructed spatial transcriptomic data, are the distribution mean μ sm reconstructed mean and variance of the reconstructed spatial metabolomic data.

[0030] The loss function in the training process of the spatial multi-omics data integration model satisfies the following formula:

[0031]

[0032] wherein, is the total loss value of the spatial multi-omics data integration model; N is the number of spatial sites; is the reconstruction loss of the spatial metabolomic data from the joint embedding z joint reconstructed spatial metabolomic data; is the distribution mean μ sm reconstruction loss of the spatial metabolomic data; is the reconstruction loss of the spatial transcriptomic data from the joint embedding z joint reconstructed spatial transcriptomic data; is the distribution mean μ st reconstruction loss of the spatial transcriptomic data; is the maximum mean discrepancy loss; is the KL divergence loss of the variational distribution.

[0033] In the step 3, if the sample data to be integrated includes spatial metabolomic data and spatial transcriptomic data, then the modal contribution degree of the current sample data to be integrated is generated according to the final joint embedding.

[0034] The modal contribution degree of the current sample data to be integrated is generated according to the final joint embedding, including:

[0035] The calculation formula of the modal contribution degree of the current sample data to be integrated is as follows:

[0036]

[0037] wherein, MCD i is the modal contribution degree of the spatial site i; normalize() represents a normalization operation, is the direction similarity of the spatial transcriptomic modal of the spatial site i on the joint embedding, which measures the angle similarity (i.e. cosine similarity) between the distribution mean μ and the ST low-dimensional representation; a direction similarity of the spatial metabolic modality of the spatial point i on the joint embedding; a distribution mean of the ST low-dimensional representation of the spatial point i; a joint embedding of the multivariate normal distribution of the spatial point i; a distribution mean of the SM low-dimensional representation of the spatial point i; || is the L2 norm (Euclidean norm) of a vector.

[0038] II. A computer device

[0039] The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the steps of the spatial transcriptome and spatial metabolome integration method based on deep learning when executing the computer program.

[0040] III. A computer readable storage medium

[0041] The computer readable storage medium stores a computer program, and the computer program implements the steps of the spatial transcriptome and spatial metabolome integration method based on deep learning when executed by a processor.

[0042] IV. A computer program product

[0043] The computer program product comprises a computer program / instruction, which implements the steps of the spatial transcriptome and spatial metabolome integration method based on deep learning when executed by a processor.

[0044] The beneficial effects of the present application are:

[0045] (1) The best ST and SM cross-modal integration effect: The present application can effectively integrate spatial transcriptome (ST) data and spatial metabolome (SM) data, and through multi-level model optimization, the fusion effect of cross-modal data is improved. The present application makes full use of the feature information of different modal data, significantly improves the accuracy and reliability of data analysis, and provides a more efficient analysis tool for spatial omics research.

[0046] (2) The best ST and SM cross-modal and cross-sample integration effect: The present application not only can efficiently integrate within the same modality, but also can enhance the correlation between different source samples through cross-sample data fusion. Through this cross-modal and cross-sample integration, the present application can provide more accurate and comprehensive biological information to help researchers discover potential biomarkers and their spatial distribution characteristics. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1Fig. 1 shows a detailed schematic diagram of each step of the method of the present application, including alignment, uniform resolution and integration of the three modules; wherein a is a schematic diagram of the alignment process of ST and SM data; b is a schematic diagram of the process of re-distribution of ST and SM data at different resolutions to a uniform resolution; c is a schematic diagram of the cross-sample and cross-modality integration analysis of ST-SM data using a variational autoencoder.

[0048] Figure 2 Fig. 2 shows the results of the alignment analysis of ST and SM, wherein a is a spatial map of the original ST data (left) and SM data (right) SM spatial points from the ccRCC (Y7_T) sample aligned with the ST spatial points through rotation, translation and non-linear distortion; b is a spatial map of the original ST data (left) and SM data (right) SM spatial points from the ccRCC (R29_T) sample aligned with the ST spatial points through rotation, translation and non-linear distortion; c is a spatial map of the original ST data (left) and SM data (right) SM spatial points from the ccRCC (Y7_T) sample aligned with the spatial histology image through rotation, translation and non-linear distortion; d is a spatial map of the original ST data (left) and SM data (right) SM spatial points from the ccRCC (R29_T) sample aligned with the spatial histology image through rotation, translation and non-linear distortion.

[0049] Figure 3 Fig. 3 shows the results of the resolution unification of ST and aligned SM, wherein a is a visualization of the aligned SM spatial points from the ccRCC (Y7_T) sample projected onto the histology image; b is a visualization of the aligned SM spatial points from the ccRCC (R29_T) sample projected onto the histology image; c is the SM data from the ccRCC (Y7_T) sample after resolution unification, having the same coordinate and resolution as the ST; d is the SM data from the ccRCC (R29_T) sample after resolution unification, having the same coordinate and resolution as the ST.

[0050] Figure 4 Fig. 4 shows the results of the cross-modality integration analysis of ST and SM from the same sample using a spatial multi-omics data integration model. From left to right, the spatial maps are from the ccRCC (Y7_T), ccRCC (R29_T) and mouse brain samples, and the different colored points represent the clustering results based on the joint embedding output by the spatial multi-omics data integration model.

[0051] Figure 5Schematic diagram of cross-sample and cross-modality integration analysis of ST and SM from multiple samples using the spatial multi-omics data integration model. A unified spatial cluster is found by cross-sample and cross-modality integration of paired ST and SM from multiple samples. Wherein, a is that the ST and SM of multiple samples have a unified spatial cluster in the spatial distribution after cross-sample and cross-modality integration; b is that the ST and SM of multiple samples have a unified spatial cluster in the low-dimensional visualization unified manifold approximation and projection after cross-sample and cross-modality integration.

[0052] Figure 6 Result graph of cross-sample and cross-modality integration analysis of ST and SM of multiple ccRCC samples using the spatial multi-omics data integration model, wherein a is the UMAP graph after dimensionality reduction visualization using the joint embedding of the spatial multi-omics data integration model, the left side is colored by sample, and the right side is colored according to the clustering result, wherein Imm, immune cluster; Endo, endothelial cluster; Stro, stromal cluster; Mal, malignant cluster; b is the visualization of the clustering result based on the joint embedding of the spatial multi-omics data integration model in the spatial graph of a single sample.

[0053] Figure 7 Result graph of cross-sample integration analysis of SM of multiple ccRCC samples using the spatial multi-omics data integration model, wherein a is the unified manifold approximation and projection (UMAP) graph after dimensionality reduction visualization using the joint embedding of the spatial multi-omics data integration model, the left side is colored by sample, and the right side is colored according to the clustering result; b is the visualization of the clustering result based on the joint embedding of the spatial multi-omics data integration model in the spatial graph of a single sample.

[0054] Figure 8 Modality contribution result graph calculated in the integration data process using the spatial multi-omics data integration model, wherein a is the spatial graph of the joint embedding obtained by cross-modality integration of ST and SM data of ccRCC using the spatial multi-omics data integration model and the contribution result of ST and SM to the joint embedding respectively; b is the spatial graph of the joint embedding obtained by cross-modality integration of ST and SM data of mouse brain samples using the spatial multi-omics data integration model and the contribution result of ST and SM to the joint embedding respectively.

[0055] Figure 9 Flow chart of the method of the application. DETAILED DESCRIPTION

[0056] The application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0057] As Figure 1 a of FIG. 1, Figure 1 b of FIG. 2, and Figure 1 c of FIG. 3, and Figure 9As shown, the spatial transcriptome and spatial metabolome integration method based on deep learning proposed by the present application comprises the following steps:

[0058] Step 1: Obtain the original spatial transcriptomic (ST) data and the original spatial metabolomic (SM) data of the biological tissue and perform data preprocessing to obtain the preprocessed biological tissue dataset; in this embodiment, the spatial transcriptomic (ST) data and the spatial metabolomic (SM) data from human clear cell renal cell carcinoma (ccRCC), glioblastoma (GBM) and mouse brain tissue samples are collected, which are derived from previous studies and published datasets. The ST dataset is generated by the 10x Genomics Visium platform, and the SM dataset is generated by the MALDI and DESI platforms.

[0059] Among them, the data preprocessing includes data format normalization processing, specifically processing the original spatial transcriptomic data into a gene expression matrix (GEM) format and processing the original spatial metabolomic data into a metabolite intensity matrix (MIM) format.

[0060] Step 2: Align the spatial metabolomic data of the same tissue section or adjacent tissue section in the preprocessed biological tissue dataset to the spatial transcriptomic data and unify the data resolution to obtain corrected spatial metabolomic data and update the biological tissue dataset;

[0061] In a feasible implementation, aligning the spatial metabolomic data of the same tissue section or adjacent tissue section to the spatial transcriptomic data comprises aligning the spatial point coordinate system in the spatial metabolomic data to the spatial point coordinate system in the spatial transcriptomic data using an affine transformation matrix A to obtain a new spatial point coordinate system of the spatial metabolomic data; or based on the spatial point coordinates in the spatial metabolomic data and the histological image in the spatial transcriptomic data, aligning the spatial point coordinate system in the spatial metabolomic data to the spatial point coordinate system in the spatial transcriptomic data using an affine transformation matrix A to obtain a new spatial point coordinate system of the spatial metabolomic data, as shown. Figure 2 Considering that ST and SM are usually derived from adjacent sections of the same sample, their spatial point distribution and tissue contour have high similarity, and ST is usually equipped with a corresponding histological image, while SM often lacks a histological image. Therefore, in general, the former data alignment method is preferred.

[0062] In a feasible implementation, aligning the spatial metabolomic data of the same tissue section or adjacent tissue section to the spatial transcriptomic data comprises aligning the spatial point coordinate system in the spatial metabolomic data to the spatial point coordinate system in the spatial transcriptomic data using an affine transformation matrix A to obtain a new spatial point coordinate system of the spatial metabolomic data, specifically comprising:

[0063] First, align the coordinate system of the spatial points in the spatial metabolomics data to the coordinate system of the spatial points in the spatial transcriptomics data according to the spatial point coordinates in the spatial metabolomics data and the spatial point coordinates in the spatial transcriptomics data Generate an initial affine transformation matrix A, the affine transformation matrix A includes a rotation matrix a translation matrix a nonlinear scaling matrix and a warping term Then, align the coordinate system of the spatial points in the spatial metabolomics data to the coordinate system of the spatial points in the spatial transcriptomics data using the current affine transformation matrix A, to obtain the aligned spatial point coordinates Coord' of the spatial metabolomics data SM , which satisfies the following formula:

[0064]

[0065] where φ() represents a warping function for nonlinear transformation of input coordinates;

[0066] Then, calculate the Euclidean distance between the aligned spatial point coordinates of the spatial metabolomics data and the spatial point coordinates in the spatial transcriptomics data; take the minimum Euclidean distance as the optimization objective, and use the stochastic gradient descent method of the Large Deformation Diffeomorphic Metric Mapping (LDDMM) model to optimize and solve the affine transformation matrix A, so that the optimization objective is minimized, to obtain the optimal affine transformation matrix A; during the optimization process, the following objective function is used, and the formula is as follows:

[0067]

[0068] where is the alignment loss function value between the spatial metabolomics spatial points and the spatial transcriptomics spatial points, and PointDistance() represents the Euclidean distance calculation function between the corresponding spatial point pairs. Finally, align the coordinate system of the spatial points in the spatial metabolomics data to the coordinate system of the spatial points in the spatial transcriptomics data using the optimal affine transformation matrix A, to obtain a new spatial point coordinate system of the spatial metabolomics data. Figure 2 a and Figure 2 b of the ST data (left) and the SM data (right) from the ccRCC (Y7_T) and the ccRCC (R29_T) samples. The spatial graph of the aligned SM spatial points through rotation, translation, and nonlinear distortion with the ST spatial points can be seen in the figure, and it can be seen that the aligned ST data and the SM spatial points present similar morphology but different spatial point resolutions.

[0069] In an implementable embodiment, based on the spatial point coordinates in the spatial metabolomics data and the histological image in the spatial transcriptomics data, an affine transformation matrix A is used to align the spatial point coordinate system in the spatial metabolomics data to the histological image coordinate system in the spatial transcriptomics data, to obtain a new spatial point coordinate system of the spatial metabolomics data, specifically including:

[0070] Firstly, the histological image in the spatial transcriptomics data is converted into a two-dimensional point set corresponding to its spatial coordinates (Coord ST ) and the two-dimensional point set of the ST image is obtained after the two are unified to the same coordinate system using the registration information provided by the ST sequencing platform Then, since the SM image (Image SM ) is usually unavailable, the spatial point coordinates in the spatial metabolomics data are extracted to obtain a two-dimensional point set of the SM tissue shape after the extraction of the tissue contour using a concave shell algorithm The tissue shape of the SM slice is approximated; an initial affine transformation matrix A is generated according to the two-dimensional point set of the ST image and the two-dimensional point set of the SM tissue shape, and the affine transformation matrix A includes a rotation matrix a translation matrix a non-linear scaling matrix and a deformation term The spatial point coordinate system in the spatial metabolomics data is aligned to the histological image coordinate system in the spatial transcriptomics data using the current affine transformation matrix A, to obtain a two-dimensional point set of the aligned SM tissue shape Image_Coord′ SM , which satisfies the following formula:

[0071]

[0072] Wherein, φ(Image_Coord′ SM ) represents the two-dimensional coordinate point set in the spatial metabolomics image after deformation and affine transformation;

[0073] Then, the Euclidean distance between the two-dimensional point set of the aligned SM tissue shape and the two-dimensional point set of the ST image is calculated; the Large Deformation Diffeomorphic Metric Mapping (LDDMM) model uses a stochastic gradient descent method to optimize and solve the current affine transformation matrix A, so that the optimization target is minimized, to obtain the optimal affine transformation matrix A; during the optimization process, the following objective function is used, and the formula is as follows:

[0074]

[0075] Wherein, represents the alignment loss function value between image level point sets. Finally, the optimal affine transformation matrix A is used to align the spatial point coordinate system in the spatial metabolomics data to the histological image coordinate system in the spatial transcriptomics data, obtaining a new spatial point coordinate system of the spatial metabolomics data. Figure 2 c and d of Figure 2 The left and right of c and d respectively show the original ST data (left) and SM data (right) from ccRCC (Y7_T) and ccRCC (R29_T) samples. The SM spatial points are aligned with the spatial histological image after rotation, translation and non-linear distortion.

[0076] The ST and SM data are derived from the same or adjacent tissue sections, and are similar in general, but differ in tissue morphology and the direction of the spatial points may be misaligned. In addition, these data sets are generated using different platforms and technologies, resulting in differences in spatial resolution. Although the SM data is similar in morphology to the ST data and the corresponding ST histological image, differences in position are observed between the SM and ST spatial points. The mouse brain (m3_FMP) sample is relatively resolved in ST and SM data, but in most samples, the resolution of the ST data is generally lower than that of the SM data. For example, in the ccRCC Y7_T sample, the original ST data contains 2018 points, while the SM data has 10145 points. The present application uses the ST coordinates to identify the KNN points in the SM data and average the MIM values, thereby generating a down-sampled SM data set containing 2018 points. If the ST data has higher spatial resolution, it also allows the ST data to be redistributed to the SM data, which provides flexibility for adapting to emerging high-resolution ST technologies.

[0077] In an implementable embodiment, the spatial metabolomics data is unified with the spatial transcriptomics data in data resolution, specifically comprising:

[0078] For the spatial metabolomics data of the same or adjacent tissue sections to the spatial transcriptomics data, the process of unifying the data resolution is as follows:

[0079] First, the spatial coordinates of points in the spatial transcriptomics data are extracted and used as the coordinates of the spatial metabolomics data in a new spatial coordinate system, denoted as the SM baseline spatial points. Then, the spatial coordinates of the preprocessed spatial metabolomics data are aligned to the spatial transcriptomics data, and the resulting spatial metabolomics data coordinates are denoted as SM aligned spatial points. Next, the K nearest SM aligned spatial points to each SM baseline spatial point are found according to Euclidean distance, with a default K value of 5. Users can adjust the K value according to their specific analysis needs. The larger the K value, the more original SM spatial points are fused to each SM baseline spatial point, resulting in a smoother distribution of the generated MIM. Spatial points that are too far from the current SM baseline spatial point are removed from the K found SM aligned spatial points, resulting in multiple SM aligned spatial points corresponding to the current SM baseline spatial point. In this embodiment, a distance threshold is set in the KNN calculation. This threshold is determined by the product of min_dist and dist_fold, where min_dist represents the minimum distance between spatial points in the spatial transcriptomics data, which can be calculated or manually set by the user; dist_fold is a scaling factor with a default value of 1.5, which can also be user-defined. Finally, the metabolic signal values ​​of these multiple SM-aligned spatial points are averaged and used as the metabolic signal (MIM) value of the current SM baseline spatial point. The remaining SM baseline spatial points are then processed to obtain their metabolic signal values, completing the correction of the spatial metabolomics data of the same or adjacent tissue slices, thereby updating the biological tissue dataset. Figure 3 a and Figure 3 b shows the visualization of the aligned SM spatial points corresponding to ccRCC(Y7_T) and ccRCC(R29_T) samples projected onto histological images. Figure 3 c and Figure 3 The d values ​​represent the SM data, which, after resolution unification, has the same coordinates and resolution as the ST.

[0080] Step 3: Select tissue slices and data modalities to be integrated from the latest biological tissue dataset and generate sample data to be integrated. Input the sample data to be integrated into the spatial multi-omics data integration model, and the model outputs the final joint embedding.

[0081] The data modality of the sample data to be integrated must include at least spatial metabolomics data. When the data modality is only spatial metabolomics data, the sample data to be integrated comes from multiple tissue samples. When the data modality is both spatial metabolomics and spatial transcriptomics data, the sample data to be integrated includes one spatial metabolomics data and one spatial transcriptomics data from the same tissue sample (the integration effect is as follows). Figure 4 (as shown), or paired spatial metabolomics and spatial transcriptomics data from multiple tissue samples (integration effect as shown).Figure 5 a and Figure 5 b, Figure 6 a and Figure 6 (as shown in b).

[0082] The spatial multi-omics data fusion model includes two batch-invariant encoders and two decoders with batch variability, with spatial metabolomics data X in the sample data to be integrated. sm After being input into the first batch of invariant encoders, the first batch of invariant encoders outputs the SM low-dimensional representation q. sm Spatial transcriptomics data X from the sample data to be integrated st After being input into the second batch invariant encoder, the second batch invariant encoder outputs the ST low-dimensional representation q. st The following formula must be satisfied:

[0083]

[0084] in, For the second batch of unchanged encoders, This is the first batch invariant encoder. The joint representation q is obtained by concatenating the outputs of the two batch invariant encoders. joint Based on joint representation q joint Generate joint embedding z of multivariate normal distribution joint It satisfies the following formula:

[0085]

[0086] Where, μ joint For joint characterization of q joint The mean, σ joint For joint characterization of q joint The root mean square of the variance, To perform the mean operation, For variance calculation, ∈ is the numerical stability factor, which is usually a small constant (default value is 1e-4) to prevent σ from being zero. It follows a multivariate normal distribution.

[0087] Jointly embed z joint As input to two batch-variable decoders, the first batch-variable decoder is based on the joint embedding z. joint Reconstructing spatial metabolomics data, a second decoder with batch variability based on joint embedding z joint Reconstruct spatial transcriptomics data; two decoders with batch variability satisfy the following formula:

[0088]

[0089] in, respectively, are the reconstructed mean and variance of the spatial transcriptomic data, and joint are the reconstructed mean and variance of the spatial metabolomic data, respectively, and is the expected expression value of the decoded ST data, is the variance of the reconstructed distribution, which is used to model the uncertainty of the reconstruction. is a gating variable ranging from 0 to 1, which is used to regulate the importance or activation level of different spatial locations or channels in the reconstruction process, and helps to improve the expressive ability and robustness of the model. These parameters are used to model the zero-inflated negative binomial distribution (ZINB), so as to reconstruct the original gene expression matrix (GEM). respectively, are the reconstructed mean and variance of the spatial transcriptomic data, and joint are the reconstructed mean and variance of the spatial metabolomic data, respectively, and and Similarly, these parameters are used to model the Gaussian distribution, so as to reconstruct the original metabolite intensity matrix (MIM). is a first batch-variant decoder, is a second batch-variant decoder, B represents batch information, which can represent samples, research projects or sequencing platforms; when the sample data to be integrated is only a single tissue slice, B is set to None; when used for cross-sample integration of multiple tissue slices, the batch information B is used to eliminate batch effects;

[0090] The distribution mean μ sm of the SM low-dimensional representation q sm is calculated and obtained, and the distribution mean μ st of the ST low-dimensional representation q st is calculated and obtained, the distribution mean μ sm of the SM low-dimensional representation q sm is input into the first batch-variant decoder, and the distribution mean μ st of the ST low-dimensional representation q st is input into the second batch-variant decoder, the first batch-variant decoder further reconstructs the spatial metabolomic data according to the distribution mean μ sm of the SM low-dimensional representation q sm , and the second batch-variant decoder further reconstructs the spatial transcriptomic data according to the distribution mean μ st of the ST low-dimensional representation q st , the two batch-variant decoders satisfy the following formula:

[0091]

[0092] wherein, Based on the distribution mean μ st The reconstructed mean, variance, and gating weights of the spatial transcriptomics data were analyzed. Based on the distribution mean μ sm The reconstructed mean and variance of spatial metabolomics data.

[0093] When the sample data to be integrated does not contain spatial transcriptomics data, the second batch-invariant encoder and the second batch-variable decoder do not work.

[0094] The two batch-invariant encoders include a variational autoencoder. The two decoders with batch variability include a conditional variational decoder.

[0095] The spatial multi-omics data fusion model is trained unsupervised, and its loss function during training satisfies the following formula:

[0096]

[0097] in, The total loss value of the spatial multi-omics data fusion model is N; N is the number of spatial sites. To embed z from the union joint Reconstruction loss of reconstructed spatial metabolomics data, which makes The resulting distribution follows a Gaussian distribution, thus allowing the reconstruction of the original metabolite intensity matrix (MIM). To characterize q from SM in low dimension sm The distribution mean μ sm Reconstruction loss of reconstructed spatial metabolomics data, which makes The resulting distribution follows a Gaussian distribution, thus allowing the reconstruction of the original metabolite intensity matrix (MIM). To embed z from the union joint Reconstruction loss of reconstructed spatial transcriptomics data. This loss results in... The resulting distribution satisfies a zero-inflated negative binomial distribution (ZINB), thus allowing the reconstruction of the original gene expression matrix (GEM). To characterize q from ST low-dimensional st The distribution mean μ st Reconstruction loss of reconstructed spatial transcriptomics data. This loss results in... The resulting distribution satisfies a zero-inflated negative binomial distribution (ZINB), thus allowing the reconstruction of the original gene expression matrix (GEM). The maximum mean difference (MMD) loss is used to ensure that the joint embeddings across samples are similar and can effectively correct for batch effects. The definition method is referenced in scArches. The KL divergence loss is a variational distribution. The weights of each loss term are customizable to adjust the optimization focus according to specific objectives. The total loss function of the spatial multi-omics data fusion model combines the total reconstruction loss of ST and SM, MMD loss, and KL divergence loss. During training, the first decoder with batch variability is based on the joint embedding z. joint Reconstructing spatial metabolomics data, the first decoder with batch variability also relies on SM low-dimensional characterization q sm The distribution mean μ sm Reconstructing spatial metabolomics data, parameters are shared between the first batch-variable decoder in both cases. The second batch-variable decoder is based on the joint embedding z. joint Reconstructing spatial transcriptomics data, a second decoder with batch variability also relies on ST low-dimensional characterization q st The distribution mean μ st Reconstruct spatial transcriptomics data, sharing parameters between the second decoder with batch variability in both cases.

[0098] That is, when integrating a single tissue slice across modalities, B = None. Two batch-invariant encoders and two batch-variant decoders are trained until training is complete, yielding the trained encoders and decoders, as well as the final joint embedding z. joint .

[0099] When integrating multiple tissue slices across modalities, B ≠ None, where B is a trainable parameter. Two batch-invariant encoders and two batch-variant decoders are trained until training is complete, yielding the trained encoders and decoders, as well as the final joint embedding z. joint .

[0100] When integrating single-modal data from multiple tissue slices, B≠None, where B is a trainable parameter. The second batch-invariant encoder and the second batch-variant decoder do not participate in training and their corresponding loss values ​​are 0. After training, the trained first batch-invariant encoder, the first batch-variant decoder, and the final joint embedding z are obtained. joint .

[0101] If the sample data to be integrated includes spatial metabolomics data and spatial transcriptomics data, then the modal contribution of the current sample data to be integrated is generated based on the final joint embedding.

[0102] The modal contribution of the current sample data to be integrated is generated based on the final joint embedding, including:

[0103] To evaluate the effect of each modality (ST and SM) on the joint embedding z jointTo assess the impact at each spatial point, this invention measures the input X by calculating the angular distance. st and X sm The formula for calculating the modal contribution of the current sample data to be integrated is as follows:

[0104]

[0105] Among them, MCD i The modal contribution of spatial point i is limited to the range [0,1], representing the relative influence of each modality in the joint embedding; normalize() represents the normalization operation. The directional similarity of spatial transcriptional modalities at spatial point i in joint embeddings is measured. The distribution mean of the low-dimensional representation of ST The similarity of the included angle between them (i.e., cosine similarity); Let be the directional similarity of the spatial metabolic modes at spatial point i in the joint embedding; Let be the distribution mean of the ST low-dimensional representation of spatial point i; For the joint embedding of the multivariate normal distribution of spatial point i; Let be the mean of the distribution of the SM low-dimensional representation of point i in space; || is the L2 norm (Euclidean norm) representing the vector.

[0106] For ccRCC samples, the spatial plot of the contribution results of their joint embedding is as follows: Figure 8 As shown in a. For mouse brain samples, the spatial plot of the contribution results of the joint embeddings is shown in Figure a. Figure 8 As shown in b, the points of different colors in the spatial map (from left to right) represent the contributions of ST or SM in the joint embedding. It can be observed that the contributions of ST and SM in ccRCC and mouse brain samples have obvious spatial distribution characteristics.

[0107] Step 4: Select and generate other sample data to be integrated from the biological tissue dataset. After processing them using the spatial multi-omics data integration model, obtain the corresponding joint embeddings, and then generate the modal contribution of each sample data to be integrated, thus completing the data integration of the biological tissue dataset.

[0108] Downstream Analysis: After completing cross-modal and cross-sample integration of ST and SM data, a series of analyses and visualizations can be performed to achieve downstream data analysis. Specific functions include spatial clustering, simultaneous identification of marker genes and metabolites, metabolite annotation and enrichment analysis, and correlation analysis between genes and metabolites. In addition, spatial trajectories of interest (TOIs) can be generated through an interactive interface to display gradients in gene expression and metabolite intensity, as well as interactive depiction and comparison of specific regions of interest (ROIs).

[0109] Figure 7 Results of SM cross-sample integration analysis of multiple ccRCC samples using the spatial multi-omics data integration model are shown. First, the SM data of multiple ccRCC samples are preprocessed and standardized, and then the cross-sample joint embedding representation is extracted by the spatial multi-omics integration model. Based on this joint embedding representation, clustering analysis is performed to identify spatial regions with shared characteristics. Finally, the integration effect is shown by uniform manifold approximation and projection (UMAP) dimensionality reduction and spatial visualization. From Figure 7 It can be seen that the joint embedding has good alignment effect between samples, and similar cell populations in different samples are clustered together and show a distribution with organizational structure characteristics in the spatial graph.

[0110] Finally, it should be noted that the above examples and explanations are only used to illustrate the technical solutions of the present application and not to limit it. Those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions disclosed by the present application, and they should be covered within the protection scope of the claims of the present application.

Claims

1. A deep learning-based method for integrating spatial transcriptomics and spatial metabolomics, characterized in that, Includes the following steps: Step 1: Obtain raw spatial transcriptomics data and raw spatial metabolomics data of biological tissues and perform data preprocessing to obtain preprocessed biological tissue datasets; Step 2: Align the spatial metabolomics data of the same or adjacent tissue slices in the preprocessed biological tissue dataset with the spatial transcriptomics data and unify the data resolution to obtain the corrected spatial metabolomics data and update the biological tissue dataset. Step 3: Select tissue samples and data modalities to be integrated from the latest biological tissue dataset and generate sample data to be integrated. Input the sample data to be integrated into the spatial multi-omics data integration model, and the model outputs the final joint embedding.

2. The method for integrating spatial transcriptomics and spatial metabolomics based on deep learning according to claim 1, characterized in that, The integration method further includes the following steps: Step 4: Select and generate other sample data to be integrated from the biological tissue dataset, process them using the spatial multi-omics data integration model, and obtain the corresponding joint embeddings.

3. The method for integrating spatial transcriptomics and spatial metabolomics based on deep learning according to claim 1, characterized in that, In step 1, data preprocessing includes data format normalization.

4. The method for integrating spatial transcriptomics and spatial metabolomics based on deep learning according to claim 1, characterized in that, In step 2, spatial metabolomics data is aligned to spatial transcriptomics data. This includes using an affine transformation matrix A to align the spatial point coordinate system in the spatial metabolomics data to the spatial point coordinate system in the spatial transcriptomics data, thereby obtaining a new spatial point coordinate system for the spatial metabolomics data; or, based on the spatial point coordinates in the spatial metabolomics data and the histological images in the spatial transcriptomics data, using an affine transformation matrix A to align the spatial point coordinate system in the spatial metabolomics data to the spatial point coordinate system in the spatial transcriptomics data, thereby obtaining a new spatial point coordinate system for the spatial metabolomics data.

5. The method for integrating spatial transcriptomics and spatial metabolomics based on deep learning according to claim 4, characterized in that, The step of aligning the spatial point coordinate system in the spatial metabolomics data to the spatial point coordinate system in the spatial transcriptomics data using an affine transformation matrix A to obtain a new spatial point coordinate system for the spatial metabolomics data specifically includes: First, an initial affine transformation matrix A is generated based on the spatial point coordinates in the spatial metabolomics data and the spatial transcriptomics data. Then, the spatial point coordinate system in the spatial metabolomics data is aligned to the spatial point coordinate system in the spatial transcriptomics data using the current affine transformation matrix A, thereby obtaining the aligned spatial point coordinates of the spatial metabolomics data. Then, the Euclidean distance between the aligned spatial point coordinates of the spatial metabolomics data and the spatial point coordinates of the spatial transcriptomics data is calculated. Taking the minimization of the Euclidean distance as the optimization objective, the large deformation differential isomorphism metric mapping model is optimized by stochastic gradient descent to solve the affine transformation matrix A, so that the optimization objective is minimized, and the optimal affine transformation matrix A is obtained. Finally, the optimal affine transformation matrix A is used to align the spatial point coordinate system of the spatial metabolomics data to the spatial point coordinate system of the spatial transcriptomics data, and a new spatial point coordinate system of the spatial metabolomics data is obtained.

6. The method for integrating spatial transcriptomics and spatial metabolomics based on deep learning according to claim 4, characterized in that, The process involves aligning the spatial point coordinates in the spatial metabolomics data to the histological image coordinate system in the spatial transcriptomics data using an affine transformation matrix A, thereby obtaining a new spatial point coordinate system for the spatial metabolomics data. Specifically, this includes: First, the histological images in the spatial transcriptomics data are converted into two-dimensional point sets corresponding to their spatial coordinates. Using registration information provided by the ST sequencing platform, both are unified to the same coordinate system to obtain the two-dimensional point set of the ST image. Next, the concave-shell algorithm is used to extract the tissue contour from the spatial point coordinates in the spatial metabolomics data, obtaining the two-dimensional point set of the SM tissue shape. An initial affine transformation matrix A is generated based on the two-dimensional point set of the ST image and the two-dimensional point set of the SM tissue shape. Finally, the current affine transformation matrix A is used to align the spatial point coordinate system in the spatial metabolomics data to the histological image in the spatial transcriptomics data. Like a coordinate system, obtain the two-dimensional point set of the aligned SM tissue shape; then calculate the Euclidean distance between the two-dimensional point set of the aligned SM tissue shape and the two-dimensional point set of the ST image; with the minimum Euclidean distance as the optimization objective, the large deformation differential isomorphism metric mapping model uses the stochastic gradient descent method to optimize the current affine transformation matrix A, so that the optimization objective is minimized, and obtain the optimal affine transformation matrix A; finally, use the optimal affine transformation matrix A to align the spatial point coordinate system in the spatial metabolomics data to the histological image coordinate system in the spatial transcriptomics data, and obtain the new spatial point coordinate system of the spatial metabolomics data.

7. The method for integrating spatial transcriptomics and spatial metabolomics based on deep learning according to claim 1, characterized in that, In step 2, the spatial metabolomics data is converted to spatial transcriptomics data with unified data resolution, specifically including: First, the spatial point coordinates in the spatial transcriptomics data are extracted and used as the coordinates of the spatial metabolomics data in the new spatial point coordinate system, denoted as the SM baseline spatial point; and the spatial metabolomics data coordinates obtained by aligning the spatial point coordinates in the preprocessed spatial metabolomics data to the spatial transcriptomics data are denoted as the SM aligned spatial point. Next, find the K nearest SM alignment points to each SM reference point; remove the K SM alignment points that are too far from the current SM reference point to obtain multiple SM alignment points corresponding to the current SM reference point; finally, calculate the average of the metabolic signal values ​​of these multiple SM alignment points and use it as the metabolic signal value of the current SM reference point; iterate through the remaining SM reference points to obtain the metabolic signal values ​​of the remaining SM reference points.

8. The method for integrating spatial transcriptomics and spatial metabolomics based on deep learning according to claim 1, characterized in that, In step 3, the data modality of the sample data to be integrated includes at least spatial metabolomics data. When the data modality is only spatial metabolomics data, the sample data to be integrated comes from multiple tissue samples. When the data modality is spatial metabolomics data and spatial transcriptomics data, the sample data to be integrated includes one spatial metabolomics data and one spatial transcriptomics data from the same tissue sample, or paired spatial metabolomics data and spatial transcriptomics data from multiple tissue samples.

9. The method for integrating spatial transcriptomics and spatial metabolomics based on deep learning according to claim 1, characterized in that, In step 3, the spatial multi-omics data integration model includes two batch invariant encoders and two decoders with batch variability. After the spatial metabolomics data in the sample data to be integrated is input into the first batch invariant encoder, the first batch invariant encoder outputs the SM low-dimensional representation. Spatial transcriptomics data from the sample data to be integrated are input into a second batch invariant encoder, which outputs a low-dimensional ST representation. The outputs of the two batch invariant encoders are then concatenated to obtain the joint representation q. joint Based on joint representation q joint Generate joint embedding z of multivariate normal distribution joint , will be jointly embedded in z joint As input to two batch-variable decoders, the first batch-variable decoder is based on the joint embedding z. joint Reconstructing spatial metabolomics data, a second decoder with batch variability based on joint embedding z joint Reconstructing spatial transcriptomics data; Two decoders with batch variability satisfy the following formula: in, Based on joint embedding z joint The reconstructed mean, variance, and gating weights of the spatial transcriptomics data were analyzed. Based on joint embedding z joint The reconstructed mean and variance of the spatial metabolomics data were analyzed. As the first decoder with batch variability, For the second decoder with batch variability, B represents batch information; when there is only a single tissue slice in the sample data to be integrated across modalities, B is set to none; when used for cross-sample integration of multiple tissue slices, the batch information B is used to eliminate batch effects. The distribution mean of the SM low-dimensional representation and the distribution mean of the ST low-dimensional representation are calculated and obtained. The distribution mean of the SM low-dimensional representation is input into a first decoder with batch variability, and the distribution mean of the ST low-dimensional representation is input into a second decoder with batch variability. The first decoder with batch variability further calculates the distribution mean μ of the SM low-dimensional representation. sm Reconstructing spatial metabolomics data, the second decoder with batch variability also relies on the distribution mean μ of the ST low-dimensional characterization. st Reconstruct spatial transcriptomics data; two decoders with batch variability satisfy the following formula: in, Based on the distribution mean μ st The reconstructed mean, variance, and gating weights of the spatial transcriptomics data were analyzed. Based on the distribution mean μ sm The reconstructed mean and variance of spatial metabolomics data.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the deep learning-based spatial transcriptome and spatial metabolome integration method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Breast cancer detection method and system based on morphological image and space transcriptome cross-graph collaborative learning

    CN121582228A