Method and device for generating multi-slice multi-omics spatial data, storage medium and computer device

By constructing a cross-slice adjacency graph and processing it with a conditional diffusion model, multi-slice multi-omics spatial data is generated, which solves the problem of insufficient cross-slice identification in existing technologies and realizes accurate integration of multi-slice data and reconstruction of spatial structure.

CN122493985APending Publication Date: 2026-07-31HANGZHOU INST FOR ADVANCED STUDY UCAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU INST FOR ADVANCED STUDY UCAS
Filing Date
2026-06-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing spatial omics data cannot effectively achieve spatial domain identification across slices, especially in continuous tissue slices and cross-modal data. Paired graph representations are insufficient to characterize the high-order spatial relationships formed by multiple adjacent spots or cells, and conventional integration methods tend to weaken the true spatial structure.

Method used

By acquiring gene expression and spatial coordinates from multiple continuous spatial transcriptome slices, a cross-slice adjacency graph is constructed to identify simple complexes with multiple levels. The graphs are then processed using a conditional diffusion model under fusion representation to generate multi-slice multi-omics spatial data. Cross-slice correction is performed using a high-order adjacency matrix and batch conditions.

Benefits of technology

It achieves accurate integration of multi-slice, multi-omics spatial data, reconstructs spatial representation and unifies potential representations, generates more accurate and effective multi-slice, multi-omics spatial data, and preserves the continuity and consistency of spatial structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493985A_ABST
    Figure CN122493985A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, storage medium, and computer equipment for generating multi-slice multi-omics spatial data, relating to the field of bioinformatics technology. Its main objective is to address the problem that existing spatial omics data cannot achieve cross-slice spatial domain identification. The method includes: acquiring gene expression and spatial coordinates of multiple consecutive spatial transcriptome slices; constructing a cross-slice adjacency graph based on gene expression and spatial coordinates, and identifying simplicities containing multiple levels based on the cross-slice adjacency graph; fusing the simplicities within the multi-level simplicities to obtain a fused representation; and processing gene expression based on a conditional diffusion model under the constraint of the fused expression to generate multi-slice multi-omics spatial data. The target variable in the conditional diffusion model is constructed from a high-order adjacency matrix based on simplicities, and the conditional diffusion process of the conditional diffusion model involves cross-slice correction through batch conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics technology, and in particular to a method, apparatus, storage medium, and computer equipment for generating multi-slice multi-omics spatial data. Background Technology

[0002] Recent advances in spatial transcriptomics (ST) and related spatial omics technologies have enabled researchers to perform molecular profiling while preserving the complete spatial coordinates of tissues. Biological systems are ordered in both spatial and temporal dimensions, but most experimental data analyze only single two-dimensional slices, providing only a local view of complex biological systems. This limitation is particularly pronounced for processes that span adjacent slices, evolve over time, or involve multiple molecular levels. To better capture these organizational patterns, an increasing number of studies are generating multi-slice and multi-omics spatial data.

[0003] Currently, existing spatial omics data can be identified through single-slice spatial domains or through multi-omics data fusion. However, these methods still represent tissue structures as ordinary graphs composed of nodes and paired edges to address issues such as local adjacency, pairwise alignment, or specific modality fusion. For continuous tissue slices, cross-modal data, and samples with complex tissue morphologies, paired graph representations are often insufficient to characterize the higher-order spatial relationships formed by multiple adjacent spots or cells. Furthermore, if integration is performed solely through static graph convolution or conventional batch correction, the true spatial structure is easily weakened, making it difficult to simultaneously obtain accurate spatial domains and stable cross-slice correspondences. Therefore, a method for generating multi-slice, multi-omics spatial data is urgently needed to address these issues. Summary of the Invention

[0004] In view of this, this application provides a method, apparatus, storage medium and computer equipment for generating multi-slice multi-omics spatial data, the main purpose of which is to solve the problem that existing spatial omics data cannot achieve cross-slice spatial domain identification.

[0005] According to one aspect of this application, a method for generating multi-slice, multi-omics spatial data is provided, comprising: Gene expression and spatial coordinates of multiple consecutive spatial transcriptome slices were obtained; Based on the gene expression and the spatial coordinates, a cross-slice adjacency graph is constructed, and based on the cross-slice adjacency graph, simple complexes containing multiple levels are determined; The simplexes in the multi-level simplex complex are fused to obtain a fused representation. Under the constraint of the fused expression, the gene expression is processed based on a conditional diffusion model to generate multi-slice multi-omics spatial data. The target variable in the conditional diffusion model is constructed based on the high-order adjacency matrix of the simplex, and the conditional diffusion process of the conditional diffusion model is cross-slice correction through batch conditions.

[0006] Furthermore, the construction of a cross-slice adjacency graph based on the gene expression and the spatial coordinates includes: Calculate the spatial autocorrelation values ​​of the gene expression and the spatial coordinates; Based on the spatial autocorrelation value, the variance between the slice sub-region and the slice is determined, the spatial information gene for alignment is determined, and the target slice and the template slice are aligned based on the spatial information gene. Based on the gene centroid coordinates of the aligned target slice and the template slice, as well as the spatial coordinates, cross-slice neighbor sites are determined, and the cross-slice neighborhood map is constructed based on the cross-slice neighbor sites.

[0007] Furthermore, determining the simple complex containing multiple levels based on the cross-slice adjacency graph includes: Extract a subset of nodes from the cross-slice adjacency graph, and construct a zero-type simplex, a first-type simplex, and a second-type simplex based on the vertices, edges, and filled triangles in the subset of nodes, respectively. The type zero simplex, the type I simplex, and the type II simplex are combined to form the simplex complex.

[0008] Furthermore, the step of fusing the simplexes in the multi-level simplex complex to obtain the fused representation includes: Construct bipartite graphs of the zero-type simplex, the first-type simplex, and the second-type simplex of each order, and perform normalization processing based on the diagonal matrix of the bipartite graphs to obtain higher-order adjacency matrices; A linear transformation is performed based on the expression feature matrix to obtain a linear expression. The linear expression and the higher-order adjacency matrix are then filtered by a preset filter to obtain a filtered linear expression. The filtered linear expressions of each order are concatenated and fused to obtain the fused representation.

[0009] Furthermore, the step of processing the gene expression based on a conditional diffusion model under the constraint of the fusion expression to generate multi-slice multi-omics spatial data includes: The target variable is constructed using the higher-order adjacency matrix, and a conditional diffusion model is constructed. The conditional diffusion process of the conditional diffusion model includes a forward diffusion process and a backward diffusion process. The drift terms in the forward diffusion process and the backward diffusion process are constrained by the fusion expression. The backward diffusion process includes a scoring neural network constrained by the fusion expression. The forward diffusion process and the reverse diffusion process are cross-slice corrected using the batch conditions, and the gene expression is processed based on the corrected conditional diffusion model to generate multi-slice multi-omics spatial data.

[0010] Furthermore, the batch conditions include batch information consisting of any one of the following: slide number, platform number, tissue source, pathological classification, experimental batch, and omics modality.

[0011] Furthermore, obtaining gene expression and spatial coordinates of multiple consecutive spatial transcriptome slices includes: Obtain the original gene expression count matrix for each spatial transcriptome slice, and perform normalization and logarithmic transformation on the original count matrix to obtain the gene expression of each slice. Based on multiple selected consecutive spatial transcriptome slices, a target number of gene expressions are selected from each slice, and the spatial coordinates corresponding to the gene expressions are determined.

[0012] According to another aspect of this application, an apparatus for generating multi-slice multi-omics spatial data is provided, comprising: The acquisition module is used to acquire gene expression and spatial coordinates of multiple consecutive spatial transcriptome slices; The determination module is used to construct a cross-slice adjacency graph based on the gene expression and the spatial coordinates, and to determine simple complexes containing multiple levels based on the cross-slice adjacency graph; The processing module is used to fuse the simplexes in the multi-level simplex complex to obtain a fused representation, and under the constraint of the fused expression, process the gene expression based on the conditional diffusion model to generate multi-slice multi-omics spatial data. The target variable in the conditional diffusion model is constructed based on the high-order adjacency matrix of the simplex, and the conditional diffusion process of the conditional diffusion model is cross-slice correction through batch conditions.

[0013] Further, the determining module is specifically used to calculate the spatial autocorrelation value of the gene expression and the spatial coordinates; determine the variance between the slice sub-region and the slice based on the spatial autocorrelation value, determine the spatial information genes for alignment, and align the target slice and the template slice based on the spatial information genes; determine cross-slice neighbor sites based on the gene centroid coordinates of the aligned target slice and the template slice and the spatial coordinates, and construct the cross-slice neighborhood map based on the cross-slice neighbor sites.

[0014] Furthermore, the determining module is specifically used to extract a subset of nodes from the cross-slice adjacency graph, and to construct a type zero simplex, a type I simplex, and a type II simplex based on the vertices, edges, and filled triangles in the subset of nodes; and to combine the type zero simplex, the type I simplex, and the type II simplex into the simplex complex.

[0015] Further, the processing module is specifically used to construct bipartite graphs of the zero-type simplex, the first-type simplex, and the second-type simplex of each order, and to perform normalization processing based on the diagonal matrix of the bipartite graph to obtain a higher-order adjacency matrix; to perform linear transformation based on the expression feature matrix to obtain a linear expression, and to filter the linear expression and the higher-order adjacency matrix through a preset filter to obtain a filtered linear expression; and to splice and fuse the filtered linear expressions of each order to obtain a fused representation.

[0016] Furthermore, the processing module is specifically used to construct a target variable using the higher-order adjacency matrix, construct a conditional diffusion model, wherein the conditional diffusion process of the conditional diffusion model includes a forward diffusion process and a backward diffusion process, the drift terms in the forward diffusion process and the backward diffusion process are constrained by the fusion expression, and the backward diffusion process includes a scoring neural network constrained by the fusion expression; the forward diffusion process and the backward diffusion process are cross-slice corrected using the batch conditions, and the gene expression is processed based on the corrected conditional diffusion model to generate multi-slice multi-omics spatial data.

[0017] Furthermore, the batch conditions include batch information consisting of any one of the following: slide number, platform number, tissue source, pathological classification, experimental batch, and omics modality.

[0018] Furthermore, the acquisition module is also used to acquire the original counting matrix of gene expression for each spatial transcriptome slice, and to perform normalization and logarithmic transformation on the original counting matrix to obtain the gene expression of each slice; based on the selected multiple consecutive spatial transcriptome slices, to select a target number of gene expressions from each slice, and to determine the spatial coordinates corresponding to the gene expressions.

[0019] According to another aspect of this application, a storage medium is provided, wherein at least one executable instruction is stored therein, the executable instruction causing a processor to perform operations corresponding to the above-described method for generating multi-slice multi-omics spatial data.

[0020] According to another aspect of this application, a terminal is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the above-described method for generating multi-slice multi-omics spatial data.

[0021] By employing the above technical solutions, the technical solutions provided in the embodiments of this application have at least the following advantages: This application provides a method, apparatus, storage medium, and computer device for generating multi-slice multi-omics spatial data. Compared with the prior art, the embodiments of this application obtain gene expression and spatial coordinates of multiple continuous spatial transcriptome slices; construct a cross-slice adjacency graph based on the gene expression and spatial coordinates, and determine simplex complexes containing multiple levels based on the cross-slice adjacency graph; fuse the simplex complexes containing multiple levels to obtain a fused representation, and process the gene expression based on a conditional diffusion model under the constraint of the fused expression to generate multi-slice multi-omics spatial data. The target variable in the conditional diffusion model is constructed based on the high-order adjacency matrix of the simplex, and the conditional diffusion process of the conditional diffusion model is cross-slice correction through batch conditions. By transforming the problem of integrating multi-slice, multi-modal spatial omics data into a conditional scoring diffusion generation problem constrained by tissue topology, spatial expression reconstruction and unified latent representation learning are achieved, and high-order spatial relationships are constructed, thereby achieving the purpose of integrating multi-slice, multi-omics spatial transcriptome data and generating more accurate and effective multi-slice multi-omics spatial data.

[0022] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description

[0023] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This paper illustrates a flowchart of a method for generating multi-slice, multi-omics spatial data according to an embodiment of this application. Figure 2 This illustration shows a schematic diagram of a paired RNA and ATAC measurement scenario provided in an embodiment of this application; Figure 3 This illustration shows a schematic diagram of an end-to-end framework for building SpaDiff according to an embodiment of this application; Figure 4 This illustration shows a schematic diagram of a topology-aware graph encoder provided in an embodiment of this application; Figure 5 This illustration shows a unified diffusion diagram provided by an embodiment of this application; Figure 6 This illustration shows a block diagram of a device for generating multi-slice multi-omics spatial data according to an embodiment of this application. Figure 7 A schematic diagram of the structure of a terminal provided in an embodiment of this application is shown. Detailed Implementation

[0024] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0026] This application provides a method for generating multi-slice, multi-omics spatial data, such as... Figure 1 As shown, the method includes: Step 101: Obtain gene expression and spatial coordinates from multiple consecutive spatial transcriptome slices.

[0027] In this embodiment, the current execution entity, acting as the processing end for generating multi-slice, multi-omics spatial data, can be a terminal device or a cloud server, etc., to obtain gene expression and spatial coordinates of multiple continuous spatial transcriptome slices. Gene expression refers to the entire process of gene (DNA) transcription into mRNA and then translation into protein, which can be represented as a gene transcription signal. Each spatial transcriptome slice corresponds to a gene expression matrix. Correspondingly, spatial coordinates refer to the unique two-dimensional location number of each detection site spot on the tissue slice. In this case, gene expression and spatial coordinates can be obtained based on gene images of tissue cells from different species obtained through different dimensional segmentation in a large-scale spatiotemporal omics image dataset (spatial transcriptomics database), or gene images of tissue cells acquired in real time. This embodiment does not impose specific limitations.

[0028] In some embodiments, a large-scale spatial transcriptome dataset may be constructed from raw transcriptome data, which may consist of a spatial transcriptomics (SRT) database. This SRT database may include data from STOmics (a spatiotemporal omics database), SOAR, SpatialDB (a spatial transcriptome resource), CROST (Comprehensive Repository of Spatial Transcriptomics), the 10x Genomics website, and the Census database. For example, it may contain 96,700,729 cells / sites from 7,367 tissue sections, covering 365 tissue cell types, including lung, skin, brain, liver, kidney, spinal cord, and embryonic tissues. It may also include normal tissues and various diseases such as pancreatic ductal adenocarcinoma, amyotrophic lateral sclerosis (ALS), non-small cell lung cancer, and hepatocellular carcinoma. The embodiments of this invention do not specifically limit the scope of the dataset.

[0029] In another embodiment of this application, for further definition and explanation, the step of obtaining gene expression and spatial coordinates of multiple consecutive spatial transcriptome slices includes: Obtain the original gene expression count matrix for each spatial transcriptome slice, and perform normalization and logarithmic transformation on the original count matrix to obtain the gene expression of each slice. Based on multiple selected consecutive spatial transcriptome slices, a target number of gene expressions are selected from each slice, and the spatial coordinates corresponding to the gene expressions are determined.

[0030] To accurately generate multi-slice, multi-omics spatial data, when obtaining gene expression data and their corresponding spatial coordinates, each spatial transcriptome slice can first be obtained from the RNA-seq dataset. The original gene expression counting matrix Furthermore, regarding the original counting matrix... Normalization and logarithmic transformation are performed. Normalization can be applied to the library size, thus normalizing the total count of the spot to 10,000. The logarithmic transformation uses... The process involves processing to obtain gene expression data for each slice. Multiple consecutive spatial transcriptome slices can be selected randomly or according to a specific sequence to select a target number of gene expressions from each slice, including but not limited to 3000, 5000, etc. The spatial coordinates of the spot sites corresponding to gene expression can be determined through molecular measurements or ST datasets.

[0031] In some embodiments, the top 3,000 hypervariable genes (HVGs) are selected from each slice, and multiple slices are selected as the union of HVGs across slices. Additionally, each gene can be z-score normalized (e.g., mean 0, variance 1) across all spots to ensure the effectiveness of gene expression.

[0032] Step 102: Construct a cross-slice adjacency graph based on the gene expression and the spatial coordinates, and determine simple complexes containing multiple levels based on the cross-slice adjacency graph.

[0033] In this embodiment, since gene expression and spatial coordinates correspond to multiple consecutive spatial transcriptome slices, the cross-slice adjacency graph refers to the adjacency matrix constructed after selecting neighboring slices from adjacent slices to achieve the purpose of cross-slice three-dimensional reconstruction. Furthermore, based on the cross-slice adjacency graph, a multi-level simplex complex is determined, where a simplex complex refers to a finite subset of nodes, and a node is a spot site. Each element in the simplex complex... It is called a simplex, which constitutes a multi-level simplex complex and is closed for any non-empty subset.

[0034] Step 103: Fuse the simplexes in the multi-level simplex complex to obtain a fused representation, and process the gene expression based on the conditional diffusion model under the constraints of the fused expression to generate multi-slice multi-omics spatial data.

[0035] In this embodiment, a multi-channel architecture is used, where each simplex order corresponds to an independent propagation channel. Therefore, to enable parallel processing of features of different orders and capture spatial backgrounds at multiple topological scales, the current execution end fuses simplexes from multiple levels to obtain a fused representation, which serves as an embedding constraint in the conditional diffusion model processing. The conditional diffusion model refers to adding constraints to a basic diffusion model to generate multi-slice, multi-omics spatial data. Here, the target variable of the conditional diffusion model is constructed from a high-order adjacency matrix based on the simplex, and the conditional diffusion process is cross-slice correction through batch conditions. Batch conditions refer to conditions used to eliminate technical or experimental differences between slices or modalities. Batch conditions include batch information corresponding to any one of the following: slice number, platform number, tissue origin, pathological classification, experimental batch, and omics modality. This embodiment does not specifically limit the batch conditions.

[0036] In another embodiment of this application, for further definition and explanation, the step of constructing a cross-slice adjacency graph based on the gene expression and the spatial coordinates includes: Calculate the spatial autocorrelation values ​​of the gene expression and the spatial coordinates; Based on the spatial autocorrelation value, the variance between the slice sub-region and the slice is determined, the spatial information gene for alignment is determined, and the target slice and the template slice are aligned based on the spatial information gene. Based on the gene centroid coordinates of the aligned target slice and the template slice, as well as the spatial coordinates, cross-slice neighbor sites are determined, and the cross-slice neighborhood map is constructed based on the cross-slice neighbor sites.

[0037] To reflect more accurate spatial relationships across slices, when constructing the cross-slice adjacency map, the spatial autocorrelation values ​​of gene expression and spatial coordinates are first calculated. At this point, equidistant grid lines can be placed along the horizontal and vertical axes to divide each tissue slice into equal-sized spatial sub-regions. Within each sub-region, Moran's algorithm is used to... Quantifying gene-wise spatial autocorrelation, specifically for tissue sections. Neutron region genes within Moran's The calculation is as follows:

[0038] in, For slices subregion Mid-spot Gene The expression, The average expression within subregion j. For spots and Spatial weights between them Let be the number of spots in subregion j, and .

[0039] Furthermore, spatial autocorrelation values ​​can be used to determine the variance between sub-regions and slices, i.e., Moran's variance can be calculated for each gene. Calculate the variance across all subregions and slices, and determine the spatial information genes used for alignment, such as selecting the slices with the largest variance. Each gene serves as spatial information for alignment. Furthermore, the target slice and template slice are aligned based on these spatial information genes. At this point, the template slice is represented as... The target slice is represented as , Slice When aligning two-dimensional coordinates... For slices The two-dimensional coordinates, used to calculate the gene centroid coordinates weighted by expression in the target slice and template slice respectively, can be represented as: .

[0040] At this point, a pre-trained multilayer perceptron (MLP) can be used to learn the mapping from the template centroid set to the target centroid set. Then, the learned transformation is applied to the spot coordinates to align the slices to a common coordinate system.

[0041] In this embodiment of the application, after alignment is completed, cross-slice neighbor sites are determined based on the gene centroid coordinates and spatial coordinates of the aligned target slice and template slice. The aligned spot coordinates are located in a shared three-dimensional space (x, y, z). To connect the spots between adjacent slices, the coordinates are projected onto the shared xy plane, and for each spot, its k nearest neighbors from adjacent slices are selected as cross-slice neighbors. Then, a cross-slice neighborhood graph is constructed based on the cross-slice neighbor sites.

[0042] In one specific embodiment, the construction of the cross-slice neighborhood graph can begin by constructing a weighted adjacency matrix. ,in This represents the total number of spots across all slices. When calculating edge weights, low-dimensional latent embeddings are obtained from the representation data using Harmony. , dimension For neighbors across slice neighborhood graphs and , ; Among them, temperature parameter After normalization, the cross-slice neighborhood graph can be represented as: ; in, This is the set of neighbors corresponding to neighbor sites across slices.

[0043] In another embodiment of this application, for further definition and explanation, the step of determining a simple complex containing multiple levels based on the cross-slice adjacency graph includes: Extract a subset of nodes from the cross-slice adjacency graph, and construct a zero-type simplex, a first-type simplex, and a second-type simplex based on the vertices, edges, and filled triangles in the subset of nodes, respectively. The type zero simplex, the type I simplex, and the type II simplex are combined to form the simplex complex.

[0044] To capture higher-order nonlinear spatial dependencies beyond pairwise connections, a simplicial complex is constructed on the merged multi-slice neighborhood graph. Specifically, a subset of nodes across the cross-slice adjacency graph is extracted, which constitutes the simplicial complex. Besides using a three-node cluster based on a k-nearest neighbor graph to form a triangle, other methods include using radius neighborhood graphs, Delaunay triangulation, alphacomplex, histological image segmentation regions, or cell type co-occurrence relationships to construct simple complexes containing higher-order simplexes. Each element... It is called a simplex if the simplex ,satisfy Then it is called - Simplex, all - The set of simplexes is denoted as ,and Furthermore, based on the vertices, edges, and filled triangles in the node subset, we construct a zero-type simplex, a first-type simplex, and a second-type simplex, respectively, which are represented as vertices corresponding to the zero-type simplex (0-simplex), edges corresponding to the first-type simplex (1-simplex), and filled triangles corresponding to the second-type simplex (2-simplex), to serve as simplex complexes.

[0045] In one embodiment, the highest order is preferably set as , to obtain from nodes ( ),side( ) and triangle ( The hierarchy, consisting of triangles formed by 3-cliques across the bottom slice adjacency graph, is able to encode higher-order spatial interactions between adjacent spots.

[0046] In another embodiment of this application, for further definition and explanation, the step of fusing the simplexes in the multi-level simplex complex to obtain the fused representation includes: Construct bipartite graphs of the zero-type simplex, the first-type simplex, and the second-type simplex of each order, and perform normalization processing based on the diagonal matrix of the bipartite graphs to obtain higher-order adjacency matrices; A linear transformation is performed based on the expression feature matrix to obtain a linear expression. The linear expression and the higher-order adjacency matrix are then filtered by a preset filter to obtain a filtered linear expression. The filtered linear expressions of each order are concatenated and fused to obtain the fused representation.

[0047] To accurately capture the high-order relationships prevalent in real-world microenvironments and identify the background of high-order spaces, thereby achieving topological relationship recognition, simplexes of various orders can be fused to obtain a fused representation that can serve as a constraint for a conditional diffusion model. First, bipartite graphs of type zero, type I, and type II simplexes of each order are constructed; that is, for each simplex order... In the node set (0-simplex) and - Set of simplex Constructing a bipartite graph between them At this point, graph neural networks, hypergraphs, cellular complexes, and combinatorial Laplacian methods can be used to construct structures that can express higher-order spatial dependencies between multiple spots or cells. The node-simplex correlation matrix can be represented as: .

[0048] Furthermore, by normalizing the diagonal matrix of the bipartite graph, a higher-order adjacency matrix is ​​obtained, and the angle matrix can be represented as:

[0049] At this point, for a simple complex, each -The simplex contains exactly There are 10 nodes, therefore... Finally, normalization is performed through two random walks: from node to the simplex, and then from node to node, resulting in the normalized higher-order adjacency matrix between nodes, expressed as: ; in, As a symmetric matrix, it can be interpreted as reflecting the... - The normalized adjacency matrix for node connectivity induced by the simplex, where different orders... Shared same set of node indexes Furthermore, it encodes different higher-order structures, thus providing complementary perspectives on spatial relationships.

[0050] In this embodiment, after obtaining the higher-order adjacency matrix, a linear transformation is performed based on the expression feature matrix to obtain a linear expression. Specifically, The input is the expression feature matrix, for example, an HVG expression or a 50-dimensional PCA feature matrix, where, The number of space spots. Let be the input feature dimension. For each order... After performing a linear transformation, the linear expression is as follows: ; in, and For trainable parameters, To hide dimensions. Furthermore, to capture multi-hop higher-order backgrounds, normalized higher-order adjacencies are used. The above is filtered using a preset filter to obtain a filtered linear expression. The preset filter is preferably a polynomial propagation filter. The filtered linear expression can be expressed as: ; in To maximize the depth of dissemination, These are the filter coefficients, used to control the contribution of different hop distances. We use hyperparameters. right Parameterization is performed to balance the preservation of initial features with the inclusion of higher-order neighborhood information: ; This design is equivalent to applying a Laplace multinomial filter and provides a stable way to aggregate information from progressively expanding higher-order neighborhoods. Finally, the filtered linear representations of each order are concatenated and fused to obtain a fused representation; that is, the multidimensional spatial background features are concatenated and input into a linear classifier to generate the final fused representation. : ; in, and For trainable parameters, by using (Edge-induced structure) and (Triangle-induced structure) Two channels are used to enhance the feature representation of spots in order to fuse the representation. It is used as a topology-aware embedding for downstream modules.

[0051] In another embodiment of this application, for further definition and explanation, the step of processing the gene expression based on a conditional diffusion model under the constraint of the fusion expression to generate multi-slice multi-omics spatial data includes: The target variable is constructed using the aforementioned higher-order adjacency matrix, and a conditional diffusion model is built. The forward diffusion process and the reverse diffusion process are cross-slice corrected using the batch conditions, and the gene expression is processed based on the corrected conditional diffusion model to generate multi-slice multi-omics spatial data.

[0052] To reduce noise in spatial omics data and improve the accuracy of multi-slice multi-omics spatial data generation, effective representation learning is performed on simple complexes to reconstruct coherent spatial signals. The current execution end constructs target variables using high-order adjacency matrices and builds a conditional diffusion model.

[0053] In a specific embodiment, the conditional diffusion process of the conditional diffusion model includes a forward diffusion process and a backward diffusion process. In the conditional diffusion module that constructs the target variable using a high-order adjacency matrix, For the higher-order adjacency matrix of each simplex order in the diffusion module The target variable to be modeled, and represented by fusion. When the condition is met, the forward diffusion process SDE is represented as: ; in, For the drift term, i.e., constrained by the fusion expression, Control noise levels. In standard Brownian motion, the forward diffusion process gradually distributes the data. Perturbation to tractable Gaussian priors The corresponding reverse-time generation dynamics are expressed as: ; in, For scoring neural networks, by neural networks Approximate determination, that is, equivalently determined in the denoising representation by Parameterization, and at the same time, the sampling process is performed by... arrive The iterative denoising is completed, thus obtaining a spatially consistent reconstruction result. Furthermore, the scoring neural network here can also be a denoising diffusion probability model, a continuous-time SDE model, a variational autoencoder, a normalized flow, or a generative adversarial network; this application does not specifically limit the implementation.

[0054] In a specific example, when using fusion representation... A single diffusion model with given conditions can integrate information from different simplicities. In this case, inspired by the consistency of score matching among different conditional representations, the forward diffusion process and backward diffusion process of the unified diffusion process can be represented as follows:

[0055]

[0056] Accordingly, the inverse dynamics are represented by fusion. As a condition.

[0057] It should be noted that, to account for technical or experimental differences between joint slices or modalities, the conditional diffusion process is extended to include batch indicator variables. The diffusion process, including forward diffusion and backward diffusion, is represented as follows: ; ; The above process involves cross-slice correction of the forward and reverse diffusion processes using batch conditions. At this point, the conditional scoring network... This enables conditional diffusion models to separate topological or spatial structure from batch-specific distortions, facilitating the coordinated integration of diffusion-based multi-tissue slices or omics layers while preserving spatial continuity. Finally, gene expression is processed based on the corrected conditional diffusion model to generate multi-slice, multi-omics spatial data.

[0058] For the model training process of the conditional diffusion model, the denoising score matching loss can be expressed as:

[0059] in, As a constraint, batch variables It can be used to align slices to their respective normal distributions, which helps to mitigate batch effects when generating characterizations. This is used to constrain the distribution of each slice to maintain as much balance as possible. Simultaneously, several hyperparameters are considered, including the hidden dimension (e.g., the number of neurons per layer), learning rate, and number of training epochs. In a specific example, the hyperparameters within a reasonable range are preferably a common default configuration to ensure fair comparison and reproducibility. For higher-order graph encoders, neural networks such as graph attention networks and GraphTransformers can be used to model a two-dimensional simplex (i.e.,...). This is sufficient to capture high-order spatial structures in spatial transcriptome data and consistently achieves good performance across various datasets. During training, priority is given to using [this method / technology]. The learning rate was adjusted and the training was performed for 500 epochs to achieve stable convergence.

[0060] For the training process, the performance of the conditional diffusion model can also be evaluated, including spatial domain identification. Specifically, spatial domain identification can be evaluated using the Adjusted Rand Index (ARI), cluster purity, and intra-group correlation coefficient (ICC). ARI measures the consistency between predicted clusters and reference annotations, and corrects for random consistency. Cluster purity measures the proportion of samples in each predicted cluster assigned to the dominant reference class. ICC measures intra-cluster similarity relative to inter-cluster variation; a higher value indicates more coherent clusters. For example, Indicates the predicted cluster, Indicates the reference group. The total number of spots. , and have and The ARI is calculated as follows: .

[0061] Furthermore, cluster purity calculation can be expressed as: .

[0062] For multi-slice spatial transcriptome data, spots from all slices can be concatenated into a unified set, and ARI / purity can be calculated on this integrated set, rather than averaging the scores of each slice, to avoid introducing bias through slice-by-slice averaging. Furthermore, spatial expression pattern preservation refers to assessing whether spatial structure is preserved during the generation process; Moran's I is used to quantify spatial autocorrelation before and after generation. ; in, For spot The expression value at that location, The average expression for all spots. For spots and Spatial weights between them Additionally, by comparing marker genes in different spatial domains... foldchanges FC) to examine its spatial domain specificity. If Moran's I and / or marker genes are denoised... If the FC increases or remains stable, it indicates that technical noise has been suppressed while preserving the true spatial gradient and spatial domain-specific expression patterns.

[0063] In some embodiments, the processed multi-slice multi-omics spatial data, such as spot, can be clustered to identify spatial domains. Two clustering strategies can be used: Mclust and Louvain. Mclust allows direct specification of the number of clusters, while Louvain controls the clustering granularity through a resolution parameter. Mclust is used for the DLPFC dataset, and Louvain is used for other datasets. For denoising, the representation matrix of the multi-slice multi-omics spatial data can be used. This allows us to obtain denoised gene expression maps, thereby reducing technical noise while enhancing spatial consistency.

[0064] In some embodiments, to characterize functional differences between spatial domains, Hallmark gene sets can be used for GeneSetVariationAnalysis (GSVA) ​​to quantify pathway activity and compare functional variations at the spatial domain level. Preferably, AUCell is used to calculate gene set activity scores and map them back to spatial coordinates to visualize the spatial functional landscape in tumor tissue. Simultaneously, for differential expression analysis, a pseudo-bulk RNA-seq counting matrix can be constructed by aggregating cluster-based spot results within each spatial domain, and DESeq2 is used for spatial domain-level differential expression testing. Furthermore, GeneOntology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) are used to perform enrichment analysis of differentially expressed genes using the R package clusterProfiler.

[0065] In a specific real-time scenario, SpaDiff can be used as an end-to-end framework for integrating multi-slice and multi-omics spatial transcriptome (ST) data. By jointly modeling spatial structure, transcriptional changes, and batch effects within a unified stochastic framework, SpaDiff can reconstruct spatially coherent organizational structures and support cross-modal signal generation. Combined with the multi-slice multi-omics spatial data generation method in the embodiments of this application, spatial multi-omics data can be integrated, for example, as... Figure 2 The paired RNA and ATAC measurements are shown. For each spatial dataset, the gene expression matrix is ​​as follows: Spatial coordinates are It can also include optional modality-specific features to define a spatial adjacency graph. In this context, vertices correspond to spots, and edges encode local spatial neighborhoods, such as... Figure 3 The diagram shows the SpaDiff framework. To capture higher-order spatial dependencies beyond pairwise interactions, a simple complex hierarchy is constructed using SpaDiff. and from a series of bipartite diagrams Learn the corresponding higher-order embeddings Each of them Node set (0-simplex) and - Set of simplex Connection. In each simple complex Above, SpaDiff used stochastic differential equations (SDEs) to define the continuous diffusion process: Forward: (Forward SDE); Reverse: ; in, Indicates time transcription signal, The learned drift function, Control noise levels. For standard Brownian motion, the likelihood is generated by the scoring function. The controlled inverse-time dynamics recovery, with the scoring function estimated by a neural network conditioned on the simplex topology, generates a consistent probabilistic representation for multi-slice observations via SpaDiff by embedding multiple simplexes into a shared diffusion space. Figure 3 As shown. Simultaneously, a topology-aware graph encoder built on top of the simplex level generates higher-order spatial embeddings, such as... Figure 4 The topology-aware graph encoder above the simplex hierarchy is shown. For each level... SpaDiff computes feature representations through graph convolution and attention, and aggregates them into a unified embedding. ,in The learned weights are used to reflect the contribution of each simplex order. Furthermore, to correct for batch effects across slices or modalities, batch / modal variables are further introduced via SpaDiff. Conditional diffusion, such as Figure 5 The unified diffusion of higher-order structures across batches is shown, and the diffusion process includes: Forward: ; Backward: ; Furthermore, conditional scoring is learned during the reverse denoising process. This allows SpaDiff to reconcile datasets from different slices or modalities while preserving topological consistency.

[0066] This application provides a method for generating multi-slice multi-omics spatial data. Compared with the prior art, this application obtains gene expression and spatial coordinates of multiple continuous spatial transcriptome slices; constructs a cross-slice adjacency graph based on the gene expression and spatial coordinates, and determines simplex complexes containing multiple levels based on the cross-slice adjacency graph; fuses the simplex complexes containing multiple levels to obtain a fused representation, and processes the gene expression based on a conditional diffusion model under the constraint of the fused expression to generate multi-slice multi-omics spatial data. The target variable in the conditional diffusion model is constructed based on the high-order adjacency matrix of the simplex, and the conditional diffusion process of the conditional diffusion model is cross-slice correction through batch conditions. By transforming the problem of integrating multi-slice, multi-modal spatial omics data into a conditional scoring diffusion generation problem constrained by tissue topology, spatial expression reconstruction and unified latent representation learning are achieved, and high-order spatial relationships are constructed, thereby achieving the purpose of integrating multi-slice, multi-omics spatial transcriptome data and generating more accurate and effective multi-slice multi-omics spatial data.

[0067] Furthermore, as a response to the above Figure 1 The implementation of the method shown in this application provides an apparatus for generating multi-slice, multi-omics spatial data, such as... Figure 6 As shown, the device includes: The acquisition module 21 is used to acquire gene expression and spatial coordinates of multiple consecutive spatial transcriptome slices; The determination module 22 is used to construct a cross-slice adjacency graph based on the gene expression and the spatial coordinates, and to determine simple complexes containing multiple levels based on the cross-slice adjacency graph; Processing module 23 is used to fuse the simplexes in the multi-level simplex complex to obtain a fused representation, and under the constraint of the fused expression, process the gene expression based on a conditional diffusion model to generate multi-slice multi-omics spatial data. The target variable in the conditional diffusion model is constructed based on the high-order adjacency matrix of the simplex, and the conditional diffusion process of the conditional diffusion model is cross-slice correction through batch conditions.

[0068] Further, the determining module is specifically used to calculate the spatial autocorrelation value of the gene expression and the spatial coordinates; determine the variance between the slice sub-region and the slice based on the spatial autocorrelation value, determine the spatial information genes for alignment, and align the target slice and the template slice based on the spatial information genes; determine cross-slice neighbor sites based on the gene centroid coordinates of the aligned target slice and the template slice and the spatial coordinates, and construct the cross-slice neighborhood map based on the cross-slice neighbor sites.

[0069] Furthermore, the determining module is specifically used to extract a subset of nodes from the cross-slice adjacency graph, and to construct a type zero simplex, a type I simplex, and a type II simplex based on the vertices, edges, and filled triangles in the subset of nodes; and to combine the type zero simplex, the type I simplex, and the type II simplex into the simplex complex.

[0070] Further, the processing module is specifically used to construct bipartite graphs of the zero-type simplex, the first-type simplex, and the second-type simplex of each order, and to perform normalization processing based on the diagonal matrix of the bipartite graph to obtain a higher-order adjacency matrix; to perform linear transformation based on the expression feature matrix to obtain a linear expression, and to filter the linear expression and the higher-order adjacency matrix through a preset filter to obtain a filtered linear expression; and to splice and fuse the filtered linear expressions of each order to obtain a fused representation.

[0071] Furthermore, the processing module is specifically used to construct a target variable using the higher-order adjacency matrix, construct a conditional diffusion model, wherein the conditional diffusion process of the conditional diffusion model includes a forward diffusion process and a backward diffusion process, the drift terms in the forward diffusion process and the backward diffusion process are constrained by the fusion expression, and the backward diffusion process includes a scoring neural network constrained by the fusion expression; the forward diffusion process and the backward diffusion process are cross-slice corrected using the batch conditions, and the gene expression is processed based on the corrected conditional diffusion model to generate multi-slice multi-omics spatial data.

[0072] Furthermore, the batch conditions include batch information consisting of any one of the following: slide number, platform number, tissue source, pathological classification, experimental batch, and omics modality.

[0073] Furthermore, the acquisition module is also used to acquire the original counting matrix of gene expression for each spatial transcriptome slice, and to perform normalization and logarithmic transformation on the original counting matrix to obtain the gene expression of each slice; based on the selected multiple consecutive spatial transcriptome slices, to select a target number of gene expressions from each slice, and to determine the spatial coordinates corresponding to the gene expressions.

[0074] This application provides an apparatus for generating multi-slice multi-omics spatial data. Compared with the prior art, this application obtains gene expression and spatial coordinates of multiple continuous spatial transcriptome slices; constructs a cross-slice adjacency graph based on the gene expression and spatial coordinates, and determines simplex complexes containing multiple levels based on the cross-slice adjacency graph; fuses the simplex complexes containing multiple levels to obtain a fused representation, and processes the gene expression based on a conditional diffusion model under the constraint of the fused expression to generate multi-slice multi-omics spatial data. The target variable in the conditional diffusion model is constructed based on the high-order adjacency matrix of the simplex, and the conditional diffusion process of the conditional diffusion model is cross-slice correction through batch conditions. By transforming the problem of integrating multi-slice, multi-modal spatial omics data into a conditional scoring diffusion generation problem constrained by tissue topology, spatial expression reconstruction and unified latent representation learning are achieved, and high-order spatial relationships are constructed, thereby achieving the purpose of integrating multi-slice, multi-omics spatial transcriptome data and generating more accurate and effective multi-slice multi-omics spatial data.

[0075] According to one embodiment of this application, a storage medium is provided, the storage medium storing at least one executable instruction that can execute the method for generating multi-slice multi-omics spatial data in any of the above method embodiments.

[0076] Figure 7 The diagram shows a structural schematic of a terminal according to one embodiment of the present application. The specific embodiments of the present application do not limit the specific implementation of the terminal.

[0077] like Figure 7 As shown, the terminal may include: a processor 302, a communications interface 304, a memory 306, and a communications bus 308.

[0078] The processor 302, communication interface 304, and memory 306 communicate with each other via communication bus 308.

[0079] Communication interface 304 is used to communicate with other network elements such as clients or other servers.

[0080] The processor 302 is used to execute program 310, specifically to execute the relevant steps in the above-described embodiment of the method for generating multi-slice multi-omics spatial data.

[0081] Specifically, program 310 may include program code that includes computer operation instructions.

[0082] Processor 302 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The terminal includes one or more processors, which may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.

[0083] Memory 306 is used to store program 310. Memory 306 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0084] Specifically, program 310 can be used to cause processor 302 to perform the following operations: Gene expression and spatial coordinates of multiple consecutive spatial transcriptome slices were obtained; Based on the gene expression and the spatial coordinates, a cross-slice adjacency graph is constructed, and based on the cross-slice adjacency graph, simple complexes containing multiple levels are determined; The simplexes in the multi-level simplex complex are fused to obtain a fused representation. Under the constraint of the fused expression, the gene expression is processed based on a conditional diffusion model to generate multi-slice multi-omics spatial data. The target variable in the conditional diffusion model is constructed based on the high-order adjacency matrix of the simplex, and the conditional diffusion process of the conditional diffusion model is cross-slice correction through batch conditions.

[0085] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0086] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for generating multi-slice, multi-omics spatial data, characterized in that, include: Gene expression and spatial coordinates of multiple consecutive spatial transcriptome slices were obtained; Based on the gene expression and the spatial coordinates, a cross-slice adjacency graph is constructed, and based on the cross-slice adjacency graph, simple complexes containing multiple levels are determined; The simplexes in the multi-level simplex complex are fused to obtain a fused representation. Under the constraints of the fused representation, the gene expression is processed based on a conditional diffusion model to generate multi-slice multi-omics spatial data. The target variable in the conditional diffusion model is constructed based on the high-order adjacency matrix of the simplex, and the conditional diffusion process of the conditional diffusion model is cross-slice correction through batch conditions.

2. The method of claim 1, wherein, The construction of a cross-slice adjacency graph based on the gene expression and the spatial coordinates includes: Calculate the spatial autocorrelation values ​​of the gene expression and the spatial coordinates; Based on the spatial autocorrelation value, the variance between the slice sub-region and the slice is determined, the spatial information gene for alignment is determined, and the target slice and the template slice are aligned based on the spatial information gene. Based on the gene centroid coordinates of the aligned target slice and the template slice, as well as the spatial coordinates, cross-slice neighbor sites are determined, and the cross-slice adjacency graph is constructed based on the cross-slice neighbor sites.

3. The method according to claim 2, characterized in that, The determination of a multi-level simple complex based on the cross-slice adjacency graph includes: Extract a subset of nodes from the cross-slice adjacency graph, and construct a zero-type simplex, a first-type simplex, and a second-type simplex based on the vertices, edges, and filled triangles in the subset of nodes, respectively. The type zero simplex, the type I simplex, and the type II simplex are combined to form the simplex complex.

4. The method according to claim 3, characterized in that, The step of fusing the simplexes in the multi-level simplex complex to obtain the fused representation includes: Construct bipartite graphs of the zero-type simplex, the first-type simplex, and the second-type simplex of each order, and perform normalization processing based on the diagonal matrix of the bipartite graphs to obtain higher-order adjacency matrices; A linear transformation is performed based on the expression feature matrix to obtain a linear expression. The linear expression and the higher-order adjacency matrix are then filtered by a preset filter to obtain a filtered linear expression. The filtered linear expressions of each order are concatenated and fused to obtain the fused representation.

5. The method according to claim 4, characterized in that, The process of processing the gene expression based on a conditional diffusion model under the constraint of the fusion expression to generate multi-slice multi-omics spatial data includes: The target variable is constructed using the higher-order adjacency matrix, and a conditional diffusion model is constructed. The conditional diffusion process of the conditional diffusion model includes a forward diffusion process and a backward diffusion process. The drift terms in the forward diffusion process and the backward diffusion process are constrained by the fusion expression. The backward diffusion process includes a scoring neural network constrained by the fusion expression. The forward diffusion process and the reverse diffusion process are cross-slice corrected using the batch conditions, and the gene expression is processed based on the corrected conditional diffusion model to generate multi-slice multi-omics spatial data.

6. The method according to claim 5, characterized in that, The batch conditions include batch information consisting of any one of the following: slide number, platform number, tissue source, pathological classification, experimental batch, and omics modality.

7. The method according to any one of claims 1-6, characterized in that, The process of obtaining gene expression and spatial coordinates from multiple consecutive spatial transcriptome slices includes: Obtain the original gene expression count matrix for each spatial transcriptome slice, and perform normalization and logarithmic transformation on the original count matrix to obtain the gene expression of each slice. Based on multiple selected consecutive spatial transcriptome slices, a target number of gene expressions are selected from each slice, and the spatial coordinates corresponding to the gene expressions are determined.

8. A device for generating multi-slice, multi-omics spatial data, characterized in that, include: The acquisition module is used to acquire gene expression and spatial coordinates of multiple consecutive spatial transcriptome slices; A determination module is used to construct a cross-slice adjacency graph based on the gene expression and the spatial coordinates, and to determine simple complexes containing multiple levels based on the cross-slice adjacency graph; The processing module is used to fuse the simplexes in the multi-level simplex complex to obtain a fused representation, and under the constraints of the fused representation, process the gene expression based on a conditional diffusion model to generate multi-slice multi-omics spatial data. The target variable in the conditional diffusion model is constructed based on the high-order adjacency matrix of the simplex, and the conditional diffusion process of the conditional diffusion model is cross-slice correction through batch conditions.

9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method of claim 1.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method of claim 1.