A three-dimensional spatial sequencing interpolation method and system based on diffusion model

Through self-supervised interpolation algorithm based on diffusion model and graph neural network, the problem of low resolution in the z-axis direction in three-dimensional spatial sequencing technology is solved, and high-precision three-dimensional gene expression reconstruction is achieved. The generated map can more accurately reflect the true shape and structure of the organ.

CN119889440BActive Publication Date: 2025-06-06BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510361237.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-06
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

Due to the limitation of slice thickness, the resolution of the z-axis direction is significantly lower than that of the xy plane. The traditional interpolation method has problems such as ghosting and blurring, which cannot effectively solve the problem of low z-axis spatial resolution in sequencing technology.

Method used

A self-supervised interpolation algorithm based on diffusion model is adopted, combined with graph neural network, a cell microenvironment graph and a random differential equation model are constructed, and the shape and expression prediction network is trained. Intermediate slices are generated through bidirectional interpolation optimization, and the interpolated slice data is integrated to generate continuous three-dimensional gene expression maps.

Benefits of technology

High-precision continuous three-dimensional gene expression reconstruction is achieved, solving the problem of low resolution in the z-axis direction. The generated three-dimensional map can more accurately reflect the true shape and structure of the organ.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119889440B_ABST
    Figure CN119889440B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional spatial sequencing interpolation method and system based on a diffusion model, the method comprising: constructing a cell microenvironment map based on spatial neighborhood relationships, establishing an association network between cell position and gene expression; modeling the continuous dynamic changes of cell position and gene expression with depth through stochastic differential equations; using graph neural networks to learn spatial gradients, respectively predicting shape drift coefficients and expression drift coefficients; using a bidirectional inference strategy to jointly optimize the interpolation results, and constraining the alignment accuracy of shape and expression through Wasserstein distance and mean square error; integrating interpolation slices to generate continuous three-dimensional transcriptional maps. The present invention achieves continuous completion of z-axis slice data by fusing the diffusion model with the graph neural network, and achieves high-resolution reconstruction of three-dimensional gene expression of tissues. It can accurately analyze complex spatial biological problems such as tumor boundaries and brain functional zoning, and provide key technical support for pathological diagnosis, developmental research and drug development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of bioinformatics, computational biology and image processing technology, and more specifically to a three-dimensional spatial sequencing interpolation method and system based on a diffusion model. Background Art

[0002] Spatial transcriptomics technology provides an important means for analyzing tissue heterogeneity, tumor microenvironment, etc. by retaining the spatial location information of gene expression. Existing three-dimensional spatial omics sequencing technology is based on continuous sectioning to characterize the expression of genes and other biological molecules in three-dimensional space. Assume that the plane of the slice is the xy plane, and the normal vector is parallel to the z-axis. Due to the objective influence of the slice thickness, the data obtained by three-dimensional spatial sequencing technology has hundreds of thousands of cells in the xy plane, and only 10 to 100 slices in the z direction. Therefore, the resolution of the tissue on the z-axis is much smaller than its resolution on the xy axis, resulting in uneven resolution along different directions.

[0003] Therefore, due to the limitation of slice thickness, the existing three-dimensional sequencing technology has a significantly lower resolution in the z-axis direction than in the xy plane.

[0004] In recent years, computational methods have been developed to complete sparse z-axis data through simple spatial proximity relationships.

[0005] 1. Sparse Vector Field Consensus (SparseVFC) is a method for sparse vector field estimation, which is particularly suitable for processing sparse vector field data with noise and outliers. It is widely used in tasks that require estimating sparse vector fields in bioinformatics (such as single-cell trajectory inference), physical modeling, and computer vision. The core idea of ​​SparseVFC is to optimize a sparse representation model while retaining the data structure and effectively eliminating the interference of outliers on vector field estimation.

[0006] However, due to the objective influence of slice thickness and the change of slice shape along the z-axis, the interpolation results are often accompanied by ghosting and cannot reflect the true shape and structure of the organ.

[0007] 2. Completing the z-axis of three-dimensional spatial omics through deep network regression method:

[0008] This method uses deep neural networks to learn the relationship between spatial coordinates and gene expression, uses sparse z-axis data to observe the results, and completes the slices at the unsequenced z values.

[0009] Spateo uses a multilayer perceptron (MLP) to learn the mapping relationship between three-dimensional coordinates (x, y, z) and gene expression g, and optimizes it through gradient descent. Similarly, because the number of observed z coordinates is too small, it is difficult for the network to generalize the sparse observation results to the entire tissue space, resulting in the interpolation results being too smooth and fuzzy, and unable to reflect the true shape and structure of the organ.

[0010] Therefore, the above two traditional interpolation methods (such as SparseVFC or deep network regression) have problems such as ghosting and blurring. They do not utilize the continuity of biological organs in space and the microenvironment information at the cellular scale, and cannot effectively interpolate sparse z-value observation data, thus failing to solve the problem of low z-axis spatial resolution in sequencing technology through computational means. Summary of the invention

[0011] In view of this, the present invention provides a three-dimensional spatial sequencing interpolation method and system based on a diffusion model. By combining the diffusion model and the graph neural network, a self-supervised interpolation algorithm is proposed to break through the bottleneck of the existing technology, solve the problem of low resolution in the z-axis direction in the existing three-dimensional spatial sequencing of spatial transcriptomics, and realize high-precision continuous three-dimensional gene expression reconstruction.

[0012] In order to achieve the above object, the present invention adopts the following technical solution:

[0013] In a first aspect, an embodiment of the present invention provides a three-dimensional spatial sequencing interpolation method based on a diffusion model, comprising the following steps:

[0014] 1) Construct a cell microenvironment map of three-dimensional spatial sequencing data: For K slices of data distributed along the z-axis, use the KNN algorithm to determine the spatial neighborhood relationship of each cell and generate a map containing spatial coordinates. and gene expression Graph structure datasets;

[0015] 2). Establishing a stochastic differential equation model: Based on the graph structure data set, a stochastic differential equation model is established in which the cell position continuously changes with the z-axis depth as a shape drift equation; and a stochastic differential equation model is established in which the gene expression continuously changes with the z-axis depth as a gene expression drift equation;

[0016] 3). Training graph neural network:

[0017] a. Build and train a shape prediction network : Input the spatial coordinates and spatial neighborhood relationship of the current slice, output the shape drift coefficient, substitute it into the shape drift equation, and predict the position offset of the cell in the next slice;

[0018] b. Build and train the expression prediction network : Input the spatial coordinates, gene expression and spatial neighborhood relationship of the current slice, output the expression drift coefficient, substitute it into the gene expression drift equation, and infer the change trend of gene expression;

[0019] 4) Bidirectional interpolation optimization: Use the trained graph neural network to perform interpolation prediction, generate intermediate slices through forward and reverse interpolation, and jointly optimize the Wasserstein distance of shape distribution and the mean square error loss of gene expression;

[0020] 5). Construct a three-dimensional transcriptional continuum: Integrate interpolated slice data to generate a continuous three-dimensional gene expression map.

[0021] Furthermore, step 1) specifically includes:

[0022] In the three-dimensional spatially resolved transcriptomics dataset, a piece of tissue is divided into K parallel sections for sequencing;

[0023] For K slice data distributed along the z axis, (x, y) is defined as the direction within each slice, and the z direction represents the tissue depth, and different slices are arranged along this direction;

[0024] For the depth A two-dimensional slice of Calculate the K nearest neighbors and generate a connection matrix A as the spatial neighborhood relationship of each cell, and form a matrix containing spatial coordinates and gene expression Graph structure dataset ;in, is the number of cells or sites in the slice, Indicates location The gene expression values.

[0025] Furthermore, in step 2):

[0026] The shape drift equation is as follows:

[0027] (1)

[0028] The gene expression drift equation is as follows:

[0029] (2)

[0030] in, and The measured depth and predicted depth The cell or location point location, and The corresponding expression characteristics; Indicates interpolation step size; drift coefficient and is a learnable gradient function that models the spatial variation of gene expression and tissue shape with depth, respectively, G is the gene depth; the diffusion term and It's white noise.

[0031] Furthermore, in step 3), a shape prediction network is constructed and trained ,include:

[0032] Using graph neural network GNN to build shape prediction network , and conduct training;

[0033] Input at depth The spatial coordinate data of the cells contained in the slice of the xy plane And the corresponding connection matrix A, predicting the position offset coefficient of each cell through spatial information , used to predict the spatial position distribution of cells at the next depth, so as to reconstruct the slice shape at the next depth;

[0034] In step 3), construct and train the expression prediction network ,include:

[0035] Using graph neural network GNN to build expression prediction network , and conduct training;

[0036] Input at depth The spatial coordinate data and gene expression of cells contained in the slices of the xy plane And the corresponding connection matrix A, predicting the gene expression deviation coefficient of each cell through spatial information , which is used to predict the gene expression of cells at different locations at the next depth, thereby reconstructing the slice space sequencing interpolation results at the next depth.

[0037] Furthermore, the optimization objective function of step 4) bidirectional interpolation optimization is as follows:

[0038]

[0039] Single layer loss :

[0040]

[0041] Where: , respectively represent the shape prediction network and expression prediction network during back diffusion; K represents the total number of slices along the z-axis; k is the index value, ∈ (1, K); The whole represents the positive loss term; represents the forward inference loss from the k-th layer slice to the k+1-th layer slice; The whole represents the reverse loss term; represents the reverse inference loss from the k-th layer slice to the k-1-th layer slice;

[0042] Represents shape alignment items; represents the Wasserstein distance, which measures the difference in the spatial distribution of cells between the predicted slice and the real slice; Indicates the depth Observed and predicted spatial distribution of cells at

[0043] represents the expression alignment item; represents the function that maps the i-th predicted cell to the observed cell with the closest spatial distance to it; represents a hyperparameter; represents the gene expression values ​​of cells in real slices; Represents the gene expression value predicted by the model.

[0044] Furthermore, step 5) specifically includes:

[0045] The interpolated slice data are mapped to voxel space, and the gene expression value of each voxel is the average of the expression values ​​of all cells it contains, generating a continuous three-dimensional volume V.

[0046] In a second aspect, an embodiment of the present invention further provides a three-dimensional spatial sequencing interpolation system based on a diffusion model, comprising:

[0047] Microenvironment map construction module, used to construct a cell microenvironment map of three-dimensional spatial sequencing data: for K slice data distributed along the z-axis, the KNN algorithm is used to determine the spatial neighborhood relationship of each cell, and a map containing spatial coordinates is generated. and gene expression Graph structure datasets;

[0048] A stochastic differential equation model module is established to model a stochastic differential equation model in which the cell position continuously changes with the z-axis depth based on the graph structure data set as a shape drift equation; and a stochastic differential equation model in which the gene expression continuously changes with the z-axis depth is modeled as a gene expression drift equation;

[0049] Train graph neural network module to build and train shape prediction networks :Input the spatial coordinates and spatial neighborhood relationship of the current slice, output the shape drift coefficient, substitute it into the shape drift equation, predict the position offset of the cell in the next slice; and build and train the expression prediction network : Input the spatial coordinates, gene expression and spatial neighborhood relationship of the current slice, output the expression drift coefficient, substitute it into the gene expression drift equation, and infer the change trend of gene expression;

[0050] The bidirectional interpolation optimization module uses the trained graph neural network to perform interpolation prediction, generates intermediate slices through forward and reverse interpolation, and jointly optimizes the Wasserstein distance of shape distribution and the mean square error loss of gene expression;

[0051] The 3D reconstruction module is used to integrate the interpolated slice data and generate a continuous 3D gene expression map.

[0052] It can be seen from the above technical solutions that compared with the prior art, the present invention has the following technical advantages:

[0053] The sequencing results of three-dimensional spatial omics will be combined to make full use of the spatial continuity and correlation of biological organs themselves, and a self-supervised interpolation algorithm based on the diffusion model will be designed to achieve continuous completion of z-axis slice data. Empowering precision medicine with three-dimensional expression data after continuous completion can assist in identifying normal-lesion tissue boundaries, finely analyzing gene expression gradients at different locations in diseased tissues, and distinguishing tissue subclasses that cannot be resolved with the original number of slices. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0055] Figure 1 A flow chart of the three-dimensional spatial sequencing interpolation method based on the diffusion model provided by the present invention.

[0056] Figure 2 This is a schematic diagram of constructing the cell microenvironment map provided by the present invention.

[0057] Figure 3 This is a schematic diagram of the predicted cell position and expression along the continuous change of the z-axis based on the diffusion model provided by the present invention.

[0058] Figure 4 The present invention provides a reconstruction accuracy comparison of existing data to determine Schematic diagram of the values.

[0059] Figure 5 A block diagram of the three-dimensional spatial sequencing interpolation system based on the diffusion model provided by the present invention. DETAILED DESCRIPTION

[0060] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0061] First, the technical terms involved in the present invention are explained as follows to facilitate a clearer understanding of the present invention:

[0062] 1. Spatial omics technology

[0063] The spatial location of transcriptomics has great academic value for in-depth analysis of complex physiological functions and pathological mechanisms. Spatial transcriptomics (ST) was selected as the 2020 method of the year by Nature Methods magazine, and has provided new technical support for research in areas such as tumor heterogeneity, gastrula embryonic development, and the principles of Alzheimer's disease. Therefore, large-scale projects such as the Human Cell Atlas and the Cell Census Network of the US Brain Project have invested heavily in using spatial transcriptomics technology to ultimately generate complete maps of large and complex tissues (such as the brain). Super-resolution reconstruction of spatial transcriptome information plays a vital role in exploring the fine expression patterns of high-throughput genes in tissues and exploring more complex joint expression relationships between genes, thereby assisting in a deeper understanding of life processes.

[0064] At the same time, three-dimensional spatial sequencing technology achieves a more complete characterization of gene expression in three-dimensional space by performing continuous slicing of an organ, tissue or organism along a certain axis. This plays an important role in revealing the three-dimensional complex structure of tissues, exploring tissue heterogeneity and accurately characterizing tumor structures.

[0065] 2. Diffusion Model

[0066] Diffusion Model Diffusion Model is a type of generative model based on probability model. In recent years, it has demonstrated strong capabilities in image generation, speech generation and other fields. The core idea of ​​this type of model is to map the complex real distribution to a simple prior distribution (such as Gaussian distribution) by gradually perturbing the data, and then gradually restore the original distribution from the simple distribution through the inverse process to generate new samples.

[0067] The diffusion model consists of two main processes:

[0068] Forward Process: Gradually add noise to the original data distribution to make it approach a known simple distribution (such as a standard Gaussian distribution).

[0069] Reverse Process: Learn the dynamics of the inverse process when training the model, simulating the path from noise to the original data distribution.

[0070] The mathematical essence of the diffusion model is built on the framework of stochastic differential equations (SDE), which provides a description of the dynamics of distribution changes and a design framework and optimization tools for modeling methods based on stochastic differential equations. The introduction of SDE enables the diffusion model to theoretically explain the training and generation process of the generative model and model it in the continuous time domain, which is more flexible than the discrete time multi-step denoising model.

[0071] The diffusion model has the following advantages:

[0072] High-quality generation: Through a step-by-step generation approach, the diffusion model is able to generate clear and diverse samples.

[0073] Stability: Using the SDE framework, model training can directly combine the dynamic characteristics of probability density function changes to avoid the pattern collapse problem that may occur in traditional generative models.

[0074] Theoretical support: Stochastic differential equations provide a strong mathematical foundation for diffusion models and have a natural connection with diffusion processes in physics.

[0075] In view of the fact that the contradiction between the "sparse observation" of three-dimensional spatial omics sequencing technology on the z-axis and the "fine observation" on the xy plane has not yet been resolved. In order to reconstruct a continuous three-dimensional expression map by modeling the continuity of biological organs and comprehensively utilizing the sequencing information of the xy plane on the basis of existing sequencing technology, using the prior knowledge of "continuity", an interpolation algorithm based on continuity modeling is crucial. Therefore, the embodiment of the present invention discloses a three-dimensional spatial sequencing interpolation method based on a diffusion model, referring to Figure 1 As shown, it includes the following steps 1)~5):

[0076] 1) Construct a cell microenvironment map of three-dimensional spatial sequencing data: For K slices of data distributed along the z-axis, use the KNN (K-Nearest Neighbor) algorithm to determine the spatial neighborhood relationship of each cell and generate a map containing spatial coordinates. and gene expression Graph structure datasets;

[0077] Reference Figure 2 As shown in the figure, in the three-dimensional spatial sequencing data containing K slices constructed by the KNN graph, each depth A graph-structured dataset of spatial neighborhood information under , where A is the connectivity matrix, which represents the neighboring relationship among cells in space; K is the number of observed slices; is the depth of the kth slice; is the number of cells contained in the k-th slice;

[0078] 2). Stochastic differential equations describing the continuity of tissue shape and gene expression;

[0079] Based on the graph structure data set, a stochastic differential equation model is modeled in which the cell position changes continuously with the z-axis depth as a shape drift equation; and a stochastic differential equation model is modeled in which the gene expression changes continuously with the z-axis depth as a gene expression drift equation;

[0080] The continuous changes of tissue morphology and gene expression in three-dimensional space are described by two SDEs: the random drift equation of morphological change is used to model the gradual transfer of cell positions, and the random drift equation of expression change is used to model the gradual change of gene expression in space.

[0081] 3). Training graph neural network:

[0082] a. Build and train a shape prediction network : Input the spatial coordinates and spatial neighborhood relationship of the current slice, and output the shape drift coefficient , substituted into the shape drift equation to predict the positional shift of the cell in the next slice;

[0083] Input at depth The spatial coordinate data of the cells contained in the slice of the xy plane And the corresponding connection matrix A, predicting the position offset coefficient of each cell through spatial information , used to predict the spatial position distribution of cells at the next depth, so as to reconstruct the slice shape at the next depth;

[0084] b. Build and train the expression prediction network : Input the spatial coordinates, gene expression and spatial neighborhood relationship of the current slice, and output the expression drift coefficient , substituted into the gene expression drift equation to infer the changing trend of gene expression;

[0085] Input at depth The spatial coordinate data and gene expression of cells contained in the slices of the xy plane And the corresponding connection matrix A, predicting the gene expression deviation coefficient of each cell through spatial information , which is used to predict the gene expression of cells at different locations at the next depth, thereby reconstructing the slice space sequencing interpolation results at the next depth.

[0086] This step uses a graph neural network (GNN) to learn the characteristic changes of cells and their neighborhoods in each slice. The present invention can effectively capture the spatial continuity of tissues and the dynamic characteristics of gene expression.

[0087] 4) Bidirectional interpolation optimization: Use the trained graph neural network to perform interpolation prediction. and reverse Interpolation generates intermediate slices, and jointly optimizes the Wasserstein distance of shape distribution and the mean square error loss of gene expression; the graph neural network is a multi-layer perceptron (MLP) or a graph convolutional network (GCN).

[0088] This step adopts a bidirectional inference strategy, that is, inferring from the beginning and end of the slice sequence respectively and comprehensively optimizing the results to further improve the accuracy and robustness of the model.

[0089] 5) Construct a three-dimensional transcriptional continuum: Integrate the interpolated slice data to generate a continuous three-dimensional gene expression map. That is, map the interpolated slice data to the voxel space, and the gene expression value of each voxel is the average of the expression values ​​of all cells it contains, generating a continuous three-dimensional volume V.

[0090] The three-dimensional spatial sequencing interpolation method based on the diffusion model proposed in this paper solves the problem that the tissue section sequencing data is sparse in the depth direction and it is difficult to construct continuous three-dimensional transcriptome features in the current spatial transcriptomics research. It realizes the accurate reconstruction of gene expression and morphology at the full depth of the tissue, providing a new technical means for the three-dimensional spatial analysis of tissues.

[0091] In a specific embodiment, in a three-dimensional (3D) spatially resolved transcriptomics dataset, a tissue is divided into K parallel slices for sequencing. Figure 3 As shown, in this embodiment, (x, y) is defined as the direction within each slice, and the z direction represents the tissue depth, and different slices are arranged along this direction. The spatial sequencing features of a two-dimensional slice are expressed as a function of the spatial position. ,in is the number of cells (or sites) in the slice, Indicates location The goal of this invention is to infer the depth of unsequenced tissue Spatial expression , thereby constructing a complete 3D transcriptional continuum for the tissue.

[0092] However, compared to the hundreds to tens of thousands of spatial locations densely measured in the (x, y) direction, the z direction typically contains only sparse depth values. This imbalance in spatial information makes it challenging to directly regress or interpolate new tissue levels.

[0093] Here, the present invention utilizes the inherent spatial continuity of tissues to convert unsequenced depth The slice at is modeled as a slice from an existing, adjacent depth The slices gradually expanded; is the interpolation step size. To achieve this, the present invention uses two stochastic differential equations (SDEs) to model tissue shape and gene expression in 3D space respectively:

[0094] The shape drift equation is as follows:

[0095] (1)

[0096] The gene expression drift equation is as follows:

[0097] (2)

[0098] in and Measured (depth ) and prediction (depth ) of the cell / point location, and and is its corresponding expression feature; drift coefficient and is a learnable gradient function that models the spatial variation of gene expression and tissue shape with depth, respectively; the superscript G is the gene expression dimension, and the superscript 3 is the dimension of the spatial coordinate; the diffusion term and It's white noise.

[0099] In small steps The first SDE models the distribution of cells in the (x, y) plane as a function of depth. The second SDE models the corresponding expression changes. Through deep accumulation of multiple steps, the SDE model gradually constructs the (near) continuous expression transition along the z-axis between the two sequenced layers, thus establishing the spatial continuum of transcriptomics.

[0100] To learn the gradient function, the present invention adopts a graph neural network (GNN) and Aggregate location Surrounding neighboring features predict tissue shape and gene expression at depth Transformation under change:

[0101] (3)

[0102] (4)

[0103] in Is Depth Position in slice All surrounding measured cells / points.

[0104] Specifically, the interpolation algorithm is as follows Figure 1 The process shown is carried out:

[0105] S1: Figure 2 As shown, according to the KNN algorithm and the spatial position relationship of each cell, the first-order jump neighbor of each cell and the corresponding connection matrix A are determined.

[0106] S2: Decomposing data into shape information and express information , respectively input into the graph neural network to obtain the prediction of the offset.

[0107] S3: Figure 1 As shown, forward prediction is performed to obtain the predicted slice position and expression after a small displacement.

[0108] S4: Figure 1 As shown, from sequenced tissue sections First, combine formulas (1) to (4) to slice Inference of organizational characteristics:

[0109] (5)

[0110] (6)

[0111] in From the depth arrive The number of steps required; m∈(1,M) is the index value; Indicates depth The spatial distribution of cells inferred at is the inferred gene expression value.

[0112] S5: Loss function design, constructing a loss function by aligning the shape and expression level of the predicted and sequenced data. Specifically, this example uses the Wasserstein distance to measure the difference in cell position distribution, and uses the mean square error (MSE) to measure the accuracy of expression features. The loss is defined as:

[0113] (7)

[0114] in, Represents a shape alignment item; shape alignment item represents the Wasserstein distance between the measured and predicted cell position distributions; and is the measured or inferred cell location distribution.

[0115] represents the expression alignment item; represents the function that maps the i-th predicted cell to the observed cell with the closest spatial distance to it; represents a hyperparameter; represents the gene expression values ​​of cells in real slices; Represents the gene expression value predicted by the model; in the expression alignment term, It is the distance inferred data position in the measured data The nearest cell location.

[0116] Hyperparameters Balance the contributions of the two terms, whose values ​​are determined empirically:

[0117] This embodiment determines the reconstruction accuracy of the existing data by comparing Take a set of mouse brain slice data as an example. Specifically, select different The 3D sequencing data of the mouse brain with 53 slices was interpolated, and only the 1st, 3rd, 5th...53rd slices were retained to interpolate and reconstruct the 2nd, 4th, 6th...52nd slices. The best value was determined according to the reconstruction accuracy. For example, we experimented with four alternative values: 10, 1, 0.1, and 0.01, and obtained the following Figure 4 The reconstruction accuracy shown (the ordinate is the mean square error, the abscissa is value), by Figure 4 It can be seen that the optimal value is 0.1 in this data set.

[0118] The total loss function is the sum of the losses of all layers. Towards The reverse inference of , the final loss function is defined as:

[0119] (8)

[0120] Where: , respectively represent the shape prediction network and expression prediction network during back diffusion; K represents the total number of slices along the z-axis; k is the index value, ∈ (1, K); The whole represents the positive loss term; represents the forward inference loss from the k-th layer slice to the k+1-th layer slice; The whole represents the reverse loss term; represents the reverse inference loss from the k-th layer slice to the k-1-th layer slice;

[0121] Optimize the above loss function, use NAdam optimizer for model optimization, and the learning rate is 0.01. To ensure accuracy and efficiency, a three-step strategy is adopted:

[0122] 1. Minimize the shape alignment term to accurately predict the spatial coordinates of the slice.

[0123] 2. Minimize the expression alignment terms to accurately match the actual transcriptome features.

[0124] 3. Minimize the overall loss function to achieve a comprehensive characterization of the slices.

[0125] S6: Continuum construction, the model generates a series of intermediate slices In order to integrate these slices into a complete 3D volume, this embodiment defines a volume V of size (L, W, H, G), where the size of each voxel is The expression value within each voxel is calculated by averaging the expression values ​​of cells located at the corresponding cubic position (x, y, z):

[0126] (9)

[0127] in, represents the collection of cells that lie within the voxel boundaries, yes The number of cells in .

[0128] The three-dimensional spatial sequencing interpolation method based on the diffusion model provided by the present invention has the following effects:

[0129] 1. Solved the problem of data sparseness in depth direction

[0130] The present invention can use a small amount of discrete depth slice data to generate high-resolution intermediate slices through the SDE model, thereby constructing continuous three-dimensional gene expression features at the full depth of the tissue. This method significantly improves the depth coverage of spatial transcriptomics research and provides important support for the comprehensive analysis of tissue morphology and function.

[0131] 2. Achieved high-precision 3D reconstruction

[0132] By introducing Wasserstein distance and mean square error as optimization targets for morphology and expression alignment, the present invention can not only accurately predict the spatial distribution of cells in unmeasured areas, but also generate gene expression features that are highly consistent with actual sequencing data. This high-precision three-dimensional reconstruction technology lays a solid foundation for subsequent biological analysis.

[0133] 3. Wide applicability and scalability

[0134] The method of the present invention is applicable to a variety of tissue types and transcriptomic data, and can be used to analyze the developmental process of different organs, pathological changes, and other spatial biological phenomena. In addition, the stochastic differential equation framework and graph neural network used in the model have good scalability and can be combined with more data features to further improve performance.

[0135] In summary, the present invention solves the problem of sparse spatial transcriptomics data in depth direction by innovatively using stochastic differential equations and graph neural network methods, and achieves high-precision reconstruction of three-dimensional gene expression characteristics. This method has significant theoretical value and application potential, and provides strong technical support for tissue function analysis and spatial biology research.

[0136] For example, the following is a detailed explanation based on the 3D reconstruction of mouse brain slice data based on 22 slices:

[0137] In order to verify the effectiveness of the present invention, we selected a set of mouse brain tissue data containing 22 slices as the research object. The data covers the expression information of 2000 genes, and in the experiment, the gene expression data was reduced to the first 50 principal components (PCs) through principal component analysis (PCA) to reduce the computational complexity while retaining the main features. The implementation steps and experimental results are described in detail below.

[0138] 1. Data preprocessing

[0139] First, the mouse brain slice data was standardized to remove technical noise and ensure the comparability of gene expression values ​​between slices. Each slice contains multiple spatial position points, each of which records the expression values ​​of 2000 genes. PCA analysis was performed on the expression matrix of all points, and the top 50 PCs were extracted as feature input to capture the main variations in gene expression.

[0140] 2. Constructing stochastic differential equation model

[0141] Based on the processed slice data, we used stochastic differential equations (SDE) to model the changes in cell position and gene expression, respectively:

[0142] 1) Cell position change: Use the morphological random drift equation to describe the gradual movement of cells in the x, y plane along the depth z direction;

[0143] 2) Changes in gene expression: The gene expression random drift equation is used to model the dynamic changes of gene expression in the tissue depth direction.

[0144] These two equations model the continuous changes of data by introducing a drift term (the gradient function learned by the graph neural network) and a diffusion term (Gaussian noise).

[0145] 3. Model training

[0146] In each slice, a graph structure containing neighboring spatial points is constructed and input into a graph neural network (GNN) to learn the spatial gradient changes of cell positions and gene expression. GNN outputs drift functions and , used to predict changes in cell location and gene expression. The model training objectives include two parts:

[0147] 1) Morphological alignment loss: through Wasserstein distance measure the difference between the actual measured and predicted cell position distributions;

[0148] 2) Expression alignment loss: The difference between actual measured and predicted gene expressions is measured by mean square error (MSE).

[0149] The optimization uses the NAdam algorithm with a learning rate of 0.001. The bidirectional inference strategy is used during training, that is, from the shallow To the Deep and reverse from deep To the shallow , and comprehensively optimize the results.

[0150] 4. Experimental Results

[0151] After the model is trained, it gradually generates intermediate slices of the unmeasured depth (z direction) area and constructs a complete three-dimensional gene expression feature. The following is the result analysis:

[0152] 1) Morphological reconstruction

[0153] The spatial distribution of cells in the experimentally generated intermediate slices is highly consistent with the actual sequenced slices. Using the Wasserstein distance as the evaluation index, the average shape error between the generated slices and the measured slices is less than 2%.

[0154] 2) Gene expression reconstruction

[0155] In the space of 50 principal components, the mean square error between the predicted expression and the actual sequencing expression was less than 0.05, indicating that the model can accurately capture the dynamic changes of gene expression.

[0156] 3) 3D volume construction

[0157] All slice data are integrated into a complete 3D volume V. The voxel size is set to The expression value of each voxel is calculated by the average gene expression of cells within the voxel. The reconstructed three-dimensional brain tissue shows the spatial continuity and dynamic changes of gene expression, which is consistent with biological laws.

[0158] Based on the same inventive concept, the present invention also provides a three-dimensional spatial sequencing interpolation system based on a diffusion model, referring to Figure 5 As shown, including:

[0159] Microenvironment map construction module, used to construct a cell microenvironment map of three-dimensional spatial sequencing data: for K slice data distributed along the z-axis, the KNN algorithm is used to determine the spatial neighborhood relationship of each cell, and a map containing spatial coordinates is generated. and gene expression Graph structure datasets;

[0160] A stochastic differential equation model module is established to model a stochastic differential equation model in which the cell position continuously changes with the z-axis depth based on the graph structure data set as a shape drift equation; and a stochastic differential equation model in which the gene expression continuously changes with the z-axis depth is modeled as a gene expression drift equation;

[0161] Train graph neural network module to build and train shape prediction networks :Input the spatial coordinates and spatial neighborhood relationship of the current slice, output the shape drift coefficient, substitute it into the shape drift equation, predict the position offset of the cell in the next slice; and build and train the expression prediction network : Input the spatial coordinates, gene expression and spatial neighborhood relationship of the current slice, output the expression drift coefficient, substitute it into the gene expression drift equation, and infer the change trend of gene expression;

[0162] The bidirectional interpolation optimization module uses the trained graph neural network to perform interpolation prediction, generates intermediate slices through forward and reverse interpolation, and jointly optimizes the Wasserstein distance of shape distribution and the mean square error loss of gene expression;

[0163] The 3D reconstruction module is used to integrate the interpolated slice data and generate a continuous 3D gene expression map.

[0164] In this embodiment, the microenvironment map construction module uses the KNN algorithm to capture the cell neighborhood relationship, and combines the graph neural network to learn the spatial gradient, which effectively solves the problem that the traditional method ignores the continuity of biological tissues, making the interpolation results closer to the real structure. The model training module introduces bidirectional interpolation (forward + reverse) and a joint loss function (Wasserstein distance + MSE), which reduces the error accumulation of unidirectional predictions through double verification and significantly improves the reconstruction accuracy. The three-dimensional reconstruction module integrates interpolated slice data to generate a continuous three-dimensional gene expression map, which can clearly analyze complex spatial features such as tumor boundaries and brain functional zoning, and support precision medicine and biological research. The system has wide applicability and scalability, is compatible with a variety of tissue types (such as brain, tumors, embryos), and supports the expansion of more data features (such as proteomics), providing a general tool for cross-organ and cross-scale spatial biology research.

[0165] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0166] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A three-dimensional spatial sequencing interpolation method based on a diffusion model, characterized in that: The following steps are involved: 1) Construct a cell microenvironment map of three-dimensional spatial sequencing data: For K slices of data distributed along the z-axis, use the KNN algorithm to determine the spatial neighborhood relationship of each cell and generate a map containing spatial coordinates. and gene expression Graph structure datasets; 2). Establish a stochastic differential equation model: Based on the graph structure data set, a stochastic differential equation model is modeled in which the cell position changes continuously with the z-axis depth as a shape drift equation; And a stochastic differential equation model of gene expression that continuously changes with z-axis depth is modeled as a gene expression drift equation; 3). Training graph neural network: a. Build and train a shape prediction network : Input the spatial coordinates and spatial neighborhood relationship of the current slice, output the shape drift coefficient, substitute it into the shape drift equation, and predict the position offset of the cell in the next slice; b. Build and train the expression prediction network : Input the spatial coordinates, gene expression and spatial neighborhood relationship of the current slice, output the expression drift coefficient, substitute it into the gene expression drift equation, and infer the change trend of gene expression; 4) Bidirectional interpolation optimization: Use the trained graph neural network to perform interpolation prediction, generate intermediate slices through forward and reverse interpolation, and jointly optimize the Wasserstein distance of shape distribution and the mean square error loss of gene expression; 5). Construct a three-dimensional transcriptional continuum: Integrate interpolated slice data to generate a continuous three-dimensional gene expression map.

2. A three-dimensional spatial sequencing interpolation method based on a diffusion model according to claim 1, characterized in that: Step 1) specifically includes: In a three-dimensional spatially resolved transcriptomic dataset, a piece of tissue is divided into K parallel sections for sequencing; For K slice data distributed along the z axis, (x, y) is defined as the direction within each slice, and the z direction represents the tissue depth, and different slices are arranged along this direction; For the depth A two-dimensional slice of Calculate the K nearest neighbors and generate a connection matrix A as the spatial neighborhood relationship of each cell, and form a matrix containing spatial coordinates and gene expression Graph structure dataset ;in, is the number of cells or sites in the slice, Indicates location The gene expression values.

3. A three-dimensional spatial sequencing interpolation method based on a diffusion model according to claim 2, characterized in that: In step 2): The shape drift equation is as follows: (1) The gene expression drift equation is as follows: (2) in, and The measured depth and predicted depth The cell or location point location, and The corresponding expression characteristics; Indicates interpolation step size; drift coefficient and is a learnable gradient function that models the spatial variation of gene expression and tissue shape with depth, respectively, G is the gene depth; the diffusion term and It's white noise.

4. The three-dimensional spatial sequencing interpolation method based on a diffusion model according to claim 3, characterized in that: In step 3), build and train the shape prediction network ,include: Using graph neural network GNN to build shape prediction network , and conduct training; Input at Depth The spatial coordinate data of the cells contained in the slice of the xy plane And the corresponding connection matrix A, predicting the position offset coefficient of each cell through spatial information , used to predict the spatial position distribution of cells at the next depth, so as to reconstruct the slice shape at the next depth; In step 3), construct and train the expression prediction network ,include: Using graph neural network GNN to build expression prediction network , and conduct training; Input at Depth The spatial coordinate data and gene expression of cells contained in the slices of the xy plane And the corresponding connection matrix A, predicting the gene expression deviation coefficient of each cell through spatial information , which is used to predict the gene expression of cells at different locations at the next depth, thereby reconstructing the slice space sequencing interpolation results at the next depth.

5. The three-dimensional spatial sequencing interpolation method based on a diffusion model according to claim 4, characterized in that: Step 4) The optimization objective function of bidirectional interpolation optimization is as follows: ; Single layer loss : ; Where: , respectively represent the shape prediction network and expression prediction network during back diffusion; K represents the total number of slices along the z-axis; k is the index value, ∈ (1, K); The whole represents the positive loss term; represents the forward inference loss from the k-th layer slice to the k+1-th layer slice; The whole represents the reverse loss term; represents the reverse inference loss from the k-th layer slice to the k-1-th layer slice; Represents shape alignment items; represents the Wasserstein distance, which measures the difference in the spatial distribution of cells between the predicted slice and the real slice; Indicates the depth Observed and predicted spatial distribution of cells at represents the expression alignment item; represents the function that maps the i-th predicted cell to the observed cell with the closest spatial distance to it; represents a hyperparameter; represents the gene expression values ​​of cells in real slices; Represents the gene expression value predicted by the model.

6. The three-dimensional spatial sequencing interpolation method based on a diffusion model according to claim 1, characterized in that: Step 5) specifically includes: The interpolated slice data are mapped to voxel space, and the gene expression value of each voxel is the average of the expression values ​​of all cells it contains, generating a continuous three-dimensional volume V.

7. A three-dimensional spatial sequencing interpolation system based on a diffusion model, characterized in that: include: Microenvironment map construction module, used to construct a cell microenvironment map of three-dimensional spatial sequencing data: for K slice data distributed along the z-axis, the KNN algorithm is used to determine the spatial neighborhood relationship of each cell, and a map containing spatial coordinates is generated. and gene expression Graph structure datasets; Establishing a stochastic differential equation model module, and modeling a stochastic differential equation model in which the cell position continuously changes with the z-axis depth based on the graph structure data set as a shape drift equation; And the stochastic differential equation model of gene expression changing continuously with z-axis depth was modeled as the gene expression drift equation; Train graph neural network module to build and train shape prediction networks : Input the spatial coordinates and spatial neighborhood relationship of the current slice, output the shape drift coefficient, substitute it into the shape drift equation, and predict the position offset of the cell in the next slice; And for building and training expression prediction networks : Input the spatial coordinates, gene expression and spatial neighborhood relationship of the current slice, output the expression drift coefficient, substitute it into the gene expression drift equation, and infer the change trend of gene expression; The bidirectional interpolation optimization module uses the trained graph neural network to perform interpolation prediction, generates intermediate slices through forward and reverse interpolation, and jointly optimizes the Wasserstein distance of shape distribution and the mean square error loss of gene expression; The 3D reconstruction module is used to integrate the interpolated slice data and generate a continuous 3D gene expression map.

Citation Information

Patent Citations

  • Tumor virtual three-dimensional transcriptome construction method and device, equipment and storage medium

    CN116129999A

  • Missing value interpolation method for single cell sequencing data

    CN118335191A