3D tissue reconstruction method and device based on space transcriptome data

By standardizing and screening spatial transcriptome data, constructing spatial kernel functions and target models, the accuracy problem of 3D gene expression map construction in existing technologies is solved, and high-precision 3D tissue reconstruction and visualization are achieved.

CN121459925APending Publication Date: 2026-02-03XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511568891.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately correct batch effects and preserve genuine biological differences in complex data, resulting in insufficient accuracy in constructing 3D gene expression maps and an inability to achieve high-precision 3D tissue reconstruction.

Method used

By acquiring spatial transcriptome data, standardizing and screening gene expression data, constructing spatial kernel functions to obtain spatial correlation strength matrices between cells or sites, combining initial parameter values ​​and weights, constructing target models for integration, and iteratively optimizing to obtain high-resolution reconstruction results.

Benefits of technology

It improves the accuracy of the integrated results, achieves high-precision 3D gene expression map reconstruction, can accurately construct spatial clusters of tissues, identify differentially expressed genes and infer cell development trajectories, and provide high-resolution visualization of tissue structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121459925A_ABST
    Figure CN121459925A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D tissue reconstruction method and device based on spatial transcriptome data. Comprising the following steps: acquiring space transcriptome data in space-time omics research, and carrying out standardization processing on the space transcriptome data; constructing a space correlation intensity matrix according to the position information of the cells distributed in the tissue space; calculating an initial value parameter of the model, constructing a target model, and determining the weight of a spatial correlation effect constraint space factor; after determining the weight of the initial parameter and the spatial correlation effect constraint space factor, inputting a target model and carrying out iterative optimization to finally obtain an integration result; and finally, performing downstream analysis, such as spatial clustering, differential gene identification and cell development trajectory inference, on the basis of an integration result (a spatial factor), and performing high-resolution reconstruction on low-resolution spatial transcriptome data (the weight of the spatial factor is constrained by the correlation effect of the spatial factor and the space), so that high-precision tissue structure 3D visualization is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics analysis technology, and in particular to a 3D tissue reconstruction method and apparatus based on spatial transcriptome data. Background Technology

[0002] With the rapid development of spatial transcriptomics technology, constructing 3D atlases based on spatial transcriptomics data has become a research hotspot in the life sciences. From a data analysis perspective, the construction of 3D atlases relies on the precise integration of spatial transcriptomics samples from consecutive tissue sections to further elucidate the correlation mechanism between the spatial distribution of cells or sites within tissues and gene expression.

[0003] Current analytical methods primarily focus on "correcting batch effects" to achieve integration. However, as the complexity of sample design increases, the proportion of batch effects and biological differences in a set of spatiotemporal omics samples fluctuates significantly. This makes "distinguishing between the two types of differences and accurately correcting them" a core challenge in integration analysis. Secondly, most integration methods model the spatial characteristics of a cell / site and several surrounding cells / sites as the spatial features of that site / cell, failing to effectively constrain the spatial correlation between cells or sites. This leads to a significant decrease in their suitability for "complex data" lacking sufficient research. Such "complex data" typically exhibits low spatial continuity, unclear tissue anatomical boundaries, and large spatial distribution differences between samples. Existing methods struggle to meet the integration requirement of "accurately removing batch effects while fully preserving true biological differences," resulting in insufficient accuracy in subsequent spatial clustering, differential gene identification, and cell development trajectory inference, making it impossible to accurately construct high-precision 3D atlases.

[0004] Therefore, how to improve the accuracy of the integration results and thus achieve precise reconstruction of 3D gene expression maps is an urgent problem to be solved. Summary of the Invention

[0005] In view of this, the 3D tissue reconstruction method and apparatus based on spatial transcriptome data provided in this application can improve the accuracy of the integration results, thereby achieving precise reconstruction of 3D gene expression maps. The 3D tissue reconstruction method and apparatus based on spatial transcriptome data provided in this application are implemented as follows: This application provides a 3D tissue reconstruction method based on spatial transcriptome data, comprising: Acquire spatial transcriptome data, which includes gene expression data and spatial location data; wherein, each sample includes multiple cells or sites, the gene expression data is used to record the expression level of different genes in each sample at the corresponding cells or sites, and the spatial location data is used to record the two-dimensional spatial coordinates of the corresponding cells or sites in each sample; The gene expression data is standardized and gene screening is performed to obtain processed gene expression data; and a spatial kernel function is constructed based on the spatial location data, and a spatial correlation strength matrix between cells or sites within the same sample is obtained based on the spatial kernel function. Based on the processed gene expression data and the spatial correlation strength matrix, a target model is constructed. The target model is used to decompose the gene expression data into the basic gene expression level of each sample, the expression weight of the gene in each sample, the spatial factor of each sample, the cross-sample common load, and the expression residual of each sample. Obtain initial parameter values, and determine the weight of the spatial correlation constraint spatial factor in each sample based on the initial parameter values; The initial parameter values ​​and the weights of the spatial correlation constraint factors in each sample are input into the target model to obtain the integrated result; Based on the integration results, the spatial transcriptome data is reconstructed at high resolution to obtain processed spatial transcriptome data, and the reconstructed 3D tissue is obtained based on the processed spatial transcriptome data.

[0006] In some embodiments, the standardization and gene screening of the gene expression data to obtain processed gene expression data; and the construction of a spatial kernel function based on the spatial location data, and the obtaining of a spatial correlation strength matrix between cells or sites within the same sample based on the spatial kernel function, include: Log-normalization and scaling are performed on the expression level of each cell or site in the gene expression data to obtain the processed gene expression data; Obtain gene expression data from processed gene expression data where spatial differences exceed a preset spatial difference significance and expression levels exceed a preset expression level; A spatial kernel function is constructed based on the spatial location data. The spatial kernel function is used to characterize the spatial interaction relationship between cells or sites within the same sample. The spatial correlation matrix between each pair of cells or sites within the same sample is calculated based on the spatial kernel function.

[0007] In some embodiments, constructing a spatial kernel function based on the spatial location data includes: Obtain the two-dimensional spatial coordinates of all cells or sites within the same sample; Calculate the spatial distance between each pair of cells or sites within the current sample, the spatial distance being determined based on two-dimensional spatial coordinates; Determine the type of kernel function, which is used to map the calculated spatial distance to the interaction strength between cells or sites; The mapping rule of the kernel function is determined, and a spatial kernel function is obtained according to the type of the kernel function and the mapping rule of the kernel function. The mapping rule is that the closer the cells or sites are, the higher the interaction strength, and the farther the cells or sites are, the lower the interaction strength.

[0008] In some embodiments, obtaining initial parameter values ​​and determining the weights of spatial correlation constraint spatial factors in each sample based on the initial parameter values ​​includes: Obtain initial parameter values, which include the basic gene expression level of each sample, the gene expression weight in each sample, the cross-sample common load, the variance of the expression residuals of each sample, the overall variance of the spatial factor, and the weight of the spatial factor constrained by the spatial correlation effect in each sample. A target model is constructed based on the initial parameter values, and the weights of the spatial factors constraining spatial correlation are determined based on the target model.

[0009] In some embodiments, inputting the initial parameter values ​​and the weights of the spatial correlation constraint factors in each sample into the target model to obtain the integrated result includes: Based on the EM algorithm, the expectation and variance of the conditional distribution are calculated as the spatial factor estimate and estimated variance; The estimated value is obtained by taking the partial derivative of the parameters with respect to the spatial factor estimate and the log-likelihood function of the expected complete data. Iteratively calculate the lower bound of the log-likelihood of the current estimated parameters. If the convergence condition is met, stop the iteration and obtain the integrated result.

[0010] In some embodiments, the high-resolution reconstruction of the spatial transcriptome data based on the integration result to obtain processed spatial transcriptome data includes: Based on the spatial factors of cells or sites in each sample, a clustering algorithm is used to group cells or sites in multiple samples simultaneously to obtain spatial clustering results. Differential gene analysis was performed based on the spatial clustering results to obtain the specific gene expression characteristics of different spatial domains; Spatial developmental trajectories are constructed based on spatial factors of each sample cell or site, and key transition nodes are identified. Processed spatial transcriptome data are obtained based on key transition nodes and specific gene expression characteristics.

[0011] This application provides a 3D tissue reconstruction device based on spatial transcriptome data, comprising: The acquisition module is used to acquire spatial transcriptome data, which includes gene expression data and spatial location data. Each sample includes multiple cells or sites. The gene expression data is used to record the expression levels of different genes in each sample at the corresponding cells or sites, and the spatial location data is used to record the two-dimensional spatial coordinates of the corresponding cells or sites in each sample. The processing module is used to standardize and screen the gene expression data to obtain processed gene expression data; and to construct a spatial kernel function based on the spatial location data, and to obtain a spatial correlation strength matrix between cells or sites within the same sample based on the spatial kernel function. The construction module is used to construct a target model based on the processed gene expression data and the spatial correlation strength matrix. The target model is used to decompose the gene expression data into the basic gene expression level of each sample, the expression weight of the gene in each sample, the spatial factor of each sample, the cross-sample common load, and the expression residual of each sample. The acquisition module is also used to acquire initial parameter values ​​and determine the weight of the spatial correlation constraint spatial factor in each sample based on the initial parameter values. The processing module is also used to input the initial parameter values ​​and the weights of the spatial correlation constraint spatial factors in each sample into the target model to obtain the integrated result; The processing module is further configured to perform high-resolution reconstruction processing on the spatial transcriptome data based on the integration result, to obtain processed spatial transcriptome data.

[0012] The computer device provided in this application includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the method described in this application.

[0013] The computer-readable storage medium provided in this application embodiment stores a computer program thereon, which, when executed by a processor, implements the method described in this application embodiment.

[0014] This application provides a 3D tissue reconstruction method and apparatus based on spatial transcriptome data. It utilizes spatial transcriptome data acquired from spatiotemporal omics studies, including gene expression data and spatial location data of cells or loci. The gene expression data is standardized and screened for spatially highly variable genes to obtain processed gene expression data. A spatial kernel function is constructed based on the spatial location data to obtain a spatial correlation strength matrix between cells or loci within the same sample. A model is constructed based on the processed gene expression data and the spatial correlation strength matrix. Initial parameter values ​​are first calculated to construct a target model and determine the weights of spatial factors constraining spatial correlation in each sample. After determining the initial parameters and the weights of the spatial factors constraining spatial correlation, the data are input into the model and iteratively optimized to obtain the final integrated result. Downstream analyses, such as spatial clustering, differential gene identification, and cell development trajectory inference, are performed based on the spatial factors in the integrated result to reconstruct low-resolution spatial transcriptome data at high resolution, thereby achieving high-precision 3D visualization of tissue structures and solving the technical problems mentioned in the background art. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 A schematic diagram illustrating the implementation process of a 3D tissue reconstruction method based on spatial transcriptome data, provided in an embodiment of this application; Figure 2 This application provides a schematic diagram of the implementation process for gene expression data and spatial location data processing. Figure 3 This is a schematic diagram of a 3D tissue reconstruction device based on spatial transcriptome data, provided in an embodiment of this application. Detailed Implementation

[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0018] The following description of some technologies involved in the embodiments of this application is provided to aid understanding and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, some descriptions of well-known functions and structures are omitted in the following description.

[0019] Figure 1 This is a schematic diagram illustrating the implementation flow of a 3D tissue reconstruction method based on spatial transcriptome data provided in this application embodiment, including steps 101 to 105. Wherein, Figure 1 This is merely one execution order shown in the embodiments of this application and does not represent the only execution order for a 3D tissue reconstruction method based on spatial transcriptome data. Where the final result can be achieved, Figure 1 The steps shown can be performed in parallel or in reverse order.

[0020] Step 101: Obtain spatial transcriptome data in spatiotemporal omics research. Spatial transcriptome data includes gene expression data and spatial location data of multiple samples. Each sample includes multiple cells or sites. Gene expression data is used to record the expression level of different genes in each sample at the corresponding cells or sites, and spatial location data is used to record the spatial coordinates of the corresponding cells or sites in each sample.

[0021] In this embodiment of the application, multiple samples to be integrated are determined according to the research objectives; wherein, each sample can identify several cells or sites, and gene expression data records the expression level of the target gene in the corresponding cell or site (i.e., the number of gene transcripts) in each sample; spatial location data uses the preset position of the sequencing vector as the origin of the coordinate system to record the two-dimensional spatial coordinates of each cell or site, which is used to characterize the spatial distribution of cells or sites in the sample tissue.

[0022] Step 102 involves standardizing and screening the gene expression data to obtain processed gene expression data; and constructing a spatial kernel function based on the spatial location data to obtain a spatial correlation strength matrix between cells or sites within the same sample based on the spatial kernel function.

[0023] In the embodiments of this application, for the gene expression levels of all cells or sites in each sample, the expression levels are first incremented by 1 (to avoid the influence of zero values), and then logarithmic transformation is performed to eliminate the interference of different sequencing depths of different samples on the expression data, so as to obtain the transformed gene expression data.

[0024] The spatial distribution differences of each gene within the sample were evaluated by spatial hypervariable gene analysis. Genes with spatial differences reaching the preset significance standard were screened and ranked according to significance. At the same time, the average expression level of the selected spatial hypervariable genes in all cells or sites was calculated and the genes were ranked. The top 3000 genes (the specific parameters need to be adjusted according to the spatial omics data type) were selected by combining the first two rankings of each screened gene in each sample to obtain the processed gene expression data.

[0025] A spatial kernel function is constructed based on the spatial location data (two-dimensional spatial coordinates of cells or sites) of each sample. Following the calculation method and parameter rules of the exponential kernel function, the spatial kernel function is obtained, transforming the two-dimensional spatial coordinates of each sample into a spatial association strength matrix between cells or sites. The core logic of this kernel function is: the closer the spatial distance between cells or sites within the same sample, the higher the corresponding spatial interaction strength; the farther the spatial distance, the lower the corresponding spatial interaction strength, thus quantifying the spatial association characteristics of cells or sites.

[0026] Step 103: Construct a model based on the processed gene expression data and spatial correlation strength matrix. This model decomposes the gene expression data into the baseline gene expression level of each sample, the gene expression weight in each sample, the spatial factor of each sample, cross-sample common loads, and the gene expression residuals of each sample. First, initial parameter values ​​are calculated to construct the target model and determine the weights of the spatial factors constrained by spatial correlation in each sample.

[0027] In this embodiment, a target model is constructed based on the processed gene expression data and the spatial correlation strength matrix of each sample. The target model decomposes the gene expression level at cells or sites into the following four parts to determine the weights of the spatial factors constrained by spatial correlation in each sample: Basal expression level of a sample: The average expression level of a gene in all cells or loci of a sample, reflecting the basal expression characteristics of the gene in that sample; The combination of sample expression weights, spatial factors, and common factor loadings: "Sample expression weights" are sample-specific coefficients used to adjust the contribution of gene expression in the sample; "spatial factors" are sample-specific variables that reflect the combined constraints of spatial correlation and non-spatial correlation effects on cells or sites in the sample; "factor loadings" are common to all samples and are used to characterize the association strength between genes and spatial factors; the combination of the three reflects the comprehensive influence of sample specificity and spatial location on gene expression. Expression residuals for each sample: the difference between the actual gene expression level and the model-predicted expression level (obtained by adding a combination term to the baseline expression level), used to reflect expression fluctuations not yet explained by the model; Spatial correlation constraint: By using the spatial correlation strength matrix, the spatial factor differences between adjacent cells or sites (with high interaction weights) are controlled within a reasonable range to ensure that the spatial factors conform to the actual spatial distribution patterns of cells or sites.

[0028] Step 104: Set the initial values ​​of the parameters of the target model to construct the target model and determine the weights of the spatial correlation constraint spatial factors in each sample; after determining the initial parameters and the weights of the spatial correlation constraint spatial factors, input them into the target model and iteratively optimize to finally obtain the integrated result.

[0029] In the embodiments of this application, the parameters of the target model include the basic gene expression level of each sample, the gene expression weight in each sample, the cross-sample common factor loading, the variance of the expression residuals of each sample, the overall variance of the spatial factors, and the weight of the spatial factors constrained by the spatial correlation effect in each sample.

[0030] The initial value of the basal expression level of the sample is set as the average expression level of the gene in the sample in the gene expression data after treatment.

[0031] The initial values ​​of the sample representation weights are set to a uniform benchmark value (such as a unit value) as the starting point for subsequent adjustments.

[0032] The initial values ​​of factor loadings are set as a matrix that reflects the initial association between genes and spatial factors. The factor loadings in the principal component analysis results of the merged gene expression matrix are used as the initial values, and the variance of each principal component is used as the total variance of the spatial factors.

[0033] The initial value of the variance of the expression residuals of each sample is set as the variance of the gene expression level in the gene expression data of each sample after treatment.

[0034] The initial weights of the spatial correlation constraint spatial factors in each sample are calculated based on the average expression of each sample. The sample with the highest average expression has a weight of 0.5, and the weights of other samples are allocated according to the proportion of their average expression value to the maximum value.

[0035] The initial parameters are input into the model to construct the target model. Based on the conditional distribution characteristics of the spatial factors, the parameter values ​​of the target model are optimized and adjusted. The weights of the spatial factors constrained by spatial correlation in each sample are iteratively estimated and continuously updated until the weights of each sample no longer change, so as to determine the weights of the spatial factors constrained by spatial correlation in the sample.

[0036] This application also designs a simplified model of the target model. The target model is used to decompose gene expression data into the basal gene expression level of each sample, the expression weight of genes in each sample, the spatial factor of each sample, the cross-sample common loading, and the expression residual of each sample. The simplified model decomposes gene expression data into the cross-sample common loading, the spatial factor of each sample, and the expression residual of each sample. Among them, the simplified model can determine the weight of the spatial factor constrained by the spatial correlation in the samples more quickly and stably.

[0037] After determining the weights of spatial factors constrained by spatial correlation in the sample, the initial parameter values ​​and these weights are input into the model. The steps of "obtaining the conditional distribution characteristics of spatial factors → updating model parameters" are repeated until the model meets the preset convergence condition (the lower bound of log-likelihood is less than 0.0005). When the convergence condition is met, the iteration stops, the model at this time is determined as the target model, and the best estimates of the spatial factors (each cell or site corresponds to a unique spatial factor value) and other parameters corresponding to the target model are obtained.

[0038] Step 105: Based on the target model and spatial factors, perform spatial clustering, differential gene identification, and cell development trajectory inference of the samples to obtain the analysis results, which include spatial clustering results, differential gene identification results, and cell development trajectory inference results.

[0039] In this embodiment of the application, based on the best estimates of the spatial factors and other parameters corresponding to the target model, the following downstream analysis is carried out to obtain comprehensive analysis results.

[0040] 1. Spatial clustering processing: The algorithm based on Gaussian finite mixture model (Mclust) is used to cluster the spatial factors of all samples simultaneously, and the cells or sites are grouped to obtain spatial clustering results, which characterize the spatial classification features of the sample tissue.

[0041] 2. Differential gene identification and processing: Based on the spatial clustering results, compare the gene expression levels of samples under different conditions or in different spatial domains, screen genes with significant differences in expression levels through the differential significance test, and associate the spatial distribution information of genes to obtain the specific gene expression characteristics of different spatial domains.

[0042] 3. Cell development trajectory inference processing: Based on the dynamic change characteristics of spatial factors, the trajectory inference algorithm is used to sort cells or sites according to developmental or differentiation logic, construct dynamic change trajectories, determine the key transition stages and corresponding characteristic genes in the trajectory, and obtain the cell development trajectory inference results.

[0043] 4. High-resolution reconstruction processing: Based on the final results of the target model, the low-resolution spatial factors are reconstructed into high-resolution spatial factors for the data obtained by low-resolution spatial transcriptomics technology (non-single-cell resolution), thereby obtaining high-resolution spatial gene expression and spatial view.

[0044] This application's embodiments constrain spatial factors by constructing a spatial kernel function, spatial correlation effects, and non-spatial correlation effects. This deeply integrates gene expression data with the two-dimensional spatial coordinates of cells or sites, fully preserving the spatial correlation characteristics of cells or sites within the tissue, clarifying the sources of differences between samples, and distinguishing between "corrected batch effects" and "preserved biological differences," thus improving the accuracy and interpretability of the integrated results. The model decomposes gene expression levels into three parts: basal expression, weight-factor loading-spatial factor combination, and residuals. This considers both sample specificity and potential correlations between samples, while reducing technical error interference through spatial constraints. Subsequent parameter iteration updates and convergence judgments further optimize the model, ensuring the accuracy of the target model and spatial factors. Simultaneously, it achieves spatial clustering, differential gene identification, and cell development trajectory inference, covering three core analytical dimensions: "spatial distribution, gene differences, and developmental dynamics." This provides comprehensive and relevant analytical evidence for interpreting tissue spatial structure, cell functional zoning, disease progression, or cell differentiation mechanisms.

[0045] Furthermore, it should be noted that compared to existing spatial transcriptome data integration methods, this invention has a wider applicability and stronger interpretability in high-precision 3D tissue reconstruction studies in spatiotemporal omics. It employs both spatial correlation and non-spatial correlation effects within samples to constrain the potential associations of cells / sites in multiple samples, flexibly estimating the degree of heterogeneity that may exist in different multi-sample designs (from low to high). For example, samples with long spatial spans within the same time point exhibit high spatial heterogeneity but similar gene expression patterns; samples within the same tissue space exhibit low spatial heterogeneity at different time points but show significant differences in the expression of time-dependent genes (such as cell cycle genes and developmental regulatory genes). While integrating and aligning spatial cells / sites within samples, it preserves potential biological differences between samples, enabling accurate interpretation of spatiotemporal omics differences and related mechanisms. Simultaneously, it considers the potential batch effect of the same gene expression in different samples, achieving more precise batch effect removal. This results in the integration of spatiotemporal omics samples with dynamically changing gene expression patterns (the roles of time-dependent genes vary at different time points), generating research results that are biologically interpretable and closer to real biological states, providing reliable data support for subsequent differential analysis and mechanism interpretation.

[0046] Compared to graph-based deep learning methods, this invention can directly generate high-resolution spatial views and transcriptome data that conform to its characteristics based on the results of integrated analysis. Currently used spatial transcriptome techniques (such as 10X Visium) have limitations such as small capture range and non-single-cell resolution. This method can further predict gene expression characteristics and spatial location of undetected locations based on gene expression data and spatial coordinate information of the captured region. Specifically, by generating spatial loci with a larger coverage area at the original resolution, or fine spatial loci within the original region at a higher resolution, a larger-scale or higher-resolution spatial expression view can be reconstructed, thereby overcoming the limitations of the original technology. Figure 1 Based on the above, this application also provides a schematic diagram of the implementation process for gene expression data and spatial location data processing, as shown below. Figure 2 As shown, steps 201 to 203 are included: Step 201: Log-normalize and scale the expression level of each cell or site in the gene expression data to obtain the converted gene expression data.

[0047] In this embodiment, for the gene expression data of each acquired sample, the interference of sequencing depth differences on the data is first eliminated. For the expression level of each gene in each cell or locus, the value is first incremented by 1, and then a logarithmic transformation is performed. The core purpose of this transformation is to convert the raw count data of gene expression levels into a numerical form that better conforms to statistical analysis rules, reducing the impact of extremely high expression values ​​on the overall data distribution. Finally, the transformed gene expression data is obtained after scaling, ensuring the comparability of gene expression levels between different samples, cells, or loci.

[0048] Step 202: Obtain the expression data of the top 3000 genes in the converted gene expression data whose spatial differences exceed the preset spatial differences and whose expression levels exceed the preset expression levels.

[0049] In this embodiment, a spatially hypervariable gene analysis method (SPARK or SPARK-X) is used to calculate the spatial significance of each gene. The calculation results are compared with a preset spatial significance standard, and genes with spatial significance exceeding the standard are retained and ranked according to significance. The expression levels of such genes show significant distribution differences in cells or sites at different spatial locations within the sample, and are more likely to carry spatially relevant biological information.

[0050] The average expression level of each spatially hypervariable gene in all cells or loci of the corresponding sample was calculated and ranked, and genes with average expression levels exceeding the threshold were retained. These genes have stronger expression signals, which can reduce the interference of background noise from low-expression genes on subsequent analyses.

[0051] Genes were ranked based on both spatial dissimilarity and expression level, and the top 3000 genes were selected. Based on this, the gene set of all samples was unified, and genes with no spatial distribution or weak expression signals were removed, ultimately yielding the processed gene expression data.

[0052] Step 203: Construct a spatial kernel function based on spatial location data. The spatial kernel function is used to characterize the spatial correlation strength between cells or sites within the same sample.

[0053] Two-dimensional spatial coordinates of all cells or loci within the same sample are obtained. A differential threshold is calculated for the expression level of each gene in each sample, and a weighted value is calculated based on the expression ratio of each gene as the bandwidth of the kernel function. The "bandwidth" or "smoothing" of the kernel is controlled to determine the extent to which each data point affects the surrounding area. The spatial kernel function is obtained according to the calculation method and parameter calculation rules of the exponential kernel function, and the two-dimensional spatial coordinates of each sample are converted into a spatial correlation strength matrix between cells or loci.

[0054] This application's embodiments construct a spatial kernel function and generate a spatial correlation strength matrix, transforming the spatial positional relationships of cells or sites into calculable values. This accurately quantifies the spatial correlation strength between cells or sites within the same sample, providing direct data support for the spatial correlation weights of constraining spatial factors in the model and solving the problem of traditional methods neglecting global spatial interactions. The first half processes gene expression data, providing high-quality input for model construction; the second half generates variables constrained by both spatial correlation and non-spatial correlation effects, providing the foundation for the model's spatial constraint logic. This effectively connects the "data preprocessing" and "model construction" stages, ensuring the coherence and logic of the entire analysis process.

[0055] In some embodiments, constructing a spatial kernel function based on spatial location data includes: obtaining the two-dimensional spatial coordinates of all cells or sites within the same sample.

[0056] Specifically, spatial location data of the acquired individual samples is extracted. This data records the two-dimensional spatial coordinates of each cell or site within the sample on the sequencing vector.

[0057] When acquiring coordinates, the origin is set at a preset fixed position on the sequencing vector. The two dimensions of the two-dimensional coordinates correspond to the horizontal and vertical directions of the vector (e.g., horizontal is the horizontal direction and vertical is the vertical direction), and the units of the coordinate values ​​are adapted to the resolution of the sequencing vector. Through this operation, the specific spatial location of each cell or site within the sample is determined, forming a "cell or site - two-dimensional coordinate" correspondence table for the sample.

[0058] Furthermore, the spatial distance between each pair of cells or sites within the current sample is calculated, and the spatial distance is determined based on two-dimensional spatial coordinates.

[0059] Specifically, based on the acquired two-dimensional spatial coordinates, the spatial distance between any two cells or sites within the current sample is calculated one by one.

[0060] For any pair of target cells or sites, their horizontal and vertical coordinates are extracted, and the spatial distance is obtained by calculating the straight-line distance between the two points. This distance directly quantifies the physical proximity of the two cells or sites in the sample tissue; the smaller the distance value, the closer they are in space; the larger the distance value, the farther they are in space.

[0061] By traversing all combinations of cells or sites within the sample, the spatial distance of each pair of cells or sites is calculated, forming a "cell or site pair - spatial distance" correspondence table for the sample.

[0062] Furthermore, the type of kernel function is determined, which is used to map the calculated spatial distance to the interaction strength between cells or sites.

[0063] Specifically, based on the analysis requirements of spatial transcriptome data, select a kernel function type that can achieve a reasonable mapping between "spatial distance and interaction strength".

[0064] In this application, the exponential kernel function is preferred as the kernel function type. The exponential kernel function has a smooth mapping curve, which can transform continuous changes in spatial distance into continuous changes in interaction strength. Furthermore, the mapping sensitivity can be controlled by adjusting parameters to adapt to the spatial distribution density of cells or sites in different samples (e.g., high-density samples can have higher mapping sensitivity, while low-density samples can have lower sensitivity). Simultaneously, the mapping results of the exponential kernel function conform to biological laws. Spatially adjacent cells or sites typically have stronger functional connections (such as signal transduction and substance exchange), corresponding to higher interaction strengths. Compared to the Gaussian kernel function, the exponential kernel function exhibits a more gradual decay rate with increasing distance when characterizing spatial interactions, and can more accurately capture weak interactions between cells at slightly greater distances. This characteristic is highly consistent with cross-distance interaction modes such as signal molecule diffusion and paracrine regulation in biological tissues, making it particularly suitable for analyzing the synergistic regulatory relationships between cells in tissue samples with complex spatial hierarchies.

[0065] Furthermore, the mapping rules of the kernel function are determined. Based on the type of kernel function and the mapping rules, the spatial kernel function is obtained. The mapping rule is that the closer the cells or sites are, the higher the interaction strength, and the farther the cells or sites are, the lower the interaction strength.

[0066] Specifically, the mapping rules of the kernel function must be clearly defined. The core logic must be strictly followed: "the closer the cells or sites are, the higher the interaction strength; the farther the cells or sites are, the lower the interaction strength." Specifically, when the spatial distance between two cells or sites approaches its minimum (i.e., they almost overlap spatially), the interaction strength approaches its maximum (e.g., approaches 1); as the spatial distance gradually increases, the interaction strength gradually decreases; when the spatial distance exceeds a preset threshold (e.g., 3 times the average distance between cells or sites within the sample), the interaction strength approaches its minimum (e.g., approaches 0).

[0067] Then, the determined kernel function type (such as Gaussian kernel function) is combined with the mapping rules. By setting the key parameters of the kernel function (such as the bandwidth parameter of Gaussian kernel function, which is used to control the rate at which the intensity decays with distance), the kernel function can accurately realize the mapping of "spatial distance → interaction intensity". For example, by adjusting the bandwidth parameter, the interaction intensity decays by a preset ratio for every unit increase in spatial distance, and the decay trend conforms to the mapping rules.

[0068] This application's embodiments transform the abstract "spatial location relationship" into a specific "spatial association strength" through step-by-step operations, and the mapping results conform to the actual association rules of cells or sites within the tissue. This provides a precise core tool for subsequently calculating the constraint weights of spatial correlation effects and non-spatial association effects, ensuring that spatial factors are associated with the true spatial association characteristics of cells or sites.

[0069] In some embodiments, the initial parameter values ​​of the calculation model are used to construct the target model and determine the weights of spatial factors constrained by spatial correlation in each sample. This includes setting the initial values ​​of the parameters of the target model, which include the basic gene expression level of each sample, the gene expression weight in each sample, the cross-sample common load, the variance of the expression residuals of each sample, the total variance of the spatial factors, and the weights of spatial factors constrained by spatial correlation in each sample.

[0070] Specifically, the initial value of the basal gene expression level in a sample is based on the obtained "processed gene expression data". The average expression level of each gene in each sample across all cells or loci is calculated, and this average value is directly used as the initial value of the "basal expression level" of that gene in that sample. For example, if the average expression level of a gene in all cells of a sample is 1.5, then the initial value of the basal expression level of that gene in that sample is set to 1.5, thus reflecting the basal expression characteristics of the gene in the sample.

[0071] To simplify the initial calculations, the initial value of the "sample expression weight" for all genes in all samples is uniformly set to a unit baseline (e.g., 1.0). This unit baseline serves as the starting point for subsequent optimizations and is continuously updated during model iterations to estimate the importance of the same gene in different samples.

[0072] Based on the preprocessed gene expression data, principal component analysis was performed after merging to take the "factor loading" as the initial value of the "factor loading" common to all samples, and the variance of each principal component was calculated as the initial value of the overall variance of all cell or site spatial factors.

[0073] Based on the processed gene expression data, the variance of the expression level of each gene in all cells or loci in each sample is calculated as the initial value of the "variance of expression residuals for each sample". For example, if the variance of the expression level of a certain gene in a sample is 0.04, then the initial value of its residual variance is set to 0.04.

[0074] Furthermore, a model is constructed based on the spatial correlation strength matrix within the same sample and the initial values ​​of the parameters.

[0075] Specifically, by combining the obtained "spatial correlation strength matrix between cells or sites within the same sample" with the set initial parameter values, the conditional distribution characteristics of spatial factors (i.e., the possible range of spatial factors and their corresponding probability distribution under the current parameter constraints) are derived through statistical analysis.

[0076] In the spatial association strength matrix, the higher the association strength of a cell or site pair (i.e., the closer they are spatially), the higher the correlation of their spatial factors should be (e.g., a greater overlap in their value ranges). For example, if the association strength between a cell and its neighboring cells is 0.8 (high association), then the possible value ranges of their spatial factors should be controlled within a similar range (e.g., both 0.7-0.9); if the weight value between a cell and a distant cell is 0.1 (low association), then the value ranges of their spatial factors can be more different (e.g., 0.7-0.9 and 0.4-0.6 respectively).

[0077] Based on spatial correlation constraints, and combined with initial parameter values ​​(such as the sample's baseline expression level and sample expression weight), a probability model for the spatial factor's values ​​is constructed. This model calculates the possible numerical range (e.g., 0.5-1.0) of the spatial factor for each cell or site, and the probability of each value within that range (e.g., a 30% probability of a value of 0.8, and 20% each of values ​​of 0.7 or 0.9). Ultimately, this forms the "value range-probability" correspondence for each spatial factor, representing its conditional distribution characteristics. This characteristic directly reflects the reasonable value pattern of the spatial factor under the current parameters.

[0078] Furthermore, based on the conditional distribution characteristics of spatial factors, the parameter values ​​of the target model are updated to obtain the updated model.

[0079] Specifically, based on the conditional distribution characteristics of spatial factors, the four parameters of the target model are optimized and updated respectively to ensure that the update results of each parameter are more in line with the actual data patterns.

[0080] First, update the common factor loadings for all samples. Based on the "correspondence between spatial factors and genes" in the conditional distribution characteristics of spatial factors (e.g., a certain spatial factor is only highly associated with a specific gene), optimize the correlation coefficients of the common factor loadings. For example, if the initial assumption is that the correlation coefficient between a gene and spatial factor A is 1.0, but the conditional distribution shows that its correlation coefficient with spatial factor A is actually 0.9 and its correlation coefficient with spatial factor B is 0.1, then update the corresponding coefficients in the common loadings to 0.9 and 0.1, so that the common loadings can truly convey the correlation information between genes and spatial factors.

[0081] The weights of gene expression in each sample are updated based on the updated spatial factor conditional distribution and shared factor loadings. If the association strength is higher than the initial hypothesis (e.g., the initial value is 1.0, but the actual association is stronger), the sample expression weights are increased (e.g., to 1.1); if the association strength is lower than the initial hypothesis, they are decreased (e.g., to 0.9), so that the weights can accurately reflect the specific impact of the sample on gene expression.

[0082] Based on the spatial factor conditional distribution, shared factor loadings, and gene expression weights, the deviation between the "actual expression level" and the "target model predicted expression level (calculated from the combination of baseline expression level + weight - factor - loading)" for each gene is calculated. The variance of this deviation is then calculated, and this variance replaces the original initial value as the updated "variance of expression residuals for each sample." For example, if the original initial value is 0.004 and the variance of the actual deviation is 0.003, it is updated to 0.003 to ensure that the residual variance matches the current prediction accuracy of the model.

[0083] Calculate the statistical mean of the difference between the combined expression value and the actual expression value of a gene in all cells or sites for each sample under the current decomposition, and replace the original initial value with this mean as the updated "basal expression level of the sample". For example, if the original initial value is 1.5 and the statistical mean under the conditional distribution is 1.52, then update it to 1.52 to make the basic expression level more closely match the actual expression pattern under spatial constraints.

[0084] By replacing the original parameter values ​​in the target model with the four updated parameter values, the updated model can be obtained.

[0085] Furthermore, if the lower bound of the evidence in the updated model satisfies the preset convergence condition, the updated model is determined as the target model, and the spatial factor corresponding to the target model is obtained.

[0086] Specifically, the lower bound of the log-likelihood of the model's current estimated parameters is used as the core indicator to determine whether the model is optimized properly. The lower bound of evidence is a key indicator for measuring the degree to which the model fits the data; the higher the value, the better the model fits the data. The preset convergence condition is that after several consecutive iterations (e.g., 5 times), the difference between the lower bound of evidence of the updated model and the previous model is less than a preset threshold (e.g., 0.0005), meaning that the model fit has stabilized and further iterations will have minimal effect on improving the fit.

[0087] Repeat the process of obtaining the spatial factor conditional distribution characteristics → updating parameters to obtain a new model. After each iteration, calculate the lower bound of the log-likelihood of the new model and compare it with the lower bound of the log-likelihood of the previous model. If the convergence condition is not met, continue to the next round of iteration; if the convergence condition is met, stop the iteration.

[0088] The updated model that meets the convergence condition is determined as the target model. Simultaneously, the spatial factor values ​​for each cell or site under this target model are extracted (at this point, the conditional distribution characteristics of the spatial factors have stabilized, and the probability of each value is concentrated at a certain value), serving as the spatial factor corresponding to the target model. This spatial factor accurately reflects the spatial location characteristics of cells or sites and their association with surrounding cells, and can be directly used for subsequent spatial clustering, differential gene identification, and other analyses.

[0089] This application's embodiments derive the conditional distribution of spatial factors by combining the spatial interaction weight matrix with initial parameter values. This ensures that the distribution characteristics conform to both spatial correlation constraints and gene expression patterns, providing an objective and reliable basis for parameter updates and avoiding the problem of spatial factor values ​​deviating from actual data. Through an iterative logic of obtaining distribution characteristics → updating parameters → determining convergence, coupled with a criterion that the lower bound of the log-likelihood meets preset convergence conditions, the model optimization process becomes quantifiable and repeatable. This ensures that the final target model has high fitting accuracy and strong stability, while the obtained spatial factors accurately reflect the spatial characteristics of cells / sites.

[0090] This application's embodiments, by focusing on four parameters—basic expression, expression weight, common loading, and residual variance—all based on the conditional distribution characteristics of spatial factors, ensure that the updates of each parameter are supported by objective data, avoiding subjective arbitrariness in parameter optimization. By refining the update logic of each parameter, the updated parameters accurately match the actual characteristics of spatial transcriptome data, significantly improving the model's fit to the data and reducing model prediction bias.

[0091] In some embodiments, cross-sample spatial clustering, differential gene identification, and cell development trajectory inference are performed based on the target model and spatial factors to obtain analysis results. The analysis results include spatial clustering results, differential gene identification results, and cell development trajectory inference results, including: grouping cells or sites in each sample using a clustering algorithm and determining spatial clustering patterns and common expression features based on the spatial factors of each sample cell or site to obtain spatial clustering results.

[0092] Specifically, using the spatial factors of each sample cell or site as the core input, a clustering algorithm is used to group the samples and mine spatial clustering features. The specific operation is as follows: Spatial factors directly reflect the spatial correlation characteristics of cells or sites within each sample and the potential non-spatial correlation characteristics between samples (the target model constrains its spatial correlation through the spatial interaction weight matrix). An algorithm based on the Gaussian finite mixture model (Mclust) is used to cluster the spatial factors of all samples simultaneously, grouping cells or sites to obtain spatial clustering results. Through this operation, the cells or sites of each sample are divided into several cluster groups, realizing spatial grouping and characterizing the spatial classification characteristics of the sample tissue.

[0093] For each cluster, the two-dimensional coordinate distribution of its cells or sites is analyzed to summarize the spatial clustering pattern (e.g., a group of cells is concentrated in the core area of ​​the tissue, or a group is distributed in a band along the tissue boundary). At the same time, the common gene expression characteristics of the cells or sites in the group are analyzed based on the target model (e.g., a cluster group has high expression of genes related to cell proliferation). The spatial clustering pattern and common expression characteristics are used together as the spatial clustering result.

[0094] Furthermore, based on the analysis of the association between spatial factors and genes using the target model, differential gene identification results are obtained.

[0095] The comparison dimensions can be selected according to the research objectives and can be divided into two categories: one is the comparison between samples of the same cluster in different time and space groups, and the other is the comparison between different clusters within all samples (such as the comparison between the core area cluster and the edge area cluster in tumor samples), to ensure that the comparison dimensions fit the core needs of the analysis.

[0096] For the set comparison dimensions, the gene expression levels of the corresponding cells or sites are extracted (the original gene expression levels or expression data corrected based on the target model after excluding residual interference). The Wilcoxon test, a common method for detecting significant differences, is used to screen genes whose expression level differences exceed a preset fold (e.g., 2-fold) and whose significance P-value is less than a preset threshold (e.g., 0.05). At the same time, combined with spatial factor association information, the spatial enrichment characteristics of differentially expressed genes are identified (e.g., a differentially expressed gene is highly expressed only in the marginal region cluster). Finally, a differential gene identification result containing a list of differentially expressed genes, expression difference magnitude, and spatial enrichment regions is formed.

[0097] Furthermore, based on spatial factors, developmental trajectories are constructed and key stages and characteristic genes are identified to obtain inference results of cell developmental trajectories.

[0098] Specifically, the distribution characteristics of spatial factors can reflect the dynamic changes in the spatial characteristics of cells or sites (such as the spatial factor gradually increasing from 0.5 to 0.9 during development); two-dimensional coordinates reflect the spatial location associations of cells or sites (such as cells being concentrated in a certain area of ​​the tissue in the early stages of development and migrating to another area in the later stages). The combination of the two provides a dual basis of features and location for trajectory construction.

[0099] A conventional cell trajectory inference algorithm (such as a trajectory inference method based on pseudo-temporal ranking) is used, with spatial factors as the core indicator, to perform pseudo-temporal ranking of cells or sites in all samples. The ranking logic is as follows: cells or sites with similar spatial factor values ​​and distribution patterns have similar pseudo-temporal values; cells or sites with spatial factor values ​​that gradually change with the developmental process have their pseudo-temporal values ​​arranged sequentially according to the trend of change, ultimately forming a continuous cell development trajectory (such as a linear trajectory of undifferentiated cells → transitional cells → mature cells, or a differentiation trajectory containing branches).

[0100] The constructed developmental trajectory is analyzed, and nodes where spatial factor values ​​change significantly (e.g., a sudden increase from 0.6 to 0.8) are identified. The cell population corresponding to these nodes is defined as a key transition stage (e.g., the transition from undifferentiated to a transitional state). At the same time, genes whose expression levels change significantly (e.g., up- or down-regulated) during key transition stages and are strongly associated with spatial factors are screened as characteristic genes of this stage (e.g., cell differentiation-related genes highly expressed during transition stages). Finally, a cell development trajectory inference result containing the developmental trajectory structure, key transition stages, and characteristic genes is formed.

[0101] Furthermore, for low-resolution spatial transcriptome data and its integration results, higher-resolution spatial views and gene expression data can be generated, achieving higher-precision 3D map construction. This process is based on the constraint that spatial factors in the same tissue space satisfy both spatial correlation and non-spatial correlation effects, and higher-resolution spatial factors can be generated based on the characteristics of conditional distribution.

[0102] While this application provides the method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-inventive labor. The order of steps listed in this embodiment is merely one possible execution order among many and does not represent the only execution order. In actual device or client product execution, the methods shown in this embodiment or the accompanying drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0103] like Figure 3As shown in the illustration, this application also provides a 3D tissue reconstruction device 300 based on spatial transcriptome data. The device includes: The acquisition module 301 is used to acquire spatial transcriptome data, which includes gene expression data and spatial location data; wherein, each sample includes multiple cells or sites, the gene expression data is used to record the expression level of different genes in each sample at the corresponding cells or sites, and the spatial location data is used to record the two-dimensional spatial coordinates of the corresponding cells or sites in each sample. Processing module 302 is used to standardize and screen the gene expression data to obtain processed gene expression data; and to construct a spatial kernel function based on the spatial location data, and to obtain a spatial correlation strength matrix between cells or sites within the same sample based on the spatial kernel function. The construction module 303 is used to construct a target model based on the processed gene expression data and the spatial correlation strength matrix. The target model is used to decompose the gene expression data into the basic gene expression level of each sample, the expression weight of the gene in each sample, the spatial factor of each sample, the cross-sample common load, and the expression residual of each sample. The acquisition module 301 is also used to acquire initial parameter values ​​and determine the weight of the spatial correlation constraint spatial factor in each sample based on the initial parameter values. The processing module 302 is also used to input the initial parameter values ​​and the weights of the spatial correlation constraint spatial factors in each sample into the target model to obtain the integrated result; The processing module 302 is further configured to perform high-resolution reconstruction processing on the spatial transcriptome data based on the integration result, so as to obtain processed spatial transcriptome data.

[0104] In some embodiments, the processing module 302 is further configured to perform logarithmic normalization and scaling on the expression level of each cell or site in the gene expression data to obtain processed gene expression data. The processing module 302 is further configured to acquire gene expression data in the processed gene expression data in which the spatial difference significance exceeds a preset spatial difference significance and the expression level exceeds a preset expression level; The processing module 302 is further configured to construct a spatial kernel function based on the spatial location data, wherein the spatial kernel function is used to characterize the spatial interaction relationship between cells or sites within the same sample; The processing module 302 is also used to calculate the strong spatial correlation matrix between each pair of cells or sites within the same sample based on the spatial kernel function.

[0105] In some embodiments, the processing module 302 is further configured to obtain the two-dimensional spatial coordinates of all cells or sites within the same sample; The processing module 302 is also used to calculate the spatial distance between each pair of cells or sites in the current sample, the spatial distance being determined based on two-dimensional spatial coordinates; The processing module 302 is further configured to determine the type of the kernel function, which is used to map the calculated spatial distance to the interaction strength between cells or sites; The processing module 302 is further configured to determine the mapping rule of the kernel function, and obtain a spatial kernel function according to the type of the kernel function and the mapping rule of the kernel function. The mapping rule is that the closer the cells or sites are, the higher the interaction strength, and the farther the cells or sites are, the lower the interaction strength.

[0106] In some embodiments, the processing module 302 is further configured to obtain initial parameter values, which include the basic gene expression level of each sample, the gene expression weight in each sample, the cross-sample common load, the variance of gene expression residuals in each sample, the overall variance of spatial factors, and the weight of spatial factors constrained by spatial correlation in each sample. The processing module 302 is further configured to construct a target model based on the initial parameter values, and determine the weights of the spatial factors constrained by spatial correlation based on the target model.

[0107] In some embodiments, the processing module 302 is further configured to calculate the expectation and variance of the conditional distribution as spatial factor estimates and estimated variances based on the EM algorithm. The processing module 302 is also used to obtain the estimated value by taking the partial derivative of the parameters based on the spatial factor estimate and the log-likelihood function of the expected complete data; The processing module 302 is also used to iteratively calculate the lower bound of the log-likelihood of the current estimated parameters, and stop the iteration when the convergence condition is met, and obtain the integrated result.

[0108] In some embodiments, the processing module 302 is further configured to use a clustering algorithm to group the cells or sites of multiple samples simultaneously based on the spatial factors of each sample cell or site to obtain spatial clustering results. The processing module 302 is also used to perform differential gene analysis based on the spatial clustering results to obtain specific gene expression characteristics of different spatial domains; The processing module 302 is also used to construct a spatial developmental trajectory based on the spatial factors of each sample cell or site and identify key transition nodes, and obtain processed spatial transcriptome data based on the key transition nodes and specific gene expression characteristics.

[0109] Some modules in the apparatus described in this application can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0110] The apparatus or module described in the above embodiments can be implemented by a computer chip or physical entity, or by a product with a certain function. For ease of description, the above apparatus is described by dividing it into various modules according to their functions. When implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.

[0111] The methods, apparatus, or modules described in this application can be implemented in a computer-readable program code manner. The controller can be implemented in any suitable manner, such as a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of a memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code manner, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included within it for implementing various functions can also be considered as structures within the hardware component. Alternatively, the device used to implement various functions can be viewed as either a software module that implements the method or a structure within a hardware component.

[0112] This application also provides an apparatus, the apparatus comprising: a processor; a memory for storing processor-executable instructions; wherein, when the processor executes the executable instructions, it implements the method described in this application.

[0113] This application also provides a non-volatile computer-readable storage medium storing a computer program or instructions thereon, which, when executed, enables the method described in this application embodiment to be implemented.

[0114] Furthermore, in the various embodiments of the present invention, each functional module can be integrated into a processing module, or each module can exist independently, or two or more modules can be integrated into a single module.

[0115] The aforementioned storage media include, but are not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), Cache, Hard Disk Drive (HDD), or Memory Card. The memory can be used to store computer program instructions.

[0116] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, or it can be embodied in the process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0117] The various embodiments described in this specification are presented in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. All or part of this application can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multiprocessor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.

[0118] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of this application.

Claims

1. A 3D tissue reconstruction method based on spatial transcriptome data, characterized in that, include: Acquire spatial transcriptome data, which includes gene expression data and spatial location data; wherein, each sample includes multiple cells or sites, the gene expression data is used to record the expression level of different genes in each sample at the corresponding cells or sites, and the spatial location data is used to record the two-dimensional spatial coordinates of the corresponding cells or sites in each sample; The gene expression data is standardized and gene screening is performed to obtain processed gene expression data; and a spatial kernel function is constructed based on the spatial location data, and a spatial correlation strength matrix between cells or sites within the same sample is obtained based on the spatial kernel function. Based on the processed gene expression data and the spatial correlation strength matrix, a target model is constructed. The target model is used to decompose the gene expression data into the basic gene expression level of each sample, the expression weight of the gene in each sample, the spatial factor of each sample, the cross-sample common load, and the expression residual of each sample. Obtain initial parameter values, and determine the weights of spatial correlation constraint spatial factors in each sample based on the initial parameter values; The initial parameter values ​​and the weights of the spatial correlation constraint factors in each sample are input into the target model to obtain the integrated result; Based on the integration results, the spatial transcriptome data is reconstructed at high resolution to obtain processed spatial transcriptome data, and the reconstructed 3D tissue is obtained based on the processed spatial transcriptome data.

2. The method according to claim 1, characterized in that, The gene expression data is standardized and gene screening is performed to obtain processed gene expression data. And based on the spatial location data, a spatial kernel function is constructed, and a spatial correlation strength matrix between cells or sites within the same sample is obtained according to the spatial kernel function, including: Log-normalization and scaling are performed on the expression level of each cell or site in the gene expression data to obtain the processed gene expression data; Obtain gene expression data from the processed gene expression data in which the spatial difference significance exceeds the preset spatial difference significance and the expression level exceeds the preset expression level; A spatial kernel function is constructed based on the spatial location data. The spatial kernel function is used to characterize the spatial interaction relationship between cells or sites within the same sample. The spatial correlation matrix between each pair of cells or sites within the same sample is calculated based on the spatial kernel function.

3. The method according to claim 2, characterized in that, The construction of the spatial kernel function based on the spatial location data includes: Obtain the two-dimensional spatial coordinates of all cells or sites within the same sample; Calculate the spatial distance between each pair of cells or sites within the current sample, the spatial distance being determined based on two-dimensional spatial coordinates; Determine the type of kernel function, which is used to map the calculated spatial distance to the interaction strength between cells or sites; The mapping rule of the kernel function is determined, and a spatial kernel function is obtained according to the type of the kernel function and the mapping rule of the kernel function. The mapping rule is that the closer the cells or sites are, the higher the interaction strength, and the farther the cells or sites are, the lower the interaction strength.

4. The method according to claim 1, characterized in that, The process of obtaining initial parameter values ​​and determining the weights of spatial correlation constraint spatial factors in each sample based on the initial parameter values ​​includes: Obtain the initial parameter values ​​of the target model. The initial parameter values ​​include the basic gene expression level of each sample, the gene expression weight in each sample, the cross-sample common load, the variance of the expression residuals of each sample, the overall variance of the spatial factors, and the weight of the spatial factors constrained by the spatial correlation in each sample. A target model is constructed based on the initial parameter values, and the weights of the spatial factors constraining spatial correlation are determined based on the target model. Calculate the initial parameters of the model, construct the target model, and determine the weights of spatial factors that constrain spatial correlation.

5. The method according to claim 1, characterized in that, The process of inputting the initial parameter values ​​and the weights of the spatial correlation constraint factors in each sample into the target model to obtain the integrated result includes: Based on the EM algorithm, the expectation and variance of the conditional distribution are calculated as the spatial factor estimate and estimated variance; The estimated value is obtained by taking the partial derivative of the parameters with respect to the spatial factor estimate and the log-likelihood function of the expected complete data. Iteratively calculate the lower bound of the log-likelihood of the current estimated parameters. If the convergence condition is met, stop the iteration and obtain the integrated result.

6. The method according to claim 1, characterized in that, The spatial transcriptome data is then reconstructed at high resolution based on the integration results to obtain processed spatial transcriptome data, including: Based on the spatial factors of cells or sites in each sample, a clustering algorithm is used to group cells or sites in multiple samples simultaneously to obtain spatial clustering results. Differential gene analysis was performed based on the spatial clustering results to obtain the specific gene expression characteristics of different spatial domains; Spatial developmental trajectories are constructed based on spatial factors of each sample cell or site, and key transition nodes are identified. Processed spatial transcriptome data are obtained based on key transition nodes and specific gene expression characteristics.

7. A 3D tissue reconstruction device based on spatial transcriptome data, characterized in that, include: The acquisition module is used to acquire spatial transcriptome data, which includes gene expression data and spatial location data. Each sample includes multiple cells or sites. The gene expression data is used to record the expression levels of different genes in each sample at the corresponding cells or sites, and the spatial location data is used to record the two-dimensional spatial coordinates of the corresponding cells or sites in each sample. The processing module is used to standardize and screen the gene expression data to obtain processed gene expression data; and to construct a spatial kernel function based on the spatial location data, and to obtain a spatial correlation strength matrix between cells or sites within the same sample based on the spatial kernel function. The construction module is used to construct a target model based on the processed gene expression data and the spatial correlation strength matrix. The target model is used to decompose the gene expression data into the basic gene expression level of each sample, the expression weight of the gene in each sample, the spatial factor of each sample, the cross-sample common load, and the expression residual of each sample. The acquisition module is also used to acquire initial parameter values ​​and determine the weight of the spatial correlation constraint spatial factor in each sample based on the initial parameter values. The processing module is also used to input the initial parameter values ​​and the weights of the spatial correlation constraint spatial factors in each sample into the target model to obtain the integrated result; The processing module is further configured to perform high-resolution reconstruction processing on the spatial transcriptome data based on the integration result, to obtain processed spatial transcriptome data.

8. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.