A method of computing RNA velocity analysis in image-based spatial transcriptome data
By analyzing image-based spatial transcriptome data, we calculated the normalized distance of transcripts using two-dimensional and three-dimensional models, corrected the cell nucleus diameter, established a model of intranuclear to perinuclear and cytoplasmic migration, and constructed an output matrix at the cell-gene level. This solved the problem that existing methods failed to fully utilize the subcellular spatial distribution information of transcripts, and achieved stable and interpretable cell state description and temporal sequencing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RENJI HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
- Filing Date
- 2026-05-13
- Publication Date
- 2026-07-31
AI Technical Summary
Existing methods fail to fully utilize the subcellular spatial distribution information of transcripts, making it difficult to construct stable and generalizable cell state description indicators. Furthermore, traditional pseudo-temporal methods ignore spatial structural information, leading to discrepancies between the inferred results and real biological processes, and lacking adaptability to different biological systems.
By analyzing image-based spatial transcriptome data, we calculated the normalized distance of transcripts using two-dimensional and three-dimensional models, corrected the cell nucleus diameter, established a model of intranuclear to perinuclear and cytoplasmic migration, constructed a cell-gene level output matrix, and used kinetic feature space for pseudo-temporal sorting to achieve gene marker-independent cell state description.
It achieves stable induction of transcript dynamic signals at the single-cell level, improves the stability and interpretability of cell state description, can adapt to the diverse needs of different biological systems, and performs cell temporal sequencing under spatial constraints.
Smart Images

Figure CN122493932A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of bioinformatics and spatial transcriptomics, specifically relating to a method for calculating RNA rate analysis in image-based spatial transcriptomics data. Background Technology
[0002] In terms of data analysis methods, RNA velocity analysis methods do not contain precise spatial location information, and their dynamic models rely on specific transcript processing assumptions, making them difficult to apply directly to spatial transcriptome data, especially in subcellular resolution scenarios. If the RNA velocity is extended by combining spatial neighborhood relationships, it can only be analyzed at a coarse-grained level at the cellular level. It lacks direct modeling of the spatial distribution differences of transcripts within the cell (such as the distribution in the nucleus and cytoplasm), and usually depends on specific genes or markers, making it difficult to construct a universal dynamic characterization system that is independent of gene selection.
[0003] In the inference of cell development trajectories, there are existing pseudo-temporal analysis methods that construct trajectory or graph structures after dimensionality reduction to sort cell states. However, these methods are mainly based on expression profile similarity, ignoring spatial structural information and the dynamic behavior of transcripts. Moreover, their results often depend on specific gene sets or dimensionality reduction methods, lacking stability and interpretability. Summary of the Invention
[0004] The technical problem solved by the present invention is:
[0005] 1. Existing methods are mostly based on expression levels or splicing information, and fail to directly and fully utilize the subcellular spatial distribution information of transcripts to characterize the transcript output process. Therefore, it is necessary to solve how to convert the spatial location of nuclear transcripts into output probability or output intensity parameters that can be used for dynamic modeling.
[0006] 2. Existing methods often rely on specific genes or model assumptions, making it difficult to construct stable and generalizable indicators for describing cell state.
[0007] 3. Traditional pseudo-temporal methods ignore spatial constraints and transcript dynamics, leading to discrepancies between the inferred results and real biological processes. Therefore, it is necessary to solve how to summarize the transcript dynamics signals at the cellular level into a small number of stable features and use these features to complete the cellular temporal sequencing under spatial constraints.
[0008] 4. Existing methods often lack adjustable prior settings (such as characteristic patterns at different stages), making it difficult to adapt to the diverse needs of different biological systems.
[0009] To address the aforementioned technical problems, the present invention provides a method for calculating RNA rate analysis in image-based spatial transcriptome data, comprising: Transcript-normalized distance is calculated using a two-dimensional model based on transcript information; Based on transcript information and transcript normalized distance, a three-dimensional model is used for correction, and each cell nucleus is approximated as a three-dimensional sphere. The true diameter of the cell nucleus corresponding to that cell nucleus is estimated by inversion, and then the corrected transcript normalized distance is calculated. Based on the transcript normalized distance or the corrected transcript normalized distance as the final transcript normalized distance, the corresponding parameters of the model describing the migration from the nucleus to the perinuclear region and the cytoplasm are calculated to obtain the output probability of any transcript in each cell nucleus. The transcript output signals of the same gene in the same cell are aggregated to obtain the cell-gene level output matrix. The dynamic characteristics of each cell are calculated based on the output matrix and mapped to a three-dimensional dynamic characteristic space to achieve a cell state description independent of gene markers; Based on the mapped three-dimensional dynamic feature space, selected and unselected cells are determined according to preset or customizable feature patterns. A pseudo-temporal model is constructed, and the final temporal ranking result is output as the pseudo-temporal result of the cells.
[0010] Preferably, the transcript information includes the specific spatial location of the transcript, cellular components, the current specific spatial location of the cell, and the gene name.
[0011] Preferably, the transcript normalized distance includes the distance from the transcript to the nuclear boundary, the transcript normalized distance, the relative perinuclear distance, or the position of the hierarchical loop.
[0012] Preferably, the transcript normalized distance The formula is as follows:
[0013] Where r represents half the maximum diameter of the cell nucleus in a certain cell, and the center of the maximum diameter is the center of the circle, while r' represents the planar distance of a certain transcript in the cell from the center of the circle.
[0014] The corrected transcript normalized distance The formula is as follows:
[0015] in, For apparent radius, This represents the true diameter of the cell nucleus.
[0016] Preferably, the calculation process for the true diameter of the cell nucleus is as follows: A mapping relationship between the radius distribution and the potential true radius distribution is constructed based on the Wicksell integral equation; The radius distribution is obtained by estimating the nuclear density based on the apparent radius. The potential true radius distribution is recovered after numerical inversion. The true nuclear radius of each cell is estimated based on the mapping relationship, and the true diameter of the cell nucleus is assigned.
[0017] Preferably, the kinetic characteristics include: overall output intensity, output gene concentration, and kinetic consistency with neighboring cells.
[0018] The present invention provides a method for calculating RNA rate analysis in image-based spatial transcriptome data. By combining the spatial distribution information of intracellular transcripts, a dynamic model of transcript output is established on a two-dimensional or three-dimensional spatial scale. Combined with multidimensional feature quantification indicators (including but not limited to transcript output intensity, specificity, and spatial consistency), a time-series inference method for cell state evolution independent of genes is constructed. This method can be applied to single-cell and spatial multi-omics data analysis, biological development process analysis, and disease process research. Attached Figure Description
[0019] Figure 1 A flowchart illustrating a method for calculating RNA rate analysis in image-based spatial transcriptome data, provided by an embodiment of the present invention; Figure 2 This is a schematic diagram of a two-dimensional model provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the three-dimensional model calculation process provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the cell-gene level output matrix calculation process provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the three-dimensional dynamic feature space construction process provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the cell pseudo-time result generation process provided in an embodiment of the present invention; Figure 7 This is a schematic diagram illustrating examples of using a two-dimensional model provided in the embodiments of the present invention; Figure 8 This is a schematic diagram illustrating examples of using a three-dimensional model provided in the embodiments of the present invention; Figure 9 This is a schematic diagram showing the calculation results of a method for calculating RNA rate analysis in image-based spatial transcriptome data provided by an embodiment of the present invention. Detailed Implementation
[0020] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.
[0021] like Figure 1 As shown in the figure, an embodiment of the present invention provides a method for calculating RNA rate analysis in image-based spatial transcriptome data, comprising the following steps: S1: Transcript information is obtained based on spatial transcriptome sequencing / imaging. Transcript information includes the specific spatial location of the transcript, cellular components (CellComp), the current specific spatial location of the cell, and gene names (used for cell construction). Gene expression matrix).
[0022] The specific spatial location of a transcript is not necessarily limited to a two-dimensional coordinate system based on the nuclear transcript; it can also be based on the nuclear boundary, nuclear membrane boundary, nuclear center point, or the geometric center of the nuclear transcript.
[0023] Cellular components can indicate whether a transcript belongs to the cell nucleus.
[0024] Transcript information also includes field of view (FOV) and cell name (used to distinguish different cells. As shown in Table 1, the cell column is the cell number, and the cell_ID column is the cell name, both of which can be used to distinguish cells), see Table 1.
[0025] Table 1 Transcription Information Table
[0026] S2: As Figure 2 As shown, the standardized distance of transcripts is calculated using the Spatial-Velo algorithm and a two-dimensional model based on transcript information.
[0027] The distance from the transcript to the nuclear boundary, the relative perinuclear distance, or the position of the hierarchical loop are used as the transcript-normalized distance.
[0028] The transcript-normalized distance is calculated by approximating each cell nucleus as a circle, calculating the maximum two-dimensional diameter 2r and the distance r' from the center of any transcript circle. The transcript-normalized distance r is then calculated. norm The formula is as follows: r norm =r' / r Where 2r is the two-dimensional maximum diameter, i.e., the maximum distance between transcripts within the same cell nucleus; the midpoint of 2r is the center of the cell nucleus. The distance r' of any transcript from the center is defined as the distance between the transcript's specific spatial location and the center of the cell nucleus.
[0029] r norm It can be directly used to infer the state of transcript exiting the cell nucleus (S4 step, i.e., S2->S4); or it can be corrected by performing a three-dimensional model (S3) to calculate r. norm,corr Then, the state of the transcript exiting the cell nucleus is inferred (S4 step, i.e., S2->S3->S4).
[0030] S3: Based on transcript information and transcript normalized distance, a three-dimensional model is used for correction, approximating each cell nucleus as a three-dimensional sphere. The true diameter R of the corresponding cell nucleus is estimated using Wicksell inversion. est Then, the normalized distance r of the corrected transcripts was calculated. norm,corr This step can reduce geometric biases caused by slice thickness, nuclear cross-section position, and two-dimensional projection, and improve the accuracy of transcript spatial location inference of output dynamics.
[0031] Corrected transcript normalized distance r norm,corr The formula is as follows:
[0032] Specific steps are as follows Figure 3 As shown: a. For each cell, based on transcript information and transcript-normalized distance, the maximum diameter 2r is measured using a two-dimensional nuclear mask. obs .
[0033] b. Approximating each cell nucleus as a three-dimensional sphere, the observed nuclear radius is interpreted as: The radius distribution of random two-dimensional cross-sections of a three-dimensional sphere is constructed based on the Wicksell integral equation. With the potential true radius distribution The mapping relationship between them is shown in the following formula: .
[0034] c. Based on apparent radius r obs Obtained by kernel density estimation The potential true radius distribution is recovered after numerical inversion. Based on the mapping relationship, the true nuclear radius of each cell is estimated, and the true nuclear diameter R is assigned. est .
[0035] The true nuclear radius of each cell is estimated based on the mapping relationship, including but not limited to expectation estimation based on conditional probability, maximum likelihood estimation, maximum distance method, mean, quantile, boundary fitting, nuclear mask fitting, estimation based on statistical distribution or sampling method.
[0036] d. Based on apparent radius r obs True diameter of cell nucleus R est The standardized distance r (in a two-dimensional model) norm Calculate the transcript normalized distance r in the three-dimensional model. norm,corr .
[0037] S4: As Figure 4 As shown, if we assume that all detected cell nuclei are of the largest diameter, then the transcript normalized distance r norm That is, the standardized distance in the two-dimensional model, which is used as the final transcript standardized distance. Alternatively, if it is assumed that the cell nucleus size follows a normal distribution, and the detected cell nucleus diameter is the result of randomly cut three-dimensional spheres, then the corrected transcript standardized distance r is used. norm,corr The normalized distance (i.e., the normalized distance in the 3D model) is used as the final transcript normalized distance. The corresponding parameters of the model describing the migration from the nucleus to the perinuclear region and the cytoplasm are calculated to obtain the output probability of any transcript in each cell nucleus. The output signals of transcripts of the same gene in the same cell are aggregated to obtain the cell-gene level output matrix (delta matrix), which is used to represent the future changes in cell expression status.
[0038] Models describing nuclear migration to the perinuclear region and cytoplasm include: diffusion-output models, partial differential equations of different forms, diffusion equations, convection-diffusion equations, stochastic process models, or Markov process models, as long as they can describe the spatial dynamics of nuclear migration to the perinuclear region and cytoplasm.
[0039] Methods for calculating the parameters of models describing intranuclear to perinuclear and cytoplasmic migration include: maximum likelihood estimation, Bayesian estimation, EM algorithm, least squares fitting, or other parameter inversion methods.
[0040] The output probability can be represented as an exponential mapping, a logistic function, a power function, a piecewise function, or another monotonic mapping function.
[0041] The methods for aggregating the output signals of transcripts of the same gene within the same cell to obtain an output matrix include: summation, weighted summation, average value, normalized integral, nuclear-internal ratio, or other equivalent aggregation methods.
[0042] The delta matrix should not be understood as the output of a fixed formula, but rather as a matrix representing future changes at the cell-gene level, inferred from a model of transcript spatial location and nuclear output dynamics.
[0043] It is important to note that we used different methods to calculate the transcript normalized distance r in S2 and S3. norm r norm,corr Both can be used for subsequent analysis. Therefore, for the sake of convenience, this section will not distinguish between r. norm r norm,corr Instead, r is used uniformly to represent the standardized distance of the final transcript, and r i This represents the normalized distance of the i-th final transcript.
[0044] Spatial-Velo uses a probability density function derived from partial differential equations (PDEs) to simulate transcript localization. This function captures the underlying biological mechanisms controlling nuclear organization to establish a diffusion-export model of migration from the nucleus to the perinuclear region and cytoplasm. Its boundary conditions represent possible interactions around the nucleus, as shown in the following formula:
[0045]
[0046] Where r represents the final transcript normalized distance. c is the concentration function, λ is the characteristic length scale associated with diffusion and degradation processes, β is the parameter of the capture boundary interaction, and c(1) is the transcript concentration distribution at the nuclear boundary.
[0047] Boundary conditions can take the form of nuclear boundary absorption, semipermeable membrane boundary, reflective boundary, or mixed boundary.
[0048] The spatial distribution pattern of nuclear transcripts is quantified using the maximum likelihood estimation (MLE) framework. For each cell with n transcripts, the normalized distance of the i-th final transcript is represented as r. i Therefore, the parameters λ and β are estimated using the following formula:
[0049] in, Let the final transcript normalized distance be the distance of the i-th transcript. It is a normalization constant, calculated using trapezoidal numerical integration:
[0050] After obtaining the parameters λ and β, we fitted the spatial distribution of transcripts for each cell based on the steady-state solution of the diffusion-output model, and the nuclear output of transcripts s(r i Expressed as a flux function, the formula is as follows:
[0051]
[0052] in, These are the zeroth and first-order modified Bessel functions of the first kind.
[0053] To establish a biologically meaningful scaling relationship between the flux function and the actual output rate, a calibration coefficient α is introduced for each cell, as shown in the following formula:
[0054] in, The total number of nuclear transcripts in the cell, ε′′=10 10 is a small constant used to prevent division by zero; this calibration ensures that the total output flux is determined according to parameter β and the total number of nuclear transcripts in the cell. Scaling should be done appropriately.
[0055] Assuming nuclear output follows a Poisson process, calculate the output probability per unit time for each transcript. for: .
[0056] For the same gene within the same cell, the output probability of different transcripts is as follows: This allows us to calculate the expected final cross-nuclear output of the gene. By integrating the future output expectations of all cells and all genes, we can calculate the cell-gene level output matrix as a delta matrix:
[0057] Where N represents the number of cells, G represents the number of genes, and v m,n This represents the expected cross-nuclear output of the nth gene in the mth cell.
[0058] We propose two models to calculate the normalized distance of transcripts from the nuclear boundary: ① The two-dimensional model idealizes the detected cell nucleus as a planar circle, ignoring the geometric deviation caused by the three-dimensional spherical geometry, and directly calculates the distance of the transcript from the cell nucleus boundary in the current cross-section. This method is simple to calculate and highly efficient.
[0059] ② The three-dimensional model is based on the calculation results of the two-dimensional model. It considers the geometric deviation caused by the geometry of the three-dimensional sphere, estimates the actual three-dimensional diameter of the cell nucleus, and finally calculates the standardized distance from the cell nucleus boundary. This method is more accurate, but the computational load is large.
[0060] S5: As Figure 5 As shown, the dynamic characteristics of each cell are calculated based on the output matrix and mapped to a three-dimensional dynamic characteristic space to achieve a cell state description independent of gene markers.
[0061] The dynamic characteristics include: overall output intensity (Hotspot), output gene concentration (Specificity), and dynamic consistency with neighboring cells (Coherence).
[0062] The overall output intensity (Hotspot) is used to characterize the overall transcript output intensity of a cell and is defined as the L2 norm of the velocity vector. The formula for calculating the hotspot of the m-th cell is as follows:
[0063] The overall output intensity Hotspot can be replaced by the L2 norm, L1 norm, weighted norm, maximum amplitude, or other indicators that characterize the overall intensity of change of the velocity vector.
[0064] Specificity of the gene concentration output by the m-th cell m The formula used to characterize whether the output signal is dominated by a few genes is as follows:
[0065]
[0066] in, This represents the positive velocity component of the nth gene in the mth cell.
[0067] Output gene concentration specificity m It can be replaced by other indicators such as the percentage of maximum values, entropy, Gini coefficient, Herfindahl-Hirschman index, kurtosis, or "the degree of dominance of minority genes".
[0068] Coherence, the dynamic consistency of a neighboring cell, is used to characterize the consistency between the dynamic direction of a cell and its spatial neighbors (i.e., the average cosine similarity of the velocity directions of the cell and its k surrounding cells), and is defined by the following formula:
[0069] in, Let m be the normalized velocity vector of the m-th cell.
[0070] Let m be a set of k nearest neighbor cells, where the velocity vector of any nearest neighbor cell p is represented by... express.
[0071] The coherence of neighborhood cell dynamics can be replaced by neighborhood cosine similarity, local correlation coefficient, graph smoothness, Laplace consistency, neighborhood orientation consistency, or other indicators that measure the coherence of spatially neighboring cell dynamics.
[0072] The essence of the above three characteristics is not a fixed mathematical formula, but rather they represent: • Overall intensity of dynamic change; • Are the dynamic changes concentrated in a few genes? • Whether the dynamic changes within the spatial neighborhood are consistent.
[0073] Therefore, any solution that achieves the same technical effect should be considered an alternative to the present invention.
[0074] S6: As Figure 6 As shown, based on the mapped three-dimensional dynamic feature space, selected and unselected cells are determined according to preset or customizable feature patterns, a pseudo-temporal model is constructed, and the final temporal sorting result is output as the cell pseudo-temporal result.
[0075] Methods for constructing pseudo-time series models include: 1. A label propagation method based on graph regularization; 2. Time propagation methods based on kNN graphs, weighted graphs, spatial adjacency graphs, or cell similarity graphs; 3. Methods based on optimal transport (OT), shortest path, diffusion mapping, manifold sorting, or spectral sorting; 4. A semi-supervised sorting method based on early / mid / late stage seed cell definitions; 5. Hierarchical ranking methods based on threshold grouping, quantile grouping, cluster center grouping, or manual prior definition.
[0076] The specific steps for constructing a pseudo-temporal model based on the graph regularization-based label propagation method are as follows: Selected cells are assigned discrete pseudo-time labels. A spatial adjacency graph (k-nearest neighbor graph G) is constructed based on the current spatial location of the cell. The adjacency weights are calculated according to the spatial distance between cells to obtain the spatial adjacency matrix W. A graph Laplacian matrix L (used to describe the spatial structural relationships between cells) is further constructed. The pseudo-time value of each cell is solved through label propagation based on graph regularization, and range constraints and normalization are applied. The final temporal ranking result is output as the cell pseudo-time result. This method achieves continuous propagation and estimation of pseudo-time using a small number of seed cells with known states, under the constraint of spatial adjacency relationships.
[0077] Preset or customizable feature patterns include dividing cells based on the quantile range of three features to determine selected and unselected cells. Selected cells include early, middle, and late states.
[0078] The division of early, middle, and late states is not limited to the 30% / 70% quantile threshold. Other percentile thresholds, clustering results, expert annotation results, or rules based on reference markers can also be used.
[0079] Discrete pseudo-time labels, for example: y=0(early),y=0.5(middle),y=1 (late) The pseudo-time value for each cell is obtained by graph regularization, and the cell pseudo-time result is output. The linear system formula for this process is as follows:
[0080] Where S is the diagonal weight matrix (which assigns greater weight to labeled cells), μ is the smoothing parameter used to control the strength of spatial consistency constraints, f is the pseudo-time value for each cell, and y is the label vector of the seed cell.
[0081] The cell segmentation, nuclear segmentation, and transcript attribution methods in this invention are not limited to any particular software or algorithm. They can be replaced with: 1. Segmentation methods based on DAPI, membrane markers, cytoplasmic markers, or combined markers; 2. Methods based on Cellpose, CellProfiler, ImageJ, deep learning segmentation networks, or other image segmentation algorithms; 3. Based on manual correction, pathological annotation assistance, or semi-automatic segmentation methods; 4. The method of attribution based on the transcript metadata files output by different spatial transcriptome platforms.
[0082] This invention is not limited to the Nanostring CosMx-SMI platform, nor is it limited to the liver, lymph node, or brain tissue samples shown herein. Its applications can be extended to: 1. Other imaging-based spatial transcriptomics platforms, such as Xenium, MERFISH, PHOTON, etc.; 2. Other tissue types, disease types, developmental stages, organoids, or in vitro culture systems; 3. Single-sample analysis or multi-sample batch analysis scenarios.
[0083] This invention can be implemented not only by software programs, but also by processors, memory, servers, workstations, cloud computing platforms, or combinations thereof. The above steps can be executed sequentially by a single processing unit, or executed in parallel by multiple modules; they can be implemented by a single integrated program, or broken down into multiple functional modules for separate execution.
[0084] The key technical points and areas to be protected in this application are as follows: (1) Transcription dynamics inference method based on subcellular spatial location By obtaining the spatial coordinates of each transcript within a single cell, especially its positional distribution within the cell nucleus, and calculating its normalized radial position relative to the nuclear center and nuclear boundary, spatial structural information is transformed into continuous variables.
[0085] Based on this, spatial diffusion and boundary output models of nuclear transcripts were established to infer the probability of each transcript being transported from the nucleus to the cytoplasm, thereby obtaining the nuclear output rate or future expression changes based on spatial information. (2) A method for modeling the probability of nuclear output based on a physical diffusion model This application constructs a steady-state diffusion model based on partial differential equations (PDEs) and introduces nuclear boundary conditions to characterize the spatial distribution of RNA within the nucleus. Furthermore, it fits the model parameters (such as λ and β) using maximum likelihood estimation and transforms the spatial distribution function into a nuclear output probability function, thereby achieving a mapping from spatial distribution to dynamic output rate. (3) Marker-independent dynamic feature construction method This application further compresses the high-dimensional delta matrix into three unsupervised dynamic features: Hotspot: Overall Transcriptional Dynamics Strength Specificity: Whether dynamic changes are concentrated in a few genes Coherence: the dynamic consistency of cells in their spatial neighborhood This feature system does not rely on any predefined set of marker genes, thereby reducing prior bias. (4) Cross-platform adapted data processing methods The method described in this application is applicable to a variety of imaging spatial transcriptome data, including but not limited to: CosMx SMI, Xenium, and MERFISH.
[0086] This application addresses the problem of existing spatial transcriptomics techniques lacking temporal dimension information and being unable to infer cellular dynamic changes from static images. It proposes a method for inferring cellular transcriptional dynamics based on the spatial location of subcellular transcripts. This approach establishes a spatial distribution model of nuclear transcripts, combines it with nuclear output diffusion mechanisms, further constructs a cellular-level future state prediction (delta matrix), and extracts label-free dynamic features for cellular temporal reconstruction, thereby recovering implicit temporal information from static spatial data.
[0087] Compared with the prior art, this application has at least the following beneficial technical effects: (1) To realize the inference of dynamic transcription process from static spatial transcriptome data Existing technologies typically rely on RNA splicing information or time-series sampling data to infer cellular dynamics. However, in imaging-based spatial transcriptome data, traditional RNA velocity methods cannot be directly applied due to the lack of splicing information.
[0088] This application establishes a nuclear diffusion-output model by incorporating spatial distribution information of transcripts within the cell nucleus. It directly infers the nuclear output tendency of transcripts from spatial cross-sectional data at a single time point, thereby obtaining a delta matrix reflecting "future expression changes." This enables the inference of cellular dynamics without the need for time-series data or splicing information. This effect is directly generated by the combination of the technical features of "subcellular spatial location modeling + nuclear output dynamics modeling."
[0089] (2) Improve the stability and generalization ability of cell state prediction Existing methods based on expression level changes or marker genes are susceptible to gene selection bias, sequencing noise, and platform differences. This application maps single-gene expression changes to nuclear output probabilities based on a physical spatial model and further aggregates them into a delta matrix. This allows the dynamic signal to originate from the overall distribution characteristics of the transcript spatial structure, rather than single-gene expression changes. This technical feature makes the model more robust to differences in gene panel differences, transcript dropout, and platform detection sensitivity, thereby significantly improving the stability and cross-platform generalization ability of cell state inference.
[0090] (3) Achieve marker-free cell kinetic characterization Existing cell trajectory analysis methods typically rely on pre-selected marker genes or sets of differentially expressed genes, making their results highly sensitive to prior knowledge. This application achieves a low-dimensional, unified characterization of cell states by compressing a high-dimensional gene-wise delta matrix into three label-free features (Hotspot, Specificity, and Coherence) based on dynamic structures.
[0091] Specifically, Hotspot represents the specificity of overall transcriptional change intensity, indicating whether the change is dominated by a few genes, while Coherence represents the consistency of dynamics within a spatial neighborhood. This feature design makes cell state characterization independent of specific marker gene sets, thereby reducing human bias and improving the consistency and interpretability of the analysis.
[0092] (4) Achieving spatially consistent cellular pseudo-temporal and trajectory reconstruction Traditional pseudo-temporal ranking methods are mostly based on expression similarity or statistical ranking, which makes it difficult to simultaneously guarantee spatial continuity and biological directionality. This application utilizes Hotspot, Specificity, and Coherence to construct a three-dimensional dynamic space, and combines it with a spatial adjacency graph for regularization propagation, thereby achieving global directional ranking while satisfying local spatial smoothness.
[0093] This method can reconstruct continuous cell state change paths within spatial tissue structures, ensuring that early, intermediate, and late states are spatially continuous, cell trajectories have clear directionality, and transcriptional dynamics are consistent with tissue structure. This effect is directly generated by a combination of techniques combining "kinetic characteristics + spatial graph regularization inference."
[0094] (5) Improve robustness to nuclear geometry differences and plateau errors. This application further introduces: a 2D to 3D nuclear structure correction model, non-circular nuclear morphology perturbation analysis, and sensitivity testing of transcript length and gene composition. This ensures that the model maintains stable output under different nuclear morphologies, tissue structures, and spatial transcriptome platform conditions. Therefore, the method proposed in this application has low sensitivity to geometric projection errors, nuclear morphology changes, and gene set variations, thereby improving the reliability for practical applications.
[0095] (6) Achieve a unified cross-platform framework for spatial transcription dynamics analysis The method presented in this application is not only applicable to Nanostring CosMx-SMI data, but can also be extended to various imaging-based spatial transcriptomics platforms such as Xenium and MERFISH. Through a unified analytical framework of "spatial location-nuclear output model-delta matrix-kinetic characteristics-quasi-temporal series," comparability and consistency analysis between data from different platforms are achieved, thereby improving the method's versatility.
[0096] In summary, this application achieves a methodological innovation by introducing subcellular transcript spatial distribution modeling and nuclear output dynamics inference mechanisms to recover the dynamic evolution of cells from static spatial transcriptome data. Furthermore, through label-free dynamic feature construction and spatially regularized temporal inference, it achieves a stable, continuous, and cross-platform consistent characterization of cell state changes. This technical solution not only overcomes the limitations of existing methods that rely on splicing information or marker genes, but also significantly improves the stability, generalization ability, and spatial consistency of cell dynamics inference, demonstrating clear technological advancement and practical application value.
[0097] Example 1. Two-dimensional model example: Based on CosMx-SMI spatial transcriptome data, the Spatial-Velo model predicted the evolution of circadian clock-related genes in liver cancer-adjacent hepatocyte data.
[0098] like Figure 7 As shown, in a liver cancer study based on CosMx-SMI spatial transcriptome data, subcellular resolution transcriptome analysis was performed on tumor and adjacent normal tissue sections from 18 hepatocellular carcinoma patients, and the Spatial-Velo algorithm was used for analysis.
[0099] like Figure 7 As shown in Figure A, the experiment covered liver cancer and adjacent normal tissues obtained at different time points, and a spatial transcriptome dataset was constructed using tissue microarrays.
[0100] like Figure 7 As shown in Figure B, Spatial-Velo first establishes a diffusion-boundary model of RNA migration from the nucleus to the outside of the nucleus based on the spatial coordinates of nuclear transcripts. It then calculates the nuclear export tendency and cellular delta changes of each transcript using a two-dimensional model, thereby inferring the future expression trend of genes.
[0101] like Figure 7 As shown in Figure C, in the circadian rhythm correlation analysis, taking CLOCK and PER1 as representative genes, it was found that the distribution of their nuclear transcripts from the nuclear boundary was significantly correlated with the time state, and showed consistent spatial gradient characteristics in different liver tissue states.
[0102] like Figure 7As shown in Figure D, further reconstruction analysis of circadian rhythm-related genes in adjacent normal tissues was performed based on the "future expression level" (i.e., the sum of current expression and nuclear output delta) predicted by Spatial-Velo. The results showed that the predicted expression pattern can more clearly distinguish different time states and is superior to traditional analysis methods based solely on cytoplasmic expression.
[0103] Example 2, 3D model example: Spatial-Velo algorithm reveals the maturation and differentiation pathway of B lymphocytes.
[0104] like Figure 8 As shown, in an analysis based on Nanostring CosMx-SMI human lymph node spatial transcriptome data, the normalized distance of transcripts was calculated using a three-dimensional model of Spatial-Velo, which was then used to model the B cell differentiation process to evaluate its ability in the dynamic analysis of multi-level intercellular dynamics and the identification of differentiation states.
[0105] like Figure 8 As shown in Figures A and B, after projecting the B cell subpopulation into the UMAP space, the cell-level transcription output rate can be inferred based on the transmembrane output dynamics of nuclear RNA. Furthermore, the directional "velocity field" corresponding to the migration of RNA from the nucleus to the outside of the nucleus is represented by arrows. The results show that different B cell states exhibit a continuous but separable distribution structure in the dynamic space.
[0106] like Figure 8 As shown in C and D, at the spatial organization level, different B cell subsets form spatial niches with structural distribution in lymph node tissue. Among them, the nuclear export hotspots inferred based on Spatial-Velo are highly consistent with the spatial location of cells and show a continuous spatial flow trend from the initial B cells to germinal center B cells and plasma cells.
[0107] like Figure 8 As shown in Figure E, in the gene-level analysis, CD20 and IGKC are typical key differentiation marker genes in early and late stages. Their expression levels and corresponding nuclear output delta signals show systematic changes in different B cell subsets. At the same time, the average distance from their transcripts to the nuclear boundary gradually changes with the differentiation process, reflecting the coupling relationship between nuclear localization and transcriptional dynamics.
[0108] like Figure 8 As shown in Figure F, by further integrating current cytoplasmic expression with predicted nuclear output delta based on Spatial-Velo to construct future expression states, compared with expression patterns based solely on cytoplasmic RNA, the spatiotemporal dynamic changes during the CD20 to IGKC conversion process can be more clearly characterized, thereby achieving high-resolution identification and continuous trajectory reconstruction of early and late differentiation states of B cells.
[0109] The method for calculating RNA rate analysis in image-based spatial transcriptome data provided in this embodiment of the invention is illustrated in the following diagram. Figure 9 As shown, the details are as follows: like Figure 9 As shown in Figure A, three core dynamic indicators are first defined, including hotspot (the total intensity of nuclear output based on the single-cell delta vector, reflecting the overall transcription output flux), coherence (reflecting the consistency of the cell and its spatial neighborhood in the dynamic direction), and specificity (reflecting the degree of concentration of transcription output in a few genes).
[0110] like Figure 9 As shown in Figure B, spatial mapping inference was performed on the entire lymph node tissue section. The results showed that the three indicators exhibited obvious spatial structural partitioning in the tissue, reflecting the local dynamic differences in different cell states.
[0111] like Figure 9 As shown in Figure C, further, by modeling the changing trends of the three indicators over time in the three-dimensional metric space, the continuous trajectory distribution of cells in the dynamic space can be obtained. The predicted relationship is highly consistent with the actual observed distribution and shows the process of evolution from low hotspot / low specificity to high specificity state.
[0112] like Figure 9 As shown in Figures D and E, the pairwise relationships of the three indices inferred from the model are highly consistent with the actual data distribution, indicating that the dynamic representation has stable structural constraints.
[0113] like Figure 9 As shown in Figure F, at the spatial level, by using coherence as the spatial height dimension and specificity as the color mapping to construct a three-dimensional tissue dynamics landscape, the continuous spatial migration paths of early, intermediate, and late cell states in the tissue can be clearly observed.
[0114] like Figure 9 As shown in G, a pseudo-time is further constructed using three dynamic indices. This pseudo-time not only exhibits continuous gradient changes in the metric space, but can also be mapped back to the real organizational space to form a structured trajectory.
[0115] like Figure 9 As shown in Figure H, monotonically upregulated or downregulated genes along pseudo-time changes can be stably identified in gene-level analysis, indicating that this method can be used to analyze dynamic gene regulation programs.
[0116] like Figure 9As shown in Figures I and J, in addition, in the UMAP embedding space, the pseudo-time based on the metric can drive the continuous migration of the cell population in the transcriptional state space and exhibit stable population composition changes in different time bins, further verifying the effectiveness and interpretability of the dynamic indicator system constructed by Spatial-Velo in reconstructing the continuous evolution of cell states.
Claims
1. A method of computing RNA velocity analysis in image-based spatial transcriptome data, characterized in that, include: Transcript-normalized distance is calculated using a two-dimensional model based on transcript information; Based on transcript information and transcript normalized distance, a three-dimensional model is used for correction, and each cell nucleus is approximated as a three-dimensional sphere. The true diameter of the cell nucleus corresponding to that cell nucleus is estimated by inversion, and then the corrected transcript normalized distance is calculated. Based on the transcript normalized distance or the corrected transcript normalized distance as the final transcript normalized distance, the corresponding parameters of the model describing the migration from the nucleus to the perinuclear region and the cytoplasm are calculated to obtain the output probability of any transcript in each cell nucleus. The transcript output signals of the same gene in the same cell are aggregated to obtain the cell-gene level output matrix. The dynamic characteristics of each cell are calculated based on the output matrix and mapped to a three-dimensional dynamic characteristic space to achieve a cell state description independent of gene markers; Based on the mapped three-dimensional dynamic feature space, selected and unselected cells are determined according to preset or customizable feature patterns. A pseudo-temporal model is constructed, and the final temporal ranking result is output as the pseudo-temporal result of the cells.
2. The method for calculating RNA rate analysis in image-based spatial transcriptome data as described in claim 1, characterized in that, The transcript information includes the transcript's specific spatial location, cellular components, current cellular spatial location, and gene name.
3. The method for calculating RNA rate analysis in image-based spatial transcriptome data as described in claim 1, characterized in that, The transcript normalized distance includes the distance from the transcript to the nuclear boundary, the relative perinuclear distance, or the location of the hierarchical loop.
4. The method for calculating RNA rate analysis in image-based spatial transcriptome data as described in claim 3, characterized in that, Transcript normalized distance The formula is as follows: Where r represents half the maximum diameter of the cell nucleus in a certain cell, and the center of the maximum diameter is the center of the circle, while r' represents the planar distance of a certain transcript in the cell from the center of the circle. The corrected transcript normalized distance The formula is as follows: in, For apparent radius, This represents the true diameter of the cell nucleus.
5. The method for calculating RNA rate analysis in image-based spatial transcriptome data as described in claim 4, characterized in that, In the three-dimensional model, the calculation process for the true diameter of the cell nucleus is as follows: A mapping relationship between the radius distribution and the potential true radius distribution is constructed based on the Wicksell integral equation; The radius distribution is obtained by estimating the nuclear density based on the apparent radius. The potential true radius distribution is recovered after numerical inversion. The true nuclear radius of each cell is estimated based on the mapping relationship, and the true diameter of the cell nucleus is assigned.
6. The method for calculating RNA rate analysis in image-based spatial transcriptome data as described in claim 1, characterized in that, The dynamic characteristics include: overall output intensity, output gene concentration, and dynamic consistency with neighboring cells.