A Spatial Gene Expression Prediction Method Based on Anchoring Correction for Whole Pathological Sections

CN122575482APending Publication Date: 2026-08-14CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0006]针对现有技术存在的上述问题,本发明要解决的技术问题是:如何能够利用全局统计稳定性来约束局部预测过程、抑制位点级测量噪声影响从而提高空间基因表达谱预测的准确性

Benefits of technology

[0046]1. 突破了传统单一局部监督的局限,通过全局统计共识实施隐式正则化,从根本上缓解了过拟合技术伪影的问题。本发明(对应权利要求1及步骤S3、S5)通过引入全切片级聚合的伪批量(Pseudo-bulk)图谱作为统计学上稳健的切片级监督信号。根据大数定律,该聚合策略相互抵消了单个测序位点因低捕获率而产生的独立随机采样噪声,将信噪比在理论上提升了倍。在联合多任务训练中,全局重构损失作为强有力的隐式正则化项,强制多实例聚合的特征空间落入稳定的生物学共识范围,从而防止模型对碎片化、带噪声的局部目标产生过拟合,保证了模型学习到的是本质的组织学与生物学模式。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575482A_ABST
    Figure CN122575482A_ABST
Patent Text Reader

Abstract

This invention discloses a spatial gene expression prediction method for whole-slice pathological tissue images based on anchor correction, belonging to the field of pathological image analysis technology. The method divides the whole-slice image into image patches, extracts initial features through a pathological basic model, and obtains topologically aware features through geometric position encoding and Nyströmformer. It utilizes variational information bottlenecks to map aggregated features to a Gaussian distribution and reparameterizes them to obtain robust global anchor points. The anchor points are then concatenated with the features of each image patch to generate site-specific scaling and translation coefficients, which are then corrected by affine transformation before predicting the gene expression profile. This invention introduces a global-to-local feedback mechanism, ensuring the consistency of the prediction results with the overall tissue environment through the anchor points. This effectively mitigates the influence of site-level technical artifacts and significantly improves prediction accuracy and robustness to random sampling noise.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of case image analysis and image processing technology, and in particular to a method for predicting spatial gene expression in whole pathological slides based on anchoring correction. Background Technology

[0002] Transcriptomics research is undergoing a paradigm shift from bulk RNA sequencing to spatial transcriptomics (ST), which preserves the precise spatial location of molecular maps within tissue architecture. This leap from "population averaging" to "high-resolution mapping" offers the possibility of in-depth analysis of cellular heterogeneity and intercellular interactions within the tumor microenvironment. Despite its enormous transformative potential, current physical sequencing-based ST technologies are limited by high costs and low throughput, hindering their widespread application in routine clinicopathology.

[0003] Spatial transcriptomics, which preserves the spatial structure of tissues, can obtain molecular maps of gene expression, providing an important tool for studying cellular heterogeneity and cell-cell interactions in the tumor microenvironment. However, physical sequencing-based spatial transcriptomics is costly and has low throughput, making it difficult to widely apply in routine clinical pathology. Therefore, in recent years, various computational methods have emerged that use whole-section pathology images to infer spatial gene expression, aiming to establish the association between morphological phenotypes and molecular characteristics using readily available and inexpensive hematoxylin and eosin (H&E) stained whole-section images.

[0004] Existing computational methods are mainly categorized into regression-based, generative modeling, and retrieval-based paradigms. Regression-based methods (such as HisToGene and TRIPLEX) model global dependencies or aggregate multi-scale tissue features through attention mechanisms like Transformers, mapping image patch features to gene expression values. Generative methods (such as STFlow) integrate spatial attention mechanisms within a conditional flow matching framework, attempting to model the high-dimensional, complex distribution of gene expression. Retrieval-based methods (such as BLEEP) achieve non-parametric inference based on reference maps through aligned bimodal embeddings. Furthermore, pathological models (such as CONCH, GigaPath, UNI, Virchow2, and UNI2) can extract domain-aligned embedding representations from massive histological data, and combined with linear or ridge regression, can predict gene expression profiles.

[0005] However, the aforementioned methods all use randomly distributed site-level measurements obtained from spatial transcriptome sequencing as direct supervisory targets. Due to low RNA capture efficiency, spatial transcriptome data generally exhibits sparsity and significant sampling variance, with individual site measurements containing high levels of random noise. Existing methods blindly pursue high consistency with these noisy local targets, leading to overfitting of deep learning models to technical artifacts and making it difficult to learn robust biological patterns. On the other hand, pseudo-batch maps obtained by aggregating expression profiles of all sites within a whole slice can offset random sampling noise based on the law of large numbers, providing a statistically robust biological benchmark. This aggregation strategy has been validated in differential expression analysis and deconvolution frameworks, but it has not yet been effectively utilized in spatial gene expression prediction research. Summary of the Invention

[0006] To address the aforementioned problems in existing technologies, the technical problem this invention aims to solve is: how to utilize global statistical stability to constrain the local prediction process and suppress the influence of site-level measurement noise, thereby improving the accuracy of spatial gene expression profile prediction.

[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: A method for predicting spatial gene expression in whole pathological slides based on anchoring correction, comprising the following steps:

[0008] S1: Data preparation and definition: Obtain the training dataset, which contains multiple H&E-stained whole-section pathological images and the ground truth values ​​of the spatial transcriptome expression profiles corresponding to each whole-section image.

[0009] S2: Construct a prediction model, which includes a pathological basic model, a geometric location encoding module, a global context modeling based on Nyströmformer, a variational global anchoring module, a local correction module for anchor points, and a prediction head.

[0010] S2-1: Randomly select a whole-slice image from the training set and divide it into sections based on the irregular spatial sequencing sites it contains. Image patches constitute an image patch set The corresponding physical coordinate set is The true value of its spatial transcriptome expression profile is expressed as: ,in This represents the i-th image patch in the selected full-slice image. This represents the physical coordinates of the i-th image patch in the selected full-slice image. This represents the i-th image patch in the selected full-slice image. Dimensional gene expression vector.

[0011] S2-2: Topology-aware feature encoding: Encoding the aforementioned... Input a frozen pathological model and extract a neighborhood-aligned initial set of image patch embeddings. Using the Geometric Position Encoding (GPE) module, according to Calculate the relative offset of the irregular neighborhood of each image patch, map it to spatial filtering coefficients, and modulate the neighborhood features to synthesize a location-coded representation that includes the local micro-environment topology. ; then The input is a Nyströmformer layer for global context modeling with linear complexity, and the output is a topology-aware feature set that integrates local topology and global semantics. ,in This represents the initial image patch embedding of the i-th image patch in the selected full-slice image. This represents the positional encoding of the i-th image patch in the selected full-slice image. This represents the topologically aware feature of the i-th image patch in the selected full-slice image.

[0012] S2-3: Variational Global Anchor Extraction: Constructing a slice-level global pseudo-bulk map with high signal-to-noise ratio. As a statistically robust biological benchmark signal; utilizing gated attention pooling to... Perform weighted aggregation to generate a deterministic full-slice global representation. Based on the variational information bottleneck (VIB) principle, a multilayer perceptron (MLP) is used to... Mapped to the mean of a multivariate Gaussian distribution With variance Random noise is introduced using reparameterization techniques. Sampling obtains specific probabilistic hidden anchor instances .

[0013] S2-4: Anchor-guided local dynamic correction and prediction: Constructing a feedback correction mechanism, utilizing... Explicitly correct the local representation of each image patch in the selected whole-slice image; for the i-th image patch in the selected whole-slice image, The image is concatenated with z, and a scaling factor specific to the i-th image patch in the selected full-slice image is generated through a dynamic feature modulation network. Translation coefficient The i-th image patch embedding representation in the selected full-slice image after affine transformation synthesis and correction. Finally, the prediction head will Mapping to the gene expression space yields the final local site-level gene expression predictions. .

[0014] S2-5: Constructing the end-to-end joint optimization objective function The objective function includes a site-level mean squared error loss that constrains local prediction fidelity. Maximize global anchor points Pseudo-batch graph Global reconstruction loss of mutual information and minimizing anchor points With the original input Mutual information to filter out KL divergence loss for local redundant noise .

[0015] The prediction model parameters are updated using backpropagation with the objective function.

[0016] S3: Traverse all slice images in the training dataset, and use methods S2-1 to S2-5 to jointly train the prediction model end-to-end with multiple tasks until the objective function is achieved. The trained prediction model converges.

[0017] S4: Forward inference: Input the new whole slice image to be predicted and its site coordinates into the trained prediction model, and directly output the spatial gene expression prediction spectrum of each corresponding sequencing site through forward inference.

[0018] As an improvement, in step S2-2, a position encoding representation is synthesized using a geometric position encoding module. The process is as follows:

[0019] First of all Normalize to a unit interval; for the i-th image patch in the selected full slice image, calculate the index set of its K nearest neighbor image patches as follows: For any and Calculate the selected image patch Relative to neighboring image patches physical offset ,in Represents the physical coordinates of the j-th image patch in the selected full-slice image; using a lightweight multilayer perceptron. The Softmax function maps the relative offset to adaptive spatial filtering coefficients. .

[0020]

[0021] Based on the above The initial neighborhood features after linear projection are weighted and modulated, and a location encoding representation with micro-environment directional topological awareness is synthesized from the selected full-slice image by layer normalization. :

[0022]

[0023] In the formula, It is a linear projection function.

[0024] The N image patches corresponding to the selected full slice image constitute .

[0025] As an improvement, in steps S2-3, hidden anchor point instances are extracted based on the variational information bottleneck (VIB) framework. Meets information compression and purification mechanisms:

[0026] Gated attention pooling Perform weighted aggregation:

[0027]

[0028]

[0029] in and These are all learnable parameters in the prediction model. , Let represent the topology-aware features of the i-th image patch and the j-th image patch, respectively. .

[0030] Aggregation features The mean of the multivariate Gaussian distribution is mapped to the mean. With variance parameter:

[0031]

[0032] Introducing reparameterization techniques to obtain specific anchor point instances ,in , This represents a random noise vector sampled from a standard multivariate Gaussian distribution;

[0033] The direct optimization objective of the variational information bottleneck is to maximize the objective function. By adjusting the coefficient Control the compression strength of information; in calculating feasible variational lower bound optimization, through the loss function To be reflected:

[0034]

[0035]

[0036] in, This indicates a global pseudo-batch prediction head. For standard Gaussian variational priors, The KL divergence is used to measure the difference between two distributions. Denotes the variational posterior distribution. Mutual information is used to measure the amount of information shared between two random variables. This represents the prior distribution.

[0037] As an improvement, in steps S2-4, the embedding representation is calculated. The steps are as follows:

[0038] Will and Perform element concatenation and input to the modulation sensing sensor. Obtain the adaptive transformation parameters:

[0039]

[0040] Subsequently, the original local embedding is dynamically modified through affine transformation, and the composite is corrected. :

[0041]

[0042] As an improvement, in steps S2-5, the objective function is jointly optimized. for:

[0043]

[0044] in, Used to ensure the fidelity of local predictions, and Each by weight and Control measures are implemented to enhance the semantic fidelity and structural robustness of global anchors.

[0045] Compared with the prior art, the present invention has at least the following advantages:

[0046] 1. This invention overcomes the limitations of traditional single-local supervision by implementing implicit regularization through global statistical consensus, fundamentally alleviating the artifact problem of overfitting techniques. The invention (corresponding to claim 1 and steps S3 and S5) introduces a pseudo-bulk map aggregated at the full slice level as a statistically robust slice-level supervision signal. According to the law of large numbers, this aggregation strategy mutually cancels out the independent random sampling noise generated by the low capture rate of individual sequencing sites, theoretically improving the signal-to-noise ratio. The global reconstruction loss is multiplied by a factor of two. In joint multi-task training, the global reconstruction loss is multiplied by a factor of two. As a powerful implicit regularization term, it forces the feature space of multi-instance aggregation to fall within a stable biological consensus range, thereby preventing the model from overfitting to fragmented and noisy local targets and ensuring that the model learns the essential histological and biological patterns.

[0047] 2. A unique "global-to-local" feedback correction mechanism was developed, achieving adaptive correction of local site representations. This invention (corresponding to claim 1 and steps S3 and S4) extracts probabilistic hidden anchors representing the core biological characteristics of the whole slice based on the variational information bottleneck (VIB) principle. Using these anchors as stable structural priors, a feedback correction module is constructed to dynamically interact the global anchors with site-specific local topological features. Explicit affine transformations of the original local representation are performed using the generated site-specific scaling and translation parameters, dynamically suppressing high-frequency updates and random technical biases that contradict the consensus of the global tissue environment, thus achieving spatially adaptive local calibration.

[0048] 3. This invention balances detailed local microenvironment topology modeling with low-complexity global context awareness, significantly improving heterogeneity capture capabilities. In the front-end processing, this invention (corresponding to claim 1 and step S2) utilizes a geometric position encoding (GPE) module to parameterize irregular physical distances between cells based on continuous relative coordinates, adaptively weighting and modulating neighborhood features to avoid quantization errors caused by rigid grid mapping, thus accurately capturing the local microenvironment topology that determines tissue specificity. Simultaneously, it employs a low-linear-complexity Nyströmformer layer for global context modeling. This architecture enables the model to significantly improve gene expression prediction accuracy and generalization robustness across multiple organs and irregular sequencing spaces while maintaining high computational efficiency. Attached Figure Description

[0049] Figure 1 This is an overview diagram of the framework of the STAR method of the present invention.

[0050] Figure 2 This is a simplified flowchart of the geometric position encoding (GPE) process in the method of the present invention.

[0051] Figure 3 To ensure robustness against monitored noise.

[0052] Figure 4 Analysis of the impact of regularization, correction mechanisms and VIB. Detailed Implementation

[0053] The present invention will now be described in further detail.

[0054] See Figure 1 and Figure 21. A method for predicting spatial gene expression in whole pathological sections based on anchoring correction, characterized by the following steps:

[0055] S1: Data preparation and definition: Obtain the training dataset, which contains multiple H&E stained pathological whole slide images and the ground truth of the spatial transcriptome expression profile corresponding to each whole slide image;

[0056] S2: Construct a prediction model, which includes a pathological basic model, a geometric location encoding module, a global context modeling based on Nyströmformer, a variational global anchoring module, a local correction module for anchor points, and a prediction head;

[0057] S2-1: Randomly select a whole-slice image from the training set and divide it into sections based on the irregular spatial sequencing sites it contains. Image patches constitute an image patch set The corresponding physical coordinate set is The true value of its spatial transcriptome expression profile is expressed as: ,in This represents the i-th image patch in the selected full-slice image. This represents the physical coordinates of the i-th image patch in the selected full-slice image. This represents the i-th image patch in the selected full-slice image. Dimensional gene expression vector (in the loss function).

[0058] S2-2: Topology-aware feature encoding: Encoding the aforementioned... Input a frozen pathological model and extract a neighborhood-aligned initial set of image patch embeddings. Using the Geometric Position Encoding (GPE) module, according to Calculate the relative offset of the irregular neighborhood of each image patch, map it to spatial filtering coefficients, and modulate the neighborhood features to synthesize a location-coded representation that includes the local micro-environment topology. ; then The input is a Nyströmformer layer for global context modeling with linear complexity, and the output is a topology-aware feature set that integrates local topology and global semantics. ,in This represents the initial image patch embedding of the i-th image patch in the selected full-slice image. This represents the positional encoding of the i-th image patch in the selected full-slice image. This represents the topologically aware feature of the i-th image patch in the selected full-slice image.

[0059] S2-3: Variational Global Anchor Extraction: Constructing a slice-level global pseudo-bulk map with high signal-to-noise ratio. As a statistically robust biological benchmark signal; utilizing gated attention pooling to... Perform weighted aggregation to generate a deterministic full-slice global representation. Based on the variational information bottleneck (VIB) principle, a multilayer perceptron (MLP) is used to... Mapped to the mean of a multivariate Gaussian distribution With variance Random noise is introduced using reparameterization techniques. Sampling obtains specific probabilistic hidden anchor instances .

[0060] S2-4: Anchor-guided local dynamic correction and prediction: Constructing a feedback correction mechanism, utilizing... Explicitly correct the local representation of each image patch in the selected whole-slice image; for the i-th image patch in the selected whole-slice image, The image is concatenated with z, and a scaling factor specific to the i-th image patch in the selected full-slice image is generated through a dynamic feature modulation network. Translation coefficient The i-th image patch embedding representation in the selected full-slice image after affine transformation synthesis and correction. Finally, the prediction head will Mapping to the gene expression space yields the final local site-level gene expression predictions. .

[0061] S2-5: Constructing the end-to-end joint optimization objective function The objective function includes a site-level mean squared error loss that constrains local prediction fidelity. Maximize global anchor points Pseudo-batch graph Global reconstruction loss of mutual information and minimizing anchor points With the original input Mutual information to filter out KL divergence loss for local redundant noise ;

[0062] The prediction model parameters are updated using backpropagation with the objective function.

[0063] S3: Traverse all slice images in the training dataset, and use methods S2-1 to S2-5 to jointly train the prediction model end-to-end with multiple tasks until the objective function is achieved. The trained prediction model converges.

[0064] S4: Forward inference: Input the new whole slice image to be predicted and its site coordinates into the trained prediction model, and directly output the spatial gene expression prediction spectrum of each corresponding sequencing site through forward inference.

[0065] Specifically, in step S2-2, the geometric position encoding module is used to synthesize the position encoding representation. The process is as follows:

[0066] First of all Normalize to a unit interval; for the i-th image patch in the selected full slice image, calculate the index set of its K nearest neighbor image patches as follows: For any and Calculate the selected image patch Relative to neighboring image patches physical offset ,in Represents the physical coordinates of the j-th image patch in the selected full-slice image; using a lightweight multilayer perceptron. The Softmax function maps the relative offset to adaptive spatial filtering coefficients. :

[0067]

[0068] Based on the above The initial neighborhood features after linear projection are weighted and modulated, and a location encoding representation with micro-environment directional topological awareness is synthesized from the selected full-slice image by layer normalization. :

[0069]

[0070] In the formula, It is a linear projection function.

[0071] The N image patches corresponding to the selected full slice image constitute .

[0072] Specifically, in steps S2-3, hidden anchor point instances are extracted based on the variational information bottleneck (VIB) framework. Meets information compression and purification mechanisms:

[0073] The core task at this stage is to embed spatial enhancement. Further compression extracts anchor instances that can characterize overall biological consensus. Traditional deterministic reconstruction mappings are often very fragile when faced with redundant information at the input, making it difficult to distinguish the core signal from background noise. To enhance the robustness of anchor points, this invention introduces the Variational Information Bottleneck (VIB) framework. This framework no longer learns a single feature vector, but instead learns... The probability distribution is designed to ensure that... Given the premise of "predictive sufficiency," redundant variance in the input data is forcibly filtered out by introducing information constraints. This mechanism ensures that the model can automatically remove technical noise that is not beneficial to the global signal, thereby refining a purer and more robust global benchmark.

[0074] In order to obtain a global representation that can represent the information of the entire slice Gated attention pooling is used for... Weighted aggregation: This mechanism can automatically identify and assign higher weights to key image patches.

[0075]

[0076]

[0077] in and These are all learnable parameters in the prediction model. , Let represent the topology-aware features of the i-th image patch and the j-th image patch, respectively. ;

[0078] Based on the design concept of Deep Variational Information Bottleneck (DeepVIB), we define the anchor point as a parameterized distribution. Instead of a single deterministic vector, this enhances the model's tolerance to noise. Specifically, it aggregates features. The mean of the multivariate Gaussian distribution is mapped to the mean. With variance parameter:

[0079]

[0080] To ensure gradient differentiability during the sampling process, a reparameterization technique is introduced to obtain specific anchor point instances. ,in , This represents a random noise vector sampled from a standard multivariate Gaussian distribution;

[0081] The direct optimization objective of the variational information bottleneck is to maximize the objective function. By adjusting the coefficient Control the compression strength of information; in calculating feasible variational lower bound optimization, through the loss function To be reflected:

[0082]

[0083]

[0084] in, This indicates a global pseudo-batch prediction head. For standard Gaussian variational priors, The KL divergence is used to measure the difference between two distributions. Denotes the variational posterior distribution. Mutual information is used to measure the amount of information shared between two random variables. This represents the prior distribution.

[0085] The item is responsible for maximizing the anchor point and Mutual information between them, forcing anchor points to lock in core biological consensus; The item is responsible for minimizing the anchor point and The mutual information between them serves as an information constraint penalty and filters out local technical noise and redundant variance that are irrelevant to the global benchmark.

[0086] Specifically, in steps S2-4, the embedding representation is calculated. The steps are as follows:

[0087] When applying global biological consistency constraints within the feature space, in order to preserve spatial heterogeneity and incorporate the local context specific to the detection site, and Perform element concatenation and input to the modulation sensing sensor. Obtain the adaptive transformation parameters:

[0088]

[0089] Subsequently, the original local embedding is dynamically modified through affine transformation, and the composite is corrected. :

[0090]

[0091] This makes the global anchor point As a stable structural prior, site-specific parameters and As a soft gating mechanism, it dynamically suppresses high-frequency updates and random technical biases that contradict the consensus of the global organizational environment. In this context, The modulation coefficients act as a stable structural prior, while the modulation coefficients function as soft gating. This mechanism effectively suppresses high-frequency noise updates that contradict global consensus, ensuring that the representation of each site retains both local specificity and global consistency.

[0092] This stage aims to utilize the purified product. Correction is performed on the site-level characterization. Although While capable of capturing fine spatial dependencies, direct supervision by high-noise site labels easily leads to overfitting to technical artifacts generated by random sampling. To mitigate this issue, this invention introduces a stable [system / mechanism] before the final prediction. Explicit correction is made to local representations. This strategy essentially imposes a global biological consistency constraint within the feature space.

[0093] Furthermore, considering the existence of spatial heterogeneity, this invention proposes a spatial adaptive calibration mechanism. Although It is shared across the entire slice, but the correction process must be customized in conjunction with the specific context of each site.

[0094] Specifically, in steps S2-5, the objective function is jointly optimized. for:

[0095]

[0096] in, Used to ensure the fidelity of local predictions, and Each by weight and Control measures are implemented to enhance the semantic fidelity and structural robustness of global anchors, respectively. During the prediction model training phase, and As an implicit regularization term, it constrains the feature space of multi-instance aggregation to fall within a stable biological consensus range; during the inference phase, the regularized global probability anchor point It is implicitly activated and explicitly participates in forward feedback, performing dynamic bias correction on the site representation of each local image patch.

[0097] Experiments and Analysis

[0098] 1. Experimental setup

[0099] We comprehensively evaluated the dataset on HEST-Benchmark, a large-scale public benchmark dataset specifically designed for spatial transcriptome (ST) prediction, containing real-world spatial transcriptome and pathological image pairings across multiple organs and technologies. In each task, the model input is... H&E image patch (corresponding) Magnification The model uses pixels and their spatial coordinates to predict the expression levels of the top 50 hypervariable genes (HVGs). Gene expression targets are pre-normalized using log1p. To evaluate the model's generalization ability, we follow a patient-stratified cross-validation protocol and use the Pearson correlation coefficient (PCC) as the primary evaluation metric to measure the consistency between the predicted results and the ground truth gene expression profiles.

[0100] To validate the effectiveness of STAR, we compared it with a series of benchmark methods, covering general...

[0101] Basic pathological models and task-specific methods.

[0102] (1) Pathological Basic Models (PFMs): Following the evaluation criteria of HEST-Benchmark, we tested five advanced basic models: CONCH, GigaPath, UNI, Virchow2, and UNI2. The extracted image patch embeddings were first dimensionality reduced by PCA ($d=256$), and then mapped to gene expression targets using Ridge Regression.

[0103] (2) Task-Specific Approaches: We selected representative models from the three mainstream paradigms discussed in relevant works. In regression-based methods, HisToGene utilizes the self-attention mechanism of Transformer for global dependency modeling, while TRIPLEX aggregates hierarchical organizational features across multiple scales. In retrieval-based methods, BLEEP achieves non-parametric inference based on reference queries by aligning bimodal embeddings. In generative benchmark methods, STFlow integrates a spatial attention mechanism within the ConditionalFlowMatching framework to model complex transcriptome distributions.

[0104] In subsequent experiments, unless otherwise specified, STAR's default hyperparameters are set to... as well as To ensure fairness in the evaluation and eliminate the influence of differences in backbone networks, we standardized the feature extraction stage for all task-specific methods, using the same PFM backbone network as STAR, as domain-relevant models capture significantly better histological semantics than standard backbone networks. We replaced the original backbone networks of these methods while strictly preserving their respective downstream modeling mechanisms.

[0105]

[0106] Table 1. Quantitative Comparison with HEST-Benchmark. This table reports the Pearson correlation coefficient (PCC) evaluation results for nine organ-specific datasets. All task-specific methods used frozen UNI2 as the feature extractor. Numerical values ​​are expressed as mean ± standard deviation.

[0107] 2. HEST-Benchmark Comparison Analysis

[0108] Overall, STAR achieved state-of-the-art (SOTA) performance with a mean Pearson correlation coefficient (PCC) of 0.4721, outperforming the best-performing benchmark method STFlow (0.4615) by more than 1.0%. Under the demanding conditions of sharing training settings and model hyperparameters across multiple subsets of datasets, task-specific benchmark methods are often affected by distribution shifts (e.g., STFlow performs exceptionally well on the SKCM dataset but shows a significant performance drop on other datasets). In contrast, STAR demonstrated superior robustness, ranking first across all datasets or maintaining a very small lead over the best-performing method.

[0109] The linear probing results in the upper half of Table 1, analyzed by method paradigms, confirm that PFM can capture rich histological semantics (e.g., UNI2 reaches 0.4166). Based on these powerful PFM features, different modeling paradigms perform differently: (1) Retrieval-based methods: BLEEP performs relatively poorly, indicating that its non-parametric reasoning ability is limited by the quality of the reference atlas. (2) Regression-based methods: Attention mechanism models represented by HisToGene and TRIPLEX perform well, confirming that global long-range modeling is an effective strategy to improve prediction performance. (3) Generative methods: STFlow has achieved considerable results by effectively combining flow matching with spatial attention mechanisms. Based on the above insights, STAR further raises the upper limit of prediction performance by introducing a global anchor point with pseudo-batch atlas as a supervision signal, which serves as a stable correction effect.

[0110]

[0111] Table 2. Performance consistency under different feature extractors. By replacing the backbone network with four different pathological baseline models (PFMs), we compared and evaluated the generalization ability of STAR with benchmark methods for each specific task.

[0112] Table 2 shows the generalization performance of different backbone features. STAR consistently outperforms all benchmark methods under four different PFM backbone networks, strongly demonstrating that its superiority stems from architectural innovation rather than dependence on a specific backbone network. Interestingly, the performance levels of each model (such as...) This is highly consistent with the linear detection trend in Table 1. It is worth noting that for STAR, we uniformly used a fixed hyperparameter ( as well as This is to align with the UNI2 configuration. Further performance gains may be achieved through fine-tuning for each backbone network individually. To visually verify these quantitative improvements...

[0113] 3. Robustness to random sampling noise

[0114] Ideally, robustness to random sampling noise should be assessed by comparison with a clean ground truth, but such ground truth is often difficult to obtain in real transcriptome (ST) data. To overcome this limitation, we designed an experiment to simulate the randomness of sequencing failures. Specifically, we introduced an element-wise Bernoulli mask into the training supervision signal and set the random dropout rate... The percentage was gradually increased from 10% to 90% to simulate sparsity under low capture efficiency. By forcing the model to learn from these fragmented targets and evaluating it on the original test set, we can effectively verify its ability to reconstruct biological patterns from sampling noise. Figure 3 As shown, we simulate different levels of sequencing failure by randomly masking the gene expression target (random inactivation rate increasing from 0 to 0.9). Different decay trajectories reveal performance differences between modeling paradigms. Traditional regression benchmark models (such as HisToGene and TRIPLEX) rely on deterministic point-to-point mappings. They inherently treat noisy sparse data as the learning target, which leads to performance issues when local gradients become unreliable. The learned mapping relationships can collapse. In contrast, the generative model STFlow fills in the information gaps by modeling the potential output distribution, demonstrating strong resilience. STAR, as an implementation of the regression paradigm, transcends the limitations of the regression paradigm itself through regularization during the training phase and correction during the inference phase. Pseudo-batch supervision acts as implicit regularization, preventing the model from overfitting to fragmented local targets; while global anchors provide explicit correction, dynamically calibrating the representation of points based on stable organizational consensus.

[0115] 4. Ablation test

[0116] To verify the rationality of the STAR architecture design, we conducted comprehensive ablation experiments and analyzed its core contributions in two dimensions: local geometric topology modeling and global statistical regularization and correction mechanism.

[0117] The impact of topology coding: We verified the role of spatial modeling by removing the position coding (PE) module from the STAR framework (see Table 3). Experimental results show that introducing topology-aware capabilities is crucial: adding GPE ( After [the experiment], the Pearson correlation coefficient (PCC) significantly increased from 0.4637 (without PE) to 0.4721. This confirms that capturing local intercellular interactions is the core foundation for resolving microenvironment heterogeneity. Regarding the encoding mechanism, GPE outperforms the grid-based APEG method proposed by TRIPLEX (0.4652). APEG forcibly maps irregularly distributed sites onto quantized grid points to fit standard convolution, while GPE parameterizes interactions through continuous relative coordinates. This allows the network to adaptively weigh neighborhood contributions based on precise physical distance, avoiding quantization errors caused by rigid mapping. Considering the increase [of certain parameters]... The value will increase the computational cost of modeling pairwise interactions, so we adopt... This is the default setting, which balances topological receptive field and computational efficiency.

[0118]

[0119] Table 3. Topological Coding Study. This experiment compared the performance differences between STAR and different benchmark schemes (including the PE-free version and the grid-based APEG method), and also analyzed the nearest neighbor number. Sensitivity analysis was performed. The values ​​are the average Pearson correlation coefficients (PCC) on the HEST-Benchmark dataset.

[0120] The impact of regularization, correction, and VIB: To decompose the contributions of each architectural component and analyze the impact of hyperparameters, we... Figure 4 The text shows the model performance as a function of anchor weights (…). The changing trend curve. The following three key findings validate our design choices:

[0121] (1) The effectiveness of implicit regularization: Once a global supervisory signal (i.e., The model performance immediately showed a leap. This validated our core insight: pseudo-bulk maps can serve as a biological benchmark for high signal-to-noise ratio (SNR). Regardless of signal strength, their very existence as supervisory signals provides a crucial regularization constraint, effectively preventing the model from overfitting to noise, thus achieving an initial performance breakthrough.

[0122] (2) Gains from explicit correction: By comparing STAR (red line) with the uncorrected version of STAR (gray line), it can be found that explicit correction and implicit regularization form a strong complement. The latter only uses global signals to constrain the feature space during the training phase through gradient sharing, while STAR achieves sustained superior performance by introducing correction during the inference phase. This confirms that actively modulating local features using global anchors can construct an effective feedback mechanism to correct random biases that cannot be solved by simply relying on auxiliary loss functions.

[0123] (3) Stabilizing effect of VIB: The deterministic variant (blue line) exhibits significant fluctuations across the entire weight range, highlighting the necessity of introducing the variational information bottleneck (VIB) to stabilize the optimization landscape. Furthermore, we observed... and There is a clear interaction between them: weaker regularization ( When the anchor point weight is appropriate ( The effect is best when it is at its highest level; while at a high level... Intervals, stronger regularization ( This makes global anchors indispensable. This indicates that as the model's reliance on global anchors increases, a more stringent information bottleneck is needed to filter redundant information and prevent feature collapse, thereby ensuring that anchors always serve as robust biological consensus rather than unstable sources of noise.

[0124] The above experiments show that:

[0125] 1. On the publicly available large-scale benchmark dataset HEST-Benchmark, the average Pearson correlation coefficient (PCC) of this invention (STAR) reaches 0.4721, which surpasses the state-of-the-art method STFlow (0.4615), achieving optimal or near-optimal performance on multiple organ-specific datasets.

[0126] 2. By simulating random sampling noise of different intensities (random inactivation rate of gene expression target from 10% to 90%), traditional regression methods break down when the noise is high. However, this invention can effectively resist noise interference and stably reconstruct biological patterns by using pseudo-batch supervised implicit regularization during the training phase and global anchor point explicit correction during the inference phase.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for predicting spatial gene expression in whole pathological slides based on anchoring correction, characterized in that: Includes the following steps: S1: Data preparation and definition: Obtain the training dataset, which contains multiple H&E stained pathological whole slide images and the ground truth of the spatial transcriptome expression profile corresponding to each whole slide image; S2: Construct a prediction model, which includes a pathological basic model, a geometric location encoding module, a global context modeling based on Nyströmformer, a variational global anchoring module, a local correction module for anchor points, and a prediction head; S2-1: Randomly select a whole-slice image from the training set and divide it into segments based on the irregular spatial sequencing sites it contains. Image patches constitute an image patch set The corresponding set of physical coordinates is The true value of its spatial transcriptome expression profile is expressed as: ,in This represents the i-th image patch in the selected full-slice image. This represents the physical coordinates of the i-th image patch in the selected full-slice image. This represents the i-th image patch in the selected full-slice image. Dimensional gene expression vector; S2-2: Topology-aware feature encoding: Encoding the... Input a frozen pathological model and extract a neighborhood-aligned initial set of image patch embeddings. ; Using the geometric position encoding module, according to Calculate the relative offset of the irregular neighborhood of each image patch, map it to spatial filtering coefficients, and modulate the neighborhood features to synthesize a location-coded representation that includes the local micro-environment topology. ; then The input is a Nyströmformer layer for global context modeling with linear complexity, and the output is a topology-aware feature set that integrates local topology and global semantics. ,in This represents the initial image patch embedding of the i-th image patch in the selected full-slice image. This represents the positional encoding of the i-th image patch in the selected full-slice image. This represents the topologically aware feature of the i-th image patch in the selected full-slice image; S2-3: Variational Global Anchor Extraction: Constructing a Slice-Level Global Pseudo-Batch Map with High Signal-to-Noise Ratio As a statistically robust biological benchmark signal; utilizing gated attention pooling to... Perform weighted aggregation to generate a deterministic full-slice global representation. Based on the variational information bottleneck (VIB) principle, a multilayer perceptron is used to... Mapped to the mean of a multivariate Gaussian distribution With variance Random noise is introduced using reparameterization techniques. Sampling obtains specific probabilistic hidden anchor instances ; S2-4: Anchor-guided local dynamic correction and prediction: Constructing a feedback correction mechanism, utilizing... Explicitly correct the local representation of each image patch in the selected whole-slice image; for the i-th image patch in the selected whole-slice image, The image is concatenated with z, and a scaling factor specific to the i-th image patch in the selected full-slice image is generated through a dynamic feature modulation network. Translation coefficient The i-th image patch embedding representation in the selected full-slice image after affine transformation synthesis and correction. Finally, the prediction head will Mapping to the gene expression space yields the final local site-level gene expression predictions. ; S2-5: Constructing the end-to-end joint optimization objective function The objective function includes a site-level mean squared error loss that constrains local prediction fidelity. Maximize global anchor points Pseudo-batch graph Global reconstruction loss of mutual information and minimizing anchor points With the original input Mutual information to filter out KL divergence loss for local redundant noise ; The prediction model parameters are updated using backpropagation with the objective function. S3: Traverse all slice images in the training dataset, and use methods S2-1 to S2-5 to jointly train the prediction model end-to-end through multiple tasks until the objective function is achieved. The trained prediction model is obtained by convergence. S4: Forward inference: Input the new whole slice image to be predicted and its site coordinates into the trained prediction model, and directly output the spatial gene expression prediction spectrum of each corresponding sequencing site through forward inference.

2. The method for predicting spatial gene expression in whole pathological slides based on anchoring correction as described in claim 1, characterized in that, In step S2-2, the position encoding representation is synthesized using the geometric position encoding module. The process is as follows: First of all Normalize to a unit interval; For the i-th image patch in the selected full-slice image, the set of indices of its K nearest neighbor image patches is denoted as follows: For any and Calculate the selected image patch Relative to neighboring image patches physical offset ,in Represents the physical coordinates of the j-th image patch in the selected full-slice image; using a lightweight multilayer perceptron. The Softmax function maps the relative offset to adaptive spatial filtering coefficients. : Based on the above The initial neighborhood features after linear projection are weighted and modulated, and a location encoding representation with micro-environment directional topological awareness is synthesized from the selected full-slice image by layer normalization. : In the formula, It is a linear projection function; The N image patches corresponding to the selected full slice image constitute .

3. The method for predicting spatial gene expression in whole pathological slides based on anchoring correction as described in claim 1, characterized in that, In steps S2-3, hidden anchor point instances are extracted based on the Variational Information Bottleneck (VIB) framework. Meets information compression and purification mechanisms: Gated attention pooling Perform weighted aggregation: in and These are all learnable parameters in the prediction model. , Let represent the topology-aware features of the i-th image patch and the j-th image patch, respectively. ; Aggregation features The mean of the multivariate Gaussian distribution is mapped to the mean. With variance parameter: Introducing reparameterization techniques to obtain specific anchor point instances ,in , This represents a random noise vector sampled from a standard multivariate Gaussian distribution; The direct optimization objective of the variational information bottleneck is to maximize the objective function. By adjusting the coefficient Control the compression strength of information; in calculating feasible variational lower bound optimization, through the loss function To be reflected: in, This indicates a global pseudo-batch prediction head. For standard Gaussian variational priors, The KL divergence is used to measure the difference between two distributions. Denotes the variational posterior distribution. Mutual information is used to measure the amount of information shared between two random variables. This represents the prior distribution.

4. The method for predicting spatial gene expression in whole pathological slides based on anchoring correction as described in claim 1, characterized in that, In steps S2-4, the embedding representation is calculated. The steps are as follows: Will and Perform element concatenation and input to the modulation sensing sensor. Obtain the adaptive transformation parameters: Subsequently, the original local embedding is dynamically modified through affine transformation, and the composite is corrected. :

5. The method for predicting spatial gene expression in whole pathological sections based on anchoring correction as described in claim 1, characterized in that, In steps S2-5, the objective function is jointly optimized. for: in, Used to ensure the fidelity of local predictions, and Each by weight and Control measures are implemented to enhance the semantic fidelity and structural robustness of global anchors.