A multi-granularity zentropy-based visual saliency prediction method and system

CN122530753BActive Publication Date: 2026-09-04TONGJI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610999989.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-04
Estimated Expiration
2046-07-07

AI Technical Summary

Technical Problem

[0003]本发明的目的在于提供一种基于多粒度Zentropy的视觉显著性预测方法及系统,以解决现有技术中存在的以下技术问题:现有视觉显著性预测方法参数量庞大导致计算开销高、难以在资源受限的临床设备和边缘计算环境中部署,纯空间回归方法缺乏对视觉不确定性的显式建模、无法提供可解释的注意力表征,现有方法在跨场景应用中泛化能力有限、难以迁移到孤独症谱系障碍注意力分析和视觉假体编码等下游任务,预测的显著性图存在过度平滑问题、边界定位精度不足的技术问题

Benefits of technology

[0021]Compared with existing technologies, the present invention has the following beneficial effects: First, by adopting a self-supervised visual Transformer encoder architecture with frozen parameters, utilizing pre-trained semantic representations, and training only the feature adaptation module, fine-grained Zentropy exploration module, and saliency decoder simultaneously, the number of trainable parameters is reduced to 2.7M, which is 99.7% lower than UniAR's 848M parameters and 98.9% lower than Temp-Sal's 242M parameters. This significantly reduces computational overhead and memory usage, supports deployment on edge devices, achieves real-time CPU inference of 82.39ms/frame, and achieves near-ceiling performance of normalized scan path saliency NSS=1.998 on the SALICON dataset, surpassing Temp-Sal's NSS=1.967 and UniAR's NSS=1.947, while maintaining a competitive correlation coefficient CC=0.911. Second, by employing a dual decomposition of object-level and decision-level granularity in the fine-grained Zentropy exploration module, hierarchical uncertainty is explicitly modeled. At the object level, local structural consistency is captured through a family of multi-scale deep convolutions, while at the decision level, discriminative uncertainty is captured through foreground-background prototype comparison. This dual-granularity complementarity enhances the accuracy of saliency prediction. Compared to the single-granularity variant, the full FGZE improves CC from 0.901 to 0.933 and NSS from 1.911 to 1.969. The learned Zentropy representation is significantly positively correlated with human gaze uncertainty (r=0.6796, p<0.001), providing interpretable attention representations. Third, by employing the boundary-aware Zentropy regularization strategy, the Sobel operator is used to extract high-frequency transition regions from the real gaze distribution, constraining the uncertainty of model learning to align with the boundary regions, reducing excessive smoothing of the saliency prediction map, and improving the accuracy of local gaze localization. After introducing regularization, the NSS increases from 1.969 to 1.998, and the uncertainty gradient magnitude of the boundary region (0.835) is significantly higher than that of the core region (0.334) and the background region (0.645). Compared with the CBAM attention mechanism, the boundary alignment error MSE of FGZE is reduced by 62%, from 0.242 to 0.092, and the L1 error is reduced by 43%, from 0.430 to 0.243. Fourth, by using the Zentropy attention gating mechanism, the multi-scale feature representation is modulated based on the total Zentropy uncertainty map, which suppresses high uncertainty responses and enhances stable gaze-related responses. Compared with direct feature attention, the NSS is improved by 8.4%, from 1.841 to 1.998. Compared with the Shannon entropy method, multi-granularity Zentropy gating improves the NSS by 2.3%, from 1.9534 to 1.9983, producing a more concentrated spatial saliency response and reducing activation leakage in the background region.Fifth, the uncertainty representation of learning demonstrated good generalization ability in cross-scenario applications. In the zero-shot transfer to the attention analysis task for autism spectrum disorder, the consistency of the predicted saliency map with the typical developmental group's gaze pattern (NSS=1.355, CC=0.786) was significantly higher than the consistency with the autism spectrum disorder group's gaze pattern (NSS=1.253, CC=0.747), with a paired test P<0.001. The zero-shot classification of TD-vs-ASD achieved an AUC=0.7125, exceeding the Gaussian center-biased baseline AUC=0.6570, an improvement of 8.5%. In the visual prosthesis coding task, the FGZE-guided gaze region resource allocation, compared to uniform downsampling, improved the recognition accuracy from 15.81% to 34.90%, an improvement of 19.09%, and the average prediction confidence from 0.0606 to 0.1016, an improvement of 67.7%, validating that regions with low Zentropy uncertainty preferentially retain semantic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122530753B_ABST
    Figure CN122530753B_ABST
Patent Text Reader

Abstract

The application discloses a visual saliency prediction method and system based on multi-granularity Zentropy, and relates to the technical field of computer vision. The method comprises the following steps: pre-processing a visual image to be analyzed to obtain a model input image; inputting the model input image into a self-supervised visual Transformer encoder with frozen parameters to extract image block-level semantic features, and mapping the image block-level semantic features into dense spatial features through a feature adaptation module; performing multi-scale context aggregation on the dense spatial features to obtain multi-scale feature representations; inputting the multi-scale feature representations into a fine-granularity Zentropy exploration module, constructing object-level granularity uncertainty through fuzzy similarity modeling, and constructing decision-level granularity uncertainty through foreground and background prototype comparison; fusing the two kinds of granularity uncertainty to obtain a total Zentropy uncertainty map, and performing attention gate modulation on the multi-scale feature representations according to the uncertainty map; and performing saliency decoding on the modulated features to generate a human eye fixation saliency prediction map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a visual saliency prediction method and system based on multi-granularity Zentropy. Background Technology

[0002] Visual saliency prediction aims to estimate regions in images that attract human attention, and has significant application value in human-computer interaction, clinical diagnosis, and assistive vision systems. Existing visual saliency prediction methods are mainly based on deep learning architectures, achieving saliency estimation through dense pixel-level regression. To capture complex multi-scale contextual dependencies, recent methods have increasingly adopted large-scale Transformer architectures; for example, UniAR uses a unified model with 848M parameters to achieve cross-domain generalization. However, these methods suffer from the following technical problems: First, the large number of model parameters leads to high computational costs, making it difficult to deploy in resource-constrained clinical devices and edge computing environments; second, pure spatial regression methods lack explicit modeling of visual uncertainty and cannot provide interpretable attention representations; third, existing methods have limited generalization ability in cross-scenario applications, making it difficult to transfer to downstream tasks such as attention analysis in autism spectrum disorder and visual prosthesis coding; fourth, the predicted saliency maps suffer from over-smoothing, resulting in insufficient boundary localization accuracy. Chinese patent CN 116258768 A discloses a first-person gaze prediction method based on attention transfer. This scheme extracts spatiotemporal features of optical flow images through a dual-stream I3D network and generates a gaze prediction map by fusing saliency images and attention images through an attention transfer mechanism. However, this scheme relies on temporal information processing, resulting in high computational complexity, and it does not explicitly model uncertainty. Chinese patent CN 118736244A discloses a visual saliency prediction method. This scheme trains a lightweight model using a knowledge distillation strategy, but it does not perform hierarchical modeling of visual uncertainty, making it difficult to provide interpretable attention representations. Therefore, a visual saliency prediction method with efficient parameters, strong interpretability, and excellent cross-scene generalization ability is needed. Summary of the Invention

[0003] The purpose of this invention is to provide a visual saliency prediction method and system based on multi-granularity Zentropy, in order to solve the following technical problems existing in the prior art: existing visual saliency prediction methods have a large number of parameters, resulting in high computational overhead and difficulty in deployment in resource-constrained clinical equipment and edge computing environments; pure spatial regression methods lack explicit modeling of visual uncertainty and cannot provide interpretable attention representations; existing methods have limited generalization ability in cross-scenario applications and are difficult to transfer to downstream tasks such as attention analysis of autism spectrum disorder and visual prosthesis coding; and the predicted saliency map has technical problems such as over-smoothing and insufficient boundary localization accuracy.

[0004] To achieve the above objectives, the present invention adopts the following technical solution: The present invention provides a visual saliency prediction method based on multi-granularity Zentropy, comprising: Step S1, acquiring a visual image to be analyzed, and performing size normalization, pixel normalization, and tensor quantization processing on the visual image to be analyzed to obtain a model input image; Step S2, inputting the model input image into a self-supervised visual Transformer encoder with frozen parameters, extracting image block-level semantic features, and mapping them to dense spatial features through a feature adaptation module; Step S3, performing multi-scale context aggregation on the dense spatial features to obtain a multi-scale feature representation; Step S4, ... The multi-scale feature representation is input into a fine-grained Zentropy exploration module. Object-level granular uncertainty is constructed through fuzzy similarity modeling, and decision-level granular uncertainty is constructed through foreground-background prototype comparison. In step S5, the object-level and decision-level granular uncertainties are fused to obtain a total Zentropy uncertainty map. An attention-gated modulation of the multi-scale feature representation is then performed based on the total Zentropy uncertainty map using a softmax gating function adjusted by a temperature parameter. In step S6, the modulated features are saliency-decoded to generate a human eye gaze saliency prediction map corresponding to the visual image to be analyzed. Optionally, step S7 further includes an abnormal attention difference assessment based on the human eye gaze saliency prediction map, performed by calculating the correlation difference with gaze distribution in typical developmental groups and autism spectrum disorder groups, or by performing gaze region resource allocation based on uncertainty-based resolution allocation.

[0005] The core technical features of this invention lie in the fine-grained Zentropy exploration module and the boundary-aware regularization strategy. In step S4, the object-level granular uncertainty Z_obj is obtained as follows: local topological features are extracted using a parallel multi-scale deep convolutional family, including 3×3, 5×5, and 7×7 deep convolutions. The normalized cosine similarity between the multi-scale feature representation and the local topological features is calculated to obtain the local membership score P_local. The object-level granular uncertainty is determined according to Z_obj=-log(P_local+ε), where ε is a stability term. The decision-level granular uncertainty Z_approx is obtained as follows: based on the feature amplitude, high-response regions are selected from the multi-scale feature representation to construct a foreground prototype, and low-response regions are selected to construct a background prototype. The similarity between each spatial location feature and the foreground prototype and the background prototype is calculated to obtain the global confidence distribution P_global. The decision-level granular uncertainty is determined according to Z_approx=-log(P_global+ε). In step S5, the attention-gated modulation is implemented as follows: the object-level granularity uncertainty and the decision-level granularity uncertainty are added to obtain the total Zentropy uncertainty map Z_total=Z_obj+Z_approx; the multi-scale feature representation is modulated by the softmax gating function of the temperature parameter τ to generate the modulated feature F_modulated, where F_modulated is the element-wise product of F_ms and softmax(-Z_total / τ), F_ms is the multi-scale feature representation, and the softmax function normalizes the total Zentropy uncertainty map along the spatial dimension.

[0006] Furthermore, the self-supervised visual Transformer encoder with frozen parameters maintains fixed parameters during training, training only the feature adaptation module, the fine-grained Zentropy exploration module, and the saliency decoder. This frozen encoder architecture utilizes pre-trained semantic representations while significantly reducing the number of trainable parameters by training only task-specific modules, achieving a balance between parameter efficiency and prediction performance.

[0007] Furthermore, the feature adaptation module projects the image block-level semantic features into a feature tensor with spatial resolution. The multi-scale context aggregation is implemented through a dilated spatial pyramid pooling module, which uses convolutional operations with different receptive fields to capture contextual dependencies. By using the dilated spatial pyramid pooling module, the receptive field can be expanded without increasing computational cost, capturing contextual information at different scales and providing multi-scale feature inputs for subsequent uncertainty modeling.

[0008] Preferably, the self-supervised visual Transformer encoder for the frozen parameters has an embedding dimension of 768, a depth of 12, and an image patch size of 14. These parameter configurations correspond to the DINOv2-ViT-B / 14 architecture, providing semantically rich image feature representations.

[0009] Specifically, the multi-scale depth convolution family includes 3×3 depth convolution, 5×5 depth convolution, and 7×7 depth convolution. These depth convolutions of different sizes are used in parallel to extract local topological features, capture the consistency of local structures at different scales, and provide multi-scale local representations for the calculation of object-level granular uncertainty.

[0010] Furthermore, the foreground prototype is obtained by averaging the features of high-response regions, and the background prototype is obtained by averaging the features of low-response regions. Using a Top-K feature amplitude selection strategy, foreground and background regions can be automatically identified, and prototype representations of the foreground and background can be constructed for calculating decision-level granular uncertainty.

[0011] Furthermore, the method also includes a training phase, which employs a boundary-aware Zentropy regularization strategy: using an edge detection operator to extract high-frequency transition regions from the true gaze distribution to obtain boundary priors. G, construct the regularization loss L_reg=λ_z·L1(normalize(Z_total), normalize( G), where λ_z is the regularization coefficient, normalize represents the normalization operation, and L1 represents the L1 distance. The regularization loss is combined with the saliency loss to form the total loss L_total = L_sal + λ_z·L_reg. This boundary-aware regularization strategy is based on the idea in rough set theory that boundary regions carry uncertainty. By aligning the uncertainty learned by the constrained model with the boundary regions of the true gaze distribution, it reduces the over-smoothing of the saliency prediction map and improves the accuracy of local gaze localization.

[0012] Preferably, the saliency loss L_sal combines KL divergence, Pearson correlation coefficient, similarity, and normalized scan path saliency. By combining multiple evaluation metrics, global distribution consistency and local gaze localization accuracy can be jointly optimized, avoiding bias from selecting a single metric.

[0013] Furthermore, the training phase employs a two-stage curriculum learning strategy: the first stage uses a hybrid dataset with data augmentation strategies, employing a cosine annealing-based learning rate decay strategy, including the regularization loss; the second stage fine-tunes on the original clean samples, using a fixed learning rate and disabling all augmentation strategies. This two-stage curriculum learning strategy achieves a coarse-to-fine optimization process: the first stage enhances representation robustness through spatial perturbation augmentation, while the second stage fine-tunes on clean data to improve positioning accuracy and stability.

[0014] Specifically, the data augmentation strategies include MixUp augmentation and Mosaic augmentation. The training phase employs the AdamW optimizer, and gradient clipping is applied in the first phase. MixUp and Mosaic augmentations can improve the model's robustness to spatial perturbations and changes in scene combinations, while gradient clipping can stabilize the early training process.

[0015] Further, the abnormal attention difference assessment in step S7 includes: calculating the normalized scan path significance and Pearson correlation coefficient between the predicted human eye gaze saliency map and the gaze distribution of the typical developmental group; calculating the normalized scan path significance and Pearson correlation coefficient between the predicted human eye gaze saliency map and the gaze distribution of the autism spectrum disorder group; and classifying the attention differences between the typical developmental group and the autism spectrum disorder group based on the differences in the normalized scan path significance and Pearson correlation coefficient through threshold judgment or a classifier. This abnormal attention difference assessment method can quantify the consistency difference between the saliency map predicted by the model and the gaze distribution of different groups, providing transferable gaze priors for the computational modeling of gaze differences related to autism spectrum disorder.

[0016] Furthermore, the gaze region resource allocation in step S7 includes: determining low-uncertainty regions and high-uncertainty regions based on the total Zentropy uncertainty map; retaining high-resolution representations for the low-uncertainty regions; and performing downsampling processing on the high-uncertainty regions to achieve gaze region resource allocation for visual prosthesis encoding. This gaze region resource allocation method can prioritize the preservation of visually information-rich regions under severe bandwidth constraints, thereby improving downstream recognition performance.

[0017] This invention also provides a visual saliency prediction system based on multi-granularity Zentropy, comprising: an image preprocessing module for performing size normalization, pixel normalization, and tensor quantization on the visual image to be analyzed to obtain a model input image; a frozen encoder module for extracting image block-level semantic features from the model input image; a feature adaptation module for mapping the image block-level semantic features to dense spatial features; a multi-scale aggregation module for performing multi-scale context aggregation on the dense spatial features to obtain a multi-scale feature representation; a fine-grained Zentropy exploration module for constructing object-level granular uncertainty and decision-level granular uncertainty respectively; a Zentropy gating module for fusing the object-level granular uncertainty and the decision-level granular uncertainty to obtain a total Zentropy uncertainty map, and performing attention-gated modulation on the multi-scale feature representation based on the total Zentropy uncertainty map; and a saliency decoding module for performing saliency decoding on the modulated features to generate a human eye gaze saliency prediction map.

[0018] Furthermore, the fine-grained Zentropy exploration module includes: an object-level granularity unit, comprising a multi-scale depth convolution and cosine similarity calculation unit, used to calculate object-level granularity uncertainty; and a decision-level granularity unit, comprising a high-low response region selection unit and a prototype comparison unit, used to construct foreground prototypes and background prototypes and calculate decision-level granularity uncertainty.

[0019] Preferably, the system is deployed in a CPU environment, supporting real-time saliency prediction. This deployment method is suitable for resource-constrained clinical devices and edge computing environments.

[0020] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described visual saliency prediction method based on multi-granularity Zentropy.

[0021] Compared with existing technologies, the present invention has the following beneficial effects: First, by adopting a self-supervised visual Transformer encoder architecture with frozen parameters, utilizing pre-trained semantic representations, and training only the feature adaptation module, fine-grained Zentropy exploration module, and saliency decoder simultaneously, the number of trainable parameters is reduced to 2.7M, which is 99.7% lower than UniAR's 848M parameters and 98.9% lower than Temp-Sal's 242M parameters. This significantly reduces computational overhead and memory usage, supports deployment on edge devices, achieves real-time CPU inference of 82.39ms / frame, and achieves near-ceiling performance of normalized scan path saliency NSS=1.998 on the SALICON dataset, surpassing Temp-Sal's NSS=1.967 and UniAR's NSS=1.947, while maintaining a competitive correlation coefficient CC=0.911. Second, by employing a dual decomposition of object-level and decision-level granularity in the fine-grained Zentropy exploration module, hierarchical uncertainty is explicitly modeled. At the object level, local structural consistency is captured through a family of multi-scale deep convolutions, while at the decision level, discriminative uncertainty is captured through foreground-background prototype comparison. This dual-granularity complementarity enhances the accuracy of saliency prediction. Compared to the single-granularity variant, the full FGZE improves CC from 0.901 to 0.933 and NSS from 1.911 to 1.969. The learned Zentropy representation is significantly positively correlated with human gaze uncertainty (r=0.6796, p<0.001), providing interpretable attention representations. Third, by employing the boundary-aware Zentropy regularization strategy, the Sobel operator is used to extract high-frequency transition regions from the real gaze distribution, constraining the uncertainty of model learning to align with the boundary regions, reducing excessive smoothing of the saliency prediction map, and improving the accuracy of local gaze localization. After introducing regularization, the NSS increases from 1.969 to 1.998, and the uncertainty gradient magnitude of the boundary region (0.835) is significantly higher than that of the core region (0.334) and the background region (0.645). Compared with the CBAM attention mechanism, the boundary alignment error MSE of FGZE is reduced by 62%, from 0.242 to 0.092, and the L1 error is reduced by 43%, from 0.430 to 0.243. Fourth, by using the Zentropy attention gating mechanism, the multi-scale feature representation is modulated based on the total Zentropy uncertainty map, which suppresses high uncertainty responses and enhances stable gaze-related responses. Compared with direct feature attention, the NSS is improved by 8.4%, from 1.841 to 1.998. Compared with the Shannon entropy method, multi-granularity Zentropy gating improves the NSS by 2.3%, from 1.9534 to 1.9983, producing a more concentrated spatial saliency response and reducing activation leakage in the background region.Fifth, the uncertainty representation of learning demonstrated good generalization ability in cross-scenario applications. In the zero-shot transfer to the attention analysis task for autism spectrum disorder, the consistency of the predicted saliency map with the typical developmental group's gaze pattern (NSS=1.355, CC=0.786) was significantly higher than the consistency with the autism spectrum disorder group's gaze pattern (NSS=1.253, CC=0.747), with a paired test P<0.001. The zero-shot classification of TD-vs-ASD achieved an AUC=0.7125, exceeding the Gaussian center-biased baseline AUC=0.6570, an improvement of 8.5%. In the visual prosthesis coding task, the FGZE-guided gaze region resource allocation, compared to uniform downsampling, improved the recognition accuracy from 15.81% to 34.90%, an improvement of 19.09%, and the average prediction confidence from 0.0606 to 0.1016, an improvement of 67.7%, validating that regions with low Zentropy uncertainty preferentially retain semantic information. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the ZenGaze framework as a whole and a schematic diagram of edge application scenarios. Figure 2 This is a schematic diagram of the overall architecture of the ZenGaze framework. Figure 3 A schematic diagram of downstream evaluation results with zero samples. Figure 4 This is a schematic diagram of the empirical analysis of Zentropy representation. Figure 5 This is a schematic diagram of the internal characterization comparison analysis. Detailed Implementation

[0023] This invention provides a visual saliency prediction method and system based on multi-granularity Zentropy. Addressing the technical problems of existing visual saliency prediction methods, such as large parameter count, difficulty in edge deployment, insufficient interpretable modeling of visual uncertainty, limited cross-scene application capabilities, and overly smoothed saliency maps, this invention achieves a parameter-efficient architecture by freezing the pre-trained visual Transformer encoder. A fine-grained Zentropy exploration module is designed to explicitly model hierarchical uncertainty through dual decomposition at both object-level and decision-level granularity. A boundary-aware Zentropy regularization strategy is introduced to constrain the model's learning uncertainty to align with the gaze boundary region. A Zentropy attention gating mechanism is employed to modulate the feature response through temperature parameters, implementing a two-stage course learning strategy to optimize the training process. This invention achieves near-ceiling performance of Normalized Scan Path Saliency (NSS) = 1.998 on the SALICON dataset, requiring only 2.7M trainable parameters, supporting real-time CPU inference at 82.39 ms / frame, and achieving an AUC of 0.7125 on the ASD attention analysis task with zero-shot transfer. In the visual prosthesis encoding task, it improves recognition accuracy by 19.09% compared to uniform downsampling. The core improvements are: enhancing saliency localization accuracy and interpretability through multi-granularity uncertainty decomposition and boundary-aware regularization; achieving a balance between parameter efficiency and prediction performance through frozen encoder architecture and Zentropy gated modulation; and enhancing robustness and cross-scenario generalization ability through two-stage learning.

[0024] Example 1 This embodiment provides a basic implementation of a visual saliency prediction method based on multi-granularity Zentropy. The method extracts semantic features through a self-supervised visual Transformer encoder with frozen parameters. These features are then processed by a feature adaptation module and multi-scale context aggregation to form a multi-scale feature representation. A fine-grained Zentropy exploration module estimates uncertainties at both the object and decision levels. These two uncertainties are fused to obtain a total Zentropy uncertainty map. Attention gating based on the total Zentropy uncertainty map is used to modulate the visual features. Finally, a saliency decoder outputs a human eye gaze saliency prediction map. (Reference) Figure 2 The overall architecture of the ZenGaze framework is shown below.

[0025] Step S1: Image Preprocessing. Obtain the visual image to be analyzed, I∈R^(H×W×3). Normalize the image size to 224×224 resolution, normalize the pixel values ​​to the [0,1] interval, and convert it to tensor format to obtain the model input image. During training and benchmarking, all input images are uniformly adjusted to 224×224 resolution to balance computational efficiency and semantic resolution. The predicted saliency map and the ground truth map are represented at 128×128 resolution. During optimization, the ground truth distribution is aligned to the network output resolution using bilinear interpolation.

[0026] Step S2: Freeze Encoder Feature Extraction. The input image to the model is fed into a self-supervised visual Transformer encoder with frozen parameters to extract image patch-level semantic features. Specifically, DINOv2-ViT-B / 14 is used as the frozen encoder, with encoder parameters kept fixed during training (requires_grad=False), producing image patch embeddings with a feature dimension D=768. The encoder depth L=12, and the image patch size P=14. The frozen encoder extracts a feature sequence F_patch∈R^(N×D), where N is the number of image patches. The image patch-level semantic features are projected into dense spatial features X∈R^(h×w×C) with spatial resolution through a feature adaptation module, where C is the number of intermediate channels output by the feature adaptation module, which can be set according to the network width; C can be 64, 128, 256, or 512.

[0027] Step S3: Multi-scale context aggregation. Apply the dilated spatial pyramid pooling module to the dense spatial feature X to perform multi-scale context aggregation, obtaining the multi-scale feature representation F_ms∈R^(h×w×d). The dilated spatial pyramid pooling module uses convolution operations with different receptive fields to capture contextual dependencies, expanding the receptive field without increasing computational cost, capturing contextual information at different scales, and providing multi-scale feature input for subsequent uncertainty modeling.

[0028] Step S4: Fine-grained Zentropy exploration module dual-granularity uncertainty modeling. The multi-scale feature representation F_ms is input into the fine-grained Zentropy exploration module to construct object-level granular uncertainty Z_obj and decision-level granular uncertainty Z_approx. The object-level granular uncertainty is constructed through fuzzy similarity modeling: a parallel multi-scale deep convolutional family K_set={3×3,5×5,7×7} is used to extract local topological features from the multi-scale feature F_ms, obtaining the local topological feature F_local. The normalized cosine similarity between F_ms and F_local is calculated to obtain the local membership score P_local=normalize(S_cos(F_ms,F_local)). The object-level granular uncertainty Z_obj=-log(P_local+ε), where ε=1×10^-8 is a stable term to prevent numerical instability.

[0029] Decision-level granular uncertainty is constructed through foreground-background prototype comparison: Foreground prototype p_fg = mean(F_ms[Ω_fg]) is constructed by selecting Top-K high-response regions from F_ms based on feature amplitude, and background prototype p_bg = mean(F_ms[Ω_bg]) is constructed by selecting Top-K low-response regions, where K is the top 10% of the total number of spatial locations in the feature map, and Ω_fg and Ω_bg are the indices of the Top-K high-response and low-response regions, respectively. The cosine similarity between each spatial location feature and the foreground and background prototypes is calculated to obtain the global confidence distribution P_global = max(S_cos(F_ms,p_fg),S_cos(F_ms,p_bg)). The decision-level granular uncertainty Z_approx = -log(P_global + ε).

[0030] Object-level granular uncertainty captures local structural consistency, while decision-level granular uncertainty captures foreground-background discrimination uncertainty; the two complement each other to form a hierarchical uncertainty representation. At the object level, multi-scale deep convolutional families capture local topological features at different scales, and fuzzy similarity is calculated to reflect the degree of consistency of local structures. At the decision level, Top-K feature amplitude selection automatically identifies foreground and background regions, constructs prototype representations of the foreground and background, and calculates the similarity between each spatial location and the prototype to reflect the confidence level of foreground-background discrimination. These two granularities of uncertainty characterize the hierarchical uncertainty in the visual search process from different perspectives.

[0031] Step S5: Zentropy Gating Modulation. The object-level and decision-level granular uncertainties are fused to obtain the total Zentropy uncertainty map Z_total = Z_obj + Z_approx. The multi-scale feature representation is then attention-gated based on the total Zentropy uncertainty map using a softmax gating function with temperature parameter τ, generating the modulated feature F_modulated = F_ms⊙softmax(-Z_total / τ), where ⊙ represents element-wise multiplication. The softmax function normalizes the total Zentropy uncertainty map along the spatial dimension. The temperature parameter τ controls the sharpness of the gating function; a smaller τ concentrates the gating in the low-uncertainty region.

[0032] Step S6: Saliency Decoding. The modulated features F_modulated are saliency decoded to generate a human eye gaze saliency prediction map S∈[0,1]^(H×W) corresponding to the visual image to be analyzed. The saliency decoder uses a progressive upsampling structure to gradually restore the spatial resolution and output a normalized saliency probability distribution.

[0033] Step S7: Application Output. Based on the human eye fixation saliency prediction map, visual attention analysis, abnormal attention difference assessment, or fixation region resource allocation can be performed. In the abnormal attention difference assessment application, the differences in attention patterns are quantified by calculating the correlation difference between the prediction saliency map and the fixation distribution of typical developmental groups and autism spectrum disorder groups. In the fixation region resource allocation application, low uncertainty regions and high uncertainty regions are identified based on the total Zentropy uncertainty map. High-resolution representation is retained for low uncertainty regions, while high uncertainty regions are downsampled.

[0034] Workflow: The input image is processed by the DINOv2 freeze encoder to extract image patch features F_patch, which are then mapped to dense spatial features X by the feature adaptation module. After ASPP multi-scale aggregation, multi-scale features F_ms are obtained. These are input to the FGZE module to calculate the object-level granular uncertainty Z_obj and the decision-level granular uncertainty Z_approx, respectively. The total uncertainty Z_total is then fused and modulated using Zentropy gating, resulting in the feature F_modulated = F_ms ⊙ softmax(-Z_total / τ). Finally, the saliency decoder outputs the predicted saliency map S. The entire process forms a complete workflow from semantic feature extraction, multi-scale context aggregation, hierarchical uncertainty modeling, uncertainty-aware feature modulation to saliency prediction.

[0035] Technical Results: Due to the adoption of a frozen encoder architecture, only the feature adaptation module, FGZE module, and saliency decoder are trained, reducing the number of trainable parameters to 2.7M and the total number of parameters to 89M. This represents a 99.7% reduction compared to UniAR's 848M parameters and a 98.9% reduction compared to Temp-Sal's 242M parameters, significantly reducing computational overhead and memory usage, and supporting deployment on edge devices. Real-time inference was achieved on an Intel Core i9-13900K CPU with an inference latency of 82.39ms / frame. Ablation experiments (Table 2) show that starting from the DINOv2 baseline, the gradual introduction of each component brings consistent performance improvements: after introducing the multi-scale encoder, NSS increased from 1.841 to 1.876; after introducing object-level granularity, CC reached 0.901 and NSS reached 1.911; after introducing decision-level granularity, foreground / background discrimination capability was further improved; and the complete FGZE (dual-granularity joint) achieved optimal distribution alignment (CC=0.933, SIM=0.798, KL=0.195). Introducing Zentropy regularization improves NSS from 1.969 to 1.998. While boundary-aware regularization encourages more concentrated activation, CC slightly decreases (0.933 → 0.911), but global distribution alignment (KL divergence) remains stable. Due to Zentropy gated modulation, compared to direct feature attention (no uncertainty modeling), NSS is improved by 8.4%, from 1.841 to 1.998. On the SALICON validation set, it achieves competitive performance with NSS = 1.998 ± 0.003, CC = 0.911 ± 0.002, SIM = 0.795 ± 0.001, and KL = 0.201 ± 0.002. NSS reaches near-ceiling levels, surpassing Temp-Sal's NSS of 1.967 and UniAR's NSS of 1.947.

[0036] Example 2 This embodiment, building upon Embodiment 1, introduces a boundary-aware Zentropy regularization strategy and a two-stage curriculum learning strategy during the training process. The boundary-aware Zentropy regularization strategy, based on the idea in rough set theory that boundary regions carry uncertainty, aligns the uncertainty of the model's learning with the gaze boundary region, reducing excessive smoothing of the saliency prediction map and improving local gaze localization accuracy. The two-stage curriculum learning strategy implements a coarse-to-fine optimization process. The first stage enhances representation robustness through spatial perturbation, while the second stage fine-tunes on clean data to improve localization accuracy and stability.

[0037] Boundary-aware Zentropy regularization: using the Sobel operator to extract high-frequency transition regions from the true gaze distribution G as boundary priors. The G. Sobel operator is a classic edge detection operator that extracts high-frequency transition regions by calculating the magnitude of the image gradient. (Boundary prior) G reflects the boundary region of the human gaze point set in the true gaze distribution. The regularization loss is constructed as L_reg = λ_z·L1(normalize(Z_total), normalize( G), where λ_z = 0.05 is the regularization coefficient, normalize represents the normalization operation, and L1 represents the L1 distance. This regularized loss constraint normalizes the total Zentropy uncertainty Z_total and the normalized boundary response. G is spatially aligned, encouraging the model to produce higher uncertainty responses in the boundary region.

[0038] The saliency loss L_sal combines multiple evaluation metrics: KL divergence (weight 10.0), Pearson correlation coefficient (CC) (weight 2.0), similarity (SIM) (weight 1.0), and normalized scan path saliency (NSS) (weight 1.0). KL divergence measures the difference between the predicted saliency distribution and the actual gaze distribution; Pearson correlation coefficient measures the linear correlation between the two; similarity measures the degree of overlap in the distributions; and normalized scan path saliency measures the response strength of the predicted saliency at the actual gaze point location. By combining multiple evaluation metrics, global distribution consistency and local gaze localization accuracy are jointly optimized, avoiding the bias of selecting a single metric. The total loss L_total = L_sal + λ_z·L_reg.

[0039] The training process is divided into two phases, totaling 20 rounds. The first phase (rounds 1-15) is the robust representation learning phase, using a mixed dataset with data augmentation strategies. These strategies include MixUp augmentation (α=0.2) and Mosaic augmentation. MixUp augmentation generates new training samples by linearly interpolating two images, while Mosaic augmentation generates new training samples by stitching four images together. These augmentation strategies improve the model's robustness to spatial perturbations and changes in scene composition. The AdamW optimizer is used, with an initial learning rate η=1×10^-3, weight decay of 1×10^-4, and a batch size of 10. Cosine annealing scheduling is used for the learning rate, T_max=15, and a minimum learning rate of 1×10^-5. The training objective includes a Zentropy regularization term L_reg. Gradient clipping is applied in the first phase, with an upper bound of L2 norm of 1.0, to stabilize the early training process.

[0040] The second phase (rounds 16-20) is the clean data fine-tuning phase, where fine-tuning is performed on the original clean SALICON samples. All data augmentation strategies are disabled, and training is conducted using only the original images. The learning rate is fixed at η = 1 × 10^-5. This phase fine-tunes the saliency distribution learned in the first phase, improving localization accuracy and stability. Through these two phases of learning, a gradual process from robust representation learning to precise localization optimization is achieved.

[0041] During validation, both the predicted saliency map P and the ground truth labeled map G were evaluated as continuous spatial distributions. Images were L1 normalized to obtain P_prob and G_prob, and then standardized with zero mean and unit variance to obtain P_norm and G_norm. Key monitored metrics included: KL divergence (calculated between P_prob and G_prob), Pearson correlation coefficient (CC) (calculated between P_norm and G_norm), and normalized scan path saliency (NSS surrogate) (estimated by weighting the standardized predicted P_norm with the continuous gaze distribution G).

[0042] To reduce model selection bias, checkpoint selection is based on the composite validation score M_comp = 2CC + NSS, balancing relevance consistency and gaze localization. Compared to single-index selection, this strategy provides a more stable balance between global distribution alignment and local gaze sensitivity. In practice, optimizing only CC often results in a smoother saliency map, while combining it with NSS improves sensitivity to high-response gaze regions.

[0043] Technical Results: After introducing boundary-aware Zentropy regularization, the NSS improved from 1.969 to 1.998, and the uncertainty gradient magnitude in the boundary region (0.835) was significantly higher than that in the core region (0.334) and the background region (0.645). Compared to the CBAM attention mechanism, FGZE reduced the boundary alignment error MSE by 62%, from 0.242 to 0.092, and the L1 error by 43%, from 0.430 to 0.243. Boundary-aware regularization encourages uncertainty to concentrate in the boundary region, reducing excessive smoothing of the saliency prediction map and improving boundary localization accuracy. The two-stage learning process reduced the standard deviation of NSS for five independent runs to 0.003, improving training stability. The first stage of data augmentation improved the model's robustness to spatial perturbations, while the second stage of clean data fine-tuning improved localization accuracy. The two stages worked together to achieve a balance between robustness and accuracy.

[0044] Example 3 This embodiment provides generalization experiments on different backbone network architectures to verify the generalization ability of the FGZE module under different backbone network architectures. The frozen encoder is replaced from DINOv2 with ResNet50 or ViT-S, and the improvement in saliency prediction performance of the FGZE module is evaluated under the same training settings. All training schedules, optimization settings, data augmentation, and evaluation protocols are kept consistent to ensure a comparative analysis.

[0045] Alternative 1: ResNet50 backbone network. A ResNet50 pre-trained network from ImageNet is used as the frozen encoder to extract convolutional feature maps. ResNet50 is a convolutional architecture that emphasizes local texture and spatial feature extraction. Baseline (ResNet50 backbone + decoder, without FGZE): CC=0.823, NSS=1.654. ​​After introducing FGZE (ResNet50 backbone + FGZE + decoder): CC=0.856 (+0.033), NSS=1.782 (+0.128). ResNet50 is suitable for scenarios with extremely limited computational resources and requires fewer parameters.

[0046] Alternative Solution 2: ViT-S Backbone Network. A lightweight Vision Transformer (ViT-S) is used as the frozen encoder, with an embedding dimension D=384. ViT-S strikes a balance between lightweightness and performance, featuring global token interaction and moderate receptive field aggregation. Baseline (ViT-S backbone + decoder, without FGZE): CC=0.867, NSS=1.789. After introducing FGZE (ViT-S backbone + FGZE + decoder): CC=0.894 (+0.027), NSS=1.901 (+0.112). ViT-S achieves a balance between lightweightness and performance.

[0047] Alternative Solution 3: DINOv2 Backbone Network (Example 1). A self-supervised pre-trained DINOv2 network is used as the frozen encoder, with an embedding dimension D=768. DINOv2 provides the strongest semantic representation and is suitable for scenarios requiring high accuracy. Baseline (DINOv2 backbone + decoder, without FGZE): CC=0.876, NSS=1.841. After introducing FGZE (DINOv2 backbone + FGZE + decoder): CC=0.911 (+0.035), NSS=1.998 (+0.157). DINOv2 provides the strongest semantic representation and is suitable for scenarios requiring high accuracy.

[0048] Technical Results: The FGZE module consistently improves saliency prediction performance across different backbone network architectures, validating the generalization ability of multi-granularity uncertainty modeling. Compared to ResNet50, the Transformer-based backbone network shows a greater improvement in NSS after introducing FGZE, possibly because the global token interaction structure of the Vision Transformer provides more compatible feature representations for hierarchical uncertainty decomposition, leading to stronger interactions between object-level and decision-level uncertainty components. Across all backbone network families, the maximum relative gain occurs in the NSS metric, consistent with ZenGaze's design goals, which emphasize improving local gaze response and high-frequency saliency transitions, rather than simply optimizing global saliency smoothness.

[0049] Example 4 This embodiment provides comparative experiments with existing methods on the SALICON and MIT1003 datasets to verify the performance advantages of the method of the present invention. The comparison methods include nine state-of-the-art methods such as Temp-Sal, UniAR, Visual Saliency Transformer, and DeepGaze IIE. The evaluation metrics include Normalized Scan Path Significance (NSS), Pearson Correlation Coefficient (CC), Similarity (SIM), and KL Divergence.

[0050] Test conditions: Dataset 1 is SALICON, with 10,000 images in the training set, 5,000 images in the validation set, and 5,000 images in the test set. Each image contains multiple human gaze annotations. Dataset 2 is MIT, with 1,003 natural scene images, using 5-fold cross-validation, and includes both on-the-fly training (MIT only) and transfer learning (SALICON→MIT) settings. The hardware environment is a Linux workstation with an Intel Core i9-13900K CPU, 128GB RAM, and a single NVIDIA GeForce RTX4090 GPU (24GB VRAM). All random operations (weight initialization, dataset partitioning, DataLoader shuffling) are initialized using a fixed random seed of 42.

[0051] Comparison methods: Temp-Sal (2023 CVPR), 242M parameters, SALICON validation set NSS=1.967, CC=0.914; UniAR (2024 NeurIPS), 848M parameters, SALICON validation set NSS=1.947, CC=0.918; VisualSaliency Transformer (2021 ICCV), SALICON validation set NSS=1.892, CC=0.897; DeepGaze IIE (2021 ICCV), SALICON validation set NSS=1.856, CC=0.883. The method of this invention, ZenGaze, has 89M parameters (2.7M trainable).

[0052] SALICON validation set results (Table 1): ZenGaze achieved NSS=1.998±0.003 (optimal), CC=0.911±0.002, SIM=0.795±0.001, and KL=0.201±0.002. The results are reported as the mean ± standard deviation of 5 independent runs, confirming performance stability. Compared to Temp-Sal, NSS improved by 1.6% (1.967→1.998), and trainable parameters decreased by 98.9% (242M→2.7M). Compared to UniAR, NSS improved by 2.6% (1.947→1.998), and trainable parameters decreased by 99.7% (848M→2.7M). ZenGaze achieved the best NSS metric, indicating that the proposed uncertainty-aware modeling strategy produces a more spatially concentrated saliency response. Qualitative checks further show a reduction in oversmoothing in the prediction graph.

[0053] Table 1. Quantitative Comparison Results of Visual Saliency Prediction on the SALICON Validation Set Note: * indicates that the DINOv2 backbone network is frozen, with only 2.7M trainable parameters; the best result is bolded, and the second-best result is underlined. The results are reported as the mean ± standard deviation of 5 independent runs.

[0054] Table 2 Ablation Experiment Results of the SALICON Validation Set Note: The components are added progressively, and the optimal result is highlighted in bold.

[0055] MIT1003 cross-domain generalization results (Table 3): To differentiate between architectural gain and pre-training performance, FGZE was evaluated under a strict 5-fold cross-validation protocol using both de novo training (MIT only) and transfer learning (SALICON→MIT) settings. In the transfer learning setting, ZenGaze achieved NSS=2.8452, CC=0.8234, and KL=0.4812, surpassing all compared methods and achieving the best overall performance. In the de novo training setting, ZenGaze achieved NSS=2.3156 and CC=0.7845, still outperforming the baseline architecture (without FGZE) with NSS=2.1023 and CC=0.7512. Transfer learning consistently improved all models, highlighting the importance of large-scale neural canonical priors.

[0056] Table 3. Cross-domain generalization results of 5-fold cross-validation on the MIT1003 dataset. Note: Under the transfer learning setting, the method of this invention achieves optimal overall performance, NSS=2.8452, KL=0.4812.

[0057] Quantitative analysis of boundary alignment (Table 4): To assess the contribution of FGZE to traditional spatial attention, an ablation study was conducted, replacing FGZE with a parameter-matched CBAM baseline. Relative to boundary prior | The G|,CBAM baseline produces an MSE of 0.242 and an L1 error of 0.430. In contrast, FGZE trained with the Zentropy regularization term L_reg reduces the MSE to 0.092 and the L1 error to 0.243. Figure 5 Qualitative comparisons show that traditional attention produces a relatively diffuse response, while FGZE reduces activation leakage to the background region by explicitly modeling uncertainty in the transition region, resulting in more spatially focused saliency predictions.

[0058] Table 4. Quantitative Analysis Results of Boundary Alignment (Compared with Boundary Priors) (Comparison of G| errors) Note: Lower error indicates better alignment with the boundary prior space; the optimal result is bolded. FGZE reduces MSE by 62% and L1 error by 43% compared to CBAM.

[0059] Validation of the correlation between uncertainty and human gaze ( Figure 4b): To examine whether the performance gain of ZenGaze is related to its uncertainty-aware modeling strategy rather than just parameter scale, the spatial features of the learned Zentropy representation were analyzed. For each image, the empirical uncertainty was estimated using the spatial Shannon entropy H_human calculated from the real gaze view, and the global average activation Z_mean of the predicted Zentropy map was computed in parallel. The two showed a positive correlation (r=0.6796, p<0.001), indicating that the learned Zentropy representation captures spatial patterns statistically consistent with the increased gaze uncertainty.

[0060] Uncertainty gradient analysis in the boundary region Figure 4 a) Inspired by the rough set formula, it is assumed that increased uncertainty tends to occur near the transition region (BND = Upper-Lower). The real view is divided into three regions: core, background, and boundary, with the boundary region extracted using the Sobel operator. Average Zentropy gradient magnitude. Z_total is highest in the boundary region (0.835), compared to the core region (0.334) and the background region (0.645). These observations suggest that the learned uncertainty representation is spatially concentrated in the object-context transition region.

[0061] Conclusion: Experimental results demonstrate that this invention significantly improves saliency localization accuracy (NSS=1.998) through multi-granularity Zentropy uncertainty modeling and boundary-aware regularization, while substantially reducing trainable parameters (2.7M) by freezing the encoder architecture, achieving a balance between parameter efficiency and prediction performance. The significant correlation between the learned uncertainty representation and human gaze uncertainty (r=0.6796) verifies the interpretability of the method. Cross-domain generalization experiments (MIT1003 NSS=2.8452) and boundary alignment analysis (MSE reduced by 62%) further demonstrate the robustness and boundary localization capability of the method. Explicit uncertainty-aware representation modeling can serve as an effective alternative to purely extended saliency architectures, offering advantages in parameter efficiency, interpretability, and cross-scene generalization ability.

[0062] Example 5 This embodiment provides a zero-shot downstream application scenario: ASD attention difference analysis. Zero-shot evaluation is performed on the Saliency4ASD dataset; the model was trained only on SALICON and not fine-tuned on ASD data. Differences in attention patterns are quantified by calculating the NSS and CC of the gaze distribution between the ZenGaze prediction saliency map and the typical developmental (TD) and autism spectrum disorder (ASD) groups, and a TD-vs-ASD zero-shot classification protocol is constructed to evaluate classification performance.

[0063] Experimental Setup: The dataset Saliency4ASD contains gaze data from children with developmental typicality (TD) and autism spectrum disorder (ASD). Evaluation methods involved calculating the NSS and CC of the ZenGaze prediction saliency map against the TD / ASD gaze distributions and performing paired statistical tests. A TD-vs-ASD zero-shot classification protocol was constructed, and classification performance was evaluated using AUC. This dataset was used only for zero-shot evaluation without additional fine-tuning. Since ZenGaze was trained only on SALICON, comparing the model predictions with the gaze distributions of the TD and ASD groups provides an indirect measure of behavioral alignment differences.

[0064] Experimental results: such as Figure 3 As shown in (a), the consistency between the ZenGaze prediction saliency map and the TD fixation pattern (NSS=1.355, CC=0.786) was significantly higher than that with the ASD fixation pattern (NSS=1.253, CC=0.747), with a paired test P<0.001. This result indicates that the learned representation is sensitive to the differences in the distribution of fixation assignments between the two groups. The violin plot shows that the distribution density of CC and NSS predicted by ZenGaze is more concentrated in the high-scoring region in the TD group, while it is more dispersed in the ASD group.

[0065] To mitigate the impact of simple center bias, a zero-shot classification protocol, TD-vs-ASD, was additionally constructed. The FGZE achieved an AUC of 0.7125, surpassing the Gaussian center prior baseline (AUC = 0.6570), representing an 8.5% improvement. These observations suggest that the learned uncertainty-perceived representation captures spatial attention patterns beyond simple position bias, even though the framework was not specifically designed for clinical diagnostic training.

[0066] Technical Significance: This study validates that the uncertainty representation of ZenGaze learning can capture spatial attention patterns beyond simple positional bias, is sensitive to differences in gaze distributions between TD and ASD, provides transferable gaze priors for computational modeling of ASD-related gaze differences, and has potential application value as a digital biomarker. This application demonstrates the generalization ability of the method in interdisciplinary fields, transferring from zero-shot standard saliency prediction tasks to gaze analysis tasks related to clinical diagnosis.

[0067] Example 6 This embodiment provides a zero-sample downstream application scenario: visual prosthetic gaze region encoding. It performs gaze-guided gaze region resource allocation under ultra-low resolution prosthetic vision settings, verifying that semantic information content is preferentially preserved in low Zentropy uncertainty regions, thus improving downstream recognition performance under severe bandwidth constraints.

[0068] Experimental setup: 1,493 valid samples were selected from the ImageNet validation set. The recognition agent was a ResNet-50 classification model. The encoding strategies were compared: uniform downsampling vs. FGZE-guided gaze region allocation. FGZE-guided allocation determined low-uncertainty and high-uncertainty regions based on the total Zentropy uncertainty map Z_total, retaining high-resolution representations for low-uncertainty regions and downsampling for high-uncertainty regions.

[0069] Experimental results: such as Figure 3 As shown in (b), uniform compression significantly reduces recognition accuracy (15.81%), while FGZE guided allocation improves accuracy to 34.90% (+19.09%), while maintaining higher prediction confidence (0.1016 vs. 0.0606, an improvement of 67.7%). These results indicate that low Zentropy regions preferentially preserve semantically rich content, improving downstream recognition performance under severe bandwidth constraints. FGZE guided coding preferentially preserves visually rich regions under severe bandwidth constraints.

[0070] Technical significance: This invention validates the principle of prioritizing the preservation of semantic information in low Zentropy uncertainty regions, supports gaze region encoding of visual prostheses under severe bandwidth constraints, and provides an uncertainty-guided resource allocation strategy for low-bandwidth visual prosthesis rendering frameworks. This application demonstrates the practical value of the method in assistive vision systems, optimizing the encoding efficiency of visual prostheses through uncertainty-aware resource allocation.

[0071] It is understood that the multi-scale deep convolutional family in the above embodiments is not limited to {3×3, 5×5, 7×7}, but can also be {3×3, 5×5} or {3×3, 7×7} or other size combinations, such as {3×3, 5×5, 7×7, 9×9} to capture local topological features with a larger receptive field. The Top-K selection strategy is not limited to feature magnitude ranking, but can also use alternative metrics such as feature activation energy, gradient norm, and attention weights for foreground-background prototype selection. The Zentropy gating function is not limited to softmax(-Z_total / τ), but can also use sigmoid gating, tanh gating, or linear gating.

[0072] Clearly, the boundary extraction operator is not limited to the Sobel operator; Canny edge detection, Laplacian operator, Prewitt operator, or learnable edge detection networks can also be used. The two-stage learning strategy is not limited to a 15-round + 5-round division; other division ratios can be used, such as 10-round + 10-round, 12-round + 8-round, or a progressive learning rate decay strategy can be used instead of a fixed-stage division. The frozen encoder is not limited to DINOv2; other self-supervised or supervised pre-trained visual encoders can also be used, such as CLIP, MAE, BEiT, and Swin Transformer.

Claims

1. A visual saliency prediction method based on multi-granularity Zentropy, characterized in that, Includes the following steps: Step S1: Obtain the visual image to be analyzed, and perform size normalization, pixel normalization, and tensor quantization on the visual image to be analyzed to obtain the model input image; Step S2: Input the model input image into a self-supervised visual Transformer encoder with frozen parameters, extract image block-level semantic features, and map them into dense spatial features through a feature adaptation module; Step S3: Perform multi-scale context aggregation on the dense spatial features to obtain multi-scale feature representations; Step S4: Input the multi-scale feature representation into the fine-grained Zentropy exploration module, construct object-level granular uncertainty through fuzzy similarity modeling, and construct decision-level granular uncertainty through foreground-background prototype comparison; The object-level granular uncertainty is obtained as follows: Local topological features are extracted using a parallel multi-scale deep convolution family, which includes deep convolutions of different sizes; the normalized cosine similarity between the multi-scale feature representation and the local topological features is calculated to obtain the local membership score P_local; The object-level granular uncertainty is determined according to Z_obj=-log(P_local+ε), where ε is a stability term; Step S5: Fuse the object-level granularity uncertainty and the decision-level granularity uncertainty to obtain a total Zentropy uncertainty map, and perform attention-gated modulation on the multi-scale feature representation based on the total Zentropy uncertainty map using a softmax gating function adjusted by temperature parameters; The decision-level granularity uncertainty is obtained as follows: Select high-response regions from the multi-scale feature representation based on feature amplitude to construct a foreground prototype, and select low-response regions to construct a background prototype; Calculate the similarity between each spatial location feature and the foreground prototype and the background prototype respectively to obtain the global confidence distribution P_global; Determine the decision-level granularity uncertainty according to Z_approx=-log(P_global+ε), where ε is a stability term; The foreground prototype is obtained by averaging the features of the high-response region, and the background prototype is obtained by averaging the features of the low-response region. The attention-gated modulation in step S5 is implemented as follows: the object-level granularity uncertainty and the decision-level granularity uncertainty are added together to obtain the total Zentropy uncertainty map Z_total=Z_obj+Z_approx; the multi-scale feature representation is modulated by the softmax gate function of the temperature parameter τ to generate the modulated feature F_modulated, where F_modulated is the element-wise product of F_ms and softmax(-Z_total / τ), F_ms is the multi-scale feature representation, and the softmax function normalizes the total Zentropy uncertainty map along the spatial dimension; Step S6: Perform saliency decoding on the modulated features to generate a human eye gaze saliency prediction map corresponding to the visual image to be analyzed; Step S7: Based on the human eye gaze saliency prediction map, perform abnormal attention difference assessment by calculating the correlation difference with the gaze distribution of typical developmental groups and autism spectrum disorder groups, or perform gaze region resource allocation by uncertainty-based resolution allocation.

2. The method according to claim 1, characterized in that, The self-supervised visual Transformer encoder with frozen parameters keeps its parameters fixed during training, and only the feature adaptation module, the fine-grained Zentropy exploration module, and the saliency decoder are trained.

3. The method according to claim 1 or 2, characterized in that, The feature adaptation module is used to project the image block-level semantic features into feature tensors with spatial resolution. The multi-scale context aggregation is implemented through the dilated spatial pyramid pooling module, which uses convolution operations with different receptive fields to capture contextual dependencies.

4. The method according to claim 1, characterized in that, The method also includes a training phase, which employs a boundary-aware Zentropy regularization strategy. The edge detection operator is used to extract high-frequency transition regions from the true gaze distribution to obtain the boundary prior. G; Construct the regularization loss L_reg = λ_z·L1(normalize(Z_total), normalize( G), where λ_z is the regularization coefficient, normalize represents the normalization operation, and L1 represents the L1 distance; The regularization loss is combined with the significance loss to form the total loss L_total = L_sal + λ_z·L_reg; The training phase employs a two-stage course learning strategy: The first stage involves training on a hybrid dataset using data augmentation strategies, employing a cosine annealing-based learning rate decay strategy, and including the regularization loss. The second stage involves fine-tuning on the original clean samples, using a fixed learning rate and disabling all augmentation strategies.

5. The method according to claim 1, characterized in that, The abnormal attention difference assessment described in step S7 includes: Calculate the normalized scan path significance and Pearson correlation coefficient between the predicted human eye fixation saliency map and the fixation distribution of a typical developmental group; Calculate the normalized scan path significance and Pearson correlation coefficient between the predicted human eye gaze saliency map and the gaze distribution in the autism spectrum disorder group; Based on the difference in the significance of the normalized scanning path and the Pearson correlation coefficient, attentional differences between the typical developmental group and the autism spectrum disorder group are classified by threshold judgment or classifier. The gaze region resource allocation in step S7 includes: Based on the total Zentropy uncertainty map, the low uncertainty region and the high uncertainty region are determined; The low-uncertainty region retains a high-resolution representation, while the high-uncertainty region undergoes downsampling processing to achieve gaze region resource allocation for visual prosthetic coding.

6. A visual saliency prediction system based on multi-granularity Zentropy, characterized in that, include: The image preprocessing module is used to perform size normalization, pixel normalization, and tensor quantization on the visual image to be analyzed, thereby obtaining the model input image; The freeze encoder module is used to extract image block-level semantic features from the input image of the model; The feature adaptation module is used to map the image block-level semantic features into dense spatial features; The multi-scale aggregation module is used to perform multi-scale context aggregation on the dense spatial features to obtain multi-scale feature representations; The fine-grained Zentropy exploration module is used to construct object-level and decision-level granular uncertainties, respectively. The object-level granular uncertainty is obtained as follows: local topological features are extracted using a parallel multi-scale deep convolutional family, which includes deep convolutions of different sizes; the normalized cosine similarity between the multi-scale feature representation and the local topological features is calculated to obtain the local membership score P_local; the object-level granular uncertainty is determined according to Z_obj=-log(P_local+ε), where ε is a stability term. The Zentropy gating module is used to fuse the object-level granular uncertainty and the decision-level granular uncertainty to obtain a total Zentropy uncertainty map, and to perform attention-gated modulation on the multi-scale feature representation based on the total Zentropy uncertainty map. The decision-level granular uncertainty is obtained as follows: based on the feature amplitude, high-response regions are selected from the multi-scale feature representation to construct a foreground prototype, and low-response regions are selected to construct a background prototype; the similarity between each spatial location feature and the foreground prototype and the background prototype is calculated to obtain the global confidence distribution P_global; the decision-level granular uncertainty is determined according to Z_approx=-log(P_global+ε), where ε is a stability term. The foreground prototype is obtained by averaging the features of the high-response region, and the background prototype is obtained by averaging the features of the low-response region. The attention-gated modulation is implemented as follows: the object-level granularity uncertainty and the decision-level granularity uncertainty are added to obtain the total Zentropy uncertainty map Z_total=Z_obj+Z_approx; the multi-scale feature representation is modulated by the softmax gating function of the temperature parameter τ to generate the modulated feature F_modulated, where F_modulated is the element-wise product of F_ms and softmax(-Z_total / τ), F_ms is the multi-scale feature representation, and the softmax function normalizes the total Zentropy uncertainty map along the spatial dimension; The saliency decoding module is used to perform saliency decoding on the modulated features and generate a human eye gaze saliency prediction map.

7. The system according to claim 6, characterized in that, The fine-grained Zentropy exploration module includes: The object-level granularity unit, containing multi-scale depthwise convolution and cosine similarity calculation units, is used to calculate object-level granularity uncertainty; The decision-level granularity unit includes high and low response region selection units and prototype comparison units, which are used to construct foreground prototypes and background prototypes and calculate decision-level granularity uncertainty.

Citation Information

Patent Citations

  • Attention transfer-based first visual angle fixation point prediction method

    CN116258768A

  • Visual saliency prediction method and system

    CN118736244A

  • Multi-granularity fusion video clip retrieval method based on audio importance perception

    CN120256674A

  • Two-way visual saliency detection method and device combining difference guidance and texture enhancement

    CN120747542A