A method for super-resolution fluorescence image quality evaluation without reference
Patent Information
- Application Number
- CN202610291927.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-11
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-03-11
AI Technical Summary
然而在荧光显微成像中,高质量参考图像往往难以获取,且成像噪声受到光子统计特性、标记密度及成像条件等多重因素影响,导致这些通用指标与实际图像的可辨识度及可用性之间存在显著差距
[0013]本发明方法无需提供参考图像,可直接针对单幅荧光显微图像输出连续质量分数,适用于真实实验中“无GT”的常见场景。
Smart Images

Figure CN122223516B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and biomedical image processing technology, specifically to a referenceless quality assessment method for super-resolution fluorescence images. Background Technology
[0002] Fluorescence microscopy has become a key tool for observing cellular substructures and dynamic processes in living cells. In recent years, advancements in super-resolution imaging techniques such as structured illumination microscopy (SIM) and their accompanying physical reconstruction and deep learning reconstruction algorithms have enabled the generation of various reconstruction results from the same set of raw imaging data. Different reconstruction algorithms and their parameters exhibit varying performance in noise suppression, detail restoration, and artifact control. Therefore, objectively evaluating the quality and structural reliability of reconstructed images is crucial for ensuring the reliability of subsequent quantitative analysis.
[0003] Currently used image quality assessment metrics, such as Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM), typically rely on high-quality reference images and are designed based on statistical assumptions about natural images. However, in fluorescence microscopy, high-quality reference images are often difficult to obtain, and imaging noise is affected by multiple factors such as photon statistical properties, label density, and imaging conditions. This leads to a significant gap between these general metrics and the discernibility and usability of actual images. On the other hand, methods commonly used in fluorescence super-resolution, such as Fourier Ring Correlation (FRC), are mainly used for resolution evaluation or system resolving power analysis, and are difficult to use for continuous and unified quantitative evaluation of the overall quality, artifact status, and structural realism of a single reconstructed image.
[0004] Therefore, there is an urgent need to develop an image quality assessment method that better fits the characteristics of fluorescence microscopy images, requires no reference image, and can stably output quality scores. Summary of the Invention
[0005] To address the aforementioned issues, this invention provides a referenceless quality assessment method for super-resolution fluorescence images. A referenceless quality assessment model is constructed and trained, and the trained model is used to evaluate the quality of the super-resolution fluorescence images. The referenceless quality assessment model includes a CNN branch network, a ViT branch network, and a cross-attention fusion module.
[0006] The training process of the no-reference quality assessment model includes the following steps:
[0007] S1. Construct a fluorescence image dataset with quality gradients; where each fluorescence image is labeled with a quality label;
[0008] S2. Input the fluorescence image into the ViT branch network to obtain enhanced global features; the ViT branch network includes a ViT backbone network and a multi-head attention module;
[0009] S3. Input the fluorescence image into a CNN branch network to obtain local features;
[0010] S4. Input the enhanced global and local features into the cross-attention fusion module to obtain the comprehensive features; extract the class embedding vector from the comprehensive features and input it into the MLP to obtain the prediction quality score;
[0011] S5. Calculate the joint loss function based on the comprehensive features, and perform backpropagation to optimize the model parameters until the model converges; introduce a comparison ranking loss into the joint loss function to keep the relative quality distance between image pairs in the feature space consistent with the difference between their true quality labels.
[0012] The beneficial effects of this invention are:
[0013] The method of this invention does not require a reference image and can directly output continuous mass fractions from a single fluorescence microscopy image, making it suitable for common "GT-free" scenarios in real experiments.
[0014] This invention integrates the local detail sensitivity of CNNs with the global structure modeling capabilities of ViT, making it more suitable for fluorescence images with repetitive structures and where local details determine recognizability. It also introduces a comparative ranking loss to ensure that the distance in the feature space aligns with the quality level, improving fine-grained quality difference discrimination and stability across annotation scales.
[0015] This invention is applicable to the screening of super-resolution reconstruction results, monitoring of delayed imaging quality (such as imaging quality degradation caused by photobleaching or phototoxicity), and early stopping judgment and quality warning during unsupervised reconstruction training. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the no-reference quality evaluation model structure of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Please see Figure 1This invention provides a referenceless quality assessment method for super-resolution fluorescence images. The method takes a single fluorescence image as input and employs a CNN branch network and a ViT branch network to capture the local high-frequency degradation features and global morphological context of the fluorescence image, respectively. Subsequently, a cross-attention mechanism is used to mine complementary information between the two stream features, achieving deep fusion of local details and global structure. Finally, an image quality prediction score is obtained through regression head mapping.
[0019] In this embodiment of the invention, a referenceless quality assessment model (which may be called SIMIQA) is constructed, such as... Figure 1 As shown, it includes a CNN branch network, a ViT branch network, a cross-attention fusion module, and a multilayer perceptron (MLP); the ViT branch network includes a ViT backbone network and a multi-head attention module; the CNN branch network includes a CNN encoder and a feature conversion unit (FCU).
[0020] In this embodiment of the invention, the training process of the no-reference quality assessment model includes the following steps:
[0021] S1. Construct a fluorescence image dataset with quality gradients; where each fluorescence image is labeled with a quality label.
[0022] Specifically, step S1 includes:
[0023] S11. Based on the original acquired data under the same field of view, a set of reconstructed images is generated by processing the data under various signal-to-noise ratio conditions and different reconstruction algorithms.
[0024] S12. Experts evaluate the quality of any two fluorescence images in the image set in pairs, and map each comparison result to an Elo battle; then, based on the Elo algorithm, the Elo rating of each fluorescence image is dynamically updated.
[0025] S13. Normalize the Elo rating scores of all fluorescence images to obtain the corresponding quality scores, and use these as quality labels.
[0026] S2. Input the fluorescence image into the ViT branch network to obtain enhanced global features.
[0027] Specifically, given a single fluorescence image Where H×W is the spatial resolution and C is the number of channels. The fluorescence image x is input into the ViT branch network to obtain enhanced global features, including:
[0028] S21. Input the fluorescence image x into the ViT backbone network to extract its global morphological structure and contextual semantic information, thereby obtaining basic global features, including:
[0029] The fluorescence image x is uniformly divided into N non-overlapping image patches using Patch Partition; where... P is the image patch size.
[0030] Each image patch is mapped to a d-dimensional feature space through a linear projection layer, generating a corresponding feature vector, which is then arranged in sequence to form a feature sequence.
[0031] Add positional encoding and learnable class embedding vectors (Class Tokens) to the feature sequences. To form the initial input sequence ;
[0032] The initial input sequence is fed into the Transformer Encoder, where it is processed through a multi-layered stacked multi-head self-attention mechanism (MSA) and a feedforward network (FFN) to effectively aggregate long-range contextual information, ultimately outputting basic global features with global representation capabilities. .
[0033] S22. To enhance the sensitivity of basic global features to quality perception and to align feature distributions before fusing with local features, this invention introduces an independent multi-head attention block (MSA-Block) after the ViT backbone network. This module uses the basic global features output by the ViT backbone network. As input, the dependencies within the basic global features are further explored through a self-attention mechanism to obtain enhanced global features with stronger semantic discriminative power.
[0034] The processing procedure of the multi-head attention module is represented as follows:
[0035] Calculate input Then calculate the query, key, and value:
[0036] ,
[0037] In the formula, Represents the basic global features, LN() represents normalization, Z represents the input; Q, K, and V represent the query matrix, key matrix, and value matrix, respectively; W Q W K W V This represents the weight matrix.
[0038] To perform multi-head operations, Q, K, and V are divided into m independent attention heads. The enhanced features are then computed using a multi-head attention mechanism.
[0039] ,
[0040] ,
[0041] In the formula, head i This represents the i=1,2,…,m attention head, where m represents the number of attention heads; Concat() represents the concatenation operation, T v This indicates the enhancement of global features. Qi, Ki, and Vi represent the query matrix, key matrix, and value matrix of the i-th attention head, respectively.
[0042] S3. Input the fluorescence image into the CNN branch network to obtain local features.
[0043] Specifically, CNN branches primarily utilize the translation invariance and local inductive bias of convolution operations to focus on extracting high-frequency degradation features in images, such as Poisson noise and edge blurring. The specific processing steps include:
[0044] S31. Input the fluorescence image into the CNN encoder, and output the deep feature map. h×w represents the resolution, and c' represents the number of channels.
[0045] Because the local spatial representation of CNNs and the global sequence representation of ViTs differ in data organization dimensions, feature fusion cannot be performed directly. This paper uses channel projection and spatial serialization operations to map the three-dimensional feature tensor to a unified vector space, thereby eliminating the representational barriers between heterogeneous branches.
[0046] S32. Channel Projection: The number of channels in the deep feature map is mapped from c' to the embedding dimension of ViT using a 1×1 convolution, resulting in the mapped features. , can be represented as:
[0047] ,
[0048] S33. Spatial Serialization: Flatten the mapped features in terms of spatial dimensions to obtain local features. , can be represented as:
[0049] .
[0050] To ensure the number of CNN tokens Consistent with the number of ViT tokens N, the network constraint CNN's total downsampling rate is equal to the ViT's patch size. At this point, It integrates rich local details extracted by CNN branches and uses them as input representations for key and value vectors to participate in subsequent cross-attention calculations.
[0051] S4. Enhance the global and local features by inputting them into the cross-attention fusion module to obtain the comprehensive features; extract the class embedding vector from the comprehensive features. Then input it into the MLP to obtain the predicted quality score.
[0052] Fluorescence microscopy images have relatively limited semantic information but rich repetitive structures, and their local quality is often non-uniform. Relying solely on local texture analysis may make it difficult to stably assess the degree of image degradation; while relying solely on global semantics may overlook fine-grained artifacts and noise. Therefore, this invention complementarily fuses the local features output by the CNN branch network with the enhanced global features output by the ViT branch network to improve the ability to characterize subtle quality differences.
[0053] This invention, through a cross-attention fusion module, leverages enhanced global structural information to achieve precise guidance in evaluating local details. Specifically, it enhances global features. As a query, with local features As keys and values, the processing includes:
[0054] ,
[0055] ,
[0056] In the formula, T v T represents enhanced global features. c Representing local features, Qca, Kca, and Vca represent the query matrix, key matrix, and value matrix, respectively; Wca Q Wca K Wca V d represents the weight matrix; k T represents the vector dimension. fused It represents comprehensive characteristics.
[0057] Through the above process, the global features extracted by the ViT branch network, while maintaining a global perspective, successfully aggregate the corresponding local high-frequency details, forming a complementary comprehensive representation.
[0058] S5. Calculate the joint loss function based on the comprehensive features, and perform backpropagation to optimize the model parameters until the model converges; introduce a comparison ranking loss into the joint loss function to keep the relative quality distance between image pairs in the feature space consistent with the difference between their true quality labels.
[0059] Regarding supervised learning strategies, addressing the issue that traditional regression losses are insufficient to effectively constrain the feature space distribution, this paper innovatively introduces ComparisonOrder Loss (COL) while maintaining the mean squared error (MSE) to ensure numerical prediction accuracy. This loss function aims to explicitly constrain the topological structure of the feature space, forcing the distances between samples in the feature space to remain consistent with the corresponding quality level differences in the label space within the same training batch. Through this structured "push-pull" mechanism, the model can more sensitively capture fine-grained quality differences while enhancing the monotonicity and robustness of the prediction results.
[0060] Specifically, compare ranking loss Represented as:
[0061] ,
[0062] In the formula, n represents the number of samples in a training batch, S(i,j) represents the feature similarity between the i-th sample and the j-th sample, and τ represents the temperature coefficient, used to control the smoothness of the similarity distribution and adjust the model's attention to easy and difficult samples. d(y i ,y j ) represents the true quality label y of the i-th sample. i Compared with the true quality label y of the j-th sample j The distance between; This indicates a dynamic sorting indicator function.
[0063] To simultaneously account for differences in local details and global structure, this invention defines the feature similarity S(i,j) between two samples as the sum of the CNN feature similarity and the ViT feature similarity, calculated as follows:
[0064] ,
[0065] In the formula, This represents the local features of the i-th sample. Let represent the enhanced global feature of the i-th sample, and sim() represents the cosine similarity calculation.
[0066] d(y i ,y j This is used to measure the difference in the true quality levels of two images. It is calculated as the difference in the absolute values of the labels, as follows:
[0067] ,
[0068] The dynamic sorting indicator function is the core of implementing sorting constraints, and is expressed as follows:
[0069] ,
[0070] As can be seen from the above, the dynamic sorting indicator function The condition for a value of 1 is: the quality difference d(y) between sample k and the reference sample i. i ,y k ) is strictly greater than the quality difference d(y) between samples j and i. i ,y j ).
[0071] Specifically, during unsupervised / zero-shot denoising or deconvolution training, the decrease in the model loss function does not always accompany a synchronous improvement in visual quality. By periodically sampling the model output from different iterations during training and inputting it into the quality score curve evaluated by this method, it can be used to effectively identify the moment when the quality reaches its peak, enabling early stopping decisions based on visual performance. Simultaneously, this method can automatically trigger alarms when high-frequency artifacts, noise overfitting, or other conditions leading to quality degradation are detected.
[0072] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A reference-free quality assessment method for super-resolution fluorescence images, characterized in that, A no-reference quality assessment model is constructed and trained, and the trained no-reference quality assessment model is used to evaluate the quality of super-resolution fluorescence images; the no-reference quality assessment model includes a CNN branch network, a ViT branch network, and a cross-attention fusion module. The training process of the no-reference quality assessment model includes the following steps: S1. Construct a fluorescence image dataset with quality gradients; where each fluorescence image is labeled with a quality label; S2. The fluorescence image is input into the ViT branch network to obtain enhanced global features; the ViT branch network includes a ViT backbone network and a multi-head attention module; S3. Input the fluorescence image into a CNN branch network to obtain local features; S4. Input the enhanced global and local features into the cross-attention fusion module to obtain the comprehensive features; extract the class embedding vector from the comprehensive features and input it into the MLP to obtain the prediction quality score; S5. Calculate the joint loss function based on the comprehensive features, and perform backpropagation to optimize the model parameters until the model converges; introduce a comparison ranking loss into the joint loss function to keep the relative quality distance between image pairs in the feature space consistent with the difference between their true quality labels.
2. The reference-free quality assessment method for super-resolution fluorescence images according to claim 1, characterized in that, Step S2 inputs the fluorescence image into the ViT branch network to obtain enhanced global features, including: S21. Input the fluorescence image into the ViT backbone network and extract basic global features, including: The fluorescence image is uniformly divided into N non-overlapping image blocks; Each image patch is mapped to a d-dimensional feature space through a linear projection layer, generating a corresponding feature vector, which is then arranged in sequence to form a feature sequence. Add positional encoding and learnable class embedding vectors to the feature sequence to form the initial input sequence; The initial feature sequence is fed into the Transformer Encoder to obtain the basic global features; S22. The basic global features are processed through a multi-head attention module to obtain enhanced global features, represented as: , , , In the formula, Represents the basic global features, LN() represents normalization, Z represents the input; Q, K, and V represent the query matrix, key matrix, and value matrix, respectively; W Q W K W V W O Represents the weight matrix; head i This represents the i=1,2,…,m attention head, where m represents the number of attention heads; Concat() represents the concatenation operation, T v This indicates an enhancement of global features.
3. The reference-free quality assessment method for super-resolution fluorescence images according to claim 1, characterized in that, Step S3 inputs the fluorescence image into the CNN branch network to obtain local features, including: S31. Input the fluorescence image into the CNN encoder and output the deep feature map; S32. Map the number of channels in the deep feature map to d dimensions using 1×1 convolution to obtain the mapped features; S33. Flatten the mapped features in terms of spatial dimensions to obtain local features.
4. The reference-free quality assessment method for super-resolution fluorescence images according to claim 1, characterized in that, Step S4 specifically includes: , , In the formula, T v T represents enhanced global features. c Representing local features, Qca, Kca, and Vca represent the query matrix, key matrix, and value matrix, respectively; Wca Q Wca K Wca V d represents the weight matrix; k T represents the vector dimension. fused It represents comprehensive characteristics.
5. The reference-free quality assessment method for super-resolution fluorescence images according to claim 1, characterized in that, Comparison and ranking loss Represented as: , , , , In the formula, n represents the number of samples in a training batch, and S(i,j) represents the feature similarity between the i-th sample and the j-th sample. This represents the local features of the i-th sample. d(y) represents the enhanced global feature of the i-th sample, sim() represents the cosine similarity calculation; τ represents the temperature coefficient, d(y) i ,y j ) represents the true quality label y of the i-th sample. i Compared with the true quality label y of the j-th sample j The distance between, This indicates a dynamic sorting indicator function.
6. The reference-free quality assessment method for super-resolution fluorescence images according to claim 1, characterized in that, Step S1 includes: S11. Based on the original acquired data under the same field of view, a set of reconstructed images is generated by processing the data under various signal-to-noise ratio conditions and different reconstruction algorithms. S12. Perform pairwise quality evaluation on any two fluorescence images in the reconstructed image set, and map each comparison result to an Elo battle; then, based on the Elo algorithm, dynamically update the Elo rating score of each fluorescence image. S13. Normalize the Elo rating scores of all fluorescence images to obtain the corresponding quality scores, and use these as quality labels.
Citation Information
Patent Citations
Image quality evaluation method based on multi-scale region self-attention fusion under meta-learning framework
CN118469930A
Non-reference image quality evaluation method based on information transfer VIT
CN119672001A