A full-reference image quality evaluation method based on Bhattacharyya distance and Tanimoto similarity
Patent Information
- Application Number
- CN202511030473.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-07-25
AI Technical Summary
对于底层的基础视觉特征,现有方法缺乏能够敏感捕捉其细微统计分布变化的、鲁棒的度量工具,例如,对由噪声或轻微模糊导致的像素级统计特性变化的感知不够精确
1、该方法通过将基础视觉层特征转化为概率分布并利用Bhattacharyya距离,该方法能敏感捕捉图像在纹理、边缘等底层特征的细微退化。其归一化概率分布处理增强了特征差异的鲁棒性,尤其适用于量化压缩、噪声等失真导致的像素级差异,提升评价结果的客观性。
Smart Images

Figure CN120852392B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image quality assessment, and more specifically to a full-reference image quality assessment method based on Bhattacharyya distance and Tanimoto similarity. Background Technology
[0002] Image quality assessment (IQA) is a key technology in the fields of image processing and computer vision, playing a central role, especially in scenarios such as video coding optimization, image communication system monitoring, medical image analysis, and image enhancement algorithm evaluation.
[0003] Traditional full-reference IQA methods primarily rely on manually designed features and mathematical models, such as those based on structural similarity (SSIM) or information fidelity criteria (VIF). While these methods perform reasonably well under specific distortion types, their core drawback lies in their limited feature representation capabilities and difficulty in generalization. Manually designed features often fail to fully capture the complex and ever-changing patterns of real-world distortion. When faced with distortion types not covered by the training data or complex scenarios, their prediction accuracy and robustness significantly decrease, limiting their application value in real-world, dynamic environments.
[0004] In recent years, deep learning-based full-reference IQA methods have demonstrated great potential, automatically learning image features using deep neural networks. However, existing deep learning methods generally suffer from a key limitation: an overemphasis on improving feature extraction capabilities while neglecting the compatibility of measurement methods with human visual perception mechanisms. Most methods simply use Euclidean distance, cosine similarity, or direct input into regression networks to predict quality. This approach fails to fully exploit the inherent statistical distribution characteristics of features and fails to effectively simulate the differentiated attention and processing mechanisms of the human visual system to information at different levels of an image, resulting in a perceptible discrepancy between model predictions and subjective human ratings.
[0005] Further analysis of existing measurement methods reveals significant shortcomings in the effective quantification of multi-level features. For low-level, fundamental visual features, existing methods lack robust measurement tools capable of sensitively capturing subtle changes in their statistical distribution; for example, they are not precise enough in perceiving pixel-level statistical characteristic changes caused by noise or slight blurring. For high-level, deep semantic features, while abstract information can be extracted, commonly used similarity measures are not intuitive enough in measuring the "overlap" or "repulsion" relationships between features, and lack effective mechanisms to focus on the key image regions most relevant to HVS, failing to effectively fuse and weight local semantic consistency with global structural information. These shortcomings at the measurement level limit the model's ability to comprehensively and accurately characterize image quality degradation.
[0006] Therefore, how to design a full-reference image quality assessment method based on Bhattacharyya distance and Tanimoto similarity that can simultaneously and accurately quantify the differences in the distribution of basic visual features of images and effectively integrate human visual perception characteristics to evaluate the similarity of deep semantic features is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0007] In view of this, the present invention provides a full-reference image quality assessment method based on Bhattacharyya distance and Tanimoto similarity. By extracting and differentially measuring the low-level visual features and high-level semantic features of the image in a hierarchical manner, and adaptively fusing the two measurement results, a more accurate, robust, and objective quality score of distorted images that is more in line with human subjective visual perception is finally achieved, thereby improving the effectiveness and reliability of the evaluation model in practical industrial applications.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] A full-reference image quality assessment method based on Bhattacharyya distance and Tanimoto similarity is proposed. This method acquires the distorted image to be evaluated and its corresponding reference image, inputs them into a pre-trained image quality assessment model for multi-module processing, and includes the following steps: S1. Based on the feature extraction module, extract the basic visual hierarchical features and deep semantic hierarchical features of the distorted image and the reference image respectively; S2. Based on the feature distribution metric module, calculate the Bhattacharyya distance of the basic visual layer features; S3. Based on the similarity calculation module, calculate the Tanimoto similarity of deep semantic layer features; S4. Based on the score prediction module, the Bhattacharyya distance and Tanimoto similarity are fused to output an objective quality score for the distorted image.
[0010] Preferably, S1 includes: using a ResNet network to extract features of the distorted image and the reference image at different network layers; defining the output of the shallow network as basic visual layer features and the output of the deep network as deep semantic layer features.
[0011] Preferably, S2 includes: S21. Perform probabilistic normalization on the basic visual layer feature maps of the distorted image and the reference image respectively to obtain the probability distribution of the normalized feature maps of the distorted image and the reference image. , ; S22. Quantize the difference between the distorted image and the reference image using the Bhattacharyya distance:
[0012] in, This represents a very small positive number to prevent numerical overflow.
[0013] Preferably, S21 includes:
[0014]
[0015] in, , These represent the positions of the distorted image and the reference image, respectively. eigenvalues, , These represent the positions of the distorted image and the reference image, respectively. eigenvalues.
[0016] Preferably, the probability distribution transformation ensures and Furthermore, position x covers all pixel coordinates of the feature map.
[0017] Preferably, S3 includes: S31. Combine the deep semantic features of the distorted image and the reference image. , Attention mechanisms are processed separately to obtain attention scores. , ; S32. Select the features corresponding to the top k highest attention scores respectively, and combine the Top-k selection mechanism and Function to obtain visual focus feature vector ; S33. Calculate global and local Tanimoto similarity. , :
[0018]
[0019] S34. Perform dynamic weighted fusion to obtain similarity calculation results. :
[0020] in, , This represents the weighting coefficient.
[0021] Preferably, S31 includes:
[0022]
[0023]
[0024]
[0025] in, , Represents the query vector. , Represents the key vector. , Represents a value vector. This represents two-dimensional convolution processing. This represents the softmax function.
[0026] Preferably, S32 includes:
[0027]
[0028]
[0029] in, This represents the k highest values selected. Indicates the corresponding index position. Represents a value vector The feature result values extracted from it.
[0030] Preferably, S32 further includes:
[0031]
[0032]
[0033] in, This represents the k highest values selected. Indicates the corresponding index position. Represents a value vector The feature result values extracted from it.
[0034] Preferably, in step S4, the objective quality score is... Represented as:
[0035] in, This represents the trainable parameter weights.
[0036] As can be seen from the above technical solution, compared with the prior art, the technical solution of the present invention has the following beneficial effects: 1. This method transforms basic visual layer features into probability distributions and utilizes Bhattacharyya distance. This approach can sensitively capture subtle degradations in low-level features such as texture and edges. Its normalized probability distribution processing enhances the robustness of feature differences, making it particularly suitable for pixel-level differences caused by distortions such as quantization compression and noise, thus improving the objectivity of evaluation results.
[0037] 2. For deep semantic features, an attention mechanism is introduced to filter key regions and calculate local and global Tanimoto similarities. This simulates the human visual system's focus on salient regions, and balances the similarity between overall structure and local semantics through dynamic weighted fusion, making the evaluation results more consistent with subjective perception.
[0038] 3. A trainable parameter adaptive fusion method is used to fuse low-level feature distance and high-level semantic similarity to form a unified objective score. This hierarchical fusion mechanism can simultaneously take into account the multi-dimensional impact of different types of distortion (such as blur and color distortion) on image quality, significantly enhancing the model's adaptability to unknown distortion scenarios. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0040] Figure 1 A flowchart of a full-reference image quality assessment method based on Bhattacharyya distance and Tanimoto similarity provided in this embodiment of the invention; Figure 2 This is a schematic diagram of the image quality evaluation model structure provided in an embodiment of the present invention; Figure 3 A flowchart of a full-reference image quality evaluation method for remote medical imaging diagnosis provided in this embodiment of the invention. Detailed Implementation
[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0042] like Figure 1 and Figure 2 As shown, this embodiment provides a full-reference image quality assessment method based on Bhattacharyya distance and Tanimoto similarity. The method involves acquiring the distorted image to be evaluated and its corresponding reference image, inputting them into a pre-trained image quality assessment model for multi-module processing, including the following steps: S1. Based on the feature extraction module, extract the basic visual hierarchical features and deep semantic hierarchical features of the distorted image and the reference image respectively; S2. Based on the feature distribution metric module, calculate the Bhattacharyya distance of the basic visual layer features; S3. Based on the similarity calculation module, calculate the Tanimoto similarity of deep semantic layer features; S4. Based on the score prediction module, the Bhattacharyya distance and Tanimoto similarity are fused to output an objective quality score for the distorted image.
[0043] This method processes image features in layers. At the basic visual layer, it uses the Bhattacharyya distance to accurately quantify the degradation differences of details such as texture and edges. At the deep semantic layer, it combines an attention mechanism to select key regions and calculates Tanimoto similarity to match the human eye's perception of structure and semantics. Finally, it dynamically fuses the measurement results of these two layers through trainable parameters, thereby significantly improving the objectivity, perceptual consistency and generalization ability of full-reference image quality evaluation for multiple distortion types.
[0044] The following provides a further detailed explanation of each step and related technical feature in the above method: The aforementioned distorted images refer to the images to be evaluated whose quality has been degraded due to processing such as compression, noise, blurring, and color distortion. The reference image is its original, undistorted, high-quality version, and the two must maintain strict spatial alignment, content consistency, and temporal synchronization. This pairing data can be obtained from publicly available image quality datasets or real-world application scenarios, ensuring that the two images input to the model differ only in the degree of distortion, thus providing a foundation for accurate measurement of subsequent feature differences.
[0045] In this embodiment S1, based on the feature extraction module, basic visual layer features and deep semantic layer features of the distorted image and the reference image are extracted respectively; including: using the ResNet network to extract features of the distorted image and the reference image at different network layers respectively; defining the output of the shallow network as basic visual layer features and the output of the deep network as deep semantic layer features.
[0046] Specifically, the shallow network refers to the output features of the first three stages of ResNet, used to capture basic visual features; the deep network refers to the output features of the fourth stage of ResNet, used to extract deep semantic features. This limitation avoids ambiguity in the hierarchical division; and the distorted image and the reference image use the same ResNet network parameters for feature extraction, ensuring consistency in the feature space.
[0047] This step uses ResNet to explicitly distinguish between shallow and deep features, solving the problems of limited feature representation capabilities of traditional manual methods and ambiguous hierarchical division in deep learning models, and providing a precise physical basis for difference measurement.
[0048] In this embodiment S2, based on the feature distribution metric module, the Bhattacharyya distance of the basic visual layer features is calculated; including: S21. Perform probabilistic normalization on the basic visual layer feature maps of the distorted image and the reference image respectively to obtain the probability distribution of the normalized feature maps of the distorted image and the reference image. , Probabilistic normalization transforms feature values into discrete probability distributions, making the Bhattacharyya distance quantifiable as the overall statistical distribution difference of the feature map, rather than just the difference at local points. S22. The difference between the distorted image and the reference image is quantified by the Bhattacharyya distance. Its geometric meaning is to measure the similarity between two probability distributions. The larger the value, the more significant the difference in feature distribution.
[0049] in, This represents a very small positive number to prevent numerical overflow.
[0050] Furthermore, S21 includes:
[0051]
[0052] in, , These represent the positions of the distorted image and the reference image, respectively. eigenvalues, , These represent the positions of the distorted image and the reference image, respectively. eigenvalues.
[0053] Furthermore, the probability distribution transformation ensures and Furthermore, position x covers all pixel coordinates of the feature map, ensuring the completeness of the probability distribution transformation.
[0054] By normalizing the feature map probability, local features are transformed into a global statistical distribution, and the Bhattacharyya distance is introduced to replace traditional metrics such as MSE / SSIM. Compared with existing methods, it has a stronger ability to capture subtle changes in statistical distribution caused by noise and blur; and avoids local biases at the pixel level, focusing on overall distribution degradation.
[0055] In this embodiment, S3, based on the similarity calculation module, the Tanimoto similarity of deep semantic layer features is calculated, including: S31. Combine the deep semantic features of the distorted image and the reference image. , Attention mechanisms are processed separately to obtain attention scores. , The attention score here is used to locate key regions in the image that are sensitive to human vision (such as high-frequency textures and prominent objects), and then filter out the local semantic features that represent the most important local features of the human visual system. S32. Select the features corresponding to the top k highest attention scores respectively, and combine the Top-k selection mechanism and Function to obtain visual focus feature vector It is used to measure set similarity (which can be regarded as the degree of overlap of feature space), and is more in line with the needs of semantic feature evaluation than cosine similarity. S33. Calculate global and local Tanimoto similarity. , :
[0056]
[0057] S34. Perform dynamic weighted fusion to obtain similarity calculation results. This is used to adaptively adjust the contribution ratio of global and local similarity based on image content.
[0058] in, , This represents the weighting coefficient.
[0059] Furthermore, S31 includes:
[0060]
[0061]
[0062]
[0063] in, , Represents the query vector. , Represents the key vector. , Represents a value vector. This represents two-dimensional convolution processing. This represents the softmax function.
[0064] Furthermore, S32 includes:
[0065]
[0066]
[0067] in, This represents the k highest values selected. Indicates the corresponding index position. Represents a value vector The feature result values extracted from it.
[0068] Furthermore, S32 also includes:
[0069]
[0070]
[0071] in, This represents the k highest values selected. Indicates the corresponding index position. Represents a value vector The feature result values extracted from it.
[0072] This step integrates the attention mechanism (Top-k selection) with Tanimoto similarity, breaking through the limitations of existing deep learning. The attention mechanism simulates the human eye's focusing characteristics on key regions, overcoming the shortcomings of traditional global feature averaging. Combined with Tanimoto similarity, which is designed specifically for set similarity, it is more suitable for evaluating the spatial overlap relationship of semantic features than cosine similarity.
[0073] In this embodiment, S4, based on the score prediction module, the Bhattacharyya distance and Tanimoto similarity are fused to output an objective quality score for the distorted image; the objective quality score... Represented as:
[0074] in, This represents the trainable parameter weights.
[0075] In specific model training, the objective quality score can be automatically updated through the backpropagation algorithm. It fits the subjective rating of the human eye; its value range is limited to [0,1]. The higher the value, the more the quality assessment depends on the differences of the underlying features. Conversely, it depends more on the semantic similarity of the higher level. It solves the problem that fixed weights are difficult to adapt to multiple distortion types in existing methods by adaptively fusing the underlying distance and the high-level similarity through trainable parameters.
[0076] The following section provides a complete explanation of the implementation process of the full-reference image quality assessment method in this implementation, using a remote medical imaging diagnostic scenario as an example.
[0077] In this scenario, a top-tier hospital needs to assess the quality degradation of MRI brain images transmitted after JPEG compression to determine whether they meet the requirements for remote diagnosis. The hospital's radiology department acquires a patient's original, high-quality brain MRI image (reference image). To improve transmission efficiency, a distorted image is generated using JPEG compression with CRF=30. After strict alignment, the two images are input into the evaluation system. The resulting texture blurring and loss of semantic information about key anatomical structures due to compression must be quantified to ensure the reliability of the assisted diagnosis.
[0078] like Figure 3 As shown, its specific implementation process includes: 1) Feature layer extraction; In this step, a pre-trained ResNet-50 model is used to process the image; the reference image and the compressed image are input into the network respectively: the output of the first three stages serves as the basic visual layer features, capturing details such as brain tissue edges and gray and white matter textures; the output of the fourth stage serves as the deep semantic layer features, extracting representations of higher-level anatomical structures such as the hippocampus and ventricles.
[0079] 2) Quantification of visual feature distribution; For the basic visual layer features, the system performs probabilistic normalization: each feature map is treated as a discrete distribution, the sum of the feature values at all pixel locations is calculated, and then the feature values at each location are divided by the sum. For example, the peak probability distribution of a certain texture region in the reference image is concentrated in the range of 0.02-0.03, while the compressed image shows a more dispersed distribution in the same region (the peak value drops to 0.01-0.02). Using the Bhattacharyya distance formula, DB = 0.37 (where ε = 1e-8), significantly higher than the 0.05 in the uncompressed image, reflecting a decrease in edge sharpness.
[0080] 3) Semantic feature similarity calculation; In the deep semantic feature layer, a self-attention mechanism is introduced to locate key regions. First, the output of the deep semantic features is convolved to generate Q / K / V vectors, which are then processed by softmax to obtain the attention score map. The high-attention zones in the reference image accurately cover the hippocampus, while the score for the same region in the compressed image drops to 0.73. Visual focus feature vectors are extracted through Top-k selection (k = 10% of total pixels) and the Gather function. ; Calculate the global and local Tanimoto similarities separately: =0.82, =0.76. Finally, the dynamic weighting (ω1=0.4, ω2=0.6) yielded TC=0.78, indicating a more significant loss of key anatomical structural information.
[0081] 4) Quality score fusion output; The score prediction module integrates two layers of metric results: setting initial trainable parameters. =0.4, converged to after backpropagation optimization. =0.35 (higher semantic weights). Final score = 0.35 0.37 + 0.65 0.78 = 0.73 (73 points out of 100), which is below the diagnostic threshold of 80 points.
[0082] The results indicate that the compression parameters resulted in excessive loss of semantic information in key regions such as the hippocampus. The system automatically triggered a warning and suggested reducing the compression intensity.
[0083] The method was validated for effectiveness in medical imaging scenarios. The Bhattacharyya distance accurately quantified texture degradation (such as blurred ventricular edges), while the attention-weighted Tanimoto similarity captured semantic-level structural loss (such as hippocampal contour distortion). The dynamic fusion mechanism made the scoring more aligned with doctors' subjective evaluations, and the hospital adjusted the transmitted CRF value to 25 accordingly. Improved to 0.86, successfully balancing efficiency and diagnostic needs.
[0084] The full-reference image quality assessment method based on Bhattacharyya distance and Tanimoto similarity provided in this embodiment significantly improves the objectivity, perceptual consistency, and generalization ability to various distortion types of image quality assessment by hierarchically processing image features, combining an attention mechanism, and dynamically fusing measurement results from different levels. This method not only performs excellently in general image quality assessment but also demonstrates strong application potential and value in specific scenarios such as remote diagnosis of medical images, providing powerful technical support for accurate image quality assessment.
[0085] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.
[0086] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for evaluating the quality of a full-reference image based on Bhattacharyya distance and Tanimoto similarity, characterized in that, Obtain the distorted image to be evaluated and its corresponding reference image, and input them into a pre-trained image quality assessment model for multi-module processing, including the following steps: S1. Based on the feature extraction module, extract the basic visual hierarchical features and deep semantic hierarchical features of the distorted image and the reference image respectively; S2. Based on the feature distribution metric module, calculate the Bhattacharyya distance of the basic visual layer features; including: S21. Perform probabilistic normalization on the basic visual layer feature maps of the distorted image and the reference image respectively to obtain the probability distribution of the normalized feature maps of the distorted image and the reference image. , ; S22. Quantize the difference between the distorted image and the reference image using the Bhattacharyya distance: in, Represents a very small positive number to prevent numerical overflow; S3. Based on the similarity calculation module, calculate the Tanimoto similarity of deep semantic layer features; including: S31. Combine the deep semantic features of the distorted image and the reference image. , Attention mechanisms are processed separately to obtain attention scores. , ; S32. Select the features corresponding to the top k highest attention scores respectively, and combine the Top-k selection mechanism and Function to obtain visual focus feature vector ; S33. Calculate global and local Tanimoto similarity. , : S34. Perform dynamic weighted fusion to obtain similarity calculation results. : in, , Indicates the weighting coefficient; S4. Based on the score prediction module, the Bhattacharyya distance and Tanimoto similarity are fused to output an objective quality score for the distorted image.
2. The full-reference image quality assessment method based on Bhattacharyya distance and Tanimoto similarity as described in claim 1, characterized in that, S1 includes: using a ResNet network to extract features of the distorted image and the reference image at different network layers; defining the output of the shallow network as basic visual layer features and the output of the deep network as deep semantic layer features.
3. The full-reference image quality assessment method based on Bhattacharyya distance and Tanimoto similarity as described in claim 1, characterized in that, S21 includes: in, , These represent the positions of the distorted image and the reference image, respectively. eigenvalues, , These represent the positions of the distorted image and the reference image, respectively. eigenvalues.
4. The full-reference image quality assessment method based on Bhattacharyya distance and Tanimoto similarity as described in claim 3, characterized in that, Probability distribution transformation ensures and Furthermore, position x covers all pixel coordinates of the feature map.
5. The full-reference image quality assessment method based on Bhattacharyya distance and Tanimoto similarity as described in claim 1, characterized in that, S31 includes: in, , Represents the query vector. , Represents the key vector. , Represents a value vector. This represents two-dimensional convolution processing. This represents the softmax function.
6. The full-reference image quality assessment method based on Bhattacharyya distance and Tanimoto similarity as described in claim 1, characterized in that, S32 includes: in, This represents the k highest values selected. Indicates the corresponding index position. Represents a value vector The feature result values extracted from it.
7. The full-reference image quality assessment method based on Bhattacharyya distance and Tanimoto similarity as described in claim 1, characterized in that, S32 further includes: in, This represents the k highest values selected. Indicates the corresponding index position. Represents a value vector The feature result values extracted from it.
8. The full-reference image quality assessment method based on Bhattacharyya distance and Tanimoto similarity as described in claim 1, characterized in that, In S4, the objective quality score Represented as: in, This represents the trainable parameter weights.
Citation Information
Patent Citations
No-reference hyperspectral image quality evaluation method and device based on sorting learning
CN117876317A
Security-related event anomaly detection
US20250168179A1