A No-Reference Quality Assessment Method for CT Images Based on Feature Fusion
By integrating edge and contrast characteristics and depth characteristics, the problem of insufficient information in the existing CT image quality evaluation methods is solved, and more accurate reference-free quality evaluation is achieved, which is suitable for medical image diagnosis.
Patent Information
- Application Number
- CN202410860043.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-06-28
AI Technical Summary
The existing CT image quality evaluation methods only focus on a single feature, insufficient information sources, resulting in insufficient evaluation, and traditional methods are time-consuming and labor-intensive, making it difficult to apply to clinical practice.
Using a feature fusion-based method, combining edge features, contrast features and depth features, the edge and contrast features of the image are extracted through the Laplace operator and grayscale symbiosis matrix, and feature fusion is performed using attention mechanism, and deep feature extraction is performed by combining the ConvNeXt model to construct a deep learning model for reference-free evaluation.
It improves the accuracy and reliability of image quality evaluation and can more comprehensively reflect the actual quality of the image, especially edge and contrast information, which is of great significance to medical image diagnosis.
Smart Images

Figure CN118735889B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and particularly to a no-reference quality assessment method for CT images based on feature fusion. Background Art
[0002] As a relatively mature imaging technology nowadays, computed tomography plays a huge role in the medical field. CT has the characteristics of high resolution, non-destructive, fast imaging, etc., so it is widely used in the field of medical clinical diagnosis. With the widespread use of CT in clinical diagnosis, people have gradually discovered the inevitable radiation dose problem brought by CT scans. Since high-dose X-rays can cause serious radiation damage to the patient's body, greatly damage their immune system, thus easily inducing metabolic disorders and increasing the possibility of cancer, the current trend is to obtain CT images using low-dose rays. However, as the ray dose is reduced, there is more noise in the CT images.
[0003] In addition, various factors in terms of equipment, technology and operation methods, such as inappropriate scanning parameters, equipment failures, calibration inaccuracies, and patient movement during image acquisition, will all lead to the generation of low-quality images. Medical CT image quality assessment algorithms can be divided into subjective image quality assessment and objective image quality assessment. Subjective image quality assessment is for clinical professional radiologists to manually evaluate the diagnostic quality of CT images. However, it has defects such as time-consuming, laborious, and being easily affected by subjective and objective factors, and it is difficult to be applied to clinical practice. Objective image quality assessment can be divided into: full-reference image quality assessment, that is, by measuring the difference between the distorted CT image and the high-quality reference image to evaluate the quality of the distorted CT image. Semi-reference image quality assessment, that is, using part of the information of the reference image to evaluate the quality of the distorted CT image. No-reference image quality assessment, where only the distorted CT image itself is used for evaluation. Since it is difficult to obtain pairs of high / low-quality CT images in clinical practice, the no-reference image quality assessment method has become the focus of CT image quality assessment research. Summary of the Invention
[0004] Aiming at the above deficiencies in the prior art, the no-reference quality assessment method for CT images based on feature fusion provided by the present invention solves the problems that the existing methods only extract a single feature from the image, the information source is less, and the features extracted by the network are not sufficient.
[0005] To achieve the above invention purpose, the technical solution adopted by the present invention is: a no-reference quality assessment method for CT images based on feature fusion, including:
[0006] S1. Collect CT images, perform distortion simulation and add quality labels to obtain CT images with known quality scores;
[0007] S2. Extract features from CT images with known quality scores to obtain image features;
[0008] S3. Use CT images with known quality scores and image features to construct and train a deep learning model to obtain a trained model;
[0009] S4. Input the CT image whose quality evaluation is to be obtained into the trained model, and obtain the quality score of the CT image whose quality evaluation is to be obtained, thus completing the no-reference quality evaluation.
[0010] Further: The said S2 includes:
[0011] S21. Extract the edge features and contrast features of the CT image with known quality scores, and obtain the intermediate result r1 according to the edge features and contrast features;
[0012] S22. Conduct deep feature extraction on the CT image with known quality scores, and obtain the intermediate result r2 according to the deep features;
[0013] S23. Perform weighted late fusion on the intermediate result r1 and the intermediate result r2 to obtain image features.
[0014] Further: In the said S21, a feature fusion module is used to obtain the intermediate result r1, which includes:
[0015] S211. Use the Lapalce operator to extract the edge features of the CT image with known quality scores;
[0016] S212. Use the gray-level co-occurrence matrix to extract the contrast features of the CT image with known quality scores;
[0017] S213. Use a feature fusion module based on the attention mechanism to fuse the edge features and contrast features to obtain the attention fusion features;
[0018] S214. Pass the attention fusion features through a classifier to obtain the intermediate result r1.
[0019] Further: The said S213 includes:
[0020] S2131. Concatenate the edge features and contrast features to obtain the concatenated features;
[0021] S2132. Respectively extract the global attention, local attention and channel attention of the concatenated features;
[0022] S2133. Use an addition operation on the global attention, local attention and channel attention of the concatenated features to obtain the first intermediate parameter;
[0023] After using the sigmoid activation function on the first intermediate parameter, perform multiplication operations with the edge feature and the contrast feature respectively to obtain a second intermediate parameter and a third intermediate parameter;
[0024] S2135. Use an addition operation on the second intermediate parameter and the third intermediate parameter to obtain an attention fusion feature.
[0025] Furthermore: In S22, a feature fusion module using depth features is used to obtain the intermediate result r2, which includes:
[0026] S221. Perform depth feature extraction on the CT image with known quality scores through ConvNeXt to obtain four layers of depth features;
[0027] S222. Fuse the four layers of depth features through a bidirectional feature pyramid to obtain four layers of fused features;
[0028] S223. Classify the four layers of fused features through a classifier to obtain four layers of classification results;
[0029] S224. Perform weighted late fusion on the four layers of classification results to obtain the intermediate result r2.
[0030] Furthermore: In S221, the method for ConvNeXt to perform depth feature extraction on the CT image includes:
[0031] S2211. Perform two-dimensional convolution on the CT image along the depth direction to obtain a fourth intermediate parameter;
[0032] S2212. Perform layer normalization on the fourth intermediate parameter to obtain a fifth intermediate parameter;
[0033] S2213. After performing two-dimensional convolution on the fifth intermediate parameter, use the GELU activation function to obtain a sixth intermediate parameter;
[0034] S2214. Perform two-dimensional convolution, anchor point scaling, and network regularization on the sixth intermediate parameter in sequence to obtain a seventh intermediate parameter;
[0035] S2215. Concatenate the fourth intermediate parameter and the seventh intermediate parameter to obtain the depth feature of the CT image.
[0036] Furthermore: In S3, the loss function L of the deep learning model is:
[0037]
[0038] Among them, ri = {r1, r2} represents the intermediate result, bi is the weight corresponding to the intermediate result, and r represents the image feature obtained by performing weighted late fusion on the intermediate result r1 and the intermediate result r2.
[0039] The beneficial effects of the present invention are as follows:
[0040] 1. Targeted extraction of traditional features: Traditional no-reference image quality assessment methods often only focus on the overall statistical characteristics of images while ignoring local characteristics, especially edge information. In medical images, edge information and contrast information are of great significance for disease diagnosis and condition analysis. Therefore, the no-reference image quality assessment method based on edge features and contrast features is innovative and can more accurately reflect the actual quality of images;
[0041] 2. Fusing traditional features with deep features can enhance the model's capabilities;
[0042] 3. Application of feature fusion: Most no-reference image quality assessment methods often only focus on a single feature, such as edges, noise, etc., while ignoring the impact of other features on image quality. By fusing multiple features together, the quality of images can be evaluated more comprehensively, improving the accuracy and reliability of the assessment;
[0043] 4. When fusing traditional features, a feature fusion method based on the attention mechanism is used, enabling the model to focus more on the important parts of the features;
[0044] 5. The loss function takes into account both intermediate results and final results, while reducing the losses of the two branch networks BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a flowchart of a no-reference quality assessment method for CT images based on feature fusion.
[0046] Figure 2 is a structural diagram of a feature fusion module for traditional features.
[0047] Figure 3 is a sub-module structural diagram of the feature fusion module for traditional features.
[0048] Figure 4 is a structural diagram of a feature fusion module for deep features.
[0049] Figure 5 is a sub-module structural diagram of the feature fusion module for deep features. DETAILED DESCRIPTION OF THE INVENTION
[0050] The following describes the specific implementation manners of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation manners. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0051] As Figure 1 shown, in an embodiment of the present invention, a no-reference quality assessment method for CT images based on feature fusion is proposed, including:
[0052] S1. Collect CT images, perform distortion simulation and add quality labels to obtain CT images with known quality scores;
[0053] In this embodiment, the scores obtained using existing full-reference image quality metrics and the scores given by doctors are used as quality labels;
[0054] S2. Extract features from the CT images with known quality scores to obtain image features;
[0055] S3. Use the CT images with known quality scores and the image features to construct and train a deep learning model to obtain a trained model;
[0056] S4. Input the CT image to be quality-assessed into the trained model to obtain the quality score of the CT image to be quality-assessed, and complete the no-reference quality assessment.
[0057] The said S2 includes:
[0058] S21. Extract the edge features and contrast features of the CT images with known quality scores, and obtain the intermediate result r1 according to the edge features and contrast features;
[0059] S22. Perform deep feature extraction on the CT images with known quality scores, and obtain the intermediate result r2 according to the deep features;
[0060] S23. Perform weighted late fusion on the intermediate result r1 and the intermediate result r2 to obtain image features.
[0061] In the said S21, a feature fusion module is used to obtain the intermediate result r1, which includes:
[0062] S211. Use the Lapalce operator to extract the edge features of the CT images with known quality scores;
[0063] S212. Use the gray-level co-occurrence matrix to extract the contrast features of the CT images with known quality scores;
[0064] S213. Use a feature fusion module based on the attention mechanism to fuse the edge features and contrast features to obtain the attention fusion features;
[0065] S214. Pass the attention fusion features through a classifier to obtain the intermediate result r1.
[0066] There are two important parameters for CT image quality, namely spatial resolution and density resolution. Spatial resolution refers to the ability to distinguish fine structures in an image, that is, the imaging ability of small-sized lesions or structures or the clarity of the image. Density resolution, also known as low-contrast resolution, refers to the ability to distinguish the smallest density differences between tissues. The Laplace operator is used to extract the edge features of the image to reflect the spatial resolution, and the contrast feature of the gray-level co-occurrence matrix is used to reflect the density resolution. After fusing the two features, the image quality score is obtained.
[0067] Generally speaking, the spatial resolution of a CT image is the ability to clearly see the specific area of a lesion. Therefore, the edge features of the image are used to reflect the small-sized structures and specific positions that a CT image can identify. Currently, the most commonly used method for detecting the edge features of an image is to use the Canny operator, which is a multi-level edge detection algorithm and is considered to be one of the best edge detection algorithms. The criteria it follows for evaluating an edge detection algorithm include fewer errors, more accurate positioning, and the detected edges being as close as possible to the true edges of the image. However, due to the strong anti-noise ability of the Canny operator, it is not suitable for image quality assessment. Compared with the edge features of ordinary images, the features required in this paper are the edge features under the influence of noise. Therefore, a strong anti-noise ability of the operator is not needed;
[0068] The Laplace operator is a second-order derivative operator. Since it is a second-order derivative, the Laplace operator is more sensitive to noise and is more suitable for this application scenario. Therefore, the Laplace operator is selected to extract the edge features of the image.
[0069] The gray-level co-occurrence matrix (GLCM) calculates its co-occurrence matrix by computing a gray-scale image, and then obtains some eigenvalue of the matrix by calculating this co-occurrence matrix to represent certain texture features of the image respectively. The gray-level co-occurrence matrix can reflect comprehensive information about the gray levels of the image in terms of direction, adjacent interval, change amplitude, etc. It is the basis for analyzing the local patterns of the image and their arrangement rules. The contrast feature of the image calculated based on the gray-level co-occurrence matrix shows the brightness and darkness difference of the image. The density resolution of a CT image refers to whether different lesions can be distinguished by the contrast difference. Therefore, the contrast feature of the image is selected to reflect the density resolution of the CT image.
[0070] As Figure 2 shown, the S213 includes:
[0071] S2131. Concatenate the edge feature and the contrast feature to obtain a concatenated feature;
[0072] S2132. Respectively extract the global attention, local attention, and channel attention of the concatenated feature;
[0073] S2133. Perform an addition operation on the global attention, local attention, and channel attention of the concatenated features to obtain a first intermediate parameter;
[0074] S2134. After applying the sigmoid activation function to the first intermediate parameter, perform multiplication operations with the edge feature and the contrast feature respectively to obtain a second intermediate parameter and a third intermediate parameter;
[0075] S2135. Perform an addition operation on the second intermediate parameter and the third intermediate parameter to obtain the attention fusion feature.
[0076] Among them, the specific method for extracting the global attention, local attention, and channel attention of the concatenated features is as Figure 3 shown.
[0077] As Figure 4 shown, in the S22, the feature fusion module using the depth feature is used to obtain the intermediate result r2, which includes:
[0078] S221. Perform depth feature extraction on the CT image with known quality scores through ConvNeXt to obtain four layers of depth features;
[0079] S222. Fuse the four layers of depth features through a bidirectional feature pyramid to obtain four layers of fused features;
[0080] S223. Classify the four layers of fused features through a classifier to obtain four layers of classification results;
[0081] S224. Perform weighted late fusion on the four layers of classification results to obtain the intermediate result r2.
[0082] Feature fusion refers to the combination of features from different levels or branches and is an omnipresent part of modern neural network architectures. It can be divided into early fusion and late fusion at the fusion stage. Among them, early fusion means fusing different features and then making predictions, which is suitable for scenarios with strong correlations between features; late fusion means fusing the prediction results of different features and is suitable for scenarios with weak correlations between features.
[0083] As Figure 5 shown, in the S221, the method for ConvNeXt to perform depth feature extraction on the CT image includes:
[0084] S2211. Perform two-dimensional convolution on the CT image along the depth direction to obtain a fourth intermediate parameter;
[0085] S2212. Perform layer normalization on the fourth intermediate parameter to obtain a fifth intermediate parameter;
[0086] After performing a two-dimensional convolution on the fifth intermediate parameter and using the GELU activation function, the sixth intermediate parameter is obtained;
[0087] S2214. Perform two-dimensional convolution, anchor point scaling, and network regularization on the sixth intermediate parameter in sequence to obtain the seventh intermediate parameter;
[0088] S2215. Concatenate the fourth intermediate parameter and the seventh intermediate parameter to obtain the depth feature of the CT image.
[0089] In S3, the loss function L of the deep learning model is:
[0090]
[0091] Among them, ri = {r1, r2} represents the intermediate result, bi is the weight corresponding to the intermediate result, and r represents the image feature obtained by weighted late fusion of the intermediate results r1 and r2.
[0092] In an embodiment of the present invention, the experimental simulation dataset is established based on the CT image low-dose simulation technology proposed by Zeng et al. on the CT image dataset publicly available at the Mayo Clinic in 2016. In Mayo2016, 16 chest and abdominal images of 10 cases, a total of 160 images, are selected. Referring to the setting of the tube current in conventional scans, all the images are simulated into 6-dose images (25, 50, 75, 100, 150, 200 mAs), and then according to the conventional quality evaluation criteria, they are divided into 480 high-quality and low-quality images respectively, jointly constituting the simulation dataset required for the experiment. The simulation method used first constructs a mathematical model based on the direct correlation between mAs (milliampere-seconds) and the incident photon flux. Then, by adjusting the data scale of the incident photon flux during high-dose scanning, the incident photon flux data under low-dose scanning conditions is simulated. After that, using the "Poisson + Gaussian" mixed statistical distribution characteristics followed by the original CT scan data, the original CT data under low-dose conditions is generated through simulation technology. Finally, using the Filtered Back Projection (FBP) algorithm, these simulated low-dose CT original data are reconstructed to obtain low-dose CT images.
[0093] For the real dataset of the experiment, according to the doctor's judgment, the data is divided into 300 high-quality and low-quality images respectively.
[0094] In the experiments described in this paper, the graphics card used in the experimental environment is an NVIDIA RTX 4060Ti with 16GB of video memory. The development tool is the professional version of PyCharm 2021.3. A model is established based on the deep learning framework PyTorch 1.7.0 with Python 3.8.10. After the same data augmentation process for all training, the AdamW optimizer is used to optimize the model. Each dataset is randomly divided into 70% of the images for training, 10% for validation, and 20% for testing.
[0095] In this embodiment, the experimental results of 4 different algorithms in recent years are collected, and the result data is summarized in Tables 1 and 2:
[0096] Results on the simulation dataset
[0097]
[0098]
[0099] Results on the real dataset
[0100]
[0101] Existing experiments have shown that this model has achieved better experimental results compared to other no-reference image quality assessments, which to a certain extent demonstrates the feasibility of this project. In summary, by extracting the texture, contrast, and depth semantic features of the image and finally using the feature fusion method, the present invention can better simulate the human eye's judgment of image distortion.
[0102] The above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A no-reference quality assessment method for CT images based on feature fusion, characterized in that, Including: S1. Collect CT images, perform distortion simulation and add quality labels to obtain CT images with known quality scores; S2. Extract features from the CT images with known quality scores to obtain image features; S3. Use the CT images with known quality scores and image features to construct and train a deep learning model to obtain a trained model; S4. Input the CT images to be evaluated for quality into the trained model, obtain the quality scores of the CT images to be evaluated for quality, and complete the no-reference quality evaluation; The S2 includes: S21. Extract the edge features and contrast features of the CT images with known quality scores, and obtain the intermediate result r1 according to the edge features and contrast features; S22. Perform deep feature extraction on the CT images with known quality scores, and obtain the intermediate result r2 according to the deep features; S23. Perform weighted late fusion on the intermediate result r1 and the intermediate result r2 to obtain image features; In the S21, a feature fusion module is used to obtain the intermediate result r1, which includes: S211. Use the Lapalce operator to extract the edge features of the CT images with known quality scores; S212. Use the gray-level co-occurrence matrix to extract the contrast features of the CT images with known quality scores; S213. Use a feature fusion module based on the attention mechanism to fuse the edge features and contrast features to obtain attention fusion features; S214. Pass the attention fusion features through a classifier to obtain the intermediate result r1; The S213 includes: S2131. Concatenate the edge features and contrast features to obtain concatenated features; S2132. Extract the global attention, local attention and channel attention of the concatenated features respectively; S2133. Use an addition operation on the global attention, local attention and channel attention of the concatenated features to obtain the first intermediate parameter; S2134. After using the sigmoid activation function on the first intermediate parameter, perform multiplication operations with the edge features and contrast features respectively to obtain the second intermediate parameter and the third intermediate parameter; S2135. Use an addition operation on the second intermediate parameter and the third intermediate parameter to obtain attention fusion features.
2. The no-reference quality assessment method for CT images based on feature fusion according to claim 1, characterized in that In the S22, a feature fusion module for deep features is used to obtain the intermediate result r2, which includes: S221. Perform deep feature extraction on the CT images with known quality scores through ConvNeXt to obtain four layers of deep features; S222. Fuse the four layers of deep features through a bidirectional feature pyramid to obtain four layers of fusion features; S223. Classify the four layers of fusion features through a classifier to obtain four layers of classification results; S224. Perform weighted late fusion on the four layers of classification results to obtain the intermediate result r2.
3. The no-reference quality assessment method for CT images based on feature fusion according to claim 2, wherein In the S221, the method for ConvNeXt to perform deep feature extraction on CT images includes: S2211. Perform two-dimensional convolution on the CT images along the depth direction to obtain the fourth intermediate parameter; S2212. Perform layer normalization on the fourth intermediate parameter to obtain the fifth intermediate parameter; S2213. After performing two-dimensional convolution on the fifth intermediate parameter, use the GELU activation function to obtain the sixth intermediate parameter; S2214. Perform two-dimensional convolution, anchor point scaling, and network regularization on the sixth intermediate parameter in sequence to obtain the seventh intermediate parameter; S2215. Concatenate the fourth intermediate parameter and the seventh intermediate parameter to obtain the depth feature of the CT image.
4. The CT image no-reference quality evaluation method based on feature fusion according to any one of claims 1-3, characterized in that, In the above S3, the loss function L of the deep learning model is: where ri = {r1, r2} represents the intermediate result, bi is the weight corresponding to the intermediate result, and r represents the image feature obtained by weighted late fusion of the intermediate results r1 and r2.
Citation Information
Patent Citations
No-reference image quality evaluation method based on deep learning
CN115272203A
Layered saliency guidance visual feature extraction model establishment and quality evaluation method
CN118154894A