Model driving and data driving fused imaging quality evaluation method
By combining the model-driven and data-driven methods, combined with lightweight convolutional neural network and gradient amplitude similarity operator, an imaging quality evaluation network is built, which solves the problems of high complexity and high computing cost of existing methods, and realizes efficient and accurate imaging quality evaluation, which is suitable for industrial monitoring scenarios.
Patent Information
- Application Number
- CN202510351555.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-18
AI Technical Summary
The existing imaging quality evaluation methods have high model complexity, high computational cost, and are difficult to achieve accurate and robust quality evaluation in complex scenarios and mixed distortion types.
A fusion model-driven and data-driven imaging quality evaluation method is designed. By extracting data-driven features, model-driven features, fusion features, feature dimensionality reduction and multi-scale feature extraction, combined with lightweight convolutional neural networks and gradient amplitude similarity operators, an ultra-light imaging quality evaluation network is constructed.
It realizes efficient and accurate imaging quality evaluation under complex scenarios and mixed distortion types, has visual quality evaluation effect close to the human eye, and is suitable for industrial monitoring scenarios with strict real-time requirements.
Smart Images

Figure FT_1 
Figure FT_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of imaging quality evaluation. Based on the imaging characteristics of objects, a data-driven convolutional neural network and a model-driven gradient magnitude similarity operator, a method for evaluating imaging quality that combines model-driven and data-driven is constructed. Background Art
[0002] With the continuous development of imaging technologies, such as infrared thermal imaging, electron microscopy imaging, hyperspectral imaging, etc., imaging quality evaluation has become an important research direction in the field of image processing and analysis. The quality of the image directly affects the application effect of the image in fields such as scientific research, industrial inspection, and medical diagnosis.
[0003] Traditional imaging quality evaluation methods are mainly divided into two categories: model-driven and data-driven. Model-driven methods are based on physical models or mathematical models, such as histogram equalization, multi-scale Retinex decomposition, etc. These methods have clear physical meanings, but are prone to problems such as detail loss and color distortion. Data-driven methods rely on a large amount of training data and automatically learn image features through deep learning models, such as convolutional neural networks (CNNs), etc. However, their performance is limited by the quality and quantity of the training data, and they have high requirements for hardware devices. In recent years, methods that combine model-driven and data-driven have gradually received attention. This method combines the physical meaning of model-driven and the automatic feature learning ability of data-driven, and can more effectively improve the accuracy and robustness of imaging quality evaluation. For example, by combining model-driven feature extraction with data-driven deep learning models, local and global features of images can be better captured, thereby achieving more accurate quality evaluation. Nevertheless, existing fusion methods still have some challenges, such as high model complexity and large computational costs. Therefore, the present invention designs a method for evaluating imaging quality that combines model-driven and data-driven, and efficiently and accurately evaluates the visual quality of images by combining a lightweight convolutional neural network driven by data and a gradient magnitude similarity operator driven by a model. Summary of the Invention
[0004] The present invention designs a method for evaluating imaging quality that combines model-driven and data-driven. This method is realized through five steps: extracting data-driven features, extracting model-driven features, fusing data-driven features and model-driven features, feature dimensionality reduction, and multi-scale feature extraction, and can significantly improve the consistency between the objective evaluation results of image quality and subjective perception. The present invention shows excellent accuracy and robustness in complex scenarios and mixed distortion types, and obtains a visual quality evaluation effect close to that of the human eye.
[0005] The present invention is realized through the following technical solutions, including the following steps:
[0006] The first step: Extract data-driven features;
[0007] The second step: Extract model-driven features;
[0008] The third step: Integrate data-driven features and model-driven features;
[0009] The fourth step: Feature dimensionality reduction;
[0010] The fifth step: Multi-scale feature extraction.
[0011] The creativity of the present invention is mainly reflected in:
[0012] (1) The imaging quality evaluation method designed by the present invention combines the advantages of the model-driven method with clear physical meaning and the data-driven method with strong automatic feature learning ability, integrates data-driven features and model-driven features, and can effectively improve the accuracy and robustness of imaging quality assessment.
[0013] (2) The present invention designs an ultra-lightweight imaging quality evaluation network. The number of parameters of the entire network model does not exceed 1k. Compared with common neural networks with millions or tens of millions of model parameters, the imaging quality evaluation network designed by the present invention can not only achieve high-accuracy quality evaluation, but also has a very high execution speed, and is suitable for industrial monitoring scenarios with strict real-time requirements. Description of the drawings
[0014] Figure 1 is a single-branch structure diagram of the imaging quality evaluation network that combines model-driven and data-driven designed by the present invention.
[0015] Figure 2 is the overall framework flowchart of the imaging quality evaluation network that combines model-driven and data-driven designed by the present invention. Detailed implementation manners
[0016] The following details the embodiments of the present invention. The embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.
[0017] Embodiment:
[0018] The first step: Extract data-driven features;
[0019] Considering the characteristic that the human eye is highly sensitive to horizontal and vertical stimuli, the present invention designs 1×5 and 5×1 bar-shaped convolutions to extract features from the input image, so as to simulate the perception of the input image by the human eye. Therefore, the present invention connects the "1×5 convolutional layer - batch normalization (BN) layer - ReLU activation layer" and the "5×1 convolutional layer - BN layer - ReLU activation layer" in parallel, and performs a concatenation operation on the output features of the two to build a data-driven module and extract the data-driven feature D F , as Figure 1 shown.
[0020] Step 2: Extract the model-driven feature;
[0021] The human visual system is very sensitive to the texture and contour changes of images. The present invention uses the gradient magnitude similarity operator, which is sensitive to the structural changes of images, can effectively capture the texture and contour of images, and adapts to various distortion types, as the model-driven module to calculate the gradient magnitude similarity of the input image, so as to realistically reflect the changes in the human eye's perception of the input image. Therefore, this operator is used to extract the model-driven feature M F .
[0022] Step 3: Fuse the data-driven feature and the model-driven feature;
[0023] To better characterize the image quality, as Figure 1 shown, the present invention first expands the dimension of the model-driven feature M F to align it with the data-driven feature D F , and then sets the weight vector W to fuse the model-driven feature and the data-driven feature to obtain the digital-analog hybrid feature MD F :
[0024] MD F = W × Exp(M F , c) + M F
[0025] where Exp(·) represents expanding the model-driven feature M F in the channel dimension c.
[0026] Step 4: Feature dimensionality reduction;
[0027] First, the present invention reduces the channel dimension of the digital-analog hybrid feature MD F to 1 through a multi-layer perceptron (MLP) to obtain the feature Then, global average pooling is used on it to obtain 1 eigenvalue S1. In other words, the global average pooling layer is used to reduce the dimension in both the width and height dimensions. It should be noted that if the network has only 1 channel, then the eigenvalue S1 is the quality score of the image.
[0028] Step 5: Multi-scale feature extraction;
[0029] To further improve the image quality assessment ability of the network, the present invention repeats the above four steps to construct five branches. As Figure 2 shown, perform Maxpooling downsampling operations with a stride of n on the reference image and the distorted image to obtain the reference image features F Ri and the distorted image features F Di at different scales. Then, input the image features at each scale into the network for data feature extraction, model feature extraction, data feature and model feature fusion, and feature dimensionality reduction operations to obtain the eigenvalue S i at each scale. In the present invention, {n = 0, 2, 4, 8, 16}, {i = 1, 2, 3, 4, 5}. Finally, use a simple single-layer fully connected layer to map the eigenvalues at each scale to the final image quality:
[0030]
[0031] where S = {S1, S2, S3, S4, S5} is the set of eigenvalues at each scale, Ψ(S) is the number of elements in S, W i and b i are the weight and bias corresponding to the i-th node in the fully connected layer respectively, is the convolution operation, and quality is the image quality score.
Claims
1. Extract data-driven features; Considering the characteristic that the human eye is highly sensitive to horizontal and vertical stimuli, the present invention designs 1×5 and 5×1 strip convolutions to extract features from the input image, so as to simulate the perception of the input image by the human eye. Therefore, the present invention connects in parallel the "1×5 convolutional layer - batch normalization (BN) layer - ReLU activation layer" and the "5×1 convolutional layer - BN layer - ReLU activation layer", and performs a concatenation operation on the output features of the two to build a data-driven module and extract the data-driven feature D F .
2. Extract model-driven features; The human visual system is very sensitive to the texture and contour changes of images. The present invention uses a gradient magnitude similarity operator that is sensitive to image structure changes, can effectively capture the texture and contour of images, and adapts to various distortion types as the model-driven module to calculate the gradient magnitude similarity of the input image, so as to realistically reflect the changes in the input image perceived by the human eye. Therefore, this operator is used to extract the model-driven feature M F .
3. Integrate data-driven features with model-driven features; To better characterize the image quality, the present invention first expands the dimension of the model-driven feature M F to align it with the data-driven feature D F and then sets the weight vector W to fuse the model-driven feature and the data-driven feature to obtain the hybrid model-data feature MD F : MD F = W × Exp(M F , c) + M F where Exp(.) represents the model-driven feature M F is extended in the channel dimension c.
4. Feature dimensionality reduction; First, the present invention reduces the channel dimension of the mixed analog-digital feature MD to 1 through a multi-layer perceptron (MLP) to obtain a feature F Then, global average pooling is applied to it to obtain a feature value S1. In other words, the global average pooling layer is used to reduce the dimension in both the width and height dimensions. It should be noted that if the network has only 1 channel, then the feature value S1 is the quality score of the image. 5. Multi-scale feature extraction; In order to further improve the image quality evaluation ability of the network, the present invention repeats the above four steps to construct five branches. A Maxpooling downsampling operation with a step size of n is performed on the reference image and the distorted image to obtain the reference image features F Ri and the distorted image features F Di at different scales. Then, the image features at each scale are input into the network for data feature extraction, model feature extraction, fusion of data features and model features, and feature dimensionality reduction operations to obtain the eigenvalue S i at each scale. In the present invention, {n = 0, 2, 4, 8, 16}, {i = 1, 2, 3, 4, 5}. Finally, a simple one-layer fully connected layer is used to map the eigenvalues at each scale to the final image quality: where \(S = \{S_1, S_2, S_3, S_4, S_5\}\) is the set of eigenvalues at each scale, \(\Psi(S)\) is the number of elements in \(S\), \(W\) i and \(b\) i are the weight and bias corresponding to the \(i\)-th node in the fully connected layer respectively, is the convolution operation, and quality is the image quality score.