Low-illumination enhanced image quality evaluation method based on image aesthetics-noise double branches
By constructing a dual-branch network for aesthetic perception and noise perception, the problem of insufficient noise modeling in low-light enhanced images is solved, enabling accurate evaluation of the quality of low-light enhanced images and improving the visual appeal and noise suppression effect of the images.
Patent Information
- Application Number
- CN202510863560.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-28
AI Technical Summary
Existing low-light image enhancement methods ignore the complex distribution characteristics and aesthetic features of noise, resulting in unsatisfactory noise suppression effects and making it difficult to comprehensively and accurately evaluate the quality of low-light enhanced images.
A dual-branch network for aesthetic perception and noise perception is constructed. Features are extracted through ResNeXt-101 and WideResNet-101. Combined with a multi-scale denoiser and a channel-aware cross-modulation module, a comprehensive quality assessment of low-light enhanced images is achieved.
Accurate assessment of low-light enhanced image quality improves the visual appeal and noise suppression of images, conforms to human visual perception, and enhances the accuracy of image quality evaluation.
Smart Images

Figure CN120852295A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and image processing technology, specifically relating to image quality assessment technology under low light conditions, and in particular a deep learning-based no-reference image quality assessment method. By constructing a dual-branch network for aesthetic perception and noise perception, it achieves joint modeling and evaluation of the subjective aesthetic quality and noise quality of low-light enhanced images, solving the shortcomings of existing methods in noise modeling and feature interaction of low-light enhanced images, and improving the consistency between quality assessment and human visual perception. Background Technology
[0002] In the widespread application of modern digital imaging technology, images acquired in low-light environments face numerous quality challenges, making the development of low-light image quality enhancement and evaluation technologies an urgent need. Under low-light conditions, the limited light captured by camera sensors results in images generally exhibiting low brightness and poor contrast, making key information in the image blurry and difficult to discern. Furthermore, low-light images are highly susceptible to noise interference; this noise, like "snowflakes," is scattered throughout the image, severely affecting its clarity and visual effect. In addition, color distortion frequently occurs; originally vibrant and realistic colors become dull and deviate from their original hues under low light, further reducing image quality.
[0003] These low-quality, low-light images have negatively impacted numerous fields that rely on image information. In security surveillance, poor image quality under low light conditions can make it difficult to identify target objects and obscure the features of criminals, significantly reducing the early warning and tracking capabilities of surveillance systems and posing a threat to public safety. In autonomous driving scenarios, if onboard cameras acquire poor image quality in low-light environments, the vehicle's visual recognition system may be unable to accurately identify road signs, pedestrians, and other vehicles, posing a serious threat to the safety of autonomous driving and easily leading to traffic accidents. In medical imaging, image quality problems caused by low light can make it difficult for doctors to accurately determine lesion characteristics, affecting early diagnosis and treatment planning, and delaying patients' treatment.
[0004] To improve the quality of low-light images, researchers have developed various low-light enhancement algorithms. Traditional methods, such as histogram equalization, enhance contrast by redistributing image grayscale values, but often over-enhance certain areas, leading to loss of image detail, blockiness, and poor noise suppression. Methods based on Retinex theory decompose the image into reflection and illumination components to enhance it; however, in practical applications, they tend to amplify noise, making the enhanced image more noisy. Furthermore, they struggle to accurately recover the natural colors and details of images in complex lighting scenes.
[0005] With the rise of deep learning technology, low-light image enhancement methods based on deep learning have gradually become mainstream. Some methods achieve enhancement by constructing deep neural networks to learn the mapping relationship between low-light and normal-light images. However, most of these methods only focus on improving image brightness and contrast, ignoring the noise introduced during the enhancement process and the changes in image aesthetic quality. For example, while some methods improve the overall visibility of the image by enhancing brightness, they also amplify noise in the image, making the image coarser; other methods generate enhanced images that lack visual appeal, with unnatural colors and inconsistent contrast, failing to conform to human visual perception habits. Moreover, existing methods have shortcomings in noise modeling, failing to fully consider the complex distribution characteristics of noise and the interaction between noise and image content, resulting in unsatisfactory noise suppression effects. In terms of feature interaction, they also fail to effectively integrate the aesthetic and noise features of the image, making it difficult to comprehensively and accurately evaluate the quality of low-light enhanced images.
[0006] Therefore, developing a low-light enhancement image quality evaluation method that can comprehensively consider image aesthetic quality and noise factors is of great significance for accurately evaluating the performance of low-light enhancement algorithms, guiding algorithm optimization, and improving the usability of low-light images in practical applications. Summary of the Invention
[0007] The core of this invention, "A Low-Light Enhanced Image Quality Evaluation Method Based on Image Aesthetics-Noise Dual Branch," lies in constructing a dual-branch structure of aesthetic perception and noise perception to achieve comprehensive and accurate quality evaluation of low-light enhanced images. In the aesthetic perception branch, the low-light enhanced image is first input, and a pre-trained ResNeXt-101 backbone network is used to extract deep aesthetic representation features. Then, global average pooling and global standard deviation pooling operations are used to extract overall statistical features of the image from the spatial dimension. These two features are concatenated to form an aesthetic feature vector that integrates global semantics and local variations. Finally, this vector is input into a dual-convolutional integral layer regression network (DCHRN) to predict the subjective aesthetic quality score. In the noise perception branch, the input image is first denoised using a multi-scale denoiser to generate a denoised image. The pixel difference between the original image and the denoised image is calculated to obtain a noise residual map, which represents the noise distribution. Then, a pre-trained WideResNet-101 backbone network is used to extract noise-related features from the residual map. To integrate aesthetic and noise features, a channel-aware cross-modulation module projects noise and aesthetic features onto the same channel dimension. Dynamically gated vectors are generated through linear mapping and GELU activation, and element-wise modulated aesthetic features are applied. Residual connections are then used to generate aesthetically modulated noise features. Finally, the modulated noise features undergo the same pooling and concatenation operations as the aesthetic branch. These are input to DCHRN regression to obtain a subjective noise perception quality score. Weighted summation is then used to fuse the aesthetic and noise perception quality scores to generate an overall quality assessment result. The noise branch prediction score is only used for loss function calculation and does not participate in the overall score weighting. Ultimately, the aesthetic quality score dominates the overall assessment, achieving a comprehensive and effective evaluation of the quality of low-light enhanced images.
[0008] In embodiments according to this disclosure, the low-light enhanced image quality assessment method based on image aesthetics-noise dual-branch includes the following steps:
[0009] Step 1: Input the low-light enhanced image into the aesthetic perception branch, use the pre-trained ResNeXt-101 backbone network to extract deep aesthetic representation features, extract the overall statistical features of the image through global average pooling and global standard deviation pooling operations, and concatenate them to form an aesthetic feature vector that integrates global semantics and local changes. Finally, input it into the dual-convolution integral layer regression network (DCHRN) to predict the subjective aesthetic quality score.
[0010] Step 2: In the noise perception branch, the input image is first denoised using a multi-scale denoiser to generate a denoised image. The pixel difference between the original image and the denoised image is calculated to obtain a noise residual map representing the noise distribution. Noise-related features are extracted from the residual map using a pre-trained WideResNet-101 backbone network.
[0011] Step 3: Semantic modulation of noise features through channel-aware cross-modulation module: Project noise features and aesthetic features to the same channel dimension, generate dynamic gating vectors through linear mapping and GELU activation, perform element-level modulation on aesthetic features, and generate aesthetically modulated noise features by combining residual connections;
[0012] Step 4: After the modulated noise features undergo the same pooling and splicing operations as the aesthetic branch, they are input into DCHRN regression to obtain the subjective noise perception quality score. The overall quality assessment result is then generated through weighted summation and fusion.
[0013] Further technical solutions, step 1 specifically includes the following steps:
[0014] Step 1.1: Input the low-light enhanced image I (dimensions (B×C×H×W), where B is the batch size, C is the number of channels, and (H, W) are the image width and height) into the aesthetic perception branch. Use a pre-trained ResNeXt-101 backbone network to extract features from it. Leveraging the network's grouped convolutions and residual structure, capture deep semantic information of the image and output deep aesthetic representation features (F). A (dimension is (B×C)) feat ×H fear ×W feat (C) feat H represents the number of feature channels. feat W feat The calculation formula is as follows (where the feature map width and height are given):
[0015] F A =ResNeXt-101(I)
[0016] Where I is the low-light enhancement image, F A It is a characteristic of deep aesthetic expression.
[0017] Step 1.2: Extract the deep aesthetic representation features F A Perform Global Average Pooling (GAP) and Global Standard Deviation Pooling (GSP) operations to aggregate image statistical features. Global Average Pooling (GAP): In the spatial dimension (H... feat ×W feat By averaging across the data, we can capture global semantic information and obtain the average feature F. A_avg The formula is:
[0018]
[0019] The output dimension is B×C feat Global Standard Deviation Pooling (GSP): First calculate feature F A The standard deviation is obtained by taking the square root of the variance in the spatial dimension, and is used to capture information about local variations. The variance is calculated as follows:
[0020]
[0021] To avoid the square root being meaningless when the variance is 0, a minimum value ∈ is introduced, and the final standard deviation characteristic FA is obtained. _std satisfy:
[0022]
[0023] The output dimension is also B×C. feat .
[0024] Step 1.3: Calculate the global average pooling result F. A_avg Compared with the global standard deviation pooling result F A_std By concatenating data along the channel dimension and integrating global semantics and local variation information, an aesthetic feature vector F is formed. A_cat Its dimension becomes B×(2×C) feat The mathematical expression is:
[0025] F A_cat =Concat(F A_avg ,F A _ std )
[0026] Step 1.4: Combine the concatenated aesthetic feature vector F A_cat The input is a dual-convolutional integral layer regression network (DCHRN). DCHRN, through its internal convolutional layers and activation functions, further transforms and regresses the features, outputting a subjective aesthetic quality score S. A That is: S A =DCHRN(F A_cat The score S A As a scalar, it is used to quantify the aesthetic quality of low-light enhanced images. The higher the value, the better the aesthetic performance of the image, thus completing the prediction of the subjective aesthetic quality of low-light enhanced images and providing a basis for subsequent overall quality evaluation.
[0027] Further technical solutions, step 2 specifically includes the following steps:
[0028] Step 2.1: The low-light enhancement image is used as input and fed into a multi-scale denoising (MSD) denoising unit. Based on its multi-scale feature extraction and noise suppression mechanism, the MSD performs noise filtering on the input image, identifying and processing noise components at different scales to generate a denoised image. Let the input low-light enhancement image be I, and the denoised image I obtained after processing by the MSD denoising unit is... denoise Mathematically, it can be represented as
[0029] I denoise =MSD(I)
[0030] Step 2.2: To obtain noise information in the image, calculate the original low-light enhanced image I and the denoised image I. denoise The pixel difference between the two images. Since the denoised image removes most of the noise, the pixel difference can reflect the distribution of the noise, thus obtaining the noise residual map N, which is calculated as N = II. denoise The noise residual map N depicts the spatial distribution and intensity information of noise in the image, providing a basis for subsequent noise feature extraction.
[0031] Step 2.3: Input the noisy residual map N into the pre-trained WideResNet-101 backbone network. WideResNet-101, leveraging its deep convolutional structure and pre-trained feature extraction capabilities, performs feature mining on the noisy residual map. During the network's forward propagation, through multi-layer convolution and pooling operations, it extracts deep-level features F related to noise from the noisy residual map. N These features include information such as the frequency and distribution pattern of the noise, and can be represented as...
[0032] F N =WideResNet-101(N)
[0033] This lays the foundation for subsequent evaluation of noise perception quality in conjunction with aesthetic features.
[0034] Further technical solutions, step 3 specifically includes the following steps:
[0035] Step 3.1: To achieve refined semantic modulation of noise features and better integrate aesthetic and noise perception information, the noise features F extracted from the noise perception branch are first... N Aesthetic features F obtained from the aesthetic perception branch A As input. Since the initial channel dimensions of the two may differ, a specially designed linear projection layer is used to transform the noise features F... N and aesthetic features F A By mapping each feature to the same target channel dimension C, we obtain noise features with unified dimensions. and aesthetic features The mathematical expression can be described as
[0036]
[0037] Among them W N W A Let b be the weight matrix of the linear projection. N b A This is a bias term.
[0038] Step 3.2: Targeting the aesthetic features after dimensional unification The vector is then transformed again using a linear mapping operation and connected to the GELU (Gaussian Error Linear Units) activation function to generate a dynamically adjustable gated vector G. The linear mapping aims to adapt the aesthetic features to different dimensions and transform information, while the GELU activation introduces non-linearity, allowing the gated vector to adaptively learn and modulate weights based on the feature content. The specific calculation is as follows:
[0039]
[0040] Here W G It is a linear mapping weight, b G For bias.
[0041] Step 3.3: After obtaining the dynamic gating vector G, based on the logic of channel-aware cross-modulation, use this gating vector to apply aesthetic features. Element-level modulation is performed. Element-level modulation involves multiplying the aesthetic features element-wise with the gating vector, allowing the aesthetic features to be adaptively enhanced or suppressed at different channel positions based on the gating information. Mathematically, this is represented as...
[0042]
[0043] Here, ⊙ represents element-wise multiplication, thereby achieving feature modulation based on aesthetic semantics.
[0044] Step 3.4: To preserve the basic information of the original noise features and avoid information loss during modulation, a residual connection mechanism is introduced. This mechanism connects the aesthetic features after element-level modulation. Compared with the original noise characteristics The fusion process involves adding the two components along the channel dimension to obtain the aesthetically modulated noise characteristics. Right now Residual connections can effectively transfer gradients and improve the stability of network training. At the same time, they allow the modulated noise features to incorporate both aesthetic and semantic information and retain the key content of the original noise features, providing a more discriminative feature representation for subsequent quality score prediction.
[0045] Further technical solutions, step 4 specifically includes the following steps:
[0046] Step 4.1: First, analyze the aesthetically modulated noise features generated in Step 3. Perform pooling and splicing operations consistent with the aesthetics branch. Specifically, first... Perform Global Average Pooling (GAP) and Global Standard Deviation Pooling (GSP): Global Average Pooling calculates the mean of the feature across the spatial dimension. To capture the overall distribution trend of noise features; to calculate the standard deviation of features in the spatial dimension through global standard deviation pooling. To characterize the degree of local variation in noise. The mathematical expressions for both are:
[0047]
[0048] Where H×W is the spatial dimension of the feature map, and ∈ represents the minimum value to avoid numerical instability during square root extraction. Subsequently, and The features are concatenated along the channel dimension to form a feature vector that integrates the global distribution and local variations of noise.
[0049] Step 4.2: Next, will The input is a Dual-Convolutional Integral Layer Regression Network (DCHRN). DCHRN first reshapes the feature vector to B×1×D (where D is the feature dimension), and then extracts non-linear features through two 3×3 convolutional layers: the first convolutional layer expands the number of channels to 16, and the second layer expands it to 32, with each layer followed by a PReLU activation function. Subsequently, the features are flattened and regressed through a fully connected layer, outputting a subjective noise-perceived quality score S. N ,Right now:
[0050] Step 4.3: Blend the aesthetic quality score S by weighted summation A With noise perception quality score S N Generate overall quality assessment result S R According to the paper's specifications, the weighting formula is: Where λ1 = 1 and λ2 = 0, that is, the prediction score S of the noise branch. N Used only for loss calculation, it does not actually participate in the weighting process of the overall score, and the final S R By S A It takes the lead in achieving a comprehensive evaluation of the quality of low-light enhanced images.
[0051] The low-light enhanced image quality evaluation method based on image aesthetics-noise dual branches provided by this invention achieves accurate evaluation of low-light enhanced image quality through innovative network architecture design and multi-dimensional feature fusion mechanism. Its beneficial effects are as follows:
[0052] 1) By employing parallel design of aesthetic perception and noise perception branches, the subjective aesthetic degradation and noise artifact features of low-light enhanced images are simultaneously captured. The aesthetic branch utilizes a pre-trained ResNeXt-101 to extract deep semantic features, combining global statistical pooling (average pooling + standard deviation pooling) to fuse global semantics and local variations, accurately characterizing the visual appeal of the image. The noise branch generates a noise residual map through a multi-scale denoiser (MSD), combining it with WideResNet-101 to extract noise distribution features, addressing the shortcomings of traditional methods in modeling complex noise. The dual-branch structure collaboratively models the potential coupling relationship between the two branches (such as the impact of noise on aesthetic perception), making the evaluation results closer to human visual perception.
[0053] 2) Introducing the CCMB module to achieve semantic modulation of noise features by aesthetic features: Two types of features are aligned to the same dimension through linear projection, and a dynamic gating vector is generated using GELU activation to adaptively enhance or suppress aesthetically semantically related components in noise features. This mechanism breaks through the limitations of simple feature concatenation in traditional methods, establishing deep associations between features of different dimensions (such as noise distribution being constrained by the semantic structure of the image), thus improving feature representation capabilities. Experiments show that removing CCMB leads to a 1.1% decrease in the SROCC index, verifying its crucial role in capturing feature interaction logic.
[0054] 3) The noise branch employs the MSD module (based on U-Net and multi-head self-attention) to achieve multi-scale noise suppression and detail preservation. A noise residual map is generated by calculating the pixel difference between the original image and the denoised image, effectively separating noise from image content. This method overcomes the limitations of traditional noise models (such as those considering only Gaussian / salt-and-pepper noise) and can capture the complex noise distribution in low-light enhanced images caused by sensor noise, gain amplification, and algorithm artifacts.
[0055] 4) The DCHRN module is designed to fuse hierarchical features with two layers of 3×3 convolutions, achieving a non-linear mapping from local features to global quality scores. Compared to traditional MLP or KAN networks, this structure maintains high accuracy while reducing the number of parameters. Its hierarchical convolutional design effectively captures the spatial dependencies of multi-scale degenerate features, solving the problem of insufficient robustness of traditional regression models to local noise and contrast changes. Attached Figure Description
[0056] Figure 1 The flowchart is a low-light enhancement image quality evaluation method based on image aesthetics-noise dual branch, which is the subject of this invention.
[0057] Figure 2 This is a schematic diagram of the network structure of the low-light enhancement image quality evaluation method based on image aesthetics-noise dual branches involved in this invention. Detailed Implementation
[0058] This invention proposes a low-light image quality evaluation method, utilizing a deep convolutional neural network to automatically evaluate low-light enhanced images. To more clearly illustrate the purpose, technical solution, and advantages of this invention, the technical solution will be described in detail below with reference to the accompanying drawings. It should be noted that the specific embodiments described below are only for explaining the principles of this invention and are not intended to limit the scope of its implementation.
[0059] This invention first provides a low-light enhancement image quality assessment method based on a dual-branch approach of image aesthetics and noise, as detailed in the flowchart. Figure 1 , including the following steps:
[0060] Step 1: Input the low-light enhanced image into the aesthetic perception branch, use the pre-trained ResNeXt-101 backbone network to extract deep aesthetic representation features, extract the overall statistical features of the image through global average pooling and global standard deviation pooling operations, and concatenate them to form an aesthetic feature vector that integrates global semantics and local changes. Finally, input it into the dual-convolution integral layer regression network (DCHRN) to predict the subjective aesthetic quality score.
[0061] Step 2: In the noise perception branch, the input image is first denoised using a multi-scale denoiser to generate a denoised image. The pixel difference between the original image and the denoised image is calculated to obtain a noise residual map representing the noise distribution. Noise-related features are extracted from the residual map using a pre-trained WideResNet-101 backbone network.
[0062] Step 3: Semantic modulation of noise features through channel-aware cross-modulation module: Project noise features and aesthetic features to the same channel dimension, generate dynamic gating vectors through linear mapping and GELU activation, perform element-level modulation on aesthetic features, and generate aesthetically modulated noise features by combining residual connections;
[0063] Step 4: After the modulated noise features undergo the same pooling and splicing operations as the aesthetic branch, they are input into DCHRN regression to obtain the subjective noise perception quality score. The overall quality assessment result is then generated through weighted summation and fusion.
[0064] Further technical solutions, step 1 specifically includes the following steps:
[0065] Step 1.1: Input the low-light enhanced image I (dimensions (B×C×H×W), where B is the batch size, C is the number of channels, and (H, W) are the image width and height) into the aesthetic perception branch. Use a pre-trained ResNeXt-101 backbone network to extract features from it. Leveraging the network's grouped convolutions and residual structure, capture deep semantic information of the image and output deep aesthetic representation features (F). A (dimension is (B×C))feat ×H feat ×W feat (C) feat H represents the number of feature channels. feat W feat The calculation formula is as follows (where the feature map width and height are given):
[0066] F A =ResNeXt-101(I)
[0067] Where I is the low-light enhancement image, F A It is a characteristic of deep aesthetic expression.
[0068] Step 1.2: Extract the deep aesthetic representation features F A Perform Global Average Pooling (GAP) and Global Standard Deviation Pooling (GSP) operations to aggregate image statistical features. Global Average Pooling (GAP): In the spatial dimension (H... feat ×W feat By averaging across the data, we can capture global semantic information and obtain the average feature F. A_avg, The formula is:
[0069]
[0070] The output dimension is B×C feat Global Standard Deviation Pooling (GSP): First calculate feature F A The standard deviation is obtained by taking the square root of the variance in the spatial dimension, and is used to capture information about local variations. The variance is calculated as follows:
[0071]
[0072] To avoid the square root being meaningless when the variance is 0, a minimum value ∈ is introduced, and the final standard deviation characteristic F is determined. A_std satisfy:
[0073]
[0074] The output dimension is also B×C. feat .
[0075] Step 1.3: Calculate the global average pooling result F. A_avg Compared with the global standard deviation pooling result F A_std By concatenating data along the channel dimension and integrating global semantics and local variation information, an aesthetic feature vector F is formed. A_cat Its dimension becomes B×(2×C) feat The mathematical expression is:
[0076] F A_cat =Concat(F A_avg , F A_std )
[0077] Step 1.4: Combine the concatenated aesthetic feature vector F A_cat The input is a dual-convolutional integral layer regression network (DCHRN). DCHRN, through its internal convolutional layers and activation functions, further transforms and regresses the features, outputting a subjective aesthetic quality score S. A That is: S A =DCHRN(F A_cat The score S A As a scalar, it is used to quantify the aesthetic quality of low-light enhanced images. The higher the value, the better the aesthetic performance of the image, thus completing the prediction of the subjective aesthetic quality of low-light enhanced images and providing a basis for subsequent overall quality evaluation.
[0078] Further technical solutions, step 2 specifically includes the following steps:
[0079] Step 2.1: The low-light enhancement image is used as input and fed into a multi-scale denoising (MSD) denoising unit. Based on its multi-scale feature extraction and noise suppression mechanism, the MSD performs noise filtering on the input image, identifying and processing noise components at different scales to generate a denoised image. Let the input low-light enhancement image be I, and the denoised image I obtained after processing by the MSD denoising unit is... denoisae Mathematically, it can be represented as
[0080] I denoiss =MSD(I)
[0081] Step 2.2: To obtain noise information in the image, calculate the original low-light enhanced image I and the denoised image I. denoise The pixel difference between the two images. Since the denoised image removes most of the noise, the pixel difference can reflect the distribution of the noise, thus obtaining the noise residual map N, which is calculated as N = II. denoise The noise residual map N depicts the spatial distribution and intensity information of noise in the image, providing a basis for subsequent noise feature extraction.
[0082] Step 2.3: Input the noisy residual map N into the pre-trained WideResNet-101 backbone network. WideResNet-101, leveraging its deep convolutional structure and pre-trained feature extraction capabilities, performs feature mining on the noisy residual map. During the network's forward propagation, through multi-layer convolution and pooling operations, it extracts deep-level features F related to noise from the noisy residual map. N These features include information such as the frequency and distribution pattern of the noise, and can be represented as...
[0083] F N =WideResNet-101(N)
[0084] This lays the foundation for subsequent evaluation of noise perception quality in conjunction with aesthetic features.
[0085] Further technical solutions, step 3 specifically includes the following steps:
[0086] Step 3.1: To achieve refined semantic modulation of noise features and better integrate aesthetic and noise perception information, the noise features F extracted from the noise perception branch are first... N Aesthetic features F obtained from the aesthetic perception branch A As input. Since the initial channel dimensions of the two may differ, a specially designed linear projection layer is used to transform the noise features F... N and aesthetic features F A By mapping each feature to the same target channel dimension C, we obtain noise features with unified dimensions. and aesthetic features The mathematical expression can be described as
[0087]
[0088] Among them W N W A Let b be the weight matrix of the linear projection. N b A This is a bias term.
[0089] Step 3.2: Targeting the aesthetic features after dimensional unification The vector is then transformed again using a linear mapping operation and connected to the GELU (Gaussian Error Linear Units) activation function to generate a dynamically adjustable gated vector G. The linear mapping aims to adapt the aesthetic features to different dimensions and transform information, while the GELU activation introduces non-linearity, allowing the gated vector to adaptively learn and modulate weights based on the feature content. The specific calculation is as follows:
[0090]
[0091] Here w G It is a linear mapping weight, b G For bias.
[0092] Step 3.3: After obtaining the dynamic gating vector G, based on the logic of channel-aware cross-modulation, use this gating vector to apply aesthetic features. Element-level modulation is performed. Element-level modulation involves multiplying the aesthetic features element-wise with the gating vector, allowing the aesthetic features to be adaptively enhanced or suppressed at different channel positions based on the gating information. Mathematically, this is represented as...
[0093]
[0094] Here, ⊙ represents element-wise multiplication, thereby achieving feature modulation based on aesthetic semantics.
[0095] Step 3.4: To preserve the basic information of the original noise features and avoid information loss during modulation, a residual connection mechanism is introduced. This mechanism connects the aesthetic features after element-level modulation. Compared with the original noise characteristics The fusion process involves adding the two components along the channel dimension to obtain the aesthetically modulated noise characteristics. Right now Residual connections can effectively transfer gradients and improve the stability of network training. At the same time, they allow the modulated noise features to incorporate both aesthetic and semantic information and retain the key content of the original noise features, providing a more discriminative feature representation for subsequent quality score prediction.
[0096] Further technical solutions, step 4 specifically includes the following steps:
[0097] Step 4.1: First, analyze the aesthetically modulated noise features generated in Step 3. Perform pooling and splicing operations consistent with the aesthetics branch. Specifically, first... Perform Global Average Pooling (GAP) and Global Standard Deviation Pooling (GSP): Global Average Pooling calculates the mean of the feature across the spatial dimension. To capture the overall distribution trend of noise features; to calculate the standard deviation of features in the spatial dimension through global standard deviation pooling. To characterize the degree of local variation in noise. The mathematical expressions for both are:
[0098]
[0099] Where H×W is the spatial dimension of the feature map, and ∈ represents the minimum value to avoid numerical instability during square root extraction. Subsequently, and The features are concatenated along the channel dimension to form a feature vector that integrates the global distribution and local variations of noise.
[0100] Step 4.2: Next, will The input is a Dual-Convolutional Integral Layer Regression Network (DCHRN). DCHRN first reshapes the feature vector to B×1×D (where D is the feature dimension), and then extracts non-linear features through two 3×3 convolutional layers: the first convolutional layer expands the number of channels to 16, and the second layer expands it to 32, with each layer followed by a PReLU activation function. Subsequently, the features are flattened and regressed through a fully connected layer, outputting a subjective noise-perceived quality score S. N ,Right now:
[0101] Step 4.3: Blend the aesthetic quality score S by weighted summationA With noise perception quality score S N Generate overall quality assessment result S R According to the paper's specifications, the weighting formula is: Where λ1 = 1 and λ2 = 0, that is, the prediction score S of the noise branch. N Used only for loss calculation, it does not actually participate in the weighting process of the overall score, and the final S R By S A It takes the lead in achieving a comprehensive evaluation of the quality of low-light enhanced images.
[0102] The above description mentions various methods of combining technical features, but does not list all possible combinations. However, it should be emphasized that as long as the combination of these technical features is reasonable in implementation and does not contradict each other, it should be considered within the scope of this specification.
Claims
1. A method for evaluating the quality of low-light enhanced images based on a dual-branch approach of image aesthetics and noise, characterized in that, The method includes the following steps: Step 1: Input the low-light enhanced image into the aesthetic perception branch, use the pre-trained ResNeXt-101 backbone network to extract deep aesthetic representation features, extract the overall statistical features of the image through global average pooling and global standard deviation pooling operations, and concatenate them to form an aesthetic feature vector that integrates global semantics and local changes. Finally, input it into the dual-convolution integral layer regression network (DCHRN) to predict the subjective aesthetic quality score. Step 2: In the noise perception branch, the input image is first denoised using a multi-scale denoiser to generate a denoised image. The pixel difference between the original image and the denoised image is calculated to obtain a noise residual map representing the noise distribution. Noise-related features are extracted from the residual map using a pre-trained WideResNet-101 backbone network. Step 3: Semantic modulation of noise features using a channel-aware cross-modulation module: Noise features and aesthetic features are projected onto the same channel dimension. A dynamic gating vector is generated through linear mapping and GELU activation. The aesthetic features are then element-wise modulated, and residual connections are used to generate aesthetically modulated noise features. Step 4: After the modulated noise features undergo the same pooling and splicing operations as the aesthetic branch, they are input into DCHRN regression to obtain the subjective noise perception quality score. The overall quality assessment result is generated by weighted summation and fusion.
2. The low-light enhanced image quality evaluation method based on image aesthetics-noise dual-branch as described in claim 1, characterized in that, Step 1 specifically includes the following steps: Step 1.1: Input the low-light enhanced image I (dimensions (B×C×H×W), where B is the batch size, C is the number of channels, and (H, W) are the image width and height) into the aesthetic perception branch. Use the pre-trained ResNeXt-101 backbone network to extract features from it. Leveraging the network's grouped convolutions and residual structure, capture the deep semantic information of the image and output the deep aesthetic representation feature (FA) (dimensions (B×C)). feat ×H feat ×W feat (C) feat H represents the number of feature channels. feat W feat The calculation formula is as follows (where the feature map width and height are given): F A =ResNeXt-101(I) Where I is the low-light enhancement image, F A It is a characteristic of deep aesthetic expression. Step 1.2: Extract the deep aesthetic representation features F A Perform Global Average Pooling (GAP) and Global Standard Deviation Pooling (GSP) operations to aggregate image statistical features. Global Average Pooling (GAP): In the spatial dimension (H... feat ×W feat By averaging across the data, we can capture global semantic information and obtain the average feature F. A_avg The formula is: The output dimension is B×C feat Global Standard Deviation Pooling (GSP): First calculate feature F A The standard deviation is obtained by taking the square root of the variance in the spatial dimension, and is used to capture information about local variations. The variance is calculated as follows: To avoid the square root being meaningless when the variance is 0, a minimum value ∈ is introduced, and the final standard deviation characteristic F is determined. A_std satisfy: The output dimension is also B×C. feat . Step 1.3: Calculate the global average pooling result F. A_avg Compared with the global standard deviation pooling result F A_std By concatenating data along the channel dimension and integrating global semantics and local variation information, an aesthetic feature vector F is formed. A_cat Its dimension becomes B×(2×C) feat The mathematical expression is: F A_cat =Concat(F A_avg ,F A_std ) Step 1.4: Combine the concatenated aesthetic feature vector F A_cat The input is a dual-convolutional integral layer regression network (DCHRN). DCHRN, through its internal convolutional layers and activation functions, further transforms and regresses the features, outputting a subjective aesthetic quality score S. A That is: S A =DCHRN(F A_cat The score S A As a scalar, it is used to quantify the aesthetic quality of low-light enhanced images. The higher the value, the better the aesthetic performance of the image, thus completing the prediction of the subjective aesthetic quality of low-light enhanced images and providing a basis for subsequent overall quality evaluation.
3. The low-light enhancement image quality evaluation method based on image aesthetics-noise dual-branch as described in claim 1, characterized in that, Step 2 specifically includes the following steps: Step 2.1: The low-light enhancement image is used as input and fed into a multi-scale denoising (MSD) denoising unit. Based on its multi-scale feature extraction and noise suppression mechanism, the MSD performs noise filtering on the input image, identifying and processing noise components at different scales to generate a denoised image. Let the input low-light enhancement image be I, and the denoised image I obtained after processing by the MSD denoising unit is... denoise Mathematically, it can be represented as I denoise =MSD(I) Step 2.2: To obtain noise information in the image, calculate the original low-light enhanced image I and the denoised image I. denoise The pixel difference between the two images. Since the denoised image removes most of the noise, the pixel difference can reflect the distribution of the noise, thus obtaining the noise residual map N, which is calculated as N = II. denoise The noise residual map N depicts the spatial distribution and intensity information of noise in the image, providing a basis for subsequent noise feature extraction. Step 2.3: Input the noisy residual map N into the pre-trained WideResNet-101 backbone network. WideResNet-101, leveraging its deep convolutional structure and pre-trained feature extraction capabilities, performs feature mining on the noisy residual map. During the network's forward propagation, through multi-layer convolution and pooling operations, it extracts deep-level features F related to noise from the noisy residual map. N These features include information such as the frequency and distribution pattern of the noise, and can be represented as... F N =WideResNet-101(N) This lays the foundation for subsequent evaluation of noise perception quality in conjunction with aesthetic features.
4. The low-light enhancement image quality evaluation method based on image aesthetics-noise dual-branch as described in claim 1, characterized in that, Step 3 specifically includes the following steps: Step 3.1: To achieve refined semantic modulation of noise features and better integrate aesthetic and noise perception information, the noise features F extracted from the noise perception branch are first... N Aesthetic features F obtained from the aesthetic perception branch A As input. Since the initial channel dimensions of the two may differ, a specially designed linear projection layer is used to transform the noise features F... N and aesthetic features F A By mapping each feature to the same target channel dimension C, we obtain noise features with unified dimensions. and aesthetic features The mathematical expression can be described as Among them W N W A Let b be the weight matrix of the linear projection. N b A This is a bias term. Step 3.2: Targeting the aesthetic features after dimensional unification The vector is then transformed again using a linear mapping operation and connected to the GELU (Gaussian Error Linear Units) activation function to generate a dynamically adjustable gated vector G. The linear mapping aims to adapt the aesthetic features to different dimensions and transform information, while the GELU activation introduces non-linearity, allowing the gated vector to adaptively learn and modulate weights based on the feature content. The specific calculation is as follows: Here W G It is a linear mapping weight, b G For bias. Step 3.3: After obtaining the dynamic gating vector G, based on the logic of channel-aware cross-modulation, use this gating vector to apply aesthetic features. Element-level modulation is performed. Element-level modulation involves multiplying the aesthetic features element-wise with the gating vector, allowing the aesthetic features to be adaptively enhanced or suppressed at different channel positions based on the gating information. Mathematically, this is represented as... Here, ⊙ represents element-wise multiplication, thereby achieving feature modulation based on aesthetic semantics. Step 3.4: To preserve the basic information of the original noise features and avoid information loss during modulation, a residual connection mechanism is introduced. This mechanism connects the aesthetic features after element-level modulation. Compared with the original noise characteristics The fusion process involves adding the two components along the channel dimension to obtain the aesthetically modulated noise characteristics. Right now Residual connections can effectively transfer gradients and improve the stability of network training. At the same time, they allow the modulated noise features to incorporate both aesthetic and semantic information and retain the key content of the original noise features, providing a more discriminative feature representation for subsequent quality score prediction.
5. The low-light enhancement image quality evaluation method based on image aesthetics-noise dual-branch as described in claim 1, characterized in that, Step 4 specifically includes the following steps: Step 4.1: First, analyze the aesthetically modulated noise features generated in Step 3. Perform pooling and splicing operations consistent with the aesthetics branch. Specifically, first... Perform Global Average Pooling (GAP) and Global Standard Deviation Pooling (GSP): Global Average Pooling calculates the mean of the feature across the spatial dimension. To capture the overall distribution trend of noise features; to calculate the standard deviation of features in the spatial dimension through global standard deviation pooling. To characterize the degree of local variation in noise. The mathematical expressions for both are: Where H×W is the spatial dimension of the feature map, and ∈ represents the minimum value to avoid numerical instability during square root extraction. Subsequently, and The features are concatenated along the channel dimension to form a feature vector that integrates the global distribution and local variations of noise. Step 4.2: Next, will The input is a Dual-Convolutional Integral Layer Regression Network (DCHRN). DCHRN first reshapes the feature vector to B×1×D (where D is the feature dimension), and then extracts non-linear features through two 3×3 convolutional layers: the first convolutional layer expands the number of channels to 16, and the second layer expands it to 32, with each layer followed by a PReLU activation function. Subsequently, the features are flattened and regressed through a fully connected layer, outputting a subjective noise-perceived quality score S. N ,Right now: Step 4.3: Combine the aesthetic quality score SA and the noise perception quality score S by weighted summation. N Generate overall quality assessment result S R According to the paper's specifications, the weighting formula is: Where λ1 = 1 and λ2 = 0, that is, the prediction score S of the noise branch. N Used only for loss calculation, it does not actually participate in the weighting process of the overall score, and the final S R By S A It takes the lead in achieving a comprehensive evaluation of the quality of low-light enhanced images.