Low light enhancement image quality evaluation method and device, and computer equipment

By decomposing illumination and reflection components through a dual-branch network and combining multi-level illumination injection and hierarchical difference perception, the problems of insufficient utilization of illumination information and simple feature fusion in existing technologies are solved, and a more accurate low-light image quality assessment consistent with human subjective perception is achieved.

CN120976046APending Publication Date: 2025-11-18CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511102636.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-18

Smart Images

  • Figure CN120976046A_ABST
    Figure CN120976046A_ABST
Patent Text Reader

Abstract

The invention relates to a computer vision technology, in particular to a low-light enhancement image quality evaluation method and device and computer equipment, and the method comprises the steps: forming a multi-scale illumination feature map according to an image illumination component; extracting multi-level original image features from the image; adjusting the illumination feature map to the same size as the corresponding original image features through bilinear interpolation, injecting the original image features at all levels through element multiplication, and calculating the difference between the original image and the enhanced image features at all levels; texture information is extracted from all levels of differences, and all levels of texture information are fused; extracting spatial information from each level of difference, and fusing each level of texture information; fusion information is input into a classifier to obtain labels, and the two labels serve as paired labels to be input into a Bradley-Terry model to be converted into a unified global quality score. According to the method, the model learns the perception difference of the pairwise enhanced image and outputs the label, and the Bradley-Terry model is utilized to simulate the human preference selection process, so that the consistency of the evaluation result and the human subjective judgment is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to computer vision technology, in particular to a low-light enhanced image quality evaluation method and device and computer equipment. BACKGROUND

[0002] There are many ways to enhance low-light images in the prior art, but the enhanced images still cannot match human subjective perception.

[0003] In terms of objective quality evaluation of low-light enhanced images, there are some representative methods. Yang et al. proposed the BEHN index, which integrates enhanced perception, structure preservation and color naturalness features through AdaBoost-RF ensemble fusion; Zhai et al. proposed the LIEQA model, which improves the adaptability to low-light environment with a four-dimensional evaluation system; Wang et al. designed a pair-wise learning network, which uses intrinsic perception feature extraction (INPFE) and impairment perception feature extraction (IMPFE) modules to extract patterns and fuse semantics, and performs better than traditional indexes in complex night scenes by aligning with human preference ranking. However, these methods have obvious shortcomings: most of them rely on enhanced images for feature extraction, and do not fully utilize key information such as illumination; simple feature fusion strategies are used, which limits the feature expression ability and leads to limited model performance; most methods assign absolute scores to each enhanced image, which cannot fully simulate the human "two-choice" preference mechanism. SUMMARY

[0004] In view of the problems of not fully utilizing illumination information, simple feature fusion strategy and difficulty in simulating human preference mechanism in the prior art of objective quality evaluation of low-light enhanced images, the present application proposes a low-light enhanced image quality evaluation method, which evaluates whether the enhanced image of the original low-light image is more suitable for human eye perception based on multi-level illumination injection and hierarchical difference perception. The specific steps include:

[0005] According to the Retinex theory, a double-branch network is used to decompose the input image into illumination components and reflection components;

[0006] The illumination component is input into the baseNet for basic feature extraction, and then a multi-scale illumination feature map is generated through pyramid pooling;

[0007] The image is input into the pre-trained MambaOut network for multi-level feature extraction to obtain multi-level original image features;

[0008] The illumination feature map is adjusted to the same size as the corresponding level feature map of the MambaOut network through bilinear interpolation, and each level of original image feature is injected through element multiplication;

[0009] Obtain the original low-light image and its enhanced image, and obtain the original image features of each level, and calculate the difference of each level feature;

[0010] Each high-resolution difference feature is extracted through a spatial-frequency domain perception branch to obtain texture information, and the extracted texture information is fused through a cross-scale fusion module; each low-resolution difference feature is extracted through a semantic perception module to obtain semantic information, and the extracted semantic information is fused through a cross-scale fusion module;

[0011] The features obtained by cross-scale fusion are input into classification to obtain paired labels, and the obtained paired labels are input into a Bradley-Terry model to convert them into a unified global quality score.

[0012] Compared with the prior art, the present application has the following beneficial effects:

[0013] 1. The present application uses multi-level illumination injection to inject illumination information hierarchically into the feature extraction process, which strengthens the role of brightness information in evaluation, and the hierarchical difference perception separately processes high-level semantic features and low-level texture features, and realizes fine utilization of semantic and texture features through cross-scale fusion, solving the problem of simple feature fusion and insufficient utilization of key information in existing methods, and improving the adaptability of the model to complex low-light scenes. Because low-light image acquisition is prone to degradation, and different low-light conditions have various noise patterns, different illumination levels and complex scene content, this sufficient feature utilization enables the model to better cope with these complex situations.

[0014] 2. The present application uses a pair-wise comparison design to enable the model to learn the perceptual difference of the pair-wise enhanced image and output a preference label, and then uses a Bradley-Terry (B-T) model to convert these pair-wise labels into a unified global quality score, simulating the human preference selection process of "two choices", rather than assigning an absolute score to each enhanced low-light image as in most existing methods. This method is more in line with the human subjective judgment mode, significantly improving the consistency of the evaluation result with human subjective judgment. BRIEF DESCRIPTION OF DRAWINGS

[0015] Figure 1 is a flowchart of the low-light enhanced image quality evaluation method of the present application;

[0016] Figure 2 is a flowchart of the multi-scale measurement feature injection of the present application;

[0017] Figure 3 is a data processing schematic diagram of the layered perception module of the present application. DETAILED DESCRIPTION

[0018] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0019] The present application provides a specific implementation method of a low-light enhanced image quality evaluation method, which specifically comprises the following steps:

[0020] S1, according to the Retinex theory, the input image is decomposed into illumination component and reflection component by using a double-branch network, wherein the reflection branch retains high-frequency texture details through the jump connection between the encoder / decoder blocks of the U-Net style, and performs multi-level feature fusion through bilinear interpolation and element addition to accurately reconstruct the surface reflectance; the illumination branch passes through the cascade up-sampling path, and then passes through the continuous 3x3 convolution and bilinear up-sampling stage, and since there is no jump connection, a smooth illumination map conforming to the physical prior is formed.

[0021] S2, in the feature extraction process, the input illumination map is first subjected to basic feature extraction by baseNet, and then subjected to pyramid pooling to form a feature map, while the original image is subjected to multi-level feature extraction by the pre-trained MambaOut network, which will sequentially pass through the initial convolution down-sampling to obtain high-resolution texture features, and the intermediate unit of gradually down-sampling and expanding the channel dimension. Then, the illumination component is hierarchically injected into the feature map of the MambaOut network through multi-level fusion, a feature pyramid from local to global features is constructed, so as to enhance the sensitivity of the model to brightness changes, and solve the problem that the existing methods mostly rely on enhancing low-light images for feature extraction, and fail to fully utilize available key information such as illumination.

[0022] S3, for the pair of enhanced images, the extracted features are hierarchically processed. For the features related to high-level semantics, the channel alignment function is used to calculate the layer-by-layer difference, and the difference features obtained are sent to the semantic branch. The channel attention weight and global context encoder in this module capture the semantic differences, reweight the key channel features, and use continuous Transformer to model long-range dependencies, focusing on semantic content changes in the image pair, such as systematic object attribute differences or scene structure changes. For features related to low-level texture information, a special spatial-frequency branch is developed to jointly process spatial and frequency domain information. In the frequency domain analysis, the real part spectrum is extracted by applying two-dimensional Fourier transform to the difference features, and deep separable convolution is used to dynamically modulate the frequency response to suppress high-frequency noise; in the spatial domain analysis, spatial attention is used to enhance the local structure sensitive area, and the channel splicing and convolution of the average pooling and maximum pooling results of the difference features are performed, and then the Sigmoid activation function is used to calculate. Through nonlinear fusion of the double-path features extracted from the input image pair, a fine-grained difference response is generated, and then these features are adaptively weighted and fused across scales to produce a preference decision, outputting a pair of preference labels, which solves the problem that the simple fusion strategy in existing methods limits feature expression and leads to limited model performance.

[0023] The embodiment proposes a specific implementation of a low-light enhanced image quality evaluation method, including a complete process of illumination decomposition, multi-level illumination injection, hierarchical difference perception, and global score generation. The network structure parameters and operations involved are reproducible specific settings. As shown in FIG. 1, the embodiment includes the following steps: Figure 1 In the embodiment, the original low-light image I j and the enhanced image I j of the original low-light image are input into the illumination decomposition network to obtain the corresponding illumination components, i.e., the original low-light image I j corresponding illumination component L j , the enhanced image I j of the original low-light image, and the corresponding illumination component L i . j , L j , I i , and L i are input into two weight-shared multi-scale brightness feature injection modules, respectively, to calculate the difference value of each scale of the two images. The high-resolution features are input into the spatial-frequency domain perception branch, and the low-resolution features are input into the semantic perception branch. The outputs of the two branches are input into the classifier to obtain two labels, respectively. The obtained labels are input into the model to obtain the quality score of the enhanced image. Next, the embodiment will explain the process in steps 1-4.

[0024] Step 1: Illumination decomposition.

[0025] Based on the Retinex theory, a double-branch network is used to decompose the input low-light enhanced image into illumination component and reflection component, and the specific structure is as follows:

[0026] The reflection branch adopts a symmetric encoder-decoder structure in the style of U-Net, the encoder includes 4 convolution blocks (each convolution block is composed of 2 3x3 convolution layers, step 1, padding 1, and the activation function is ReLU), the decoder is symmetric with the encoder, and the i-th layer of the encoder and the 4-i-th layer of the decoder are connected by a jump connection (when the feature map sizes are inconsistent, the bilinear interpolation is used to adjust to the same size and then the element addition is performed) to realize multi-level feature fusion. The reflection component is output through the branch , H and W are the height and width of the input image, and the default size in the embodiment is 224x224;

[0027] The illumination branch adopts a cascaded up-sampling path, including 3 up-sampling units, each up-sampling unit is composed of 1 3x3 convolution layer (step 1, padding 1, ReLU activation) and 1 bilinear up-sampling layer (the scaling factor of the up-sampling layer is set to 2), and there is no jump connection to ensure the smoothness of the illumination map. The illumination component is output through the branch ;

[0028] The above decomposition process satisfies input image = illumination component Reflection component, is element multiplication. Step 2: Multi-level illumination injection.

[0029] The embodiment takes the image I i as an example to illustrate the multi-level illumination injection process, and the illumination component I obtained in step 1 is injected hierarchically.

[0030] The illumination feature is extracted from the illumination component I , including: inputting the illumination feature I into baseNet for basic feature extraction, baseNet is composed of 2 3x3 convolution layers (step 1, padding 1) and 1 max pooling layer (2x2, step 2); then multi-scale illumination feature maps are generated through pyramid pooling, the pyramid pooling includes 4 pooling scales, which are 1x1, 2x2, 4x4 and 8x8 respectively, and after each pooling size is pooled, the channel is compressed to 64 through 1x1 convolution to obtain the injection feature map;

[0031] The original image feature is extracted from the input image, and the pre-trained MambaOut network is used for feature extraction in the embodiment, and the MambaOut network adopts four-level feature extraction, specifically including:

[0032] MambaOut stage 1 includes 1 3x3 convolutional layer (stride 2, padding 1), the output channel of MambaOut stage 1 is 64, the features of image I i are extracted by MambaOut stage 1, the injection head 1 is used to match the size of the injection feature map with the feature map obtained by MambaOut stage 1, and then the Hadamard product is performed on the two feature maps to obtain the first-level feature of image I i ;

[0033] MambaOut stage 2 includes 1 3x3 convolutional layer (stride 2, padding 1), the channel of MambaOut stage 2 is expanded to 128, the first-level feature of image I i is used as the input of MambaOut stage 2, the injection head 2 is used to match the size of the injection feature map with the feature map obtained by MambaOut stage 2, and then the Hadamard product is performed on the two feature maps to obtain the second-level feature of image I i ;

[0034] MambaOut stage 3 includes 1 3x3 convolutional layer (stride 2, padding 1), the channel of MambaOut stage 2 is expanded to 256, the second-level feature of image I i is used as the input of MambaOut stage 3, the injection head 3 is used to match the size of the injection feature map with the feature map obtained by MambaOut stage 3, and then the Hadamard product is performed on the two feature maps to obtain the third-level feature of image I i ;

[0035] MambaOut stage 4 includes 1 3x3 convolutional layer (stride 2, padding 1), the channel of MambaOut stage 2 is expanded to 512, the third-level feature of image I i is used as the input of MambaOut stage 4, the injection head 4 is used to match the size of the injection feature map with the feature map obtained by MambaOut stage 4, and then the Hadamard product is performed on the two feature maps to obtain the fourth-level feature of image I i ;

[0036] As an optional embodiment, the injection head adjusts the illumination feature map to the same size as the corresponding level feature map of MambaOut by bilinear interpolation, and injects the first to fourth level feature maps by element multiplication to obtain the multi-level image features of image I i ; , represents the k-level feature of image I i ; similarly, the multi-level image features of image I j can be obtained , represents image I jThe k-th level feature.

[0037] Step 3: Hierarchical difference perception.

[0038] The input of this invention is the original low-light image I. j and its enhanced diagram I i The image pairs formed by the images are processed through steps 1 and 2 to obtain image I. j Multi-level image features and image I i Multi-level image features Perform stratified difference processing, such as Figure 3 ,include:

[0039] First, the images are divided into a high-resolution set and a low-resolution set, each set including at least one image feature. Those skilled in the art perform the set division according to the resolution. In this embodiment, four-level feature extraction is used, with the first and second level image features assigned to the high-resolution set and the third and fourth level image features assigned to the low-resolution set.

[0040] For the features of the high-resolution set, the difference of each level of features is calculated separately, i.e. (k={1,2}) Represents the original low-light image Its corresponding enhanced image The difference in the k-th level features between them;

[0041] right (k={1,2}) Perform a two-dimensional Fourier transform to extract the real part spectrum, and obtain the frequency domain features by modulating the frequency response through a 3×3 depth separable convolution with 128 output channels;

[0042] right (k={1,2}) undergo 2×2 average pooling and max pooling. The pooling results are concatenated along the channel dimension and then used to generate spatial attention weights through a 3×3 convolution with 1 output channel and sigmoid activation. Spatial attention weights and Perform element-wise multiplication to obtain the spatial domain characteristics;

[0043] After unifying the frequency domain and spatial domain features through a 1×1 convolution and then adding the elements together, texture difference features are obtained.

[0044] The fusion is achieved through adaptive weights (calculated based on feature variance, with a weight sum of 1). , The obtained features are then fed into a binary classifier (a 2-layer multilayer perceptron MLP with a 2-dimensional output) to generate the first preference label. (when A value of 1 indicates the original low-light image. Superior quality, when A value of 2 indicates that the enhanced image corresponds to the original low-light image. Superior quality;

[0045] For the characteristics of low-resolution sets, alignment is performed using a channel alignment function, and then the differences are calculated. The alignment function used in this embodiment is... A 1×1 convolution is used to unify the channels of the feature map to 256 dimensions before calculating the differences. (k={3,4});

[0046] right Global average pooling is performed on k={3,4}. A 256-dimensional weight vector is generated using a cascaded multilayer perceptron (MLP) with 64-dimensional hidden layers and sigmoid activation. This weight vector is then used in conjunction with... Perform element-wise multiplication (k={3,4}) to obtain the weighted features;

[0047] We use a cascaded two-layer Transformer (each layer includes 8 attention heads and 512 hidden dimensions) to extract long-term dependencies from weighted features to obtain semantic differential features;

[0048] The fusion is achieved through adaptive weights (calculated based on feature variance, with a weight sum of 1). , The obtained features are then fed into a binary classifier (a 2-layer multilayer perceptron MLP with a 2-dimensional output) to generate a second preference label. (when A value of 1 indicates the original low-light image I. j Superior quality, when A value of 2 indicates that the enhanced image I corresponds to the original low-light image. i Superior quality.

[0049] Step 4: Bradley-Terry model.

[0050] The Bradley-Terry model is a commonly used probability model for comparing the relative abilities or preferences between two or more items, in which each item is assigned a capability parameter representing its relative ability size, and by observing the comparison results between items, these capability parameters can be estimated and predictions can be made, for example, the BradleyTerry2 package developed by Tencent Cloud can handle binary comparison data, that is, each comparison result involves only two items. The present application inputs the first preference label and the second preference label into the Bradley-Terry (B-T) model as a label pair, and converts these paired labels into a unified global quality score through the B-T model. Specifically, for each scene, first construct a preference matrix, where each entry represents the empirical probability that image i is better than image j according to the collected comparison results. The B-T model estimates the underlying quality scores by assuming that the preference probability follows a logistic form, and these scores are optimized by maximum likelihood estimation to ensure global consistency for all comparisons. Finally, the resulting scores are normalized to obtain the final quality evaluation metric. In this way, the human "two-choice" preference mechanism is simulated, replacing the traditional evaluation method of assigning an absolute score to each enhanced low-light image, and the problem that existing methods cannot fully simulate the human preference mechanism is solved.

[0051] Table 1 and Table 2 show the performance of the present application and other existing image quality evaluation methods on the LE dataset and the RNTIEQA dataset. The present embodiment evaluates the effectiveness of the model from two dimensions: prediction accuracy (Pearson Linear Correlation Coefficient, PLCC) and monotonicity (Spearman Rank-Order Correlation Coefficient, SRCC).

[0052] Table 1

[0053] SRCC PLCC NIQE 0.0632 0.0828 BRISQUE 0.4689 0.3169 CNNIQA 0.5211 0.5349 DBCNN 0.6726 0.6903 HyperIQA 0.7772 0.7851 VCRNet 0.6514 0.6526 StairIQA 0.7575 0.7718 VIPNet 0.6846 0.6708 TemPQT 0.7018 0.7115 PINet 0.7822 0.7822 AGAIQA 0.8127 0.8021 SGDNet 0.7817 0.7825 The present invention 0.8144 0.8208

[0054] Table 2

[0055] SRCC PLCC NIQE 0.0094 0.0061 IL-NIQE 0.0245 0.0162 NI-QMC 0.0576 0.0377 SNP-NIQE 0.1333 0.088 PIQE 0.2058 0.1376 NUIQ 0.3826 0.2656 NBIQA 0.4342 0.3008 BlIInds-II 0.4391 0.3041 BRISQUE 0.464 0.3209 FRIQUEE 0.475 0.3296 CNNIQA 0.5425 0.3851 VCRNet 0.6308 0.4513 DBCNN 0.6491 0.4684 The present invention 0.6696 0.6722

[0056] Among them, the Natural Image Quality Evaluator (NIQE) comes from the paper “Making a “Completely Blind” Image Quality Analyzer”; the blind / referenceless image spatial quality evaluator BRISQUE comes from the paper “BLIND / REFERENCELESS IMAGE SPATIAL QUALITY EVALUATOR”; the convolutional neural network for no-reference image quality assessment CNNIQA comes from the paper “Convolutional Neural Networks for No-Reference Image Quality Assessment”; the application of a deep bilinear convolutional neural network DBCNN in image quality assessment comes from the paper “Blind Image Quality Assessment Using a Deep Bilinear Convolutional Neural Network”; the quality evaluation method HyperIQA comes from the paper “Blindly Assess Image Quality in the Wild Guided by A Self-Adaptive Hyper Network”; the deep learning model VCRNet for no-reference image quality assessment comes from the paper “VCRNet: Visual Compensation Restoration Network for No-Reference Image Quality Assessment”; StairIQA is an image quality evaluation tool; the image quality evaluation method VIPNet comes from the paper “Visual Interaction Perceptual Network for Blind Image Quality Assessment”; the image quality evaluation method TemPQT comes from the paper “Blind Image Quality Assessment via Transformer Predicted Error Map and Perceptual Quality Token”; AGAIQA comes from the paper “Blind Image Quality Assessment via Adaptive Graph Attention”;Image quality evaluation method SGDNet is from the paper "Blind image quality index with high-level Semantic Guidance and low-level fine-grained Representation"; image quality evaluation method IL-NIQE is from the paper "A Feature-Enriched Completely Blind Image Quality Evaluator"; image quality evaluation method NI-QMC is from the paper "No-Reference Quality Metric of Contrast-Distorted Images Based on Information Maximization"; image quality evaluation method SNP-NIQE is from the paper "Unsupervised Blind Image Quality Evaluation via Statistical Measurements of Structure, Naturalness, and Perception"; image quality evaluation method PIQE is from the paper "BLIND IMAGE QUALITY EVALUATION USING PERCEPTION BASED FEATURES"; image quality evaluation method NUIQ is from the paper "Underwater Image Enhancement Quality Evaluation: Benchmark Dataset and Objective Metric"; image quality evaluation method NBIQA is from the paper "A NOVEL BLIND IMAGE QUALITY ASSESSMENT METHOD BASED ON REFINED NATURAL SCENE STATISTICS"; image quality evaluation method BlIInds-II is from the paper "Blind Image Quality Assessment: A Natural Scene Statistics Approach in the DCT Domain"; image quality evaluation method BRISQUE is from the paper "BLIND / REFERENCELESS IMAGE SPATIAL QUALITY EVALUATOR";Image quality assessment method FRIQUEE is from the paper “Perceptual Quality Prediction on Authentically Distorted Images Using a Bag of Features Approach”.

[0057] Table 1 and Table 2 are respectively the PLCC, SRCC indicators of the present application and other existing image quality assessment models on the LE dataset and the RNTIEQA dataset. As shown in Table 2, on the RNTIEQA dataset, the SRCC of the method of the present application is 0.6696, and the PLCC is 0.6722, while the SRCC of other methods such as NIQE is only 0.0094, and the PLCC is 0.0061, the SRCC of BRISQUE is 0.464, and the PLCC is 0.3209, and even the better-performing DBCNN, its SRCC is only 0.6491, and the PLCC is 0.4684; as shown in Table 1, on the RNTIEQA dataset, the SRCC of the present application is 0.6696, and the PLCC is 0.6722, which is better than methods such as CNNIQA (SRCC is 0.5425, and PLCC is 0.3851), VCRNet (SRCC is 0.6308, and PLCC is 0.4513), and more accurately reflects the subjective quality.

[0058] Ablation experiments verify the effectiveness of each module by sequentially removing or modifying specific components under consistent experimental settings. When step 2 (i.e., multi-level illumination injection) is removed, MambaOut is used directly for original feature extraction, resulting in a slight decrease in SRCC and KRCC scores and a significant decrease in PLCC. When step 3 (hierarchical difference perception) is removed, the multi-stage features obtained from MambaOut are concatenated and then processed through global average pooling and input into the MLP layer for preference label prediction. The result is a decrease in all indicators, demonstrating the key role of the two steps in improving evaluation performance.

[0059] Although embodiments of the present application have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and alterations can be made thereto without departing from the principles and spirit of the present application, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for evaluating the quality of low-light enhanced images, characterized in that, The evaluation of whether an enhanced image of the original low-light image is more suitable for human visual perception than the original low-light image, based on multi-level illumination injection and hierarchical difference perception, specifically includes the following steps: Based on Retinex theory, a dual-branch network is used to decompose the input image into illumination components. The lighting components are input into baseNet for basic feature extraction, and then multi-scale lighting feature maps are generated through pyramid pooling. The image input is processed through a pre-trained MambaOut network to extract features at multiple levels, resulting in multi-level original image features. The illumination feature map is adjusted to the same size as the feature map of the corresponding layer of the MambaOut network through bilinear interpolation, and then the original image features at each level are injected through element-wise multiplication. Obtain the original low-light image and its enhanced image at each level of original image features, and calculate the differences between the features at each level; Texture information is extracted from each high-resolution differential feature through a spatial-frequency domain sensing branch, and then the extracted texture information is fused through a cross-scale fusion module; semantic information is extracted from each low-resolution differential feature through a semantic sensing module, and then the extracted semantic information is fused through a cross-scale fusion module. The features obtained from cross-scale fusion are input into the classification system to obtain paired labels. The paired labels are then input into the Bradley-Terry model to convert them into a unified global quality score.

2. The low-light enhancement image quality assessment method according to claim 1, characterized in that, The process of using a dual-branch network to decompose the input image into illumination and reflection components includes: Surface reflectivity is reconstructed using a U-Net module, in which skip connections are used between the encoder and decoder, and multi-level feature fusion is performed through bilinear interpolation and element-wise addition. A smooth illumination map is obtained by cascading upsampling paths, which consist of cascaded 3×3 convolutions and bilinear upsampling.

3. The low-light enhancement image quality assessment method according to claim 1, characterized in that, The baseNet consists of two 3×3 convolutional layers with a stride of 1 and padding of 1, and a 2×2 max pooling layer with a stride of 2, cascaded together. The pyramid pooling includes four pooling sizes cascaded in sequence: 1×1, 2×2, 4×4, and 8×8. After each pooling, the channels are compressed to 64 by a 1×1 convolution.

4. The low-light enhancement image quality assessment method according to claim 3, characterized in that, The MambaOut network consists of four cascaded 3×3 convolutional layers with a stride of 2 and padding of 1. The number of channels in the first to fourth levels are 64, 28, 256, and 512, respectively.

5. The low-light enhancement image quality assessment method according to claim 1, characterized in that, The calculation of the difference of the i-th level feature includes: unifying the i-th level feature of the original low-light image and the i-th level feature of the corresponding enhanced image of the original low-light image to 256 channels through a channel alignment function, and then taking the difference between the features of the unified original low-light image and its enhanced image as the difference of the i-th level feature.

6. The low-light enhancement image quality assessment method according to claim 1, characterized in that, Extracting texture information through the spatial-frequency domain perceptual branch includes: Two-dimensional Fourier transforms are performed on each differential feature to extract the real part spectrum. The frequency domain features are obtained by outputting a 3×3 depth-separable convolutional modulation frequency response with 128 output channels. Each differential feature is subjected to 2×2 average pooling and max pooling respectively. After the pooling results are concatenated, they are sequentially passed through 3×3 convolution and Sigmoid activation of output channel 1 to generate spatial attention weights. The spatial attention weights are then used to weight the corresponding differential features to obtain spatial features. After unifying the frequency domain and spatial domain features through a 1×1 convolution and then adding the elements together, texture difference features are obtained. The texture difference features corresponding to each difference feature are fused together using adaptive weights with a weight sum of 1 to obtain texture information.

7. The low-light enhancement image quality assessment method according to claim 1, characterized in that, Semantic information extracted through the semantic awareness module includes: Global average pooling is performed on the differential features at each level. After pooling, a weight vector is generated by a two-layer cascaded multilayer perceptron. The corresponding differential features are then weighted by element-wise multiplication using the weight vector. The weighted differential features are input into two cascaded Transformer layers to extract long-term dependency features as the semantic features of the current differential features; The semantic information is obtained by fusing all the different semantic features through an adaptive weight that sums to 1.

8. A low-light enhanced image quality assessment device, characterized in that, To implement the low-light enhancement image quality assessment method of claim 1, the method includes: A dual-branch network is used to decompose the input image into illumination and reflection components; The system employs a weight-sharing approach, comprising a first multi-scale metric feature injection and a second multi-scale metric feature injection. The first multi-scale metric feature injection processes the original low-light image, while the second multi-scale metric feature injection processes the enhanced image corresponding to the original low-light image. Each multi-scale metric feature injection includes a multi-scale illumination feature map extraction module and a pre-trained MambaOut network. The multi-scale illumination feature map extraction module includes baseNet and pyramid pooling. baseNet extracts basic features from the illumination components, and pyramid pooling generates multi-scale illumination feature maps based on the basic features. The pre-trained MambaOut network performs multi-level feature extraction on the input image to obtain multi-level original image features. The illumination feature map is then adjusted to the same size as the corresponding level feature map of the MambaOut network using bilinear interpolation. Element-wise multiplication is used to inject the original image features at each level, and the fused features are used as the output of the multi-scale metric feature injection. The feature difference module is used to calculate the differences in various scale features between the original low-light image and its enhanced image; Spatial-frequency domain aware branch is used to extract texture information from high-resolution differences; A semantic awareness module is used to extract semantic information from low-resolution differences; A classifier is used to classify images based on their texture and semantic information, and to use the classification results of a set of images as paired labels. The Bradley-Terry model is used to convert pairwise labels into a uniform global quality score.

9. A computer device, characterized in that, The computer device includes: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the low-light enhancement image quality assessment method as described in claim 1.