Low-light image enhancement method based on YUV color space

Through the CSLNet network based on YUV color space, the brightness and chromaticity channels are separated for targeted processing, which solves the problems of color distortion and noise amplification in low-light image enhancement, achieves the improvement of brightness and details, and improves the visual effect of the image.

CN120339099APending Publication Date: 2025-07-18CHONGQING UNIV OF TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510462247.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18

Smart Images

  • Figure CN120339099A_ABST
    Figure CN120339099A_ABST
Patent Text Reader

Abstract

The invention discloses a low-light image enhancement method based on a YUV color space, and relates to the technical field of image processing. According to the method, a color space illumination enhancement network (CSLNet is established, brightness (Y) and chrominance (U, V) information is separated, and brightness and chrominance channels are respectively subjected to targeted processing, so that the image brightness and details are improved, the problems of color distortion and noise amplification are avoided, the CSLNet makes full use of the advantages of a YUV color space, and the image quality is improved. Wherein the Y channel represents brightness information, the U and V channels represent chrominance information, brightness enhancement and color correction can be more accurately controlled through separation processing and the CSLNet, a color space enhancement module CSE Block (Color Space Enhancement Block) is introduced into the CSLNet, the CSE Block is fused with a U-shaped network structure and a multi-head self-attention mechanism, and optimization processing is specially carried out on chrominance channels U and V of a low-illumination image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a low-light image enhancement method based on a YUV color space. Background Art

[0002] Due to force majeure factors such as insufficient natural light, performance limitations of photographic equipment, and differences in operating skills, the quality of images taken in low-light environments is often unsatisfactory. For example, a large amount of image details are easily lost, making it difficult to identify the originally rich textures and tiny structural features; at the same time, the noise level is significantly enhanced, and these random pixel interferences not only destroy the clarity of the image, but may also cover up important visual information; in addition, the accurate presentation of colors is also seriously affected, and color distortion is common, resulting in a large deviation between the image color and the actual scene. These factors work together to greatly reduce the recognition of image content, bringing many inconveniences to subsequent image processing, analysis, and application.

[0003] Compared with methods that directly learn enhancement results in an end-to-end network, deep Retinex-based methods have attracted much attention due to their physical interpretability. Based on the Retinex theory, these methods decompose the image into illumination and reflection components and enhance them separately, thereby effectively solving the problems of uneven illumination and loss of details. For example, RetinexNet contains two modules, Decom-Net decomposes the image into reflection and illumination components to separate the influence of illumination; Enhance-Net optimizes the illumination component to achieve enhancement. This modular design improves the interpretability and flexibility of the system, and is improved by introducing new constraints and network designs to better adapt to different lighting conditions. To reduce the computational burden, Li et al. proposed a lightweight model LightenNet consisting of only four layers to quickly estimate the illumination map suitable for real-time enhancement. By dividing the input image by the illumination map, an enhanced image can be obtained, which effectively eliminates the influence of uneven illumination. Zhang et al. developed the KinD architecture, which contains three sub-networks to handle layer decomposition, reflectivity recovery, and illumination adjustment respectively. The improved KinD++ introduces a multi-scale illumination attention module to further improve its adaptability to complex illumination. The Retinex-based method improves the interpretability and flexibility of the enhancement effect by decomposing and processing the illumination and reflection components separately, especially showing significant advantages in dealing with problems such as uneven illumination, loss of details and noise interference, and can meet the enhancement needs in different scenarios.

[0004] The Transformer architecture was initially successful in the field of natural language processing as it can effectively capture the global dependencies in sequential data, and then was introduced into the field of computer vision. The Vision Transformer (ViT) treats an image as a sequence of patches, thus improving the performance of tasks such as image classification and object detection. In the field of image restoration, Restormer improved the design of ViT, abandoned the segmentation method of non-overlapping image patches, retained the spatial continuity of the image, and reduced the computational complexity. LLFormer applied the Transformer model to low-light image enhancement, adopted an axis-based multi-head self-attention mechanism, and further simplified the model complexity. The Transformer architecture brought new ideas to the field of low-light image enhancement and provided new directions for other tasks in the field of computer vision. Retinexformer proposed the first Transformer algorithm combining the Retinex theory. The algorithm formulates a one-stage Retinex-based framework (ORF), modifies the original Retinex model by introducing a perturbation term, and the one-stage framework estimates the illumination information and uses it to illuminate the low-light image. Then, the Illumination-Guided Transformer (IGT) is used to suppress noise, artifacts, exposure problems, and color distortion. Different from previous Retinex-based deep learning frameworks, ORF is trained in a one-stage end-to-end manner. The key of IGT lies in the illumination-guided multi-head self-attention mechanism, which uses the illumination representation to guide the self-attention calculation and enhances the interaction between regions with different exposure levels.

[0005] Restoring an image to a color effect that conforms to human visual perception is as challenging as enhancing the image brightness. A color image usually consists of three channels: red (R), green (G), and blue (B). Due to the limitation of available image information, accurately restoring the data of these three channels simultaneously is a daunting task. In addition, most deep learning networks use L1 or L2 loss functions during training, and these functions mainly focus on the accurate correspondence restoration of image pixels, often ignoring the mutual correlation between channels. Therefore, there are problems of image color distortion in the images restored by many methods.

[0006] In the low-light image enhancement task, although many algorithms can effectively improve the image brightness, compared with the reference image, there are color distortion problems in some images, which affect the visual effect and subsequent tasks. The YUV color space separates the brightness information and chrominance information of the image, where Y represents the brightness component, and U and V represent the chrominance components. This separation enables separate processing of brightness and chrominance when enhancing the image, avoiding the problem of mutual coupling between brightness distortion and noise in the RGB color space, thus more precisely controlling color changes. However, attention still needs to be paid to the possible color deviation problems in the image after enhancement processing. Summary of the Invention

[0007] The purpose of the present invention is to provide a low-light image enhancement method based on the YUV color space to solve the possible color deviation problems in the image after enhancement processing using the YUV color space, and further optimize the brightness and detail performance of the image.

[0008] To achieve the above object, the present invention provides the following technical solution: A low-light image enhancement method based on the YUV color space, at least including the following steps:

[0009] S1: Build a color space illumination enhancement network, the color space illumination enhancement network is the CSLNet, and the CSLNet includes but is not limited to a Transformer enhancement module, a color space enhancement module, and a fusion module. The Transformer enhancement module is the TransformerBlock, the color space enhancement module is the CSE Block, and the fusion module is the Fusion Block;

[0010] S2: Input an RGB three-channel image into the CSLNet, perform image preprocessing and channel separation to obtain a Y channel, a U channel, and a V channel;

[0011] S3: For the Y channel, perform brightness channel enhancement processing;

[0012] S4: For the U channel and the V channel, perform chrominance channel enhancement processing;

[0013] S5: Perform feature fusion and multi-level feature extraction through the Fusion Block;

[0014] S6: Finally, perform image reconstruction and noise reduction to complete low-light image enhancement.

[0015] Further, the S2 at least includes the following steps:

[0016] Convert the RGB three-channel image input to the CSLNet into a YUV three-channel image. This conversion provides a basis for subsequent processing of luminance and chrominance information respectively. Decomposing the image into luminance (Y) and chrominance (U, V) channels can more accurately perform targeted processing on different information, avoiding the problem of luminance distortion and noise coupling in the RGB color space, thereby more effectively improving image quality and color fidelity. Refer to the following formula:

[0017] Y, U, V = RGB to YUV(I input )

[0018] In the formula, RGB to YUV represents the operation of converting the image from RGB format to YUV format.

[0019] Furthermore, the S3 at least includes the following steps:

[0020] First, apply a 3×3 convolutional layer to extract the preliminary features of the image. The preliminary features are the basic elements, which include but are not limited to edges and textures;

[0021] Subsequently, introduce non-linearity through the ReLU activation function so that the network can learn more complex patterns;

[0022] Then, use the max pooling layer to reduce the spatial size of the feature map and appropriately increase the number of channels at the same time to retain key features and reduce redundant information;

[0023] Subsequently, use the Transformer Block self-attention mechanism to capture the long-range dependencies between features, which can model the relationships between different regions, and thus better understand the global structure of the image. The Transformer Block can more comprehensively understand the image content, especially under low-light conditions, and can more effectively restore the details and texture information of the image;

[0024] The upsampling step is to restore the feature map processed by pooling and the Transformer Block to the original spatial size, ensuring that the enhanced image is consistent with the input image in spatial resolution and avoiding detail loss caused by resolution reduction. The specific operation is as follows:

[0025] Y′ = ReLU(Conv3x3(Y))

[0026] Y″ = MaxPool(Y′)

[0027] Y″′ = Transformer(Y″)

[0028] Y″′ = UpSample(Y″′) + Y

[0029] In the formula, Transformer represents being processed by the Transformer Block; UpSample represents performing the upsampling operation.

[0030] Furthermore, the S4 at least includes the following steps:

[0031] Firstly, it is processed by the CSE Block. The CSE Block integrates the U-shaped network structure and the multi-head self-attention mechanism, and is specifically optimized for the chrominance channel of low-light images.

[0032] Through the multi-scale feature extraction and skip connections of the U-shaped network, the CSE Block can effectively retain the detailed information of the image. At the same time, by using the multi-head self-attention mechanism, it enhances the model's ability to understand complex spatial structures, accurately restores the chrominance details and suppresses noise. This multi-dimensional design not only solves the limitations of traditional methods in processing two-dimensional visual data, but also significantly improves the effect of low-light image enhancement, especially in reducing color distortion and noise.

[0033] The processed chrominance information is then input into the CSE Block for color information fusion processing. The specific operation is as follows:

[0034] Firstly, perform convolution processing on the chrominance information:

[0035] F = CR n (F in )

[0036] CR n represents performing n times of Conv3×3 + ReLU operations, where n is set to 3; that is, perform three 3×3 convolution operations on F in and apply the ReLU activation function after each convolution. The purpose is to extract more complex feature information, especially in terms of chrominance information, to help the model capture more detailed information of the image.

[0037] Transformer-Block processing:

[0038] F′ = Transformer(F)

[0039] The feature map F after convolution processing is input into the Transformer Block;

[0040] Perform the upsampling operation to increase the size of the image:

[0041] F up = UpSample(F′)

[0042] Convolution fusion operation:

[0043] F″ = Conv3x3(F up ) + F m

[0044] The feature map F after upsampling up will undergo a 3×3 convolution operation to obtain a new feature map, and then, this feature map is added to the input F m for the purpose of fusing the high-level features after upsampling with the low-level detailed information, thereby enhancing the detailed performance of the image;

[0045] Final output:

[0046] F out ' = Tanh(Conv3x3(F″))

[0047] F″ after convolution processing will pass through a 3×3 convolutional layer and apply the Tanh activation function.

[0048] Furthermore, the Fusion Block combines the edge texture enhancement mechanism with the multi-path feature fusion strategy and cooperates with the multi-level feature extraction module to achieve better enhancement and extraction effects;

[0049] The edge texture enhancement mechanism is based on the edge texture enhancement block, which is the ETE Block;

[0050] The multi-path feature fusion strategy is to introduce the Laplacian operator and the Sobel operator, so as to be able to extract the edge and texture features of the image respectively, effectively capture the key detailed information of the image, and the dual-path processing strategy using the Laplacian operator and the Sobel operator solves the deficiency of the traditional single feature extraction method in detail retention, enabling the model to comprehensively integrate the complementary information from different input paths;

[0051] The Fusion Block uses learnable convolutional layers and activation functions to perform non-linear transformation on the features, further optimizing the feature combination and enhancing the network's ability to understand complex scenes;

[0052] The enhanced luminance information and the fused chrominance information are sent to the multi-level feature extraction module to screen and fuse the luminance and chrominance features, retaining more effective feature information and further enhancing the detailed performance of the image. The specific operations are as follows:

[0053] First feature fusion:

[0054] F u ″ / F v ' = LReLU(Conv1×3(F u / F v′))

[0055] Process the F u / F v ′ feature through a 3×3 convolution operation, and then apply the Leaky ReLU activation function, aiming to extract more effective features through convolution and maintain a certain non-linearity through the activation function;

[0056] Second feature fusion:

[0057] F u ″ / F v ″ = F u ′ / F v ′ + ETE(F u ′ / F v ′)

[0058] In this operation, directly add the results of F u ′ / F v ′ and F u ′ / F v ′ after being processed by the ETE Block. The ETE Block provides further feature enhancement and contains detailed information;

[0059] Third feature fusion:

[0060] F u ″ / F v ″′ = F u ′ / F v ′ + F u ″ / F v ″ + ETE(F u ″ / F v ″)

[0061] Combines the original F u ′ / F v ′, F u ″ / F v ″ after the second feature fusion, and F u ″ / F v ″ after being processed by ETE, further strengthening the detailed and effective feature information;

[0062] Final feature output:

[0063] F out ″ = CDoncat(F u ″, F v ″, F v ″′)

[0064] Wherein, Concat represents the concatenate operation; ETE represents being processed by the ETE Block.

[0065] Further, the specific operations of the ETE Block at least include the following steps:

[0066] Laplacian operation:

[0067] L = Laplacian(F m ) + F m

[0068] Laplacian represents the Laplacian Operator processing, which can extract the detailed information in the image. By performing the Laplacian Operator processing on the feature map F of the input image m and adding the processed F after Laplacian processing m to the original F m can enhance the detailed features of the image;

[0069] Sobel operation:

[0070] S = Sobel(F m )

[0071] Sobel represents being processed by the Sobel Operator. By performing the Sobel Operator processing on F m can extract the edge details of the image;

[0072] Fusion processing of Laplacian and Sobel:

[0073] L′ / S′ = LReLU(Conv1×1(L / S))

[0074] Performing 1×1 convolution on the results of Laplacian and Sobel operations and applying the Leaky ReLU activation function to further fuse these two features;

[0075] Deeper convolution operation:

[0076] L″ / S″ = LReLU(Conv3×3(L′ / S′))

[0077] Applying 3×3 convolution and the Leaky ReLU activation function to the result of L′ / S′ to further extract features and enhance details;

[0078] Final Laplacian and Sobel features:

[0079] L″′ = LReLU(Conv3×3(L″ + L′)) + L′ + L″

[0080] S″′ = LReLU(Conv3×3(S″ + S′))

[0081] Output of the ETE Block:

[0082] F out ″′ = F m + L″′ + S″′.

[0083] Furthermore, the S6 at least includes the following steps:

[0084] Fuse the processed luminance and chrominance information, and combine the fused image with the original image to reduce the impact caused by the loss of some original information during the processing.

[0085] Then, perform noise reduction processing on the output image through the enhanced image denoising module to further optimize the image quality.

[0086] The enhanced image denoising module is the EID Block. The enhanced image denoising module uses the method of dense residual connection, effectively removing noise while retaining the important details of the image, and significantly improving the visual quality of the image. The specific operations are as follows:

[0087] Concat operation:

[0088] F concat = Concat(Y″′, UV)

[0089] Convolution and addition with the original input low-light image:

[0090] I enhanced = Conv3×3(F concat ) + I low

[0091] where I low is the original input low-light image;

[0092] EID Block processing:

[0093] I out = EID(I enhanced )

[0094] where EID represents processing through the EID Block.

[0095] Compared with the prior art, the beneficial effects of the present invention are:

[0096] 1. The present invention constructs a Color Space Lightenhancement Network (abbreviated as CSLNet). By separating the luminance (Y) and chrominance (U, V) information, and performing targeted processing on the luminance and chrominance channels respectively, it can improve the image luminance and details while avoiding the problems of color distortion and noise amplification. Moreover, CSLNet makes full use of the advantages of the YUV color space. The Y channel represents luminance information, and the U and V channels represent chrominance information. Through separate processing, CSLNet can more precisely control luminance enhancement and color correction;

[0097] 2. The present invention introduces a Color Space Enhancement Block (CSE Block) into CSLNet. The CSE Block combines a U-shaped network structure with a multi-head self-attention mechanism, and specifically optimizes the chrominance channels U and V of low-light images. Through the multi-scale feature extraction and skip connections of the U-shaped network, the CSE Block can effectively retain the detail information of the image. At the same time, the multi-head self-attention mechanism is used to enhance the model's understanding ability of complex spatial structures, accurately restore chrominance details and suppress noise. In addition, a simple and efficient upsampling method further improves the efficiency of feature map restoration and reduces the computational complexity. This multi-dimensional design not only solves the limitations of traditional methods in processing two-dimensional visual data, but also significantly improves the effect of low-light image enhancement, especially in reducing color distortion and noise, providing an efficient and accurate solution for the low-light image enhancement task;

[0098] 3. The present invention further introduces a Fusion Block into CSLNet. The Fusion Block is used to combine the edge texture enhancement mechanism with the multi-path feature fusion strategy and apply it to the feature integration in low-light image enhancement. By introducing the Laplacian operator and the Sobel operator in the Fusion Block, the edge and texture features of the image can be extracted respectively, effectively capturing the key detail information of the image. This dual-path processing strategy solves the deficiency of traditional single feature extraction methods in detail retention, enabling the model to comprehensively integrate complementary information from different input paths. The Fusion Block uses learnable convolutional layers and activation functions to perform non-linear transformation on the features, further optimizing the feature combination and enhancing the network's understanding ability of complex scenes. Brief Description of the Drawings

[0099] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0100] Figure 1 It is a schematic diagram of the overall structure of the CSLNet of the present invention;

[0101] Figure 2 It is a schematic diagram of the structure of the CSE Block of the present invention

[0102] Figure 3 It is a schematic diagram of the structure of the Fusion Block of the present invention;

[0103] Figure 4 It is a schematic diagram of the structure of the ETE Block of the present invention. Detailed implementation manners

[0104] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments.

[0105] A low-light image enhancement method based on the YUV color space includes at least the following steps:

[0106] S1: Build a color space illumination enhancement network (please refer to Figure 1 ), the color space illumination enhancement network is the CSLNet, and the CSLNet includes but is not limited to a Transformer enhancement module, a color space enhancement module (please refer to Figure 2 ) and a fusion module (please refer to Figure 3 ). The Transformer enhancement module is the TransformerBlock, the color space enhancement module is the CSE Block, and the fusion module is the Fusion Block;

[0107] S2: Input an RGB three-channel image into the CSLNet, perform image preprocessing and channel separation to obtain a Y channel, a U channel, and a V channel;

[0108] S3: For the Y channel, perform brightness channel enhancement processing;

[0109] S4: For the U channel and the V channel, perform chrominance channel enhancement processing;

[0110] S5: Perform feature fusion and multi-level feature extraction through the Fusion Block;

[0111] S6: Finally, perform image reconstruction and noise reduction to complete low-light image enhancement.

[0112] S2 includes at least the following steps:

[0113] Convert the RGB three-channel image input to the CSLNet into a YUV three-channel image. This conversion provides a basis for subsequent processing of luminance and chrominance information respectively. Decomposing the image into luminance (Y) and chrominance (U, V) channels can more accurately perform targeted processing on different information, avoiding the problems of luminance distortion and noise coupling in the RGB color space, thereby more effectively improving image quality and color fidelity. Refer to the following formula:

[0114] Y, U, V = RGB to YUV(I input )

[0115] In the formula, RGB to YUV represents the operation of converting the image from RGB format to YUV format.

[0116] S3 includes at least the following steps:

[0117] First, apply a 3×3 convolutional layer to extract the preliminary features of the image. The preliminary features are basic elements, including but not limited to edges and textures;

[0118] Subsequently, introduce non-linearity through the ReLU activation function to enable the network to learn more complex patterns;

[0119] Next, use the max pooling layer to reduce the spatial size of the feature map while appropriately increasing the number of channels to retain key features and reduce redundant information;

[0120] Subsequently, use the Transformer Block self-attention mechanism to capture the long-range dependencies between features, which can model the relationships between different regions, and thus better understand the global structure of the image. The Transformer Block can more comprehensively understand the image content, especially under low-light conditions, and can more effectively restore the details and texture information of the image;

[0121] The upsampling step is to restore the feature map processed by pooling and the Transformer Block to the original spatial size, ensuring that the enhanced image is consistent with the input image in spatial resolution and avoiding detail loss caused by resolution reduction. The specific operation is as follows:

[0122] Y′ = ReLU(Conv3x3(Y))

[0123] Y″ = MaxPool(Y′)

[0124] Y‴ = Transformer(Y″)

[0125] Y‴ = UpSample(Y‴) + Y

[0126] Wherein, Transformer represents being processed by the Transformer Block; UpSample represents performing an upsampling operation.

[0127] S4 includes at least the following steps:

[0128] First, it is processed by the CSE Block. The CSE Block integrates the U-shaped network structure and the multi-head self-attention mechanism, and is specifically optimized for the chrominance channel of low-light images;

[0129] Through the multi-scale feature extraction and skip connections of the U-shaped network, the CSE Block can effectively retain the detailed information of the image. At the same time, it uses the multi-head self-attention mechanism to enhance the model's understanding ability of complex spatial structures, accurately restore chrominance details and suppress noise; This multi-dimensional design not only solves the limitations of traditional methods in processing two-dimensional visual data, but also significantly improves the effect of low-light image enhancement, especially in reducing color distortion and noise;

[0130] The processed chrominance information is then input into the CSE Block for color information fusion processing. The specific operations are as follows:

[0131] First, perform convolution processing on the chrominance information:

[0132] F = CR n (F in )

[0133] CR n represents performing n Conv3×3 + ReLU operations, where n is set to 3; that is, performing three 3×3 convolution operations on F in and applying the ReLU activation function after each convolution. The purpose is to extract more complex feature information, especially in terms of chrominance information, to help the model capture more detailed information of the image;

[0134] Transformer-Block processing:

[0135] F′ = Transformer(F)

[0136] The feature map F after convolution processing is input into the Transformer Block;

[0137] Perform an upsampling operation to increase the size of the image:

[0138] F up = UpSample(F′)

[0139] Convolution fusion operation:

[0140] F″ = Conv3x3(F up ) + F m

[0141] The feature map F after upsampling up will go through a 3×3 convolution operation to obtain a new feature map, and then, this feature map is added to the input F m (The addition operation is to fuse the high-level features after upsampling with the low-level detailed information, so as to enhance the detailed performance of the image;

[0142] Final output:

[0143] F out ′ = Tanh(Conv3x3(F″))

[0144] F″ after convolution processing will pass through a 3×3 convolution layer and apply the Tanh activation function.

[0145] The Fusion Block combines the edge texture enhancement mechanism with the multi-path feature fusion strategy, and cooperates with the multi-level feature extraction module to achieve better enhancement and extraction effects;

[0146] The edge texture enhancement mechanism is based on the edge texture enhancement block, and the edge texture enhancement block is the ETE Block (please refer to Figure 4 );

[0147] The multi-path feature fusion strategy is to introduce the Laplacian operator and the Sobel operator, so as to be able to extract the edge and texture features of the image respectively, effectively capture the key detailed information of the image, and the dual-path processing strategy of the Laplacian operator and the Sobel operator solves the deficiency of the traditional single feature extraction method in detail retention, enabling the model to comprehensively integrate the complementary information from different input paths;

[0148] The Fusion Block uses learnable convolution layers and activation functions to perform non-linear transformations on the features, further optimizing the feature combination and enhancing the network's ability to understand complex scenes;

[0149] The enhanced brightness information and the fused chromaticity information are sent to the multi-level feature extraction module to screen and fuse the brightness and chromaticity features, retain more effective feature information, and further improve the detailed performance of the image. The specific operations are as follows:

[0150] The first feature fusion:

[0151] F u ″ / F v ′ = LReLU(Conv1×3(F u / F v ′))

[0152] Process the F u / F v ′ feature through a 3×3 convolution operation, and then apply the Leaky ReLU activation function, aiming to extract more effective features through convolution and maintain a certain non-linearity through the activation function;

[0153] Second feature fusion:

[0154] F u ″ / F v ″ = F u ′ / F v ′ + ETE(F u ′ / F v ′)

[0155] In this operation, directly add the results of F u ′ / F v ′ and F u ′ / F v ′ after being processed by the ETE Block. The ETE Block provides further feature enhancement and contains detailed information;

[0156] Third feature fusion:

[0157] F u ″ / F v ′″ = F u ′ / F v ′ + F u ″ / F v ″ + ETE(F u ″ / F v ″)

[0158] Combined with the original F u ′ / F v ′, F u ″ / F v ″ after the second feature fusion, and F u ″ / F v ″ after being processed by ETE, further strengthening the detailed and effective feature information;

[0159] Final feature output:

[0160] F out ″ = Concat(F u ″, F v ″, Fv ″′)

[0161] In the formula, Concat represents the concatenate operation; ETE represents being processed by the ETE Block.

[0162] The specific operations of the ETE Block include at least the following steps:

[0163] Laplacian operation:

[0164] L = Laplacian(F m ) + F m

[0165] Laplacian represents the processing by the Laplacian Operator, which can extract the detailed information in the image. By performing the Laplacian Operator processing on the feature map F of the input image m and adding the processed F after Laplacian m to the original F m can enhance the detailed features of the image;

[0166] Sobel operation:

[0167] S = Sobel(F m )

[0168] Sobel represents the processing by the Sobel Operator. By performing the Sobel Operator processing on F m can extract the edge details of the image;

[0169] Fusion processing of Laplacian and Sobel:

[0170] L′ / S′ = LReLU(Conv1×1(L / S))

[0171] Perform 1×1 convolution on the results of Laplacian and Sobel operations and apply the Leaky ReLU activation function to further fuse these two features;

[0172] Deeper convolution operation:

[0173] L″ / S″ = LReLU(Conv3×3(L′ / S′))

[0174] Apply 3×3 convolution and the Leaky ReLU activation function to the result of L′ / S′ to further extract features and enhance details;

[0175] Final Laplacian and Sobel features:

[0176] L″′ = LReLU(Conv3×3(L″ + L′)) + L′ + L″

[0177] S″′ = LReLU(Conv3×3(S″ + S′))

[0178] Output of the ETE Block:

[0179] F out ″′ = F m + L′" + S′".

[0180] S6 includes at least the following steps:

[0181] Fuse the processed luminance and chrominance information, and combine the fused image with the original image to reduce the impact caused by the loss of some original information during the processing.

[0182] Then, denoise the output image through the enhanced image denoising module to further optimize the image quality.

[0183] The enhanced image denoising module is the EID Block. The enhanced image denoising module uses the method of dense residual connection, effectively removing noise while retaining the important details of the image, significantly improving the visual quality of the image. The specific operations are as follows:

[0184] Concat operation:

[0185] F concat = Concat(Y′", UV)

[0186] Convolution and addition with the original input low-light image:

[0187] I enhanced = Conv3×3(F concat ) + I low

[0188] where I low is the original input low-light image;

[0189] EID Block processing:

[0190] I out = EID(I enhanced )

[0191] where EID represents processing through the EID Block.

[0192] In summary:

[0193] The present invention makes full use of the advantages of the YUV color space, which is helpful for low-light image enhancement tasks because it can clearly distinguish luminance (Y) and chrominance (U and V) information. With this color space, we can specifically improve the visibility and detail clarity of images under low-light conditions while avoiding adverse effects on color information. Given the sensitivity of the human visual system to luminance changes, focusing on adjusting the Y channel can achieve a more natural and visually appealing image enhancement effect.

[0194] Moreover, the core of CSLNet in the present invention lies in the color space enhancement module, which is specifically used to process the chrominance channels (U and V). Through the U-shaped network structure and the multi-head self-attention mechanism, it effectively restores chrominance details and suppresses noise. At the same time, the Transformer enhancement module is used to process the luminance channel (Y), capturing the global dependencies of the image through the self-attention mechanism to enhance the overall illumination effect of the image. In addition, the fusion module combines the texture features extracted by the Laplacian and Sobel operators, retaining more edge information and further enhancing the detail performance of the image. The multi-level feature extraction module is responsible for screening and integrating multi-scale features to optimize the detail performance of the image. Finally, the enhanced image denoising module uses dense residual connections to denoise the enhanced image, effectively removing noise while retaining the important details of the image, significantly improving the visual quality of the image. Through this design, CSLNet not only performs excellently in enhancing image luminance and details but also has significant advantages in color fidelity and visual naturalness.

[0195] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.

Claims

1. A low-light image enhancement method based on the YUV color space, characterized in that: At least include the following steps: S1: Build a color space illumination enhancement network, which is the CSLNet. The CSLNet includes, but is not limited to, a Transformer enhancement module, a color space enhancement module, and a fusion module. The Transformer enhancement module is the TransformerBlock, the color space enhancement module is the CSE Block, and the fusion module is the FusionBlock; S2: Input an RGB three-channel image into the CSLNet, perform image preprocessing and channel separation to obtain a Y channel, a U channel, and a V channel; S3: For the Y channel, perform brightness channel enhancement processing; S4: For the U channel and the V channel, perform chrominance channel enhancement processing; S5: Perform feature fusion and multi-level feature extraction through the Fusion Block; S6: Finally, perform image reconstruction and noise reduction to complete low-light image enhancement.

2. The low-light image enhancement method based on the YUV color space according to claim 1, characterized in that: The S2 at least includes the following steps: Convert the RGB three-channel image input into the CSLNet into a YUV three-channel image. This conversion provides a basis for subsequent processing of brightness and chrominance information respectively. Decomposing the image into brightness (Y) and chrominance (U, V) channels can more accurately perform targeted processing on different information, avoiding the problems of brightness distortion and noise coupling in the RGB color space, thereby more effectively improving image quality and color fidelity. Refer to the following formula: Y, U, V = RGB to YUV(I input ) In the formula, RGB to YUV represents the operation of converting the image from the RGB format to the YUV format.

3. A low-light image enhancement method based on the YUV color space according to claim 1, characterized in that: The S3 at least includes the following steps: First, apply a 3×3 convolutional layer to extract the initial features of the image. The initial features are basic elements, and the basic elements include, but are not limited to, edges and textures; Subsequently, introduce non-linearity through the ReLU activation function to enable the network to learn more complex patterns; Next, use the max pooling layer to reduce the spatial size of the feature map while appropriately increasing the number of channels to retain key features and reduce redundant information; Subsequently, use the self-attention mechanism of the Transformer Block to capture the long-range dependence relationships between features, which can model the mutual relationships between different regions, and thus better understand the global structure of the image. The Transformer Block can more comprehensively understand the image content, especially under low-light conditions, and can more effectively restore the details and texture information of the image; The upsampling step is to restore the feature map processed by pooling and the Transformer Block to the original spatial size, ensuring that the enhanced image is consistent with the input image in spatial resolution and avoiding detail loss caused by resolution reduction. The specific operation is as follows: Y′ = ReLU(Conv3x3(Y)) Y″ = MaxPool(Y′) Y″′ = Transformer(Y″) Y″′ = UpSample(Y″′)+Y Wherein, Transformer represents being processed by the Transformer Block; UpSample represents performing an upsampling operation.

4. The low-light image enhancement method based on the YUV color space according to claim 3, wherein: S4 at least includes the following steps: First, it is processed by the CSE Block. The CSE Block integrates the U-shaped network structure and the multi-head self-attention mechanism, and is specifically optimized for the chrominance channel of low-light images. Through the multi-scale feature extraction and skip connections of the U-shaped network, the CSE Block can effectively retain the detailed information of the image. At the same time, it uses the multi-head self-attention mechanism to enhance the model's understanding ability of complex spatial structures, accurately restore chrominance details and suppress noise. The processed chrominance information is then input into the CSE Block for color information fusion processing. The specific operations are as follows: First, perform convolution processing on the chrominance information: F = CR n (F in ) CR n Indicates that the Conv3×3 + ReLU operation is performed n times, where n is set to 3; that is, for F in Three 3×3 convolution processes are carried out, and the ReLU activation function is applied after each convolution. The purpose is to extract more complex feature information, especially in terms of chromaticity information, to help the model capture more image details; Transformer-Block processing: F′ = Transformer(F) The feature map F after convolution processing is input into the Transformer Block. Perform an upsampling operation to increase the size of the image: F up = UpSample(F′) Convolution fusion operation: F″ = Conv3x3(F up ) + F m The feature map F after upsampling up will undergo a 3×3 convolution operation to obtain a new feature map. Then, this feature map is added to the input F m to fuse the high-level features after upsampling with the low-level detailed information, thereby enhancing the detailed performance of the image; Final output: F out ′ = Tanh(Conv3x3(F″)) The F″ after convolution processing will pass through a 3×3 convolutional layer and apply the Tanh activation function.

5. A low-light image enhancement method based on the YUV color space according to claim 4, characterized in that: The Fusion Block combines the edge texture enhancement mechanism and the multi-path feature fusion strategy, and cooperates with the multi-level feature extraction module to achieve better enhancement and extraction effects. The edge texture enhancement mechanism is based on the edge texture enhancement block, which is the ETE Block. The multi-path feature fusion strategy is to introduce the Laplacian operator and the Sobel operator, so as to be able to extract the edge and texture features of the image respectively, effectively capture the key detailed information of the image. The dual-path processing strategy of the Laplacian operator and the Sobel operator solves the deficiency of the traditional single feature extraction method in detail retention, enabling the model to comprehensively integrate complementary information from different input paths. The Fusion Block uses learnable convolutional layers and activation functions to perform non-linear transformation on the features, further optimizing the feature combination and enhancing the network's understanding ability of complex scenes. Send the enhanced luminance information and the fused chrominance information into the multi-level feature extraction module to screen and fuse the luminance and chrominance features, retain more effective feature information, and further improve the detail performance of the image. The specific operations are as follows: First feature fusion: F u ″ / F v ′=LReLU(Conv1×3(F u / F v ′)) Process the F through a 3×3 convolution operation u / F v ' features, and then apply the Leaky ReLU activation function. The purpose is to extract more effective features through convolution and maintain a certain degree of non-linearity through the activation function; Second feature fusion: F u ″ / F v ″ = F u ′ / F v ′ + ETE(F u ′ / F v ′) In this operation, directly add F u ′ / F v ′ and the result of F u ′ / F v ′ after being processed by the ETE Block. The ETE Block provides further feature enhancement and contains detailed information; Third feature fusion: F u ″ / F v ″′=F u ′ / F v ′+F u ″ / F v ″+ETE(F u ″ / F v ″) Combined with the original F u ′ / F v ′, F u ″ / F v ″ after the second feature fusion, and F u ″ / F v ″ after ETE processing, further enhancing details and effective feature information; Final feature output: F out " = Concat(F u ", F v ", F v ") Wherein, Concat represents the concatenate operation; ETE represents being processed by the ETE Block.

6. The low-light image enhancement method based on the YUV color space according to claim 5, wherein: The specific operations of the ETE Block at least include the following steps: Laplacian operation: L = Laplacian(F m ) + F m The Laplacian represents the processing of the Laplacian Operator, which can extract the detailed information in the image. By applying it to the feature map F of the input image m performing the Laplacian Operator processing, and then adding the processed F m to the original F m can enhance the detailed features of the image; Sobel operation: S = Sobel(F m ) Sobel indicates that after being processed by the Sobel Operator, by applying the Sobel Operator to F m it is possible to extract the edge details of the image after being processed by the Sobel Operator; Fusion processing of Laplacian and Sobel: L′ / S′ = LReLU(Conv1×1(L / S)) Perform 1×1 convolution on the Laplacian and Sobel operation results, and apply the Leaky ReLU activation function to further fuse these two features; Deeper convolutional operations: L″ / S″ = LReLU(Conv3×3(L′ / S′)) Apply 3×3 convolution and Leaky ReLU activation function to the L′ / S′ results to further extract features and enhance details; Final Laplacian and Sobel features: L″′ = LReLU(Conv3×3(L″ + L′)) + L′ + L″ S″′ = LReLU(Conv3×3(S″ + S′)) ETE Block output: F out ″′ = F m + L″′ + S″′.

7. A low-light image enhancement method based on the YUV color space according to claim 5, characterized in that: The S6 at least includes the following steps: Fuse the processed luminance and chrominance information, and combine the fused image with the original image to reduce the impact caused by the loss of some original information due to the processing process; Then, perform noise reduction processing on the output image through the enhanced image noise reduction module to further optimize the image quality, The enhanced image noise reduction module is the EID Block. The enhanced image noise reduction module adopts the method of dense residual connection, effectively removing noise while retaining the important details of the image, and significantly improving the visual quality of the image. The specific operations are as follows: Concat operation: F concat = Concat(Y″′, UV) Convolution and addition with the original input low-light image: I enhanced = Conv3×3(F concat ) + I low where I low is the original input low-light image; EID Block processing: I out = EID(I enhanced ) Where EID is represented as processed by the EID Block.

Citation Information

Cited By

  • Multi-stage information interactive weak light image enhancement method based on two-channel standardized flow

    CN120912472A

  • Multi-stage information interactive weak light image enhancement method based on double-channel normalized flow

    CN120912472B

  • Greenhouse plant growth state recognition method and planting management and control system

    CN121191002A

  • Greenhouse plant growth state recognition method and planting management system

    CN121191002B

  • Image color fidelity optimization method and system based on color perception feature mapping

    CN121883619A