A low-light image enhancement method based on multi-scale global joint local

Through a multi-scale global combined local low-light image enhancement method, the multi-scale prior module, local enhancement module and adaptive core selection module are used, combined with the student-teacher network, the global and local degradation problems in low-light image enhancement are solved, and the image enhancement effect and model generalization ability are improved.

CN119919772BActive Publication Date: 2025-07-11NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510398481.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-11
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The existing low-light image enhancement methods have shortcomings in dealing with global and local degradation problems. Traditional methods are difficult to effectively deal with local details loss and noise interference. The deep learning model has high computational complexity in resource-constrained scenarios. The existing visual SSM model fails to effectively maintain the spatial locality of the image, and the physical prior guidance mechanism is rigid.

Method used

A multi-scale global combined local low-light image enhancement method is adopted, and a multi-scale prior module, a local enhancement module and an adaptive core selection module are used to fine-tune the student-teacher network to improve the generalization ability of the network, and the multi-scale hollow convolution and channel and spatial attention are used, and the dynamic pseudo-label generation mechanism is used to enhance the image enhancement effect.

Benefits of technology

Significantly expand the receptive field, strengthen feature representation, restore image details, reduce noise interference, improve the generalization ability of the model in real scenes, reduce dependence on manual labeled data, and improve image enhancement effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919772B_ABST
    Figure CN119919772B_ABST
Patent Text Reader

Abstract

The present invention proposes a low-light image enhancement method based on multi-scale global joint local. This method constructs a multi-scale prior extraction module, takes a low-light image as input for pre-training, and extracts the prior features of the image. By introducing multi-scale prior extraction and combining local enhancement, it effectively captures complex global and local dependencies, further improving the quality of image enhancement. By introducing an adaptive kernel selection module to dynamically select features using spatially varying operations, it achieves flexible adaptation to different inputs. This method also introduces a dynamic pseudo-label generation framework, which improves the generalization ability of the network through pseudo-label generation, dynamic confidence evaluation, and knowledge distillation. The present invention also demonstrates excellent performance in practical application scenarios such as low-light object detection and has broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of low-light image enhancement, and specifically to a low-light image enhancement method based on multi-scale global joint local. Background Art

[0002] Low-light image enhancement is a key technology in the field of computer vision, aiming to improve the quality of visual data captured under insufficient lighting conditions. Due to the dual effects of the physical limitations of sensors and the complexity of environmental lighting conditions, low-light images generally suffer from problems such as reduced global contrast, loss of local details, noise interference, and color distortion, which directly affect human visual perception and the performance of downstream tasks. Traditional enhancement methods mainly rely on global mapping strategies such as histogram equalization and gamma correction, but it is difficult to effectively handle local degradation phenomena, resulting in problems such as blurred details and artifacts in the enhancement results.

[0003] In recent years, significant progress has been made in low-light enhancement methods based on deep learning. Among them, algorithms guided by the Retinex theory optimize the reflection component and the illumination component of the image independently, showing advantages in enhancing global brightness. However, traditional Retinex models lack the ability to explicitly model noise and local degradation, and rely on complex multi-branch network structures and physical constraints, resulting in a significant increase in model complexity and computational cost. Convolutional neural networks (CNNs) perform excellently in detail restoration through local feature aggregation capabilities, but their fixed receptive fields limit the ability to model global degradation. Although the Transformer architecture realizes long-range dependence modeling through the self-attention mechanism, its quadratic computational complexity limits its application in resource-constrained scenarios.

[0004] As an emerging long-sequence modeling tool, the state space model (SSM) has emerged in vision tasks with its linear computational complexity and global information capture ability. However, existing visual SSM models mostly follow the one-way scanning strategy and fail to effectively maintain the spatial locality of two-dimensional images, resulting in insufficient extraction of local degradation features. For example, although typical visual SSM models can establish global dependencies, they lack targeted designs for local optimization tasks such as noise suppression and edge preservation. In addition, existing enhancement methods generally face the problem of rigid physical prior guidance mechanisms and are difficult to achieve dynamic adaptive fusion of illumination components and reflection components.

[0005] In summary, although existing low-light image enhancement methods have made some progress, there are still deficiencies in simultaneously handling global and local degradation problems. Summary of the Invention

[0006] In view of the limitations faced by existing convolutional neural networks and Transformers in processing low-light images, as well as the problem of network overfitting commonly encountered in small-sample datasets, the present invention proposes a low-light image enhancement method based on multi-scale global joint local, which makes full use of global and local information to enhance the model's understanding. At the same time, a student-teacher network is introduced for fine-tuning in the second stage to improve the generalization ability of the network in the face of real scenes and enhance the effect of image enhancement, so as to solve the problems raised in the above background technology. The technical solutions provided by the present invention are as follows:

[0007] A low-light image enhancement method based on multi-scale global joint local, comprising the following steps:

[0008] Step 1, construct a multi-scene low-light image training dataset composed of multiple low-light images, and the low-light images are used to train an image enhancement network;

[0009] Step 2, the image enhancement network is a U-Net-like network, which consists of a multi-scale prior extraction module, multiple local enhancement modules and an adaptive kernel selection module. The multi-scale prior extraction module consists of two parts. The first part is the multi-scale convolution part, which takes three branch networks and a feature fusion module as the main body. The input low-light image extracts features with different receptive fields through the three branch networks, and the extracted features are integrated in the feature fusion module to obtain the processing result of the first part. The second part inputs the processing result from the first part, and performs parallel processing on the channel and spatial two branches to learn the importance weights on the channel and spatial two branches, and obtains the important feature information on the channel and spatial two branches through weighted operations. Integrate the important feature information on the channel and spatial two branches to obtain the prior feature F p ; input the low-light image into the multi-scale prior extraction module to extract the prior feature F of the image p ; obtain the low-light image feature F through convolution operation on the low-light image l ;

[0010] Step 3, input the low-light image feature F l and the prior feature F p into the local enhancement module. The local enhancement module consists of two branches. In the first branch, the low-light image feature F l successively undergoes operations of the LayerNorm layer, the Linear layer, the DWConv layer, the 2DSSM module and the Layer Norm layer to obtain the feature F1; the prior feature F p successively undergoes processing of the Conv layer, the Linear layer and the SiLU layer to obtain the feature F2; multiply the feature F1 and the feature F2, and obtain the local enhancement feature F after processing by the Linear layer out1 , and output the local enhancement feature F out1 ;

[0011] Step 4: local enhancement feature F out1 and the prior features F p Input adaptive kernel selection module, which consists of two branches. The first branch contains Conv, SK-1, SK-2 and kernel selection module, and the second branch contains DWConv, Sigmoid and Chunk layers. This module uses the prior feature F p As input, after being processed by the second branch, it is input into the kernel selection module together with the local enhanced feature F out1 After being processed by the first branch, the final output feature F is obtained through the operations of DWConv, GELU and Conv. out2 , output enhanced feature F out2 ;

[0012] Step 5: Enhance the feature F out2 and the prior features F p Repeating steps 3 and 4 in the image enhancement network through downsampling and upsampling to train the image enhancement network;

[0013] In step 6, real low-light images in different scenes are taken as input and fine-tuned using the second-stage student-teacher network to obtain an image enhancement network with the ability to generalize to real scenes.

[0014] Preferably, in step 1, the specific process of constructing a multi-scene low-light image training dataset is: selecting low-light images of multiple scenes from existing public low-light datasets, including a LOL dataset consisting of pairs of low-light images and their corresponding normal-light images, a DPED dataset containing real photos taken from three different mobile phones and a high-end SLR camera, and a Dark Face dataset focusing on face detection under low-light conditions; collecting a real low-light dataset and eliminating some blurred images, and merging the public low-light dataset images and the real low-light images to obtain a dataset for network training.

[0015] Preferably, in the multi-scale convolution part, the first branch uses a 1x1 convolution kernel, does not change the spatial scale, and directly extracts features; the second branch uses a 3x3 convolution kernel with a void ratio of 12 to expand the receptive field to capture a wider range of contextual information; the third branch uses global average pooling to extract global contextual features to enhance the model's ability to understand the overall layout.

[0016] Preferably, on the channel branch, first, a global pooling operation is performed to obtain the global information of each channel, and then through two fully connected layers, with ReLU and Sigmoid used as activation functions respectively for the two fully connected layers to learn the importance weights on the channel branch; on the spatial branch, the input features are subjected to 1×1 convolution and Sigmoid activation function operations to learn the importance weights on the spatial branch.

[0017] Preferably, the 2DSSM module enhances the original state space model SSM through enhanced local deviation, and this module is expressed as:

[0018]

[0019] y t =Ch t +Dx t +E

[0020] where the one-dimensional sequence x t is projected into a new one-dimensional sequence y t through the hidden state h t , and A, B, C, D represent the state, input projection, output projection, and feedthrough parameters respectively, and E is the feature enhancement information.

[0021] Preferably, the student-teacher network consists of a student network, a teacher network, a pseudo-label quality score, and a high-quality pseudo-label pool. The student network and the teacher network have the same network structure. The teacher network uses non-enhanced real low-light images as input, and the student network uses strongly enhanced images as input; the labels output by the teacher network are discriminated through the pseudo-label quality score to form a high-quality pseudo-label pool for the student network to perform high-quality training.

[0022] Preferably, the pseudo-label quality score is a hybrid scoring mechanism that fuses a physical model and a domain adaptation discriminator, specifically as follows:

[0023] Assume that the low-light noise follows a Poisson-Gaussian mixture model, and its observed noise variance is decomposed as: σ 2 (y)=α·y+β 2 ;

[0024] where y is the ideal noise-free signal, α is the Poisson noise intensity, and β is the Gaussian noise standard deviation;

[0025] For each pixel block p, calculate its actual noise variance:

[0026] By minimizing the difference between the actual variance and the model-predicted variance, the residual confidence score is defined:

[0027]

[0028] where is the mean of the pseudo - labels within the block, and λ is the control sensitivity;

[0029] Train the binary discriminator D to distinguish real normal light images from low - light images. Input the pseudo - label I output by the teacher network into the discriminator D and calculate its naturalness score: S enhanced = D(I domain ) enhanced )

[0030] Perform weighted fusion on the final confidence score of pixel (i,j): S(i,j)=γ·S noise (i,j)+(1 - γ)·S domain (i,j)

[0031] where γ is a balancing factor with a value between 0 and 1;

[0032] Set the threshold τ to generate a binary mask M and screen the high - confidence regions for student network training:

[0033]

[0034] The pseudo - label pool is dynamically updated according to the pseudo - label scores to ensure that the labels used for student network training are high - quality labels.

[0035] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0036] The present invention adopts a multi - scale prior module. By introducing multi - scale dilated convolutions, the receptive field can be significantly expanded without losing resolution, enabling the network to extract features at different spatial scales; introducing channel attention, by evaluating the importance of each channel and making corresponding adjustments, useful feature channels can be strengthened and unimportant channels can be suppressed, enabling the network to focus on the information of interest, thereby improving the quality of feature representation; introducing spatial attention, paying attention to the importance of specific regions in the image, which helps the model to focus on processing the local regions that have the greatest impact on the final task.

[0037] The present invention adopts a local enhanced spatial module. Through an enhanced local deviation mechanism, local 2D dependencies are retained during the two - dimensional selective scanning process, improving the original state - space model, enabling the network to better capture non - causal relationships when processing visual data, and avoiding the problem of increased distance of neighborhood data caused by fixed scanning methods.

[0038] The present invention adopts an adaptive kernel selection module, which adaptively adjusts the illumination intensity according to the specific situation of the input image, avoiding complex structural designs and constraints to estimate physical priors. At the same time, through residual connections, depth convolution, and the GELU function, this module can effectively restore the texture and color details in the image, especially under complex degradation conditions such as low illumination.

[0039] The present invention also adopts a dynamic pseudo-label generation mechanism. By evaluating the performance of the current model and storing the best results in the label pool, the pseudo-labels are dynamically updated to ensure that the optimal results are used, reducing noise and inconsistency. By generating high-quality pseudo-labels, the unlabeled data is efficiently utilized, reducing the dependence on manually labeled data and improving the training efficiency of the model. At the same time, the generalization ability of the model is improved, enabling the model to better adapt to real-world data and perform well on unseen data. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0041] Figure 1 is a schematic diagram of the overall network structure of the method of the present invention;

[0042] Figure 2 is a schematic diagram of the multi-scale prior extraction network proposed by the present invention;

[0043] Figure 3 is a schematic diagram of the student-teacher network proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] Please refer to Figures 1-3 , the present invention provides the following technical solutions:

[0046] Embodiment 1: A low-light image enhancement method based on multi-scale global joint local, as Figure 1 shown, includes the following steps:

[0047] Step 1, construct a multi-scene low-light image training dataset composed of multiple low-light images, and the low-light images are used to train the image enhancement network.

[0048] Specifically, low-light images of multiple scenarios are selected from existing public low-light datasets, including the LOL dataset composed of pairs of low-light images and their corresponding normal-light images, the DPED dataset containing real photos taken from three different mobile phones and a high-end single-lens reflex camera, and the DarkFace dataset focusing on face detection under low-light conditions; a real low-light dataset is collected and some blurred images are removed, and the images in the public low-light dataset and the real low-light images are merged to obtain a dataset for network training.

[0049] In this embodiment, 485 pairs of low-light / normal-light images in the LOL dataset are used for training, and 15 images are used for testing; 500 pairs of low-light / normal-light images in the DPED dataset are used for training, and 20 images are used for testing; 600 pairs of low-light / normal-light images in the Dark Face dataset are used for training, and 25 images are used for testing; the real low-light dataset includes 50 low-light images.

[0050] Step 2, the image enhancement network is a U-Net-like network, which consists of a multi-scale prior extraction module, multiple local enhancement modules and an adaptive kernel selection module. The low-light image is input into the multi-scale prior extraction module to extract the prior feature F of the image. p ; The low-light image is subjected to a convolution operation to obtain the low-light image feature F. l .

[0051] The multi-scale prior extraction module is as Figure 2 shown, and consists of two parts. The first part is the multi-scale convolution part, which takes three branch networks and a feature fusion module as the main body. The input low-light image extracts features with different receptive fields through the three branch networks, and the extracted features are integrated in the feature fusion module to obtain the processing result of the first part; the second part inputs the processing result from the first part, and performs parallel processing on the channel and spatial branches, learns the importance weights on the channel and spatial branches, and obtains the important feature information on the channel and spatial branches through a weighted operation; integrates the important feature information on the channel and spatial branches to obtain the prior feature F. p .

[0052] Specifically, in the multi-scale convolution part, the first branch uses a 1x1 convolution kernel, does not change the spatial scale, and directly extracts features; the second branch uses a 3x3 convolution kernel with a dilation rate of 12 to expand the receptive field to capture more extensive context information; the third branch uses global average pooling to extract global context features and enhance the model's understanding ability of the overall layout.

[0053] On the second part of the channel branch, first, a global pooling operation is performed to obtain the global information of each channel. Then, through two fully connected layers, ReLU and Sigmoid are respectively used as activation functions in the two fully connected layers to learn the importance weights on the channel branch. On the spatial branch, the input features are subjected to 1×1 convolution and Sigmoid activation function operations to learn the importance weights on the spatial branch.

[0054] The above process is described as:

[0055] F s1 = concat(Conv 1×1 (x), Conv 3×3 (x), AvgPool(x))

[0056]

[0057] Step 3, input the low-light image feature F l and the prior feature F p into the local enhancement module, and output the local enhancement feature F out1 .

[0058] As shown in the LEM part of Figure 1 , the local enhancement module consists of two branches. In the first branch, the low-light image feature F l is successively subjected to operations of Layer Norm layer, Linear layer, DWConv layer, 2DSSM module and Layer Norm layer to obtain the feature F1; the prior feature F p is successively processed by Conv layer, Linear layer and SiLU layer to obtain the feature F2; the feature F1 and the feature F2 are multiplied point by point and processed by the Linear layer to obtain the local enhancement feature F out1 . The above process is described as:

[0059] F1 = LN(2DSSM(DWConv(Linear(LN(F l )))))

[0060] F2 = SiLU(Linear(Conv(F p )))

[0061]

[0062] Furthermore, the local enhancement module of this embodiment enhances the original SSM by maintaining local 2D dependencies. Structured state space sequence models and SSMs such as Mamba can be regarded as continuous linear time-invariant systems. Given a one-dimensional sequence x(t), it is projected into a new one-dimensional sequence y(t) through the hidden state h(t), and the whole system can be defined as a linear ordinary differential equation:

[0063] h′(t) = Ah(t) + Bx(t)

[0064] y(t) = Ch(t) + Dx(t)

[0065] Where A, B, C, and D represent the state, input projection, output projection, and feedthrough parameters respectively. The LEM module enhances the original SSM by introducing enhanced local deviation, and the entire system is described as:

[0066]

[0067] y t = Ch t + Dx t + E

[0068] Where E is the feature enhancement information. Therefore, the calculation of this model is simple. Given the low-light image feature F l and the prior feature F p , the Layernorm and LEM modules are used to integrate the spatial long-term dependence.

[0069] Step 4, input the local enhanced feature F out1 and the prior feature F p into the adaptive kernel selection module, and output the enhanced feature F out2 .

[0070] As Figure 1 shown in the AKSM part of p , the adaptive kernel selection module consists of two branches. The first branch includes Conv, SK-1, SK-2, and the kernel selection module, and the second branch includes DWConv, Sigmoid, and the Chunk layer. This module takes the prior feature F out1 as the input, and after being processed by the second branch, it is input into the kernel selection module. Together with the local enhanced feature F out2 after being processed by the first branch and then through the operations of DWConv, GELU, and Conv, the final output feature F

[0071] F k = F out1 , F k+1 = DWConv(F k )

[0072] {S1, S2} = Chunk(Sigmoid(Conv(F p )))

[0073]

[0074] Step 5, input the enhanced feature Fout2 and the prior feature F p Repeat the operations of step 3 and step 4 in the image enhancement network through downsampling and upsampling to train the image enhancement network.

[0075] Embodiment 2: Further, a student-teacher network is used to fine-tune the model trained in step 5. Real-world low-light images are introduced, combined with strong geometric data augmentation, to further optimize the model, and high-quality pseudo-labels are generated through a label generation framework to improve the generalization ability and practical application effect of the model.

[0076] The student-teacher network is as Figure 3 shown. This network consists of a student network, a teacher network, a pseudo-label quality score, and a high-quality pseudo-label pool. The student network and the teacher network have the same network structure and share weights through exponential moving average. Specifically, the teacher network uses non-enhanced real low-light images as input, and the student network uses strongly enhanced images as input. After the inference of the network, the following results are generated:

[0077] I enhanced = f tea (I L ), I pre = f stu (A(I L ))

[0078] where I L is the low-light image of the real scene, I enhanced is the output result of the teacher network, and I pre is the output result of the student network.

[0079] This embodiment uses a hybrid scoring mechanism that combines a physical model and a domain adaptation discriminator, specifically as follows:

[0080] Assume that the low-light noise follows a Poisson-Gaussian mixture model, and its observed noise variance is decomposed as: σ 2 (y) = α·y + β 2 ;

[0081] where y is the ideal noise-free signal, α is the Poisson noise intensity, and β is the Gaussian noise standard deviation;

[0082] For each pixel block p, calculate its actual noise variance:

[0083] By minimizing the difference between the actual variance and the model-predicted variance, define the residual confidence score:

[0084]

[0085] where is the mean value of the pseudo-label within the block, and λ is the control sensitivity;

[0086] Train the binary discriminator D to distinguish between real normal light images and low-light images, and input the pseudo-label I output by the teacher network enhanced into the discriminator D, and calculate its naturalness score close to normal light: S domain = D(I enhanced )

[0087] Perform weighted fusion on the final confidence score of pixel (i, j): S(i, j) = γ·S noise (i, j) + (1 - γ)·S domain (i, j)

[0088] where γ is a balance factor with a value of 0 - 1;

[0089] Set the threshold τ, generate a binary mask M, and screen the high-confidence regions for the student network training:

[0090]

[0091] The pseudo-label pool is dynamically updated according to the pseudo-label scores to ensure that the labels used for the student network training are high-quality labels.

[0092] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A low-light image enhancement method based on multi-scale global joint local, characterized in that, The steps include: Step 1, constructing a multi-scene low-light image training dataset consisting of multiple low-light images, wherein the low-light images are used to train an image enhancement network; Step 2: The image enhancement network is a U-Net-like network, which consists of a multi-scale prior extraction module, multiple local enhancement modules, and an adaptive kernel selection module. The multi-scale prior extraction module consists of two parts. The first part is the multi-scale convolution part, which mainly includes three branch networks and a feature fusion module. The input low-light image undergoes feature extraction with different receptive fields through the three branch networks, and the extracted features are integrated in the feature fusion module to obtain the processing result of the first part. The second part takes the processing result from the first part and processes it in parallel in two branches, namely the channel branch and the spatial branch, to learn the importance weights on the two branches. Through weighted operations, important feature information on the channel and spatial branches is obtained. The important feature information on the channel and spatial branches is integrated to obtain the prior feature F p ; The low-light image is input into the multi-scale prior extraction module to extract the prior feature F of the image p ; The low-light image undergoes a convolution operation to obtain the low-light image feature F l ; Step 3, input the low-light image feature F l and the prior feature F p into the local enhancement module. The local enhancement module consists of two branches. In the first branch, the low-light image feature F l is sequentially operated through a LayerNorm layer, a Linear layer, a DWConv layer, a 2DSSM module, and a Layer Norm layer to obtain the feature F1; the prior feature F p is sequentially processed through a Conv layer, a Linear layer, and a SiLU layer to obtain the feature F2; the feature F1 and the feature F2 are multiplied element-wise and then processed through a Linear layer to obtain the local enhancement feature F out1 , and output the local enhancement feature F out1 ; Step 4, input the local enhanced feature F out1 and the prior feature F p into the adaptive kernel selection module. The adaptive kernel selection module consists of two branches. The first branch includes Conv, SK-1, SK-2, and the kernel selection module. The second branch includes DWConv, Sigmoid, and the Chunk layer. This module takes the prior feature F p as the input. After being processed by the second branch, it is input into the kernel selection module. Together with the local enhanced feature F out1 , after being processed by the first branch, it undergoes operations of DWConv, GELU, and Conv to obtain the final output feature F out2 , and outputs the enhanced feature F out2 ; Step 5, take the enhanced feature F out2 and the prior feature F p Through downsampling and upsampling, repeat the operations of Step 3 and Step 4 in the image enhancement network to train the image enhancement network; In step 6, real low-light images in different scenes are taken as input and fine-tuned using the second-stage student-teacher network to obtain an image enhancement network with the ability to generalize to real scenes.

2. The low-light image enhancement method based on multi-scale global joint local according to claim 1, wherein In step 1, the specific process of constructing a multi-scene low-light image training dataset is as follows: select low-light images of various scenes from existing public low-light datasets, including the LOL dataset consisting of pairs of low-light images and their corresponding normal-light images, the DPED dataset containing real photos taken from three different mobile phones and a high-end SLR camera, and the Dark Face dataset that focuses on face detection under low-light conditions; collect a real low-light dataset and remove some blurred images, and merge the public low-light dataset images and the real low-light images to obtain a dataset for network training.

3. A low-light image enhancement method based on multi-scale global joint local according to claim 1, characterized in that, In the multi-scale convolution part, the first branch uses a 1x1 convolution kernel, does not change the spatial scale, and directly extracts features; The second branch uses a 3x3 convolution kernel with a void ratio of 12 to expand the receptive field to capture a wider range of contextual information. The third branch uses global average pooling to extract global contextual features and enhance the model's ability to understand the overall layout.

4. A low-light image enhancement method based on multi-scale global joint local according to claim 3, characterized in that On the channel branch, a global pooling operation is first performed to obtain the global information of each channel, and then two fully connected layers are passed. The two fully connected layers use ReLU and Sigmoid as activation functions respectively to learn the importance weights on the channel branch; On the spatial branch, the input features are subjected to 1×1 convolution and Sigmoid activation function operations to learn the importance weights on the spatial branch.

5. A low-light image enhancement method based on multi-scale global joint local according to claim 1, characterized in that The 2DSSM module enhances the original state space model SSM by enhancing the local deviation, and the module is expressed as: y t = Ch t + Dx t + E where the one-dimensional sequence x t is projected into a new one-dimensional sequence y t through the hidden state h t , where A, B, C, and D respectively represent the state, input projection, output projection, and feedthrough parameters, and E is the feature enhancement information.

6. A low-light image enhancement method based on multi-scale global joint local according to any one of claims 1-5, characterized in that, The student-teacher network consists of a student network, a teacher network, a pseudo-label quality score and a high-quality pseudo-label pool. The student network and the teacher network have the same network structure. The teacher network uses non-enhanced real low-light images as input, and the student network uses strongly enhanced images as input. The output of the teacher network is a label that is discriminated by the pseudo-label quality score to form a high-quality pseudo-label pool for high-quality training of the student network.

7. A low-light image enhancement method based on multi-scale global joint local according to claim 6, characterized in that, The pseudo-label quality score is a hybrid scoring mechanism that integrates the physical model and the domain adaptation discriminator, as follows: Assume that the low-light noise follows a Poisson-Gaussian mixture model, and its observation noise variance is decomposed as: σ 2 (y) = α·y + β 2 ; Where y is the ideal noise-free signal, α is the Poisson noise intensity, and β is the Gaussian noise standard deviation; For each pixel block p, calculate its actual noise variance: The residual confidence score is defined by minimizing the difference between the actual variance and the model predicted variance: where is the mean of the pseudo-labels within the block, and λ is the control sensitivity; Train the binary discriminator D to distinguish between real normal light images and low-light images, and input the pseudo-label I output by the teacher network enhanced into the discriminator D and calculate its naturalness score: S domain = D(I enhanced ) Perform weighted fusion on the final confidence score of pixel (i, j): S(i, j) = γ · S noise (i, j) + (1 - γ) · S domain (i, j), where γ is a balancing factor with a value between 0 and 1; Set the threshold τ, generate a binary mask M, and filter high confidence areas for student network training: The pseudo-label pool is dynamically updated according to the pseudo-label scores to ensure that the labels used for student network training are high-quality labels.

Citation Information

Patent Citations

  • Low-illumination image enhancement method based on multi-scale and context learning network

    CN114998145A

  • Image enhancement method in low-light scene

    CN118195978A