Method for suppressing glare of low-illumination image during night driving

By using an improved Uformer network structure, combined with gated position attention and local context-aware Transformer modules, the problem of accurate localization of glare areas and restoration of local structural details in low-light images of nighttime driving is solved, improving the clarity and perception reliability of nighttime images, and making it suitable for intelligent driving assistance systems.

CN120953103APending Publication Date: 2025-11-14UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511068545.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify glare areas and recover local structural details in low-light images of vehicles driving at night, and they also require high computational resources, making it difficult to meet the dual requirements of real-time performance and accuracy.

Method used

An improved Uformer network structure is adopted, which combines a gated position attention module and a local context-aware Transformer module. Image sampling is performed through a depthwise separable convolution and pixel rearrangement module. The background loss, glare loss and reconstruction loss function are combined for joint training to achieve accurate localization of glare areas and restoration of background details.

Benefits of technology

It effectively restores road details obscured by glare, improves the clarity and perception reliability of nighttime images, and is suitable for electronic rearview mirrors and forward perception systems in intelligent driving assistance systems, ensuring the driver's visual recognition ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953103A_ABST
    Figure CN120953103A_ABST
Patent Text Reader

Abstract

The invention discloses a night driving low-illumination image glare suppression method, and relates to the technical field of image processing, and the method comprises the following steps: collecting a night road scene and a glare image, and fusing the night road scene and the glare image to generate a data set; according to the method, an improved Uform network structure is adopted, efficient sampling is achieved through a deep separable convolution and pixel rearrangement (Pixel Shuffle) module, and local details and structure information are extracted by means of local context awareness (LCAWin-Transform); a glare area is accurately positioned through a gating position attention module in combination with position codes and context information; and adopting background loss, glare loss and reconstruction loss to jointly train the network, recovering background details and predicting glare. The problem of glare interference caused by a strong light source under the complex illumination condition at night can be effectively solved, the image definition and perception reliability are remarkably improved in an intelligent driving assistance system, driving safety is guaranteed, and the method has good engineering practical value and popularization prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method for suppressing glare in low-light images of vehicles driving at night. Background Technology

[0002] With the continuous development of intelligent driving technology, electronic rearview mirrors, as a new generation of intelligent sensing terminals, have been widely used in driver assistance systems. By capturing images of the rear and sides of the vehicle in real time through cameras and projecting them onto in-vehicle display devices, electronic rearview mirrors effectively overcome the limitations of traditional optical mirrors, such as the limited field of view and weather conditions, greatly improving driving safety and information perception capabilities. However, in practical applications, due to interference from complex light sources such as streetlights, vehicle headlights, and billboards, images in nighttime road scenes often exhibit strong glare areas, severely affecting the image readability and driving assistance effects of the system. Traditional image processing methods, such as histogram equalization and Retinex theory, can enhance visibility by increasing the overall brightness of the image, but when faced with localized strong light interference, they easily lead to loss of detail and blurred edges, and are difficult to effectively locate and suppress irregularly shaped glare areas. In addition, while image decomposition-based methods can separate the glare layer from the scene layer to some extent, they rely on fixed rules and are difficult to adapt to complex glare distributions of different shapes and locations, resulting in a significant performance degradation in real complex scenes. In recent years, deep learning-based methods have gradually become a research hotspot, such as glare removal from single RGB images through physical modeling and semi-synthetic data training. However, these methods have limited generalization ability in extreme lenses or special reflection modes, and distortion may occur in some areas. Meanwhile, the introduction of the Transformer architecture has provided new ideas for image enhancement and restoration tasks, but its standard structure still suffers from limitations such as insufficient perception range, weak feature contrast, and high computational cost when facing irregular glare and multi-scale interference in nighttime road images, making it difficult to meet the dual requirements of real-time performance and accuracy in vehicular scenarios.

[0003] In summary, existing technologies have not yet been able to simultaneously achieve accurate identification of glare regions and effective restoration of local structural details in nighttime images under low computational resources. There is an urgent need for a lightweight deglare method that can restore glare background details to improve the perception reliability and safety of nighttime autonomous driving systems. This invention aims to address the above problems by proposing a glare suppression method for low-light nighttime driving images based on an improved Uformer network structure. This method accurately locates glare regions through a gated position attention module and enhances feature representation capabilities by combining a local context-aware structure, thereby restoring the realistic scene content while preserving detailed structure. Furthermore, a joint loss mechanism optimizes the glare suppression and background restoration effects, further improving image quality and structural consistency. Summary of the Invention

[0004] This invention addresses the shortcomings of existing technologies in glare suppression for low-light images of nighttime driving, such as inaccurate glare region localization, insufficient restoration of local structural details, and excessive computational resource requirements. It proposes a glare suppression method for low-light images of nighttime driving based on an improved Uformer network structure. This method solves the problems of local overexposure and glare interference caused by strong light sources under complex lighting conditions through a technical approach involving data acquisition and fusion, extraction of local details and structural information, precise glare region localization, and background detail restoration.

[0005] To achieve the above objectives, the present invention is implemented according to the following technical solution:

[0006] This invention includes the following steps:

[0007] S1: Collect nighttime road scene images and glare images; fuse the collected data and save them as a low-light, glare-interference image dataset.

[0008] S1-1: Use a high dynamic range camera to capture glare-free nighttime road background and glare images in natural nighttime scenes. The background images cover urban roads, highways, and rural roads, while the glare images are from glare generated by streetlights, car taillights, and car high beams.

[0009] S1-2: Perform gamma correction on the background image and glare image respectively;

[0010] S1-3: Control the fusion ratio of the background image and the glare image to synthesize an image containing glare;

[0011] S1-4: Perform inverse gamma correction on the synthesized image;

[0012] S1-5: Crop, scale, and trim the boundaries of the synthesized image to ensure that the generated image has a uniform size and limit the pixel values ​​to the legal range [0,1] to obtain a dataset of background images with glare.

[0013] S2: The Uformer anti-glare framework employs a depthwise separable convolution and pixel shuffle module to achieve efficient sampling operations, and extracts local details and structural information through a local context-aware Transformer.

[0014] The lightweight Uformer anti-glare frame includes:

[0015] An encoder-decoder framework is adopted, with one Transformer block used in each encoding and decoding layer;

[0016] The encoding stage introduces a lightweight downsampling module (L-Downsampling), which combines depthwise separable convolution with pointwise convolution (PWConv).

[0017] The decoding stage introduces a residual pixel shuffle upsampling module (L-Upsampling), which includes the main path and the residual path.

[0018] S3: Utilizing the gated position attention module of the Uformer model, combined with position encoding and local context information, the glare area is accurately located;

[0019] The implementation of the gating position attention module includes the following steps:

[0020] S3-1: Construct a normalized position encoding matrix, where each pixel position contains its horizontal and vertical coordinate information;

[0021] S3-2: Concatenate the input feature map with the location code along the channel dimension;

[0022] S3-3: Perform 1×1 convolution to reduce the dimensionality of the fused feature map;

[0023] S3-4: Local spatial context information is extracted through depthwise separable convolution and depthwise separable dilated convolution, the receptive field is expanded and a wider range of background information is captured. The output features of the two are concatenated by channels and then fused by 1×1 convolution.

[0024] S3-5: Design a gated branch. By applying convolution and sigmoid activation to the original feature map, gated weights are generated to adjust the importance of features in different regions. The enhanced features are then gated and fused with the original input features to obtain the final output features.

[0025] S4: The Uformer model is jointly trained using background loss, glare loss and reconstruction loss functions to achieve background detail recovery and glare prediction, thereby completing glare suppression in low-light images.

[0026] The joint training of background loss, glare loss, and reconstruction loss functions specifically includes:

[0027] Calculate the background loss term:

[0028]

[0029] Where L1 represents L1 loss, L vgg Represents the VGG loss, where I0 is the background image. The image is a glare-free background image predicted by the network.

[0030] Calculate the glare loss term:

[0031]

[0032] Where L1 represents L1 loss, L vgg Represents the VGG loss, where F is the true glare image. The glare image predicted by the network;

[0033] Introducing the reconstruction loss term L rec :

[0034]

[0035] in, Represents element addition. The background image predicted by the network without glare. I is the glare image predicted by the network, and Clip(·) means cropping the result to the interval [0,1] to ensure that the pixel values ​​are within the valid range;

[0036] Construct the total loss function

[0037]

[0038] Among them, L B L represents the background loss. F Indicates glare loss, L rec The value represents the reconstruction loss, where w1, w2, and w3 represent the background loss, glare loss, and weighting coefficients of the reconstruction loss, respectively.

[0039] The beneficial effects of this invention are:

[0040] By employing an improved Uformer network structure, the problem of accurately locating glare areas in low-light nighttime images is solved, and road detail information in glare-obscured areas is effectively recovered. The gated position attention module in this method can accurately locate glare areas caused by strong light sources such as vehicle headlights and streetlights, while the Local Context Aware Feedforward Neural Network (LCAFFN) in the Local Context Aware Transformer module enhances the extraction of local texture and structural information through a convolutional structure combined with a channel attention mechanism. Furthermore, the lightweight upsampling and downsampling module reduces upsampling artifacts while maintaining model computational efficiency, balancing real-time performance and accuracy requirements. This method is applicable to electronic rearview mirrors and forward-looking perception systems in intelligent driving assistance systems, significantly improving image clarity and perception reliability in nighttime environments, ensuring driver visual recognition capabilities, and enhancing driving safety. It has good engineering practical value and promising prospects for widespread application. Attached Figure Description

[0041] Figure 1 This is a general flowchart of an embodiment of the present invention;

[0042] Figure 2 This is a diagram of a nighttime road image glare suppression model according to an embodiment of the present invention;

[0043] Figure 3 This is a lightweight Uformer structure according to an embodiment of the present invention;

[0044] Figure 4 This is a local context-aware Transformer module in an embodiment of the present invention;

[0045] Figure 5 This is a gating position attention module according to an embodiment of the present invention;

[0046] Figure 6 This is a schematic diagram of the original nighttime image according to an embodiment of the present invention;

[0047] Figure 7 This is a schematic diagram of the original nighttime image of an embodiment of the present invention, showing the effect of glare suppression. Detailed Implementation

[0048] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. The illustrative embodiments and descriptions herein are used to explain the present invention, but are not intended to limit the present invention.

[0049] This invention provides a nighttime road glare removal method based on a Uformer structure, wherein: L-Downsampling: a lightweight downsampling module; L-Upsampling: a residual pixel rearrangement (PixelShuffle) upsampling module; WMSA: a window-based multi-head self-attention mechanism; LCAFFN: a local context-aware feedforward neural network; the flowchart is as follows. Figure 1 As shown, the deep learning model framework is as follows: Figure 2 As shown. The specific implementation method is as follows:

[0050] S1: Acquire nighttime road scene images and glare images. The acquired data is then fused and saved as a low-light, glare-affected image dataset. The specific steps for acquiring and fusing nighttime road scene images and glare images include the following:

[0051] S1-1: Use a high dynamic range camera to acquire glare-free nighttime road background and glare images in natural nighttime scenes. The background images cover urban roads, highways, and rural roads; the glare images are from streetlights, car taillights, and glare generated by car high beams.

[0052] S1-2: For background image I b and glare image I f Gamma correction was performed separately to simulate the nonlinear effect of lighting conditions on image brightness. The correction formula is as follows:

[0053]

[0054] The gamma value γ is randomly sampled from a Gaussian distribution, γ ~ μ(1.8, 2.2), to enhance the diversity and realism of the synthesized images. b Background image I b I f This is a glare image.

[0055] S1-3: Controls the fusion ratio between the background image and the glare image. The method for generating a composite image containing glare is as follows:

[0056]

[0057] Where M f This is a binarized glare mask used to control the shape and position of the glare area. α is the sampling mixing coefficient, α ~ μ (0.5, 1.0). and These are the background image and glare image after gamma correction, respectively.

[0058] S1-4. Perform inverse gamma correction on the synthesized image to restore linear brightness:

[0059]

[0060] Among them, I b-f It is the synthesized background image with glare, where γ is randomly sampled from a Gaussian distribution, γ ~ μ(1.8, 2.2).

[0061] S1-5: Crop, scale, and trim the boundaries of the synthesized image to ensure uniform image size and limit pixel values ​​to the legal range [0,1]. Obtain a dataset of background images with glare. Save the background image, glare image, and synthesized image with glare interference as an image dataset.

[0062] S2: The Uformer anti-glare framework employs a depthwise separable convolution and pixel rearrangement module to achieve efficient sampling operations, and extracts local details and structural information through a local context-aware Transformer.

[0063] The constructed Uformer anti-glare framework is used for sampling, and local details and structural information of the image are extracted, such as... Figure 2 As shown, the specific steps include the following:

[0064] S2-1: Input the image dataset obtained in step S1 into the constructed lightweight Uformer anti-glare framework. Use an encoder-decoder framework, with a Transformer block used in each encoding and decoding layer to reduce the computational complexity of the model.

[0065] S2-2: As Figure 3 As shown, in the lightweight Uformer architecture, a lightweight downsampling module (L-Downsampling) is introduced in the encoding stage. This module combines depthwise separable convolution and pointwise convolution (PWConv) to achieve efficient feature downsampling, preserve spatial information, and extract image features, such as... Figure 3 The lightweight downsampling module (L-Downsampling) is shown in the image.

[0066] L-Downsampling=PWConv(DWConv 3×3 (X))

[0067] Where X is the input, DWConv 3×3 (·) represents a 3×3 depthwise separable convolution, and PWConv(·) represents a pointwise convolution.

[0068] S2-3: The decoding stage introduces a residual pixel shuffle upsampling module (L-Upsampling) to replace the deconvolution operation. This module contains two paths: the main path and the residual path.

[0069] The main path first uses 3×3 convolution to extract local features, and then uses pixel shaving to achieve spatial upsampling:

[0070] U m =PixelShuffle(Conv 3×3 (X))

[0071] After adjusting the channels using 1×1 convolution in the residual path, upsampling is performed using bilinear interpolation:

[0072] U S =Bilinear(Conv 1×1 (X))

[0073] Finally, the output is merged, with the main path and residual path added together and then activated using ReLU:

[0074] U = ReLU(U m +U s )

[0075] Among them, Conv 3×3 (·) represents a convolution with a kernel size of 3×3, PixelShuffle(·) represents a pixel rearrangement operation, and Conv 1×1 (·) represents a convolution with a kernel size of 1×1, Bilinear(·) represents bilinear interpolation upsampling, ReLU represents the activation function, and U m Indicates the output result of the main path, U S This indicates the output result of the residual path.

[0076] S2-4: As Figure 4 As shown, the Transformer module, which is aware of local context, accurately extracts key structural features. This module first processes the input X... l-1 Layer Normalization (LN) is performed to standardize the feature distribution within each channel. The normalized result is then input into a window-based multi-head self-attention (WMSA) mechanism to extract global attention features within a local window. These features are then added to the input via residual connections to obtain the intermediate enhanced features.

[0077] X l =WMSA(LN(X) l-1))+X l-1

[0078] Wherein, WMSA represents a window-based multi-head self-attention mechanism, X l-1 It is the input of the current layer, and LN represents layer normalization.

[0079] S2-5: Enhanced feature X l The data is then normalized again and fed into a Local Context-Aware Feedforward Neural Network (LCAFFN). This module mainly consists of the following structure:

[0080] F1 = GELU(Conv 1×1 (X))

[0081]

[0082] Among them, Conv 1×1 It is a convolution with a kernel size of 1×1, GELU represents the GELU activation function, SE(·) represents the SE channel attention mechanism, and DWConv 3×3 (·) indicates a depthwise separable convolution with a kernel size of 3×3. This represents element-wise addition.

[0083] S2-6: Add the residuals of the result processed by the local channel enhancement feedforward network and the intermediate enhancement features:

[0084] X l =LCAFFN(LN(X) l ′))+X l ′

[0085] Where LCAFFN represents a Local Context-Aware Feedforward Neural Network, X l ' represents intermediate layer features, LN represents layer normalization, X l This represents the feature map after processing by the entire local context-aware Transformer.

[0086] S3: Locate the glare area. Using a gated position attention module, combined with position encoding and local contextual information, the glare area is precisely located.

[0087] In this example, step S3, which locates the glare area, specifically includes the following sub-steps:

[0088] S3-1: As Figure 5 As shown, the feature map processed by the local context-aware Transformer is used to locate the positions of different glare by passing through gated position attention via skip connections.

[0089] S3-2: Construct a normalized position encoding matrix P from the feature map, where each pixel position contains its horizontal and vertical coordinate information. This position encoding is generated using linear interpolation, calculated as follows:

[0090] P=Concat(linspace(-1,1,H),linspace(-1,1,W))

[0091] Where H is the feature map height, W is the feature map width, linspace(-1,1,H) represents generating a linear sequence of length H in the vertical direction to generate normalized row position information, and linspace(-1,1,W) represents generating a linear sequence of length W in the horizontal direction to generate normalized column position information. The final position embedding matrix P is obtained.

[0092] S3-3: Concatenate the input feature map X with the location code P along the channel dimension to generate a feature map X that incorporates location information. p :

[0093] X p =Concat(X,P)

[0094] Where Concat(·) represents channel concatenation, X represents the input feature map, and P represents position encoding.

[0095] S3-4: Fusing feature map X p Perform 1×1 convolutional dimensionality reduction to reduce the number of channels and compress redundant information, obtaining the dimensionality-reduced feature X. r :

[0096] X r =Conv 1×1 (X p )

[0097] Among them, X p Represents the fused feature map, Conv 1×1 It is a convolution with a kernel size of 1×1.

[0098] S3-5: Local spatial context information is extracted through depthwise separable convolution and depthwise separable dilated convolution, expanding the receptive field and capturing a wider range of background information. The output features of both are concatenated by channels and then fused using a 1×1 convolution.

[0099] X d =DWConv 3×3 (X r )

[0100] X k =DDWConv 3×3 (X r )

[0101] X f =Conv 1×1 (Concat(X d ,X k ))

[0102] Among them, X r It is a dimensionality reduction feature, DWConv 3×3 (·) indicates a depthwise separable convolution with a kernel size of 3×3, DDConv 3×3 (·) represents a 3×3 depthwise separable dilated convolution with a kernel size of 3×3. Concat(·) represents channel concatenation. Conv 1×1 It is a convolution with a kernel size of 1×1, X d It is the feature map after depthwise separable convolution processing, X k It is the feature map after depthwise separable dilated convolution, X f It is an enhanced feature resulting from the fusion of two branches.

[0103] S3-6: Design a gated branch. By applying convolution and sigmoid activation to the original feature map X, gate weights G are generated to adjust the importance of features in different regions, thus enhancing feature X. f The final output features are obtained by gating and fusing them with the original input features:

[0104] G=σ(Conv 1×1 (X))

[0105] Y = X⊙G + X f ⊙(1-G)

[0106] Where X is the input, Conv 1×1 It is a convolution with a kernel size of 1×1, where G is the gate weight and X is the weight. f It is the result of merging two branches, and ⊙ represents element-wise multiplication.

[0107] S4: Restore image background details. The network is jointly trained using background loss, glare loss, and reconstruction loss to achieve background detail restoration and glare prediction.

[0108] In this example, the restoration of image background details specifically includes the following sub-steps:

[0109] S4-1: Calculate the background loss term:

[0110]

[0111] Where L1 represents L1 loss, L vgg Represents the VGG loss, where I0 is the background image. The image is a glare-free background image predicted by the network.

[0112] S4-2: Calculate the glare loss term:

[0113]

[0114] Where L1 represents L1 loss, L vgg Represents the VGG loss, where F is the true glare image. The image shows the glare predicted by the network.

[0115] S4-3: Introducing the reconstruction loss term L rec This is used to ensure the physical consistency of the images predicted by the network, i.e.:

[0116]

[0117] in, Represents element addition. The background image predicted by the network without glare. I is the glare image predicted by the network, and Clip(·) means cropping the result to the interval [0,1] to ensure that the pixel values ​​are within the valid range.

[0118] S4-4: Constructing the total loss function The three types of losses are weighted and combined to simultaneously optimize glare suppression, background restoration, and image consistency:

[0119]

[0120] Among them, L B L represents the background loss. F Indicates glare loss, L rec The value represents the reconstruction loss, where w1, w2, and w3 represent the background loss, glare loss, and weighting coefficients of the reconstruction loss, respectively.

[0121] In this embodiment, the effect of nighttime road images after glare suppression is as follows: Figure 6 , Figure 7 As shown. Thanks to the lightweight Uformer glare suppression network proposed in this invention, glare is suppressed, and details obscured by glare in the road are significantly restored. In particular, the gated position attention module designed in this method can accurately locate glare areas caused by strong light sources such as car headlights and streetlights in the image. At the same time, the Local Context Aware Feedforward Neural Network (LCAFFN) in the proposed Local Context Aware Transformer effectively improves the ability to extract local texture and structural information through convolutional structure combined with channel attention mechanism; while the lightweight upsampling module balances efficiency and preservation of spatial details, reducing upsampling artifacts while maintaining model computational efficiency. Figure 6 , 7 As can be seen, the method of the present invention achieves the suppression of glare in nighttime road images and restores the true structure of the road in the glare-obscured parts.

[0122] In summary, this invention proposes a glare suppression method for low-light images of vehicles driving at night. Based on an improved Uformer network structure, it effectively solves the problems of local overexposure and glare interference caused by strong light sources under complex lighting conditions. This method utilizes a gated positional attention module to guide the network to focus on areas of strong light interference and achieve targeted suppression. Simultaneously, it enhances feature representation capabilities through a local context-aware structure, restoring realistic scene content while preserving detailed structure. Furthermore, this invention employs a joint loss mechanism to jointly supervise the predicted background, glare components, and image reconstruction, enabling the model to explicitly learn glare component separation and scene content restoration tasks during training, further improving image quality and structural consistency.

[0123] The method of this invention is particularly applicable to electronic rearview mirrors and forward perception systems in intelligent driving assistance systems. It effectively improves image clarity and perception reliability in nighttime environments, ensures the driver's visual recognition ability, enhances driving safety, and has good engineering practical value and promotion prospects.

[0124] The technical solutions of the present invention are not limited to the specific embodiments described above. Any technical modifications made in accordance with the technical solutions of the present invention fall within the protection scope of the present invention.

Claims

1. A method for suppressing glare in low-light images during nighttime driving, characterized in that, Includes the following steps: S1: Collect nighttime road scene images and glare images; fuse the collected data and save them as a low-light, glare-interference image dataset. S2: The Uformer anti-glare framework employs a depthwise separable convolution and pixel shuffle module to achieve efficient sampling operations, and extracts local details and structural information through a local context-aware Transformer. S3: Utilizing the gated position attention module of the Uformer model, combined with position encoding and local context information, the glare area is accurately located; S4: The Uformer model is jointly trained using background loss, glare loss and reconstruction loss functions to achieve background detail recovery and glare prediction, thereby completing glare suppression in low-light images.

2. The method for suppressing glare in low-light images during nighttime driving according to claim 1, characterized in that, Step S1 includes the following steps: S1-1: Use a high dynamic range camera to capture glare-free nighttime road background and glare images in natural nighttime scenes. The background images cover urban roads, highways, and rural roads, while the glare images are from glare generated by streetlights, car taillights, and car high beams. S1-2: Perform gamma correction on the background image and glare image respectively; S1-3: Control the fusion ratio of the background image and the glare image to synthesize an image containing glare; S1-4: Perform inverse gamma correction on the synthesized image; S1-5: Crop, scale, and trim the boundaries of the synthesized image to ensure that the generated image has a uniform size and limit the pixel values ​​to the legal range [0,1] to obtain a dataset of background images with glare.

3. The method for suppressing glare in low-light images during nighttime driving according to claim 1, characterized in that, The lightweight Uformer anti-glare frame in step S2 includes: An encoder-decoder framework is adopted, with one Transformer block used in each encoding and decoding layer; The encoding stage introduces a lightweight downsampling module (L-Downsampling), which combines depthwise separable convolution with pointwise convolution (PWConv). The decoding stage introduces a residual pixel shuffle upsampling module (L-Upsampling), which includes the main path and the residual path.

4. The method for suppressing glare in low-light images during nighttime driving according to claim 1, characterized in that, The implementation of the gating position attention module in step S3 includes the following steps: S3-1: Construct a normalized position encoding matrix, where each pixel position contains its horizontal and vertical coordinate information; S3-2: Concatenate the input feature map with the location code along the channel dimension; S3-3: Perform 1×1 convolution to reduce the dimensionality of the fused feature map; S3-4: Local spatial context information is extracted through depthwise separable convolution and depthwise separable dilated convolution, the receptive field is expanded and a wider range of background information is captured. The output features of the two are concatenated by channels and then fused by 1×1 convolution. S3-5: Design a gated branch. By applying convolution and sigmoid activation to the original feature map, gated weights are generated to adjust the importance of features in different regions. The enhanced features are then gated and fused with the original input features to obtain the final output features.

5. The method for suppressing glare in low-light images during nighttime driving according to claim 1, characterized in that, The joint training of background loss, glare loss, and reconstruction loss function in step S4 specifically includes: Calculate the background loss term: Where L1 represents L1 loss, L vgg Represents the VGG loss, where I0 is the background image. The image is a glare-free background image predicted by the network. Calculate the glare loss term: Where L1 represents L1 loss, L vgg Represents the VGG loss, where F is the true glare image. The glare image predicted by the network; Introducing the reconstruction loss term L rec : in, Represents element addition. The background image predicted by the network without glare. I is the glare image predicted by the network, and Clip(·) means cropping the result to the interval [0,1] to ensure that the pixel values ​​are within the valid range; Construct the total loss function Among them, L B L represents the background loss. F Indicates glare loss, L rec The value represents the reconstruction loss, where w1, w2, and w3 represent the background loss, glare loss, and weighting coefficients of the reconstruction loss, respectively.

Citation Information

Cited By

  • Low-light enhancement and glare processing method and device for image data and electronic equipment

    CN121724882A

  • Low-light enhancement and glare processing method and device of image data and electronic equipment

    CN121724882B