An Image Enhancement Method and Device Based on Multi-Perspective Input and Structure Guidance

Through the multi-view input and structure-guided image enhancement method, combined with cross-channel attention feature fusion and structural information guidance, the problem of poor image enhancement effect in the prior art is solved, the image details and visual effects are improved, and the calculation complexity is reduced.

CN119399087BActive Publication Date: 2025-07-08SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411553764.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-03
Publication Date
2025-07-08
Estimated Expiration
2044-11-03

AI Technical Summary

Technical Problem

Existing image enhancement technologies are difficult to effectively deal with complex image quality problems, such as degradation of image quality under noise, blur or low light conditions, and deep learning models are not effective in global structure and long-distance dependence, resulting in loss of local details and high computational complexity.

Method used

The image enhancement method of multi-view input and structure guidance is adopted, combined with multi-view input enhancement network and structure guidance enhancement network, through cross-channel attention feature fusion, channel group axial Transformer module and structural information guidance, the color accuracy and local contrast of the image are improved, details and boundary information are preserved, and calculation complexity is reduced.

Benefits of technology

Improve image details and visual effects, enhance image robustness and computing efficiency, especially in complex scenes or low-quality image processing, reducing noise and artifacts, and improving image contrast and sharpness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399087B_ABST
    Figure CN119399087B_ABST
Patent Text Reader

Abstract

The present invention discloses a structure-guided multi-prior input image enhancement method and apparatus, including the following steps: Step 1: Collect low-quality images and corresponding high-quality images to form a data set; divide the data set into a training set, a validation set, and a test set; Step 2: Construct a multi-view input enhancement network, which is used to perform preliminary enhancement on low-quality images; and design a multi-view enhancement joint loss function L mve , specify the number of training epochs, train and optimize the multi-view input enhancement network model, and save the weights of the optimal model; Step 3: Construct a structure-guided enhancement network, which performs secondary enhancement on image data based on the structure information extracted by the structure extraction module; and design a structure-guided enhancement joint loss function L sge , specify the number of training epochs, train and optimize the structure-guided enhancement network model, and save the weights of the optimal model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image enhancement or restoration classification, and particularly to an image enhancement method and device based on multi-view input and structure guidance. Background Art

[0002] Image enhancement technology processes image data to improve its visual quality and information content, thereby enhancing the effects of subsequent processing tasks. Image enhancement technology plays a crucial role in the field of image processing and is widely used in multiple fields such as medical imaging, remote sensing images, video surveillance, etc. to optimize the visibility, contrast, and detail performance of images.

[0003] In the prior art, the following problems exist:

[0004] 1. Traditional image enhancement algorithms usually use various filters and transformation methods to improve image quality. For example, histogram equalization can enhance the contrast of an image, while a high-pass filter can enhance the edge details of an image. However, these traditional methods are often limited to rule-based processing methods and may not be able to effectively handle complex image quality problems such as noise, blur, or image quality degradation under low-light conditions;

[0005] 2. In recent years, deep learning methods have also made remarkable progress in the field of image enhancement. Deep learning models such as convolutional neural networks (CNNs) and Transformers can achieve more advanced image enhancement by learning complex features of image data. However, convolutional neural networks (CNNs) mainly focus on local information and are difficult to capture global context, resulting in poor performance when dealing with global structures and long-range dependencies; although existing Vision Transformers (ViTs) have a global receptive field, their image patch embedding part may cause local detail loss and artifacts or incoherence at the image patch boundaries, affecting the final restoration quality, and the computational complexity of ViTs is O(n 2 ), which leads to high computational and storage requirements, thus resulting in efficiency issues;

[0006] 3. In deep neural networks, although rich semantic information is usually extracted in the deeper layers of the network, this process often leads to the abstraction and loss of important information such as structural features, edges, and textures, thereby affecting the effect of image enhancement. Summary of the Invention

[0007] In view of the above problems, the present invention provides an image enhancement method and device that combines multi-view input and structure guidance.

[0008] This method comprehensively improves the color accuracy and local contrast of images by adopting multi-perspective input. By combining different traditional image enhancement algorithms, it can more effectively handle various image problems, enhance the details and visibility of images, and avoid the deficiencies that may be brought by a single input method, thereby improving the image quality and visual effects.

[0009] The proposed channel group axial Transformer module in this method has no image patch embedding part, which can retain all the details of the original image, avoid information loss or boundary effects that may occur during image patch division, and can more effectively preserve and restore detail and boundary information. Moreover, the channel group axial Transformer has both a global receptive field and a linear computational complexity.

[0010] This method uses structural information for secondary guided enhancement to supplement and strengthen the network's extraction of structural information, which helps to make up for the deficiencies of deep networks in processing structural information and ensures that the structure and details of the image are better retained and enhanced.

[0011] The present invention adopts the following technical solutions:

[0012] An image enhancement method based on multi-perspective input and structure guidance, comprising the following steps:

[0013] S1: Collect low-quality images and their corresponding high-quality images to form a dataset; divide the dataset into a training set, a validation set, and a test set.

[0014] S2: Construct a multi-perspective input enhancement network, which is used to initially enhance the low-quality images; and design a multi-perspective enhancement joint loss function L mve , specify the number of training epochs, train and optimize the multi-perspective input enhancement network model, and save the weights of the optimal model.

[0015] The multi-perspective input enhancement network includes multi-perspective input (the original input and its corresponding white balance and contrast-limited adaptive histogram equalization images), a cross-channel attention feature fusion (CCAF, Channel-Cross Attention Fusion) module, and a channel group axial Transformer (CGAT, Channel-Group Axial Transformer) module; the multi-perspective enhancement joint loss function L mve includes L1 loss and perceptual loss.

[0016] S3: Construct a structure-guided enhancement network, which performs secondary enhancement on the image data based on the structural information extracted by the structure extraction module; and design a structure-guided enhancement joint loss function L sge, specify the number of training rounds, train and optimize the structure-guided enhancement network model, and save the weights of the optimal model;

[0017] The structure-guided enhancement network includes a structure extraction module and a structure-guided enhancement module (SGEM, Structure Guided Enhancement Module); the structure-guided enhancement joint loss function Ls sge includes L1 loss, perceptual loss, and structural similarity index loss;

[0018] S4: Construct an enhancement network that combines multi-view input and structure guidance. This network combines the trained multi-view input enhancement network and the structure-guided enhancement network to obtain the final image enhancement result;

[0019] S5: Test the model performance.

[0020] As a preferred solution of the present invention, an image enhancement method based on multi-view input and structure guidance, the multi-view input enhancement network in S2:

[0021] The multi-view input enhancement network consists of two parts, namely multi-view input feature fusion and an asymmetric encoder-decoder structure network.

[0022] In the multi-view input feature fusion part, the multi-view input includes the original input and its corresponding WB and CLAHE images. The three inputs of the multi-view input data are mapped to a new feature space through the same convolutional layer, and the three groups of feature maps obtained are fused using a cross-channel attention feature fusion module to obtain the feature fusion result FU;

[0023] The encoder part has three layers of networks, and the feature fusion result FU is used as the input of the three layers of the encoder network. Among them, each layer of the network is composed of a CGAT module, and the output of each layer will be element-wise added to the feature fusion result FU, and the result is respectively subjected to two downsampling operations. Each layer of the network thus obtains three different-scale feature maps, and then the feature maps of the same scale in different layers are connected in the channel dimension to obtain three groups of different-scale feature maps, which are FS1, FS2, and FS3 in increasing order of scale;

[0024] The decoder part has three layers. Each layer of the decoder consists of a Squeeze-and-Excitation Network and a CGAT module. The output feature maps of the encoder at three different scales will be used as the inputs of the three-layer network of the decoder. That is, FS1, FS2, and FS3 are the inputs of the bottom layer, the middle layer, and the top layer of the decoder respectively. Each group of feature maps will first pass through the Squeeze-and-Excitation Network of the corresponding layer to obtain three feature maps FS11, FS21, and FS31. Starting from the bottom layer, FS11 undergoes an upsampling operation after passing through the CGAT module. The result is concatenated with FS21 in the channel dimension. The resulting data is put into the CGAT module in the middle layer for upsampling, and its result is then concatenated with FS31 in the channel dimension. After passing through the CGAT module and the convolutional layer in the top layer, the multi-view input enhancement result OUT is obtained. mve 。

[0025] As a preferred embodiment of the present invention, an image enhancement method based on multi-view input and structure guidance, the cross-channel attention feature fusion module in S2:

[0026] The cross-channel attention feature fusion module can capture the relationship between different channels of multiple input feature maps based on the channel self-attention mechanism.

[0027] Specifically, N multi-view input feature maps with the shape of (C, H, W) are concatenated in the channel dimension to obtain a feature map S1 with the shape of (N×C, H, W). Among them, C is the number of channels, and H and W are the height and width of the feature map respectively. After the feature map S1 passes through a 1×1 convolution and a 3×3 depth convolution, its shape is reshaped into (N×C, H×W) to obtain a query matrix Q, a key matrix K, and a value matrix V. The Q matrix and the K matrix are multiplied to obtain a feature matrix with the shape of (N×C, N×C). Then, this feature matrix is multiplied with V, and the result is added to the feature map S1 element by element. Finally, after passing through a 1×1 convolution and Channel Shuffle, the feature fusion result is obtained.

[0028] As a preferred embodiment of the present invention, an image enhancement method based on multi-view input and structure guidance, the channel group axial Transformer module in S2:

[0029] The channel group axial Transformer module consists of Channel Shuffle, channel grouping, and axial Transformer. The axial Transformer consists of multiple layers, including an instance normalization layer, an axial self-attention layer, and a spatial gating feed-forward network layer.

[0030] The channel group axial Transformer module generates multiple feature maps after channel shuffling and channel grouping of the input of this module. Each feature map will pass through the axial Transformer respectively, and the results will be concatenated in the channel dimension to obtain the output feature map.

[0031] The spatial gated feed-forward network is a feed-forward network designed to utilize local context information and global spatial feature information, and realizes information fusion through three paths. This mechanism consists of three paths: one is the local context information path, one is the local context information gating path, and one is the global spatial feature information gating path. The feature map first passes through a 1×1 convolution to expand the number of channels to 3 times that of the input, and then passes through a 3×3 depth convolution and is divided into three groups of feature maps F1, F2, and F3. After F2 passes through the Gaussian Error Linear Unit (GELU) and then through the Sigmoid function to map the values to the interval (0, 1), the result is multiplied element-wise with F1 to obtain F11. After F3 passes through GELU, spatial attention calculation is performed. Specifically, the maximum value and average value of each pixel on all channels are calculated respectively, and the results are added to obtain the feature map FS3. FS3 uses the Sigmoid function to obtain the spatial weight. After F11 and FS3 are multiplied element-wise, the result passes through a 1×1 convolution to keep the number of channels consistent with the input channels and is used as the output of the CGAT module.

[0032] As a preferred solution of the present invention, an image enhancement method based on multi-view input and structure guidance, in the S2, the multi-view enhancement joint loss function L mve :

[0033] The joint loss function L MVE mainly consists of two parts: L1 loss and perceptual loss, as shown in formula (1). Among them, the L1 loss is used to measure the absolute difference between the predicted value and the actual value. It takes the absolute value of the error of each pixel and then calculates the average value of these absolute errors, as shown in formula (2). The perceptual loss is based on the intermediate layer features of a pre-trained deep network (such as VGG) and is used to measure the difference between the generated image and the target image in the feature space, rather than the pixel-level difference. It emphasizes the high-level semantics and details of the image, making the generated image visually closer to the real image, as shown in formula (3).

[0034] Loss mve =α1Loss per +α2Loss l1 (1)

[0035]

[0036] In the formula, Loss peris the perceptual loss, Loss l1 is the L1 loss, φ is the VGG16 network pre-trained on the ImageNet dataset, and H and W are the height and width of the feature map respectively. is the reconstructed image, and I is the ground truth label. and are the values of the corresponding pixels of the predicted image and the ground truth label in the pre-trained VGG16 network respectively. and I(m,n) are the pixel values of the predicted image and the ground truth label respectively. Where α1 and α2 are the weights of the L1 loss and the perceptual loss in the joint loss function L mve respectively.

[0037] As a preferred embodiment of the present invention, an image enhancement method based on multi-view input and structure guidance, the structure-guided enhancement network in S3:

[0038] The structure-guided enhancement network consists of two parts, namely a structure extraction module and a U-shaped network with an encoder-decoder structure.

[0039] In the structure extraction module part, the module first uses the white balance WB (White Balance) and the contrast-limited adaptive histogram equalization CLAHE (Contrast Limited Adaptive Histogram Equalization) algorithms to process the low-quality image, so that the color consistency, contrast, etc. of the image are improved to a certain extent, and then uses an edge detection algorithm (such as Canny edge detection) to extract the structure information of the image.

[0040] In the U-shaped network part, there are two inputs, namely the structure information extracted by the structure extraction module and the multi-view input enhancement result OUT mve . In the encoder part, there are a total of three network structures. Each layer of the network includes a convolutional block and a downsampling module, but the bottom layer only contains a structure-guided enhancement module; correspondingly, the decoder part also has three network structures. Each layer of the network includes a structure-guided enhancement module and an upsampling module, and the top layer only contains a structure-guided enhancement module. The decoder will start from the bottom layer and use the structural feature information for guided enhancement. Specifically, each layer will connect the input feature map with the output feature map of the corresponding layer of the encoder in the channel dimension, and the result is combined with the structure information through the structure-guided enhancement module and the upsampling module to obtain the output of this layer. Finally, the structure-guided enhancement result OUT sge is obtained.

[0041] As a preferred embodiment of the present invention, an image enhancement method based on multi-view input and structure guidance, the structure-guided enhancement module in S3:

[0042] The structure-guided enhancement module consists of multiple layers, including a structure-guided self-attention layer, an instance normalization layer, and a spatial gating feed-forward network layer, and has two inputs, namely the input feature map and the structure information.

[0043] For the structure-guided self-attention layer, specifically, the input feature map is first normalized by instance normalization to obtain the feature map S, and then the query matrix Q and the key matrix K are obtained through a 3×3 depth convolution; the structural feature information passes through a multilayer perceptron (MLP) to obtain two groups of feature maps α and γ, and the numerical matrix V is calculated through the formula S×(1 + α)+γ. The Q matrix and the K matrix are multiplied to obtain the feature matrix, and then the feature matrix is multiplied by V, and the result is reshaped in shape to be consistent with the input feature Figure 1 After that, it is added element-wise to the input feature map to obtain the output.

[0044] As a preferred embodiment of the present invention, an image enhancement method based on multi-view input and structure guidance, the structure-guided enhancement joint loss function L in S3 sge :

[0045] The joint loss function L sge mainly consists of three parts: L1 loss, perceptual loss, and structural similarity index loss, as shown in formula (4). Among them, the L1 loss is used to measure the absolute difference between the predicted value and the actual value. It takes the absolute value of the error of each pixel and then calculates the average of these absolute errors, as shown in formula (2). The perceptual loss is based on the intermediate layer features of a pre-trained deep network (such as VGG) and is used to measure the difference between the generated image and the target image in the feature space, rather than the pixel-level difference. It emphasizes the high-level semantics and details of the image, making the generated image visually closer to the real image, as shown in formula (3). The structural similarity index loss is a loss function that evaluates the similarity of images by comparing the brightness, contrast, and structural information of local regions of the images. The main goal is to ensure that the generated image is consistent with the reference image in terms of structure, brightness, and contrast, thereby improving the quality and visual effect of the image, as shown in formula (6).

[0046] Loss sge =β1Loss per +β2Loss l1 +β3Loss ssim (4)

[0047]

[0048] In the formula, Loss per is the perceptual loss, Loss l1is the L1 loss, Loss ssim is the structural similarity loss, where x represents each pixel value, and are the mean and standard deviation of the corresponding patch of the predicted image. Similarly, μ I (x) and σ I (x) are the mean and standard deviation of the image patch corresponding to the ground truth label I, while is the cross-covariance. Here, c1 and c2 are stable constants, and β1, β2, and β3 are the weights of the L1 loss, perceptual loss, and structural similarity index loss in the joint loss function L1, respectively.

[0049] As a preferred embodiment of the present invention, an image enhancement method based on multi-view input and structure guidance, the enhancement network based on multi-view input and structure guidance in S4:

[0050] The enhancement network based on multi-view input and structure guidance consists of two parts, namely the multi-view input enhancement network and the structure guidance enhancement network.

[0051] Specifically, the multi-view input enhancement network and the structure guidance enhancement network respectively adopt the corresponding optimal model weights saved after their training optimization. The original input and its corresponding WB and CLAHE images pass through the multi-view input enhancement network to obtain the multi-view input enhancement result OUT mve , the original input and OUT mve pass through the structure guidance enhancement network to obtain the structure guidance enhancement result OUT sge , OUT sge is the final image enhancement result.

[0052] Preferably, an image enhancement device based on multi-view input and structure guidance includes at least one processor and a memory; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any one of the above image enhancement methods based on multi-view input and structure guidance.

[0053] The beneficial effects of the present invention are:

[0054] (1) The present invention uses multi-view input to provide additional context information and constraints. By combining different perspective knowledge, it can specifically help the enhancement model to more accurately restore details, reduce noise and artifacts, and improve the contrast, clarity, and visual effects of the image, etc. Especially when dealing with complex scenes or low-quality images, it can enhance the robustness of the model;

[0055] (2) The proposed Cross-Channel Attention Fusion Module (CCAFM) in the present invention uses the self-attention calculation method in the channel dimension for multi-view inputs, enabling better fusion of multiple features. Finally, through Channel Shuffle, it improves the feature expression ability and enhances the generalization ability of the model.

[0056] (3) The proposed Channel Group Axial Transformer (CGAT) module in the present invention reduces the computational complexity in the self-attention calculation process by performing grouped calculations on the channel dimension, thereby accelerating the inference speed. And through Channel Shuffle, it prevents the possibility of locality in the channel combination. At the same time, the Spatial Gated Feed-Forward Network (SGFFN) effectively makes up for the lack of attention to the image spatial information during the calculation of axial self-attention.

[0057] (4) The proposed Structure-Guided Enhancement Network (SGEN) in the present invention performs secondary guidance and enhancement on the image through the structural information of the image. Using the structural information can enhance the details of the image, especially for regions with low signal-to-noise ratio, helping to generate clear and realistic results. Moreover, the edge information helps to distinguish different regions and establish better connections between these regions. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present invention and do not limit the present invention.

[0059] Figure 1 It is a flowchart of the method proposed by the present invention;

[0060] Figure 2 It is a network structure diagram of the method proposed by the present invention;

[0061] Figure 3 It is a structure diagram of the Channel Group Axial Transformer module of the method proposed by the present invention;

[0062] Figure 4 It is a structure diagram of the Structure-Guided Enhancement module of the method proposed by the present invention;

[0063] Figure 5 It is a structure diagram of the Spatial Gated Feed-Forward Network of the method proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the embodiments of the present invention with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0065] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0066] As Figure 1 , an image enhancement method based on multi-view input and structure guidance includes the following steps:

[0067] S1: Using the UIEB underwater image enhancement benchmark dataset, divide 800 pairs of images in the dataset into the training set, and divide the remaining 90 pairs of images into the test set;

[0068] S2: Construct a multi-view input enhancement network, which is used to preliminarily enhance low-quality images; and design a multi-view enhancement joint loss function L mve , specify the number of training rounds, train the multi-view input enhancement network model for 600 rounds, and save the weights of the optimal model.

[0069] As Figure 2 , the multi-view input enhancement network includes multi-view input (the original input and its corresponding white balance and contrast-limited adaptive histogram equalization image), a cross-channel attention feature fusion (CCAF, Channel-CrossAttention Fusion) module, and a channel-group axial Transformer (CGAT, Channel-Group AxialTransformer) module; the multi-view enhancement joint loss function L mve includes L1 loss and perceptual loss;

[0070] The multi-view input enhancement network consists of two parts, namely multi-view input feature fusion and an asymmetric encoder-decoder structure network.

[0071] The multi-view input feature fusion part. The multi-view input is used to provide additional context information and constraints to improve the image quality. Aiming at the problems of serious color imbalance and low contrast in underwater images, by combining two traditional enhancement algorithms, WB and CLAHE, the contrast, color, and visual effects of the image are improved. The input of 256×256×3 in these three dimensions is mapped to a new feature space through the same convolutional layer, and three groups of feature maps S1, S2, and S3 with the shape of (64, 256, 256) are obtained. S1, S2, and S3 use a cross-channel attention feature fusion module for feature fusion to obtain the feature fusion result FU;

[0072] Such as Figure 3 , in the encoder part, the feature fusion result FU is used as the input of the three-layer network in the encoder. Each layer of the network is composed of a channel group axial Transformer. The output of each layer will be element-wise added to the corresponding layer's feature maps S1, S2, and S3. The result is downsampled twice, and thus three different-scale feature maps are obtained for each layer. Then, the feature maps of the same scale in different layers are concatenated in the channel dimension to obtain three groups of feature maps of different scales, which are FS1, FS2, and FS3 in increasing order of scale, with the shapes of (384, 64, 64), (192, 128, 128), and (192, 256, 256) respectively;

[0073] In the decoder part, each layer of the decoder is composed of a Squeeze-and-Excitation Network and a channel group axial Transformer module. The three different-scale feature maps FS1, FS2, and FS3 of the encoder are the inputs of the bottom layer, middle layer, and top layer of the decoder respectively. Each feature map will first pass through the corresponding layer's Squeeze-and-Excitation Network to obtain three feature maps FS11, FS21, and FS31. Starting from the bottom layer, FS11 undergoes an upsampling operation after passing through the CGAT module, and its result is concatenated with FS21 in the channel dimension. The result is put into the CGAT module in the middle layer for upsampling, and its result is then concatenated with FS31 in the channel dimension. After passing through the CGAT module and convolutional layer in the top layer, the first-stage enhancement result is obtained.

[0074] The cross-channel attention feature fusion module is based on the channel self-attention mechanism, which can capture the relationships between different channels of multiple input feature maps, thereby enhancing the model's ability to express features.

[0075] Specifically, three multi-view input feature maps with the shape of (64, 256, 256) are concatenated in the channel dimension to obtain a feature map S1 with the shape of (192, 256, 256). After passing through a 1×1 convolution and a 3×3 depth convolution, the shape of the feature map is reshaped to (192, 65536) to obtain the Q, K, and V matrices. The Q matrix and the K matrix are multiplied to obtain a feature matrix with the shape of (192, 192), and then this feature matrix is multiplied with V. The result is added element-wise to the feature map S1, and finally, after a 1×1 convolution and Channel Shuffle, the feature fusion result is obtained. The channel self-attention mechanism can dynamically adjust the weights of different channels, enabling the model to better process complex input data and improving the model's expressive ability and performance; the purpose of Channel Shuffle is to improve the feature expression ability and enhance the generalization ability of the model.

[0076] The channel group axial self-attention module consists of Channel Shuffle, channel grouping, and an axial transformer. The axial transformer consists of multiple layers, including an instance normalization layer, an axial self-attention layer, and a spatial gating feed-forward network layer. Channel Shuffle aims to avoid locality caused by channel grouping, and channel grouping is to reduce the computational complexity during self-attention calculation.

[0077] Specifically, the channel group axial self-attention module generates multiple feature maps after passing the output through Channel Shuffle and channel grouping. Each feature map will perform axial self-attention calculation separately, and the results will be concatenated in the channel dimension to obtain the output feature map.

[0078] Such as Figure 5 The spatial gating feed-forward network module is a feed-forward network designed to utilize local context information and global spatial feature information, and it realizes information fusion through three paths. This mechanism consists of three paths: one is the local context information path, one is the local context information gating path, and one is the global spatial feature information gating path. In a transformer, the traditional feed-forward network maps the input features to a higher-dimensional feature space through its stacked fully connected layers to extract and represent complex patterns and features. However, relatively speaking, the traditional feed-forward network has no direct perception ability for the spatial layout and local features in the input data, while for image enhancement, both the spatial layout and local feature information are of crucial reference significance. The three paths of the spatial gating feed-forward network module aim to focus on the spatial layout and local feature information.

[0079] Specifically, after the feature map passes through a 3×3 depth convolution, it is divided into three groups of feature maps F1, F2, and F3. After F2 passes through the Gaussian Error Linear Unit (GELU), the Sigmoid function is used, and the result is multiplied element-wise with F1 to obtain F11. After F3 passes through GELU, spatial attention calculation is performed. Specifically, the maximum value and average value of each pixel across all channels are calculated separately, and the results are added to obtain the feature map FS3. FS3 uses the Sigmoid function to obtain the spatial weight. After F11 and FS3 are multiplied element-wise, the result is used as the output of the CGAT module.

[0080] The multi-view enhanced joint loss function L mve mainly consists of two parts: L1 loss and perceptual loss, as shown in Equation (1). Among them, the L1 loss is used to measure the absolute difference between the predicted value and the actual value. It takes the absolute value of the error of each pixel and then calculates the average of these absolute errors, as shown in Equation (2). The perceptual loss is based on the intermediate layer features of a pre-trained deep network (such as VGG) and is used to measure the difference between the generated image and the target image in the feature space, rather than the pixel-level difference. It emphasizes the high-level semantics and details of the image, making the generated image visually closer to the real image, as shown in Equation (3).

[0081] Loss mve = α1Loss per + α2Loss l1 (1)

[0082]

[0083] In the formula, Loss per is the perceptual loss, Loss l1 is the L1 loss, φ is the VGG16 network pre-trained on the ImageNet dataset, and H and W are the height and width of the feature map respectively. is the reconstructed image, and I is the ground truth label. Among them, α1 and α2 are the weights of the L1 loss and the perceptual loss in the joint loss function L1 respectively, and take the values of 0.5 and 1 in this embodiment.

[0084] S3: Construct a structure-guided enhancement network. The structure-guided enhancement network performs secondary enhancement on the image data based on the structure information extracted by the structure extraction module; and design the structure-guided enhancement joint loss function L sge , specify the number of training epochs, train the structure-guided enhancement network model for 600 epochs, and save the weights of the optimal model.

[0085] The described structure-guided enhancement network consists of two parts, namely the structure extraction module part and a U-shaped network with an encoder-decoder structure, including a structure extraction module and a structure-guided enhancement module (SGEM, Structure Guided Enhancement Module); the structure-guided enhancement joint loss function L sge includes L1 loss, perceptual loss, and structural similarity index loss;

[0086] The structure extraction module first uses the white balance WB (White Balance) and contrast-limited adaptive histogram equalization CLAHE (Contrast Limited Adaptive Histogram Equalization) algorithms to process the low-quality image, so as to improve the color consistency, contrast, etc. of the image to a certain extent, and then uses Canny edge detection to extract the structural information of the image.

[0087] For the U-shaped network part, there are two inputs, namely the structural information extracted by the structure extraction module and the multi-view input enhancement result OUT mve . In the encoder part, there are a total of three network structures. Each layer of the network includes a convolutional block and a downsampling module, but the bottom layer only contains the structure-guided enhancement module; correspondingly, the decoder part also has three network structures. Each layer of the network includes a structure-guided enhancement module and an upsampling module, and the top layer only contains the structure-guided enhancement module. The decoder will start from the bottom layer and use the structural feature information for guided enhancement. Specifically, each layer will connect the input feature map with the output feature map of the corresponding layer of the encoder in the channel dimension, and the result is combined with the structural information and passed through the structure-guided enhancement module and the upsampling module to obtain the output of this layer. Finally, the structure-guided enhancement result OUT is obtained from the top layer of the decoder sge .

[0088] Such as Figure 4 , the structure-guided enhancement module consists of multiple layers, including a structure-guided self-attention layer, an instance normalization layer, and a spatial gating feed-forward network layer, and has two inputs, namely the input feature map and the structural information. Specifically, the feature map first undergoes instance normalization (Instance Normalization) to obtain the feature map S, and then Q and K matrices are obtained through 3×3 depth convolution; the structural feature information passes through a multilayer perceptron MLP (Multilayer Perceptron) to obtain two groups of feature maps α and γ, and the V matrix is calculated through the formula S×(1 + α)+γ. The Q, K, and V matrices then obtain the output result through the self-attention calculation method.

[0089] The structure-guided enhancement joint loss function L sgeIt mainly consists of three parts: L1 loss, perceptual loss, and structural similarity index loss, as shown in Equation (4). Among them, the L1 loss is used to measure the absolute difference between the predicted value and the actual value. It takes the absolute value of the error for each pixel and then calculates the average of these absolute errors, as shown in Equation (2). The perceptual loss is based on the intermediate layer features of a pre-trained deep network (such as VGG) and is used to measure the difference between the generated image and the target image in the feature space rather than at the pixel level. It emphasizes the high-level semantics and details of the image, making the generated image visually closer to the real image, as shown in Equation (3). The structural similarity index loss is a loss function that evaluates the similarity of images by comparing the brightness, contrast, and structural information of local regions of the images. The main goal is to ensure that the generated image is consistent with the reference image in terms of structure, brightness, and contrast, thereby improving the quality and visual effect of the image, as shown in Equation (6).

[0090] Losssge = β1Loss per + β2Loss l1 + β3Loss ssim (4)

[0091]

[0092] where Loss per is the perceptual loss, Loss l1 is the L1 loss, Loss ssim is the structural similarity loss, x represents each pixel value, and are the mean and standard deviation of the predicted image patch. Similarly, μ I (x) and σ I (x) are the mean and standard deviation of the corresponding image patch of the ground truth label I, and is the cross-covariance. Among them, c1 and c2 are stable constants, which take the values of 0.02 and 0.03 respectively in this embodiment, and β1, β2, and β3 are the weights of the L1 loss, perceptual loss, and structural similarity index loss in the joint loss function L1 respectively, which take the values of 10, 1, and 1 respectively in this embodiment.

[0093] S4: Construct an enhancement network based on multi-view input and structure guidance. This network consists of two parts, namely the multi-view input enhancement network and the structure guidance enhancement network.

[0094] Specifically, the multi-view input enhancement network and the structure guidance enhancement network respectively adopt the corresponding optimal model weights saved after their training and optimization. The original input and its corresponding WB and CLAHE images pass through the multi-view input enhancement network to obtain the multi-view input enhancement result OUT mve , and the original input and OUT mve pass through the structure guidance enhancement network to obtain the structure guidance enhancement result OUTsge , OUT sge is the final image enhancement result.

[0095] S5: Test the model performance.

Claims

1. An image enhancement method based on multi-view input and structure guidance, characterized in that It includes the following steps: S1: Collect low-quality images and their corresponding high-quality images to form a dataset; divide the dataset into a training set, a validation set, and a test set; S2: Construct a multi-view input enhancement network, which is used to preliminarily enhance the low-quality images; And design a multi-view enhanced joint loss function L mve , specify the number of training rounds, train and optimize the multi-view input enhancement network model, and save the weights of the optimal model; The multi-view input enhancement network includes multi-view input, a cross-channel attention feature fusion module, and a channel group axial Transformer module; the multi-view enhanced joint loss function L mve includes an L1 loss and a perceptual loss; S3: Construct a structure-guided enhancement network. The structure-guided enhancement network performs secondary enhancement on the image data based on the structure information extracted by the structure extraction module; and design a structure-guided enhancement joint loss function L sge , specify the number of training epochs, train and optimize the structure-guided enhancement network model, and save the weights of the optimal model; The structure-guided enhancement network includes a structure extraction module and a structure-guided enhancement module; the structure-guided enhancement joint loss function L sge includes L1 loss, perceptual loss, and structural similarity index loss; S4: Construct an enhancement network that combines multi-view input and structure guidance. This network combines the trained multi-view input enhancement network and the structure-guided enhancement network to obtain the final image enhancement result; S5: Test the model performance; In S2, the multi-view input enhancement network is: The multi-view input enhancement network consists of two parts, namely multi-view input feature fusion and an asymmetric encoder-decoder structure network; For the multi-view input feature fusion part, the multi-view input includes the original input and its corresponding WB and CLAHE images. The multi-view input data is mapped to a new feature space through the same convolutional layer, and the three groups of feature maps obtained are fused using a cross-channel attention feature fusion module to obtain the feature fusion result FU; The encoder part has three layers of networks, and the feature fusion result FU is used as the input of the three-layer network of the encoder; among them, each layer of the network consists of a CGAT module, and the output of each layer will be element-wise added to the feature fusion result FU. The result is respectively subjected to two downsampling operations, so that each layer of the network obtains three feature maps of different scales. Then, the feature maps of the same scale in different layers are connected in the channel dimension to obtain three groups of feature maps of different scales, which are FS1, FS2, and FS3 in ascending order of scale; The decoder part has three layers. Each layer of the decoder consists of a compression activation network and a CGAT module. The output feature maps of three different scales of the encoder will be used as the inputs of the three-layer network of the decoder. That is, FS1, FS2, and FS3 are the inputs of the bottom layer, the middle layer, and the top layer of the decoder respectively. Each group of feature maps will first pass through the compression activation network of the corresponding layer to obtain three feature maps FS11, FS21, and FS31. Starting from the bottom layer, FS11 undergoes an upsampling operation after passing through the CGAT module. Its result is concatenated with FS21 in the channel dimension. The result is put into the CGAT module in the middle layer for upsampling, and its result is then concatenated with FS31 in the channel dimension. After passing through the CGAT module and the convolutional layer in the top layer, the multi-view input enhancement result OUT is obtained mve 。 2. The image enhancement method based on multi-view input and structure guidance according to claim 1, wherein In S2, the cross-channel attention feature fusion module is: The cross-channel attention feature fusion module is based on the channel self-attention mechanism, which can capture the relationship between different channels of the multi-input feature maps; N multi-view input feature maps with the shape of "C, H, W" are connected in the channel dimension to obtain a feature map S1 with the shape of "N×C, H, W", where C is the number of channels, and H and W are the height and width of the feature map respectively; the feature map S1 is reshaped into the shape of "N×C, H×W" after 1×1 convolution and 3×3 depth convolution to obtain the query matrix Q, the key matrix K, and the value matrix V. The Q matrix and the K matrix are multiplied to obtain a feature matrix with the shape of "N×C, N×C", and then this feature matrix is multiplied by the V matrix. The result is element-wise added to the feature map S1, and finally, after 1×1 convolution and channel shuffling, the feature fusion result is obtained.

3. A method for image enhancement based on multi-view input and structure guidance according to claim 1, characterized in that In S2, the channel group axial Transformer module is: The channel group axial self-attention module consists of channel shuffling, channel grouping, and axial Transformer. The axial Transformer consists of multiple layers of structures, including an instance normalization layer, an axial self-attention layer, and a spatial gating feed-forward network layer; The channel group axial self-attention module generates multiple feature maps after channel shuffling and channel grouping of the input of this module. Each feature map will pass through the axial Transformer respectively, and the results are then connected in the channel dimension to obtain the output feature map; The spatial gated feed-forward network is a feed-forward network that utilizes local context information and global spatial feature information, and achieves information fusion through three paths; this mechanism consists of three paths: one is the local context information path, one local context information gating path, and one global spatial feature information gating path; the feature map first undergoes 1×1 convolution to expand the number of channels to three times that of the input, and then after 3×3 depth convolution, it is divided into three groups of feature maps F1, F2, and F3. After F2 passes through the Gaussian error linear activation function GELU and then through the Sigmoid function to map the values to the interval (0, 1), the result is multiplied element-wise with F1 to obtain F11. After F3 passes through GELU, spatial attention calculation is performed. The maximum value and average value of each pixel on all channels are calculated separately, and the results are added to obtain the feature map FS3; FS3 uses the Sigmoid function to obtain the spatial weight. After F11 is multiplied element-wise with FS3, the result passes through 1×1 convolution to keep the number of channels consistent with the input channels and is used as the output of the CGAT module.

4. An image enhancement method based on multi-perspective input and structure guidance according to claim 1, characterized in that, In S2, the multi-view enhanced joint loss function L mve is as follows: Joint loss function L mve It mainly consists of two parts: L1 loss and perceptual loss, as shown in Equation (1); among them, the L1 loss is used to measure the absolute difference between the predicted value and the actual value. It takes the absolute value of the error for each pixel and then calculates the average of these absolute errors, as shown in Equation (2); the perceptual loss is based on the intermediate layer features of a pre-trained deep network and is used to measure the difference between the generated image and the target image in the feature space rather than at the pixel level. It emphasizes the high-level semantics and details of the image, making the generated image visually closer to the real image, as shown in Equation (3). Loss mve = α1Loss per + α2Loss l1 (1) where Loss per is the perceptual loss, Loss l1 is the L1 loss, φ is the VGG16 network pre-trained on the ImageNet dataset, and H and W are the height and width of the feature map, respectively; is the reconstructed image, and I is the ground truth label; and are the values of the corresponding pixels of the predicted image and the ground truth in the pre-trained VGG16 network, respectively; and I(m,n) are the pixel values of the predicted image and the ground truth, respectively; where α1 and α2 are the weights of the L1 loss and the perceptual loss in the joint loss function L mve respectively.

5. The image enhancement method based on multi-view input and structure guidance according to claim 1, wherein In S3, the structure-guided enhancement network is as follows: The structure-guided enhancement network consists of two parts, namely the structure extraction module and a U-shaped network with an encoder-decoder structure. The structure extraction module first uses the white balance and contrast-limited adaptive histogram equalization algorithms to process the low-quality image, which improves the color consistency, contrast, etc. of the image to a certain extent, and then uses the edge detection algorithm to extract the structure information of the image. U-shaped network part, with two inputs, namely the structure information extracted by the structure extraction module and the multi-view input enhancement result OUT mve ; In the encoder part, there are a total of three network structures. Each layer of the network includes a convolutional block and a downsampling module, but the bottom layer only contains a structure-guided enhancement module; correspondingly, the decoder part also has three network structures. Each layer of the network includes a structure-guided enhancement module and an upsampling module, and the top layer only contains a structure-guided enhancement module; The decoder will start from the bottom layer and use the structural feature information for guidance enhancement. Each layer will concatenate the input feature map with the output feature map of the corresponding layer of the encoder in the channel dimension. The result, together with the structural information, passes through the structure-guided enhancement module and the upsampling module to obtain the output of this layer. Finally, the structure-guided enhancement result OUT is obtained from the top layer of the decoder. sge 。 6. The image enhancement method based on multi-view input and structure guidance according to claim 1, wherein In S3, the structure-guided enhancement module is as follows: The structure-guided enhancement module consists of multiple layers, including the structure-guided self-attention layer, instance normalization layer, and spatial gated feed-forward network layer. The structure-guided enhancement module has two inputs, namely the input feature map and the structure information. For the structure-guided self-attention layer, specifically, the input feature map first undergoes instance normalization to obtain the feature map S, and then through 3×3 depth convolution, the query matrix Q and the key matrix K are obtained. The structure feature information passes through a multi-layer perceptron to obtain two groups of feature maps α and γ, and the numerical matrix V is calculated through the formula S×(1 + α)+γ; the Q matrix and the K matrix are multiplied to obtain the feature matrix, and then the feature matrix is multiplied with V, and the result is reshaped to be consistent with the input feature map and then added element-wise with the input feature map to obtain the output.

7. A method for image enhancement based on multi - perspective input and structure guidance according to claim 1, wherein In S3, the structure guides the enhanced combined loss function L sge which is: Joint loss function L sge It mainly consists of three parts: L1 loss, perceptual loss, and structural similarity index loss, as shown in Equation (4); among them, L1 loss is used to measure the absolute difference between the predicted value and the actual value. It takes the absolute value of the error for each pixel and then calculates the average of these absolute errors, as shown in Equation (2); perceptual loss is based on the intermediate layer features of a pre-trained deep network and is used to measure the difference between the generated image and the target image in the feature space rather than at the pixel level. It emphasizes the high-level semantics and details of the image, making the generated image visually closer to the real image, as shown in Equation (3); the structural similarity index loss is a loss function that evaluates the similarity of images by comparing the brightness, contrast, and structural information in the local regions of the images. The main goal is to ensure that the generated image is consistent with the reference image in terms of structure, brightness, and contrast, thereby improving the quality and visual effect of the image, as shown in Equation (6). Loss sge = β1Loss per + β2Loss l1 + β3Loss ssim (4) where Loss per is the perceptual loss, Loss l1 is the L1 loss, Loss ssim is the structural similarity loss, x represents each pixel value, and are the mean and standard deviation of the corresponding patch of the predicted image, μ I (x) and σ I (x) are the mean and standard deviation of the corresponding patch of the ground truth label I, while is the cross-covariance; where c1 and c2 are stable constants, and β1, β2, and β3 are the weights of the L1 loss, perceptual loss, and structural similarity index loss in the joint loss function L sge respectively.

8. A method for image enhancement based on multi-view input and structure guidance according to claim 1, characterized in that In S4, the enhancement network based on multi-view input and structure guidance is as follows: The enhancement network based on multi-view input and structure guidance consists of two parts, namely the multi-view input enhancement network and the structure-guided enhancement network. The multi-view input enhancement network and the structure-guided enhancement network respectively adopt the corresponding optimal model weights saved after their training optimization. The original input and its corresponding white balance and contrast-limited adaptive histogram equalization images pass through the multi-view input enhancement network to obtain the multi-view input enhancement result OUT mve , the original input and OUT mve pass through the structure-guided enhancement network to obtain the structure-guided enhancement result OUT sge , OUT sge is the final image enhancement result.

9. An image enhancement device based on multi-view input and structure guidance, characterized in that It includes at least one processor and a memory; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a method for enhancing a multi-view input image based on structure guidance as described in any one of claims 1 to 8.