Image enhancement processing method based on deep learning

Through deep learning-based image enhancement processing methods, multi-scale features are extracted and fused, and content-aware weight maps are generated by the generation of adversarial networks, the problem of traditional image enhancement technology lacks adaptability and correlation, and the image enhancement effect with higher quality and aesthetics is achieved.

CN120147180AActive Publication Date: 2025-06-13江西省科技基础条件平台中心(江西省计算中心)
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510626771.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-13
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

Traditional image enhancement processing technology lacks adaptability and is difficult to enhance targetedly according to the different image contents. It ignores the correlation and interaction between different regions in the image, resulting in the inadequate and coherent enhancement effect.

Method used

A deep learning-based image enhancement processing method is proposed. By extracting multi-scale features, recalibration and fusion, nonlinear transformation, decoder network reconstruction and generation of adversarial network generation content-aware weight maps, identifying key areas and associated areas, and further enhancing the overall effect of the image through interaction and collaboration between regions.

Benefits of technology

It achieves higher targeting and effectiveness of image enhancement, improves the overall effect and beauty of the image, and the generated images are clearer, more natural and more visual.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147180A_ABST
    Figure CN120147180A_ABST
Patent Text Reader

Abstract

The invention provides an image enhancement processing method based on deep learning. The method comprises the following steps: extracting multi-scale features of an input image; re-calibrating and fusing the multi-scale features to obtain a comprehensive feature map; carrying out nonlinear transformation; reconstructing the transformed features into a preliminary enhanced image by using a decoder network; a generative adversarial network is designed, a content perception weight map is generated, and each pixel value in the weight map reflects the importance of the corresponding position on the content; and according to the content perception weight map, identifying a key region and an associated region, and further enhancing the overall effect of the image through interaction and collaboration among the regions. According to the method, the pertinence and effectiveness of image enhancement are improved by extracting the multi-scale features and carrying out re-calibration and fusion, then recognition of different areas is guided according to the content perception weight map, collaborative enhancement among the areas is achieved, the overall effect of the image is improved, and the method is suitable for popularization and application. And the enhanced image is more natural and vivid in the aspects of details, colors, contrast and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an image enhancement processing method based on deep learning. Background Art

[0002] Image enhancement processing technology is an image processing means in the field of computer vision. Its main objective is to improve the image quality by processing the input image, thereby enhancing the recognition and detection performance of the computer vision system. It can be applied to various computer vision tasks, such as image recognition, image classification, object detection, etc.

[0003] Traditional image enhancement processing technologies usually have the following defects: 1. Lack of adaptability, it is difficult to perform targeted enhancement according to different image contents, and the enhancement process often ignores the importance differences of different regions in the image content, resulting in insufficiently fine enhancement effects; 2. Often ignores the relevance and interactivity between different regions in the image, resulting in localized and incoherent enhancement effects, making it difficult to achieve global optimization in the enhancement process, and thus difficult to improve the overall effect and aesthetic feeling of the image. Summary of the Invention

[0004] Based on this, the objective of the present invention is to propose an image enhancement processing method based on deep learning to solve the above-mentioned problems.

[0005] According to the image enhancement processing method based on deep learning proposed by the present invention, the method includes:

[0006] Extract multi-scale features of the input image;

[0007] Perform recalibration and fusion on the multi-scale features to obtain a comprehensive feature map;

[0008] Perform a non-linear transformation on the comprehensive feature map;

[0009] Use a decoder network to reconstruct the transformed features into a preliminary enhanced image;

[0010] Design a generative adversarial network, using the preliminary enhanced image as the input to generate a content-aware weight map, where each pixel value in the content-aware weight map reflects the importance of the corresponding position in terms of content;

[0011] According to the content-aware weight map, identify key regions and associated regions, and further enhance the overall effect of the image through the interaction and cooperation between regions.

[0012] Furthermore, the step of designing a generative adversarial network, using the preliminary enhanced image as the input to generate a content-aware weight map, includes:

[0013] Determine the structure of the generative adversarial network, including a generator and a discriminator;

[0014] Construct a training set based on a large amount of historical data. Each sample data in the training set contains a preliminary enhanced image and the corresponding true content-aware weight map to adaptively adjust the parameters of the generator to generate an image that better meets the expectations of the discriminator. The content-aware weight map is marked with labels of content importance;

[0015] Train the generator based on the generator loss;

[0016] Train the discriminator based on the discriminator loss;

[0017] Through the adversarial training of the generator and the discriminator, adaptively adjust the parameters of the generator to generate an image that better meets the expectations of the discriminator;

[0018] After the training is completed, input the preliminary enhanced image into the trained generator and generate a content-aware weight map with the same size as the preliminary enhanced image.

[0019] Furthermore, the step of training the generator based on the generator loss includes:

[0020] Randomly extract a preliminary enhanced image from the training set;

[0021] Use the generator to generate a content-aware weight map;

[0022] Input the generated content-aware weight map into the discriminator and calculate the generator loss. The calculation formula is:

[0023]

[0024] where, is the generator loss, is the content-aware weight map generated by the generator, z is the input of the generator, is random noise, is the output of the discriminator for indicating the probability that the discriminator believes is a true content-aware weight map;

[0025] Use the backpropagation algorithm to update the network parameters of the generator.

[0026] Furthermore, the step of training the discriminator based on the discriminator loss includes:

[0027] Initialize the network parameters of the generator and the discriminator;

[0028] Randomly extract a preliminary enhanced image and a marked true content-aware weight map from the training set;

[0029] Use the generator to generate a forged content-aware weight map;

[0030] Input the real and forged content perception weight maps into the discriminator, and calculate the discriminator loss. The calculation formula is as follows:

[0031] ,

[0032] where, is the discriminator loss, p is the real content perception weight map, and are the outputs of the discriminator for p and G(z) respectively, representing the probabilities that the discriminator believes p and G(z) are real content perception weight maps;

[0033] Update the network parameters of the discriminator through the backpropagation algorithm.

[0034] Furthermore, the steps of identifying the key regions and associated regions according to the content perception weight map and enhancing the overall effect of the image through the interaction and cooperation between regions include:

[0035] Define the key regions and associated regions according to the weight values in the content perception weight map;

[0036] Perform binarization processing on the content perception weight map to obtain a binary image containing only the key regions and associated regions;

[0037] Segment the binary image into different regions and assign unique region labels;

[0038] Identify the key regions and associated regions from the content perception weight map according to the region labels;

[0039] Calculate the similarity between two key regions with a connection relationship and the associated regions;

[0040] If the similarity is higher than the first threshold, strengthen the visual connection between the two regions;

[0041] If the similarity is lower than the second threshold, weaken or eliminate the visual interference between the two regions.

[0042] Furthermore, the steps of defining the key regions and associated regions according to the weight values in the content perception weight map include:

[0043] Define the regions with weight values higher than the third threshold as key regions;

[0044] Define the regions with weight values lower than the third threshold and connected to the key regions as associated regions.

[0045] Furthermore, the steps of recalibrating and fusing the multi-scale features to obtain the comprehensive feature map include:

[0046] Input multi-scale features into a deep clustering network, divide each scale of features into corresponding feature clusters, and recalibrate each scale of features according to the importance weights of the corresponding feature clusters;

[0047] Fuse the recalibrated features of each scale to obtain a comprehensive feature map.

[0048] Furthermore, the step of inputting multi-scale features into a deep clustering network, dividing each scale of features into corresponding feature clusters, and recalibrating each scale of features according to the importance weights of the corresponding feature clusters includes:

[0049] Input the extracted multi-scale features into the encoder of the deep clustering network to obtain the low-dimensional representations of each scale of features;

[0050] In the low-dimensional representation space, use a clustering algorithm to divide each scale of features into corresponding feature clusters respectively;

[0051] Predict the importance weights of each feature cluster through an MLP model;

[0052] Recalibrate each scale of features according to the importance weights corresponding to the feature clusters divided by each scale of features to obtain the recalibrated features of each scale. The formula is:

[0053]

[0054] where, is the recalibrated feature of the i-th scale, is the feature of the scale The importance weight corresponding to the divided feature cluster;

[0055] The step of fusing the recalibrated features of each scale to obtain a comprehensive feature includes:

[0056] Fuse the recalibrated features of each scale to obtain a fused feature X. The formula is:

[0057]

[0058] where, are the recalibrated features of each scale respectively, and M is the number of scale features.

[0059] Furthermore, the step of predicting the importance weights of each feature cluster through an MLP model includes:

[0060] Obtain the representative feature Y of each feature cluster from the deep clustering network and construct an input matrix , where, , n is the number of feature clusters, and d is the dimension of the representative feature of each feature cluster;

[0061] Label each representative feature in the input matrix according to the true label data related to the importance of the feature clusters;

[0062] Perform forward propagation on the input features through the MLP model to obtain the output prediction values;

[0063] Use the loss function to calculate the error between the prediction values and the true labels;

[0064] Calculate the gradient of the loss function with respect to the model parameters through the chain rule, and update the model parameters using the gradient descent method;

[0065] For the new feature clusters, extract their representative features and input them into the trained MLP model to predict the importance weight results of each new feature cluster through the model output.

[0066] In summary, the image enhancement processing method based on deep learning of the present invention extracts multi-scale features, enabling the detailed information in the image to be fully retained and enhanced, and recalibrates and fuses the multi-scale features to obtain a comprehensive feature map. Recalibration is to ensure that important features are strengthened while secondary features are appropriately suppressed, thereby improving the pertinence and effectiveness of image enhancement. Then, perform a non-linear transformation on the comprehensive feature map to enhance the contrast and color saturation of the image. Then use the decoder network to reconstruct the transformed features into a preliminary enhanced image to achieve an automatic and accurate mapping from the feature space to the image space. The decoder network can optimize the efficiency and quality of image reconstruction, enabling the preliminary enhanced image to have a better visual effect while maintaining high fidelity. To further enhance the overall effect of the image, use a generative adversarial network to generate a content-aware weight map. The content-aware weight map can accurately reflect the importance of each pixel in the image in terms of content, providing guidance for further image enhancement and making the image enhancement process more intelligent and adaptive. Then divide the key regions and associated regions according to the values of the content-aware weight map, and achieve global optimization through the interaction and cooperation between regions, so that the key regions are enhanced with emphasis and the associated regions are appropriately adjusted, making each part of the image more harmonious and unified, thereby enhancing the overall effect and beauty of the image and obtaining a higher-quality image.

[0067] The additional aspects and advantages of the present invention will be partly given in the following description, partly will become obvious from the following description, or can be understood through the embodiments of the present invention. Brief Description of the Drawings

[0068] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:

[0069] Figure 1Flowchart of the image enhancement processing method based on deep learning according to Embodiment 1 of the present invention;

[0070] Figure 2 System block diagram of the image enhancement processing system based on deep learning according to Embodiment 2 of the present invention. Detailed implementation manners

[0071] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided so that the disclosure of the present invention is thorough and comprehensive.

[0072] It should be noted that when an element is referred to as being "fixedly provided on" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are only for the purpose of illustration.

[0073] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0074] Embodiment 1

[0075] Please refer to Figure 1 , the present invention proposes an image enhancement processing method based on deep learning, and this method includes steps S101 to S106:

[0076] S101, extract multi-scale features of the input image.

[0077] It should be noted that extracting the multi-scale features of the input image is to capture different details and structures in the image, including edges, textures, shapes, etc., and provide rich feature information for subsequent image enhancement. Before extracting features, the input image can be preprocessed, such as normalization and denoising processing, etc.

[0078] A convolutional neural network (CNN) can be used to extract multi-scale features of the input image. When constructing a CNN model, CNN architectures such as ResNet and VGG can be selected and parameters adjusted. For example, to extract multi-scale features, different-sized convolutional kernels are needed to capture different-scale information in the image. Then the preprocessed image is input into the CNN model, and the feature maps of each layer are calculated through forward propagation. When extracting multi-scale features, pay attention to the feature maps output by different convolutional layers. The feature maps output by shallow convolutional layers usually contain more detailed information (such as edges, textures, etc.), while the feature maps output by deep convolutional layers focus more on the overall structure and semantic information of the image. After feature extraction, multi-scale features are output, and these features will be used for image enhancement processing after re-calibration and fusion.

[0079] S102, Re-calibrate and fuse the multi-scale features to obtain a comprehensive feature map.

[0080] It should be noted that before fusing the multi-scale features, it is necessary to re-calibrate the features of each scale to adjust the importance of different-scale features, highlight the important features, suppress the secondary features, and fuse the re-calibrated different-scale features to obtain a comprehensive feature map. This feature map integrates information from multiple scales and has stronger expressive power and robustness.

[0081] S103, Perform a non-linear transformation on the comprehensive feature map.

[0082] It should be noted that a non-linear transformation is performed on the fused comprehensive feature map to enhance the expressive power and flexibility of the features, making subsequent image reconstruction and enhancement more accurate and diverse.

[0083] It can be achieved through activation functions (such as ReLU, Sigmoid, etc.) to perform non-linear mapping on the features. In the image enhancement task of this embodiment, the ReLU function, which is simple and effective and has easy-to-calculate gradients, can be used. The comprehensive feature map obtained in step S102 is used as the input of the non-linear transformation function, and the non-linear transformation function is applied to each element in the comprehensive feature map to obtain the transformed features. During the training process, the effect of the non-linear transformation can be evaluated according to the change of the loss function. If the loss function value continues to decrease, it indicates that the non-linear transformation helps to improve the model performance; otherwise, the non-linear transformation function or parameters need to be adjusted.

[0084] S104, Use the decoder network to reconstruct the transformed features into a preliminary enhanced image.

[0085] It should be noted that the decoder network is used to map the information in the feature space back to the image space to obtain a preliminary enhanced image.

[0086] Construct a decoder network. Since the decoder network is used to map features back to the image space, a structure symmetric to the encoder network (e.g., the CNN in step S101) can be selected. The decoder network can be composed of multiple upsampling layers and convolutional layers, which are used to gradually restore the spatial resolution and detailed information of the image. Then, initialize the parameters of the decoder network (such as weights and biases), and then input the features after non-linear transformation into the decoder network, ensuring that the format of the input features meets the input requirements of the decoder network. Through the forward propagation of the decoder network, calculate the feature maps layer by layer to gradually restore the spatial resolution and detailed information of the image. After the layer-by-layer calculation of the decoder network, finally output a preliminary enhanced image, ensuring that the shape and type of the output image meet the requirements. For example, the pixel values of the image can be normalized to the range of 0 to 255 and converted to the RGB format.

[0087] S105. Design a generative adversarial network that takes the preliminary enhanced image as input and generates a content-aware weight map, where each pixel value in the content-aware weight map reflects the importance of the corresponding position in terms of content.

[0088] It should be noted that a generative adversarial network is designed to generate a content-aware weight map based on the content of the preliminary enhanced image, reflecting the importance of each pixel in terms of content.

[0089] The generative adversarial network (GAN) consists of a generator and a discriminator. The generator is responsible for generating the content-aware weight map, and the discriminator is responsible for judging whether the generated weight map is real. Since the GAN continuously optimizes the generator and the discriminator during the training process to generate more realistic images, its ability to understand the content of the images is also improved during this process. Compared with traditional image processing models, such as convolutional neural networks and autoencoders, it shows higher accuracy in content awareness.

[0090] Based on the content-aware weight map generated by the GAN, more accurate enhancement processing can be performed on the image. For example, in image super-resolution reconstruction, the key regions of the image can be enhanced targeted according to the weight map, thereby improving the quality of the reconstructed image. Traditional image enhancement methods often enhance based on global statistical features or simple local features, making it difficult to accurately capture the key information and details in the image.

[0091] The generated content-aware weight map can be applied to a variety of image processing tasks, such as image inpainting, image style transfer, image segmentation, etc., and has broad application prospects in the field of image processing.

[0092] The recalibration and fusion in step S102 act on the feature level to optimize feature representation, which can improve the accurate expression ability and robustness of features, thereby indirectly improving the image enhancement effect. Then, by generating a content-aware weight map and acting on the image pixel level to guide the further enhancement of the image, the overall quality and visual effect of the image can be improved, making the processed image clearer, more natural, and with better visual effects.

[0093] Further optionally, the step of designing a generative adversarial network to generate a content-aware weight map with the preliminarily enhanced image as the input includes:

[0094] Determine the structure of the generative adversarial network, including a generator and a discriminator;

[0095] Construct a training set based on a large amount of historical data. Each sample data in the training set includes a preliminarily enhanced image and the corresponding real content-aware weight map. Adaptively adjust the parameters of the generator to generate an image that better meets the expectations of the discriminator. The content-aware weight map is marked with labels of content importance;

[0096] Train the generator based on the generator loss;

[0097] Train the discriminator based on the discriminator loss;

[0098] Through the adversarial training of the generator and the discriminator, adaptively adjust the parameters of the generator to generate an image that better meets the expectations of the discriminator;

[0099] After training is completed, input the preliminarily enhanced image into the trained generator and generate a content-aware weight map with the same size as the preliminarily enhanced image.

[0100] It can be understood that through the adversarial training of the generator and the discriminator, GAN can learn the deep features of the image content. This deep learning ability enables GAN to more accurately capture the key information and details in the image when generating the content-aware weight map, thus generating a more accurate content-aware weight map.

[0101] The adversarial training mechanism of GAN enables it to adaptively adjust the parameters of the generator to generate an image that better meets the expectations of the discriminator. This adaptability enables GAN to more flexibly adapt to different types of image content and styles when generating the content-aware weight map. By training a large amount of data, GAN can learn the general laws of image content and can better generalize to unseen image data when generating the content-aware weight map.

[0102] Further optionally, the step of training the generator based on the generator loss includes:

[0103] Randomly extract a preliminarily enhanced image from the training set;

[0104] Use a generator to generate a content-aware weight map;

[0105] Input the generated content-aware weight map into the discriminator and calculate the generator loss. The calculation formula is:

[0106] ,

[0107] where, is the generator loss, is the content-aware weight map generated by the generator, z is the input of the generator, which is random noise, is the output of the discriminator for indicating the probability that the discriminator believes is a real content-aware weight map;

[0108] Use the backpropagation algorithm to update the network parameters of the generator.

[0109] It can be understood that through the adversarial training process, the generator can generate more and more realistic content-aware weight maps that are difficult to distinguish by the discriminator. Among them, the generator and the discriminator compete with each other and co-evolve, and finally the generator can generate high-quality forged images.

[0110] The formula of the generator loss indicates that the generator hopes that the discriminator will misidentify the generated weight map G(z) as real, that is, D(G(z)) is as close to 1 as possible. Therefore, the generator loss is the negative logarithm of D(G(z)), and when D(G(z)) is close to 1, the loss is close to 0.

[0111] Further optionally, the steps of training the generator based on the discriminator loss include:

[0112] Initialize the network parameters of the generator and the discriminator;

[0113] Randomly extract preliminary enhanced images and labeled real content-aware weight maps from the training set;

[0114] Use the generator to generate forged content-aware weight maps;

[0115] Input the real and forged content-aware weight maps into the discriminator and calculate the discriminator loss. The calculation formula is:

[0116] ,

[0117] where, is the discriminator loss, p is the real content-aware weight map, and They are the outputs of the discriminator for p and G(z) respectively, representing the probabilities that the discriminator believes p and G(z) are real content-aware weight maps;

[0118] Update the network parameters of the discriminator through the backpropagation algorithm.

[0119] It is understandable that by training the discriminator to distinguish between real and forged content-aware weight maps, and at the same time using this adversarial training process to indirectly train the generator so that it can generate more realistic and difficult-to-distinguish content-aware weight maps by the discriminator.

[0120] The formula for the discriminator loss indicates that the discriminator hopes that the probability D(p) of judging the real weight map p as true is as close to 1 as possible, and at the same time the probability of judging the fake weight map G(z) as false is also as close to 1 as possible. Therefore, the discriminator loss is the sum of the negative logarithms of D(p) and ...

[0121] S106. According to the content-aware weight map, identify the key regions and associated regions, and further enhance the overall effect of the image through the interaction and cooperation between the regions.

[0122] It should be noted that the content-aware weight map is used to identify the key regions and associated regions in the image, and the overall effect and aesthetic feeling of the image are enhanced through the interaction and cooperation between the regions.

[0123] The image is divided into key regions and associated regions according to the values of the weight map. And according to the gap between the two regions, adaptive enhancement processing can be carried out, such as focusing on enhancing the key regions, such as increasing brightness and color saturation, and making appropriate adjustments to the associated regions to maintain the overall harmony and unity of the image. The interaction and cooperation between the regions can also be achieved through methods such as fusion, transition or collaborative enhancement between the regions.

[0124] Further optionally, the step of identifying the key regions and associated regions according to the content-aware weight map and enhancing the overall effect of the image through the interaction and cooperation between the regions includes:

[0125] Define the key regions and associated regions according to the weight values in the content-aware weight map;

[0126] Perform binarization processing on the content-aware weight map to obtain a binary image containing only the key regions and associated regions;

[0127] Segment the binary image into different regions and assign unique region labels;

[0128] Identify the key regions and associated regions from the content-aware weight map according to the region labels;

[0129] Calculate the similarity between two key regions with a connection relationship and the associated region;

[0130] If the similarity is higher than the first threshold, strengthen the visual connection between the two regions;

[0131] If the similarity is lower than the second threshold, weaken or eliminate the visual interference between the two regions.

[0132] It is understandable that the content-aware weight map provides information about the importance or relevance of each region in the image. By analyzing the weight values, key regions and associated regions can be identified. Key regions are usually important or prominent parts of the image, while associated regions are regions that have a connection relationship with the key regions and may have an impact on them.

[0133] Convert the weight map into a binary image that only contains key regions and associated regions through binarization. In the binary image, key regions and associated regions are clearly distinguished, facilitating subsequent segmentation and recognition. Then use an image segmentation algorithm to segment the binary image into different regions and assign unique region labels, providing a basis for subsequent region recognition and similarity calculation. The following method can be used for region labeling:

[0134] Use a connected component labeling algorithm (such as the scan line algorithm, etc.) to label the binary image. Among them, each connected component is assigned a unique region label. During or after the labeling process, traverse the image and record each region and its adjacent regions. Specifically, an adjacency matrix can be used to represent the connection relationship between regions. Among them, the adjacency matrix is a two-dimensional array, where the rows and columns represent different regions respectively, and the elements in the matrix indicate whether two regions are adjacent. When calculating the similarity, for each key region, traverse its adjacency matrix to find an associated region adjacent to it, and calculate the similarity between the two regions.

[0135] After region labeling, when calculating the similarity, for each key region, traverse its adjacency matrix to find an associated region adjacent to it. Then, feature descriptors (such as color histograms, texture features, shape features, etc.) can be extracted and converted into feature vectors to quantify the similarity between regions. Based on the feature vectors, calculate the similarity between the two regions through similarity calculation methods (such as Euclidean distance, cosine similarity, etc.). When the similarity is higher than the first threshold, it means that the two regions are highly correlated in content. Therefore, strengthening the connection between them can enhance the overall coordination and unity of the image. If the similarity is lower than the second threshold, it means that the two regions are quite different in content. Therefore, weakening or eliminating the interference between them can make the image more harmonious.

[0136] The key area is the most important part of the image and has a direct impact on the overall effect of the image. The importance of the key area can be highlighted by strengthening the visual connection between the key area and similar related areas. If the related area has a high correlation with the key area, it can assist or set off the key area. The color, texture and other features of the related area can be adjusted to make it more coordinated with the key area, thereby enhancing the overall effect of the image. For related areas with low similarity to the key area, their visual interference can be reduced or eliminated through smoothing.

[0137] This embodiment uses a content-aware weight map to identify and process key areas and related areas in an image, and enhances the overall effect of the image by strengthening the visual connection between similar areas and weakening the visual interference between dissimilar areas. This image enhancement processing method takes into account both the importance of the area and the correlation between the areas.

[0138] Further optionally, the step of defining the key area and the associated area according to the weight values ​​in the content-aware weight map includes:

[0139] Determine the area with a weight value higher than the third threshold as a key area;

[0140] The area whose weight value is lower than the third threshold and connected to the key area is defined as the associated area.

[0141] Further optionally, the step of recalibrating and fusing the multi-scale features to obtain a comprehensive feature map includes:

[0142] Input the multi-scale features into the deep clustering network, divide the scale features into corresponding feature clusters, and recalibrate the scale features according to the importance weight of the corresponding feature cluster;

[0143] The recalibrated features of each scale are fused to obtain a comprehensive feature map.

[0144] Understandably, each scale feature is divided into corresponding feature clusters through the deep clustering network, and each scale feature is recalibrated according to the importance weight of the corresponding feature cluster to emphasize important features and suppress unimportant features. The importance weight can be used to adjust the weight of each scale feature to improve the feature representation before feature fusion.

[0145] In this embodiment, the deep clustering network is used to divide the features of each scale into corresponding feature clusters, and each feature cluster represents a specific pattern, structure or semantic information in the image. According to the importance weights of the corresponding feature clusters, the features of each scale are recalibrated. The importance weights of the feature clusters can reflect the contribution degree or significance of the feature clusters in specific tasks (such as classification, detection, segmentation, etc.). The recalibration process may involve scaling, weighting or selecting the features, so as to enhance the influence of important features and weaken the influence of unimportant features. The recalibrated features of each scale are fused, and the obtained comprehensive feature map not only fuses the information of multi-scale features, but also considers the importance of the feature clusters (that is, the contribution degree or significance of different feature clusters in the current task), so it has stronger expressive ability and discriminability.

[0146] Further optionally, the step of inputting the multi-scale features into the deep clustering network, dividing the features of each scale into corresponding feature clusters, and recalibrating the features of each scale according to the importance weights of the corresponding feature clusters includes:

[0147] Input the extracted multi-scale features into the encoder of the deep clustering network to obtain the low-dimensional representations of the features of each scale;

[0148] In the low-dimensional representation space, use the clustering algorithm to divide the features of each scale into corresponding feature clusters;

[0149] Predict the importance weights of each feature cluster through the MLP model;

[0150] Recalibrate the features of each scale according to the importance weights corresponding to the feature clusters divided by the features of each scale to obtain the recalibrated features of each scale. The formula is:

[0151] ,

[0152] where, is the recalibrated feature of the i-th scale, is the feature of the scale The importance weight corresponding to the divided feature cluster.

[0153] Further optionally, the step of fusing the recalibrated features of each scale to obtain the comprehensive features includes:

[0154] Fuse the recalibrated features of each scale to obtain the fused feature X. The formula is:

[0155] ,

[0156] where, , , … are the recalibrated features of each scale respectively, and M is the number of scale features.

[0157] It is understandable that by assigning importance weights through a deep clustering network and recalibrating multi-scale features, the information of multi-scale features can be fully utilized and the importance of feature clusters can be considered, thereby obtaining a more expressive and discriminative comprehensive feature map.

[0158] The goal of the clustering algorithm is to make the features within the same feature cluster as similar as possible and the features between different feature clusters as different as possible. K-means, spectral clustering, etc. can be used.

[0159] Further optionally, predicting the importance weight of each feature cluster through the MLP model includes:

[0160] Obtain the representative feature Y of each feature cluster from the deep clustering network and construct an input matrix , where , n is the number of feature clusters, d is the dimension of the representative feature of each feature cluster, and the representative feature can be the centroid, mean of the feature cluster, or the feature representation after a certain transformation;

[0161] Label each representative feature in the input matrix according to the ground truth label data related to the importance of the feature cluster;

[0162] Perform forward propagation on the input features through the MLP model to obtain the output prediction value;

[0163] Use the loss function to calculate the error between the prediction value and the ground truth label;

[0164] Calculate the gradient of the loss function with respect to the model parameters through the chain rule and use the gradient descent method to update the model parameters;

[0165] For a new feature cluster, extract its representative feature and input it into the trained MLP model to predict the importance weight of each new feature cluster through the model output.

[0166] It is understandable that predicting the importance weight of each feature cluster through the MLP (Multi-Layer Perceptron) model realizes the automation, efficiency, and accuracy of importance weight determination. And through the training of a large amount of data, the MLP model can learn the potential laws and patterns in the data. These laws and patterns not only apply to the training data but also can be generalized to new feature clusters, thereby accurately predicting the importance weight of new data feature clusters.

[0167] In summary, the image enhancement processing method based on deep learning of the present invention extracts multi-scale features, enabling the full retention and enhancement of detailed information in the image, and recalibrates and fuses the multi-scale features to obtain a comprehensive feature map. Recalibration is to ensure that important features are strengthened while secondary features are appropriately suppressed, thereby improving the pertinence and effectiveness of image enhancement. Then, a non-linear transformation is performed on the comprehensive feature map to enhance the contrast and color saturation of the image. Next, a decoder network is used to reconstruct the transformed features into a preliminary enhanced image to achieve an automatic and accurate mapping from the feature space to the image space. The decoder network can optimize the efficiency and quality of image reconstruction, enabling the preliminary enhanced image to have better visual effects while maintaining high fidelity. To further enhance the overall effect of the image, a generative adversarial network is utilized to generate a content-aware weight map. The content-aware weight map can accurately reflect the importance of each pixel in the image in terms of content, providing guidance for further image enhancement and making the image enhancement process more intelligent and adaptive. Then, according to the values of the content-aware weight map, key regions and associated regions are divided, and global optimization is achieved through the interaction and cooperation between regions, enabling key regions to be enhanced with emphasis and associated regions to be appropriately adjusted, making each part of the image more harmonious and unified, thereby enhancing the overall effect and aesthetic feeling of the image and obtaining a higher-quality image.

[0168] Embodiment 2

[0169] Please refer to Figure 2 , the image enhancement processing system based on deep learning proposed by the present invention, which includes:

[0170] Feature extraction module: used to extract multi-scale features of the input image;

[0171] Feature fusion module: used to recalibrate and fuse multi-scale features to obtain a comprehensive feature map;

[0172] Feature transformation module: used to perform non-linear transformation on the comprehensive feature map;

[0173] Preliminary enhancement module: used to reconstruct the transformed features into a preliminary enhanced image using a decoder network;

[0174] Weight map module: used to design a generative adversarial network, taking the preliminary enhanced image as the input to generate a content-aware weight map, where each pixel value in the content-aware weight map reflects the importance of the corresponding position in terms of content;

[0175] Cooperative enhancement module: used to identify key regions and associated regions according to the content-aware weight map, and further enhance the overall effect of the image through the interaction and cooperation between regions.

[0176] Further optionally, the weight map module is further used for:

[0177] Determine the structure of the generative adversarial network, including the generator and the discriminator;

[0178] Construct a training set based on a large amount of historical data. Each sample data in the training set contains a preliminary enhanced image and the corresponding true content-aware weight map. Adaptively adjust the parameters of the generator to generate an image that better meets the expectations of the discriminator. The content-aware weight map is marked with labels of content importance;

[0179] Train the generator based on the generator loss;

[0180] Train the discriminator based on the discriminator loss;

[0181] Through the adversarial training of the generator and the discriminator, adaptively adjust the parameters of the generator to generate an image that better meets the expectations of the discriminator;

[0182] After the training is completed, input the preliminary enhanced image into the trained generator and generate a content-aware weight map with the same size as the preliminary enhanced image.

[0183] Further optionally, the weight map module is further configured to:

[0184] Randomly extract a preliminary enhanced image from the training set;

[0185] Use the generator to generate a content-aware weight map;

[0186] Input the generated content-aware weight map into the discriminator and calculate the generator loss. The calculation formula is:

[0187] ,

[0188] where, is the generator loss, is the content-aware weight map generated by the generator, z is the input of the generator, is random noise, is the output of the discriminator for indicating the probability that the discriminator believes is a true content-aware weight map;

[0189] Use the backpropagation algorithm to update the network parameters of the generator.

[0190] Further optionally, the weight map module is further configured to:

[0191] Initialize the network parameters of the generator and the discriminator;

[0192] Randomly extract a preliminary enhanced image and a marked true content-aware weight map from the training set;

[0193] Use a generator to generate a forged content-aware weight map;

[0194] Input the real and forged content-aware weight maps into the discriminator, and calculate the discriminator loss. The calculation formula is:

[0195] ,

[0196] where, is the discriminator loss, p is the real content-aware weight map, and are the outputs of the discriminator for p and G(z) respectively, representing the probabilities that the discriminator thinks p and G(z) are real content-aware weight maps;

[0197] Update the network parameters of the discriminator through the backpropagation algorithm.

[0198] Further optionally, the collaborative enhancement module is also used for:

[0199] Define the key regions and associated regions according to the weight values in the content-aware weight map;

[0200] Perform binarization processing on the content-aware weight map to obtain a binary image containing only the key regions and associated regions;

[0201] Segment the binary image into different regions and assign unique region labels;

[0202] Identify the key regions and associated regions from the content-aware weight map according to the region labels;

[0203] Calculate the similarity between two key regions and the associated region with a connection relationship;

[0204] If the similarity is higher than the first threshold, strengthen the visual connection between the two regions;

[0205] If the similarity is lower than the second threshold, weaken or eliminate the visual interference between the two regions.

[0206] Further optionally, the collaborative enhancement module is also used for:

[0207] Define the regions with weight values higher than the third threshold as key regions;

[0208] Define the regions with weight values lower than the third threshold and connected to the key regions as associated regions.

[0209] Further optionally, the feature fusion module is also used for:

[0210] Input the multi-scale features into the deep clustering network, divide each scale feature into the corresponding feature clusters, and recalibrate each scale feature according to the importance weights of the corresponding feature clusters;

[0211] Fuse the re-calibrated multi-scale features to obtain a comprehensive feature map.

[0212] Further optionally, the feature fusion module is further configured to:

[0213] Input the extracted multi-scale features into the encoder of the deep clustering network to obtain the low-dimensional representations of the multi-scale features;

[0214] In the low-dimensional representation space, use a clustering algorithm to divide the multi-scale features into corresponding feature clusters respectively;

[0215] Predict the importance weights of each feature cluster through an MLP model;

[0216] Re-calibrate the multi-scale features according to the importance weights corresponding to the feature clusters divided by the multi-scale features, and obtain the re-calibrated multi-scale features. The formula is:

[0217] ,

[0218] where, is the re-calibrated i-th scale feature, is the importance weight corresponding to the feature cluster divided by the scale feature ;

[0219] Further optionally, the feature fusion module is further configured to:

[0220] Fuse the re-calibrated multi-scale features to obtain a fused feature X. The formula is:

[0221] ,

[0222] where, , , … , are the re-calibrated multi-scale features respectively, and M is the number of scale features.

[0223] Further optionally, the feature fusion module is further configured to:

[0224] Obtain the representative feature Y of each feature cluster from the deep clustering network, and construct an input matrix , where, , n is the number of feature clusters, and d is the dimension of the representative feature of each feature cluster;

[0225] Label each representative feature in the input matrix according to the ground truth label data related to the importance of the feature cluster;

[0226] Perform forward propagation on the input features through an MLP model to obtain an output prediction value;

[0227] Calculate the error between the predicted value and the true label using a loss function;

[0228] Calculate the gradient of the loss function with respect to the model parameters using the chain rule, and update the model parameters using gradient descent;

[0229] For the new feature clusters, extract their representative features and input them into the trained MLP model to predict the importance weight results of each new feature cluster through the model output.

[0230] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.

Claims

1. An image enhancement processing method based on deep learning, characterized in that: The method comprises: Extract multi-scale features of the input image; Recalibrate and fuse multi-scale features to obtain a comprehensive feature map; Perform nonlinear transformation on the comprehensive feature map; Use the decoder network to reconstruct the transformed features into a preliminary enhanced image; Design a generative adversarial network that takes the preliminary enhanced image as input and generates a content-aware weight map, where each pixel value in the content-aware weight map reflects the importance of the corresponding position in the content; According to the content-aware weight map, key areas and related areas are identified, and the overall effect of the image is further enhanced through interaction and collaboration between regions.

2. The image enhancement processing method based on deep learning according to claim 1, characterized in that: The steps of designing a generative adversarial network to generate a content-aware weight map using a preliminary enhanced image as input include: Determine the structure of the generative adversarial network, including the generator and the discriminator; A training set is built based on a large amount of historical data. Each sample data in the training set contains a preliminary enhanced image and a corresponding real content-aware weight map to adaptively adjust the parameters of the generator to generate an image that better meets the expectations of the discriminator. The content-aware weight map is marked with labels of content importance. Train the generator based on the generator loss; Train the discriminator based on the discriminator loss; Through adversarial training between the generator and the discriminator, the parameters of the generator are adaptively adjusted to generate images that better meet the discriminator's expectations; After the training is completed, the preliminary enhanced image is input into the trained generator, and a content-aware weight map with the same size as the preliminary enhanced image is generated.

3. The image enhancement processing method based on deep learning according to claim 2 is characterized in that: The step of training the generator based on the generator loss includes: Randomly extract preliminary enhanced images from the training set; Generate a content-aware weight map using a generator; The generated content-aware weight map is input into the discriminator and the generator loss is calculated as: in, is the generator loss, The content-aware weight map generated for the generator, is the input of the generator, is random noise, For the discriminator The output of is the probability of a true content-aware weight map; The back-propagation algorithm is used to update the generator network parameters.

4. The image enhancement processing method based on deep learning according to claim 2, characterized in that: The step of training the generator based on the discriminator loss includes: Initialize the network parameters of the generator and discriminator; Randomly extract preliminary enhanced images and labeled true content-aware weight maps from the training set; Generate fake content-aware weight maps using a generator; The real and fake content-aware weight maps are fed into the discriminator and the discriminator loss is calculated as: in, is the discriminator loss, is the true content-aware weight map, and are the discriminator pairs and The output of and is the probability of a true content-aware weight map; The network parameters of the discriminator are updated through the back-propagation algorithm.

5. The image enhancement processing method based on deep learning according to claim 1, characterized in that: The step of identifying key areas and associated areas according to the content-aware weight map and enhancing the overall effect of the image through interaction and collaboration between the areas comprises: defining key areas and associated areas according to weight values ​​in a content-aware weight map; Binarize the content-aware weight map to obtain a binary image containing only the key area and the associated area; Segment the binary image into different regions and assign unique region labels; Identify key regions and associated regions from the content-aware weight map based on region labels; Calculate the similarity between two key regions with a connection relationship and the associated region; If the similarity is higher than a first threshold, the visual connection between the two regions is strengthened; If the similarity is lower than the second threshold, the visual interference between the two areas is reduced or eliminated.

6. The image enhancement processing method based on deep learning according to claim 5, characterized in that: The step of defining the key area and the associated area according to the weight values ​​in the content-aware weight map comprises: Determine the area with a weight value higher than the third threshold as a key area; The area whose weight value is lower than the third threshold and connected to the key area is defined as the associated area.

7. The image enhancement processing method based on deep learning according to claim 1, characterized in that: The step of recalibrating and fusing the multi-scale features to obtain a comprehensive feature map includes: Input the multi-scale features into the deep clustering network, divide the scale features into corresponding feature clusters, and recalibrate the scale features according to the importance weight of the corresponding feature cluster; The recalibrated features of each scale are fused to obtain a comprehensive feature map.

8. The image enhancement processing method based on deep learning according to claim 7, characterized in that: The multi-scale features are input into the deep clustering network, each scale feature is divided into a corresponding feature cluster, and each scale feature is recalibrated according to the importance weight of the corresponding feature cluster, including: The extracted multi-scale features are input into the encoder of the deep clustering network to obtain the low-dimensional representation of each scale feature; In the low-dimensional representation space, clustering algorithms are used to divide the scale features into corresponding feature clusters; Predict the importance weight of each feature cluster through the MLP model; The scale features are recalibrated according to the importance weights corresponding to the feature clusters divided by the scale features to obtain the recalibrated scale features. The formula is: in, After recalibration scale characteristics, It is a scale feature Importance weights corresponding to the divided feature clusters; The step of fusing the recalibrated scale features to obtain comprehensive features includes: The recalibrated scale features are fused to obtain the fused features. , the formula is: in, , are the scale features after recalibration, is the number of scale features.

9. The image enhancement processing method based on deep learning according to claim 8, characterized in that: The MLP model is used to predict the importance weight of each feature cluster, including: Get the representative feature Y of each feature cluster from the deep clustering network and construct the input matrix ,in, , is the number of feature clusters, is the dimension of the representative feature of each feature cluster; Labeling each representative feature in the input matrix according to true label data related to the importance of feature clusters; The input features are forward propagated through the MLP model to obtain the output prediction value; Use the loss function to calculate the error between the predicted value and the true label; The gradient of the loss function with respect to the model parameters is calculated by the chain rule, and the model parameters are updated using the gradient descent method; For new feature clusters, their representative features are extracted and input into the trained MLP model to output the importance weight prediction results of each new feature cluster through the model.

Citation Information

Patent Citations

  • Underwater image enhancement method based on multi-scale attention mechanism fusion

    CN115034982A

  • Pelvic image restoration method based on occlusion perception and multi-scale feature fusion

    CN118429225A

  • Unsupervised domain adaptive medical image segmentation method and device based on multi-scale features

    CN119649038A

  • Image super-resolution reconstruction method based on deep learning

    CN119809933A

  • Medical image segmentation method based on u-net

    US20220309674A1