Artistic photo image enhancement processing method

By introducing a hybrid domain attention HDAB module and an adaptive contrast enhancement ACEM module in art photo processing, combined with deep learning models and feature extraction technology, the problem that the existing technology is difficult to improve visual effects and retain artistic style when processing artistic photos, and achieve efficient image enhancement and detail retention effects.

CN120182115AActive Publication Date: 2025-06-20SHANDONG HAOXIN INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510263003.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-20
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

When processing artistic photos, it is difficult for the prior art to effectively improve the visual effect of the image, while retaining the unique style and details of the art work, especially in the processing of details and local features.

Method used

By introducing a hybrid domain attention HDAB module and an adaptive contrast enhancement ACEM module, combining deep learning models and feature extraction technology, the frequency and spatial domain features of art photos are extracted, weighted fusion and enhancement are carried out, and the parameters of the model are optimized through multi-objective optimization strategies of content loss, style loss and high-frequency loss.

Benefits of technology

It significantly improves the visual expression and detail clarity of artistic photos, ensures the structural consistency, style matching and detailed expression of the image, and effectively preserves the original style and details of the artistic photos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182115A_ABST
    Figure CN120182115A_ABST
Patent Text Reader

Abstract

The invention provides an art photo image enhancement processing method, and belongs to the field of image enhancement. According to the method, a dynamically generated multi-scale fusion coefficient and a cross-domain interaction mechanism are utilized, a hybrid domain attention (HDAB) module is constructed to extract frequency domain and spatial domain feature information of an art photo image, and fusion and enhancement of high-frequency, low-frequency and multi-scale features are realized; through a normalization and contrast stretching method, in combination with a region saliency enhancement strategy, an adaptive contrast enhancement (ACEM) module is constructed, contrast enhancement is performed on artistic photo image features, weight information is dynamically generated, target region features are highlighted, and the method improves the visual effect of the artistic photo and enhances the artistic expressive force and detail definition of the artistic photo.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of image enhancement, and in particular relates to an artistic photo image enhancement processing method. Background Art

[0002] With the continuous development of digital imaging technology, image enhancement, as an important image processing technology, has been widely used in photography, medical imaging, satellite remote sensing, art restoration and other fields. Especially in the processing of artistic photos, how to effectively enhance the visual effect of the image while retaining the unique style and details of the artwork has become a major challenge in the current field of image enhancement.

[0003] Traditional image enhancement methods, such as histogram equalization, contrast adjustment and image filtering, can improve the brightness and contrast of images, but they often have some obvious shortcomings when processing artistic photos. First, traditional methods are prone to image color distortion. Especially in artistic works, the richness and gradient of colors are very important. Any excessive enhancement or adjustment of colors may destroy the style of the original work. Second, traditional technologies usually perform image enhancement in a global manner and cannot effectively process details and local features, resulting in deficiencies in the expression of image details and textures. Especially for complex artistic photos, the loss of details and local contrast is particularly serious.

[0004] In recent years, with the development of deep learning and convolutional neural networks (CNN), image enhancement technology has made significant progress; deep learning-based methods have achieved effective improvement in image quality by automatically learning image features, especially in detail recovery, noise removal and contrast enhancement. However, existing deep learning methods still have some problems when processing artistic photos. Since artistic photos often contain rich details, special textures and unique color combinations, existing deep learning methods sometimes over-smooth images or do not fully retain the original artistic style, affecting the expressiveness and artistic sense of the image.

[0005] Therefore, how to improve the visual effects of artistic photos while retaining their original artistic style and details during the image enhancement process is a difficult problem that needs to be solved urgently in the field of image processing. Current technology has not been able to effectively combine the special requirements of artistic photos, both to enhance image quality and to ensure that the artistic style and details are not distorted. Therefore, proposing an image processing method that can enhance the characteristics of artistic photos has become an important research task in image enhancement technology. Summary of the invention

[0006] The present invention provides an artistic photo image enhancement processing method, which aims to improve the quality of artistic photo images and enhance the details and clarity of the images through a deep learning model and feature extraction enhancement technology.

[0007] The present invention aims to propose an art photo image enhancement processing method, which specifically includes the following steps.

[0008] S1. Collect art photos. After preprocessing and annotating features, divide them into a training set, a validation set, and a test set according to the ratios of 70%, 20%, and 10% for training the art photo image enhancement model.

[0009] S2. Construct a Hybrid Domain Attention (HDAB) module through dynamically generated multi-scale fusion coefficients and the proposed cross-domain interaction mechanism. Use the HDAB module to extract the frequency-domain and spatial-domain information of the art photo image features, perform weighted fusion on high-frequency, low-frequency, and multi-scale features, and enhance the frequency-domain and spatial-domain features through the cross-domain interaction mechanism.

[0010] S3. Construct an Adaptive Contrast Enhancement Module (ACEM) through the normalization and contrast stretching method and the proposed region saliency enhancement strategy. Use the ACEM module to enhance the art photo image features, generate attention weights according to the region saliency enhancement strategy, highlight the key information regions, and suppress redundant information.

[0011] S4. Design content loss, style loss, and high-frequency loss, construct a loss function and a training strategy module, and optimize the parameters of the model by adjusting the weights and learning rate of the loss function to complete the training process of the art photo image enhancement model.

[0012] S5. Construct an art photo image enhancement model, including input, image embedding, Hybrid Domain Attention (HDAB) module, Adaptive Contrast Enhancement Module (ACEM), fully connected layer, and output. Use the loss function and training strategy to guide the training of the art photo image enhancement model.

[0013] S6. Obtain the art photo image that needs to be enhanced, input it into the art photo image enhancement model for processing, and output a high-resolution art photo image.

[0014] Preferably, in step S2, constructing the Hybrid Domain Attention (HDAB) module specifically includes the following steps: Step S21. First, extract the frequency-domain features from the input art photo image features X, where X ∈ R H×W×C , and H, W, and C respectively represent the height, width, and number of channels of X; use the Fourier transform FFT to extract the frequency-domain information of the art photo image features: F freq (X) = FFT(X), Design dynamic filters H H 、H L Separate the high-frequency and low-frequency information of F freq (X), and the filters H H 、HL Dynamically generated from global information, and the generation formula is: H H ,H L = Sigmoid(MLP(GAP(X))), where GAP is global average pooling and Sigmod is the activation function, The high-frequency information is represented as: The low-frequency information is represented as: where IFFT is the inverse operation of the Fourier transform. The high-frequency features retain the texture details and edge information in the artistic photo, and the low-frequency features highlight the global consistency, corresponding to the overall style and color distribution in the artistic photo. Then, the separated high-frequency and low-frequency features are non-linearly transformed using 1×1 convolution. The specific formula is: Conv 1×1 is 1×1 convolution.

[0015] Step S22: Extract spatial domain features from the artistic photo image feature X, and use convolution kernels of different sizes to extract multi-scale local features. where Conv k represents a convolution kernel of size k×k, and k ∈ {3, 5, 7}.

[0016] Step S23: Introduce the dynamically generated multi-scale fusion coefficient w k , and perform weighted summation on the multi-scale features to generate the spatial domain fused artistic photo image feature: are the multi-scale local features extracted using convolution kernels of different sizes, and the multi-scale fusion coefficient w k , and the calculation formula is: where is the feature mean of the feature map , indicating the global mean calculation of the multi-scale convolution features , is the feature standard deviation of the feature map , indicating the calculation of the local variation intensity of the multi-scale convolution features , μ kis the dynamic offset, representing the prior preference of a specific scale k, which is extracted from multi-scale convolution features through a non-linear transformation. The formula is: Extracted from Linear k (·) is a linear transformation operation, C is the number of channels, ∈ is a small value, and ∈ prevents the denominator from being zero.

[0017] Step S24: Generate dynamic weights through global average pooling and cross-domain interaction to dynamically adjust the importance of frequency-domain and spatial-domain features. First, generate the global artistic photo image feature weights for the frequency domain and spatial domain by global average pooling GAP and MLP, denoted as and Softmax is a normalization function; Based on the global weights, add the interaction information between the frequency domain and the spatial domain: First, generate the importance weights for the spatial domain from the frequency-domain features: W spatial = Softmax(MLP(GAP(F′ freq (X)))), Then generate the importance weights for the frequency domain from the spatial-domain features: W freq = Softmax(MLP(GAP(F′ spatial (X)))), Finally, combine the global weights and cross-domain weights to generate the final weights: λ1, λ2, λ3, λ4 are learnable parameters.

[0018] Step S25: Output the artistic photo image feature Z through weighted fusion of the frequency-domain and spatial-domain features, Z = Weight freq ·F′ freq (X) + Weight spatial ·F′ spatial (X).

[0019] Preferably, in step S2, for the feature extraction module, for the Hybrid Domain Attention Block (HDAB) module, through the dynamic decomposition and fusion of frequency-domain and spatial-domain features, combined with the multi-scale feature extraction and weight adjustment mechanism, the model's ability to capture global information and local details is effectively enhanced; the dynamically generated multi-scale fusion coefficients and the adaptive weight mechanism further strengthen the focus on key information regions while suppressing redundant features, thereby significantly improving the visual expressiveness and detail clarity of artistic photos; in addition, the module design takes into account the balance between computational efficiency and performance and is applicable to image enhancement tasks in multiple scenarios.

[0020] Preferably, in step S3, construct an Adaptive Contrast Enhancement Module (ACEM), which specifically includes the following steps: Step S31, perform feature contrast enhancement. Input the artistic photo image features X processed by the Hybrid Domain Attention Block (HDAB) module, where X ∈ R H ×W×C , where H, W, and C respectively represent the height, width, and number of channels of X; First, normalize the input artistic photo image features to enhance the contrast of the artistic photo image features. The normalized feature map X norm has a distribution with a mean of 0 and a standard deviation of 1: X norm is expressed as: mean(X) and std(X) respectively represent the mean and standard deviation of the feature map X, and ∈ is a small value to prevent the denominator from being zero.

[0021] Step S32, perform contrast stretching on X norm to map the feature value range to a larger dynamic range and enhance the contrast of the features. Specifically: X min and X max are the minimum and maximum values of the feature map X norm , and Clip(x, a, b) means restricting the value of x within the interval [a, b]. This operation sets the values less than a to a and the values greater than b to b.

[0022] Step S33, enhance the features of specific regions in the artistic photo image through the regional saliency enhancement strategy to make the target regions more prominent: First, extract the global artistic photo image feature information through average pooling and max pooling: F avg = AvgPool(X enhanced ), F max = MaxPool(X enhanced ), AvgPool represents average pooling, and MaxPool represents max pooling. Fuse features through convolution operations to generate attention weights. First, fuse features through a 7×7 convolution: F conv =Conv 7×7 ([F avg ,F max ) Conv 7×7 is a 7×7 convolution. [F avg ,F max represents feature concatenation. Use the trigonometric functions sin and cos combined with dynamic adjustment coefficients α and β to generate attention weights: A spatial =sin(α·F conv )+cos(β·F conv ) where α is used to dynamically adjust the frequency of sin, β is used to dynamically adjust the frequency of cos, and the calculation formulas for α and β are: Mean(F conv ) and Std(F conv ) respectively represent the mean and standard deviation of the feature F conv . ∈ is a small value to prevent the denominator from being zero. Then normalize A spatial to ensure that the numerical range is within [0,1]. The specific formula is: max(A spatial ) and min(A spatial ) respectively represent the maximum and minimum values of the feature A spatial . Use the attention weight A spatial to weight the input feature X enhanced point by point. The specific formula is: X spatial =X enhanced ·A spatial .

[0023] Preferably, in step S3, for the Adaptive Contrast Enhancement Module (ACEM), by combining normalization and contrast stretching techniques, the attention weights are dynamically generated and normalized, significantly improving the model's focusing ability on key regions of the image and its ability to suppress redundant information. At the same time, by combining average pooling and max pooling to extract global features and then using dynamic trigonometric functions to generate weights, the model is given stronger expression and adaptive capabilities, effectively enhancing the detail expressiveness and overall visual effect of artistic photos. In addition, this method is flexible, applicable to a variety of complex scenarios, and balances local details and global consistency.

[0024] Preferably, in step S4, a loss function and training strategy module are constructed, which specifically includes the following steps: To ensure that the generated artistic photo image meets the desired effects in both structure and style, we will use three key losses: Content Loss, Style Loss, and High-Frequency Loss: The content loss is calculated using the intermediate layer features of the VGG-19 network. These layers capture the high-level structural features of the image. The formula for the content loss is: and represent the feature maps of the target artistic image and the enhanced image processed by the Adaptive Contrast Enhancement Module (ACEM) at the i-th layer, respectively. N is the total number of elements in the feature map, N = C × H × W, that is, the number of channels × height × width, and ||·|| 2 represents the L2 norm; The Gram matrix is used to capture the style features of the image. The Gram matrix represents the correlation between different features. By comparing the Gram matrices of the target artistic image and the generated image, the style consistency is optimized. The formula for the style loss is: and represent the Gram matrices of the target artistic image and the enhanced image processed by the Adaptive Contrast Enhancement Module (ACEM) at the i-th layer, respectively. M is the total number of elements in the feature map; The high-frequency loss enhances the performance of edges and details by extracting the high-frequency components of the enhanced image and the target artistic image and comparing their differences in the high-frequency space. Specifically, the Laplacian operator is used to extract the high-frequency components (edges and textures) of the image. The formula for the high-frequency loss is: L highfreq = ||Laplacian(G(X)) - Laplacian(X)|| 2 , Let \(X\) denote the target artistic image, and \(G(X)\) denote the enhanced image, i.e., the output after being processed by the HDAB and ACEM modules; Laplacian(\(\cdot\)) represents the Laplacian operator, which is used to extract high-frequency information (edges and textures); To simultaneously optimize the objectives in terms of content, style, and high-frequency loss, a total loss function is designed: \(L\) total \(=\alpha L_{content}\) content \(+\beta L_{style}\) style \(+\gamma L_{high - frequency}\) highfreq where \(\alpha\), \(\beta\), and \(\gamma\) respectively represent the weight coefficients of content loss, style loss, and high - frequency loss. Meanwhile, during the training process, the weights are dynamically adjusted. In the initial stage: \(\alpha>\beta\), mainly optimizing the content loss to ensure the accurate structure of the generated image; in the middle stage: gradually increase \(\beta\) to optimize style consistency; in the later stage: increase \(\gamma\) to enhance the authenticity of the image; emphasizing content loss in the initial stage and increasing the weights of style loss and high - frequency loss in the later stage to optimize the overall artistry of the generated image.

[0025] Preferably, in step S4, for the loss function and training strategy module, the three parts of content loss, style loss, and high - frequency loss are comprehensively used. Through joint optimization, the model's ability to generate artistic photos is improved; content loss focuses on structural consistency, style loss captures the details of style features through the Gram matrix, and high - frequency loss uses the Laplacian operator to strengthen the ability to capture edges and details; the dynamic adjustment strategy of weight parameters further balances the content, style, and detail performance of the generated image, making the optimization process more flexible and efficient, and ultimately enhancing the overall artistry and detail expressiveness of the generated artistic photos.

[0026] Compared with the prior art, the present invention has the following technical effects: By introducing the hybrid - domain attention HDAB module to extract multi - scale features in the frequency domain and spatial domain, combining with the adaptive contrast enhancement ACEM module to dynamically adjust feature contrast, and adopting a multi - objective optimization strategy of content loss, style loss, and high - frequency loss, the present invention realizes the accurate expression and enhancement of the features of artistic photo images; through the dynamic weight generation and optimization mechanism, the present invention effectively highlights the key information regions, suppresses redundant features, and significantly improves the structural consistency, style matching, and detail expressiveness of the generated images. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 FIG. is a flowchart of a method for enhancing an artistic photo image provided by the present invention.

[0028] Figure 2 FIG. is a structural diagram of the hybrid - domain attention HDAB module provided by the present invention.

[0029] Figure 3 It is the structural diagram of the Adaptive Contrast Enhancement Module (ACEM) provided by the present invention.

[0030] Figure 4 It is the low-resolution artistic photo image provided by the present invention.

[0031] Figure 5 It is the high-resolution effect image obtained by processing the low-resolution artistic photo image provided by the present invention. Detailed implementation manners

[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0033] Please refer to the attached Figures 1-5 , the present invention provides a method for enhancing and processing artistic photo images.

[0034] As shown in the attached Figure 1 , the present invention proposes a method for enhancing and processing artistic photo images. The specific implementation manners include the following steps:

[0035] S1. Collect artistic photos. After preprocessing and annotating features, divide them into a training set, a validation set, and a test set according to the ratio of 70%, 20%, and 10% for training the artistic photo image enhancement model.

[0036] Further, in step S1, collect artistic photo images to make a data set, ensuring that the data source has the characteristics of high resolution and no compression damage. Secondly, perform standardized preprocessing on all pictures, uniformly adjust the resolution to 640×640 pixels, and remove existing compression artifacts or noises. Then, according to the artistic style and visual features of the images, manually annotate the target enhancement requirements, including retaining texture details, enhancing color saturation, and style consistency features, to generate an enhanced target image matching each input picture. Finally, divide the data set into a training set, a validation set, and a test set according to the ratio of 70%, 20%, and 10%, ensuring balanced data distribution within each subset.

[0037] S2. Construct a Hybrid Domain Attention Block (HDAB) by using dynamically generated multi-scale fusion coefficients and the proposed cross-domain interaction mechanism. Use the HDAB to extract the frequency-domain and spatial-domain information of the features of artistic photos, perform weighted fusion on high-frequency, low-frequency, and multi-scale features, and enhance the frequency-domain and spatial-domain features through the cross-domain interaction mechanism.

[0038] Further, in step S2, for the Hybrid Domain Attention Block (HDAB), please refer to Figure 2 as shown.

[0039] Step S21: First, extract the frequency domain features from the input artistic photo image feature X, where X ∈ R H×W×C , and H, W, and C represent the height, width, and number of channels of X respectively; use the Fast Fourier Transform (FFT) to extract the frequency domain information of the artistic photo image feature: F freq (X) = FFT(X), Design dynamic filters H H and H L to separate the high-frequency and low-frequency information of F freq (X). The filters H H and H L are dynamically generated by the global information, and the generation formula is: H H , H L = Sigmoid(MLP(GAP(X))), where GAP is the global average pooling and Sigmod is the activation function. The high-frequency information is expressed as: The low-frequency information is expressed as: where IFFT is the inverse operation of the Fourier transform. The high-frequency features retain the texture details and edge information in the artistic photo, and the low-frequency features highlight the global consistency, corresponding to the overall style and color distribution in the artistic photo. Then, perform a non-linear transformation on the separated high-frequency and low-frequency features using a 1×1 convolution. The specific formula is: Conv 1×1 is a 1×1 convolution.

[0040] Step S22: Extract the spatial domain features from the artistic photo image feature X, and use convolution kernels of different sizes to extract multi-scale local features. where Conv k represents a convolution kernel of size k×k, and k ∈ {3, 5, 7}.

[0041] Step S23: Introduce the dynamically generated multi-scale fusion coefficient w k , and perform a weighted sum on the multi-scale features to generate the spatially fused artistic photo image feature: are multi-scale local features extracted using convolution kernels of different sizes, and the multi-scale fusion coefficient w k , and the calculation formula is: where, is the feature mean of the feature map , indicating the global mean calculation of the multi-scale convolution features , is the feature standard deviation of the feature map , indicating the calculation of the local change intensity of the multi-scale convolution features , μ k is the dynamic offset, indicating the prior preference for a specific scale k, and is extracted from the multi-scale convolution features through a non-linear transformation. The formula is: Linear k (·) is a linear transformation operation, C is the number of channels, ∈ is a small value, and ∈ prevents the denominator from being zero.

[0042] Step S24, generate dynamic weights through global average pooling and cross-domain interaction, and dynamically adjust the importance of frequency domain and spatial domain features. First, generate the global artistic image feature weights of the frequency domain and spatial domain by global average pooling GAP and MLP, which are respectively represented as and Softmax is a normalization function; Based on the global weights, add the interaction information between the frequency domain and the spatial domain: first generate the importance weights of the spatial domain from the frequency domain features: W spatial = Softmax(MLP(GAP(F′ freq (X)))), Then generate the importance weights of the frequency domain from the spatial domain features: W freq = Softmax(MLP(GAP(F′ spatial (X)))), Finally, combine the global weights and cross-domain weights to generate the final weights: λ1, λ2, λ3, λ4 are learnable parameters.

[0043] Step S25: Output the artistic photo image feature Z through weighted fusion of the frequency domain and spatial domain features, Z = Weight freq ·F′ freq (X)+Weight spatial ·F′ spatial (X).

[0044] S3: Construct an Adaptive Contrast Enhancement Module (ACEM) through the normalization and contrast stretching method and the proposed regional saliency enhancement strategy. Use the ACEM module to enhance the artistic photo image features, generate attention weights according to the regional saliency enhancement strategy, highlight the key information areas and suppress redundant information.

[0045] Furthermore, in step S3, for the Adaptive Contrast Enhancement Module (ACEM), please refer to Figure 3 as shown.

[0046] Step S31: Perform feature contrast enhancement. Input the artistic photo image feature X processed by the Hybrid Domain Attention Block (HDAB) module, X ∈ R H×W×C , where H, W, and C represent the height, width, and number of channels of X respectively; First, normalize the input artistic photo image feature to enhance the contrast of the artistic photo image feature. The normalized feature map X norm has a distribution with a mean of 0 and a standard deviation of 1: X norm is expressed as: mean(X) and std(X) represent the mean and standard deviation of the feature map X respectively, and ∈ is a small value to prevent the denominator from being zero.

[0047] Step S32: Perform contrast stretching on X norm to map the feature value range to a larger dynamic range and enhance the contrast of the feature. Specifically: X min and X max are the minimum and maximum values of the feature map X norm . Clip(x, a, b) means restricting the value of x within the interval [a, b]. This operation sets the values less than a to a and the values greater than b to b, where a = 0 and b = 1.

[0048] Step S33: Enhance the features of specific regions in the artistic photo image through the regional saliency enhancement strategy to make the target regions more prominent: First, extract the global artistic photo image feature information through average pooling and max pooling: F avg= AvgPool(X enhanced ), F max = MaxPool(X enhanced ), AvgPool represents average pooling, and MaxPool represents max pooling. Fuse features through a convolution operation to generate attention weights. First, fuse features through a 7×7 convolution: F conv = Conv 7×7 ([F avg , F max ), Conv 7×7 is a 7×7 convolution, and [F avg , F max represents feature concatenation. Use the trigonometric functions sin and cos to dynamically adjust the coefficients α and β to generate attention weights: A spatial = sin(α·F conv ) + cos(β·F conv ), where α is used to dynamically adjust the frequency of sin, β is used to dynamically adjust the frequency of cos, and the calculation formulas for α and β are: Mean(F conv ) and Std(F conv ) respectively represent the mean and standard deviation of the feature F conv . ∈ is a small value to prevent the denominator from being zero. Then normalize A spatial to ensure that the numerical range is within [0, 1]. The specific formula is: max(A spatial ) and min(A spatial ) respectively represent the maximum and minimum values of the feature A spatial . Use the attention weight A spatial to weight the input feature X enhanced point by point. The specific formula is: X spatial = X enhanced ·A spatial .

[0049] S4. Design content loss, style loss, and high-frequency loss, construct a loss function and training strategy module, and optimize the parameters of the model by adjusting the weights and learning rate of the loss function to complete the training process of the artistic photo image enhancement model.

[0050] Furthermore, in step S4, for the loss function and training strategy module, the following steps are specifically included: To ensure that the generated artistic photo images achieve the desired effects in both structure and style, we will use three key losses: Content Loss, Style Loss, and High-Frequency Loss: Calculate the content loss using the intermediate layer features of the VGG-19 network. These layers capture the high-level structural features of the image. The formula for the content loss is: and represent the feature maps of the target artistic image and the enhanced image processed by the Adaptive Contrast Enhancement Module (ACEM) at the i-th layer respectively. N is the total number of elements in the feature map, N = C × H × W, that is, the number of channels × height × width. ||·|| 2 represents the L2 norm; Use the Gram matrix to capture the style features of the image. The Gram matrix represents the correlation between different features. Optimize the style consistency by comparing the Gram matrices of the target artistic image and the generated image. The formula for the style loss is: and represent the Gram matrices of the target artistic image and the enhanced image processed by the Adaptive Contrast Enhancement Module (ACEM) at the i-th layer respectively. M is the total number of elements in the feature map; The high-frequency loss enhances the performance of edges and details by extracting the high-frequency components of the enhanced image and the target artistic image and comparing their differences in the high-frequency space. Specifically, use the Laplacian operator to extract the high-frequency components (edges and textures) of the image. The formula for the high-frequency loss is: L highfreq = ||Laplacian(G(X)) - Laplacian(X)|| 2 , X represents the target artistic image, G(X) represents the enhanced image, that is, the output after processing by the HDAB and ACEM modules; Laplacian(·) represents the Laplacian operator, which is used to extract high-frequency information (edges and textures); To simultaneously optimize the objectives in terms of content, style, and high-frequency loss, design the total loss function: L total = αL content + βL style + γL highfreq , α, β, and γ represent the weight coefficients of content loss, style loss, and high-frequency loss respectively. During the training process, the weight coefficients of content loss α, style loss β, and high-frequency loss γ are dynamically adjusted to optimize the effect of artistic image enhancement. The specific adjustment strategy is as follows: Initial stage (the first 0 - 30% of training iterations): Set α = 10, β = 1, γ = 0. Mainly optimize the content loss to ensure the accurate structure of the generated image and avoid premature loss of original features;

[0051] Middle stage (30% - 70% of training iterations): Gradually adjust the weights to enhance the influence of style loss. Set α = 5, β = 5, γ = 1. On the basis of maintaining the content structure, enhance style consistency, and at the same time introduce high-frequency loss to improve the quality of local details;

[0052] Final stage (70% - 100% of training iterations): Further adjust the weights to strengthen style transfer and high-frequency information, make the artistic style more prominent, and the image more realistic and delicate. Set α = 2, β = 8, γ = 5, reduce the influence of content loss, and make the style features and local edges clearer.

[0053] S5. Build an artistic photo image enhancement model, including input, image embedding, Hybrid Domain Attention Block (HDAB) module, Adaptive Contrast Enhancement Module (ACEM), fully connected layer, and output. Use the loss function and training strategy to guide the training of the artistic photo image enhancement model.

[0054] Furthermore, in step S5, for the artistic photo image enhancement model, input a low-resolution artistic photo image I, I ∈ R H×W×C , where H, W, and C represent the height, width, and number of channels of I respectively. In this embodiment, H = 640, W = 640, C = 3, that is, the input image is I ∈ R 640×640×3 . Input I into the image embedding. The image embedding contains a 3×3 convolution to obtain the low-resolution artistic photo image feature X0. Input X0 into the Hybrid Domain Attention Block (HDAB) module to obtain X1. Input X1 into the Adaptive Contrast Enhancement Module (ACEM) to obtain X2. Input X2 into a 3×3 convolution for processing. The processed artistic photo image feature is input into the fully connected layer, and finally output to obtain I + , I + ∈R 640×640×3 .

[0055] Further, in step S5, the artistic photo image enhancement model writes code using the Pycharm application and the Python language, uses the Pytorch framework, and trains with a low-resolution artistic photo image with an input resolution of 640×640×3. The model is trained from scratch for 300 epochs.

[0056] S6. Obtain the artistic photo image that needs to be subjected to image enhancement operations, input it into the artistic photo image enhancement model for processing, and output a high-resolution artistic photo image.

[0057] Further, as Figure 4 and Figure 5 shown, Figure 4 shows a low-resolution artistic photo image, Figure 5 shows a high-resolution artistic photo image obtained by processing with the artistic photo image enhancement model. It can be seen that the clarity and detail performance of its image in the parts of the house, trees, and background have been significantly improved, and the details are presented more clearly.

[0058] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the creative concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A method for enhancing an artistic photo image, characterized in that: The following steps are involved: S1. Collect art photos, pre-process and annotate features, and divide them into training set, validation set and test set in the ratio of 70%, 20% and 10% for training art photo image enhancement model; S2. Through the dynamically generated multi-scale fusion coefficients and the proposed cross-domain interaction mechanism, a hybrid domain attention HDAB module is constructed, and the HDAB module is used to extract the frequency domain and spatial domain information of the artistic photo image features, and the high-frequency, low-frequency and multi-scale features are weightedly fused, and the frequency domain and spatial domain features are enhanced through the cross-domain interaction mechanism; S3. By using normalization and contrast stretching methods and proposing a regional saliency enhancement strategy, an adaptive contrast enhancement ACEM module is constructed. The ACEM module is used to enhance the image features of artistic photos. The attention weight is generated according to the regional saliency enhancement strategy to highlight the key information area and suppress redundant information. S4. Design content loss, style loss and high-frequency loss, construct loss function and training strategy module, and optimize model parameters by adjusting the weight and learning rate of loss function to complete the training process of artistic photo image enhancement model; S5. Construct an artistic photo image enhancement model, including input, image embedding, mixed domain attention HDAB module, adaptive contrast enhancement ACEM module, fully connected layer and output, and use loss function and training strategy to guide the training of the artistic photo image enhancement model; S6. Obtain an artistic photo image that needs to be enhanced, input it into an artistic photo image enhancement model for processing, and output a high-resolution artistic photo image.

2. The method for enhancing an artistic photo image according to claim 1, characterized in that: In step S2, a mixed domain attention HDAB module is constructed. The specific method is as follows: S21, first extract frequency domain features from the input art photo image feature X, X∈R H×W×C , H, W and C represent the height, width and number of channels of X respectively; the frequency domain information of the artistic photo image features is extracted using Fourier transform FFT: F freq (X)=FFT(X), Design dynamic filter H H , H L Separation F freq (X) high-frequency and low-frequency information, filter H H , H L Dynamically generated by global information, the generation formula is: H H ,H L =Sigmoid(MLP(GAP(X))), GAP is the global average pooling, Sigmod is the activation function, High frequency information is represented as: The low-frequency information is represented as: Among them, IFFT is the inverse operation of Fourier transform. Then, the separated high-frequency and low-frequency features are transformed nonlinearly using 1×1 convolution. The specific formula is: Conv 1×1 It is a 1×1 convolution; S22, extract spatial domain features from the artistic photo image feature X, and use convolution kernels of different sizes to extract multi-scale local features. Among them, Conv k represents a convolution kernel of size k×k, k∈{3, 5, 7}; S23, introduce dynamically generated multi-scale fusion coefficient ω k , for multi-scale features Weighted summation to generate spatial domain fusion art photo image features: The multi-scale local features extracted using convolution kernels of different sizes, the multi-scale fusion coefficient ω k , the calculation formula is: in, It is a feature map The feature mean of Calculate the global mean. It is a feature map The characteristic standard deviation of The local variation intensity is calculated. μ k is a dynamic offset, which indicates the prior preference for a specific scale k, and is obtained from multi-scale convolution features through nonlinear transformation. Extracted from, the formula is: Linear k (·) is the linear transformation operation, C is the number of channels, ∈ is a small value, ∈ prevents the denominator from being zero; S24. Generate dynamic weights through global average pooling and cross-domain interaction to dynamically adjust the importance of frequency domain and spatial domain features. First, the global artistic photo image feature weights in the frequency domain and spatial domain are generated by global average pooling GAP and MLP, which are expressed as Softmax is a normalization function; On the basis of global weight, the interactive information between frequency domain and spatial domain is added: first, the importance weight of spatial domain is generated by frequency domain features: W spatial =Softmax(MLP(GAP(F′ freq (X)))), Then the importance weights in the frequency domain are generated through the spatial domain features: W freq =Softmax(MLP(GAP(F′ spatial (X)))), Finally, the global weight and cross-domain weight are combined to generate the final weight: λ1, λ2, λ3, and λ4 are learnable parameters; S25, outputting artistic photo image feature Z by weighted fusion of frequency domain and space domain features, Z=Weight freq ·F′ freq (X)+Weight spatial ·F′ spatial (X)。 3. The method for enhancing an artistic photo image according to claim 2, characterized in that: In the step S3, an adaptive contrast enhancement ACEM module is constructed, and the specific method is as follows: S31, perform feature contrast enhancement, input the artistic photo image feature X processed by the mixed domain attention HDAB module, X∈R H×W×C , H, W, and C represent the height, width, and number of channels of X, respectively; First, the input art photo image features are normalized, and the normalized feature map X norm A distribution with mean 0 and standard deviation 1: norm It is expressed as: mean(X) and std(X) represent the mean and standard deviation of feature map X, respectively. ∈ is a small value to prevent the denominator from being zero. S32, X norm Perform contrast stretching to enhance the contrast of features, specifically: X min and X max is the feature map X norm The minimum and maximum values ​​of , Clip(x,a,b) means limiting the value of x to the interval [a,b]. This operation sets the value less than a to a and the value greater than b to b; S33. Enhance the features of specific regions in the artistic photo image through regional saliency enhancement strategy to make the target region more prominent: First, extract the global artistic photo image feature information through average pooling and maximum pooling: F avg =AvgPool(X enhanced ),F max =MaxPool(X enhanced ), AvgPool means average pooling, MaxPool means maximum pooling, The features are fused through convolution operations to generate attention weights. The features are first fused through 7×7 convolution: F conv =Conv 7×7 ([F avg ,F max ]), Conv 7×7 is a 7×7 convolution, [F avg , F max ] indicates feature concatenation, Use the trigonometric functions sin and cos combined with dynamic adjustment coefficients α, β to generate attention weights: A spatial =sin(α·F conv )+cos(β·F conv ), Among them, α is used to dynamically adjust the frequency of sin, and β is used to dynamically adjust the frequency of cos. The calculation formulas of α and β are: Mean(F conv ) and Std(F conv ) represent the features F conv The mean and standard deviation of ,∈ is a small value,∈ prevents the denominator from being zero, Then for A spatial Normalize to ensure that the value range is within [0,1]. The specific formula is: max(A spatial ) and min(A spatial ) represent the features A spatial The maximum and minimum values ​​of Use the attention weight A spatial For the input feature X enhanced Point-by-point weighting, the specific formula is: X spatial =X enhanced ·A spatial 。 4. The method for enhancing an artistic photo image according to claim 3, characterized in that: In the step S4, a loss function and a training strategy module are constructed, and the specific method is as follows: The content loss is calculated using the intermediate layer features in the VGG network. These layers capture the high-level structural features of the image. The calculation formula for the content loss is: and They represent the feature maps of the target art image and the enhanced image processed by the adaptive contrast enhancement ACEM module at the i-th layer, respectively. N is the total number of elements in the feature map, N = C × H × W, that is, the number of channels × height × width, ||·|| 2 represents the L2 norm; The Gram matrix is ​​used to capture the style features of the image. The Gram matrix represents the correlation between different features. The style consistency is optimized by comparing the Gram matrices of the target art image and the generated image. The calculation formula of the style loss is: and They represent the Gram matrices of the target art image and the enhanced image processed by the adaptive contrast enhancement ACEM module at the i-th layer, respectively, and M is the total number of elements in the feature map; The high-frequency loss extracts the high-frequency components of the enhanced image and the target art image and compares their differences in the high-frequency space, thereby enhancing the performance of edges and details. Specifically, the Laplacian operator is used to extract the high-frequency components (edges and textures) of the image. The calculation formula of the high-frequency loss is: L highfreq =|Laplacian(G(X))-Laplacian(X)|| 2 , X represents the target art image, G(X) represents the enhanced image, that is, the output after processing by the HDAB and ACEM modules; Laplacian(·) represents the Laplacian operator, which is used to extract high-frequency information (edges and textures); In order to optimize the three objectives of content, style and high-frequency loss at the same time, the total loss function is designed: L total =αL content +βL style +γL highfreq , α, β, and γ represent the weight coefficients of content loss, style loss, and high-frequency loss, respectively. At the same time, the weights are dynamically adjusted during the training process. In the early stage: α>β, mainly optimizing content loss to ensure the accuracy of the structure of the generated image; in the middle stage: gradually increase β to optimize style consistency; Late stage: Improve γ to enhance the authenticity of the image; emphasize content loss in the early stage, increase the weight of style loss and high-frequency loss in the late stage, and optimize the overall artistry of the generated image.

Citation Information

Patent Citations

  • Retinal blood vessel segmentation method based on long-range dependency relationship and multi-scale input

    CN116563232A

  • Infrared weak and small target detection method based on adaptive contrast enhancement

    CN118840635A

  • Contextual visual-based SAR target detection method and apparatus, and storage medium

    US20230184927A1

  • Object-level infrared-and-visible-light image fusion method based on fully convolutional neural network

    WO2024174488A1

Cited By

  • Abdominal tuberculosis image processing method

    CN120612251A