Multi-scale image style migration optimization method based on contrast loss and residual connection
By introducing residual connection and contrast loss mechanisms and combining multi-layer feature fusion, the existing image style transfer technology has solved the high computational cost and style and content balance problems in high-resolution image processing, achieving efficient and real-time style transfer effect.
Patent Information
- Application Number
- CN202510312632.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-29
AI Technical Summary
The existing image style transfer technology has high calculation costs and slow speed when processing high-resolution and rich in detail images, making it difficult to meet the real-time processing needs, and it is difficult to achieve the best effect in the balance between style and content features, resulting in insufficient visual hierarchy and structural integrity of the generated images.
By introducing residual connection and contrast loss mechanisms, combining multi-layer feature fusion, image features are extracted using a pre-trained convolutional neural network, and the calculation amount is reduced through 1×1 convolution, the contrast, color and hierarchy of the generated images are optimized to achieve a balance between style and content.
Effectively balance style expression and detail retention, improves the computing efficiency and visual expression of the image, and is suitable for real-time application scenarios.
Smart Images

Figure CN120388180A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of deep learning and computer vision, and particularly relates to a multi-scale image style transfer optimization method based on contrastive loss and residual connection. Background Art
[0002] Current image style transfer technologies are widely applied in fields such as image editing, video processing, and artistic creation. Since the method of using a deep neural network (DNN) combined with a VGG network for style transfer was proposed, this method separates the image style and content through the combination of content loss and style loss, greatly improving the visual quality of style transfer. Subsequently, technologies such as generative adversarial networks (GANs) have also been introduced into style transfer, further enhancing the effect and stability of style transfer. These methods rely on the powerful feature extraction ability of convolutional neural networks (CNNs), gradually extract and combine features from the low level to the high level, making the generated stylized images significantly improved in overall style and detail consistency, and gradually developing towards a delicate and realistic style transfer effect.
[0003] Although style transfer technologies have developed rapidly, existing methods still have some bottlenecks. Existing deep learning models usually have high computational costs and slow speeds when processing high-resolution and detail-rich images, making it difficult to meet the requirements of real-time processing. Secondly, the problem of detail retention in style transfer is also relatively prominent. Existing methods are lacking in fineness, contrast, and edge sharpness, especially in high-resolution scenarios, where blurring or structural distortion will occur. In addition, it is difficult to control the balance between style and content features in traditional style transfer models. Usually, a trade-off needs to be made between style expression and content retention, and this trade-off often cannot achieve the best effect. Moreover, due to the complexity of multi-level feature integration, the model is prone to gradient disappearance or instability during the training process, resulting in the difficulty of stabilizing the quality of style transfer. Existing methods have large computational amounts and slow speeds in high-resolution image processing, making it difficult to meet the requirements of real-time applications. At the same time, it is difficult to achieve an ideal balance between style expression and detail retention, resulting in insufficient visual hierarchy and structural integrity of the generated images. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to provide a multi-scale image style transfer optimization method based on contrastive loss and residual connection. By introducing residual connection and contrastive loss mechanisms, through multi-level feature fusion, it aims to improve training stability, enhance the detail expressiveness of images, and achieve a better balance between style and content fidelity, thereby providing an efficient style transfer solution suitable for real-time applications.
[0005] Technical Solution: A multi-scale image style transfer optimization method based on contrastive loss and residual connection of the present invention includes the following steps:
[0006] Step S1: Scale and crop the content image and the style image to a unified resolution to obtain the preprocessed image;
[0007] Step S2: Use a pre-trained convolutional neural network to extract multi-level features from the preprocessed image to capture the global and local information of the image, and fuse the content and style features;
[0008] Step S3: Introduce a residual connection module in the convolutional neural network to enable the transfer of content features;
[0009] Step S4: Use 1×1 convolution to reduce the number of channels of the feature map, ensuring the integrity of feature information while reducing the computational load;
[0010] Step S5: Introduce a contrastive loss in the loss function. By comparing the detail differences between the generated image and the style image at different feature layers, optimize the contrast, color, and sense of hierarchy of the generated image, and restore the features layer by layer through the decoder to obtain the stylized image.
[0011] Furthermore, step S1 specifically includes the following steps:
[0012] S101: Uniformly scale the content image and the style image to 512×512 pixels to retain the global information of the image;
[0013] S102: Perform central cropping on the scaled image to obtain a fixed size of 256×256 pixels;
[0014] S103: Normalize the pixel values of the cropped image to the range of [0, 1] to unify the data distribution.
[0015] Furthermore, step S2 is specifically as follows: Adopt the pre-trained convolutional neural network VGG-19 to extract multi-level features from the content image and the style image to capture the global and local information of the image and provide a feature basis for style and content fusion. The specific steps are as follows:
[0016] S201: Select to use the pre-trained convolutional neural network VGG-19 and load its pre-trained weights to perform feature extraction operations on the content image I c and the style image I s The convolutional neural network captures different levels of information of the image through a multi-layer structure;
[0017] S202: Use different layers of the convolutional neural network to extract features from the content image I c to capture information from low-level textures to high-level structures. Let the content feature be extracted from the l-th layer of the network Then there is:
[0018]
[0019] Among them, represents the content feature of the l-th layer, and E l represents the feature extraction operation of the l-th layer of the network. The selected extraction layers include ReLU3_1, ReLU4_1, and ReLU5_1 layers, which are used to obtain the details, contours, and structures of the content;
[0020] S203: Extract features from the style image I s at different levels to obtain the style information, specifically:
[0021]
[0022] Among them represents the style feature of the l-th layer. The selected extraction layers include ReLU3_1, ReLU4_1, and ReLU5_1 layers to extract style information at different scales to cover texture, color, and overall style;
[0023] S204: To achieve style expression, further calculate the statistical information from the style features extracted from each layer ;
[0024] The mean calculation formula is as follows:
[0025]
[0026] Among them, N is the total number of elements in the feature map, represents the i-th eigenvalue in the style feature map;
[0027] Capture the texture and color similarity of the style features through the Gram matrix. The Gram matrix calculation formula is as follows:
[0028]
[0029] Among them represents the Gram matrix of the style feature of the l-th layer, and represents the texture and color information by calculating the inner product of the feature maps;
[0030] S205: Fuse the content features and style features to generate the feature representation of the stylized image. The features extracted from each layer are weighted and then superimposed to provide multi-level information for style transfer. The fusion process is expressed as:
[0031]
[0032] Among them, α l and β l are the weight factors of the content and style features of the l-th layer, which are used to adjust the contributions of different feature layers in the final style transfer.
[0033] Further, step S3 specifically includes the following steps:
[0034] S301: Add a residual connection module to the deep convolutional network to ensure the effective transmission of content features. For a given feature F and a convolutional layer operation H(F), the basic calculation method of the residual connection is:
[0035] F out = H(F) + F (6)
[0036] where F out represents the output feature after the residual connection, and H(F) is the convolutional operation on the input feature F;
[0037] S302: In the transmission of content features, the residual connection avoids the problem of gradient disappearance in the deep network. Let be the content feature extracted in the l-th layer. After the residual connection, the feature of the (l + 1)-th layer is:
[0038]
[0039] S303: The introduction of the residual connection enables the network to avoid the problem of gradient disappearance during training by maintaining the transmission of the original features in each layer. For the gradient under the residual connection, its gradient transmission is:
[0040]
[0041] where the additional +1 term comes from the direct path of the residual connection, ensuring that the gradient is passed forward in the deep network.
[0042] Further, step S4 specifically includes the following steps:
[0043] S401: Use a 1×1 convolution to reduce the number of channels of the feature map. For the input feature map F ∈ R H×W×C where the height is H, the width is W, and the number of channels is C. The output feature map of the 1×1 convolution is calculated as:
[0044]
[0045] where are the weight parameters of the 1×1 convolution kernel, and c1 is the number of output channels. This operation only performs weighted summation in the channel dimension and does not change the spatial dimensions H and W;
[0046] S402: By learning the convolution kernel parameters, the network optimizes the feature expression while reducing the number of channels, improving the operation efficiency. Specifically:
[0047] Fc = conv1×1(F) (10)
[0048] Among them, F c represents the feature map after dimensionality reduction by 1×1 convolution, retaining the core content of the feature information.
[0049] Furthermore, step S5 specifically includes the following steps:
[0050] S501: In style transfer, perform mean-variance normalization on the feature map to ensure that the content features and style features can be compared and fused on the same scale. The mean μ(F l ) and standard deviation σ(F l ) are used to normalize the feature map F l to eliminate the differences in feature intensity between different channels and make it have a more consistent numerical distribution. The calculation processes of the mean and standard deviation are as follows:
[0051] μ(F l ) is the mean of the feature map F l , representing the average feature intensity across all channels:
[0052]
[0053] Among them, F l is the original feature of the feature map at layer l, and N is the total number of elements in the feature map F l ;
[0054] σ(F l ) is the standard deviation of the feature map F l , representing the distribution range of the feature intensity:
[0055]
[0056] Among them, represents the square of the difference between the i-th element and the mean of the feature map, used to measure the degree of dispersion of the element relative to the mean. Through mean-variance normalization, the feature map F l is converted into a feature with zero mean and unit variance to reduce the scale differences between different feature channels and make it more comparable. The normalization formula is as follows:
[0057]
[0058] Among them, represents the mean-variance normalized feature of the content image at feature layer l.
[0059] S502: Content loss L cBy calculating the Euclidean distance between the mean-variance normalized features of the generated image and the content image at different feature layers l, it is ensured that the generated image retains the structural information of the content image, and the specific definition is as follows:
[0060]
[0061] Among them, represents the mean-variance normalized feature of the generated image at feature layer l;
[0062] S503: The definition of the style loss is as follows:
[0063]
[0064] where each φ represents the feature map of the layer in the encoder used to calculate the style loss, using the Relu_1_1, Relu_2_1, Relu_3_1, Relu_4_1, and Relu_5_1 layers with equal weights;
[0065] S504: When W f , W g and W h are fixed as the identity matrix, each position in the content feature map can be transformed into the semantically nearest feature in the style feature map. To simultaneously consider the global statistics and semantic local mapping between the content feature and the style feature, an identity loss function is defined as:
[0066] L id1 = ||I cc - I c ||2 + ||I ss - I s ||2 (16)
[0067]
[0068] Among them, I cc represents the content image in the generated image, that is, the prediction result of the model on the content image; I c represents the real content image; I ss represents the style image in the generated image, that is, the prediction result of the model on the style image; I s represents the real style image; N l represents the number of feature layers, that is, the number of layers used to extract content and style features in the deep convolutional network. φ i represents a feature extraction function, referring to the output of a certain convolutional neural network layer. This is used to obtain the feature representation of the image at the i-th layer; φ i (I) is the feature map extracted at the i-th layer on the input image I.
[0069] The entire network is optimized by minimizing the following function:
[0070] L = λ c L c + λ s L s + λ id1 L id1 + λ id2 L id2 (18)
[0071] where L represents the entire optimization objective function, which is a weighted sum of different loss terms, and the network is trained by minimizing L; L c represents the content loss; I s represents the style loss; λ c 、λ s 、λ id1 、λ id2 represent weight factors used to adjust the importance of each loss term in the optimization objective.
[0072] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method of the present invention.
[0073] The present invention also discloses a computer-readable storage medium, on which a computer program / instruction is stored, and when the computer program / instruction is executed by a processor, the steps of the method of the present invention are implemented.
[0074] The present invention also discloses a computer program product, including a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps of the method of the present invention are implemented.
[0075] Advantageous effects: Compared with the prior art, the present invention has the following remarkable advantages:
[0076] 1. Effectively balance style expression and detail retention: By combining contrastive loss and residual connections, the method achieves a good balance between style expression and detail retention. The introduction of residual connections ensures the effective transmission of content features, enabling the generated image to maintain the structure and details of the original image while being stylized, and avoiding information loss caused by over-stylization.
[0077] 2. Efficient transmission and computational optimization: The combined application of residual connections and 1x1 convolutions helps the content features to be efficiently transmitted in the multi-layer network, avoiding the problem of gradient disappearance. At the same time, it reduces the number of feature channels, effectively reduces the computational amount, and improves the computational efficiency of the model. The 1x1 convolution for dimensionality reduction further improves the feature extraction efficiency, making the method suitable for the efficient processing requirements in practical applications.
[0078] 3. Enhance the image's layering and contrast effect: By introducing contrast loss, the details differences between the generated image and the style image at multiple levels of features are optimized, resulting in a significant improvement in the color performance, contrast, and layering of the generated stylized image. This method can better capture the fine texture features of the style image, making the generated result more realistic, natural, and visually expressive. Description of the Drawings
[0079] Figure 1 It is a diagram of the existing style transfer framework;
[0080] Figure 2 It is a diagram of the overall framework of the present invention;
[0081] Figure 3 It is a flowchart of the present invention;
[0082] Figure 4 Effect diagram of style transfer. Detailed Implementation Manner
[0083] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0084] To solve the problem of it being difficult to achieve an ideal balance between style expression and detail retention, the present invention proposes an optimized method for multi-scale image style transfer based on contrast loss and residual connection.
[0085] The present invention first combines the introduction of residual connection and 1×1 convolution in the network to ensure the efficient transmission of image features in the multi-layer convolutional network, avoiding the common problems of gradient disappearance and feature information loss in traditional deep networks. The residual connection enables the content features to be retained, while the 1×1 convolution significantly reduces the computational amount while reducing the number of feature channels, enabling the model to have efficient computing capabilities. In addition, by introducing contrast loss, the model is optimized in capturing style details and enhancing the image layering, resulting in a significant improvement in the color performance, contrast, and detail richness of the generated stylized image. This method not only enhances the naturalness of the image in style expression but also achieves the best fusion of style and details while maintaining the content structure, with broad potential for real-time applications.
[0086] To meet the above requirements, the technical solution adopted by the present invention is as follows, as Figure 3 shown, which specifically includes the following steps:
[0087] S1: Scale the content image and the style image to a unified resolution (such as 512 pixels) and crop them to a size of 256×256 to ensure the consistency of the input data and enhance the adaptability of the model to different images.
[0088] S2: Use a pre-trained convolutional neural network (such as VGG-19) to extract multi-level features from the content and style images to capture the global and local information of the images, providing a rich feature basis for the fusion of style and content.
[0089] S3: Introduce a residual connection module during the feature extraction process to ensure the efficient transmission of content features in the deep convolutional network, avoid gradient vanishing, and thus improve the stability of the network and the content retention effect.
[0090] S4: Use 1×1 convolution to reduce the number of channels of the feature map, reduce the computational amount while ensuring the integrity of the feature information, and improve the operation efficiency of the model.
[0091] S5: Introduce a contrastive loss in the loss function. By comparing the detail differences between the generated image and the style image at different feature layers, optimize the contrast, color, and sense of hierarchy of the generated image, making the style transfer result more visually expressive.
[0092] S6: Input the fused content and style features into the decoder, and gradually restore the features to a stylized image, ensuring that the generated image presents an ideal style effect while maintaining the content structure.
[0093] S7: Use the Adam optimizer to train the model. By dynamically adjusting the weights of the content and style losses, find the best balance point between style expression and content retention, and improve the detail fidelity and structural consistency of the generated image.
[0094] Preferably, in step S! above, to preprocess the content image and the style image and ensure their consistency in size and resolution, the present invention performs normalization processing on the input images to improve the adaptability and effect of the model during the style transfer process. The specific method is as follows:
[0095] S101: Uniformly scale the content image and the style image to 512×512 pixels to retain the global information of the images, and at the same time ensure the same resolution of images from different sources, providing rich input information for the model.
[0096] S102: Perform central cropping on the scaled image to obtain a fixed size of 256×256 pixels. Central cropping can retain the main information of the image to the greatest extent, avoid the loss of edge information, and ensure that the processed image meets the input requirements of the model.
[0097] S103: Normalize the pixel values of the cropped image to the range of [0,1], unify the data distribution, which helps the model to converge quickly and improve the generalization ability for different images.
[0098] Preferably, in step S2, to achieve multi-level feature extraction, the present invention employs a pre-trained convolutional neural network (VGG-19) to extract multi-level features from the content image and the style image, so as to fully capture the global and local information of the images and provide a rich feature basis for style and content fusion. The specific method is as follows:
[0099] S201: Select and use a pre-trained convolutional neural network (VGG-19) and load its pre-trained weights to perform feature extraction operations on the content image I c and the style image I s . The convolutional neural network captures different levels of information of the image through a multi-layer structure.
[0100] S202: Use different layers of the convolutional neural network to extract features from the content image I c to capture information from low-level textures to high-level structures. Assume that the content features are extracted from the l-th layer of the network then there is:
[0101]
[0102] wherein, represents the content features of the l-th layer, and E l represents the feature extraction operation of the l-th layer of the network. Usually, layers such as ReLU3_1, ReLU4_1, and ReLU5_1 can be selected as the extraction layers to obtain the details, contours, and structures of the content.
[0103] S203: Extract features from different levels of the style image I s to obtain rich style information, specifically: where
[0104]
[0105] represents the style features of the l-th layer. Layers such as ReLU3_1, ReLU4_1, and ReLU5_1 are selected to extract style information at different scales to cover textures, colors, and overall styles.
[0106] S204: To achieve style expression, the statistical information of the style features extracted from each layer needs to be further calculated.
[0107] The mean calculation formula is as follows:
[0108]
[0109] where N is the total number of elements in the feature map, represents the i-th eigenvalue in the style feature map.
[0110] The Gram matrix is used to capture the texture and color similarities of style features. The calculation formula of the Gram matrix is as follows:
[0111]
[0112] Here represents the Gram matrix of the style features at the l-th layer, and the texture and color information are represented by calculating the inner product of the feature maps.
[0113] S205: The content features and style features are fused in subsequent steps to generate the feature representation of the stylized image. The features extracted from each layer are weighted and then superimposed respectively to provide multi-level information for style transfer. The fusion process can be expressed as:
[0114]
[0115] Among them, α l and β l are the weight factors of the content and style features at the l-th layer, which are used to adjust the contributions of different feature layers in the final style transfer.
[0116] Preferably, in step S3, to ensure the effective transmission of content features in the deep convolutional network, the present invention introduces a residual connection module in the feature extraction process to achieve efficient feature transmission and enhance the stability and content retention effect of the model. The specific steps are as follows:
[0117] S301: Add a residual connection module in the deep convolutional network to ensure the effective transmission of content features. For a given feature F and a convolutional layer operation H(F), the basic calculation method of the residual connection is:
[0118] F out = H(F) + F (6)
[0119] Among them, F out represents the output feature after the residual connection, and H(F) is the convolutional operation on the input feature F. By directly adding, the original feature is directly transmitted to the subsequent layer, maintaining the consistency of the feature in the deep network.
[0120] S302: In the transmission of content features, the residual connection can effectively avoid the problem of gradient disappearance in the deep network. Let be the content feature extracted at the l-th layer. After the residual connection, the feature of the (l + 1)-th layer is:
[0121]
[0122] This structure can maintain the consistency of content information in different layers and improve the retention effect of content features in the subsequent style transfer process.
[0123] S303: The introduction of residual connection enables the network to avoid the vanishing gradient problem during training and improve the model's stability by maintaining the transmission of original features across layers. For the gradient Under the residual connection, its gradient transmission is as follows:
[0124]
[0125] Among them, the additional +1 term comes from the direct path of the residual connection, ensuring that the gradient is transmitted more stably forward in the deep network.
[0126] Preferably, in step S4, to improve the operation efficiency of the model and ensure the integrity of feature information, the present invention uses 1×1 convolution for dimensionality reduction. The specific steps are as follows:
[0127] S401: Use 1×1 convolution to reduce the number of channels of the feature map, so as to reduce the computational amount while maintaining the integrity of feature information. For the input feature map F ∈ R H×W×C (height H, width W, number of channels C), the output feature map of 1×1 convolution The calculation formula is:
[0128]
[0129] Among them [[ID=2,8]]are the weight parameters of the 1×1 convolution kernel, and c1 is the number of output channels. This operation only performs weighted summation in the channel dimension and does not change the spatial dimensions H and W.
[0130] S402: 1×1 convolution is equivalent to selectively compressing the channels of each pixel point, and retaining the most important information in a lower dimension. By learning the convolution kernel parameters, the network can optimize the feature representation while reducing the number of channels, thereby improving the operation efficiency. Specifically:
[0131] F c = conv1×1(F) (10)
[0132] Among them, F c represents the feature map after dimensionality reduction by 1×1 convolution, retaining the core content of the feature information.
[0133] Preferably, to enhance the visual expressiveness of the style transfer result, the present invention introduces a contrast loss into the loss function to optimize the contrast, color, and layering of the generated image. The specific steps are as follows:
[0134] S501: In style transfer, in order to ensure that the content features and style features can be compared and fused on the same scale, we perform mean-variance normalization on the feature map. Specifically, the mean μ(Fl ) and standard deviation σ(F l ) are used to normalize the feature map F l to eliminate the differences in feature intensities across different channels and make it have a more consistent numerical distribution. The calculation processes of the mean and standard deviation are as follows:
[0135] μ(F l ) is the mean of the feature map F l and represents the average feature intensity across all channels:
[0136]
[0137] where F l is the original feature of the feature map at layer l, and N is the total number of elements in the feature map F l .
[0138] σ(F l ) is the standard deviation of the feature map F l and represents the distribution range of feature intensities:
[0139]
[0140] Perform mean-variance normalization to transform the feature map F l into a feature with zero mean and unit variance to reduce the scale differences between different feature channels and make them more comparable. The normalization formula is as follows:
[0141]
[0142] S502: Content loss L c Ensures that the generated image retains the structural information of the content image by calculating the Euclidean distance between the mean-variance normalized features of the generated image and the content image at different feature layers l. The specific definition is as follows:
[0143]
[0144] where: represents the mean-variance normalized feature of the content image at feature layer l, represents the mean-variance normalized feature of the generated image at feature layer l.
[0145] S503: The style loss is defined as follows:
[0146]
[0147] where each φ represents the feature map of the layer in the encoder used to calculate the style loss. We use the layers Relu_1_1, Relu_2_1, Relu_3_1, Relu_4_1, and Relu_5_1 with equal weights.
[0148] S504: When W f , W g and W h are fixed as the identity matrix, each position in the content feature map can be transformed into the semantically nearest feature in the style feature map. In this case, the system cannot parse out sufficient style features. In MSA, although W f , W g and W h are learnable matrices, our style transfer model can be trained by only considering the global statistics of the style loss Ls. To simultaneously consider the global statistics and semantic local mapping between the content features and style features, we define an identity loss function as
[0149] L id1 = ||I cc - I c ||2 + ||I ss - I s ||2 (16)
[0150]
[0151] The entire network is optimized by minimizing the following function:
[0152] L = λ c L c + λ s L s + λ id1 L id1 + λ id2 L id2 (18)
[0153] Example 1: Comparative Test on Cross-Dataset Style Transfer Effect
[0154] This example compares the difference in style transfer accuracy between the multi-scale image style transfer optimization method (MSF-Net) based on contrast loss and residual connection of the present invention and traditional style transfer methods. Through large-scale tests on multiple style transfer datasets, the advantages of the present invention in enhancing style detail performance and content retention are verified.
[0155] 8000 images were randomly selected and combined from multiple standard style transfer datasets such as the VGG19 style image set, WikiArt, and COCO as the test set. These images cover a variety of styles (oil painting, impressionism, abstract art, etc.) and contain different resolutions and scenes to ensure the diversity and representativeness of the dataset.
[0156] In this experiment, PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index), and Style Detail Retention (SDR) were used as evaluation metrics to calculate the difference between the generated image and the style image. To ensure fairness, MSF-Net was compared with traditional VGG-based style transfer methods, and the same evaluation criteria and parameter settings were adopted.
[0157] Each test image was input into MSF-Net and divided into content and style feature streams after data preprocessing. The feature extraction module included a combination of residual connections and 1×1 convolutions. By optimizing the contrast, color performance, and detail retention of the generated image through contrast loss, the best balance between style and detail was ensured.
[0158] The experimental results show that the MSF-Net of the present invention has a PSNR of 34.2 dB on the test set, which is 3.4 dB higher than that of the traditional method (30.8 dB); the SSIM value is 0.91, while that of the traditional method is 0.83, indicating a significant improvement in the structural consistency of the generated image. The Style Detail Retention (SDR) has increased by 15%, especially in complex style images, where MSF-Net can better capture details and textures.
[0159] Example 2: Test on the Optimization Effect of Low-Resolution Image Style Transfer
[0160] This example verified the style transfer effect of the multi-scale image style transfer optimization method based on contrast loss and residual connection of the present invention on low-resolution images, especially the optimization in terms of style detail retention.
[0161] To simulate the style transfer problem of low-resolution images, in this experiment, the content image and the style image were scaled to 128×128 pixels and 256×256 pixels respectively to test the style transfer effect of low-resolution images. By comparing MSF-Net with traditional VGG-based style transfer methods, the advantages of the present invention under low-resolution conditions were verified. All images were scaled to 512 pixels and cropped to 128×128 pixels and 256×256 pixels to simulate low-resolution images. The VGG-19 network was used to extract multi-level features of the content image and the style image, the feature transfer was optimized through residual connections, and 1×1 convolutions were used to reduce the computational amount. A contrast loss was introduced to optimize the style transfer effect in low-resolution images and ensure the natural expression of color, contrast, and details. On the 128×128 pixel image, the PSNR of MSF-Net was 30.5dB, which was 3.8dB higher than that of the traditional method (26.7dB). The SSIM was 0.85, compared with 0.78 of the traditional method, indicating that MSF-Net could maintain a better image structure in low-resolution images. The style detail retention of the present invention was improved by 12% on low-resolution images. Especially in the performance of color and texture, MSF-Net showed significant advantages.
[0162] This embodiment verified the significant advantages of MSF-Net in the style transfer of low-resolution images, especially in the retention of style details and the optimization of contrast and color performance. The generated stylized images had stronger expressiveness than traditional methods.
[0163] Example 3: Application of Style Transfer in Real-Time Image Streams
[0164] This embodiment verified the application of the style transfer optimization method of the present invention in real-time image streams, especially the real-time style transfer performance in virtual reality (VR) and augmented reality (AR). MSF-Net was applied in the real-time video stream, and 30 frames were processed per second for style transfer. The PSNR, SSIM, and the retention of style details in the real-time image stream were evaluated. Among them, in the 30fps video stream, the PSNR of MSF-Net was 32.8dB, which was significantly higher than 29.1dB of the traditional method. The SSIM was 0.88, which performed better than 0.82 of the traditional method. The improvement in the retention of style details indicated that MSF-Net could efficiently perform style transfer in real-time image streams, and the style details were natural and rich.
[0165] This embodiment proved the efficiency and accuracy of MSF-Net in real-time image streams, and it could achieve high-quality style transfer without reducing real-time performance and fluency, making it suitable for real-time application scenarios such as VR / AR.
[0166] Example 4: Application of Style Transfer in Art Education
[0167] This embodiment verifies the application of the multi-scale image style transfer optimization method (MSF-Net) based on contrastive loss and residual connection of the present invention in the field of art education, especially its performance in real-time art creation and style analysis. The effect diagrams are as Figure 4 shown.
[0168] During the art education process, students usually need to learn the characteristics of different art styles and create artworks imitating specific styles through creation. The method of the present invention can use style transfer technology to convert ordinary photos into artworks with specific art styles in real time, helping students more intuitively understand the characteristics of different art styles.
[0169] In this experiment, works of styles such as Impressionism, Post-Impressionism, Cubism, and Abstract Art were selected from WikiArt as the target style image set. 5000 natural landscape and portrait photos were selected as content images for generating stylized works. The content images were input into MSF-Net and fused with the selected style images. The generated stylized images were used for art style learning and creative practice. It can not only provide efficient technical support for the learning and analysis of art styles, but also stimulate students' creative interest and provide innovative practical tools for art education.
Claims
1. An optimization method for multi-scale image style transfer based on contrastive loss and residual connection, characterized in that, It includes the following steps: Step S1: Scale and crop the content image and the style image to a unified resolution to obtain the preprocessed image; Step S2: Use a pre-trained convolutional neural network to extract multi-level features from the preprocessed image to capture the global and local information of the image, and fuse the content and style features; Step S3: Introduce a residual connection module in the convolutional neural network to forward the content features; Step S4: Use 1×1 convolution to reduce the number of channels of the feature map, reducing the computational amount while ensuring the integrity of the feature information; Step S5: Introduce a contrastive loss in the loss function. By comparing the detail differences between the generated image and the style image at different feature layers, optimize the contrast, color, and sense of hierarchy of the generated image, and restore the features layer by layer through the decoder to obtain the stylized image.
2. An optimization method for multi-scale image style transfer based on contrastive loss and residual connection according to claim 1, characterized in that Step S1 specifically includes the following steps: S101: Uniformly scale the content image and the style image to 512×512 pixels to retain the global information of the image; S102: Perform central cropping on the scaled image to obtain a fixed size of 256×256 pixels; S103: Normalize the pixel values of the cropped image to the range of [0,1] to unify the data distribution.
3. An optimization method for multi-scale image style transfer based on contrastive loss and residual connection according to claim 1, characterized in that, Step S2 is specifically: Adopt the pre-trained convolutional neural network VGG-19 to extract multi-level features from the content image and the style image to capture the global and local information of the image, and provide a feature basis for the fusion of style and content. The specific steps are as follows: S201: Select to use the pre-trained convolutional neural network VGG-19 and load its pre-trained weights to perform feature extraction operations on the content image I c and the style image I s The convolutional neural network captures different levels of information of the image through a multi-layer structure; S202: For the content image I c Extract features using different layers of the convolutional neural network to capture information from low-level textures to high-level structures. Let the content features be extracted from the l-th layer of the network Then we have: Among them, represents the content feature of the l-th layer, and E l represents the feature extraction operation of the l-th layer of the network. The selection includes the ReLU3_1, ReLU4_1, and ReLU5_1 layers as the extraction layers, which are used to obtain the details, outlines, and structures of the content; S203: Extract features from the style image I s at different levels to obtain the style information, specifically: Among them, represents the style feature of the l-th layer. The selection includes ReLU3_1, ReLU4_1, and ReLU5_1 layers to extract style information at different scales to cover texture, color, and overall style. S204: To achieve style expression, the style features extracted from each layer Further calculate its statistical information; The mean calculation formula is as follows: where N is the total number of elements in the feature map, represents the i-th eigenvalue in the style feature map; Capture the texture and color similarity of the style features through the Gram matrix. The Gram matrix calculation formula is as follows: Among them, represents the Gram matrix of the style features of the l-th layer, and represents the texture and color information by calculating the inner product of the feature maps; S205: Fuse the content features and the style features to generate a feature representation of the stylized image. The features extracted from each layer are weighted and then superimposed to provide multi-level information for style transfer. The fusion process is expressed as: Among them, α l and β l are the weight factors of the content and style features of the l-th layer, which are used to adjust the contributions of different feature layers in the final style transfer.
4. An optimization method for multi-scale image style transfer based on contrastive loss and residual connection according to claim 1, characterized in that Step S3 specifically includes the following steps: S301: Add a residual connection module in the deep convolutional network to ensure the effective transfer of the content features. For the given feature F and the convolutional layer operation H(F), the basic calculation method of the residual connection is: F out = H(F) + F (6) Among them, F out represents the output feature after the residual connection, and H(F) is the convolution operation on the input feature F; S302: In the transmission of content features, residual connections avoid the problem of gradient vanishing in deep networks. Let be the content features extracted in the l-th layer. After passing through the residual connection, the features of the (l + 1)-th layer are: S303: The introduction of residual connections enables the network to avoid the problem of gradient vanishing during training by maintaining the transmission of the original features through each layer. For the gradient Under residual connections, its gradient transmission is as follows: Among them, the additional +1 term comes from the direct path of the residual connection to ensure that the gradient is forwarded in the deep network.
5. A multi-scale image style transfer optimization method based on contrastive loss and residual connection according to claim 1, characterized in that, Step S4 specifically includes the following steps: S401: Reduce the number of channels of the feature map using 1×1 convolution. For the input feature map F ∈ R H×W×C where the height is H, the width is W, and the number of channels is C. The output feature map of the 1×1 convolution The calculation formula is: Among them, is the weight parameter of the 1×1 convolution kernel, c1 is the number of output channels, and this operation only performs weighted summation in the channel dimension without changing the spatial dimensions H and W; S402: Through the learning of the convolutional kernel parameters, the network optimizes the feature expression while reducing the number of channels, improving the operation efficiency. Specifically: F c = conv1×1(F) (10) Among them, F c represents the feature map after 1×1 convolution dimensionality reduction, retaining the core content of the feature information.
6. A multi-scale image style transfer optimization method based on contrastive loss and residual connection according to claim 1, characterized in that, Step S5 specifically includes the following steps: S501: In style transfer, perform mean-variance normalization on the feature map to ensure that content features and style features can be compared and fused on the same scale. The mean μ(F l ) and standard deviation σ(F l ) are used to normalize the feature map F l to eliminate the differences in feature intensities across different channels and make it have a more consistent numerical distribution. The calculation processes of the mean and standard deviation are as follows: μ(F l ) is the mean of the feature map F l , representing the average feature intensity across all channels: Among them, F l is the original feature of the feature map on layer l, and N is the total number of elements in the feature map F l ; σ(F l ) is the standard deviation of the feature map F l , representing the distribution range of the feature intensity: Among them, represents the square of the difference between the i-th element and the mean of the feature map, which is used to measure the degree of dispersion of the element relative to the mean. Through mean-variance normalization, the feature map F l is converted into features with zero mean and unit variance to reduce the scale difference between different feature channels and make them more comparable. The normalization formula is as follows: Among them, represents the mean-variance normalized feature of the content image on the feature layer l; S502: Content loss L c By calculating the Euclidean distance between the mean-variance normalized features of the generated image and the content image on different feature layers l, it is ensured that the generated image retains the structural information of the content image, and the specific definition is as follows: Among them, represents the mean-variance normalized feature of the generated image on the feature layer l; S503: The definition of the style loss is as follows: Where each φ represents the feature map of the layer used to calculate the style loss in the encoder, using the Relu_1_1, Relu_2_1, Relu_3_1, Relu_4_1, and Relu_5_1 layers with equal weights; S504: When W f , W g and W h are fixed as the identity matrix, each position in the content feature map can be transformed into the semantically nearest feature in the style feature map. To simultaneously consider the global statistics and semantic local mapping between the content feature and the style feature, an identity loss function is defined as: L id1 = ||I cc -I c ||2 + ||I ss -I s ||2 (16) Among them, I cc represents the content image in the generated image, that is, the prediction result of the model on the content image; I c represents the real content image; I ss represents the style image in the generated image, that is, the prediction result of the model on the style image; I s represents the real style image; N l represents the number of feature layers, that is, the number of layers used to extract content and style features in the deep convolutional network, φ i represents a feature extraction function, which refers to the output of a certain convolutional neural network layer and is used to obtain the feature representation of the image at the i-th layer; φ i (I) is the feature map extracted at the i-th layer on the input image I; The entire network is optimized by minimizing the following function: L = λ c L c + λ s L s + λ id1 L id1 + λ id2 L id2 (18) Among them, \(L\) represents the entire optimization objective function, which is the weighted sum of different loss terms, and the network is trained by minimizing \(L\); \(L\) c represents the content loss; \(L\) s represents the style loss; \(\lambda\) c , \(\lambda\) s , \(\lambda\) id1 , \(\lambda\) id2 represent weight factors used to adjust the importance of each loss term in the optimization objective.
7. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method described in claim 1.
8. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method described in claim 1 are implemented.
9. A computer program product comprising computer programs / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method described in claim 1 are implemented.
Citation Information
Patent Citations
Method for realizing real image style migration based on global information guide network
CN113570500A
Image style migration method and system based on deep learning
CN114581341A