Image style processing framework based on deep learning

By introducing Lora-based style migration module and GAN network repaint network module in the image style processing framework, the problems of insufficient model expansion, high computational complexity, inflexible style control and low image details in image style migration are solved, and more efficient style expansion and detail control are achieved.

CN120013747APending Publication Date: 2025-05-16GIANT MOBILE TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510090408.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient model scalability, high computational complexity, inflexible style control and low image details in image style transfer.

Method used

The image style processing framework based on deep learning is adopted, including image data processing module, style transfer module and redraw network module. The style transfer module is based on Lora's training method, adding style matching loss and time domain-based texture matching loss to form a total loss function; the redrawing network module adopts the GAN network structure design, and the generator and discriminator generate rich detailed high-resolution images through adversarial training.

Benefits of technology

It improves the scalability and computing efficiency of the model, enhances the flexibility of style control and the expressiveness of image details, and solves the problems of insufficient scalability of the model, high computational complexity, inflexible style control and low image details in traditional frameworks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013747A_ABST
    Figure CN120013747A_ABST
Patent Text Reader

Abstract

The invention provides an image style processing framework based on deep learning, and relates to the field of image style processing, the image style processing framework comprises an image data processing module, a style migration module and a redrawing network module, the image data processing module comprises data cleaning, data preprocessing and vllm-based style classification; the style migration module adds style matching loss and texture matching loss based on a time domain based on a Lora training mode, then forms a total loss function based on the style matching loss and the texture matching loss based on the time domain, and trains a Lora inserted low-rank matrix by adding a corresponding weight coefficient. According to the method, a low-rank adaptation model and a redrawing network module are introduced into an existing style migration framework, and the problems of insufficient model expansibility, high calculation complexity, inflexible style control and low image details in a traditional framework are solved by realizing efficient expansion of styles and fine control of details.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image style processing, and in particular to an image style processing framework based on deep learning. Background Art

[0002] "Giant Mojing" is a cloud-based AI painting platform for the pan-entertainment industry, focusing on providing efficient and convenient creation tools and resources for art creators. The platform integrates an advanced model reasoning framework to support a variety of complex AI functions, such as style transfer, image generation, and intelligent completion. The reasoning framework is optimized for the special needs of cloud painting, and can dynamically allocate computing resources, improve concurrent processing capabilities, and ensure low-latency reasoning performance, bringing users a smooth creation experience.

[0003] Image style transfer is a technique that applies a specific artistic style to an input image. Its core is to extract the features of the target style through a neural network and apply these features in the image generation process. However, existing technologies still face the following problems:

[0004] 1. Insufficient model scalability: Traditional neural networks often need to retrain the entire model when adding new styles, which is costly.

[0005] 2. High computational complexity: Especially in high-resolution image migration, inference efficiency becomes a bottleneck.

[0006] 3. Inflexible style control: Existing methods have limited support for regionalized style application and style intensity adjustment.

[0007] 4. Low image details: Existing methods have limited performance in image details. Summary of the invention

[0008] In order to make up for the above shortcomings, the present invention provides an image style processing framework based on deep learning, aiming to improve the problems of insufficient model scalability, high computational complexity, inflexible style control and low image details in the prior art.

[0009] The present invention is achieved in that:

[0010] The present invention provides an image style processing framework based on deep learning, comprising:

[0011] Image data processing module, which includes data cleaning, data preprocessing and VLLM-based style classification;

[0012] The style transfer module is based on the training method of Lora, adding style matching loss and time-domain-based texture matching loss. Then, based on the style matching loss and time-domain-based texture matching loss, a total loss function is formed, and the corresponding weight coefficient is added to train the low-rank matrix inserted by Lora, where;

[0013] L total =L task +λ1L perceptual +λ2L adv +λ3L style +λ4L texture ,

[0014] is the total loss function; L task is the perceived loss; L perceptual is against loss; L adv is the style matching loss; Lt exture is the frequency domain texture matching loss;

[0015] The redrawing network module adopts the GAN network structure design. The generator is responsible for mapping the input low-resolution or blurred image to a high-resolution image with rich details; the discriminator is used to judge the difference between the generated image and the real image. It guides the generator to improve the optimization model capabilities by comparing the generated image with the real image and feeding back the loss value.

[0016] Preferably, the data cleaning is to clean and standardize the image based on a denoising and enhancement algorithm.

[0017] Preferably, the data preprocessing comprises the following steps:

[0018] S1. Read the image: First, use OpenCV to load the original image;

[0019] S2. Calculate target resolution: determine the fixed resolution required for model input;

[0020] S3, image resizing: resizing the original image to the target resolution by image scaling, using an interpolation method that maintains the aspect ratio;

[0021] S4, fill or crop: When the adjusted size of the image does not match the aspect ratio of the target resolution, choose to crop the image to ensure that the final size fully meets the input requirements and fill it by adding a background color;

[0022] S5. Normalize pixel values: Normalize the pixel values ​​of the image and convert the pixel values ​​from [0, 255] to [0, 1] to adapt to the input requirements of the model.

[0023] Preferably, the vllm-based style classification comprises the following steps:

[0024] S1. Use ResNet to extract image features. The image is input into the network and processed through multiple convolutional layers and pooling layers, and finally the feature representation is output in the fully connected layer;

[0025] S2. After obtaining the feature vector, if its dimension is too high, use PCA dimensionality reduction technology to reduce the computational complexity and remove redundant information, and then input the extracted high-dimensional image features into the decision tree for category prediction;

[0026] S3. Use cross-validation evaluation methods to verify the performance of the model, and adjust the parameters and strategies of the classification model based on the results to optimize the prediction effect;

[0027] S4. Adopt a hierarchical classification strategy, first use the large category model for coarse classification, and then use the sub-model for detailed classification.

[0028] Preferably, in the S4 process of the vllm-based style classification step:

[0029] The coarse classification of the large category model is to extract high-dimensional features through a pre-trained large convolutional neural network, and input it into a relatively simple multi-layer perceptron classifier for coarse classification.

[0030] Preferably, in the S4 process of the vllm-based style classification step:

[0031] The sub-model performs fine classification by using the style dataset to fine-tune the pre-trained model, adding style-related labels to improve classification accuracy; and using multiple GPUs or the distributed reasoning framework Triton to improve the classification speed in concurrent situations.

[0032] Preferably, the perceptual loss is defined as follows:

[0033]

[0034] Where: G is the generated image and T is the target image.

[0035] Preferably, the adversarial loss is defined as follows:

[0036]

[0037] Preferably, the style matching loss is defined as follows:

[0038]

[0039] Preferably, the frequency domain texture matching loss is defined as follows:

[0040]

[0041] The beneficial effects of the present invention are:

[0042] The present invention introduces a low-rank adaptation model and a redrawing network module into the existing style transfer framework, and solves the problems of insufficient model scalability, high computational complexity, inflexible style control and low image details in the traditional framework by achieving efficient style expansion and refined control of details. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0044] Figure 1 It is a structural diagram of an image style processing framework based on deep learning provided by an embodiment of the present invention;

[0045] Figure 2 It is a data preprocessing flow chart in a deep learning-based image style processing framework provided by an embodiment of the present invention;

[0046] Figure 3 It is a style classification flowchart based on vllm in a deep learning-based image style processing framework provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0048] Example

[0049] Reference Figure 1-Figure 3 , an image style processing framework based on deep learning, including an image data processing module, a style transfer module and a redrawing network module.

[0050] Among them, the image data processing module includes data cleaning, data preprocessing and vlllm-based style classification. Its efficient image data processing process is as follows:

[0051] 1) Data cleaning. Remove blurry, duplicate or irrelevant images to ensure the purity of the data set, and apply denoising and enhancement algorithms to clean and standardize the images. Denoising and enhancement algorithms such as the CLAHE algorithm are image processing algorithms that are mainly used to enhance the contrast of images. It works by dividing the image into multiple small areas (called tiles), and then performing histogram equalization on each small area. Unlike ordinary histogram equalization, CLAHE limits the degree of contrast enhancement to avoid problems such as over-amplifying the noise in the image.

[0052] 2) Data preprocessing. Normalize the image size to a fixed resolution to adapt to the model input. Improve the efficiency and effect of model training and reasoning. The specific steps are as follows:

[0053] Reading the image: First, load the original image using OpenCV.

[0054] Compute target resolution: Determines the fixed resolution required for the model input.

[0055] Resize Image: Resize the original image to the target resolution by scaling the image. Use an interpolation method that maintains the aspect ratio.

[0056] Fill or Crop: If the image is not in the right aspect ratio after resizing, choose to crop the image to ensure that the final size is exactly the same as the input size. Fill by adding a background color, such as black.

[0057] Normalize pixel values: Finally, normalize the pixel values ​​of the image (convert the pixel values ​​from [0, 255] to [0, 1] to fit the input requirements of the model.

[0058] 3) VLLM-based style classification. Select appropriate multi-modal extraction of high-dimensional feature vectors of images and classify them. Use ResNet to extract image features. The features extracted in the convolutional layer can usually effectively capture the abstract information in the image. In order to obtain high-dimensional feature vectors from these models, the image is input into the network and processed through multiple convolutional layers and pooling layers, and finally the feature representation is output in the fully connected layer. When the feature vector is obtained, if its dimension is too high, the PCA dimensionality reduction technique is used to reduce the computational complexity and remove redundant information for subsequent processing. Finally, the extracted high-dimensional image features are input into the decision tree for category prediction. After training, the cross-validation evaluation method is used to verify the performance of the model, and the parameters and strategies of the classification model are adjusted according to the results to optimize the prediction effect.

[0059] A hierarchical classification strategy is adopted. First, the system uses a large category model for rough classification, such as abstract and realistic. For example, the image is classified according to its abstract degree, and the image is roughly divided into two categories: "abstract" and "realistic". In this stage, a pre-trained large convolutional neural network is used to extract high-dimensional features and input them into a relatively simple multi-layer perceptron classifier for rough classification. The purpose of this step is to decompose complex classification tasks into relatively simple subtasks through rough category division, thereby reducing the difficulty of subsequent fine classification.

[0060] Then use sub-models for fine-tuning, such as impressionism and cubism. At the same time, for modules with poor recognition effects. Use style datasets to fine-tune the pre-trained model and add style-related labels to improve classification accuracy. And use multi-GPU or distributed reasoning framework Triton to improve classification speed in high concurrency. In order to solve the performance bottleneck problem when reasoning on large-scale datasets, you can use multi-GPU or NVIDIA Triton distributed reasoning framework. Triton is a high-performance inference server that supports multi-model parallel reasoning and distributed reasoning. By assigning classification tasks to multiple GPUs for parallel processing, the classification speed in high concurrency can be greatly improved. Triton supports different hardware accelerations, such as GPU and TensorRT, and can automatically select the best execution strategy to improve performance when reasoning. By integrating Triton, the response time of the classification system can be effectively compressed to milliseconds, which is especially suitable for real-time application scenarios.

[0061] The style transfer module is based on the traditional Lora training method, and adds style matching loss and time-domain-based texture matching loss during the Lora training process. Style matching loss is used to measure the matching degree between an image or generated content and a specific style. Time-domain texture matching loss is used to measure texture consistency in the time domain, where the time domain, such as video or sequence data, mainly focuses on the smooth transition of texture between frames. This loss function can reduce the discontinuity of texture within the time span of the image generation process by comparing the texture features of the image or video sequence.

[0062] In image generation or image conversion tasks, the style matching loss calculates the style difference between the generated image and the target style image. Commonly used metrics include the Gram matrix or similarity based on convolutional features.

[0063] Among them, Lora is a lightweight model adaptation method that greatly reduces the amount of training parameters and storage overhead by inserting a low-rank matrix update method in a specific layer of the model. In the style transfer network, the goal of Lora is to inject specific style information into the basic generative model without retraining the entire model. Insert the Lora module into the key layers of the base model, such as the self-attention layer, convolution layer, or decoder part. We first collect a dataset of multi-style images, such as impressionist or abstract styles, and annotate style labels for each image. Most of the weights of the base generative model are frozen during training, and only the low-rank matrix inserted into the Lora module is trained.

[0064] When designing the total loss function, it includes perceptual loss, adversarial loss, style matching loss, and frequency domain texture matching loss. Perceptual loss: ensures the consistency of high-level semantic features between the generated image and the target style image; adversarial loss: enhances the realism and artistry of the generated image; style matching loss: ensures style texture consistency through Gram matrix comparison.

[0065] In the frequency domain, a frequency domain texture matching loss function is designed by calculating the difference in amplitude spectrum and phase spectrum between the reference image and the generated image. The input image is converted from the spatial domain to the frequency domain using the fast Fourier transform, and the high-frequency component (detailed texture) and low-frequency component (overall structure) of the image are separated. The learning ability of the network is further optimized through the supervision of the high-frequency components in the frequency domain, and the network's control over texture details is improved.

[0066] 1. Perceptual loss aims to ensure the consistency of high-level semantic features between the generated image and the target image. The high-level features of the image are extracted by using a pre-trained deep convolutional neural network, where the activation difference between the generated image and the target image in some convolutional layers is used as the loss.

[0067] The perceptual loss is defined as follows:

[0068]

[0069] Where: G is the generated image and T is the target image.

[0070] 2. Adversarial loss is used to enhance the authenticity of generated images, and improves the quality of generated images through the game between the discriminator and the generator in the generative adversarial network. The discriminator tries to distinguish between real and generated images, while the generator is constantly optimized to "fool" the discriminator.

[0071] The adversarial loss is defined as:

[0072]

[0073] 3. Style matching loss ensures style consistency by calculating the difference between the Gram matrix of the generated image and the target style image. The Gram matrix describes the local features and texture information of the image and can be used to measure the style of the image.

[0074] The style matching loss is defined as follows:

[0075]

[0076] 4. Frequency domain texture matching loss. In the frequency domain, texture details are usually manifested as high-frequency components, while the overall structure of the image is mainly determined by low-frequency components. In order to enhance the detailed texture of the generated image, the image is converted from the spatial domain to the frequency domain using fast Fourier transform, the high-frequency and low-frequency components of the image are separated, and the high-frequency texture consistency of the generated image is supervised in the frequency domain.

[0077] The design idea of ​​the frequency domain texture matching loss is as follows: apply FFT to the generated image and the target image to obtain their amplitude spectrum and phase spectrum; calculate the difference between the amplitude spectrum and the phase spectrum, especially in the high-frequency area, which corresponds to the detail texture part.

[0078] The frequency domain texture matching loss is defined as follows:

[0079]

[0080] Adjust the weights of the magnitude spectrum and phase spectrum in the total loss.

[0081] 5. The total loss function is a combination of all loss functions and adds the corresponding weight coefficients:

[0082] L total =L task +λ1L perceptual +λ2L adv +λ3L style +λ4L texture For different styles, you only need to train the Lora module for the new style without modifying the basic model or the existing style weights.

[0083] In order to generate images with detailed generation capabilities, the system has a detail redrawing network that focuses on enhancing and redrawing the details in the image during the image generation process. Through adversarial training of the generator and the discriminator, the generated image is not only similar to the real image in global structure, but also has a higher sense of reality in details. Among them, the redrawing network module adopts the GAN network structure design, and the generator is responsible for mapping the input low-resolution or blurred image to a high-resolution image with rich details.

[0084] In order to focus on the restoration of details, the discriminator is used to judge the difference between the generated image and the real image. It guides the generator to improve the optimization model capabilities by comparing the generated image with the real image and feeding back the loss value.

[0085] In the design of the GAN network structure, the collaborative work of the generator and the discriminator is crucial. The task of the generator is to map the input low-resolution or blurred image to a high-resolution image with rich details. The generator learns how to restore the high-frequency details of the image through a series of convolution, deconvolution, activation function and other layers.

[0086] In order to ensure that the generated image has higher detail quality, the discriminator is responsible for judging the difference between the generated image and the real image. The discriminator compares the generated image with the real image, calculates and returns a loss value to guide the optimization of the generator. The optimization process is as follows:

[0087] First, the generator generates a fake image through noise input or low-resolution images, and the discriminator receives the generated image and the real image as input to judge its "authenticity". Through this adversarial process, the generator continuously adjusts its weights to generate more realistic images, while the discriminator continuously improves its ability to identify generated images.

[0088] Secondly, the generator and the discriminator continuously optimize through mutual competition, so that the generated image is not only close to the real image in overall structure, but also matches the high-resolution image in terms of detail texture, clarity, etc.

[0089] Furthermore, during the training process, the discriminator calculates the difference between the generated image and the real image through the cross entropy loss function, and optimizes the network parameters through back propagation, while the generator improves the image quality by minimizing the error rate of the discriminator.

[0090] Ultimately, the GAN network continues to improve by combining the high-resolution image output of the generator with the feedback of the discriminator until the generated image reaches a very realistic and high-quality level in both detail and structure.

[0091] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A deep learning-based image style processing framework, characterized in that: include: Image data processing module, which includes data cleaning, data preprocessing and VLLM-based style classification; The style transfer module is based on the training method of Lora, adding style matching loss and time-domain-based texture matching loss. Then, based on the style matching loss and time-domain-based texture matching loss, a total loss function is formed, and the corresponding weight coefficient is added to train the low-rank matrix inserted by Lora, where; L total =L task +λ1L perceptual +λ2L adv +λ3L style +λ4L texture , is the total loss function; L task is the perceived loss; L perceptual is against loss; L style is the style matching loss; L texture is the frequency domain texture matching loss; The redrawing network module adopts the GAN network structure. The generator is responsible for mapping the input low-resolution or blurred image to a high-resolution image with rich details; the discriminator is used to judge the difference between the generated image and the real image. It guides the generator to improve the optimization model by comparing the generated image with the real image and feeding back the loss value.

2. The image style processing framework based on deep learning according to claim 1, characterized in that: The data cleaning is to clean and standardize the image based on the denoising and enhancement algorithm.

3. The image style processing framework based on deep learning according to claim 1, characterized in that: The data preprocessing comprises the following steps: S1. Read the image: First, use OpenCV to load the original image; S2. Calculate target resolution: determine the fixed resolution required for model input; S3, image resizing: resizing the original image to the target resolution by image scaling, using an interpolation method that maintains the aspect ratio; S4, fill or crop: When the adjusted size of the image does not match the aspect ratio of the target resolution, choose to crop the image to ensure that the final size fully meets the input requirements and fill it by adding a background color; S5. Normalize pixel values: Normalize the pixel values ​​of the image and convert the pixel values ​​from [0, 255] to [0, 1] to adapt to the input requirements of the model.

4. The image style processing framework based on deep learning according to claim 1, characterized in that: The vllm-based style classification includes the following steps: S1. Use ResNet to extract image features. The image is input into the network and processed through multiple convolutional layers and pooling layers, and finally the feature representation is output in the fully connected layer; S2. When the feature vector is obtained, the dimension is too high. PCA dimensionality reduction technology is used to reduce the computational complexity and remove redundant information. The extracted high-dimensional image features are then input into the decision tree for category prediction. S3. Use cross-validation evaluation methods to verify the performance of the model, and adjust the parameters and strategies of the classification model based on the results to optimize the prediction effect; S4. Adopt a hierarchical classification strategy, first use the large category model for coarse classification, and then use the sub-model for detailed classification.

5. The image style processing framework based on deep learning according to claim 4, characterized in that: In the S4 process of the vllm-based style classification step: The coarse classification of the large category model is to extract high-dimensional features through a pre-trained large convolutional neural network, and input it into a relatively simple multi-layer perceptron classifier for coarse classification.

6. The image style processing framework based on deep learning according to claim 4, characterized in that: In the S4 process of the vllm-based style classification step: The sub-model performs fine classification by using the style dataset to fine-tune the pre-trained model, adding style-related labels to improve classification accuracy; and using multiple GPUs or the distributed reasoning framework Triton to improve the classification speed in concurrent situations.

7. The image style processing framework based on deep learning according to claim 1, characterized in that: The perceptual loss is defined as follows: Where: G is the generated image and T is the target image.

8. The image style processing framework based on deep learning according to claim 1, characterized in that: The adversarial loss is defined as follows:

9. The image style processing framework based on deep learning according to claim 1, characterized in that: The style matching loss is defined as follows:

10. The image style processing framework based on deep learning according to claim 1, characterized in that: The frequency domain texture matching loss is as follows:

Citation Information

Cited By

  • Training-independent stylized abstraction method and device based on VLLM scaling and cross-domain rectification current inversion during reasoning

    CN121544452A

  • A training-independent stylized abstraction method and apparatus based on inference-time VLLM scaling and cross-domain rectified flow inversion.

    CN121544452B