Image animation style migration method and system based on Puff-Net

By introducing CBAM attention mechanism and edge loss function in Puff-Net, the problems of edge clarity and content structure retention in image anime style transfer are solved, and a higher quality and efficient style transfer effect is achieved.

CN120198276APending Publication Date: 2025-06-24NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510345320.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing Puff-Net method has potential optimization spaces in feature extraction and style transfer in image anime style transfer, especially in terms of clarity at the edge of the image and the retention of content structure.

Method used

By introducing CBAM attention mechanism and edge loss function based on Sobel operator, the Puff-Net network architecture is improved, the content and style characteristics are enhanced, and the quality and efficiency of image animation style migration is improved.

Benefits of technology

The generated anime-style images are clearer and more natural, and the content structure is well preserved, which avoids the problems of blurred and distorted image details and improves the efficiency and robustness of style transfer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198276A_ABST
    Figure CN120198276A_ABST
Patent Text Reader

Abstract

The invention discloses a Puff-Net-based image animation style migration method and system, and relates to the technical field of computer vision, and the method comprises the steps: obtaining an image data set, carrying out the preprocessing of the image data set, obtaining a processed image data set, and storing the processed image data set in a database; the processed image data set is input into a pre-established Puff-Net-based image animation style migration model for training, a trained Puff-Net-based image animation style migration model is output, and the image data set comprises an animation style image and a real photographing image; a real image to be migrated is received, the real image to be migrated is input into the trained image animation style migration model based on the Puff-Net, a corresponding animation style migration image is obtained through output, the image animation style migration model based on the Puff-Net comprises two feature extractors and a converter only comprising an encoder, and the two feature extractors are connected with the converter. The animation style image meeting the artistic style requirement can be effectively generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and specifically, to an image anime style transfer method and system based on Puff-Net. Background Art

[0002] Image style transfer is an emerging technology in the field of deep learning. Since the concept of image style is very abstract, and the computer only processes some pixel points during the image processing, and cannot distinguish different styles like humans, people expect to solve this problem by extracting the style features of the image.

[0003] Anime is a very popular art form now, and has a wide range of applications in multiple aspects such as film and television works, advertisements, etc. However, the production of comics usually requires professionals to complete, and it takes a lot of time and manpower. With the rapid development of deep learning, using generative adversarial networks to learn the abstract style features of images for style transfer has achieved good results, but there are problems with the difficulty of convergence in training GAN models, especially in terms of detail retention and style consistency. These methods often ignore the balance between the content structure and style features, resulting in the generated images not being precise enough in expressing the artistic style while retaining the content, or the style conversion being too rigid.

[0004] In recent years, with the continuous progress of deep learning in the field of image generation, Puff-Net has been proposed as an innovative image style transfer method. Puff-Net realizes the effective decoupling of content and style by combining an efficient Transformer encoder and a lightweight feature extraction network, and uses multi-level feature fusion to generate higher-quality stylized images. Puff-Net significantly improves the quality and efficiency of image style transfer by separately extracting and finely synthesizing content features and style features.

[0005] Although Puff-Net has achieved good results in the style transfer task, there are still some potential optimization spaces in the existing Puff-Net methods during the feature extraction and style transfer processes, especially in terms of the clarity of the image edges and the retention of the content structure. Summary of the Invention

[0006] To solve the deficiencies mentioned in the above background art, the purpose of the present invention is to provide an image anime style transfer method and system based on Puff-Net, which can effectively generate anime style images that meet the requirements of artistic styles.

[0007] In a first aspect, the purpose of the present invention can be achieved by the following technical solutions: An image anime style transfer method based on Puff-Net, the method comprising the following steps:

[0008] Obtain an image dataset, preprocess the image dataset to obtain a processed image dataset, and input the processed image dataset into a pre-established Puff-Net-based image anime style transfer model for training to output a trained Puff-Net-based image anime style transfer model. Among them, the image dataset includes anime style images and real-life photographed images;

[0009] Receive a real image to be transferred, input the real image to be transferred into the trained Puff-Net-based image anime style transfer model, and output a corresponding anime style transferred image. Among them, the Puff-Net-based image anime style transfer model includes two feature extractors and a converter that only contains an encoder.

[0010] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the preprocessing of the image dataset includes resizing the picture, tensor conversion, and batch dimension increase.

[0011] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the process of resizing the picture includes:

[0012] resize = transforms.Resize(H, W)

[0013] img2 = resize(img)

[0014] Among them, transforms.Resize represents a function for resizing the picture, temporarily stored in resize, H and W respectively represent the width and height of the adjusted image, img represents the input image, and img2 represents the output image after resizing;

[0015] Tensor conversion includes:

[0016] to_tensor = transforms.ToTensor()

[0017] img_tensor = to_tensor(img2)

[0018] Among them, transforms.ToTensor represents an operation for converting an image into a tensor, temporarily stored in to_tensor, img2 represents the resized image, and img_tensor represents the image after being converted into a tensor;

[0019] Batch dimension increase includes:

[0020] img_batch = img_tensor.unsqueeze(0)

[0021] Among them, unsqueeze(0) means adding a batch dimension to the first dimension of the image, converting a single image into a tensor containing one batch. img_tensor is the image converted to the tensor format, and img_batch is the image after adding the batch dimension.

[0022] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: during the process of inputting the image dataset into the pre-established image anime style transfer model based on Puff-Net for training, an optimized loss function design is adopted, including: adding a pixel-level L1 loss function and an edge loss function based on the Sobel operator.

[0023] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the two feature extractors and a transformer containing only an encoder are specifically:

[0024] The content feature extractor is constructed based on a reversible neural network, and efficiently extracts content features through multiple layers of reversible layers, and introduces the CBAM attention mechanism after the preprocessing stage to obtain content features;

[0025] The style feature extractor is constructed based on lightweight Transformer blocks, and introduces the CBAM attention mechanism after the preprocessing stage to obtain style features;

[0026] The transformer containing only an encoder is used to fuse the content features and style features, and encodes the content features and style features through a multi-head self-attention mechanism and a feed-forward network to generate stylized image features.

[0027] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the process of introducing the CBAM attention mechanism by the content feature extractor and the style feature extractor after the preprocessing stage includes:

[0028] The CBAM attention mechanism performs weighted processing on the input features. The CBAM attention mechanism is composed of a channel attention mechanism and a spatial attention mechanism respectively;

[0029] The channel attention module performs global average pooling and max pooling operations on the input feature map along the channel dimension. After the pooling results obtained are subjected to feature mapping through two layers of convolutional networks, the channel-level attention weights are obtained through the Sigmoid activation function. The process is as follows:

[0030]

[0031] F max (X1) = max h,w = X1(c, h, w)

[0032] SE avg (X1) = W2 × Relu × (W1 × F avg (X1) + b1) + b2

[0033] SE max (X1) = W2 × Relu × (W1 × F max (X1) + b1) + b2

[0034] X ′ 1 = σ(SE max (X1) + SE max (X1))

[0035] where X1 ∈ R B×C×H×W represents the input feature map, B is the batch size, C represents the number of channels of the feature map, H represents the height of the feature map, W represents the width of the feature map, F avg and F max are respectively the outputs after global average pooling operation and the outputs after max pooling operation, SE avg and SE max are respectively the results after processing the results of global average pooling and max pooling operations through two convolutional layers, w1 ∈ R C / G×C , w2 ∈ R C×C / G are learnable weight matrices, G represents the number of split channels, b1, b2 are learnable biases, X ’ 1 represents the channel attention feature output by the channel attention module, σ represents the Sigmoid activation function;

[0036] The spatial attention module is complementary to the channel attention. It performs spatial domain operations on the feature map X ′ 1. By performing max pooling and average pooling on the input feature map, the importance of each pixel point is calculated, and then the spatial attention weight is generated through a convolutional operation. The process includes:

[0037] Concat(x2) = [F avg (X2), F max (X2)]

[0038] X ′ 2 = σ × (W × Concat(X2)) + b

[0039] where X2 ∈ R B×C×H×W represents X ‘ 1 output by the channel attention module, Concat(x2) ∈ RB×2×H×W denotes concatenating the results after global average pooling and max pooling operations along the channel dimension, X ′ 2 represents the spatial attention features output by the spatial attention module, W is a learnable weight matrix, b is a learnable bias, and σ represents the Sigmoid activation function.

[0040] Combined with the first aspect, in some implementations of the first aspect, the method further includes: The newly added pixel-level L1 loss function and the edge loss function based on the Sobel operator are as follows:

[0041] The newly added pixel-level L1 loss function calculates by comparing the generated image and the content image at the pixel level, and the calculation process is as follows:

[0042]

[0043] where y ′ is the generated stylized image, y is the original content image, serving as the target image, N is the total number of pixels in the image, y ′ i , y i are respectively the i-th pixel values of the generated stylized image and the original content image;

[0044] In the edge loss function based on the Sobel operator, the Sobel operator is used. The Sobel operator calculates the gradients in the horizontal and vertical directions by performing a convolution operation on the input image. The process of using the Sobel operator to extract the edge information of the image includes:

[0045]

[0046] where Sobel x , Sobel y are respectively used to calculate the gradients of the edges in the horizontal and vertical directions, edge(*) is the edge feature of the image, Edge loss is the edge loss between the generated stylized image and the original content image, and N is the total number of pixels in the image.

[0047] Combined with the first aspect, in some implementations of the first aspect, the method further includes: The total loss function used during the process of inputting the processed image dataset into the pre-established image anime style transfer model based on Puff-Net for training is:

[0048] L = λ c L c + λ s L s + λ fe L fe + λ id1L id1 + λ id2 L id2 + λ edge L edge + λ L1 L L1

[0049]

[0050] L fe = L cc + L cs + L sc + L ss

[0051] L id1 = ‖I cc - I c ‖² + ‖I ss - I s ‖²

[0052]

[0053] Among them, L represents the total loss function, L C represents the content loss, L S represents the style loss, L fe represents the feature extractor loss, L id1 and L id2 represent the identity loss, L edge represents the edge loss, L L1 represents the L1 loss, λ c , λ s , λ fe , λ id1 , λ id2 , λ edge , λ L1 are the weight coefficients of each loss respectively; I O represents the generated image, I C represents the content image, I S represents the style image, ψ i (*) represents the feature extraction function of the i-th layer in the pre-trained VGG network, N1 represents the number of layers of feature extraction, μ(*) and σ(*) represent the mean and variance of the feature map respectively; L CC and L CS represent the content loss between the input and output of the content feature extractor for the content image and the style image respectively, L SC and L SS represent the content loss between the input and output of the style feature extractor for the content image and the style image respectively; I CC represents the image reconstructed from the content features extracted from the content image, I SSAn image reconstructed from the style features extracted from the style image.

[0054] In a second aspect, to achieve the above object, the present invention discloses an image anime style transfer system based on Puff-Net, including:

[0055] A model training module, configured to obtain an image data set, preprocess the image data set to obtain a processed image data set, and input the processed image data set into a pre-established image anime style transfer model based on Puff-Net for training, and output a trained image anime style transfer model based on Puff-Net, wherein the image data set includes anime style images and real-life photographed images;

[0056] A style transfer module, configured to receive a real image to be transferred, input the real image to be transferred into the trained image anime style transfer model based on Puff-Net, and output a corresponding anime style transferred image, wherein the image anime style transfer model based on Puff-Net includes two feature extractors and a converter including only an encoder.

[0057] In another aspect of the present invention, to achieve the above object, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The computer program stored in the memory is such that when the processor loads and executes the computer program, a method for transferring an image anime style based on Puff-Net as described above is adopted.

[0058] Advantages of the present invention:

[0059] By introducing the CBAM attention mechanism and the edge loss function based on the Sobel operator, the present invention improves the Puff-Net network architecture, enabling the content and style features in the image anime style transfer process to be more precisely fused, thereby generating clearer and more natural anime style images. Compared with traditional style transfer methods, the present invention effectively improves the retention of the content structure and the edge sharpness of the generated images, and can avoid problems such as blurred and distorted image details. The image style transfer method of the present invention has high efficiency and robustness, is applicable to various image style transfer tasks, especially in the field of anime style conversion, and can achieve a natural and harmonious style conversion while retaining content details. Description of the Drawings

[0060] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings;

[0061] Figure 1 It is a schematic flowchart of the method of the present invention;

[0062] Figure 2 It is a schematic workflow diagram of the present invention;

[0063] Figure 3 It is the network structure diagram of the Puff-Net model in the embodiment of the present invention;

[0064] Figure 4 It is the flowchart of introducing the CBAM attention mechanism in the feature extractor in the embodiment of the present invention;

[0065] Figure 5 It is a schematic diagram of the system structure of the present invention. Specific embodiments

[0066] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0067] Embodiment 1:

[0068] As Figure 1 shown, an image anime style transfer method based on Puff-Net, the method includes the following steps:

[0069] S101: Obtain an image dataset, preprocess the image dataset to obtain a processed image dataset, and input the processed image dataset into a pre-established image anime style transfer model based on Puff-Net for training, and output a trained image anime style transfer model based on Puff-Net, wherein the image dataset includes anime style images and real-life photographed images;

[0070] The preprocessing of the image dataset includes adjusting the picture size, tensor conversion, and batch dimension increase;

[0071] The process of adjusting the picture size includes:

[0072] resize = transforms.Resize(H,W)

[0073] img2 = resize(img)

[0074] Among them, transforms.Resize represents a function for resizing images, temporarily stored in resize, H and W respectively represent the width and height of the resized image, img represents the input image, and img2 represents the output image after resizing;

[0075] Tensor conversion includes:

[0076] to_tensor = transforms.ToTensor()

[0077] img_tensor = to_tensor(img2)

[0078] Among them, transforms.ToTensor represents an operation for converting an image to a tensor, temporarily stored in to_tensor, img2 represents the resized image, and img_tensor represents the image after being converted to a tensor;

[0079] Batch dimension addition includes:

[0080] img_batch = img_tensor.unsqueeze(0)

[0081] Among them, unsqueeze(0) means adding a batch dimension to the first dimension of the image, converting a single image into a tensor containing one batch, img_tensor is the image converted to the tensor format, and img_batch is the image after adding the batch dimension.

[0082] During the process of inputting the image dataset into a pre-established image anime style transfer model based on Puff-Net for training, an optimized loss function design is adopted, which includes: adding a pixel-level L1 loss function and an edge loss function based on the Sobel operator.

[0083] The added pixel-level L1 loss function and the edge loss function based on the Sobel operator are as follows:

[0084] The added pixel-level L1 loss function is used to compare the generated image and the content image at the pixel level, and the calculation process is as follows:

[0085]

[0086] where y ′x is the generated stylized image, y is the original content image as the target image, N is the total number of pixels in the image, y ′ i , y i are the i-th pixel values of the stylized generated image and the original content image respectively;

[0087] The Sobel operator is used in the edge loss function based on the Sobel operator. The Sobel operator calculates the gradients in the horizontal and vertical directions by performing a convolution operation on the input image. The process of using the Sobel operator to extract the edge information of the image includes:

[0088]

[0089] where Sobel x , Sobel y are respectively used to calculate the gradients of the edges in the horizontal and vertical directions, edge(*) is the edge feature of the image, Edge loss is the edge loss between the stylized generated image and the original content image, and N is the total number of pixels in the image.

[0090] S102: Receive the real image to be migrated, input the real image to be migrated into the trained image anime style migration model based on Puff-Net, and output the corresponding anime style migration image. Among them, the image anime style migration model based on Puff-Net includes two feature extractors and a transformer that only contains an encoder.

[0091] As Figure 4 shown, the two feature extractors and a transformer that only contains an encoder are specifically:

[0092] Content feature extractor, constructed based on a reversible neural network, realizes the efficient extraction of content features through multiple reversible layers, and introduces the CBAM attention mechanism after the preprocessing stage to obtain content features;

[0093] Style feature extractor, constructed based on lightweight Transformer blocks, introduces the CBAM attention mechanism after the preprocessing stage to obtain style features;

[0094] Transformer that only contains an encoder, used to fuse the content features and style features, encodes the content features and style features through the multi-head self-attention mechanism and the feed-forward network, and generates the stylized image features.

[0095] The process of introducing the CBAM attention mechanism by the content feature extractor and the style feature extractor after the preprocessing stage includes:

[0096] The CBAM attention mechanism performs weighted processing on the input features. The CBAM attention mechanism consists of a channel attention mechanism and a spatial attention mechanism respectively;

[0097] The channel attention module calculates the pooling results by performing global average pooling and max pooling operations on the input feature map along the channel dimension. After the feature mapping of the calculated pooling results through two convolutional networks, the channel-level attention weights are obtained through the Sigmoid activation function. The process is as follows:

[0098]

[0099] F max (X1) = max h,w = X1(c, h, w)

[0100] SE avg (X1) = W2 × Relu × (W1 × F avg (X1) + b1) + b2

[0101] SE max (X1) = W2 × Relu × (W1 × F max (X1) + b1) + b2

[0102] X ′ 1 = σ(SE max (X1) + SE max (X1))

[0103] Among them, X1 ∈ R B×C×H×W represents the input feature map, B is the batch size, C represents the number of channels of the feature map, H represents the height of the feature map, W represents the width of the feature map, F avg and F max are the outputs after the global average pooling operation and the max pooling operation respectively, SE avg and SE max are the results after processing the results of the global average pooling and max pooling operations through two convolutional layers respectively, w1 ∈ R C / G×C , w2 ∈ R C×C / G are learnable weight matrices, G represents the number of split channels, b1, b2 are learnable biases, X ’ 1 represents the channel attention features output by the channel attention module, and σ represents the Sigmoid activation function;

[0104] The spatial attention module is complementary to the channel attention. It performs spatial domain operations on the feature map X ′ 1 after the channel attention processing. By performing max pooling and average pooling on the input feature map, the importance of each pixel point is calculated, and then the spatial attention weights are generated through convolutional operations. The process includes:

[0105] Concat(x2) = [F avg (X2), F max (X2)]

[0106] X ′ 2 = σ×(W×Concat(X2)) + b

[0107] where X2 ∈ R B×C×H×W represents X output by the channel attention module ‘ 1, and Concat(x2) ∈ R B×2×H×W represents the result after global average pooling and max pooling operations concatenated along the channel dimension, and X ′ 2 represents the spatial attention feature output by the spatial attention module, W is a learnable weight matrix, b is a learnable bias, and σ represents the Sigmoid activation function

[0108] During the process of inputting the processed image dataset into the pre-established Puff-Net-based image anime style transfer model for training, the total loss function used is:

[0109] L = λ c L c + λ s L s + λ fe L fe + λ id1 L id1 + λ id2 L id2 + λ edge L edge + λ L1 L L1

[0110]

[0111] L fe = L cc + L cs + L sc + L ss

[0112] L id1 = ‖I cc - I c ‖2 + ‖I ss - I s ‖2

[0113]

[0114] where L represents the total loss function, and L C represents the content loss, and LS Denote the style loss, L fe Denote the feature extractor loss, L id1 and L id2 Denote the identity loss, L edge Denote the edge loss, L L1 Denote the L1 loss, λ c , λ s , λ fe , λ id1 , λ id2 , λ edge , λ L1 are the weight coefficients of each loss respectively; I O Denote the generated image, I C Denote the content image, I S Denote the style image, ψ i (*) denotes the feature extraction function of the i-th layer in the pre-trained VGG network, N1 denotes the number of feature extraction layers, and μ(*) and σ(*) denote the mean and variance of the feature map respectively; L CC and L CS Denote the content loss between the content image and the style image respectively between the input and output of the content feature extractor, L SC and L SS Denote the content loss between the content image and the style image respectively between the input and output of the style feature extractor; I CC Denote the image reconstructed from the content features extracted from the content image, I SS Denote the image reconstructed from the style features extracted from the style image.

[0115] Based on the first embodiment, in order to verify its beneficial effects, experimental simulation data of the anime image style transfer method based on Puff-Net of the present invention is provided.

[0116] Data preparation: Collect the anime image datasets Shinkai, Hayao, and SummerWar containing different anime styles and the content image dataset MS-COCO containing daily scenes. Each of the 3 anime style datasets contains 1000 anime images, and the content image dataset MS-COCO contains 2000 daily life scene images.

[0117] Model construction: Use Puff-Net as the backbone network and improve the network structure: Add a CBAM attention module after the preprocessing of the content feature extractor and the style feature extractor.

[0118] Model Training: Select a style image dataset, construct a training dataset that sequentially inputs a random image from the content image dataset and a random image from the style image dataset. All input images are randomly cropped to a resolution of 256×256, and the output is an anime image training dataset.

[0119] Training Parameters: Set the learning rate to 0.0005, the training batch_size to 1, and train for 100,000 epochs.

[0120] Add an edge loss function based on the Sobel operator and a pixel-level L1 loss function to the original loss function, update the network parameters, and obtain a trained style transfer model.

[0121] Input the content image to be transferred and generate an anime image after style transfer.

[0122] To verify the optimization effect of the CBAM attention mechanism, Sobel edge loss, and L1 loss on the Puff-Net model, this paper designed two ablation experiments to compare the performance differences of different module combinations on three anime style datasets: Shinkai, Hayao, and SummerWar. The first experiment added the CBAM attention mechanism to the feature extractor part based on the original Puff-Net model, and the new model is denoted as: Puff-Net+CBAM. The second experiment also added the Sobel operator and L1 loss during the training process based on the original Puff-Net model to detect the edge loss, and the new model is denoted as: Puff-Net+Sobel+L1. This paper uses LPIPS and FID as quantitative evaluation indicators, and the best experimental results are bolded, as shown in Table 1.

[0123] Table 1 Comparison Table of Ablation Experiment Results

[0124]

[0125] Example Two: Second, as Figure 5 shown, to achieve the above purpose, the present invention discloses an image anime style transfer system based on Puff-Net, including:

[0126] A model training module 11, which is used to obtain an image dataset, preprocess the image dataset to obtain a processed image dataset, and input the processed image dataset into a pre-established image anime style transfer model based on Puff-Net for training, and output a trained image anime style transfer model based on Puff-Net, where the image dataset includes anime style images and real-life photographed images;

[0127] The style transfer module 12 is configured to receive a real image to be transferred, input the real image to be transferred into a trained image anime style transfer model based on Puff-Net, and output a corresponding anime style transferred image. Among them, the image anime style transfer model based on Puff-Net includes two feature extractors and a transformer including only an encoder.

[0128] Based on the same inventive concept, the present invention further provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions. Specifically, it is used to load and execute one or more instructions in the computer storage medium to implement the above method.

[0129] It should be further noted that, based on the same inventive concept, the present invention further provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.

[0130] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0131] The above shows and describes the basic principles, main features and advantages of the present disclosure. Those skilled in the art of this industry should understand that the present disclosure is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements all fall within the scope of the present disclosure claimed.

Claims

1. A method for image animation style transfer based on Puff-Net, characterized in that: The method comprises the following steps: Acquire an image data set, preprocess the image data set to obtain a processed image data set, input the processed image data set into a pre-established Puff-Net-based image anime style transfer model for training, and output a trained Puff-Net-based image anime style transfer model, wherein the image data set includes anime-style images and real-life photographed images; A real image to be transferred is received, the real image to be transferred is input into a trained Puff-Net-based image animation style transfer model, and a corresponding animation style transfer image is output, wherein the Puff-Net-based image animation style transfer model includes two feature extractors and a converter that only includes an encoder.

2. The method for image animation style transfer based on Puff-Net according to claim 1, characterized in that: The preprocessing of the image dataset includes adjusting the image size, tensor conversion, and increasing the batch dimension.

3. The method for image animation style transfer based on Puff-Net according to claim 2, characterized in that: The process of adjusting the image size includes: resize=transforms.Resize(H,W) img2 = resize(img) Among them, transforms.Resize represents the function of adjusting the image size, which is temporarily stored in resize, H and W represent the width and height of the adjusted image respectively, img represents the input image, and img2 represents the output image after resizing; Tensor transformations include: to_tensor=transforms.ToTensor() img_tensor = to_tensor(img2) Among them, transforms.ToTensor represents the operation of converting an image into a tensor, which is temporarily stored in to_tensor, img2 represents the resized image, and img_tensor represents the image after conversion into a tensor; Batch dimension additions include: img_batch=img_tensor.unsqueeze(0) Among them, unsqueeze(0) means adding a batch dimension to the first dimension of the image and converting a single image into a tensor containing a batch. img_tensor is the image converted to tensor format, and img_batch is the image after the batch dimension is added.

4. The method for image animation style transfer based on Puff-Net according to claim 1, characterized in that: In the process of inputting the image dataset into the pre-established Puff-Net-based image animation style transfer model for training, an optimized loss function design is adopted, including: a new pixel-level L1 loss function and an edge loss function based on the Sobel operator.

5. The method for image animation style transfer based on Puff-Net according to claim 1, characterized in that: The two feature extractors and one encoder-only converter are specifically: The content feature extractor is built based on a reversible neural network. It achieves efficient extraction of content features through multiple reversible layers and introduces the CBAM attention mechanism after the preprocessing stage to obtain content features. The style feature extractor is built based on the lightweight Transformer block and introduces the CBAM attention mechanism after the preprocessing stage to obtain style features; The converter that only contains the encoder is used to fuse the content features and style features, encode the content features and style features through the multi-head self-attention mechanism and feedforward network, and generate stylized image features.

6. The method for image animation style transfer based on Puff-Net according to claim 5, characterized in that: The process of introducing the CBAM attention mechanism into the content feature extractor and the style feature extractor after the preprocessing stage includes: The CBAM attention mechanism performs weighted processing on the input features. The CBAM attention mechanism consists of a channel attention mechanism and a spatial attention mechanism. The channel attention module performs global average pooling and maximum pooling operations on the input feature map along the channel dimension. After the calculated pooling result is feature mapped through a two-layer convolutional network, the channel-level attention weight is obtained through the Sigmoid activation function. The process is as follows: F max (X1)=max h,w =X1(c,h,w) HE avg (X1)=W2×Relu×(W1×F avg (X1)+b1)+b2 HE max (X1)=W2×Relu×(W1×F max (X1)+b1)+b2 X ′ 1=σ(IF max (X1)+IF max (X1)) Where X1∈R B×C×H×W represents the input feature map, B is the batch size, C is the number of feature map channels, H is the height of the feature map, W is the width of the feature map, and F avg and F max are the output after the global average pooling operation and the output after the maximum pooling operation, SE avg and SE max are the results of global average pooling and maximum pooling after processing through two convolutional layers, w1∈R C / G×C ,w2∈R C×C / G is a learnable weight matrix, G represents the number of split channels, b1, b2 are learnable biases, X ’ 1 represents the channel attention feature output by the channel attention module, σ represents the Sigmoid activation function; The spatial attention module is complementary to the channel attention module and processes the feature map X after channel attention. ′ 1 Perform spatial domain operations, calculate the importance of each pixel by performing maximum pooling and average pooling on the input feature map, and then generate spatial attention weights through convolution operations. The process includes: Concat(x2)=[F avg (X2),F max (X2)] X ′ 2=σ×(W×Concat(X2))+b Where X2∈R B×C×H×W X represents the output of the channel attention module ‘ 1, Concat(x2)∈R B×2×H×W Indicates that the results of global average pooling and maximum pooling operations are concatenated along the channel dimension, X ′ 2 represents the spatial attention feature output by the spatial attention module, W is the learnable weight matrix, b is the learnable bias, and σ represents the Sigmoid activation function.

7. The method for image animation style transfer based on Puff-Net according to claim 4, characterized in that: The newly added pixel-level L1 loss function and the edge loss function based on the Sobel operator are as follows: The new pixel-level L1 loss function compares the generated image with the content image at the pixel level. The calculation process is as follows: where y ′ is the generated stylized image, y is the original content image, as the target image, N is the total number of pixels in the image, y ′ i ,y i are the i-th pixel values ​​of the stylized generated image and the original content image respectively; The Sobel operator is used in the edge loss function based on the Sobel operator. The Sobel operator calculates the gradients in the horizontal and vertical directions by performing a convolution operation on the input image. The process of using the Sobel operator to extract the edge information of the image includes: Among them, Sobel x ,Sobel y They are used to calculate the gradient of the edge in the horizontal and vertical directions respectively. Edge(*) is the edge feature of the image. loss is the edge loss between the stylized generated image and the original content image, and N is the total number of pixels in the image.

8. The method for image animation style transfer based on Puff-Net according to claim 1, characterized in that: The total loss function used in the process of inputting the processed image dataset into the pre-established Puff-Net-based image animation style transfer model for training is: L=λ c L c +λ s L s +λ fe L fe +λ id1 L id1 +λ id2 L id2 +λ edge L edge +λ L1 L L1 L fe =L cc +L cs +L sc +L ss L id1 =‖I cc -I c ‖2+‖I ss -I s ‖2 Among them, L represents the total loss function, L C Indicates content loss, L S represents the style loss, L fe represents the feature extractor loss, L id1 and L id2 represents identity loss, L edge represents the marginal loss, L L1 represents L1 loss, λ c ,λ s ,λ fe ,λ id1 ,λ id2 ,λ edge ,λ L1 are the weight coefficients of each loss; I O Represents the generated image, I C Represents the content image, I S represents the style image, ψ i (*) represents the feature extraction function of the i-th layer in the pre-trained VGG network, N1 represents the number of feature extraction layers, μ(*) and σ(*) represent the mean and variance of the feature map, respectively; L CC and L CS represents the content loss between the content image and the style image, respectively, between the input and output of the content feature extractor, L SC and L SS I represents the content loss between the input and output of the style feature extractor for the content image and style image, respectively; CC represents the image reconstructed from the content features extracted from the content image, I SS represents an image reconstructed from style features extracted from the style image.

9. A Puff-Net-based image animation style transfer system, characterized in that: include: A model training module is used to obtain an image data set, preprocess the image data set to obtain a processed image data set, input the processed image data set into a pre-established Puff-Net-based image animation style transfer model for training, and output a trained Puff-Net-based image animation style transfer model, wherein the image data set includes anime-style images and real-life photographed images; The style transfer module is used to receive a real image to be transferred, input the real image to be transferred into a trained Puff-Net-based image animation style transfer model, and output a corresponding animation style transfer image, wherein the Puff-Net-based image animation style transfer model includes two feature extractors and a converter that only includes an encoder.

10. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, a Puff-Net-based image animation style transfer method according to any one of claims 1 to 8 is adopted.