A style transfer image processing method and device, electronic equipment and storage medium

By acquiring style reference images and text descriptions, extracting style features using preset image size processing and a preset style-aware encoder, and combining these with a preset image generation model to generate target style images, the problem of coupling style and content information is solved, achieving high-quality style transfer results.

CN118447262BActive Publication Date: 2026-02-06SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410535954.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2026-02-06
Estimated Expiration
2044-04-30

AI Technical Summary

Technical Problem

Existing style transfer techniques tend to couple style and content information, making it difficult to effectively capture and express high-level style features, resulting in generated images containing unnecessary reference image content.

Method used

By acquiring style reference images and text descriptions, using preset image size processing and a preset style-aware encoder to extract style feature information, and processing based on a preset image generation model, a target style image is generated, reducing the coupling between style and content information.

Benefits of technology

It enables the generation of high-quality stylized images based on user text information, reduces the introduction of reference image content into the target image, and improves the accuracy and quality of style transfer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118447262B_ABST
    Figure CN118447262B_ABST
Patent Text Reader

Abstract

The application discloses a style transfer image processing method and device, electronic equipment and storage medium, and is applied to the field of image processing, wherein the method comprises the following steps: acquiring a style reference picture and text description information, and processing the style reference picture into a to-be-processed picture according to a preset image size; extracting style feature information of the to-be-processed picture according to a preset style perception encoder; and processing the style feature information and the text description information based on a preset picture generation model to generate a target style picture. The embodiment of the application can realize the migration of the picture style, generate a picture with a corresponding style of the reference picture according to the user's text information, can reduce the coupling degree of the style and the picture content information, and can introduce the content of the reference picture into the target picture as little as possible in the generation process, thereby providing a high-quality stylized image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and in particular to a style transfer image processing method and device, electronic equipment and storage medium. BACKGROUND

[0002] Style transfer is an image processing technique that aims to apply the specific visual style of one reference image to another image, thereby creating a new image with a specific style. Style transfer technology is widely used in various fields, including but not limited to camera filters, artistic creation, image editing, and visual effects. In recent years, large text-to-image models based on diffusion models have demonstrated excellent image generation capabilities, gradually leading the development of style transfer technology. Currently, common style transfer technology mainly relies on a contrastive language-image pre-training (CLIP) model as a style encoder, but extracting uniform semantic features through the CLIP model can easily lead to coupling of style and content information, resulting in content leakage during style transfer and making it difficult to effectively capture and express high-level style features. To address this issue, there is an urgent need for a style transfer image processing method that can effectively identify and extract style features from a reference image while minimizing unnecessary content information, thereby ensuring that the generated image only reflects the desired artistic style and does not include the content of the reference image. SUMMARY

[0003] The present application provides a style transfer image processing method, device, electronic equipment and storage medium to realize the transfer of picture style, generate a picture with the corresponding style of the reference picture according to the user's text information, reduce the coupling degree of style and picture content information, and minimize the introduction of the content of the reference picture into the target picture during the generation process, thereby providing high-quality stylized images.

[0004] According to an aspect of the present application, a style transfer image processing method is provided, wherein the method comprises:

[0005] Obtaining a style reference picture and text description information, and processing the style reference picture into a to-be-processed picture according to a preset image size;

[0006] Extracting style feature information of the to-be-processed picture according to a preset style perception encoder;

[0007] Processing the style feature information and the text description information based on a preset picture generation model to generate a target style picture.

[0008] According to another aspect of the present application, a style transfer image processing device is provided, wherein the device comprises:

[0009] a preprocessing module, configured to acquire a style reference picture and text description information, and process the style reference picture into a to-be-processed picture according to a preset image size;

[0010] a style extraction module, configured to extract style feature information of the to-be-processed picture according to a preset style perception encoder;

[0011] a picture generation module, configured to process the style feature information and the text description information based on a preset picture generation model to generate a target style picture.

[0012] According to another aspect of the present application, an electronic device is provided, which comprises:

[0013] at least one processor; and a memory connected to the at least one processor in communication; wherein

[0014] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the style transfer image processing method according to any one of the embodiments of the present application.

[0015] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to implement the style transfer image processing method according to any one of the embodiments of the present application when executed.

[0016] The technical solution of the embodiments of the present application realizes the migration of picture style, generates a picture of a reference picture corresponding style according to user text information, can reduce the coupling degree of style and picture content information, and can introduce the content of the reference picture into the target picture as little as possible in the generation process, thereby providing high-quality stylized images.

[0017] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description.

[0018] BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to make the technical solutions in the embodiments of the present application clearer, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.

[0020] Figure 1 is a flow chart of a style transfer image processing method according to an embodiment of the present application;

[0021] Figure 2 is a flow chart of another style transfer image processing method according to an embodiment of the present application;

[0022] Figure 3 is a flow chart of another style transfer image processing method according to an embodiment of the present application;

[0023] Figure 4 is a residual block structure diagram according to an embodiment of the present application;

[0024] Figure 5 is a structure diagram of a Transformer block according to an embodiment of the present application;

[0025] Figure 6 is a flow chart of another style transfer image processing method according to an embodiment of the present application;

[0026] Figure 7 is an example diagram of a style transfer image processing method according to an embodiment of the present application;

[0027] Figure 8 is a structure diagram of a style transfer image processing device according to an embodiment of the present application;

[0028] Figure 9 is a structure diagram of an electronic device implementing a style transfer image processing method according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to make the technical solutions in the embodiments of the present application clearer, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.

[0030] It should be noted that the terms "first", "second", and the like in the description and claims of the application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in other than the order illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a list of steps or units does not necessarily limit those steps or units to the clearly listed ones, but can include other steps or units not clearly listed or inherent to such processes, methods, products or devices.

[0031] Embodiment one

[0032] Figure 1 A flowchart of a style transfer image processing method is provided for the first embodiment of the application. The present embodiment can be applied to the case of generating a corresponding style picture according to user text information and reference pictures. The method can be executed by a style transfer image processing device, which can be realized in the form of hardware and / or software. The device can be configured in a server or a user terminal. As shown in the figure, the method comprises: Figure 1

[0033] Step 110, obtaining style reference pictures and text description information, and processing the style reference pictures into the to-be-processed pictures according to the preset image size.

[0034] The style reference pictures can be reference pictures used to generate target style pictures. The style reference pictures can provide styles for the target style pictures. The styles provided by the style reference pictures can be representative appearances of pictures in overall presentation. The styles provided by the style reference pictures can include but are not limited to classicism, romanticism, cubism, expressionism, abstract art, etc. The text description information can describe the image content of the to-be-processed pictures in the form of text. The language of the text description information can not be limited. The text description information can be collected and generated in the form of text or voice, etc. The preset image size can be a preset size for processing the style reference pictures. The preset image size can be a specific size set. The style reference pictures can be segmented according to each size in the preset image size, so as to generate multiple groups of to-be-processed pictures with different sizes.

[0035] ​In the embodiment of the present application, the style reference picture and the text description information can be obtained, which can be input by the user or pre-configured by the system. The style reference picture and the text description information can be input at the same time or not. After obtaining the style reference picture, the pre-configured preset image size can be extracted, and the style reference picture can be segmented according to the preset image size, thereby generating a group of pictures to be processed corresponding to the size. It can be understood that each picture to be processed can belong to a part of the style reference picture, and the number of the pictures to be processed generated by the preset image size segmentation can be one or more.

[0036] Step 120, extracting style feature information of the picture to be processed according to the preset style perception encoder.

[0037] The preset style perception encoder can be a pre-trained model for extracting the style in the picture to be processed. The preset style perception encoder can include at least a residual block and an attention mechanism module. The preset style perception encoder can extract features of the picture to be processed through the residual block, and focus on important parts of the picture to be processed through the attention mechanism module, thereby improving the accuracy of extracting style feature information. The style feature information can be the style feature extracted by the preset style perception encoder, and the style feature information can exist in the form of a feature vector.

[0038] In the embodiment of the present application, the pre-trained preset style perception encoder can be obtained, and the generated picture to be processed can be input into the preset style perception encoder for processing. The feature vector processed by the preset style perception encoder can be used as the style feature information.

[0039] Step 130, processing the style feature information and the text description information based on the preset picture generation model to generate a target style picture.

[0040] The preset picture generation model can be a model for diffusing relevant feature information to a noise picture based on the style feature information and the text description information to generate a target style picture. The preset picture generation model can input the style feature information and the text description information as model input, and the target style picture can be output by the preset picture generation model.

[0041] In the embodiment of the present application, the picture generation model can be used to generate pictures based on style feature information and text description information, and the preset picture generation model can be used to obtain the attention output corresponding to the style feature information and the text description information, so that the target style picture can be generated in the picture iteration generation process. For example, the preset picture generation model can include a diffusion model, a multi-modal representation learning model, a cross-modal image generation model, etc. For example, the preset picture generation model can include a CLIP model, which can perform contrastive learning on the style feature information and the text description information of the picture to be processed, so as to generate a target style picture. Alternatively, the preset picture generation model can include a DALLE model, which can generate a target style picture by processing the style feature information and the text description information.

[0042] In the embodiment of the present application, the style reference picture and the text description information are obtained, and the style parameter picture is processed into pictures to be processed with different preset image sizes. The style feature information of the picture to be processed is extracted by the preset style perception encoder, and the style feature information and the text description information are processed by the preset picture generation model to generate a target style picture. The embodiment of the present application realizes the migration of picture style, generates a picture with a reference picture corresponding to the style according to user text information, reduces the coupling degree of style and picture content information, and introduces the content of the reference picture into the target picture as little as possible in the generation process, thereby providing a high-quality stylized image.

[0043] Embodiment two

[0044] Figure 2 is a flowchart of another style migration image processing method according to the second embodiment of the present application. The embodiment of the present application is a specific embodiment based on the above-mentioned application embodiment, and the preprocessing process of the style reference picture is described in detail. Referring to Figure 2 The method provided by the embodiment of the present application specifically includes the following steps:

[0045] Step 210, receiving a user-input style reference picture and text description information.

[0046] In the embodiment of the present application, the style reference picture and the text description information can be input by the user. For example, the user can upload the style reference picture and the text description information in the application software interface or the webpage interface, and the device executing the method can receive the user-uploaded style reference picture and text description information.

[0047] Step 220, processing the style reference picture into at least two groups of pictures to be processed according to preset image sizes, wherein the preset image sizes include at least two groups of picture specifications.

[0048] In the embodiment of the present application, in order to extract style information at different levels of detail, the image to be processed needs to include multiple different sizes, therefore, the preset image size can include at least two different image specifications, the style reference picture can be processed according to different image specifications in the preset image size respectively, thereby generating at least two groups of to-be-processed pictures, and the picture specifications of different groups of to-be-processed pictures can be different. For example, the style reference picture can be divided into image blocks with an original image side length of 1 / 4, 1 / 8 and 1 / 16, and these image blocks can be taken as to-be-processed pictures respectively, so as to facilitate subsequent extraction of style information at different levels of granularity.

[0049] Step 230, extracting style feature information of the to-be-processed picture according to the preset style perceptual encoder.

[0050] Step 240, processing the style feature information and the text description information based on the preset picture generation model to generate a target style picture.

[0051] In the embodiment of the present application, the style reference picture and the text description information input by the user are acquired, the style parameter picture is processed into at least two groups of to-be-processed pictures according to the picture specifications of the preset image size, the style feature information of the to-be-processed picture is extracted according to the preset style perceptual encoder, and the style feature information and the text description information are processed based on the preset picture generation model to generate a target style picture, so that the picture can be generated according to the user's text information and the reference picture, the picture style migration is realized, the decoupling of the style information and the picture content information of the reference picture is realized, the information content of the reference picture is not introduced in the process of migrating the style, and the generation quality of the stylized picture is high.

[0052] Embodiment three

[0053] Figure 3 is a flowchart of another style migration image processing method provided by the embodiment three of the present application, and the embodiment of the present application is a specific embodiment based on the above-mentioned embodiment of the application, and the extraction of the style feature information of the style reference picture is described in detail, referring to Figure 3 The method provided by the embodiment of the present application specifically includes the following steps:

[0054] Step 310, acquiring a style reference picture and text description information, and processing the style reference picture into a to-be-processed picture according to a preset image size.

[0055] Step 320, processing the to-be-processed picture of different preset image sizes by a residual block of a corresponding granularity in a preset style perceptual encoder respectively, to generate image encoding information corresponding to the preset image size.

[0056] The image coding information can be generated by a residual block processing the to-be-processed picture, and the image coding information can include feature information of the to-be-processed picture. Figure 4 The residual block can be composed of a convolution layer, a normalization layer, an activation function, and residual feedback. The number of data channels of residual blocks of different granularities can be different, and residual blocks of different granularities can process to-be-processed pictures of different image specifications.

[0057] In the embodiment of the present application, a plurality of residual blocks can be configured in the preset style perceiver, each residual block can have a different number of data channels, and for a to-be-processed picture of each picture specification, the to-be-processed picture can be input to the residual block of the corresponding granularity of the preset style perceiver, and the to-be-processed picture is processed by the residual block matching the picture specification of the to-be-processed picture, and the image coding information of the corresponding to-be-processed picture is generated by the residual block.

[0058] Step 330, the image coding information and the preset style feature information of the preset style perception encoder are combined as fusion information.

[0059] The preset style feature information can be style feature information pre-configured in the preset style perception encoder, and the preset style feature information can be generated by training an initialization parameter generated at random. The value of the preset style feature information can be determined in the training process of the preset style perception encoder. It can be understood that the preset style feature information can be an input parameter when the preset style perception encoder is trained.

[0060] In the embodiment of the present application, for each image coding information, the corresponding preset style feature information can be extracted in the preset style perception encoder, and the image coding information and the extracted preset style feature information can be fused. The fusion method can not be limited, for example, the fusion method can include but is not limited to merging the feature matrix of the image coding information and the preset style feature information as the fusion information, or generating the feature matrix product result of the image coding information and the preset style feature information as the fusion information, etc.

[0061] Step 340, calling the preset attention mechanism processing module of the preset style perception encoder to process the fusion information to generate style feature information, wherein the preset attention mechanism processing module at least includes a pre-trained Transformer block.

[0062] The preset attention mechanism processing module can be a neural network model based on an attention mechanism for extracting style features from the fusion information. The fusion information in the form of sequence data can be processed to focus on different positions in the fusion information, and different weights are assigned to different positions to output a weighted position vector as style feature information. The preset attention mechanism processing module can adopt an encoder-decoder architecture. The fusion information is encoded into a high-dimensional feature vector by an encoder, and the high-dimensional feature vector is decoded into style feature information by a decoder. The encoder and the decoder in the preset attention mechanism processing module can each be composed of multiple layers, each layer can include a multi-attention mechanism and a feedforward neural network model.

[0063] In the embodiment of the present application, the fusion information can be input into the preset attention mechanism processing module in the preset style perception encoder. The fusion information can be processed by the preset attention mechanism processing module to output the information from the preset attention mechanism processing module as style feature information. Specifically, the preset attention mechanism processing module can specifically adopt a pre-trained Transformer block. Referring to Figure 5 The Transformer block can include an encoder and a decoder. The encoder can include an Add&Norm normalization layer, a feedback layer, and a multi-head attention mechanism layer. The decoder can also include an Add&Norm normalization layer, a feedback layer, and a multi-head attention mechanism layer. The decoder also needs to input the previous output result into one input head of the multi-head attention mechanism during processing. It can be understood that, in addition to the Transformer block, the preset attention mechanism processing module used in the embodiment of the present application can also adopt other models, including but not limited to GPT model, BERT model, etc.

[0064] Step 350: processing the style feature information and the text description information based on the preset picture generation model to generate a target style picture.

[0065] In the embodiment of the present application, the style reference picture and the text description information are obtained, the style parameter picture is processed into a to-be-processed picture of different preset image sizes, the to-be-processed picture of different preset image sizes is input into a residual block processing of corresponding granularity to generate image coding information, the image coding information and the corresponding preset style feature information are jointly used as fusion information, the fusion information is input into a pre-trained Transformer block for processing to generate style feature information, and the style feature information and the text description information are processed according to a preset picture generation model to generate a target style picture. The embodiment of the present application realizes the migration of the picture style, generates a picture of a reference picture corresponding style according to user text information, can reduce the coupling degree of the style and the picture content information, and can introduce the content of the reference picture into the target picture as little as possible in the generation process, thereby providing a high-quality stylized image.

[0066] On the basis of the above-mentioned embodiment of the application, the preset style feature information at least includes input parameter information when the preset style perceptual encoder is trained.

[0067] In the embodiment of the present application, the preset style feature information can be configured in the preset style perceptual encoder. The preset style feature information can be parameter information with a general style. The preset style feature information can be generated in the training process of the preset style perceptual encoder. The preset style feature information can be input parameter information when the preset style perceptual encoder is trained. It can be understood that the initial state of the input parameter information can be randomly generated parameters. The input parameter information can be continuously adjusted and updated in the training process of the preset style perceptual encoder. When the preset style perceptual encoder is finally trained, the input parameter information at this time can be used as the preset style feature information of the preset style perceptual encoder.

[0068] Embodiment four

[0069] Figure 6 is a flowchart of another style migration image processing method provided according to the fourth embodiment of the present application. The embodiment of the present application is a specific embodiment on the basis of the above-mentioned embodiment of the application. The generation process of the target style picture is described in detail. Referring to Figure 6 , the style migration image processing method provided by the embodiment of the present application specifically includes the following steps:

[0070] In step 410, the style reference picture and the text description information are obtained, and the style reference picture is processed into a to-be-processed picture according to a preset image size.

[0071] In step 420, the style feature information of the to-be-processed picture is extracted according to the preset style perceptual encoder.

[0072] Step 430, obtain a preset vector mapping function, and call the preset vector mapping function to determine a style feature query vector, a style feature key vector and a style feature value vector of the style feature information, and determine a text description query vector, a text description key vector and a text description value vector of the text description information.

[0073] The preset vector mapping function can be a function of mapping the style feature information or the text description information, for example, the style feature query vector, the style feature key vector and the style feature value vector can be generated by mapping the style feature information through the preset vector mapping function, and the style feature query vector, the style feature key vector and the style feature value vector can correspond to the same or different mapping functions in the preset vector mapping function.

[0074] In the embodiment of the application, the preset vector mapping function for processing the style feature information or the text description information can be obtained in advance, and the style feature information or the text description information can be mapped through various corresponding preset vector mapping functions respectively, so as to generate the style feature query vector, the style feature key vector and the style feature value vector of the style feature information, and the text description query vector, the text description key vector and the text description value vector of the text description information.

[0075] Step 440, merging the text description query vector, the text description key vector and the text description value vector of the text description information and the style feature query vector, the style feature key vector and the style feature value vector of the style feature information into cross-attention output based on a parallel attention module of a preset picture generation model.

[0076] The parallel attention module can include two parallel attention mechanism processing units, and each attention mechanism processing unit can perform attention mechanism processing on the text description information and the style feature information respectively.

[0077] In the embodiment of the application, the parallel attention module of the preset picture generation model can be obtained, and the text description query vector, the text description key vector and the text description value vector of the text description information and the style feature query vector, the style feature key vector and the style feature value vector of the style feature information can be input into different attention mechanism processing units of the parallel attention module for processing, so as to generate cross-attention output, which can be generated by the parallel attention module processing the information corresponding to the text description information and the style feature information.

[0078] Step 450, processing in a denoising diffusion network of the preset picture generation model according to the cross-attention output and initial noise coding of the preset picture generation model.

[0079] The denoising diffusion network can include a forward process and a reverse process. In the forward process, the denoising diffusion network can diffuse information in the image and finally become random noise. The reverse process can be the reverse operation of the forward process, and the original image can be reconstructed from the random noise. The denoising diffusion network can include, but is not limited to, a U-net network, a Unet++ network, a Deeplab V3+ network, and the like. The initial noise encoding can be a Gaussian noise picture generated by the preset picture generation module in the forward process.

[0080] In the embodiment of the present application, the cross-attention output can be injected into the initial noise encoding to set different attention weights for different positions of the initial noise encoding. The initial noise encoding is processed in the denoising diffusion network, and the initial noise encoding is reconstructed into the target style picture encoding by removing the initial noise encoding in the denoising diffusion network.

[0081] Step 460, obtaining the target style picture encoding generated after feature diffusion, and calling the decoder of the preset picture generation model to process the target style picture encoding into the target style picture.

[0082] In the embodiment of the present application, the target style picture encoding generated by the denoising diffusion network can be decoded and reconstructed by the decoder configured in the preset picture generation model to process the target style picture encoding into the target style picture.

[0083] In the embodiment of the present application, the style reference picture and the text description information are obtained, and the style parameter picture is processed into a to-be-processed picture of different preset image sizes. The style feature information of the to-be-processed picture is extracted according to the preset style perception encoder. The style feature query vector, the style feature key vector and the style feature value vector corresponding to the style feature information and the text description information, and the text description query vector, the text description key vector and the text description value vector are determined according to the preset vector mapping function. The cross-attention output is determined based on the parallel attention module based on the style feature query vector, the style feature key vector and the style feature value vector, and the text description query vector, the text description key vector and the text description value vector. The cross-attention output is injected into the initial noise encoding, and the target style picture encoding is generated in the denoising diffusion network. The decoder is called to generate the target style picture corresponding to the target style picture encoding. The embodiment of the present application realizes the migration of the picture style, generates the picture of the reference picture corresponding to the style according to the user's text information, reduces the coupling degree of the style and the picture content information, and introduces the content of the reference picture into the target picture as much as possible in the generation process, thereby providing a high-quality stylized image.

[0084] On the basis of the above-mentioned embodiment of the application, the parallel attention model based on the preset picture generation model combines the text description query vector, the text description key vector and the text description value vector of the text description information and the style feature query vector, the style feature key vector and the style feature value vector of the style feature information into cross-attention output, which comprises:

[0085] The text description query vector, the text description key vector and the text description value vector are determined by the text attention module of the parallel attention model based on a first preset formula to obtain the text attention output corresponding to the text description information.

[0086] The style feature query vector, the style feature key vector and the style feature value vector are determined by the style feature attention module of the parallel attention model based on a second preset formula to obtain the style feature attention output corresponding to the style feature information.

[0087] The sum of the text attention output and the style feature attention output is taken as the cross-attention output.

[0088] Among the first preset formula and the second preset formula, at least one is as follows:

[0089] Among them, Attention(Q, K, V) represents the attention output, Q represents the style feature query vector or the text description query vector, K represents the style feature key vector or the text description key vector, V represents the style feature value vector or the text description value vector, d k represents a preset scaling factor, T represents vector transposition, and Softmax() represents normalizing a numerical vector into a probability distribution vector.

[0090] In the embodiment of the application, the text attention output and the style attention output can be determined by the first preset formula and the second preset formula in the parallel attention module, respectively. The first preset formula or the second preset formula can be the same or different, and at least one of the first preset formula or the second preset formula can take the following form:

[0091] Among them, Attention(Q, K, V) represents the attention output, Q represents the style feature query vector or the text description query vector, K represents the style feature key vector or the text description key vector, V represents the style feature value vector or the text description value vector, d k represents a preset scaling factor, T represents vector transposition, and Softmax() represents normalizing a numerical vector into a probability distribution vector.

[0092] On the basis of the above-mentioned embodiment of the application, the cross-attention output and the initial noise coding of the preset picture generation model are subjected to feature diffusion in the denoising diffusion network of the preset picture generation model, which comprises:

[0093] extracting initial noise encoding in the preset picture generation model; injecting the cross attention output into the initial noise encoding, and inputting the initial noise encoding into a denoising diffusion network for processing, wherein the denoising diffusion network at least includes a U-net network; judging whether the intermediate image encoding output by the denoising diffusion network satisfies a preset loss function value; if not, updating a preset scaling factor of a text attention module in the parallel attention model, and regenerating the cross attention output; injecting the cross attention output into the intermediate image encoding output by the denoising diffusion network; calling the denoising diffusion network to reprocess the intermediate image encoding, and re-determining whether the intermediate image encoding output by the denoising diffusion network satisfies the preset loss function value; if yes, taking the intermediate image encoding as the target style picture encoding.

[0094] In the embodiment of the application, the initial noise encoding can be pre-configured in the preset picture generation model, and the initial noise encoding can include picture encoding of a randomly generated Gaussian noise picture or a pre-trained generated Gaussian noise picture. The cross attention output generated in the foregoing step can be injected into the initial noise encoding to give different attention weights to different positions of the initial noise encoding. The initial noise encoding is processed by the denoising diffusion network including the U-net network, and the intermediate image encoding output by the denoising diffusion network is extracted. The intermediate image encoding can be generated by the denoising diffusion network by processing the initial noise encoding. It can be determined whether the intermediate image encoding output by the denoising diffusion network satisfies the preset loss function value configured by the denoising diffusion network. If yes, the intermediate image encoding is taken as the target style picture encoding. If not, the preset scaling factor of the text attention module in the parallel attention model is updated, and the cross attention output is regenerated. The cross attention output is injected into the intermediate image encoding output by the denoising diffusion network. The denoising diffusion network is called to reprocess the intermediate image encoding, thereby generating a new intermediate image encoding. It is determined whether the new intermediate image encoding satisfies the preset loss function. The process of determining the intermediate image encoding can be repeated until the generated intermediate image encoding satisfies the preset loss function.

[0095] Embodiment five

[0096] Figure 7 is an example diagram of style transfer picture processing according to the embodiment five of the application, referring to Figure 7The style transfer image processing method provided by the embodiment of the present application can include a style-aware encoder, style feature injection, and a style-aware dataset. The style-aware encoder can extract style features of a reference picture for style transfer. The style-aware encoder first pre-processes the reference picture, divides the reference picture into image blocks with lengths of 1 / 4, 1 / 8, and 1 / 16 of the original image, so as to extract style information at different levels of details. After the pre-processing, the embodiment of the present application provides three residual blocks with different depths: a deep residual block, a middle residual block, and a shallow residual block. The three residual blocks are respectively used to extract image block embeddings of different levels of style. Specifically, for the image block with the largest size, the deep residual block is used to extract an image block embedding rich in high-level style information, and for the image block with a smaller size, the middle residual block and the shallow residual block are respectively used to capture other low-level style embeddings. After the style embeddings are obtained in the image blocks with different sizes, the style embeddings and the generated learnable style embedding fs are input into a Transformer block to integrate style features of different levels.

[0097] In the style injection process, the extracted style features are injected into the denoising network of the stable diffusion model through a parallel cross-attention module, so as to realize the transfer of any style without style fine-tuning during testing. First, the style embedding fs is projected onto the key K' and the value V' through an independent learnable mapping function. Then, the key K' and the value V' are used to calculate the cross-attention output of the style embedding with the noise feature projection Q of the latent space. Similarly, the text embedding ft can also be projected into the key K and the value V, and the attention output of the text embedding is calculated with Q. Finally, the two parallel attention outputs are combined and input into the subsequent blocks of the stable diffusion model. During training, the text cross-attention module of the stable diffusion model is kept frozen, and only the style cross-attention module is trained to adapt to the style of the reference picture. Exemplarily, the attention output can be determined by the following formula:

[0098]

[0099] wherein Attention(Q, K, V) represents the attention output, Q represents the noise feature portrait vector of the latent space, K represents the key vector, V represents the value vector, d k denotes a preset scaling factor, T represents vector transposition, and Softmax() represents normalizing a numerical vector into a probability distribution vector.

[0100] In the embodiment of the application, the style-aware encoder can be generated by training a style-aware dataset, wherein the style-aware dataset can include a dataset JourneyDB, WIKIART, and a subset of stylized images selected from LAION-Aesthetics. The style distribution in the style-aware dataset is balanced, which improves the generalization ability of the model. However, the text prompt corresponding to the image also contains style descriptions. Since the pre-trained stable diffusion model has good responsiveness to the text condition, these style descriptions may hinder the ability of the model to learn style features from the reference image. Therefore, in the embodiment of the application, all style-related descriptions in the text prompt are removed, and only content-related context is retained, decoupling style and content information into image and text modalities, and enhancing the ability of the model to learn style features from the style-aware dataset.

[0101] Embodiment six

[0102] Figure 8 is a structural schematic diagram of a style transfer image processing device provided according to Embodiment six of the application, as shown in the figure, the device comprises: Figure 8

[0103] The preprocessing module 510 is configured to obtain a style reference picture and text description information, and process the style reference picture into a to-be-processed picture according to a preset image size.

[0104] The style extraction module 520 is configured to extract style feature information of the to-be-processed picture according to a preset style-aware encoder.

[0105] The picture generation module 530 is configured to process the style feature information and the text description information based on a preset picture generation model to generate a target style picture.

[0106] In the embodiment of the application, the preprocessing module obtains a style reference picture and text description information, and processes the style parameter picture into a to-be-processed picture of different preset image sizes. The style extraction module extracts style feature information of the to-be-processed picture according to a preset style-aware encoder. The picture generation module processes the style feature information and the text description information based on a preset picture generation model to generate a target style picture. The embodiment of the application realizes the migration of the picture style, generates a picture corresponding to the style of the reference picture according to the user text information, reduces the coupling degree of the style and the picture content information, and introduces the content of the reference picture into the target picture as little as possible in the generation process, thereby providing a high-quality stylized image.

[0107] On the basis of the above-mentioned embodiments of the application, the preprocessing module 510 comprises:

[0108] ​An information receiving unit is configured to receive the style reference picture and the text description information input by a user.

[0109] A picture processing unit is configured to process the style reference picture into at least two groups of the to-be-processed pictures according to preset image sizes.

[0110] On the basis of the above-mentioned embodiments, the style extraction module 520 comprises:

[0111] A residual extraction unit is configured to process the to-be-processed pictures of different preset image sizes by residual blocks of corresponding granularities in the preset style perception encoder to generate image encoding information corresponding to the preset image sizes.

[0112] An input processing unit is configured to jointly use the image encoding information and preset style feature information of the preset style perception encoder as fusion information.

[0113] A style fusion unit is configured to use a preset attention mechanism processing module of the preset style perception encoder to process the fusion information to generate the style feature information, wherein the preset attention mechanism processing module at least comprises a pre-trained Transformer block.

[0114] On the basis of the above-mentioned embodiments, the device further comprises a preset style unit configured to randomly generate initialization parameter information as the preset style feature information of the preset style perception encoder.

[0115] On the basis of the above-mentioned embodiments, the picture generation module 530 comprises:

[0116] A vector mapping unit is configured to obtain a preset vector mapping function, and use the preset vector mapping function to determine a style feature query vector, a style feature key vector and a style feature value vector of the style feature information, and to determine a text description query vector, a text description key vector and a text description value vector of the text description information.

[0117] An attention unit is configured to use a parallel attention model of the preset picture generation model to merge the text description query vector, the text description key vector and the text description value vector of the text description information, and the style feature query vector, the style feature key vector and the style feature value vector of the style feature information, to obtain a cross-attention output.

[0118] A denoising processing unit is configured to use the cross-attention output and initial noise encoding of the preset picture generation model to process in a denoising diffusion network of the preset picture generation model.

[0119] An image reconstruction unit is configured to acquire the target style picture code generated by the processing, and call a decoder of the preset picture generation model to process the target style picture code into the target style picture.

[0120] On the basis of the above-mentioned embodiments, the attention unit is specifically configured to determine, based on a first preset formula, a text attention output corresponding to the text description information by the text description query vector, the text description key vector and the text description value vector in the text attention module of the parallel attention model;

[0121] determine, based on a second preset formula, a style feature attention output corresponding to the style feature information by the style feature query vector, the style feature key vector and the style feature value vector in the style feature attention module of the parallel attention model;

[0122] sum the text attention output and the style feature attention output as the cross-attention output;

[0123] wherein at least one of the first preset formula and the second preset formula is:

[0124] wherein Attention(Q, K, V) represents the attention output, Q represents the style feature query vector or the text description query vector, K represents the style feature key vector or the text description key vector, V represents the style feature value vector or the text description value vector, d k represents a preset scaling factor, T represents vector transposition, and Softmax() represents normalizing a numerical vector into a probability distribution vector.

[0125] In some other embodiments, the denoising processing unit is specifically configured to extract the initial noise code in the preset picture generation model;

[0126] inject the cross-attention output into the initial noise code, and input the initial noise code into the denoising diffusion network for processing, wherein the denoising diffusion network at least includes a U-net network;

[0127] determine whether the intermediate image code output by the denoising diffusion network satisfies a preset loss function value;

[0128] If not, update the preset scaling factor of the text attention module in the parallel attention model, and regenerate the cross-attention output; inject the cross-attention output into the intermediate image encoding output by the denoising diffusion network; call the denoising diffusion network to reprocess the intermediate image encoding, and re-determine whether the intermediate image encoding output by the denoising diffusion network satisfies the preset loss function value;

[0129] If yes, the intermediate image encoding is taken as the target style picture encoding.

[0130] The style transfer image processing device provided by the embodiment of the present application can execute the style transfer image processing method provided by any embodiment of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0131] Embodiment seven

[0132] Figure 9 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit implementations of the present application described and / or claimed in this document.

[0133] As shown in Figure 9 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11, wherein the memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0134] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0135] The processor 11 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the style transfer image processing method.

[0136] In some embodiments, the style transfer image processing method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the style transfer image processing method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the style transfer image processing method by any other appropriate means, such as by means of firmware.

[0137] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0138] Computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program, when executed, can cause instructions defined in the flow charts and / or block diagrams to be implemented. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, and partially on a remote machine or entirely on a remote machine or server.

[0139] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0140] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0141] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0142] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0143] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the technical solutions of the present disclosure are achieved, and the present disclosure is not limited herein.

[0144] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A style transfer image processing method, characterized in that, The method comprises: obtaining a style reference picture and text description information, and processing the style reference picture into a to-be-processed picture according to a preset image size; extracting style feature information of the to-be-processed picture according to a preset style perception encoder, comprising: processing the to-be-processed picture of different preset image sizes by residual blocks of corresponding granularity in the preset style perception encoder respectively to generate image encoding information corresponding to the preset image size; jointly taking each image encoding information and preset style feature information of the preset style perception encoder as fusion information; calling a preset attention mechanism processing module of the preset style perception encoder to process the fusion information to generate the style feature information, wherein the preset attention mechanism processing module at least comprises a pre-trained Transformer block; processing the style feature information and the text description information based on a preset picture generation model to generate a target style picture.

2. The method of claim 1, wherein, The obtaining of the style reference picture and the text description information, and the processing of the style reference picture into the to-be-processed picture according to the preset image size comprises: receiving the style reference picture and the text description information input by a user; processing the style reference picture into at least two groups of to-be-processed pictures according to the preset image size, wherein the preset image size comprises at least two groups of picture specifications.

3. The method of claim 1, wherein, The preset style feature information at least comprises input parameter information when the preset style perception encoder is trained.

4. The method of claim 1, wherein, The processing of the style feature information and the text description information based on the preset picture generation model to generate the target style picture comprises: obtaining a preset vector mapping function, and calling the preset vector mapping function to determine a style feature query vector, a style feature key vector and a style feature value vector of the style feature information, and to determine a text description query vector, a text description key vector and a text description value vector of the text description information; merging the text description query vector, the text description key vector and the text description value vector of the text description information and the style feature query vector, the style feature key vector and the style feature value vector of the style feature information into cross-attention output based on a parallel attention model of the preset picture generation model; processing according to the cross-attention output and initial noise coding of the preset picture generation model in a denoising diffusion network of the preset picture generation model; obtaining target style picture coding generated after processing, and calling a decoder of the preset picture generation model to process the target style picture coding into the target style picture.

5. The method of claim 4, wherein, The merging of the text description query vector, the text description key vector and the text description value vector of the text description information and the style feature query vector, the style feature key vector and the style feature value vector of the style feature information into cross-attention output based on the parallel attention model of the preset picture generation model comprises: The text description query vector, the text description key vector, and the text description value vector are input into a text attention module of the parallel attention model to determine a text attention output corresponding to the text description information based on a first preset formula; The style feature query vector, the style feature key vector, and the style feature value vector are input into a style feature attention module of the parallel attention model to determine a style feature attention output corresponding to the style feature information based on a second preset formula; The sum of the text attention output and the style feature attention output is taken as the cross-attention output. At least one of the first preset formula and the second preset formula is as follows: wherein, denotes attention output, Q denotes the style feature query vector or the text description query vector, K denotes the style feature key vector or the text description key vector, V denotes the style feature value vector or the text description value vector, denotes a preset scaling factor, T denotes vector transposition, and Softmax() denotes normalizing a numerical vector into a probability distribution vector.

6. The method of claim 4, wherein, The initial noise encoding is extracted from the preset picture generation model; The cross-attention output is injected into the initial noise encoding, and the initial noise encoding is input into the denoising diffusion network for processing, wherein the denoising diffusion network at least includes a U-net network; It is determined whether the intermediate image encoding output by the denoising diffusion network satisfies a preset loss function value; If not, a preset scaling factor of a text attention module in the parallel attention model is updated, and the cross-attention output is regenerated; the cross-attention output is injected into the intermediate image encoding output by the denoising diffusion network; the denoising diffusion network is called to reprocess the intermediate image encoding, and it is determined again whether the intermediate image encoding output by the denoising diffusion network satisfies the preset loss function value; If yes, the intermediate image encoding is taken as the target style picture encoding. The device comprises:

7. A style transfer image processing apparatus, characterized by comprising: A preprocessing module configured to obtain a style reference picture and text description information, and process the style reference picture into a to-be-processed picture according to a preset image size; A style extraction module configured to extract style feature information of the to-be-processed picture according to a preset style perception encoder; A picture generation module configured to process the style feature information and the text description information based on a preset picture generation model to generate a target style picture. The style extraction module comprises: A residual extraction unit configured to process the to-be-processed picture of different preset image sizes by residual blocks of corresponding granularity in the preset style perception encoder to generate image encoding information corresponding to the preset image sizes; An input processing unit configured to jointly take the image encoding information and preset style feature information of the preset style perception encoder as fusion information; A style fusion unit configured to process the fusion information by a preset attention mechanism processing module of the preset style perception encoder to generate the style feature information, wherein the preset attention mechanism processing module at least includes a pre-trained Transformer block. The electronic device comprises:

8. An electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein ​ The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the style transfer image processing method in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the style transfer image processing method in any one of claims 1-6 when executed.

Citation Information

Patent Citations

  • Stylized image generation method and device, computer equipment and storage medium

    CN116012488A

  • Image generation method and device, electronic equipment, storage medium and program product

    CN116958323A