Image style migration method, apparatus and device, and computer program product
By adopting an IP-Adapter-based diffusion model in the image style transfer technology, combining the prompt word prediction model and the enhanced image generation model, the problems of insufficient detail control, lack of flexibility and high computing resource consumption in the existing technology are solved, and image style transfer with higher accuracy and flexibility are achieved.
Patent Information
- Application Number
- CN202510101486.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
AI Technical Summary
The existing image style migration technology has problems such as insufficient detail control, lack of flexibility and high computing resource consumption.
Using an IP-Adapter-based diffusion model, combining the prompt word prediction model and the enhanced image generation model, image style transfer is performed through conditional embedding and unconditional embedding.
It realizes higher-precision image style migration, improves flexibility and control capabilities in the image generation process, and reduces the consumption of computing resources.
Smart Images

Figure CN119941494A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image style transfer method, device and equipment, and a computer program product. Background Art
[0002] Image style transfer technology is an image processing technology that applies the artistic style of one image to another image while retaining the content information of the latter. With the rapid development of Internet technology, image style transfer technology has received widespread attention in the field of computer vision and has been widely used in the fields of artistic creation, image enhancement, etc., thereby achieving image style conversion and bringing changes in visual effects.
[0003] In recent years, image style transfer methods based on deep learning have achieved remarkable results, among which the generative adversarial network (GAN) and neural network style transfer algorithms are the most representative.
[0004] The existing image style transfer technology has at least the following problems:
[0005] 1) Insufficient detail control: Existing style transfer methods cannot accurately control the details of the generated images, and the style and content of the generated images may deviate from the user's expectations;
[0006] 2) Lack of flexibility: Most style transfer methods rely on a single neural network architecture, which is difficult to flexibly adjust according to different style requirements;
[0007] 3) High computational resource consumption: Existing style transfer methods are inefficient when processing complex and high-resolution images, resulting in high resource consumption and slow generation speed. Summary of the invention
[0008] In order to solve at least one of the technical problems mentioned above, the embodiments of the present application provide an image style transfer method, an apparatus and equipment, and a computer program product to improve the quality and flexibility of image style transfer.
[0009] The present application embodiment adopts the following technical solutions:
[0010] In a first aspect, an embodiment of the present application provides an image style transfer method, the image style transfer method comprising:
[0011] Obtain the image to be style transferred and the style guide image;
[0012] According to the style guidance image, using a prompt word prediction model to predict prompt words, to obtain prompt words corresponding to the style guidance image;
[0013] Generating conditional embedding and unconditional embedding of a diffusion model according to the image to be style-transferred and the style-guided image, and constructing a diffusion model based on IP-Adapter according to the conditional embedding and unconditional embedding of the diffusion model;
[0014] According to the image to be style transferred, using an enhanced image generation model to generate image generation enhancement information corresponding to the image to be style transferred;
[0015] According to the prompt words corresponding to the style guiding image and the image enhancement information corresponding to the image to be style transferred, the style transfer image is generated by using the diffusion model based on IP-Adapter.
[0016] Optionally, the predicting a prompt word using a prompt word prediction model according to the style guidance image to obtain a prompt word corresponding to the style guidance image includes:
[0017] Preprocessing the style guide image to obtain a preprocessed style guide image;
[0018] According to the preprocessed style guidance image, using a prompt word prediction model and a prompt word library to perform prompt word prediction to obtain an initial prompt word prediction result;
[0019] The initial prompt word prediction result is filtered to obtain a final prompt word corresponding to the style guidance image.
[0020] Optionally, generating conditional embedding and unconditional embedding of a diffusion model according to the image to be style transferred and the style guiding image comprises:
[0021] Extracting image embedding features of the style guide image from the style guide image using a CLIP model;
[0022] Extracting face embedding features of the image to be style transferred from the image to be style transferred using a face recognition model;
[0023] The conditional embedding and unconditional embedding of the diffusion model are generated according to the image embedding features of the style-guided image and the face embedding features of the image to be style-transferred.
[0024] Optionally, generating conditional embedding and unconditional embedding of a diffusion model according to the image to be style transferred and the style guiding image comprises:
[0025] If the face recognition model fails to extract the face embedding feature from the image to be style transferred, extracting the image embedding feature of the image to be style transferred from the image to be style transferred using the CLIP model;
[0026] The conditional embedding and unconditional embedding of the diffusion model are generated according to the image embedding features of the style-guiding image and the image embedding features of the image to be style-migrated.
[0027] Optionally, constructing a diffusion model based on IP-Adapter according to the conditional embedding and unconditional embedding of the diffusion model includes:
[0028] The conditional embedding and unconditional embedding of the diffusion model are integrated into the cross attention layer of the diffusion model to obtain the diffusion model based on IP-Adapter.
[0029] Optionally, the image generation enhancement information corresponding to the image to be style transferred includes a line drawing corresponding to the image to be style transferred, and the generating the image generation enhancement information corresponding to the image to be style transferred by using the enhanced image generation model according to the image to be style transferred includes:
[0030] Performing a first preprocessing on the image to be style-transferred to obtain the image to be style-transferred after the first preprocessing, wherein the first preprocessing includes at least one of adjusting the resolution and normalizing processing;
[0031] According to the first preprocessed image to be style-transferred, using a ControlNet model to generate an initial line drawing corresponding to the image to be style-transferred;
[0032] A first post-processing is performed on an initial line drawing corresponding to the image to be style-transferred to obtain a final line drawing corresponding to the image to be style-transferred, wherein the first post-processing includes linear interpolation.
[0033] Optionally, the image generation enhancement information corresponding to the image to be style transferred includes a depth map corresponding to the image to be style transferred, and the generating the image generation enhancement information corresponding to the image to be style transferred by using the enhanced image generation model according to the image to be style transferred includes:
[0034] Performing a second preprocessing on the image to be style-transferred to obtain the image to be style-transferred after the second preprocessing, wherein the second preprocessing includes at least one of adjusting the resolution and regularization processing;
[0035] According to the second preprocessed image to be style transferred, using the ControlNet model to generate an initial depth map corresponding to the image to be style transferred;
[0036] Performing a second post-processing on the initial depth map corresponding to the image to be style-transferred to obtain a final depth map corresponding to the image to be style-transferred, wherein the second post-processing includes at least one of linear interpolation and normalization.
[0037] In a second aspect, an embodiment of the present application further provides an image style transfer device, the image style transfer device comprising:
[0038] An acquisition unit, used for acquiring an image to be style transferred and a style guiding image;
[0039] A prediction unit, configured to predict prompt words according to the style guidance image using a prompt word prediction model to obtain prompt words corresponding to the style guidance image;
[0040] A construction unit, used to generate conditional embedding and unconditional embedding of a diffusion model according to the image to be style-transferred and the style-guided image, and to construct a diffusion model based on the IP-Adapter according to the conditional embedding and unconditional embedding of the diffusion model;
[0041] An enhancement unit, configured to generate image enhancement information corresponding to the image to be style transferred using an enhanced image generation model according to the image to be style transferred;
[0042] A generating unit is used to generate a style transfer image by using the diffusion model based on the IP-Adapter according to the prompt word corresponding to the style guidance image and the image enhancement information corresponding to the image to be style transferred.
[0043] In a third aspect, an embodiment of the present application further provides a device, including:
[0044] A processor; and a memory arranged to store computer executable instructions, wherein when the executable instructions are executed, the processor performs any of the aforementioned image style transfer methods.
[0045] In a fourth aspect, an embodiment of the present application further provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements any of the aforementioned image style transfer methods.
[0046] At least one of the above-mentioned technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: the image style transfer method in the embodiments of the present application first obtains the image to be style transferred and the style guide image; then, based on the style guide image, a prompt word prediction model is used to predict the prompt word to obtain the prompt word corresponding to the style guide image; then, conditional embedding and unconditional embedding of the diffusion model are generated according to the image to be style transferred and the style guide image, and a diffusion model based on IP-Adapter is constructed according to the conditional embedding and unconditional embedding of the diffusion model; then, based on the image to be style transferred, an enhanced image generation model is used to generate image generation enhancement information corresponding to the image to be style transferred; finally, based on the prompt word corresponding to the style guide image and the image enhancement information corresponding to the image to be style transferred, a style transfer image is generated using the diffusion model based on IP-Adapter. The image style transfer method of the embodiment of the present application utilizes IP-Adapter to perform more refined feature extraction and processing, so that the style transfer result is more in line with expectations, and utilizes the enhanced control capability of the enhanced image generation model to ensure the consistency of the details and structure of the generated image with the original image. By combining IP-Adapter with the enhanced image generation model, a new image style transfer solution is provided, which not only achieves higher-precision image style transfer, but also significantly improves the flexibility and control capability in the image generation process. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0048] Figure 1 A schematic diagram of a process of image style transfer method in an embodiment of the present application;
[0049] Figure 2 This is a schematic diagram of a prompt word prediction process in an embodiment of the present application;
[0050] Figure 3 A schematic diagram of a process of building a diffusion model based on IP-Adapter in an embodiment of the present application;
[0051] Figure 4 This is a schematic diagram of a line drawing generation process in an embodiment of the present application;
[0052] Figure 5 This is a schematic diagram of a depth map generation process in an embodiment of the present application;
[0053] Figure 6 This is a schematic diagram of a style transfer image generation process in an embodiment of the present application;
[0054] Figure 7 This is a schematic diagram of the structure of an image style transfer device in an embodiment of the present application;
[0055] Figure 8 This is a schematic diagram of the structure of a device in an embodiment of the present application. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0057] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.
[0058] The main technical terms involved in this application include:
[0059] 1) IP-Adapter (Image Prompt Adapter): A modular technology for image generation that guides the generation results by introducing image prompts. IP-Adapter helps the generation model understand the features in the image prompts, generate images related to the prompts, and improve the accuracy and diversity of the generated images. As a lightweight network, IP-Adapter plays an important guiding and controlling role in the field of image generation, making the generated images more in line with user needs.
[0060] 2) ControlNet: A neural network architecture for enhancing the control capabilities of text-to-image generation models. It allows adding multiple conditional controls, such as edge and depth information, to generate images with diversity and accuracy. ControlNet combined with the diffusion model can achieve more sophisticated image control and is suitable for a variety of application scenarios.
[0061] 3) Insightface: It is a deep learning framework for face recognition based on deep neural network technology. It can accurately identify and match different facial features by extracting features from face images.
[0062] 4) CLIP (Contrastive Language-Image Pretraining): Embed images and text into the same feature space through contrastive learning. It can understand the relationship between images and text, thereby achieving cross-modal tasks such as image classification, image generation, and text-guided image generation.
[0063] 5) Hires.Fix: It is a post-processing technology for image generation, which is mainly used to enhance the resolution and details of the generated image. By enlarging and supplementing the details of low-resolution images, Hires.Fix can improve the visual quality of the image, especially in detail restoration and clarity improvement.
[0064] The present application embodiment provides an image style transfer method, such as Figure 1 As shown, a flow chart of an image style transfer method in an embodiment of the present application is provided, and the image style transfer method at least includes the following steps S110 to S150:
[0065] Step S110, obtaining the image to be style transferred and the style guide image.
[0066] When performing image style transfer, you need to first obtain the image to be style transferred and the style guide image. The image to be style transferred can be regarded as the target image, that is, the image whose style the user wants to change, and the style guide image can be regarded as the source image, that is, the style the user wants to apply to the target image. These two images serve as the basis for subsequent image style transfer processing.
[0067] Step S120 , predicting prompt words using a prompt word prediction model according to the style guide image, to obtain prompt words corresponding to the style guide image.
[0068] For style-guided images, the embodiment of the present application uses a pre-trained prompt word prediction model to analyze the content, color, texture and other features of the style-guided images, and predicts a series of prompt words that can describe the style of the image. Pre-trained prompt word prediction models can use ViLBERT, UNITER, etc. These models combine visual and language information, and can process image and text data at the same time, so as to predict prompt words related to the image content. Of course, the specific prompt word prediction model to be used can be flexibly selected by technicians in this field according to actual needs, and no specific limitation is made here.
[0069] The embodiment of the present application uses the prompt word prediction model to reversely predict the prompt words from the style guidance image, which can convert the style guidance image into a richer and more diverse guide phrase group, and can generate more comprehensive prompt words compared to the manual labeling method. This not only improves the accuracy of the prompt words, but also greatly optimizes the output effect of the subsequent model, and can more accurately control the style transfer process of the image, especially in detail processing and style consistency.
[0070] Step S130, generating conditional embedding and unconditional embedding of a diffusion model according to the image to be style-transferred and the style-guided image, and constructing a diffusion model based on IP-Adapter according to the conditional embedding and unconditional embedding of the diffusion model.
[0071] Conditional embedding refers to the introduction of additional conditional information into the diffusion model to guide the sample generation process. In the task of image style transfer, the image style information of the style guide image can be introduced according to the task requirements of the image style transfer, thereby guiding the subsequent model to generate an image of a specified style. Unconditional embedding means that the diffusion model does not rely on any additional conditional information when generating an image, but only on the image distribution. The embodiment of the present application can generate conditional embedding and unconditional embedding of the diffusion model based on the image embedding information of the image to be style transferred and the style guide image.
[0072] IP-Adapter is a tool for enhancing image generation capabilities, and is particularly suitable for pre-trained text-to-image diffusion models. The embodiment of the present application is based on the IP-Adapter tool, using conditional embedding and unconditional embedding as inputs to the pre-trained text-to-image diffusion model to form an IP-Adapter-based diffusion model, and gradually generates images with a target style through the process of learning and simulating image style transfer based on the IP-Adapter diffusion model.
[0073] Step S140: generating image enhancement information corresponding to the image to be style transferred using an enhanced image generation model according to the image to be style transferred.
[0074] In order to ensure that the target details in the final generated style transfer image are consistent with the target details in the image to be style transferred, the embodiment of the present application also uses an enhanced image generation model to further analyze the content of the image to be style transferred, and generates some additional information for enhanced image generation, such as the contour and layering of the character target, etc. This information helps to maintain the details and clarity of the image during the style transfer process.
[0075] The above-mentioned enhanced image generation model is mainly used to extract detailed information of the target in the image, and can be implemented by using the ControlNet model, for example. Of course, those skilled in the art can flexibly select which enhanced image generation model to use according to actual needs, and no specific limitation is made here.
[0076] Step S150 , generating a style transfer image using the IP-Adapter-based diffusion model according to the prompt words corresponding to the style guidance image and the image enhancement information corresponding to the image to be style transferred.
[0077] Based on the prompt words of the style guidance image obtained in the previous steps, the enhanced information of the image to be style-transferred, and the diffusion model based on IP-Adapter, the prompt words of the style guidance image and the enhanced information of the image to be style-transferred are used as the input of the diffusion model based on IP-Adapter to guide the diffusion model to generate the final style-transferred image. The generated style-transferred image not only retains the basic content of the image to be style-transferred, but also integrates the style features of the style guidance image.
[0078] It should be noted that there is no strict sequence between the above steps S110 to S130 and they can be executed in parallel.
[0079] The image style transfer method of the embodiment of the present application utilizes IP-Adapter to perform more refined feature extraction and processing, so that the style transfer result is more in line with expectations, and utilizes the enhanced control capability of the enhanced image generation model to ensure the consistency of the details and structure of the generated image with the original image. By combining IP-Adapter with the enhanced image generation model, a new image style transfer solution is provided, which not only achieves higher-precision image style transfer, but also significantly improves the flexibility and control capability in the image generation process.
[0080] In some embodiments of the present application, the method of predicting prompt words using a prompt word prediction model based on the style guide image to obtain prompt words corresponding to the style guide image includes: preprocessing the style guide image to obtain a preprocessed style guide image; predicting prompt words using a prompt word prediction model and a prompt word library based on the preprocessed style guide image to obtain an initial prompt word prediction result; and filtering the initial prompt word prediction result to obtain a final prompt word corresponding to the style guide image.
[0081] Combination Figure 2 , provides a schematic diagram of a prompt word prediction process in an embodiment of the present application. When predicting the prompt word of a style guidance image, the style guidance image may be preprocessed first, and the preprocessing includes, for example, adjusting the image size, filling the white background, etc., in order to make the image more suitable for subsequent processing steps and improve the accuracy of prompt word prediction.
[0082] The preprocessed style-guided image is input into the prompt word prediction model, which can analyze the image content, color, texture and other features, and select the most matching words or phrases from the prompt word library. The prompt word library is a database containing a large number of words or phrases related to the image style, which are used to describe different features of the image.
[0083] Considering that the initial cue word prediction results may contain some inaccurate, redundant, illegal or irrelevant words to the image style, it is necessary to filter these results to remove these words and retain the cue words that can most accurately describe the image style.
[0084] The embodiment of the present application can optimize the image quality by preprocessing the style-guided image, making it more suitable for processing by the prompt word prediction model, which helps to improve the accuracy of the prediction results and enables the generated prompt words to more accurately reflect the style characteristics of the image. Using the prompt word prediction model and the prompt word library for prediction can ensure that the generated prompt words are closely related to the image style, which helps to improve the effect of subsequent generation of style transfer images. By filtering the initial prompt word prediction results, redundant and irrelevant words can be removed, interference from irrelevant information can be avoided, the amount of data for subsequent processing can be reduced, and the accuracy of style transfer can be improved.
[0085] In some embodiments of the present application, the conditional embedding and unconditional embedding of the diffusion model generated based on the image to be style transferred and the style guide image includes: using the CLIP model to extract image embedding features of the style guide image from the style guide image; using a face recognition model to extract face embedding features of the image to be style transferred from the image to be style transferred; generating conditional embedding and unconditional embedding of the diffusion model based on the image embedding features of the style guide image and the face embedding features of the image to be style transferred.
[0086] Combination Figure 3 , a flow chart of constructing a diffusion model based on IP-Adapter in an embodiment of the present application is provided. When generating conditional embedding and unconditional embedding of the diffusion model, the CLIP model can be used to extract image embedding features from the style-guided image. The CLIP model is a multimodal visual and textual learning model that can understand and associate images with text descriptions. Here, it is used to capture style features in the style-guided image, which will be used to guide the style transformation of the image to be style-transferred.
[0087] Use face recognition models such as the InsightFace model to extract face embedding features from the image to be style transferred. The purpose of this step is to identify and extract the key features of the face in the image to ensure that the characteristic information of the face is preserved during the style transfer process.
[0088] The conditional embedding and unconditional embedding of the diffusion model are generated by combining the image embedding features of the style-guided image and the face embedding features of the image to be style-transferred. The conditional embedding contains the specific style information required for style transfer and the key face information of the image to be style-transferred, while the unconditional embedding focuses more on the global statistical characteristics or more general style features of the image, which are not specific to a certain style-guided image or image to be style-transferred.
[0089] By combining the style features extracted by the CLIP model and the facial features extracted by the face recognition model, a more accurate and personalized style transfer can be achieved. The style features ensure that the transferred image meets the specified style requirements, while the facial features ensure the accurate communication of the identity of the face in the image. By generating conditional embeddings and unconditional embeddings, the versatility and flexibility of the diffusion model in the style transfer task are enhanced. Conditional embeddings allow the model to adjust according to specific style guidance images and images to be style transferred, while unconditional embeddings provide a wider range of style representations, allowing the model to adapt to a variety of different style transfer requirements.
[0090] In some embodiments of the present application, the conditional embedding and unconditional embedding of the diffusion model generated according to the image to be style transferred and the style guiding image includes: if the face embedding features are not extracted from the image to be style transferred using the face recognition model, then using the CLIP model to extract the image embedding features of the image to be style transferred from the image to be style transferred; generating the conditional embedding and unconditional embedding of the diffusion model according to the image embedding features of the style guiding image and the image embedding features of the image to be style transferred.
[0091] When performing embedding conversion on the image to be style transferred, the face recognition model can be used to extract face embedding features from the image to be style transferred. If the face recognition model fails to extract face embedding features (i.e., no face is detected or the quality of the extracted features is insufficient), the CLIP model is used to extract image embedding features from the image to be style transferred to capture the overall content characteristics of the image to be style transferred.
[0092] Combine the image embedding features of the style guide image and the image embedding features of the image to be style transferred (in this case, the embedding features extracted by the CLIP model) to generate the conditional embedding and unconditional embedding of the diffusion model. The conditional embedding contains the specific style information required for style transfer and the content information of the image to be style transferred, which is used to guide the style transfer process. The unconditional embedding may focus more on the global statistical characteristics or more general style features of the image, which is not specific to a certain style guide image or image to be style transferred, and is used to provide diversity in style transfer.
[0093] By introducing the CLIP model as an alternative to failed face recognition, the style transfer task is no longer limited to images containing faces, which greatly expands the scope of application of style transfer and enables it to be applied to a wider range of image types. Even in the absence of recognizable faces, the accuracy of style transfer can still be maintained by leveraging the image embedding features extracted by the CLIP model. By providing two feature extraction methods, face recognition and CLIP model, the robustness and flexibility of the model in processing different types of images are enhanced, which helps the model better adapt to various complex style transfer scenarios.
[0094] In some embodiments of the present application, constructing the IP-Adapter-based diffusion model according to the conditional embedding and unconditional embedding of the diffusion model includes: integrating the conditional embedding and unconditional embedding of the diffusion model into the cross-attention layer of the diffusion model to obtain the IP-Adapter-based diffusion model.
[0095] Based on IP-Adapter, the conditional embedding and unconditional embedding obtained in the above embodiment are integrated into the cross-attention layer of the diffusion model. The cross-attention layer is a key component in the diffusion model, which allows the model to focus on different parts of the input data when processing sequence data (such as image pixel sequences) and generate output based on this attention information. During the integration process, conditional embedding and unconditional embedding are added to the cross-attention layer of the existing diffusion model as additional input information, thereby influencing and guiding the model's understanding of image content and the processing of style transfer.
[0096] IP-Adapter uses a decoupled cross-attention mechanism to separate text features and image features, allowing them to be more flexibly combined and matched during the generation process. This decoupling mechanism greatly reduces computing time and resource requirements, and improves the efficiency of style transfer. While ensuring high-quality generation, style transfer can be completed at a faster speed.
[0097] In some embodiments of the present application, the image generation enhancement information corresponding to the image to be style transferred includes a line drawing corresponding to the image to be style transferred, and the image generation enhancement information corresponding to the image to be style transferred is generated by using an enhanced image generation model based on the image to be style transferred, including: performing a first preprocessing on the image to be style transferred to obtain a first preprocessed image to be style transferred, the first preprocessing including at least one of adjusting the resolution and normalizing the image; generating an initial line drawing corresponding to the image to be style transferred using a ControlNet model based on the image to be style transferred after the first preprocessing; performing a first post-processing on the initial line drawing corresponding to the image to be style transferred to obtain a final line drawing corresponding to the image to be style transferred, the first post-processing including linear interpolation.
[0098] Combination Figure 4 , a schematic diagram of a line drawing generation process in an embodiment of the present application is provided. In order to make the image to be style-transferred more suitable for subsequent processing steps, it can be preprocessed first. The preprocessing may include adjusting the resolution and normalization processing. Adjusting the resolution can ensure that the image resolution meets the input requirements of the ControlNet model, and normalization processing helps to standardize the pixel values of the image to make it more suitable for model training and processing.
[0099] The preprocessed image to be style transferred is input into the ControlNet model, which automatically generates a line drawing based on the image content. The line drawing reflects the main structure and edge information of the target in the image.
[0100] Finally, the initial line drawing is further optimized to obtain a more refined and accurate final line drawing. Post-processing operations include processing steps such as linear interpolation, which can be used to smooth the edges of the line drawing, reduce the jagged effect, and make the lines more smooth and natural.
[0101] By generating and using line drawings of the image to be style-transferred, the main structure and edge information of the image can be retained, thereby reducing information loss and deformation during the style transfer process and improving the accuracy of the transfer results. As an abstract representation of an image, the line drawing can highlight the structural and morphological features of the image. In the process of style transfer, combining the line drawing can enable the transferred image to better integrate the target style while retaining the structure of the original image, thereby enhancing the artistic expression of the image. In addition, through pre-processing and post-processing steps, it can be ensured that the image to be style-transferred and the generated line drawing meet the processing requirements of the model, thereby improving the efficiency and flexibility of the entire style transfer process.
[0102] In some embodiments of the present application, the image generation enhancement information corresponding to the image to be style transferred includes a depth map corresponding to the image to be style transferred, and the image generation enhancement information corresponding to the image to be style transferred using an enhanced image generation model based on the image to be style transferred includes: performing a second preprocessing on the image to be style transferred to obtain the image to be style transferred after the second preprocessing, the second preprocessing including at least one of adjusting the resolution and regularization processing; based on the image to be style transferred after the second preprocessing, using a ControlNet model to generate an initial depth map corresponding to the image to be style transferred; performing a second post-processing on the initial depth map corresponding to the image to be style transferred to obtain a final depth map corresponding to the image to be style transferred, the second post-processing including at least one of linear interpolation and normalization processing.
[0103] Combination Figure 5 , a schematic diagram of a depth map generation process in an embodiment of the present application is provided. In order to make the image to be style-transferred more suitable for the input requirements of the ControlNet model to generate a depth map, the image to be style-transferred is also preprocessed first. Preprocessing may include adjusting the resolution and regularization processing, etc. Adjusting the resolution ensures that the image size matches the input size of the ControlNet model. Regularization processing involves smoothing the image to reduce noise, which helps the model to estimate depth information more accurately.
[0104] The preprocessed image to be style transferred is input into the ControlNet model. The model uses visual clues in the image (such as texture, shadow, occlusion relationship, etc.) to estimate the depth value of each pixel to generate an initial depth map.
[0105] Finally, the initial depth map can be further optimized to obtain a more refined and accurate final depth map. Post-processing can include linear interpolation and normalization, etc. Linear interpolation is used to smooth discontinuous areas in the depth map, reduce depth jumps, and make the depth information more continuous and smooth. Normalization standardizes the depth value to a certain range (such as 0 to 1), which is helpful for subsequent processing and application in the style transfer process.
[0106] By introducing the depth map as image generation enhancement information, the style transfer process is able to take into account the three-dimensional spatial information in the image, thereby maintaining or enhancing the depth and stereoscopic sense of the scene while transferring the style, making the transfer result more realistic and natural. The depth map provides information about the distance and hierarchy of objects in the image, which allows the style transfer algorithm to make more detailed adjustments to the style based on the depth information while maintaining the original image structure. For example, different style strengths or styles are applied between the foreground and background. In addition, the pre-processing and post-processing steps ensure that both the input image and the generated depth map meet the processing requirements of the model, which helps to improve the efficiency and stability of the entire style transfer process. Regularization and normalization processing help reduce noise and outliers, making the model output more reliable.
[0107] Combination Figure 6 , a schematic diagram of a style transfer image generation process in an embodiment of the present application is provided, the prompt words of the style guide image and the enhancement information of the image to be style transferred are used as the input of the diffusion model based on IP-Adapter, and the pre-trained LoRA model can also be loaded on the basis of the diffusion model based on IP-Adapter. The LoRA model is a lightweight model adjustment method, which fine-tunes the behavior of the diffusion model based on IP-Adapter by adding a low-rank matrix, thereby further enhancing the details of the style transfer image generation. Furthermore, the Hires.Fix model can also be used to amplify the style transfer image and further enhance the details.
[0108] In summary, the key points of the image style transfer method of this application are mainly:
[0109] 1) Innovatively combining IP-Adapter with ControlNet provides a new image style transfer solution. This combined technology not only achieves higher-precision image style transfer, but also significantly improves the flexibility and control capabilities in the image generation process. Compared with traditional style transfer methods, the IP-Adapter in this application makes the style transfer results more in line with expectations through more sophisticated feature extraction and processing, while ControlNet ensures the consistency of the details and structure of the generated image with the original image by enhancing control capabilities.
[0110] 2) Through the reverse prediction of prompt words, the image information is converted into multimodal prompt words to guide the subsequent image generation process. Combining the diffusion model of IP-Adapter and the line drawings and depth maps generated by ControlNet, highly flexible style control can be achieved to meet the needs of different application scenarios. Whether it is high-resolution images with complex details or scenes requiring high style consistency, it can be effectively handled.
[0111] This application is based on the ControlNet model and cleverly combines the prompt word reverse prediction model with the IP-Adapter preprocessor to achieve higher quality, more labor-saving, and faster style transfer drawing, and has achieved at least the following technical effects:
[0112] 1) Innovative application of the cue word reverse prediction model: Through the cue word reverse prediction model, the style guidance image can be converted into a richer and more diverse guidance phrase, and more comprehensive cue words can be generated compared to manual labeling. This not only improves the accuracy of the cue words, but also greatly optimizes the output effect of the model, and can more accurately control the image generation process, especially in detail processing and style consistency.
[0113] 2) ControlNet’s advantage in detail control: ControlNet provides powerful detail control capabilities for image generation by generating fine line drawings and depth maps. Compared with traditional methods, the application of ControlNet in this application ensures the accuracy of every detail in the generated image, and can more effectively maintain the overall composition and style consistency of the image when processing complex images. The quality of image migration has been significantly improved, which can better meet the needs of high-precision image generation.
[0114] 3) IP-Adapter's efficient preprocessing and decoupling mechanism: IP-Adapter uses a decoupled cross-attention mechanism to separate text features and image features, allowing them to be more flexibly combined and matched during the generation process. This decoupling mechanism greatly reduces computing time and resource requirements, and improves the efficiency of style transfer. While ensuring high-quality generation, style transfer can be completed at a faster speed.
[0115] The present application embodiment also provides an image style transfer device 700, such as Figure 7 As shown, a schematic diagram of the structure of an image style transfer device in an embodiment of the present application is provided, wherein the image style transfer device 700 at least includes: an acquisition unit 710, a prediction unit 720, a construction unit 730, an enhancement unit 740, and a generation unit 750, wherein:
[0116] An acquisition unit 710 is used to acquire an image to be style transferred and a style guiding image;
[0117] A prediction unit 720 is used to predict prompt words according to the style guidance image using a prompt word prediction model to obtain prompt words corresponding to the style guidance image;
[0118] A construction unit 730 is used to generate a conditional embedding and an unconditional embedding of a diffusion model according to the image to be style-transferred and the style-guided image, and to construct a diffusion model based on the IP-Adapter according to the conditional embedding and the unconditional embedding of the diffusion model;
[0119] An enhancement unit 740 is configured to generate image enhancement information corresponding to the image to be style transferred using an enhanced image generation model according to the image to be style transferred;
[0120] The generating unit 750 is configured to generate a style transfer image by using the diffusion model based on the IP-Adapter according to the prompt word corresponding to the style guiding image and the image enhancement information corresponding to the image to be style transferred.
[0121] In some embodiments of the present application, the prediction unit 720 is specifically used to: preprocess the style guidance image to obtain a preprocessed style guidance image; perform prompt word prediction based on the preprocessed style guidance image using a prompt word prediction model and a prompt word library to obtain an initial prompt word prediction result; filter the initial prompt word prediction result to obtain a final prompt word corresponding to the style guidance image.
[0122] In some embodiments of the present application, the construction unit 730 is specifically used to: extract image embedding features of the style guiding image from the style guiding image using the CLIP model; extract face embedding features of the image to be style transferred from the image to be style transferred using the face recognition model; and generate conditional embedding and unconditional embedding of the diffusion model based on the image embedding features of the style guiding image and the face embedding features of the image to be style transferred.
[0123] In some embodiments of the present application, the construction unit 730 is specifically used to: if the face recognition model fails to extract the face embedding features from the image to be style transferred, then use the CLIP model to extract the image embedding features of the image to be style transferred from the image to be style transferred; and generate the conditional embedding and unconditional embedding of the diffusion model according to the image embedding features of the style-guiding image and the image embedding features of the image to be style transferred.
[0124] In some embodiments of the present application, the construction unit 730 is specifically used to: integrate the conditional embedding and unconditional embedding of the diffusion model into the cross attention layer of the diffusion model to obtain the IP-Adapter-based diffusion model.
[0125] In some embodiments of the present application, the image generation enhancement information corresponding to the image to be style transferred includes a line drawing corresponding to the image to be style transferred, and the enhancement unit 740 is specifically used to: perform a first preprocessing on the image to be style transferred to obtain a first preprocessed image to be style transferred, the first preprocessing including at least one of adjusting the resolution and normalizing the image; based on the first preprocessed image to be style transferred, using a ControlNet model to generate an initial line drawing corresponding to the image to be style transferred; perform a first post-processing on the initial line drawing corresponding to the image to be style transferred to obtain a final line drawing corresponding to the image to be style transferred, the first post-processing including linear interpolation.
[0126] In some embodiments of the present application, the image generation enhancement information corresponding to the image to be style transferred includes a depth map corresponding to the image to be style transferred, and the enhancement unit 740 is specifically used to: perform a second preprocessing on the image to be style transferred to obtain the image to be style transferred after the second preprocessing, and the second preprocessing includes at least one of adjusting the resolution and regularization processing; based on the image to be style transferred after the second preprocessing, using the ControlNet model to generate an initial depth map corresponding to the image to be style transferred; perform a second post-processing on the initial depth map corresponding to the image to be style transferred to obtain a final depth map corresponding to the image to be style transferred, and the second post-processing includes at least one of linear interpolation and normalization processing.
[0127] It can be understood that the above-mentioned image style transfer device can implement each step of the image style transfer method provided in the aforementioned embodiment, and the relevant explanations about the image style transfer method are applicable to the image style transfer device, which will not be repeated here.
[0128] Figure 8 Schematic diagram of the structure of a device in the embodiment of the present application. Figure 8 As shown, the device includes one or more processors (or processing units), may further include one or more memories coupled to the processors, and may further include a communication module coupled to the processors.
[0129] The communication module can be used to communicate with other devices or apparatuses, such as the transmission or reception of data and / or signals. The communication module can have at least one communication module for communication. The communication module can include any interface necessary for communicating with other devices. Exemplarily, the communication module can be a transceiver, a circuit, a bus, a module, or other types of communication modules.
[0130] The processor may include, but is not limited to, at least one of the following: a general-purpose computer, a special-purpose computer, a microcontroller, a digital signal controller (DSP), or one or more of a controller-based multi-core controller architecture. The device may have multiple processors, such as application-specific integrated circuit chips, which are time-dependent and synchronized with a clock of a main processor.
[0131] The memory may include one or more non-volatile memories and one or more volatile memories. Examples of non-volatile memories include, but are not limited to, at least one of the following: read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, hard disk, compact disc (CD), digital video disc (DVD), or other magnetic storage and / or optical storage. Examples of volatile memories include, but are not limited to, at least one of the following: random access memory (RAM), or other volatile memories that do not persist during the duration of a power outage.
[0132] The computer program includes computer executable instructions executed by an associated processor. The program can be stored in ROM. The processor can perform any suitable actions and processes by loading the program into RAM.
[0133] The possible implementation of the present application can be implemented by means of a program, so that the communication device can perform any process discussed in the above embodiments. The possible implementation of the present application can also be implemented by hardware or by a combination of software and hardware.
[0134] In some embodiments, the program may be tangibly contained in a computer-readable storage medium, which may be included in the device (such as in a memory) or other storage device accessible by the device. The program may be loaded from the computer-readable storage medium to the RAM for execution. The computer-readable storage medium may include any type of tangible non-volatile memory, such as ROM, EPROM, flash memory, hard disk, CD, DVD, etc.
[0135] The present application embodiment also provides a computer-readable storage medium, on which computer instructions or program codes are stored, and when the processor runs the instructions or the program codes, the processor executes the methods and functions involved in any of the above embodiments. Computer-readable media can be any tangible medium containing or storing programs for or related to instruction execution systems, devices or equipment. Computer-readable media can be computer-readable signal media or computer-readable storage media. Computer-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or devices, or any suitable combination thereof. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrations. More detailed examples of computer-readable storage media include electrical connections with one or more wires, magnetic media (e.g., disks, floppy disks, hard disks, tapes, magnetic storage devices), optical media (e.g., optical storage devices, DVDs), semiconductor media (e.g., solid-state hard drives), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), or any suitable combination thereof, etc.
[0136] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The embodiment of the present application also provides at least one computer program product tangibly stored on a non-temporary computer-readable storage medium. The computer program product includes one or more computer executable instructions, such as instructions included in a program module, which are executed in a device on a real or virtual processor of the target to perform the process, method and function involved in any of the above embodiments. When the computer program instruction is loaded and executed on a computer, a process or function according to an embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instruction can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instruction can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center.
[0137] The present application embodiment also proposes a computer program product, including a computer program or instruction, when the computer program or instruction is run on a computer, the computer is made to perform the process, method and function in the above-mentioned embodiment. Usually, a program module includes routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or realize specific abstract data types. In various embodiments, the functions of program modules can be combined or divided between program modules as needed. Machine executable instructions for program modules can be executed in local or distributed devices. In distributed devices, program modules can be located in local and remote storage media.
[0138] In general, various embodiments of the present application may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software, which may be performed by a controller, microprocessor, or other computing device. Although various aspects of the embodiments of the present disclosure are shown and described as block diagrams, flow charts, or using some other graphical representations, it should be understood that the boxes, devices, systems, techniques, or methods described herein may be implemented as, for example, non-limiting examples, hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.
[0139] It should be noted that although the embodiments of the present application are described above in conjunction with the accompanying drawings, the above embodiments are not independent of each other, and they can also be combined to obtain other embodiments. The division of the modes, situations, categories and embodiments in the embodiments of the present application is only for the convenience of description and should not constitute a special limitation. The features in the various modes, categories, situations and embodiments can be combined with each other in a logical manner. The various implementation methods of the present application can be combined arbitrarily to achieve different technical effects. The embodiments of the present application no longer list various combinations.
[0140] In addition, although the operation of the method of the present disclosure is described in a particular order in the accompanying drawings, this does not require or imply that these operations must be performed in this particular order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flow chart can change the order of execution. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution. It should also be noted that the features and functions of two or more devices according to the present disclosure can be embodied in one device. Conversely, the features and functions of a device described above can be further divided into being embodied by multiple devices.
[0141] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0142] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A method for image style transfer, characterized in that: The image style transfer method comprises: Obtain the image to be style transferred and the style guide image; According to the style guidance image, using a prompt word prediction model to predict prompt words, to obtain prompt words corresponding to the style guidance image; Generating conditional embedding and unconditional embedding of a diffusion model according to the image to be style-transferred and the style-guided image, and constructing a diffusion model based on IP-Adapter according to the conditional embedding and unconditional embedding of the diffusion model; According to the image to be style transferred, using an enhanced image generation model to generate image generation enhancement information corresponding to the image to be style transferred; According to the prompt words corresponding to the style guiding image and the image enhancement information corresponding to the image to be style transferred, the style transfer image is generated by using the diffusion model based on IP-Adapter.
2. The image style transfer method according to claim 1, characterized in that: The step of predicting prompt words using a prompt word prediction model according to the style guidance image to obtain prompt words corresponding to the style guidance image includes: Preprocessing the style guide image to obtain a preprocessed style guide image; According to the preprocessed style guidance image, using a prompt word prediction model and a prompt word library to perform prompt word prediction to obtain an initial prompt word prediction result; The initial prompt word prediction result is filtered to obtain a final prompt word corresponding to the style guidance image.
3. The image style transfer method according to claim 1, characterized in that: The conditional embedding and unconditional embedding of generating a diffusion model according to the image to be style transferred and the style guiding image comprises: Extracting image embedding features of the style guide image from the style guide image using a CLIP model; Extracting face embedding features of the image to be style transferred from the image to be style transferred using a face recognition model; The conditional embedding and unconditional embedding of the diffusion model are generated according to the image embedding features of the style-guided image and the face embedding features of the image to be style-transferred.
4. The image style transfer method according to claim 3, characterized in that: The conditional embedding and unconditional embedding of generating a diffusion model according to the image to be style transferred and the style guiding image comprises: If the face recognition model fails to extract the face embedding feature from the image to be style transferred, extracting the image embedding feature of the image to be style transferred from the image to be style transferred using the CLIP model; The conditional embedding and unconditional embedding of the diffusion model are generated according to the image embedding features of the style-guiding image and the image embedding features of the image to be style-migrated.
5. The image style transfer method according to claim 1, characterized in that: The step of constructing a diffusion model based on IP-Adapter according to the conditional embedding and unconditional embedding of the diffusion model includes: The conditional embedding and unconditional embedding of the diffusion model are integrated into the cross attention layer of the diffusion model to obtain the diffusion model based on IP-Adapter.
6. The image style transfer method according to claim 1, characterized in that: The image generation enhancement information corresponding to the image to be style-transferred includes a line drawing corresponding to the image to be style-transferred, and the image generation enhancement information corresponding to the image to be style-transferred is generated by using an enhanced image generation model according to the image to be style-transferred, including: Performing a first preprocessing on the image to be style-transferred to obtain the image to be style-transferred after the first preprocessing, wherein the first preprocessing includes at least one of adjusting the resolution and normalizing processing; According to the first preprocessed image to be style-transferred, using a ControlNet model to generate an initial line drawing corresponding to the image to be style-transferred; A first post-processing is performed on an initial line drawing corresponding to the image to be style-transferred to obtain a final line drawing corresponding to the image to be style-transferred, wherein the first post-processing includes linear interpolation.
7. The image style transfer method according to claim 1, characterized in that: The image generation enhancement information corresponding to the image to be style transferred includes a depth map corresponding to the image to be style transferred, and the image generation enhancement information corresponding to the image to be style transferred is generated by using an enhanced image generation model according to the image to be style transferred, including: Performing a second preprocessing on the image to be style-transferred to obtain the image to be style-transferred after the second preprocessing, wherein the second preprocessing includes at least one of adjusting the resolution and regularization processing; According to the second preprocessed image to be style transferred, using the ControlNet model to generate an initial depth map corresponding to the image to be style transferred; Performing a second post-processing on the initial depth map corresponding to the image to be style-transferred to obtain a final depth map corresponding to the image to be style-transferred, wherein the second post-processing includes at least one of linear interpolation and normalization.
8. An image style transfer device, characterized in that: The image style transfer device comprises: An acquisition unit, used for acquiring an image to be style transferred and a style guiding image; A prediction unit, configured to predict prompt words according to the style guidance image using a prompt word prediction model to obtain prompt words corresponding to the style guidance image; A construction unit, used to generate conditional embedding and unconditional embedding of a diffusion model according to the image to be style-transferred and the style-guided image, and to construct a diffusion model based on the IP-Adapter according to the conditional embedding and unconditional embedding of the diffusion model; An enhancement unit, configured to generate image enhancement information corresponding to the image to be style transferred using an enhanced image generation model according to the image to be style transferred; A generating unit is used to generate a style transfer image by using the diffusion model based on the IP-Adapter according to the prompt word corresponding to the style guidance image and the image enhancement information corresponding to the image to be style transferred.
9. A device comprising: processor; and a memory arranged to store computer executable instructions, wherein when the executable instructions are executed, the processor executes the image style transfer method according to any one of claims 1 to 7.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the image style transfer method according to any one of claims 1 to 7 is implemented.