Portrait-based controllable style migration and image art enhancement method and system

By combining image segmentation and diffusion models with cascaded generative adversarial networks, the problems of insufficient accuracy and poor generation quality in portrait style transfer in existing technologies are solved, achieving high-precision segmentation and multi-style fusion, thereby improving the artistic expression of images and user experience.

CN120976037APending Publication Date: 2025-11-18HUASHU (ZHEJIANG) TECH CO LTD

Patent Information

Application Number
CN202511099116.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing image style transfer technologies suffer from insufficient accuracy in style transfer, weak personalization capabilities, and poor image quality when processing portrait content. In particular, they suffer from problems such as blurred edges between the subject and background, inconsistent style fusion, limited style types and poor customization flexibility, and loss of facial details.

Method used

We employ a portrait-based controllable style transfer and image art enhancement method. We accurately segment the subject through an image segmentation model, and combine a diffusion model and a cascaded generative adversarial network for style transfer and detail enhancement. We then use the ComfyUI workflow to fuse the images, achieving high-precision segmentation, multi-style fusion, and facial detail restoration.

Benefits of technology

It achieves high-precision portrait segmentation, supports multi-style condition fusion, enhances the artistic expression and generation quality of images, shortens processing time, and improves the real-time nature of user interaction and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976037A_ABST
    Figure CN120976037A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image style migration, in particular to a portrait-based controllable style migration and image art enhancement method and system, and the method comprises the steps: carrying out the preprocessing of an obtained content image, enhancing the contrast and definition of the image, and obtaining a preprocessed image; based on a preset image segmentation model, pixel-level semantic segmentation is carried out on the preprocessed image, and a figure main body image is output; obtaining style information, based on a preset style generation model, performing style migration on the character main body image according to the style information, and generating a stylized character image; obtaining a target style image, and fusing the figure image and the target style image based on a preset ComfyUI workflow to obtain a final processing image; the image processing method and the image processing device have the advantages that on the premise of high customization flexibility, the quality of the image after style migration is improved, and the coordination between the portrait main body and the background is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of image style transfer, and in particular to a method and system for controllable style transfer and image art enhancement based on human portraits. Background Technology

[0002] Currently, image style transfer techniques are mainly divided into traditional algorithms and deep learning-based methods.

[0003] Traditional algorithms apply the style of the source image to the target image by manually designing image features and transformation rules. For example, texture synthesis-based methods achieve style transfer by matching and copying the texture structure of the style image to the content image. Methods based on traditional algorithm optimization, on the other hand, continuously adjust the pixel values ​​of the target image by constructing a loss function to make it close to the original image and style image in terms of content and style.

[0004] Deep learning-based methods, such as classic models like CycleGAN and StyleGAN, train generative adversarial networks using large amounts of image data to achieve image style transfer.

[0005] When the aforementioned image style transfer techniques are used to process portrait content, they generally suffer from insufficient accuracy in style transfer and weak personalization capabilities.

[0006] First, the edges between the subject and the background are blurred, and the style blending is inconsistent. Traditional techniques often use global style transfer, which results in abrupt transitions between the subject's edges and the background during style transfer. For example, the style characteristics of the person's hair and the background scenery are mixed, resulting in blurred or misaligned edges, which disrupts the overall visual consistency of the image.

[0007] Secondly, the styles are limited and customization is inflexible. Existing technologies mostly rely on preset style templates (such as fixed art styles), which prevent users from freely defining styles according to their needs and make it difficult to meet diverse creative scenarios.

[0008] Furthermore, the generated image quality is poor. Due to the lack of targeted detail optimization mechanisms, style transfer often results in problems such as facial feature distortion, loss of hair texture, and unnatural lighting transitions. For example, in anime style transfer, the outline details of the character's eyes are oversimplified, leading to distorted facial features and affecting the image's artistic expression. Summary of the Invention

[0009] In order to improve the image quality after style transfer and the harmony between the subject and the background while maintaining high customization flexibility, this application provides a method and system for controllable style transfer and image art enhancement based on portraits.

[0010] Firstly, this application provides a method for controllable style transfer and image artistic enhancement based on human portraits, employing the following technical solution:

[0011] A method for controlled style transfer and image art enhancement based on portraits, comprising:

[0012] The acquired content image is preprocessed to enhance its contrast and clarity, resulting in a preprocessed image.

[0013] Based on a preset image segmentation model, pixel-level semantic segmentation is performed on the preprocessed image to output the main image of the person.

[0014] Obtain style information, generate a stylized image of the person based on a preset style generation model, perform style transfer on the main image of the person according to the style information;

[0015] The target style image is obtained, and based on the preset ComfyUI workflow, the portrait image and the target style image are merged to obtain the final processed image.

[0016] Obtain adjustment parameters and update the final processed image based on the gradient algorithm.

[0017] In one embodiment: the network structure of the image segmentation model includes an encoder and a decoder. The encoder extracts image features through convolution operations, and the decoder restores the image size through transposed convolution. The encoder and decoder fuse features from different levels through skip connections.

[0018] The image segmentation model described is an improved U-Net model that incorporates multi-head self-attention layers. Multi-head self-attention layers are introduced in both the encoder and decoder, and the attention weights are calculated using the following formula:

[0019] ;

[0020] In the formula, Q and K are the query vector and key vector of the feature mapping, respectively.

[0021] In one embodiment: the method for obtaining the main image of the person specifically includes:

[0022] Based on the image segmentation model, each pixel of the preprocessed image is classified. Pixels that are identified as human regions are labeled with the first value, and pixels that are identified as background regions are labeled with the second value, thus forming a binarized mask.

[0023] The main image of the person is extracted by multiplying the mask and the original image pixel by pixel.

[0024] In one embodiment: the style transfer of the style generation model is implemented based on a diffusion model, and the training method of the diffusion model is as follows:

[0025] Acquire training data, which includes original portrait images and style information;

[0026] The original portrait image is noise-added to generate noisy images with different noise levels;

[0027] The noisy image, noise level parameters, and style labels are input together into the denoising network of the diffusion model, so that the denoising network learns to remove noise while learning the style features corresponding to the style labels.

[0028] The parameters of the denoising network are iteratively updated using a loss function that includes style constraints until the value of the loss function drops to a preset threshold.

[0029] In one embodiment: the loss function including style constraints is:

[0030] ;

[0031] In the formula, x0 is the original image of the main person. The noise is random, and t is the number of diffusion steps. For noise reduction networks.

[0032] In one embodiment: if the style information includes multiple style features, a corresponding style image is generated based on each style feature in the style information through the style generation model;

[0033] Obtain the fusion weights, and then perform weighted fusion on the generated multiple style images based on the fusion weights to obtain a stylized portrait image.

[0034] In one embodiment: the step of obtaining the target style image, based on a preset ComfyUI workflow, by fusing the person image and the target style image to obtain the final processed image specifically includes:

[0035] The target style image is obtained, and based on the style transfer algorithm, the feature differences between the target style image and the image of the person are analyzed. Based on the feature differences, the global style consistency of the image is adjusted.

[0036] Based on a cascaded generative adversarial network, texture restoration and detail enhancement are performed on the facial region of a globally style-enhanced portrait image to obtain the final portrait image.

[0037] The final image of the person and the target style image for global style enhancement are merged to obtain the final processed image.

[0038] In one embodiment: the steps of acquiring the target style image, analyzing the feature differences between the target style image and the person image based on a style transfer algorithm, and adjusting the global style consistency of the image based on the feature differences specifically include:

[0039] Based on a preset feature extraction model, the target style image and the person image are extracted to obtain the corresponding feature map;

[0040] Calculate the corresponding Gram matrix based on the feature maps of the two images;

[0041] Based on the style transfer algorithm, the style loss of the stylized image is calculated using the Gram matrix;

[0042] Adjusting global style consistency of images based on minimizing style loss.

[0043] In one embodiment: the step of performing texture restoration and detail enhancement on the facial region of a globally style-enhanced portrait image based on a cascaded generative adversarial network to obtain the final portrait image specifically includes:

[0044] Facial features are obtained based on a facial landmark detection model;

[0045] Based on a pre-defined generative adversarial network, facial features are textured and enhanced to obtain the final image of the person.

[0046] The generator loss function of the generative adversarial network is:

[0047] ;

[0048] In the formula, D is the discriminator, G is the generator, x is the real image, and z is the noise vector.

[0049] Secondly, this application provides a controllable style transfer and image art enhancement system based on human portraits, employing the following technical solution:

[0050] A controllable style transfer and image art enhancement system based on human portraits, characterized in that it includes:

[0051] The image acquisition and preprocessing module is used to preprocess the acquired content images, enhance the contrast and clarity of the images, and obtain the preprocessed images.

[0052] The portrait segmentation module is used to perform pixel-level semantic segmentation on the preprocessed image based on a preset image segmentation model, and output the main image of the person.

[0053] The AIGC style generation module is used to acquire style information, perform style transfer on the main image of the person based on the preset style generation model, and generate a stylized image of the person.

[0054] The ComfyUI processing module is used to acquire the target style image and, based on the preset ComfyUI workflow, merge the person image and the target style image to obtain the final processed image.

[0055] The user interaction and parameter adjustment module is used to obtain adjustment parameters and update the final processed image based on the gradient algorithm.

[0056] In summary, this application has the following beneficial effects:

[0057] 1. By leveraging the powerful feature extraction and semantic understanding capabilities of the Unet model, and further optimizing the network structure and adjusting parameter configurations, high-precision portrait segmentation results were achieved. It can accurately capture the details of a person's outline, clearly distinguishing the subject from the background, whether it's flowing hair or complex clothing textures.

[0058] 2. The AIGC style generation model based on the diffusion model improves the denoising objective function by introducing style information. Combined with cloud-preset anime style, pixel style and other style models, it not only supports users to select a single style, but also supports multi-style conditional fusion mechanism. Moreover, users can adjust the fusion weight in real time, which greatly expands the style types and meets the growing demand for personalized image creation. It breaks through the bottleneck of the limited style types of traditional technology.

[0059] 3. In the ComfyUI workflow, the Flux image generation node uses the Gram matrix to calculate the style loss. By minimizing this loss, global style enhancement is achieved, ensuring that the generated image is highly consistent with the selected style in terms of color, texture, and composition.

[0060] 4. The facial detail optimization node is based on facial key point detection and cascaded generative adversarial network. It performs texture restoration and detail enhancement on the facial region. It uses the generator loss function for adversarial training, which can effectively restore facial details, such as making the eyes more expressive and the hair more delicate. Compared with traditional methods, the integrity of facial details is improved by more than 30%, which effectively solves the problems of facial detail loss and structural deformation in existing technologies, and achieves high-quality image artistic enhancement effect.

[0061] 5. User interaction and parameter adjustment are achieved by acquiring adjustment parameters. Through a real-time gradient adjustment algorithm, when the user adjusts the parameters, the system quickly calculates and generates the gradient of the image with respect to the parameters, achieving real-time image updates with a response time controllable within 0.5 seconds. Compared to traditional techniques that require recalculating and generating images, this significantly improves the real-time performance and smoothness of parameter adjustment, allowing users to instantly see the adjustment effects, fully unleashing creative expression space, and significantly enhancing the user experience.

[0062] 6. The system deeply integrates portrait segmentation algorithms, AIGC style generation algorithms, and ComfyUI platform processing algorithms. Each module collaborates through optimized data transmission and interaction mechanisms, forming a complete closed loop from image acquisition and preprocessing to final output. Compared to the relatively independent processing methods in existing technologies, this invention reduces redundant data processing and improves overall processing efficiency. When processing a single image, the total processing time is reduced by approximately 40% compared to traditional solutions, demonstrating greater practicality and application value. Attached Figure Description

[0063] Figure 1 This is a framework diagram of the controllable style transfer and image art enhancement system based on human portraits in this embodiment;

[0064] Figure 2 This is a flowchart of the controllable style transfer and image art enhancement method based on human portrait in this embodiment;

[0065] Figure 3 This is the network architecture of the image segmentation model in this embodiment.

[0066] In the diagram, 10 is the image acquisition and preprocessing module; 20 is the portrait segmentation module; 30 is the AIGC style generation module; 40 is the ComfyUI processing module; and 50 is the user interaction and parameter adjustment module. Detailed Implementation

[0067] The present application will be further described in detail below with reference to the accompanying drawings.

[0068] To better understand the purpose, technical solutions, and advantages of this application, it has been described and illustrated below with reference to the accompanying drawings and embodiments. However, those skilled in the art should understand that this application can be implemented without these details. In some cases, to avoid obscuring various aspects of this application due to unnecessary description, well-known methods, processes, systems, components, and / or circuits already described at a higher level will not be elaborated upon. It will be apparent to those skilled in the art that various modifications can be made to the embodiments disclosed in this application, and the general principles defined in this application can be applied to other embodiments and application scenarios without departing from the principles and scope of this application. Therefore, this application is not limited to the illustrated embodiments, but conforms to the broadest scope consistent with the scope of protection claimed in this application.

[0069] It should be noted that the descriptions of these embodiments are for the purpose of aiding understanding the present invention, but do not constitute a limitation thereof. Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0070] In the description of this application, "several" means one or more, "more than" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. The use of "first" and "second" in the description is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.

[0071] In the description of this application, the terms "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples.

[0072] like Figure 1As shown, this application discloses a controllable style transfer and image art enhancement system based on human portraits, which includes an image acquisition and preprocessing module 10, a human portrait segmentation module 20, an AIGC style generation module 30, a ComfyUI processing module 40, and a user interaction and parameter adjustment module 50. In this system, the human portrait segmentation algorithm, the AIGC style generation algorithm, and the ComfyUI platform processing algorithm are deeply integrated. The modules work collaboratively through optimized data transmission and interaction mechanisms, forming a complete closed loop from image acquisition and preprocessing to final output.

[0073] The image acquisition and preprocessing module 10 acquires RGB content images through devices such as an all-in-one camera. The preprocessing module preprocesses the acquired content images to enhance their contrast and clarity, and then inputs the preprocessed image into the portrait segmentation module 20.

[0074] The portrait segmentation module 20 is used to perform pixel-level semantic segmentation on the preprocessed image based on a preset image segmentation model, output the main image of the person, and input the main image of the person into the AIGC style generation module 30.

[0075] AIGC style generation module 30 is used to acquire style information, perform style transfer on the main image of the person based on the preset style generation model, generate a stylized image of the person, and input the image of the person into ComfyUI processing module 40.

[0076] ComfyUI processing module 40 is used to acquire the target style image, and based on the preset ComfyUI workflow, it merges the person image and the target style image to obtain the final processed image.

[0077] The user interaction and parameter adjustment module 50 is used to obtain adjustment parameters and update the final processed image based on the gradient algorithm.

[0078] like Figure 2 As shown, a method for controllable style transfer and image artistic enhancement based on portraits includes the following steps:

[0079] S100. Preprocess the acquired content image to enhance its contrast and clarity, and then obtain the preprocessed image.

[0080] In this step, the preprocessing process involves first converting the content image into a grayscale image, then using an adaptive noise suppression algorithm based on the bilateral filtering principle for noise reduction, and finally using the contrast-limited adaptive histogram equalization (CLAHE) algorithm to improve the visual effect of the image through local histogram statistics and contrast stretching, thereby achieving contrast enhancement.

[0081] The filtering formula for bilateral filtering is as follows:

[0082] ;

[0083] Where f(i,j) is the range weight, d(i,j) is the domain weight, and I(i,j) is the original image pixel value. This algorithm can remove noise while preserving image edge details.

[0084] S200: Based on a preset image segmentation model, perform pixel-level semantic segmentation on the preprocessed image and output the main image of the person.

[0085] In the above steps, the image segmentation model is an improved U-Net model that incorporates a multi-head self-attention layer, such as... Figure 3 As shown, its network structure includes an encoder and a decoder. The encoder consists of multiple convolutional and pooling layers, which gradually reduce the image resolution and extract features. During this process, the number of channels gradually increases, focusing on key information in the image.

[0086] The decoder recovers the image size through deconvolution and upsampling operations, during which the number of channels is gradually reduced, and features are fused with the corresponding layer of the encoder through skip connections.

[0087] The arrows in the diagram clearly indicate the flow of image data from input to output, as well as the processing paths in each layer of the network, highlighting how the improvement mechanism enhances and utilizes feature information.

[0088] The objective function for training the image segmentation model described above is the cross-entropy loss function:

[0089] ;

[0090] In the formula, N is the number of samples, C is the number of categories, and y ic p is the true label of sample i. ic Let c be the probability of the predicted category by the model, where the number of categories is divided into two categories: people and background.

[0091] Furthermore, in this embodiment, a multi-head self-attention layer is introduced in the encoder and decoder, and the attention weight calculation formula is as follows:

[0092] ;

[0093] In the formula, Q and K are the query vector and key vector of the feature mapping, respectively. An attention mechanism is used to focus on key parts of the subject, improving segmentation accuracy.

[0094] In this step, the main image of the person is obtained in the following way:

[0095] Based on the image segmentation model, each pixel of the preprocessed image is classified. Pixels that are identified as human regions are labeled with the first value, and pixels that are identified as background regions are labeled with the second value, thus forming a binarized mask.

[0096] The main image of the person is extracted by multiplying the mask and the original image pixel by pixel.

[0097] By minimizing the cross-entropy loss function, each pixel of the preprocessed image can be accurately classified, and a binary mask is output. The mask has a value of 1 in the human region and 0 in the background region.

[0098] The main subject is extracted by multiplying the mask element-by-element with the original image using the following formula:

[0099] ;

[0100] In the formula, I1(x,y) is the pixel coordinate of the preprocessed image, and M(x,y) is the mask corresponding to the pixel coordinate.

[0101] S300: Obtain style information, and based on the preset style generation model, perform style transfer on the main image of the person according to the style information to generate a stylized image of the person.

[0102] The style transfer in this step of the style generation model is based on the diffusion model. The user selects style information through the interactive module, and the diffusion model completes the style transfer through iterative denoising. Through this conditional training, the diffusion model can generate corresponding style images according to different style information.

[0103] The training method for the above diffusion model is as follows:

[0104] Acquire training data, which includes original portrait images and style information;

[0105] The original portrait image is noise-added to generate noisy images with different noise levels;

[0106] The noisy image, noise level parameters, and style labels are input together into the denoising network of the diffusion model, so that the denoising network learns to remove noise while learning the style features corresponding to the style labels.

[0107] The parameters of the denoising network are iteratively updated using a loss function that includes style constraints until the value of the loss function drops to a preset threshold.

[0108] In this embodiment, the loss function that includes style constraints is:

[0109] ;

[0110] In the formula, x0 is the original image of the main person. The noise is random, and t is the number of diffusion steps. For noise reduction networks.

[0111] Furthermore, if the style information includes multiple style features, a corresponding style image is generated based on each style feature in the style information using the style generation model. Then, the fusion weights are obtained, and the generated multiple style images are weighted and fused based on the fusion weights to obtain a stylized portrait image.

[0112] Taking two style features as an example, the generated results of different style models are obtained by weighted fusion:

[0113] ;

[0114] In the formula, I style1 and I style2 Images generated for models of different styles. This is to integrate weights. Users can adjust these weights in real time through the user interaction and parameter adjustment modules to achieve personalized style customization.

[0115] The loss function described above is used to measure the difference between the noise predicted by the denoising network and the actual noise, while also constraining the consistency between the denoising result and the style label.

[0116] S400: Obtain the target style image. Based on the preset ComfyUI workflow, merge the person image and the target style image to obtain the final processed image.

[0117] In one embodiment, the steps of obtaining the target style image, fusing the person image and the target style image based on a preset ComfyUI workflow, and obtaining the final processed image specifically include:

[0118] The target style image is obtained, and based on the style transfer algorithm, the feature differences between the target style image and the image of the person are analyzed. Based on the feature differences, the global style consistency of the image is adjusted.

[0119] Based on a cascaded generative adversarial network, texture restoration and detail enhancement are performed on the facial region of a globally style-enhanced portrait image to obtain the final portrait image.

[0120] The final image of the person and the target style image for global style enhancement are merged to obtain the final processed image.

[0121] In this embodiment, the specific steps for adjusting the global style consistency of the image are as follows:

[0122] Based on a preset feature extraction model, the target style image and the person image are extracted to obtain the corresponding feature map;

[0123] Calculate the corresponding Gram matrix based on the feature maps of the two images;

[0124] Based on the style transfer algorithm, the style loss of the stylized image is calculated using the Gram matrix;

[0125] Adjusting global style consistency of images based on minimizing style loss.

[0126] The formula for calculating style loss using the Gram matrix in the above steps is:

[0127] ;

[0128] In the formula, N i and M i These represent the number of channels and the size of the feature map, respectively, G s and G c These are Gram matrices for the style image and the content image, respectively.

[0129] In another embodiment, the steps described above, based on a cascaded generative adversarial network, to perform texture restoration and detail enhancement on the facial region of a globally style-enhanced portrait image to obtain the final portrait image, specifically include:

[0130] Based on the facial landmark detection model, facial features are obtained, which are the coordinates of 68 key points on the face.

[0131] Based on a pre-defined generative adversarial network, facial features are textured and enhanced to obtain the final image of the person.

[0132] The generator loss function of the above generative adversarial network is:

[0133] ;

[0134] In the formula, D is the discriminator, G is the generator, x is the real image, and z is the noise vector. Facial details are optimized through local texture synthesis and restoration algorithms.

[0135] S500: Obtain adjustment parameters and update the final processed image based on gradient algorithm.

[0136] In this step, parameters such as style intensity, color saturation, and blending weights are adjusted. The final processed image is then updated until it meets the user's needs, and finally saved as an output. The user can choose to print or export it.

[0137] The embodiments described in this specific implementation are preferred embodiments of this application and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A method and system for controllable style transfer and image art enhancement based on human portraits, characterized in that, include: The acquired content image is preprocessed to enhance its contrast and clarity, resulting in a preprocessed image. Based on a preset image segmentation model, pixel-level semantic segmentation is performed on the preprocessed image to output the main image of the person. Obtain style information, generate a stylized image of the person based on a preset style generation model, perform style transfer on the main image of the person according to the style information; The target style image is obtained, and based on the preset ComfyUI workflow, the portrait image and the target style image are merged to obtain the final processed image. Obtain adjustment parameters and update the final processed image based on the gradient algorithm.

2. The method and system for controllable style transfer and image art enhancement based on human portraits according to claim 1, characterized in that: The network structure of the image segmentation model includes an encoder and a decoder. The encoder extracts image features through convolution operations, and the decoder restores the image size through transposed convolution. Furthermore, the encoder and decoder fuse features from different levels through skip connections. The image segmentation model described is an improved U-Net model that incorporates multi-head self-attention layers. Multi-head self-attention layers are introduced in both the encoder and decoder, and the attention weights are calculated using the following formula: ; In the formula, Q and K are the query vector and key vector of the feature mapping, respectively.

3. The method and system for controllable style transfer and image art enhancement based on human portraits according to claim 1 or 2, characterized in that, The method for obtaining the main image of the person specifically includes: Based on the image segmentation model, each pixel of the preprocessed image is classified. Pixels that are identified as human regions are labeled with the first value, and pixels that are identified as background regions are labeled with the second value, thus forming a binarized mask. The main image of the person is extracted by multiplying the mask and the original image pixel by pixel.

4. The method and system for controllable style transfer and image art enhancement based on human portraits according to claim 1, characterized in that: The style transfer of the style generation model is based on a diffusion model, and the training method of the diffusion model is as follows: Acquire training data, which includes original portrait images and style information; The original portrait image is noise-added to generate noisy images with different noise levels; The noisy image, noise level parameters, and style labels are input together into the denoising network of the diffusion model, so that the denoising network learns to remove noise while learning the style features corresponding to the style labels. The parameters of the denoising network are iteratively updated using a loss function that includes style constraints until the value of the loss function drops to a preset threshold.

5. The method and system for controllable style transfer and image art enhancement based on human portraits according to claim 4, characterized in that, The loss function that includes style constraints is: ; In the formula, x0 is the original image of the main person. The noise is random, and t is the number of diffusion steps. For noise reduction networks.

6. The method and system for controllable style transfer and image art enhancement based on human portraits according to claim 1, 4, or 5, characterized in that: If the style information includes multiple style features, the style generation model generates a corresponding style image based on each style feature in the style information. Obtain the fusion weights, and then perform weighted fusion on the generated multiple style images based on the fusion weights to obtain a stylized portrait image.

7. The method and system for controllable style transfer and image art enhancement based on human portraits according to claim 1, characterized in that, The steps of obtaining the target style image, based on a preset ComfyUI workflow, and fusing the portrait image and the target style image to obtain the final processed image specifically include: The target style image is obtained, and based on the style transfer algorithm, the feature differences between the target style image and the image of the person are analyzed. Based on the feature differences, the global style consistency of the image is adjusted. Based on a cascaded generative adversarial network, texture restoration and detail enhancement are performed on the facial region of a globally style-enhanced portrait image to obtain the final portrait image. The final image of the person and the target style image for global style enhancement are merged to obtain the final processed image.

8. The method and system for controllable style transfer and image art enhancement based on human portraits according to claim 7, characterized in that, The steps of acquiring the target style image, analyzing the feature differences between the target style image and the person image based on the style transfer algorithm, and adjusting the global style consistency of the image based on the feature differences specifically include: Based on a preset feature extraction model, the target style image and the person image are extracted to obtain the corresponding feature map; Calculate the corresponding Gram matrix based on the feature maps of the two images; Based on the style transfer algorithm, the style loss of the stylized image is calculated using the Gram matrix; Adjusting global style consistency of images based on minimizing style loss.

9. The method and system for controllable style transfer and image art enhancement based on human portraits according to claim 1, characterized in that, The steps for performing texture restoration and detail enhancement on the facial region of a globally style-enhanced portrait image based on a cascaded generative adversarial network to obtain the final portrait image specifically include: Facial features are obtained based on a facial landmark detection model; Based on a pre-defined generative adversarial network, facial features are textured and enhanced to obtain the final image of the person. The generator loss function of the generative adversarial network is: ; In the formula, D is the discriminator, G is the generator, x is the real image, and z is the noise vector.

10. A controllable style transfer and image art enhancement system based on human portraits, characterized in that, include: The image acquisition and preprocessing module (10) is used to preprocess the acquired content image, enhance the contrast and clarity of the image, and obtain the preprocessed image. The portrait segmentation module (20) is used to perform pixel-level semantic segmentation on the preprocessed image based on a preset image segmentation model and output the main image of the person. The AIGC style generation module is used to acquire style information, perform style transfer on the main image of the person based on the preset style generation model, and generate a stylized image of the person. The ComfyUI processing module (40) is used to obtain the target style image, and based on the preset ComfyUI workflow, it merges the character image and the target style image to obtain the final processed image. The user interaction and parameter adjustment module (50) is used to obtain adjustment parameters and update the final processed image based on the gradient algorithm.

Citation Information

Patent Citations

  • A semantic style migration method and system fusing depth features

    CN109712081A

  • Image processing method and device, equipment and storage medium

    CN110956654A

  • Image cross-domain migration method with controllable stylization degree, computer equipment, readable storage medium and program product

    CN115100025A

  • Stylized image generation method, system and device and medium

    CN115272146A

  • Method for artistic style migration based on diffusion model, computer equipment, readable storage medium and program product

    CN118505497A

Cited By

  • Generation method and system of Q-version digital head portrait, and storage medium

    CN121437673A

  • Method for finely adjusting color difference of adjacent pictures in text H5 scene

    CN121582359A