Method and device for generating specified model chart data based on ComfyUI

Through the image generation method based on ComfyUI, combined with multiple algorithms and lighting optimization technologies, the problem of difficulty in controlling model characteristics and real texture in the existing technology is solved, and a highly personalized and customized model image generation is achieved.

CN120070619APending Publication Date: 2025-05-30ZIXUN TECHNOLOGY (FUJIAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108220.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing image generation algorithms have difficulty accurately controlling the specific appearance characteristics of models in the generated image, such as facial features, body proportions, or specific postures, while it is difficult to achieve real-textured lighting and skin texture.

Method used

Using a ComfyUI-based method, by obtaining the front face portrait data and model reference pose charts required by users, combining Ipadapter, InstantId and openpose algorithms, model images with specified facial features, body proportions and poses are generated, and lighting optimization is performed through IC-Light to achieve realistic skin texture texture.

Benefits of technology

Accurate control of model images is achieved, and images with realistic lighting effects and realistic skin texture are generated, which significantly improves the personalization and customization of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070619A_ABST
    Figure CN120070619A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for generating specified model chart data based on a ComfyUI. The method comprises the following steps: acquiring front face portrait data of an uploaded picture; acquiring an uploaded model reference posture map; front face portrait data is processed through an Ipayer algorithm and an InstantId algorithm, a model reference posture graph is processed through an openpose algorithm, and then results obtained through processing are transmitted into a first graph generation node of a ComfyUI in a unified mode to generate portrait data; face segmentation and eye detection are carried out on the portrait data, face restoration and eye restoration are carried out through mask redrawing of the second image generation nodes, and a restoration image is obtained; transmitting the repaired image into IC-Light, and carrying out illumination optimization according to a set prompt word to obtain an optimized image; a model image with specified facial features, a body proportion and a specific posture can be accurately generated according to the requirements of a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and device for generating specified model image data based on ComfyUI. Background Art

[0002] In the field of AIGC (Artificial Intelligence Generated Content), generating fixed models often involves training on a large amount of data and generating specified model data by controlling specified prompt words or other parameters; it faces some challenges and problems:

[0003] 1. Existing image generation algorithms often cannot precisely control specific appearance features of the models in the generated images, such as facial features, body proportions, or specific poses, etc.

[0004] 2. In the generated images, it is often difficult to achieve realistic textures such as lighting, shadows, and skin textures. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method and device for generating specified model image data based on ComfyUI, which can not only accurately generate model images with specified facial features, body proportions, and specific poses according to user needs, but also generate realistic lighting effects and realistic skin texture sensations, greatly improving the personalization and customization level of image generation.

[0006] In the first aspect, the present invention provides a method for generating specified model image data based on ComfyUI, including the following steps:

[0007] Step 1: Obtain the frontal portrait data of the uploaded picture;

[0008] Step 2: Obtain the uploaded model reference pose picture;

[0009] Step 3: ComfyUI calls the Ipadapter algorithm, adjusts the weight type to ease out, and sets the weight parameter weight of the Ipadapter algorithm to 0.5;

[0010] Step 4: ComfyUI calls the InstantId algorithm and sets the weight parameter weight of the InstantId algorithm to 0.6;

[0011] Step 5: ComfyUI calls the openpose algorithm, and the openpose algorithm is used for body key point detection of the model reference pose picture and outputs a key point picture;

[0012] Step 6: The frontal portrait data is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference pose map is processed by the openpose algorithm. Then, the processed results are uniformly input into the first image generation node of ComfyUI to generate portrait data;

[0013] Step 7: The portrait data is subjected to face segmentation and eye detection, and face repair and eye repair are performed through mask redrawing of the second image generation node to obtain a repaired image;

[0014] Step 8: The repaired image is input into IC-Light, and lighting optimization is performed according to the set prompt words to obtain an optimized image.

[0015] In a second aspect, the present invention provides a device for generating specified model image data based on ComfyUI, including:

[0016] A frontal face acquisition module for acquiring the frontal portrait data of the uploaded image;

[0017] A pose acquisition module for acquiring the model reference pose map uploaded;

[0018] A first algorithm setting module, where ComfyUI calls the Ipadapter algorithm, adjusts the weight type to ease out, and sets the weight parameter weight of the Ipadapter algorithm to 0.5;

[0019] A second algorithm setting module, where ComfyUI calls the InstantId algorithm and sets the weight parameter weight of the InstantId algorithm to 0.6;

[0020] A third algorithm setting module, where ComfyUI calls the openpose algorithm, and the openpose algorithm is used for body key point detection of the model reference pose map and outputs a key point map;

[0021] An image generation module, where the frontal portrait data is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference pose map is processed by the openpose algorithm. Then, the processed results are uniformly input into the first image generation node of ComfyUI to generate portrait data;

[0022] An image repair module for performing face segmentation and eye detection on the portrait data, and performing face repair and eye repair through mask redrawing of the second image generation node to obtain a repaired image;

[0023] An image optimization module for inputting the repaired image into IC-Light and performing lighting optimization according to the set prompt words to obtain an optimized image.

[0024] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:

[0025] The present invention can not only accurately generate model images with specified facial features, body proportions and specific postures according to the needs of users, but also generate realistic lighting effects and realistic skin texture, greatly improving the personalization and customization level of image generation; enabling merchants to quickly apply the generated images to their product promotion, greatly increasing the promotion speed of users and reducing the promotion costs of merchants.

[0026] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. Brief Description of the Drawings

[0027] The present invention will be further described below with reference to the drawings in conjunction with embodiments.

[0028] Figure 1 It is a flowchart of the method in Embodiment 1 of the present invention;

[0029] Figure 2 It is a structural schematic diagram of the device in Embodiment 2 of the present invention. Detailed Embodiments

[0030] By providing a method and device for generating specified model image data based on ComfyUI in the embodiments of the present application, the technical problem that the existing image generation algorithm cannot accurately control the specific appearance features of the model in the generated image is solved.

[0031] The general idea of the technical solution in the embodiments of the present application is as follows:

[0032] (1) Upload clear frontal portrait data of the specified model as the facial reference for generating the specified fixed model

[0033] (2) Upload a reference pose map of the model as the proportion reference and pose reference for generating the model image (to generate an image with a specific pose, just upload an image of that pose)

[0034] (3) Call the Ipadapter algorithm, adjust the weight type to "ease out", and adjust the weight parameter weight to 0.5. The main purpose of this parameter is to control the influence intensity of the Ipadapter algorithm. (In ComfyUI, the ipadapter node has a controllable weight parameter, which is mainly used to regulate the influence intensity of the algorithm)

[0035] (4) Call the InstantId algorithm and adjust the weight parameter weight to 0.6. This parameter is the same as the weight parameter in the Ipadapter algorithm and is used to control the influence intensity of the algorithm. (In ComfyUI, the InstantId node has an adjustable weight parameter, which is mainly used to control the influence intensity of the algorithm.)

[0036] (5) Call the openpose algorithm to detect the body key points of the model's pose in the generated pose reference image. Here, the 18-joint key point model in openpose is used.

[0037] (6) Input the above weight and the pose preprocessed image into the image-to-image node in ComfyUI to generate preliminary portrait data. The parameter settings in the image-to-image are as follows: the prompt correlation coefficient is set to 4 - 9, preferably 7, and the redrawing amplitude is set to 0.75.

[0038] (7) Perform face segmentation and eye detection on the portrait data generated by the image-to-image (face segmentation directly calls the face segmentation model and can be directly loaded and called). Use a small amount of image-to-image mask redrawing for face repair and eye repair. Here, the parameter settings of the image-to-image are as follows: the prompt correlation coefficient is set to 1, the redrawing steps are set to 4, and the redrawing amplitude is only 0.35. These parameters are in the image-to-image module. The prompt correlation coefficient has been mentioned above, and the redrawing amplitude is a parameter that affects the degree of change in the image-to-image. The value of this parameter ranges from 0 to 1. The redrawing steps are the number of iterations when generating images in the image-to-image and are also a parameter that controls the similarity of the generated images. All of these can be directly adjusted

[0039] (8) Input the repaired portrait data into iclight to optimize the image lighting using the prompt.

[0040] (9) Through a custom node, Python code will be used in the node to encapsulate the algorithm in the way of Photoshop to adjust the tone of the portrait face and the overall image, mainly involving brightness, contrast, and detail enhancement, etc., to obtain the data of the specified model we need in the end. It can be manually operated with Photoshop or a custom node can be encapsulated and the operation can be performed using code to encapsulate the node.

[0041] InstantId: InstantId is an image generation technology based on the diffusion model, focusing on achieving zero-shot identity-preserving personalized image synthesis. This technology allows users to generate personalized images in multiple styles using only one facial image, while ensuring high fidelity. It extracts the identity embedding of the reference facial image to maintain the face details in the generated image, and at the same time uses a decoupled cross-attention mechanism to support the image as a visual cue.

[0042] Ipadapter: An Ipadapter is an adapter used to add image prompting capabilities to a pre-trained text-to-image diffusion model. It embeds image features into the pre-trained text-to-image diffusion model by using a decoupled cross-attention mechanism, thus realizing the image prompting function. The key design of the Ipadapter is the decoupled cross-attention mechanism, which separates the cross-attention layers of text features and image features, allowing the model to more effectively capture fine-grained features from image prompts.

[0043] Openpose: Openpose is a key point detection framework that can detect the key points of the human body in an image or video stream, including key points of the body, face, hands, and even feet.

[0044] iclight: The full name is "Imposing Consistent Light", which can perform relighting operations on the input image. Specifically, it includes generating corresponding background-matched lighting after removing the background, customizing lighting maps and illumination, and matching lighting for a fixed background image and foreground. It is particularly effective for realistic images. Currently, iclight supports two methods: text-guided and background-image-guided.

[0045] Example 1

[0046] As Figure 1 shown, this embodiment provides a method for generating specified model image data based on ComfyUI, including the following steps:

[0047] Step 1: Obtain the frontal portrait data of the uploaded image;

[0048] Step 2: Obtain the uploaded model reference pose image;

[0049] Step 3: ComfyUI calls the Ipadapter algorithm, adjusts the weight type to ease out, and sets the weight parameter weight of the Ipadapter algorithm to 0.5;

[0050] Step 4: ComfyUI calls the InstantId algorithm and sets the weight parameter weight of the InstantId algorithm to 0.6;

[0051] Step 5: ComfyUI calls the openpose algorithm, which is used for detecting the body key points of the model reference pose image and outputs a key point image;

[0052] Step 6: The frontal portrait data is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference pose map is processed by the openpose algorithm. Then, the processed results are uniformly input into the first image-to-image generation node of ComfyUI to generate portrait data;

[0053] Step 7: Perform face segmentation and eye detection on the portrait data, and perform face repair and eye repair through the mask redrawing of the second image-to-image generation node to obtain a repaired image;

[0054] Step 8: Input the repaired image into IC-Light for lighting optimization according to the set prompt words to obtain an optimized image.

[0055] In this embodiment, preferably, Step 6 is specifically that the frontal portrait data is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference pose map is processed by the openpose algorithm. Then, the processed results are uniformly input into the first image-to-image generation node of ComfyUI to generate portrait data; the first image-to-image generation node uses the img2img algorithm, and the set prompt word correlation coefficient is set to 4-9, and the redrawing amplitude is set to 0.75.

[0056] In this embodiment, preferably, Step 7 is specifically: perform face segmentation and eye detection on the portrait data, and perform face repair and eye repair through the mask redrawing of the second image-to-image generation node to obtain a repaired image; the second image-to-image generation node uses the img2img algorithm, the set prompt word correlation coefficient is set to 1, the redrawing steps are set to 4, and the redrawing amplitude is only 0.35.

[0057] In this embodiment, preferably, it further includes; Step 9: Set a custom node in ComfyUI, and the custom node is used to adjust brightness, transparency, and contrast; adjust the optimized image according to the custom node to obtain the final output image.

[0058] Based on the same inventive concept, the present application also provides a device corresponding to the method in Embodiment 1, as detailed in Embodiment 2.

[0059] Embodiment 2

[0060] As Figure 2 shown, in this embodiment, a device for generating specified model map data based on ComfyUI is provided, including:

[0061] A frontal face acquisition module for acquiring the frontal portrait data of the uploaded image;

[0062] A pose acquisition module for acquiring the model reference pose map of the uploaded image;

[0063] Set the first algorithm module. ComfyUI calls the Ipadapter algorithm, adjusts the weight type to ease out, and sets the weight parameter weight of the Ipadapter algorithm to 0.5;

[0064] Set the second algorithm module. ComfyUI calls the InstantId algorithm and sets the weight parameter weight of the InstantId algorithm to 0.6;

[0065] Set the third algorithm module. ComfyUI calls the openpose algorithm, which is used for detecting the body key points of the model reference pose map and outputs a key point map;

[0066] Generate image module. The frontal portrait data is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference pose map is processed by the openpose algorithm. Then the processed results are uniformly input into the first image generation node of ComfyUI to generate portrait data;

[0067] Image repair module. The portrait data is subjected to face segmentation and eye detection, and face repair and eye repair are performed through the mask redrawing of the second image generation node to obtain a repaired image;

[0068] Optimize image module. The repaired image is input into IC-Light, and lighting optimization is performed according to the set prompt words to obtain an optimized image.

[0069] In this embodiment, preferably, the generating image module is specifically: the frontal portrait data is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference pose map is processed by the openpose algorithm. Then the processed results are uniformly input into the first image generation node of ComfyUI to generate portrait data; the first image generation node adopts the img2img algorithm, and the set prompt word correlation coefficient is set to 4-9, and the redrawing amplitude is set to 0.75.

[0070] In this embodiment, preferably, the image repair module is specifically: the portrait data is subjected to face segmentation and eye detection, and face repair and eye repair are performed through the mask redrawing of the second image generation node to obtain a repaired image; the second image generation node adopts the img2img algorithm, the set prompt word correlation coefficient is set to 1, the redrawing steps are set to 4, and the redrawing amplitude is only 0.35.

[0071] In this embodiment, preferably, it further includes an image adjustment module. A custom node is set in ComfyUI, and the custom node is used to adjust brightness, transparency, and contrast; the optimized image is adjusted according to the custom node to obtain the final output image.

[0072] Since the device described in the second embodiment of the present invention is the device adopted for implementing the method of the first embodiment of the present invention, based on the method described in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, and thus will not be elaborated herein. Any device adopted for the method of the first embodiment of the present invention falls within the scope of protection of the present invention.

[0073] Although the specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A method for generating specified model image data based on ComfyUI, characterized in that: The steps include: Step 1: Get the front face portrait data of the uploaded image; Step 2: Obtain the uploaded model reference pose diagram; Step 3. ComfyUI calls the Ipadapter algorithm, adjusts the weight type to ease out, and sets the weight parameter weight of the Ipadapter algorithm to 0.5; Step 4. ComfyUI calls the InstantId algorithm and sets the weight parameter weight of the InstantId algorithm to 0.6; Step 5, ComfyUI calls the openpose algorithm, which is used to detect the key points of the body of the model with reference to the pose diagram and output the key point diagram; Step 6: The frontal portrait data is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference posture diagram is processed by the openpose algorithm, and then the processed results are uniformly passed to the first image generation node of ComfyUI to generate portrait data; Step 7: Perform facial segmentation and eye detection on the portrait data, perform facial repair and eye repair by redrawing the mask of the second image raw node, and obtain a repaired image; Step 8: Transfer the repaired image to IC-Light, perform lighting optimization according to the set prompt words, and obtain the optimized image.

2. A method for generating designated model image data based on ComfyUI according to claim 1, characterized in that: The step 6 specifically includes processing the frontal portrait data through the Ipadapter algorithm and the InstantId algorithm respectively, processing the model reference posture diagram through the openpose algorithm, and then transferring the processed results to the first image generation node of ComfyUI to generate portrait data; the first image generation node adopts the img2img algorithm, sets the prompt word correlation coefficient to 4-9, and sets the redrawing amplitude to 0.

75.

3. The method for generating designated model image data based on ComfyUI according to claim 1, characterized in that: The step 7 is specifically as follows: performing facial segmentation and eye detection on the portrait data, and performing facial and eye repair by mask redrawing of the second image raw node to obtain a repaired image; the second image raw node adopts the img2img algorithm, sets the prompt word correlation coefficient to 1, sets the redrawing step number to 4, and the redrawing amplitude is only 0.

35.

4. The method for generating designated model image data based on ComfyUI according to claim 1, characterized in that: It also includes: Step 9, setting a custom node in ComfyUI, where the custom node is used to adjust brightness, transparency and contrast; adjusting the optimized graph according to the custom node to obtain a final output graph.

5. A device for generating designated model image data based on ComfyUI, characterized in that: include: Get the front face module to obtain the front face portrait data of the uploaded image; Get the posture module to get the uploaded model reference posture diagram; Set the first algorithm module, ComfyUI calls the Ipadapter algorithm, adjusts the weight type to ease out, and sets the weight parameter weight of the Ipadapter algorithm to 0.5; Set the second algorithm module, ComfyUI calls the InstantId algorithm, and set the weight parameter weight of the InstantId algorithm to 0.6; Setting a third algorithm module, ComfyUI calls the openpose algorithm, the openpose algorithm is used to detect the key points of the body of the model with reference to the pose diagram, and outputs the key point diagram; Generate image module. The frontal portrait data is processed by Ipadapter algorithm and InstantId algorithm respectively. The model reference posture diagram is processed by openpose algorithm. Then the processed results are uniformly passed to the first image generation node of ComfyUI to generate portrait data. The image repair module performs facial segmentation and eye detection on the portrait data, and performs facial and eye repair by redrawing the mask of the second image generation node to obtain a repaired image; The image optimization module transfers the repaired image to IC-Light, performs lighting optimization according to the set prompt words, and obtains the optimized image.

6. The device for generating designated model image data based on ComfyUI according to claim 5, characterized in that: The image generation module is specifically as follows: the frontal portrait data is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference posture diagram is processed by the openpose algorithm, and then the processed results are uniformly passed to the first image generation node of ComfyUI to generate portrait data; the first image generation node adopts the img2img algorithm, sets the prompt word correlation coefficient to 4-9, and sets the redrawing amplitude to 0.

75.

7. The device for generating designated model image data based on ComfyUI according to claim 5, characterized in that: The image repair module specifically includes: performing facial segmentation and eye detection on the portrait data, performing facial repair and eye repair through mask redrawing of the second image generation node, and obtaining a repaired image; the second image generation node adopts the img2img algorithm, sets the prompt word correlation coefficient to 1, sets the redrawing step number to 4, and the redrawing amplitude is only 0.

35.

8. The device for generating designated model image data based on ComfyUI according to claim 5, characterized in that: It also includes an image adjustment module, which sets a custom node in ComfyUI, and the custom node is used to adjust brightness, transparency and contrast; the optimized image is adjusted according to the custom node to obtain a final output image.