Method, device and equipment for repairing face of model chart and medium

By detecting, enlarging, repainting and supplementing details of the model picture faces generated by the Stable Diffusion model, the problem of lack of face details and insufficient realism is solved, and the effect of improving the detailed performance and realism of the model picture face is achieved.

CN120070265APending Publication Date: 2025-05-30ZIXUN TECHNOLOGY (FUJIAN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510108247.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the model drawings generated by the Stable Diffusion model often lack details, resulting in insufficient realism and poor practicality.

Method used

The face area in the model image is intercepted through the face detection algorithm, enlarged to the most suitable resolution, and repainted and supplemented with the Stable Diffusion model and the face detail Lora model for redrawing and details. Finally, the repaired face image is reduced and pasted back to the original image, and processed through an image blur filter to improve the overall effect.

Benefits of technology

It effectively improves the detailed performance and realism of the faces of the model picture, so that the repaired pictures can be used directly, reducing the user's cost of use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070265A_ABST
    Figure CN120070265A_ABST
Patent Text Reader

Abstract

The invention provides a method, a device, equipment and a medium for repairing the face of a model chart, and the method comprises the steps: detecting a face region in the model chart through a face detection algorithm, obtaining a face rectangular frame, and intercepting a face region image from the model chart according to the face rectangular frame; magnifying the face region image to a set resolution to obtain a face magnified image; setting cue words, inputting the amplified face image into a Stable Diffusion diffusion model for redrawing, setting a redrawing amplitude, and performing image face generation control by adopting a face detail Lora model to obtain a face generation image; reducing the face generated image to be consistent with the face region image to obtain a face reduced image; pasting the reduced face image back to a set position in the model image to generate a fit image; and processing the fit image through an image fuzzy filter to obtain a result image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method, device, equipment and medium for repairing the face of a model picture. Background Art

[0002] In the prior art, the Stable Diffusion diffusion model is used to generate model pictures, and the Stable Diffusion diffusion model has an optimal resolution for drawing. Different models have different optimal resolutions for drawing. Since the face requires more details, when drawing the whole human body, the face often lacks details. E-commerce merchants often have poor practicability due to the lack of realism and skin details in the generated model pictures. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for repairing the face of a model picture, which can improve the detail performance and realism of the model face by post-processing the model pictures generated by the stablediffusion model.

[0004] In a first aspect, the present invention provides a method for repairing the face of a model picture, which is used to process the model picture generated by a diffusion model, and specifically includes the following steps:

[0005] Step 1: Detect the face area in the model picture through a face detection algorithm to obtain a face rectangle, and intercept the face area image from the model picture according to the face rectangle; Enlarge the face area image to a set resolution to obtain an enlarged face image;

[0006] Step 2: Set a prompt, input the enlarged face image into the Stable Diffusion diffusion model for redrawing, set the redrawing amplitude, and use the face detail Lora model to control the generation of the face in the image to obtain a face generation image;

[0007] Step 3: Reduce the face generation image to be the same as the face area image to obtain a reduced face image; Paste the reduced face image back to a set position in the model picture to generate a fitting picture;

[0008] Step 4: Process the fitting picture through an image blur filter to obtain a result picture.

[0009] In a second aspect, the present invention provides a device for repairing the face of a model picture, which is used to process the model picture generated by a diffusion model, and includes:

[0010] The cropping and enlarging module detects the face region in the model image through a face detection algorithm to obtain a face rectangular frame, and crops the face region image from the model image according to the face rectangular frame; enlarges the face region image to a set resolution to obtain an enlarged face image;

[0011] The face generation module sets a prompt, inputs the enlarged face image into the Stable Diffusion diffusion model for redrawing, sets the redrawing amplitude, and uses the face detail Lora model to control the image face generation to obtain a face generated image;

[0012] The fitting module shrinks the face generated image to be the same as the face region image to obtain a shrunk face image; pastes the shrunk face image back to a set position in the model image to generate a fitted image;

[0013] The processing module processes the fitted image through an image blur filter to obtain a result image.

[0014] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect is implemented.

[0015] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect is implemented.

[0016] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:

[0017] The present invention uses the Stable Diffusion diffusion model to draw the face region at the most suitable resolution, and combines the prompt and the Lora model to supplement the facial details, which can effectively solve most of the problems of lack of facial realism and insufficient details; enabling the repaired pictures to be directly used, reducing the user's usage cost.

[0018] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically gives the specific embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.

[0020] Figure 1 It is a flowchart in the method of Embodiment 1 of the present invention;

[0021] Figure 2 This is a schematic structural diagram of the device in the second embodiment of the present invention. Detailed implementation mode

[0022] The overall idea of the technical solution in the embodiment of this application is as follows:

[0023] The Stable Diffusion diffusion model, Latent is the latent representation space of the image in the Stable Diffusion diffusion model, and Controlnet is the control generation network in the Stable Diffusion diffusion model, which guides the generation of corresponding images;

[0024] The Dlib face detection algorithm is a face detection algorithm based on feature recognition, which distinguishes faces from other objects by extracting features such as the texture, shape, and color of the face from the image;

[0025] The Lanczos algorithm is a commonly used image scaling algorithm, and its basic principle is to calculate the pixel values of the target image according to the pixel values of the original image through certain rules;

[0026] The Lora model, full name Low-Rank Adaptation of Large Language Models, is a low-rank adaptation technology for fine-tuning large language models. In StableDiffusion, it is often used to fine-tune the output results of large models to improve details.

[0027] The implementation is as follows:

[0028] (1) The input image in this embodiment is an unrealistic face image generated by the StableDiffusion model. Here, the input image is called Image One. The Dlib face detection algorithm is used to detect the regional position of the face. Four coordinate points are detected to determine the area that needs to be redrawn. By using the method of image cropping, the face area image that needs to be drawn is cropped from the original image, and the cropped face area image is called Image Two.

[0029] (2) Image Two is enlarged to a resolution of 640*640 using the Lanczos algorithm to obtain Image Three.

[0030] (3) Image Three is redrawn using the StableDiffusion model. The redrawing amplitude is selected as 0.55, and the drawing prompts are high quality, best quality, 4K, professional, facial feature. Select to import the self-developed Lora for adding details to further add details, and obtain Image Four through generation.

[0031] (4) Resize Image Four to the size of Image Two using the Lanczos algorithm to obtain Image Five.

[0032] (5) Paste Image Five back onto Image One according to four coordinate points to obtain Image Six.

[0033] (6) For the edges of Image Six, use the Gaussian blur algorithm to process and cover up the traces of pasting back to obtain the final result image.

[0034] Example One

[0035] As Figure 1 shown, this example provides a method for repairing the face of a model image, which is used to process the model image generated by a diffusion model, and specifically includes the following steps:

[0036] Step 1: Detect the face area in the model image through a face detection algorithm to obtain a face rectangle, and intercept the face area image from the model image according to the face rectangle; Enlarge the face area image to a set resolution to obtain an enlarged face image;

[0037] Step 2: Set a prompt, input the enlarged face image into the Stable Diffusion diffusion model for redrawing, set the redrawing amplitude, and use the face detail Lora model to control the generation of the face in the image to obtain a face generation image;

[0038] Step 3: Shrink the face generation image to be the same as the face area image to obtain a shrunk face image; Paste the shrunk face image back to a set position in the model image to generate a pasted image;

[0039] Step 4: Process the pasted image through an image blur filter to obtain a result image.

[0040] In this example, preferably, Step 1 is specifically: Detect the face area in the model image through the Dlib face detection algorithm to obtain a face rectangle, the face rectangle includes four coordinate points, and intercept the face area image from the model image according to the four coordinate points; Enlarge the face area image to a resolution of 640*640 using the Lanczos algorithm to obtain an enlarged face image.

[0041] In this example, preferably, Step 3 is specifically: Shrink the face generation image to be the same as the face area image using the Lanczos algorithm to obtain a shrunk face image; Paste the shrunk face image back to a set position in the model image according to the four coordinate points to generate a pasted image.

[0042] In this example, preferably, the face detail Lora model is specifically:

[0043] Obtain the picture data of a set model, and perform data cleaning to obtain cleaned data. The picture data includes close-up face pictures;

[0044] If the number of pictures in the cleaned data is greater than the set threshold, and the cleaned data includes: close-up face pictures, half-body pictures, and full-body pictures; otherwise, perform data augmentation on the cleaned data to obtain augmented data. The augmented data includes: close-up face pictures, half-body pictures, and full-body pictures, and the quantity ratio is 2:1:1; the number of pictures in the augmented data is greater than the set threshold;

[0045] Process each picture in the cleaned data or the augmented data to ensure that the long side of each picture is a set value to obtain training data;

[0046] Format and label the pictures in the training data through a multi-modal model to form a label file. The formatted format is: gender, skin, eye expression, hairstyle, lip shape, expression, accessories, clothing, and background;

[0047] Set the number of training rounds, learning rate, and training image pixels, and then input the training data and the label file for model training to obtain a face detail Lora model;

[0048] The data augmentation includes: generating model pictures. The specific method for generating model pictures is:

[0049] Obtain a close-up face image from the picture data, and this close-up face image is a frontal portrait; obtain the uploaded model reference half-body image or the model reference half-body image; ComfyUI calls the Ipadapter algorithm, adjusts the weight type to ease out, and sets the weight parameter weight of the Ipadapter algorithm to 0.5; ComfyUI calls the InstantId algorithm and sets the weight parameter weight of the InstantId algorithm to 0.6; ComfyUI calls the openpose algorithm, and the openpose algorithm is used for body key point detection of the model reference half-body image or the model reference half-body image, and outputs a key point image; the close-up face image is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference half-body image or the model reference half-body image is processed by the openpose algorithm, and then the processed results are uniformly input into the first image-to-image generation node of ComfyUI to generate portrait data; the first image-to-image generation node uses the img2img algorithm, sets the prompt word correlation coefficient to 7, and the redrawing amplitude to 0.75; perform face segmentation and eye detection on the portrait data, and perform face repair and eye repair through mask redrawing of the second image-to-image generation node to obtain a repaired image; the second image-to-image generation node uses the img2img algorithm, sets the prompt word correlation coefficient to 1, the redrawing steps to 20, and the redrawing amplitude to only 0.45; input the repaired image into IC-Light, and perform lighting optimization according to the set prompt word to obtain an optimized image; set a custom node in ComfyUI, and the custom node is used to adjust brightness, transparency, and contrast; adjust the optimized image according to the custom node to obtain the required half-body image or full-body image.

[0050] Based on the same inventive concept, the present application also provides an apparatus corresponding to the method in Embodiment 1. For details, see Embodiment 2.

[0051] Embodiment 2

[0052] As Figure 2 shown, in this embodiment, an apparatus for repairing the face of a model image is provided, which is used to process the model image generated by a diffusion model, and includes:

[0053] A cropping and enlarging module, which detects the face area in the model image through a face detection algorithm to obtain a face rectangular frame, and crops the face area image from the model image according to the face rectangular frame; enlarges the face area image to a set resolution to obtain an enlarged face image;

[0054] Generate a face module, set a prompt, input the enlarged face image into the Stable Diffusion diffusion model for redrawing, set the redrawing amplitude, and use the face detail Lora model to control the image face generation to obtain a face generation image;

[0055] A fitting module that reduces the face generation image to be the same as the face area image to obtain a reduced face image; pastes the reduced face image back to a set position in the model image to generate a fitted image;

[0056] A processing module that processes the fitted image through an image blur filter to obtain a result image.

[0057] In this embodiment, preferably, the cropping and enlarging module is specifically: detecting the face area in the model image through the Dlib face detection algorithm to obtain a face rectangular frame, the face rectangular frame includes four coordinate points, and cropping the face area image from the model image according to the four coordinate points; enlarging the face area image to a resolution of 640*640 through the Lanczos algorithm to obtain an enlarged face image.

[0058] In this embodiment, preferably, the fitting module is specifically: reducing the face generation image to be the same as the face area image through the Lanczos algorithm to obtain a reduced face image; pasting the reduced face image back to a set position in the model image according to the four coordinate points to generate a fitted image.

[0059] In this embodiment, preferably, the face detail Lora model is specifically:

[0060] Obtain the picture data of a set model and perform data cleaning to obtain cleaned data, the picture data includes face close-up pictures;

[0061] If the number of pictures in the cleaned data is greater than the set threshold, and the cleaned data includes: face close-up pictures, half-body pictures, and full-body pictures; otherwise, perform data augmentation on the cleaned data to obtain augmented data, the augmented data includes: face close-up pictures, half-body pictures, and full-body pictures, and the quantity ratio is 2:1:1; the number of pictures in the augmented data is greater than the set threshold;

[0062] Process each picture in the cleaned data or the augmented data to ensure that the long side of each picture is a set value to obtain training data;

[0063] Format and label the pictures in the training data through a multi-modal model to form a label file, and the formatted format is: gender, skin, eye expression, hairstyle, lip shape, expression, accessories, clothing, and background;

[0064] Set the number of training epochs, learning rate, and the pixel size of training images. Then, input the training data and label files to perform model training and obtain a face details Lora model.

[0065] The data augmentation includes: generating a model image, and specifically, the generation of the model image is as follows:

[0066] Obtain a face close-up image from the picture data, and this face close-up image is a frontal portrait; obtain the uploaded model reference half-body image or model reference half-body image; ComfyUI calls the Ipadapter algorithm, adjusts the weight type to ease out, and sets the weight parameter weight of the Ipadapter algorithm to 0.5; ComfyUI calls the InstantId algorithm and sets the weight parameter weight of the InstantId algorithm to 0.6; ComfyUI calls the openpose algorithm, and the openpose algorithm is used for body key point detection of the model reference half-body image or model reference half-body image, and outputs a key point image; the face close-up image is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference half-body image or model reference half-body image is processed by the openpose algorithm. Then, the processed results are uniformly input into the first image generation node of ComfyUI to generate portrait data; the first image generation node uses the img2img algorithm, sets the prompt word correlation coefficient to 7, and the redrawing amplitude to 0.75; perform face segmentation and eye detection on the portrait data, and perform face repair and eye repair through mask redrawing of the second image generation node to obtain a repaired image; the second image generation node uses the img2img algorithm, sets the prompt word correlation coefficient to 1, the redrawing steps to 20, and the redrawing amplitude to only 0.45; input the repaired image into IC-Light for lighting optimization according to the set prompt word to obtain an optimized image; set a custom node in ComfyUI, and the custom node is used to adjust brightness, transparency, and contrast; adjust the optimized image according to the custom node to obtain the required half-body image or full-body image.

[0067] Since the device introduced in the second embodiment of the present invention is the device used to implement the method of the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated here. Any device used in the method of the first embodiment of the present invention falls within the scope of protection of the present invention.

[0068] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to the first embodiment, as shown in detail in the third embodiment.

[0069] Embodiment Three

[0070] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, any implementation manner in Embodiment 1 can be implemented.

[0071] Since the electronic device introduced in this embodiment is the device used to implement the method in Embodiment 1 of this application, based on the method introduced in Embodiment 1 of this application, those skilled in the art can understand the specific implementation manner and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of this application will not be introduced in detail here. As long as the device used by those skilled in the art to implement the method in the embodiments of this application belongs to the scope protected by this application.

[0072] Based on the same inventive concept, this application provides a storage medium corresponding to Embodiment 1, as detailed in Embodiment 4.

[0073] Embodiment 4

[0074] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any implementation manner in Embodiment 1 can be implemented.

[0075] The technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0076] In this embodiment, the Stable Diffusion diffusion model is used to draw the face area at the most suitable resolution, and the prompt words and the Lora model are used to supplement the facial details, which can effectively solve most of the problems of lack of facial realism and insufficient details; the repaired pictures can be directly used, reducing the user's usage cost.

[0077] The face detail Lora model in this embodiment not only significantly enhances the consistency of the generated character images, ensures the stability and coherence of the model characteristics, but also greatly improves the clarity of the images, making the detail performance more accurate and vivid.

[0078] This embodiment improves the training strategy and parameter adjustment, enhances the generalization ability of the model in diverse scenarios, enables the model to adapt to a wider range of application requirements, and improves the practicality and flexibility of the model.

[0079] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0080] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0081] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0082] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0083] Although the specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope of the claims of the present invention.

Claims

1. A method for repairing a model's face, characterized in that: The model graph used to process the diffusion model generation includes the following steps: Step 1: Detect the face area in the model image by using a face detection algorithm to obtain a face rectangular frame, and capture a face area image from the model image according to the face rectangular frame; enlarge the face area image to a set resolution to obtain an enlarged face image; Step 2: Set the prompt word, input the face enlargement image into the Stable Diffusion diffusion model for redrawing, set the redrawing amplitude, and use the face detail Lora model to control the image face generation to obtain the face generation map; Step 3, reducing the face generation image to be consistent with the face area image to obtain a reduced face image; pasting the reduced face image back to a set position in the model image to generate a pasted image; Step 4: Process the fitted image through an image blur filter to obtain a result image.

2. A method for repairing a model's face according to claim 1, characterized in that: The step 1 is specifically as follows: the face area in the model image is detected by the Dlib face detection algorithm to obtain a face rectangular frame, wherein the face rectangular frame includes four coordinate points, and a face area image is captured from the model image according to the four coordinate points; the face area image is enlarged to a resolution of 640*640 by the Lanczos algorithm to obtain an enlarged face image.

3. The method for repairing a model's face according to claim 1, characterized in that: The step 3 is specifically as follows: reducing the face generation image to be consistent with the face area image through the Lanczos algorithm to obtain a reduced face image; pasting the reduced face image back to the set position in the model image according to four coordinate points to generate a pasted image.

4. The method for repairing a model's face according to claim 1, characterized in that: The face detail Lora model is specifically: Acquire image data of a set model, and perform data cleaning to obtain cleaned data, wherein the image data includes a close-up image of a face; If the number of images in the cleaned data is greater than the set threshold, and the cleaned data includes: close-up images of faces, half-body images, and full-body images; if not, the cleaned data is enhanced to obtain enhanced data, the enhanced data includes: close-up images of faces, half-body images, and full-body images, and the ratio of the number of images is 2:1:1; the number of images in the enhanced data is greater than the set threshold; Process each image in the cleaned data or enhanced data to ensure that the long side of each image is the set value, and obtain training data; The images in the training data are formatted and labeled through the multimodal model to form a label file, wherein the format is: gender, skin, eyes, hairstyle, mouth shape, expression, accessories, clothing and background; Set the number of training rounds, learning rate, and training image pixels, then input the training data and label file to perform model training to obtain the face detail Lora model; The data enhancement includes: generating a model graph, wherein generating the model graph is specifically: A close-up face image is obtained from the image data, which is a frontal portrait; an uploaded model reference half-length image or a model reference half-length image is obtained; ComfyUI calls the Ipadapter algorithm, adjusts the weight type to ease out, and sets the weight parameter weight of the Ipadapter algorithm to 0.5; ComfyUI calls the InstantId algorithm, and sets the weight parameter weight of the InstantId algorithm to 0.6; ComfyUI calls the openpose algorithm, which is used for body key point detection of the model reference half-length image or the model reference half-length image, and outputs a key point image; the close-up face image is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference half-length image or the model reference half-length image is processed by the openpose algorithm, and then the processed results are uniformly passed to the first image generation node of ComfyUI to generate portrait data; the first The first image raw node adopts the img2img algorithm, sets the prompt word correlation coefficient to 7, and the redrawing amplitude to 0.75; performs facial segmentation and eye detection on the portrait data, and performs facial and eye repair through mask redrawing of the second image raw node to obtain a repaired image; the second image raw node adopts the img2img algorithm, sets the prompt word correlation coefficient to 1, the redrawing step number to 20, and the redrawing amplitude is only 0.45; the repaired image is passed into IC-Light, and the lighting is optimized according to the set prompt words to obtain an optimized image; a custom node is set in ComfyUI, and the custom node is used to adjust the brightness, transparency and contrast; the optimized image is adjusted according to the custom node to obtain the required half-body or full-body image.

5. A device for repairing a model's face, characterized in that: Model graphs generated for processing diffusion models, including: The interception and enlargement module detects the face area of ​​the model image by using a face detection algorithm to obtain a face rectangular frame, intercepts a face area image from the model image according to the face rectangular frame; enlarges the face area image to a set resolution to obtain an enlarged face image; Generate a face module, set prompt words, input the face enlargement image into the Stable Diffusion diffusion model for redrawing, set the redrawing amplitude, and use the face detail Lora model to control the image face generation to obtain the face generation map; A fitting module is used to reduce the face generation image to be consistent with the face region image to obtain a reduced face image; and the reduced face image is pasted back to a set position in the model image to generate a fitted image; The processing module processes the fitting image through an image blur filter to obtain a result image.

6. The device for repairing the face of a model according to claim 5, characterized in that: The capture and magnification module is specifically as follows: the face area in the model image is detected by the Dlib face detection algorithm to obtain a face rectangular frame, wherein the face rectangular frame includes four coordinate points, and the face area image is captured from the model image according to the four coordinate points; the face area image is enlarged to a resolution of 640*640 by the Lanczos algorithm to obtain an enlarged face image.

7. The device for repairing the face of a model according to claim 5, characterized in that: The fitting module specifically comprises: reducing the face generation image to be consistent with the face area image through the Lanczos algorithm to obtain a reduced face image; pasting the reduced face image back to the set position in the model image according to four coordinate points to generate a fitting image.

8. The device for repairing the face of a model according to claim 5, characterized in that: The face detail Lora model is specifically: Acquire image data of a set model, and perform data cleaning to obtain cleaned data, wherein the image data includes a close-up image of a face; If the number of images in the cleaned data is greater than the set threshold, and the cleaned data includes: close-up images of faces, half-body images, and full-body images; if not, the cleaned data is enhanced to obtain enhanced data, the enhanced data includes: close-up images of faces, half-body images, and full-body images, and the ratio of the number of images is 2:1:1; the number of images in the enhanced data is greater than the set threshold; Process each image in the cleaned data or enhanced data to ensure that the long side of each image is the set value, and obtain training data; The images in the training data are formatted and labeled through the multimodal model to form a label file, wherein the format is: gender, skin, eyes, hairstyle, mouth shape, expression, accessories, clothing and background; Set the number of training rounds, learning rate, and training image pixels, then input the training data and label file to perform model training to obtain the face detail Lora model; The data enhancement includes: generating a model graph, wherein generating the model graph is specifically: A close-up face image is obtained from the image data, which is a frontal portrait; an uploaded model reference half-length image or a model reference half-length image is obtained; ComfyUI calls the Ipadapter algorithm, adjusts the weight type to ease out, and sets the weight parameter weight of the Ipadapter algorithm to 0.5; ComfyUI calls the InstantId algorithm, and sets the weight parameter weight of the InstantId algorithm to 0.6; ComfyUI calls the openpose algorithm, which is used for body key point detection of the model reference half-length image or the model reference half-length image, and outputs a key point image; the close-up face image is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference half-length image or the model reference half-length image is processed by the openpose algorithm, and then the processed results are uniformly passed to the first image generation node of ComfyUI to generate portrait data; the first The first image raw node adopts the img2img algorithm, sets the prompt word correlation coefficient to 7, and the redrawing amplitude to 0.75; performs facial segmentation and eye detection on the portrait data, and performs facial and eye repair through mask redrawing of the second image raw node to obtain a repaired image; the second image raw node adopts the img2img algorithm, sets the prompt word correlation coefficient to 1, the redrawing step number to 20, and the redrawing amplitude is only 0.45; the repaired image is passed into IC-Light, and the lighting is optimized according to the set prompt words to obtain an optimized image; a custom node is set in ComfyUI, and the custom node is used to adjust the brightness, transparency and contrast; the optimized image is adjusted according to the custom node to obtain the required half-body or full-body image.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Cited By

  • Texture map generation method and device

    CN120563703A