Method, device and equipment for training exclusive model by Lora and medium

By cleaning and enhancing the data of Lora training exclusive models, and using multimodal models for formatting and marking, the problems of poor portrait generation and poor consistency in the existing technology are solved, and high-quality and stable image generation is achieved.

CN120070628APending Publication Date: 2025-05-30ZIXUN TECHNOLOGY (FUJIAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108342.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When training exclusive models of characters through Lora, there are problems such as incomplete data sets, irregular labels, poor portrait generation effects, overfitting models, distortion of portraits, poor consistency between portrait generation and training data, insufficient clarity, and poor generalization of models.

Method used

By obtaining the image data of the set model, data cleaning and enhancement are performed to ensure the integrity and consistency of the data set. Then, the processed data is formatted and marked by multimodal model, appropriate training parameters are set, model training is performed to obtain an exclusive model model.

Benefits of technology

It significantly enhances the consistency of generated character images, ensures the stability and coherence of model features, greatly improves the clarity of the image, makes the details more accurate and vivid, and improves the generalization ability and practicality of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070628A_ABST
    Figure CN120070628A_ABST
Patent Text Reader

Abstract

The invention provides a method, a device, equipment and a medium for Lora training of an exclusive model, and the method comprises the steps: obtaining the picture data of a set model, and carrying out the data cleaning, and obtaining the cleaning data; if the number of the pictures in the cleaning data is greater than a set threshold value, entering the next step; if the number of the pictures in the cleaned data is smaller than or equal to a set threshold value, data enhancement is conducted on the cleaned data, enhanced data is obtained, and the number of the pictures in the enhanced data is larger than the set threshold value; each picture in the enhanced data or the cleaned data is processed, it is guaranteed that the long edge of each picture is 1024, and training data is obtained; formatting and labeling pictures in the training data through a multi-modal model to form a label file; according to the method, the training round number, the learning rate and the training image pixels are set, the training data and the label file are input, model training is carried out, and the exclusive model is obtained, so that the consistency of the generated character image is remarkably enhanced, and the stability and the continuity of model features are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method, device, equipment and medium for training an exclusive model for Lora. Background Art

[0002] With the rapid development of the Internet, image processing technology has become increasingly mature. In order to efficiently draw and process digital image works with personalized features, a large number of artificial intelligence technologies have emerged to assist humans in completing image drawing work, and image automatic generation technology has received more and more attention and research. The generation quality has met the needs of many different industries for content production. Especially in the portrait generation scenario, users hope to be able to automatically generate specific images with the help of image generation technology while maintaining their own facial features.

[0003] When training an exclusive model for a person through Lora currently, there are often the following problems: 1. The dataset is not perfect and there is no complete set of pictures for training; 2. The labels are not standardized and the dataset is not unified; 3. The generated portrait effect is poor, the model is overfitted, and the portrait is distorted; 4. The consistency between the generated portrait and the training data is not strong, and the clarity is poor; 5. The generalization ability of the model is poor, and the model fails to achieve the expected effect when combined with other usage scenarios, and many other problems. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for training an exclusive model for Lora, which significantly enhances the consistency of the generated human images, ensures the stability and coherence of the model features, greatly improves the clarity of the images, and makes the detail performance more accurate and vivid.

[0005] In a first aspect, the present invention provides a method for training an exclusive model for Lora, including the following steps:

[0006] Step 1: Obtain the picture data of a set model and perform data cleaning to obtain cleaned data, where the picture data includes close-up pictures of the face;

[0007] Step 2: If the number of pictures in the cleaned data is greater than a set threshold, then enter Step 3; if the number of pictures in the cleaned data is less than or equal to the set threshold, then perform data augmentation on the cleaned data to obtain augmented data, where the number of pictures in the augmented data is greater than the set threshold, and enter Step 3;

[0008] Step 3: Process each picture in the augmented data or the cleaned data to ensure that the long side of each picture is a set value to obtain training data;

[0009] Step 4: Format and label the pictures in the training data through a multi-modal model to form a label file;

[0010] Step 5: Set the number of training epochs, learning rate, and training image pixels. Then input the training data and label file to perform model training and obtain an exclusive model.

[0011] In a second aspect, the present invention provides an apparatus for training an exclusive model using LoRA, including:

[0012] A data cleaning module that acquires picture data of a set model and performs data cleaning to obtain cleaned data. The picture data includes close-up face pictures;

[0013] An enhanced data module. If the number of pictures in the cleaned data is greater than a set threshold, it enters the data processing module. If the number of pictures in the cleaned data is less than or equal to the set threshold, the cleaned data is enhanced to obtain enhanced data. The number of pictures in the enhanced data is greater than the set threshold, and then it enters Step 3;

[0014] A data processing module that processes each picture in the enhanced data or the cleaned data to ensure that the long side of each picture is a set value, thereby obtaining training data;

[0015] A labeling module that formats and labels the pictures in the training data through a multi-modal model to form a label file;

[0016] A training module that sets the number of training epochs, learning rate, and training image pixels. Then it inputs the training data and the label file to perform model training and obtain an exclusive model.

[0017] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method described in the first aspect.

[0018] In a fourth aspect, the present invention provides a computer-readable storage medium with a computer program stored thereon. When the program is executed by a processor, it implements the method described in the first aspect.

[0019] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:

[0020] The present invention not only significantly enhances the consistency of generated character images, ensures the stability and coherence of model features, but also greatly improves the clarity of images, making the detail performance more accurate and vivid. In addition, the present invention improves the training strategy and parameter adjustment, enhances the generalization ability of the model in diverse scenarios, enables the model to adapt to a wider range of application requirements, and improves the practicality and flexibility of the model.

[0021] The above description is only an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention, it can be implemented according to the content of the specification. And in order to make the above and other objects, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. Brief Description of the Drawings

[0022] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.

[0023] Figure 1 It is a flowchart of the method in Embodiment 1 of the present invention;

[0024] Figure 2 It is a structural schematic diagram of the device in Embodiment 2 of the present invention. Detailed Embodiments

[0025] The overall idea of the technical solution in the embodiments of the present application is as follows:

[0026] Data cleaning: Data cleaning is an important step in the training of deep learning models. The quality of data cleaning plays a decisive role in the effect of the model. Especially in the Lora fine-tuning model, poor data is extremely likely to cause the model to be unfitted.

[0027] Image enhancement: It is a commonly used technology in machine learning and deep learning, especially in the field of image processing. It involves creating new, modified versions from the original training data to increase the diversity and quantity of the dataset, thereby improving the generalization ability and robustness of the model.

[0028] Lossless image scaling: It refers to the process of adjusting the image size without losing any image details and quality. Compared with traditional scaling methods, lossless scaling technology aims to maintain the original visual information of the image and avoid blurring, pixelation or other distortion phenomena caused by scaling.

[0029] Formatting and labeling: Formatting and labeling is a method of organizing and managing data. In the fields of machine learning and deep learning, it involves labeling data with specific categories or labels so that the model can recognize and learn.

[0030] For data cleaning, screen the set of image data, and screen the data with relatively high facial consistency. Especially pay attention to screening out the data with large changes in facial features and expressions even for the same person to ensure high consistency of facial features. Data cleaning can be carried out manually, and there must be a close-up of a face in the data, and there is a frontal portrait of the face in this close-up of the face;

[0031] Data augmentation: For the portrait data used in training, it is necessary to ensure that it covers close-up face pictures, half-body pictures, and full-body pictures, and the ratio of the number of pictures approaches 2:1:1. If the data is insufficient, data augmentation is performed through means such as cropping, flipping, and image filling to ensure that the number of training data images is greater than 50.

[0032] During the training process, conventional graphics cards only support the training of images with a size of 1024. To ensure the quality of the training data, which are all carefully selected high-definition data, a written ps automation script is required to perform bicubic sharpening interpolation scaling on the images. When shrinking the images, it can better maintain the sharpness of the images, achieving lossless scaling to a certain extent, ensuring that the training data has a long side of 1024. In this way, during the model training process, the automatic scaling in the training script will not be triggered, guaranteeing the high quality of the training images.

[0033] Call the large language model or multi-modal model to format and label the scaled images. The specific format is: gender, skin, eye expression, hairstyle, lip shape, expression, accessories, clothing, background. Formatting and labeling can greatly enhance the role of prompts in the use of the LoRA model, and the finer the labeling, the better.

[0034] Model training parameter setting and adjustment: Set the total number of training epochs to 6000, the learning rate to 0.0001, the training image size to 1024*1024, set the model path, dataset configuration file path, and model output path, etc. Run the script to start the model training with one key. The LoRA exclusive model for the model is obtained for use when generating pictures with the diffusion model, which can ensure that the face in the generated picture is a specific face.

[0035] Example 1

[0036] As Figure 1 shown, this example provides a method for training an exclusive model for LoRA, including the following steps:

[0037] Step 1: Obtain the picture data of a set model and perform data cleaning to obtain cleaned data. The picture data includes close-up face pictures.

[0038] Step 2: If the number of pictures in the cleaned data is greater than the set threshold, proceed to Step 3; if the number of pictures in the cleaned data is less than or equal to the set threshold, perform data augmentation on the cleaned data to obtain augmented data. The number of pictures in the augmented data is greater than the set threshold, and then proceed to Step 3.

[0039] Step 3: Process each picture in the augmented data or cleaned data to ensure that the long side of each picture is a set value to obtain training data. This set value is 1024.

[0040] Step 4: Format and label the pictures in the training data through a multi-modal model to form a label file;

[0041] Step 5: Set the number of training epochs, learning rate, and training image pixels, and then input the training data and the label file for model training to obtain an exclusive model.

[0042] In this embodiment, preferably, Step 2 is specifically as follows: If the number of pictures in the cleaned data is greater than the set threshold, and the cleaned data includes: face close-up pictures, half-body pictures, and full-body pictures, then proceed to Step 3; otherwise, perform data augmentation on the cleaned data to obtain augmented data, where the augmented data includes: face close-up pictures, half-body pictures, and full-body pictures, and the quantity ratio is 2:1:1; if the number of pictures in the augmented data is greater than the set threshold, proceed to Step 3.

[0043] In this embodiment, preferably, the data augmentation includes: generating model pictures, and the specific method for generating model pictures is:

[0044] Obtain a close-up face image from the picture data, and this close-up face image is a frontal portrait; obtain the uploaded model reference half-body image or the model reference half-body image; ComfyUI calls the Ipadapter algorithm, adjusts the weight type to ease out, and sets the weight parameter weight of the Ipadapter algorithm to 0.5; ComfyUI calls the InstantId algorithm and sets the weight parameter weight of the InstantId algorithm to 0.6; ComfyUI calls the openpose algorithm, and the openpose algorithm is used for body key point detection of the model reference half-body image or the model reference half-body image, and outputs a key point image; the close-up face image is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference half-body image or the model reference half-body image is processed by the openpose algorithm, and then the processed results are uniformly input into the first image generation node of ComfyUI to generate portrait data; the first image generation node adopts the img2img algorithm, sets the prompt word correlation coefficient to 7, and sets the redrawing amplitude to 0.75; perform face segmentation and eye detection on the portrait data, and perform face repair and eye repair through mask redrawing of the second image generation node to obtain a repaired image; the second image generation node adopts the img2img algorithm, sets the prompt word correlation coefficient to 1, sets the redrawing steps to 20, and the redrawing amplitude is only 0.45; input the repaired image into IC-Light, and perform light optimization according to the set prompt word to obtain an optimized image; set a custom node in ComfyUI, and the custom node is used to adjust brightness, transparency and contrast; adjust the optimized image according to the custom node to obtain the required half-body image or full-body image; through the above method, a half-body image or a full-body image can be generated for later model training;

[0045] In this embodiment, preferably, step 4 is specifically: formatting and tagging the pictures in the training data through a multi-modal model to form a tag file, and the formatting format is: gender, skin, eye expression, hairstyle, lip shape, expression, accessories, clothing, and background.

[0046] Based on the same inventive concept, the present application also provides a device corresponding to the method in Embodiment 1, as detailed in Embodiment 2.

[0047] Embodiment 2

[0048] As Figure 2 shown, in this embodiment, a device for a dedicated model trained by Lora is provided, including:

[0049] A data cleaning module, which obtains the picture data of a set model and performs data cleaning to obtain cleaned data, and the picture data includes a close-up face image;

[0050] Enhanced data module. If the number of pictures in the cleaned data is greater than the set threshold, enter the processed data module. If the number of pictures in the cleaned data is less than or equal to the set threshold, perform data enhancement on the cleaned data to obtain enhanced data. The number of pictures in the enhanced data is greater than the set threshold, and enter step 3.

[0051] Processed data module. Process each picture in the enhanced data or the cleaned data to ensure that the long side of each picture is the set value, and obtain training data. The set value is 1024.

[0052] Labeling module. Format and label the pictures in the training data through a multi-modal model to form a label file.

[0053] Training module. Set the number of training rounds, learning rate, and training image pixels, and then input the training data and the label file for model training to obtain a dedicated model of the model.

[0054] In this embodiment, preferably, the enhanced data module is specifically: if the number of pictures in the cleaned data is greater than the set threshold, and the cleaned data includes: close-up face pictures, half-body pictures, and full-body pictures, enter the processed data module; otherwise, perform data enhancement on the cleaned data to obtain enhanced data. The enhanced data includes: close-up face pictures, half-body pictures, and full-body pictures, and the quantity ratio is 2:1:1. The number of pictures in the enhanced data is greater than the set threshold, and enter step 3.

[0055] In this embodiment, preferably, the data enhancement includes: generating model pictures, and the specific method for generating model pictures is:

[0056] Obtain a close-up face image from the picture data, and this close-up face image is a frontal portrait; obtain the uploaded model reference half-body image or the model reference half-body image; ComfyUI calls the Ipadapter algorithm, adjusts the weight type to ease out, and sets the weight parameter weight of the Ipadapter algorithm to 0.5; ComfyUI calls the InstantId algorithm and sets the weight parameter weight of the InstantId algorithm to 0.6; ComfyUI calls the openpose algorithm, and the openpose algorithm is used for body key point detection of the model reference half-body image or the model reference half-body image, and outputs a key point image; the close-up face image is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference half-body image or the model reference half-body image is processed by the openpose algorithm, and then the processed results are uniformly input into the first image generation node of ComfyUI to generate portrait data; the first image generation node adopts the img2img algorithm, sets the prompt word correlation coefficient to 7, and the redrawing amplitude to 0.75; perform face segmentation and eye detection on the portrait data, and perform face repair and eye repair through mask redrawing of the second image generation node to obtain a repaired image; the second image generation node adopts the img2img algorithm, sets the prompt word correlation coefficient to 1, the redrawing steps to 20, and the redrawing amplitude to only 0.45; input the repaired image into IC-Light, and perform lighting optimization according to the set prompt word to obtain an optimized image; set a custom node in ComfyUI, and the custom node is used to adjust brightness, transparency, and contrast; adjust the optimized image according to the custom node to obtain the required half-body image or full-body image.

[0057] In this embodiment, preferably, the labeling module is specifically: formatting and labeling the pictures in the training data through a multimodal model to form a label file, and the formatting format is: gender, skin, eye expression, hairstyle, mouth shape, expression, accessories, clothing, and background.

[0058] Since the device introduced in the second embodiment of the present invention is the device adopted for implementing the method in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and deformation of the device, so it will not be elaborated here. Any device adopted for the method in the first embodiment of the present invention falls within the scope of protection of the present invention.

[0059] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to Embodiment 1, as detailed in Embodiment 3.

[0060] Embodiment 3

[0061] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, any implementation manner in Embodiment 1 can be realized.

[0062] Since the electronic device introduced in this embodiment is the device used to implement the method in Embodiment 1 of this application, based on the method introduced in Embodiment 1 of this application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the implementation of how this electronic device realizes the method in the embodiments of this application will not be introduced in detail here. As long as the device used by those skilled in the art to implement the method in the embodiments of this application belongs to the scope protected by this application.

[0063] Based on the same inventive concept, this application provides a storage medium corresponding to Embodiment 1, as detailed in Embodiment 4.

[0064] Embodiment 4

[0065] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any implementation manner in Embodiment 1 can be realized.

[0066] The technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0067] This embodiment not only significantly enhances the consistency of the generated character images, ensures the stability and coherence of the model characteristics, but also greatly improves the clarity of the images, making the detail performance more accurate and vivid. In addition, this embodiment improves the training strategy and parameter adjustment, enhances the generalization ability of the model in diverse scenarios, enables the model to adapt to a wider range of application requirements, and improves the practicality and flexibility of the model; and this embodiment enhances the training data by providing the generated model images, provides sufficient training data for the model, and ensures the quality of training.

[0068] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0069] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0070] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0071] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0072] Although the specific embodiments of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative only and not used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope of the claims of the present invention.

Claims

1. A method for training exclusive models of Lora, characterized by: The steps include: Step 1, obtaining image data of a set model, and performing data cleaning to obtain cleaned data, wherein the image data includes a close-up image of a face; Step 2: If the number of images in the cleaned data is greater than the set threshold, proceed to step 3; If the number of images in the cleaned data is less than or equal to the set threshold, the cleaned data is enhanced to obtain enhanced data. If the number of images in the enhanced data is greater than the set threshold, the process proceeds to step 3. Step 3: Process each image in the enhanced data or cleaned data to ensure that the long side of each image is a set value, and obtain training data; Step 4: Format and label the images in the training data through the multimodal model to form a label file; Step 5: Set the number of training rounds, learning rate, and training image pixels, then input the training data and label files to perform model training to obtain a dedicated model.

2. A method for training exclusive models of Lora according to claim 1, characterized in that: The specific step 2 is: if the number of images in the cleaned data is greater than the set threshold, and the cleaned data includes: close-up images of faces, half-body images and full-body images, then proceed to step 3; otherwise, the cleaned data is enhanced to obtain enhanced data, and the enhanced data includes: close-up images of faces, half-body images and full-body images, and the ratio of the number is 2:1:1; the number of images in the enhanced data is greater than the set threshold, and then proceed to step 3.

3. A method for training exclusive models of Lora according to claim 1 or 2, characterized in that: The data enhancement includes: generating a model graph, wherein generating the model graph is specifically: A close-up face image is obtained from the image data, which is a frontal portrait; an uploaded model reference half-length image or a model reference half-length image is obtained; ComfyUI calls the Ipadapter algorithm, adjusts the weight type to ease out, and sets the weight parameter weight of the Ipadapter algorithm to 0.5; ComfyUI calls the InstantId algorithm, and sets the weight parameter weight of the InstantId algorithm to 0.6; ComfyUI calls the openpose algorithm, which is used for body key point detection of the model reference half-length image or the model reference half-length image, and outputs a key point image; the close-up face image is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference half-length image or the model reference half-length image is processed by the openpose algorithm, and then the processed results are uniformly passed to the first image generation node of ComfyUI to generate portrait data; the first The first image raw node adopts the img2img algorithm, sets the prompt word correlation coefficient to 7, and the redrawing amplitude to 0.75; performs facial segmentation and eye detection on the portrait data, and performs facial and eye repair through mask redrawing of the second image raw node to obtain a repaired image; the second image raw node adopts the img2img algorithm, sets the prompt word correlation coefficient to 1, the redrawing step number to 20, and the redrawing amplitude is only 0.45; the repaired image is passed into IC-Light, and the lighting is optimized according to the set prompt words to obtain an optimized image; a custom node is set in ComfyUI, and the custom node is used to adjust the brightness, transparency and contrast; the optimized image is adjusted according to the custom node to obtain the required half-body or full-body image.

4. The method for training exclusive models of Lora according to claim 1, characterized in that: The step 4 is specifically as follows: formatting and labeling the images in the training data through a multimodal model to form a label file, wherein the formatting format is: gender, skin, eyes, hairstyle, mouth shape, expression, accessories, clothing and background.

5. A device for training exclusive models of Lora, characterized by: include: A data cleaning module, which obtains image data of a set model and performs data cleaning to obtain cleaned data, wherein the image data includes a close-up image of a face; Enhanced data module: if the number of images in the cleaned data is greater than the set threshold, the data processing module is entered. If the number of images in the cleaned data is less than or equal to the set threshold, the cleaned data is enhanced to obtain enhanced data. If the number of images in the enhanced data is greater than the set threshold, step 3 is entered; The data processing module processes each image in the enhanced data or cleaned data to ensure that the long side of each image is the set value to obtain training data; The labeling module formats and labels the images in the training data through a multimodal model to form a label file; In the training module, set the number of training rounds, learning rate, and training image pixels, then input the training data and label files to perform model training and obtain a dedicated model.

6. The device for training exclusive models of Lora according to claim 5, characterized in that: The enhanced data module is specifically as follows: if the number of images in the cleaned data is greater than a set threshold, and the cleaned data includes: close-up images of faces, half-body images, and full-body images, then the data processing module is entered; otherwise, the cleaned data is enhanced to obtain enhanced data, and the enhanced data includes: close-up images of faces, half-body images, and full-body images, and the ratio of the number of the images is 2:1:1; if the number of images in the enhanced data is greater than the set threshold, step 3 is entered.

7. A device for training exclusive models of Lora according to claim 5 or 6, characterized in that: The data enhancement includes: generating a model graph, wherein generating the model graph is specifically: A close-up face image is obtained from the image data, which is a frontal portrait; an uploaded model reference half-length image or a model reference half-length image is obtained; ComfyUI calls the Ipadapter algorithm, adjusts the weight type to ease out, and sets the weight parameter weight of the Ipadapter algorithm to 0.5; ComfyUI calls the InstantId algorithm, and sets the weight parameter weight of the InstantId algorithm to 0.6; ComfyUI calls the openpose algorithm, which is used for body key point detection of the model reference half-length image or the model reference half-length image, and outputs a key point image; the close-up face image is processed by the Ipadapter algorithm and the InstantId algorithm respectively, and the model reference half-length image or the model reference half-length image is processed by the openpose algorithm, and then the processed results are uniformly passed to the first image generation node of ComfyUI to generate portrait data; the first The first image raw node adopts the img2img algorithm, sets the prompt word correlation coefficient to 7, and the redrawing amplitude to 0.75; performs facial segmentation and eye detection on the portrait data, and performs facial and eye repair through mask redrawing of the second image raw node to obtain a repaired image; the second image raw node adopts the img2img algorithm, sets the prompt word correlation coefficient to 1, the redrawing step number to 20, and the redrawing amplitude is only 0.45; the repaired image is passed into IC-Light, and the lighting is optimized according to the set prompt words to obtain an optimized image; a custom node is set in ComfyUI, and the custom node is used to adjust the brightness, transparency and contrast; the optimized image is adjusted according to the custom node to obtain the required half-body or full-body image.

8. The device for training exclusive models of Lora according to claim 5, characterized in that: The labeling module specifically formats and labels the images in the training data through a multimodal model to form a label file, and the formatting format is: gender, skin, eyes, hairstyle, mouth shape, expression, accessories, clothing and background.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.