Controllable method, device, storage medium and equipment for generating portraits of specific people
By combining IPAdapter, Stable-DiffusionXL and Lora models, processing clothing and posture images and text descriptions, the problem of the Stable-Diffusion model lacks control when generating image content, achieving high-quality specific character image generation and fine control effects.
Patent Information
- Application Number
- CN202410422573.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-09
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-04-09
AI Technical Summary
The Stable-Diffusion model lacks control when image content generation, especially in terms of color confusion and character feature retention.
By combining the pre-trained IPAdapter model, Stable-DiffusionXL model and Lora model, clothing images, posture images and text descriptors are obtained and these inputs are processed to generate specific person images with specific clothing and postures.
It realizes the generation of high-quality digital clones for specific characters, and has fine control over the figure and clothing, with high restoration, good controllability and easy operation.
Smart Images

Figure CN118447110B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a controllable method, device, storage medium and equipment for generating a portrait of a specific person. Background Art
[0002] The Stable-Diffusion model can output stunning results in a few seconds using appropriate prompt words without complicated post-processing. In addition, the Stable-Diffusion model can simulate various styles, such as the style of Van Gogh by inputting relevant prompt words. Although the Stable-Diffusion generation capability is so powerful, it still lacks control over the image content generation, such as color confusion and character feature retention. Summary of the invention
[0003] The present application provides a controllable method, device, storage medium and equipment for generating a portrait of a specific person, which is used to solve the problem that the Stable-Diffusion model lacks control over the generation of image content. The technical solution is as follows:
[0004] According to a first aspect of the present application, a controllable method for generating a portrait of a specific person is provided, the method comprising:
[0005] Acquire a clothing image, a body posture image, text description words and random noise, wherein the clothing image is used to define the clothing of a specific person, the body posture image is used to define the body posture of the specific person, and the text description words are used to define the portrait style;
[0006] The clothing image and the random noise are processed using a pre-trained IPAdapter model, a Stable-DiffusionXL model, and a Lora model to obtain a first feature vector, wherein the Lora model is a model trained for the specific person, and the first feature vector is used to represent a specific person with specific clothing;
[0007] Using the pre-trained ControlNet model, the Stable-DiffusionXL model and the Lora model, the posture image, the text description word and the random noise are processed to obtain a second feature vector, where the second feature vector is used to represent a specific person with a specific posture;
[0008] The first feature vector and the second feature vector are processed by using the Stable-DiffusionXL model and the Lora model to obtain a specific person image with the specific clothing and the specific posture.
[0009] In a possible implementation, the pre-trained IPAdapter model, Stable-DiffusionXL model, and Lora model are used to process the clothing image and the random noise to obtain a first feature vector, including:
[0010] Processing the clothing image using the IPAdapter model to obtain a first intermediate feature;
[0011] The first intermediate feature and the random noise are processed by using the Stable-DiffusionXL model and the Lora model to obtain a first feature vector.
[0012] In a possible implementation, the pre-trained ControlNet model, the Stable-DiffusionXL model, and the Lora model are used to process the posture image, the text description word, and the random noise to obtain a second feature vector, including:
[0013] Using the ControlNet model to process the posture image, the text description word, and the random noise to obtain a second intermediate feature;
[0014] The second intermediate feature, the text description word and the random noise are processed by using the Stable-DiffusionXL model and the Lora model to obtain a second feature vector.
[0015] In a possible implementation, the method further includes:
[0016] Create the Lora model;
[0017] Acquire a first training sample set, each set of first training samples includes an original specific person image and a first text description, where the first text description is used to describe the original specific person image;
[0018] Processing the original specific person image and the first text description using the Stable-DiffusionXL model and the Lora model to obtain a predicted specific person image;
[0019] Calculating the loss function of the Lora model according to the original specific person image and the predicted specific person image;
[0020] The Lora model is trained according to the loss function of the Lora model.
[0021] In a possible implementation, the method further includes:
[0022] Create a ControlNet model;
[0023] Acquire a second training sample set, each set of second training samples includes a posture image, a first original person image, and a second text description, wherein the posture image is extracted from the first original person image, and the second text description is used to describe the first original person image;
[0024] Processing the posture image and the second text description using the Stable-DiffusionXL model and the ControlNet model to obtain a first predicted person image;
[0025] Calculating a loss function of the ControlNet model according to the first original character image and the first predicted character image;
[0026] The ControlNet model is trained according to the loss function of the ControlNet model.
[0027] In a possible implementation, the method further includes:
[0028] Create IPAdapter model;
[0029] Acquire a third training sample set, each set of the third training samples includes a clothing image, a second original character image, and a third text description, wherein the third text description is used to describe the second original character image;
[0030] Processing the clothing image and the third text description using the Stable-DiffusionXL model and the IPAdapter model to obtain a second predicted character image;
[0031] Calculating a loss function of the IPAdapter model according to the second original character image and the second predicted character image;
[0032] The IPAdapter model is trained according to the loss function of the IPAdapter model.
[0033] In a possible implementation, the method further includes:
[0034] Get the original character image;
[0035] The original person image is extracted using a Densepose algorithm to obtain the posture image.
[0036] According to a second aspect of the present application, a controllable device for generating a portrait of a specific person is provided, the device comprising:
[0037] an acquisition module, used to acquire clothing images, posture images, text description words and random noise, wherein the clothing images are used to define the clothing of a specific person, the posture images are used to define the body posture of a specific person, and the text description words are used to define the portrait style;
[0038] A first processing module is used to process the clothing image and the random noise using a pre-trained IPAdapter model, a Stable-DiffusionXL model and a Lora model to obtain a first feature vector, wherein the Lora model is a model trained for the specific person, and the first feature vector is used to represent a specific person with specific clothing;
[0039] A second processing module is used to process the posture image, the text description word and the random noise by using the pre-trained ControlNet model, the Stable-DiffusionXL model and the Lora model to obtain a second feature vector, wherein the second feature vector is used to represent a specific person with a specific posture;
[0040] A generation module is used to process the first feature vector and the second feature vector by using the Stable-DiffusionXL model and the Lora model to obtain a specific person image with the specific clothing and the specific posture.
[0041] According to a third aspect of the present application, a computer-readable storage medium is provided, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the controllable specific character portrait generation method as described above.
[0042] According to a fourth aspect of the present application, a computer device is provided, wherein the computer device includes the above-mentioned controllable specific person portrait generation device.
[0043] The beneficial effects of the technical solution provided by this application include at least:
[0044] Based on the StableDiffusionXL model, combined with the Lora model, ControlNet model and IPAdapter model, it is possible to create high-quality digital avatars of specific people based on small data sets, and to control the posture, clothing, etc. of specific people, with high restoration, good controllability and easy operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0046] Figure 1 It is a schematic diagram of the network structure of the Stable-DiffsuionXL model and the Lora model;
[0047] Figure 2 It is a schematic diagram of the network structure of the ControlNet model and the StableDiffusionXL model;
[0048] Figure 3 It is a schematic diagram of posture image extraction;
[0049] Figure 4 This is a schematic diagram of the network structure of the IPAdapter model and the StableDiffusionXL model;
[0050] Figure 5 It is a flow chart of a controllable method for generating a portrait of a specific person;
[0051] Figure 6 It is a schematic diagram of a controllable method for generating a portrait of a specific person;
[0052] Figure 7 It is a flow chart of a controllable method for generating a portrait of a specific person;
[0053] Figure 8 It is a structural block diagram of a controllable specific person portrait generation device. DETAILED DESCRIPTION
[0054] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings.
[0055] The following first introduces the four models involved in this application.
[0056] (1) Stable-DiffusionXL model
[0057] The Stable-DiffusionXL model is an upgraded version of the Stable-Diffsuion model. It is a basic large model that can simulate various styles. For example, by inputting relevant prompt words, you can simulate the style of Van Gogh.
[0058] (2) Lora model
[0059] For some special scenarios, it is often necessary to fine-tune the large multi-modal generative model. For example, in order to maintain the characteristics of the characters, we use the low-rank adaptation large model fine-tuning technology (Low Rank Adaptation) to train a dedicated small character model Lora model for each person to maintain the character ID identity.
[0060] Specifically, we can obtain 30-50 images of a specific person, and fine-tune the Lora model based on the Stable-DiffsuionXL model and these images to obtain the generation capability of the specific person. Among them, each Lora model corresponds to a specific person. In other words, we need to train different Lora models for different specific people.
[0061] Figure 1 The network structure diagram of the Stable-DiffsuionXL model and the Lora model is shown, where: Figure 1 .a is a training diagram of the Lora model. Figure 1 .b is a schematic diagram of the network layer structure after the Lora model and the Stable-DiffsuionXL model are fused. Figure 1 As shown in .a, under the premise of freezing the weights of the Stable-DiffsuionXL model, we can fine-tune a Lora model through 30-50 single-person images of a specific person XXX. Specifically, create a Lora model; obtain a first training sample set, each set of first training samples includes an original specific person image and a first text description, and the first text description is used to describe the original specific person image; use the Stable-DiffusionXL model and the Lora model to process the original specific person image and the first text description to obtain a predicted specific person image; calculate the loss function of the Lora model based on the original specific person image and the predicted specific person image; train the Lora model based on the loss function of the Lora model.
[0062] The trained Lora model is combined with the Stable-DiffuseXL model to generate specific character images in different scenes, different actions, and different costumes under appropriate prompt words. Figure 1 As shown in .b, X represents input, Y represents input, and the Lora model adds a bypass to the computational layer W of the original Stable-DiffsuionXL model, inserts a low-rank matrix, and fine-tunes the change ΔW of the specific person XXX. During the training process of the Lora model, W needs to be kept unchanged, and ΔW is adjusted to fit and generate the specific person XXX.
[0063] In this embodiment, small data can be used to fine-tune the Lora model of a specific person, with a short training cycle and low consumption of computing resources.
[0064] (3) ControlNet model
[0065] The images generated by the Stable-DiffusionXL model guided by prompt words cannot accurately control the image details. The ControlNet model is introduced in this application. The ControlNet model chooses to introduce other control conditions, such as contour edges, line shapes, texture tones, depth normals, human postures, etc. Based on the pre-trained Stable-DiffusionXL model, the ControlNet model adds an additional image input, such as a posture map, line draft, sketch, etc. This additional image input serves as a control condition of the StableDiffusionXL model to control the image generated by the StableDiffusionXL model to meet the image features we input.
[0066] Figure 2 Figure 2 is a schematic diagram of the network structure of the ControlNet model and the StableDiffusionXL model. Figure 2 As shown in the figure, the ControlNet model and the StableDiffusionXL model are connected in parallel. The right side shows the branch of the ControlNet model with updated weights, whose input control conditions include posture images and text descriptions. The left side shows the branch of the StableDiffusionXL model with frozen weights, whose input includes text descriptions and random noise x. t , output result x t-1 is a person image with a specific posture. Specifically, a ControlNet model is created; a second training sample set is obtained, each set of second training samples includes a posture image, a first original person image and a second text description, the posture image is extracted from the first original person image, and the second text description is used to describe the first original person image; the posture image and the second text description are processed by using the Stable-DiffusionXL model and the ControlNet model to obtain a first predicted person image; the loss function of the ControlNet model is calculated according to the first original person image and the first predicted person image; and the ControlNet model is trained according to the loss function of the ControlNet model.
[0067] During the training process, the ControlNet model uses the weights of the SD encoding layer and SD middle layer corresponding to the StableDiffusionXL model as the initial values, while the StableDiffusionXL model branch retains the original weights of the basic model. Therefore, using a small amount of data to guide can ensure that the pre-constraints can be fully learned while retaining the learning ability of the original diffusion model itself. Among them, only the SD encoding layer and SD middle layer of the ControlNet model branch receive text descriptions; the intermediate results of the ControlNet model branch flow to the SD middle layer and SD decoding layer of the StableDiffusionXL model branch after the convolutional network layer. The extraction method of posture images in the ControlNet model can be obtained by Figure 3 The method shown is obtained. Figure 3 As shown in the figure, the original character image can obtain a semantic representation of the human body through the dense pose algorithm DensePose, and different colors correspond to different parts of the body.
[0068] In this embodiment, the ControlNet model can not only control body movements, but also control body shape.
[0069] (4)IPAdapter Model
[0070] The IPAdapter model is another image-controlled generation technique that extracts feature vectors from images and feeds them into an additional cross-attention layer in Unet to guide the generation of images. The key design of the IPAdapter model is to decouple the cross-attention mechanism, which separates the cross-attention layers of text features and image features, so that image prompts and text prompts can better cooperate to achieve multimodal image generation. In order to achieve fine control over the color, texture, and shape of clothing, we used 50K text-clothing-model triples to train the IPAdapter model.
[0071] Figure 4The diagram is a network structure diagram of the IPAdapter model and the StableDiffusionXL model. The upper part is the IPAdapter model, the lower part is the StableDiffusionXL model, and the snowflakes represent weight freezing. The IPAdapter model guides the generation result of the StableDiffusionXL model to approach a specific clothing image by decoupling cross attention, and the final output result is a model wearing a specific clothing. Specifically, create an IPAdapter model; obtain a third training sample set, each set of third training samples includes a clothing image, a second original character image and a third text description, and the third text description is used to describe the second original character image; use the Stable-DiffusionXL model and the IPAdapter model to process the clothing image and the third text description to obtain a second predicted character image; calculate the loss function of the IPAdapter model based on the second original character image and the second predicted character image; train the IPAdapter model based on the loss function of the IPAdapter model.
[0072] During the training process, only the network weights of the IPAdapter model need to be adjusted to inject image features through cross-attention.
[0073] In this embodiment, the IPAdapter model is friendly to clothing and accessories, and generates clothing and accessories by reference as much as possible, and can also change some attributes of clothing and accessories in combination with text prompt words.
[0074] Although the Lora model, ControlNet model and IPAdapter model are trained separately, all three models use the same StableDiffusionXL model. Therefore, the StableDiffusionXL model, the Lora model, the ControlNet model and the IPAdapter model are not independent, but are connected in parallel to form a "big model". They can be used and controlled separately at the same time to achieve the generation of body movements and clothing for specific characters.
[0075] like Figure 5 As shown, it shows a method flow chart of a controllable specific person portrait generation method provided by an embodiment of the present application, and the controllable specific person portrait generation method can be applied to a computer device. The controllable specific person portrait generation method may include:
[0076] Step 501, acquiring clothing images, body image, text description words and random noise, wherein the clothing images are used to define the clothing of a specific person, the body image is used to define the body posture of a specific person, and the text description words are used to define the portrait style.
[0077] like Figure 6 As shown, the clothing image is a red vest, the posture image is an image of a sideways standing posture, the text description word is "a photo of a xxx woman", and the random noise is randomly generated noise. Among them, xxx is the trigger word of a specific person.
[0078] Step 502, using the pre-trained IPAdapter model, Stable-DiffusionXL model and Lora model to process the clothing image and random noise to obtain a first feature vector. The Lora model is a model trained for a specific person, and the first feature vector is used to represent a specific person with specific clothing.
[0079] The training process of the IPAdapter model, Stable-DiffusionXL model, and Lora model is described above and will not be repeated here.
[0080] Step 503 , using the pre-trained ControlNet model, Stable-DiffusionXL model and Lora model to process the posture image, text description words and random noise to obtain a second feature vector, where the second feature vector is used to represent a specific person with a specific posture.
[0081] The training process of the ControlNet model is described above and will not be repeated here.
[0082] Step 504: Use the Stable-DiffusionXL model and the Lora model to process the first feature vector and the second feature vector to obtain a specific person image with specific clothing and specific posture.
[0083] like Figure 6 As shown, the first eigenvector and the second eigenvector are processed using the Stable-DiffusionXL model and the Lora model, and the processing results are input into the VAE (Variational Autoencoder) decoder for processing to obtain an image of a woman wearing a red vest and standing sideways.
[0084] To sum up, the controllable specific person portrait generation method provided in the embodiment of the present application is based on the StableDiffusionXL model, combined with the Lora model, ControlNet model and IPAdapter model, to achieve the production of high-quality digital avatars of specific persons based on small data sets, and can achieve control over the posture, clothing, etc. of the specific person, with high restoration, good controllability and easy operation.
[0085] like Figure 7 As shown, it shows a flow chart of a controllable specific person portrait generation method provided by an embodiment of the present application, and the controllable specific person portrait generation method can be applied to a computer device. The controllable specific person portrait generation method may include:
[0086] Step 701, acquiring clothing images, body image, text description words and random noise, wherein the clothing images are used to define the clothing of a specific person, the body image is used to define the body posture of a specific person, and the text description words are used to define the portrait style.
[0087] Among them, the posture image is extracted from the original character image, such as Figure 3 Specifically, obtain the original character image; use the Densepose algorithm to extract the original character image to obtain the posture image. Among them, xxx is the trigger word of a specific character.
[0088] like Figure 6 As shown, the clothing image is a red vest, the posture image is an image of a sideways standing posture, the text description word is "a photo of a xxx woman", and the random noise is randomly generated noise.
[0089] Step 702, use the IPAdapter model to process the clothing image to obtain a first intermediate feature; use the Stable-DiffusionXL model and the Lora model to process the first intermediate feature and random noise to obtain a first feature vector, the Lora model is a model trained for a specific person, and the first feature vector is used to represent a specific person with specific clothing.
[0090] like Figure 4 As shown in the figure, the clothing image is first encoded using the picture encoder in the IPAdapter model, and then the image features are extracted from the encoded features using MLP (Multilayer Perceptron), and finally the image features are processed using the cross attention module to obtain the first intermediate features. The first intermediate features and random noise are input into the Stable-DiffusionXL model and the Lora model for processing to obtain the first feature vector.
[0091] Step 703, using the ControlNet model to process the posture image, text description words and random noise to obtain a second intermediate feature; using the Stable-DiffusionXL model and the Lora model to process the second intermediate feature, text description words and random noise to obtain a second feature vector, which is used to represent a specific person with a specific posture.
[0092] like Figure 2As shown, the convolutional network, SD encoding layer, SD intermediate layer and two convolutional networks in the ControlNet model are first used to process the posture image, text description words and random noise in turn to obtain the second intermediate features; the second intermediate features output by the two convolutional networks are input into the Stable-DiffusionXL model and the Lora model for processing to obtain the second feature vector.
[0093] Step 704: Use the Stable-DiffusionXL model and the Lora model to process the first feature vector and the second feature vector to obtain a specific person image with specific clothing and specific posture.
[0094] like Figure 6 As shown, the first eigenvector and the second eigenvector are processed using the Stable-DiffusionXL model and the Lora model, and the processing results are input into the VAE decoder for processing to obtain an image of a woman wearing a red vest and standing sideways.
[0095] To sum up, the controllable specific person portrait generation method provided in the embodiment of the present application is based on the StableDiffusionXL model, combined with the Lora model, ControlNet model and IPAdapter model, to achieve the production of high-quality digital avatars of specific persons based on small data sets, and can achieve control over the posture, clothing, etc. of the specific person, with high restoration, good controllability and easy operation.
[0096] like Figure 8 As shown, it shows a structural block diagram of a controllable specific person portrait generation device provided by an embodiment of the present application, and the controllable specific person portrait generation device can be applied to a computer device. The controllable specific person portrait generation device may include:
[0097] An acquisition module 810 is used to acquire clothing images, body posture images, text description words and random noise, wherein the clothing images are used to define the clothing of a specific person, the body posture images are used to define the body posture of a specific person, and the text description words are used to define the portrait style;
[0098] A first processing module 820 is used to process the clothing image and random noise using the pre-trained IPAdapter model, the Stable-DiffusionXL model and the Lora model to obtain a first feature vector. The Lora model is a model trained for a specific person. The first feature vector is used to represent a specific person with specific clothing.
[0099] The second processing module 830 is used to process the posture image, the text description words and the random noise using the pre-trained ControlNet model, the Stable-DiffusionXL model and the Lora model to obtain a second feature vector, where the second feature vector is used to represent a specific person with a specific posture;
[0100] The generating module 840 is used to process the first feature vector and the second feature vector by using the Stable-DiffusionXL model and the Lora model to obtain a specific person image with specific clothing and specific posture.
[0101] In an optional embodiment, the first processing module 820 is further configured to:
[0102] The clothing image is processed using the IPAdapter model to obtain the first intermediate feature;
[0103] The first intermediate feature and random noise are processed using the Stable-DiffusionXL model and the Lora model to obtain the first feature vector.
[0104] In an optional embodiment, the second processing module 830 is further configured to:
[0105] The ControlNet model is used to process the posture image, text description words and random noise to obtain the second intermediate feature;
[0106] The second intermediate feature, text description words and random noise are processed using the Stable-DiffusionXL model and the Lora model to obtain the second feature vector.
[0107] In an optional embodiment, the device further includes a first training module, which is used to:
[0108] Create the Lora model;
[0109] Acquire a first training sample set, each set of first training samples includes an original specific person image and a first text description, and the first text description is used to describe the original specific person image;
[0110] The original specific person image and the first text description are processed using the Stable-DiffusionXL model and the Lora model to obtain a predicted specific person image;
[0111] Calculate the loss function of the Lora model based on the original specific person image and the predicted specific person image;
[0112] The Lora model is trained according to the loss function of the Lora model.
[0113] In an optional embodiment, the device further includes a second training module, which is used to:
[0114] Create a ControlNet model;
[0115] Acquire a second training sample set, each set of second training samples includes a posture image, a first original person image, and a second text description, the posture image is extracted from the first original person image, and the second text description is used to describe the first original person image;
[0116] The posture image and the second text description are processed using the Stable-DiffusionXL model and the ControlNet model to obtain a first predicted person image;
[0117] Calculate the loss function of the ControlNet model according to the first original character image and the first predicted character image;
[0118] Train the ControlNet model according to its loss function.
[0119] In an optional embodiment, the device further includes a third training module, which is used to:
[0120] Create IPAdapter model;
[0121] Acquire a third training sample set, each set of the third training samples includes a clothing image, a second original character image and a third text description, and the third text description is used to describe the second original character image;
[0122] The clothing image and the third text description are processed using the Stable-DiffusionXL model and the IPAdapter model to obtain a second predicted character image;
[0123] Calculate the loss function of the IPAdapter model according to the second original character image and the second predicted character image;
[0124] Train the IPAdapter model according to its loss function.
[0125] In an optional embodiment, the acquisition module 810 is further configured to:
[0126] Get the original character image;
[0127] The Densepose algorithm is used to extract the original character image to obtain the posture image.
[0128] To sum up, the controllable specific person portrait generation device provided in the embodiment of the present application is based on the StableDiffusionXL model, combined with the Lora model, the ControlNet model and the IPAdapter model, to achieve the production of high-quality digital avatars of specific persons based on small data sets, and can control the posture, clothing, etc. of the specific person, with high restoration, good controllability and easy operation.
[0129] An embodiment of the present application provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the controllable specific character portrait generation method as described above.
[0130] An embodiment of the present application provides a computer device, which includes any of the above-mentioned controllable specific person portrait generation devices.
[0131] It should be noted that: the controllable specific person portrait generation device provided in the above embodiment only uses the division of the above functional modules as an example when performing the portrait control generation of a specific person. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the controllable specific person portrait generation device is divided into different functional modules to complete all or part of the functions described above. In addition, the controllable specific person portrait generation device provided in the above embodiment and the controllable specific person portrait generation method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0132] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0133] The above description is not intended to limit the embodiments of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the protection scope of the embodiments of the present application.
Claims
1. A controllable method for generating a portrait of a specific person, characterized in that: The method comprises: Acquire a clothing image, a body posture image, text description words and random noise, wherein the clothing image is used to define the clothing of a specific person, the body posture image is used to define the body posture of the specific person, and the text description words are used to define the portrait style; The clothing image and the random noise are processed using a pre-trained IPAdapter model, a Stable-DiffusionXL model, and a Lora model to obtain a first feature vector, wherein the Lora model is a model trained for the specific person, and the first feature vector is used to represent a specific person with specific clothing; Using the pre-trained ControlNet model, the Stable-DiffusionXL model and the Lora model, the posture image, the text description word and the random noise are processed to obtain a second feature vector, where the second feature vector is used to represent a specific person with a specific posture; Processing the first feature vector and the second feature vector by using the Stable-DiffusionXL model and the Lora model to obtain a specific person image with the specific clothing and the specific posture; The using the pre-trained IPAdapter model, Stable-DiffusionXL model and Lora model to process the clothing image and the random noise to obtain a first feature vector includes: using the IPAdapter model to process the clothing image to obtain a first intermediate feature; using the Stable-DiffusionXL model and the Lora model to process the first intermediate feature and the random noise to obtain a first feature vector; The method of using the pre-trained ControlNet model, the Stable-DiffusionXL model and the Lora model to process the posture image, the text description words and the random noise to obtain a second feature vector includes: using the ControlNet model to process the posture image, the text description words and the random noise to obtain a second intermediate feature; and using the Stable-DiffusionXL model and the Lora model to process the second intermediate feature, the text description words and the random noise to obtain a second feature vector.
2. The controllable specific person portrait generation method according to claim 1, characterized in that: The method further comprises: Create Lora model; Acquire a first training sample set, each set of first training samples includes an original specific person image and a first text description, where the first text description is used to describe the original specific person image; Processing the original specific person image and the first text description using the Stable-DiffusionXL model and the Lora model to obtain a predicted specific person image; Calculating the loss function of the Lora model according to the original specific person image and the predicted specific person image; The Lora model is trained according to the loss function of the Lora model.
3. The controllable method for generating a specific person portrait according to claim 1, characterized in that: The method further comprises: Create a ControlNet model; Acquire a second training sample set, each set of second training samples includes a posture image, a first original person image, and a second text description, wherein the posture image is extracted from the first original person image, and the second text description is used to describe the first original person image; Processing the posture image and the second text description using the Stable-DiffusionXL model and the ControlNet model to obtain a first predicted person image; Calculating a loss function of the ControlNet model according to the first original character image and the first predicted character image; The ControlNet model is trained according to the loss function of the ControlNet model.
4. The controllable method for generating a specific person portrait according to claim 1, characterized in that: The method further comprises: Create IPAdapter model; Acquire a third training sample set, each set of the third training samples includes a clothing image, a second original character image, and a third text description, wherein the third text description is used to describe the second original character image; Processing the clothing image and the third text description using the Stable-DiffusionXL model and the IPAdapter model to obtain a second predicted character image; Calculating a loss function of the IPAdapter model according to the second original character image and the second predicted character image; The IPAdapter model is trained according to the loss function of the IPAdapter model.
5. The controllable method for generating a specific person portrait according to any one of claims 1 to 4, characterized in that: The method further comprises: Get the original character image; The original person image is extracted using a Densepose algorithm to obtain the posture image.
6. A controllable device for generating a portrait of a specific person, characterized in that: The device comprises: an acquisition module, used to acquire clothing images, posture images, text description words and random noise, wherein the clothing images are used to define the clothing of a specific person, the posture images are used to define the body posture of a specific person, and the text description words are used to define the portrait style; A first processing module is used to process the clothing image and the random noise using a pre-trained IPAdapter model, a Stable-DiffusionXL model and a Lora model to obtain a first feature vector, wherein the Lora model is a model trained for the specific person, and the first feature vector is used to represent a specific person with specific clothing; A second processing module is used to process the posture image, the text description word and the random noise by using the pre-trained ControlNet model, the Stable-DiffusionXL model and the Lora model to obtain a second feature vector, wherein the second feature vector is used to represent a specific person with a specific posture; A generating module, configured to process the first feature vector and the second feature vector by using the Stable-DiffusionXL model and the Lora model, so as to obtain a specific person image having the specific clothing and the specific posture; The first processing module is further used to: process the clothing image using the IPAdapter model to obtain a first intermediate feature; process the first intermediate feature and the random noise using the Stable-DiffusionXL model and the Lora model to obtain a first feature vector; The second processing module is also used to: use the ControlNet model to process the posture image, the text description words and the random noise to obtain a second intermediate feature; use the Stable-DiffusionXL model and the Lora model to process the second intermediate feature, the text description words and the random noise to obtain a second feature vector.
7. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the controllable specific person portrait generation method as described in any one of claims 1 to 5.
8. A computer device, characterized in that: The computer device includes: the controllable specific person portrait generation device described in claim 6.
Citation Information
Patent Citations
Specific scene portrait generation method and device, storage medium and equipment
CN116977461A
Picture generation method and device, storage medium and computing equipment
CN117036546A