Method, device and equipment for generating model commodity graph based on mannequin graph and medium
By combining image processing technology and Flux_fill model, using deep learning and machine vision technology, the real model product picture is quickly generated from human-stage model images, solving the problems of high cost of shooting traditional clothing models and lagging timeliness, and achieving rapid promotion and monetization of the e-commerce industry.
Patent Information
- Application Number
- CN202510108335.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional clothing models shoot requires hiring professionals and high shooting costs. At the same time, there is a lag in timeliness, which is difficult to meet the needs of rapid promotion and monetization in the e-commerce industry.
By combining image processing technology with Flux_fill model, deep learning models and machine vision technologies such as Redux, Flux_fill, Ipadapter, Openpose and Lineart can be used to quickly generate real model product images from human-stage model images to achieve efficient conversion.
This greatly reduces the cost of clothing shooting in the e-commerce industry, improves timeliness, and allows merchants to quickly update product displays, adapt to market changes, and improve market response speed.
Smart Images

Figure CN120070661A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of generating pictures, and particularly relates to a method, device, equipment and medium for generating model product pictures based on mannequin pictures. Background Art
[0002] In the e-commerce industry, merchants need to display clothing. Many supply chains or sellers display clothing in the form of mannequin pictures. However, for users, the experience of mannequin pictures is poor. Therefore, when merchants promote, using mannequin pictures has a poor effect; on the contrary, using real model pictures to display clothing has a better experience for users and is more conducive to merchant promotion. Therefore, merchants still need to use real people to take pictures of clothing to bring more traffic to their product links and increase the product conversion rate.
[0003] Traditional clothing model shooting requires hiring professional personnel such as models, photographers, and makeup artists, renting or building a photography studio, and then requires post-production retouching, which has a high shooting cost and strong lag in timeliness, which is not conducive to merchant promotion and monetization. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for generating model product pictures based on mannequin pictures, which can quickly generate the required product pictures without any delay, facilitating users to quickly go live.
[0005] In a first aspect, the present invention provides a method for generating model product pictures based on mannequin pictures, including the following steps:
[0006] Step 1: Upload a mannequin picture of a set of clothing and a real person picture with a set exposed skin area; obtain a clothing mask picture from the mannequin picture;
[0007] Step 2: Input the real person picture into the Redux model, input the clothing mask picture, the mannequin picture and the Redux model into the collector of the Flux_fill model, and then perform picture-to-picture generation through the Flux_fill model to obtain a preliminary portrait picture;
[0008] Step 3: Perform matte extraction on the preliminary portrait picture to obtain a portrait main body picture, and synthesize the portrait main body picture with a white background picture to obtain a white background portrait picture;
[0009] Step 4: Extract the edge information of the white background portrait picture through the lineart model; extract the pose information of the white background portrait picture through the openpose model;
[0010] Step 5: Input the set reference portrait into the ipadapter model to extract the portrait features; input the edge information, pose information, white-background portrait, and portrait features into the diffusion model to generate the real model product image.
[0011] In a second aspect, the present invention provides an apparatus for generating a model product image based on a mannequin image, including:
[0012] An upload module that uploads a mannequin image of a set of clothing and a real person image with a set exposed skin area; obtains a clothing mask image from the mannequin image;
[0013] A preliminary generation module that inputs the real person image into the Redux model, inputs the clothing mask image, the mannequin image, and the Redux model into the collector of the Flux_fill model, and then performs image generation through the Flux_fill model to obtain a preliminary portrait image;
[0014] A matte extraction module that extracts the matte of the preliminary portrait image to obtain the portrait main body image, and synthesizes the portrait main body image with a white-background image to obtain a white-background portrait image;
[0015] A guidance information module that extracts the edge information of the white-background portrait image through the lineart model; extracts the pose information of the white-background portrait image through the openpose model;
[0016] A final generation module that inputs the set reference portrait into the ipadapter model to extract the portrait features; inputs the edge information, pose information, white-background portrait, and portrait features into the diffusion model to generate the real model product image.
[0017] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method described in the first aspect is implemented.
[0018] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in the first aspect is implemented.
[0019] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:
[0020] By combining image processing technology with the Flux_fill model, the present invention realizes an efficient conversion from a mannequin model to a real model, greatly reducing the clothing shooting cost in the e-commerce industry and improving the timeliness.
[0021] The present invention utilizes a deep learning model, combines the principles of machine vision and graphics, and can automatically identify the contours of mannequin models and the poses of human figures, creating highly realistic model images. This enables merchants to no longer need to pay the costs of hiring real models, photographers, and post-processing, further shortening the product listing cycle and improving the market response speed.
[0022] The present invention supports large-scale image processing, enabling e-commerce sellers to quickly update product displays and rapidly adapt to market changes, thereby gaining an advantage in the fierce market competition.
[0023] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the content of the specification. And in order to make the above and other objects, features, and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. Brief Description of the Drawings
[0024] The following further describes the present invention with reference to the accompanying drawings in conjunction with embodiments.
[0025] Figure 1 It is a flowchart of the method in Embodiment 1 of the present invention;
[0026] Figure 2 It is a structural schematic diagram of the device in Embodiment 2 of the present invention. Detailed Embodiments
[0027] Embodiments of the present application provide a method, device, equipment, and medium for generating model product images based on mannequin images. By combining image processing technology with the Flux_fill model, efficient conversion from mannequin models to real models is achieved, greatly reducing the clothing shooting costs in the e-commerce industry and improving timeliness.
[0028] The overall idea of the technical solution in the embodiments of the present application is as follows:
[0029] Redux algorithm: The redux model is a tool class model under the flux series, mainly used to generate variants of images, allowing the generated images to refer to the appearance, material, style features, etc. of the uploaded images.
[0030] Flux_fill model: The fill model is a tool class model under the flux series, and its main function is to perform local redrawing.
[0031] Ipadapter: Ipadapter is an adapter used to add image prompting capabilities to a pre-trained text-to-image diffusion model. It embeds image features into the pre-trained text-to-image diffusion model by using a decoupled cross-attention mechanism, thus realizing the image prompting function. The key design of Ipadapter is the decoupled cross-attention mechanism, which separates the cross-attention layers of text features and image features, allowing the model to more effectively capture fine-grained features from image prompts.
[0032] Openpose: Openpose is a keypoint detection framework that can detect the keypoints of the human body in an image or video stream, including keypoints of the body, face, hands, and even feet.
[0033] Lineart: The Lineart algorithm can convert complex color images into simple line drawings. These line drawings emphasize the contours and structures of the images, removing or simplifying color and texture information, and playing a role in controlling edge contour information in the algorithm.
[0034] The main implementation is as follows:
[0035] (1) Upload a mannequin image as the initial image to be replaced.
[0036] (2) Upload a human limb image with as large an exposed skin area as possible as a reference image for the redux algorithm.
[0037] (3) Call the redux algorithm and set the redux weight control parameter to 0.9. Redux will learn the appearance, material, pose, and style features of the image.
[0038] (4) Pass the mannequin image, the clothes mask image, and the redux algorithm into the sampler and call the flux_fill model for image generation from image. Here, it mainly generates a preliminary portrait image. At this time, the mannequin texture of the mannequin image has been converted into a real human texture. The specific parameters are as follows: the sampler is euler; the scheduler is simple; the redrawing amplitude is 0.9; the sampling steps are 15, and the prompt word correlation coefficient is 2.
[0039] (5) Call the portrait segmentation algorithm to perform preliminary portrait cropping and paste the portrait on a white background image to avoid the influence of the background of the image generated in the first step on the next image generation from image.
[0040] (6) Call the lineart algorithm to extract the edge information of the white background portrait image. The model weight parameter is set to 0.6, which plays a role in controlling the edges of the clothes in image generation from image.
[0041] (7) Call the OpenPose algorithm to extract the pose information of the white-background portrait image. The model weight parameter is set to 0.9, which plays a role in guiding the model to generate poses in image-to-image generation.
[0042] (8) Finally, call the iPadapter algorithm and upload a reference portrait image to be used as a reference for generating the model's face. The model weight parameter here can be set to 1, so that the generated model has a high similarity to the reference portrait image.
[0043] (9) Input the above three algorithms into the SD model for image-to-image generation to generate the final realistic model effect diagram. The specific parameters for this image-to-image generation are as follows: The sampler is dpmpp_2m; the scheduler is Karras; the redraw amplitude is 0.7; the sampling steps are 20, and the prompt word correlation coefficient is 7.
[0044] Example 1
[0045] As Figure 1 shown, this example provides a method for generating model product images based on a mannequin image, including the following steps:
[0046] Step 1: Upload a mannequin image of a set of clothing and a real-person image with a set exposed skin area; obtain the clothing mask image from the mannequin image.
[0047] Step 2: Input the real-person image into the Redux model, input the clothing mask image, the mannequin image, and the Redux model into the collector of the Flux_fill model, and then perform image-to-image generation through the Flux_fill model to obtain a preliminary portrait image.
[0048] Step 3: Cut out the preliminary portrait image to obtain the main body image of the portrait, and synthesize the main body image of the portrait with the white-background image to obtain a white-background portrait image.
[0049] Step 4: Extract the edge information of the white-background portrait image through the lineart model; extract the pose information of the white-background portrait image through the OpenPose model.
[0050] Step 5: Input the set reference portrait image into the iPadapter model to extract the portrait features; input the edge information, pose information, white-background portrait image, and portrait features into the diffusion model to generate the real model product image.
[0051] In this embodiment, preferably, step 2 is specifically as follows: Input the real person image into the Redux model, and set the weight control parameter of the Redux model to 0.9. Input the clothing mask image, the first set of prompt words, the mannequin image, and the Redux model into the collector of the Flux_fill model, and then perform image generation from image through the Flux_fill model to obtain a preliminary portrait image; The parameters in the Flux_fill model are set as follows: the sampler is euler; the scheduler is simple; the redrawing amplitude is 0.9; the sampling steps are 15, and the prompt word correlation coefficient is 2; The first set of prompt words is used to describe limb details.
[0052] In this embodiment, preferably, step 4 is specifically as follows: Extract the edge information of the white-background portrait image through the lineart model, and set the weight parameter of the lineart model to 0.6; Extract the pose information of the white-background portrait image through the openpose model, and set the weight parameter of the openpose model to 0.9;
[0053] Step 5 is specifically as follows: Input the set reference portrait image into the ipadapter model to extract portrait features; Input the edge information, pose information, the second set of prompt words, the white-background portrait image, and the portrait features into the diffusion model to generate a real model product image. The parameters of the diffusion model are set as follows: the sampler is dpmpp_2m; the scheduler is Karras; the redrawing amplitude is 0.7; the sampling steps are 20, and the prompt word correlation coefficient is 7.
[0054] In this embodiment, preferably, it further includes step 6. When it is necessary to generate a real model product picture of a set face, picture data of a set model is obtained and data cleaning is performed to obtain cleaned data. The picture data includes a face close-up picture. If the number of pictures in the cleaned data is greater than a set threshold, and the cleaned data includes: face close-up pictures, half-body pictures, and full-body pictures, then data enhancement processing is not required. If the number of pictures in the cleaned data is less than or equal to the set threshold, the cleaned data is enhanced to obtain enhanced data. The enhanced data includes: face close-up pictures, half-body pictures, and full-body pictures, and the quantity ratio is 2:1:1. The number of pictures in the enhanced data is greater than the set threshold. Each picture in the enhanced data or the cleaned data is processed to ensure that the long side of each picture is a set value to obtain training data. The pictures in the training data are formatted and tagged through a multi-modal model to form a tag file. The formatted format is: gender, skin, eye expression, hairstyle, lip shape, expression, accessories, clothing, and background. Set the number of training rounds, learning rate, and training image pixels, and then input the training data and the tag file to perform Lora model training to obtain an exclusive model. This exclusive model not only significantly enhances the consistency of the generated human face, ensures the stability and coherence of the model features, but also greatly improves the clarity of the image, making the detail performance more accurate and vivid.
[0055] Input the edge information, pose information, second prompt, and white-background portrait into the diffusion model, and be guided by the exclusive model to generate a real model product picture. The parameter settings of the diffusion model are: the sampler is dpmpp_2m; the scheduler is Karras; the redrawing amplitude is 0.7; the number of sampling steps is 20, and the prompt correlation coefficient is 7.
[0056] Based on the same inventive concept, the present application also provides a device corresponding to the method in Embodiment 1. See Embodiment 2 for details.
[0057] Embodiment 2
[0058] As Figure 2 shown, in this embodiment, a device for generating a model product picture based on a mannequin picture is provided, including:
[0059] An upload module that uploads a mannequin picture of a set of clothing and a real person picture with a set exposed skin area; obtains a clothing mask picture from the mannequin picture;
[0060] A preliminary generation module that inputs the real person picture into the Redux model, inputs the clothing mask picture, the mannequin picture, and the Redux model into the collector of the Flux_fill model, and then performs picture-to-picture generation through the Flux_fill model to obtain a preliminary portrait picture;
[0061] The matting module performs matting on the preliminary portrait image to obtain the main body image of the portrait, and synthesizes the main body image of the portrait with the white background image to obtain the portrait image with a white background;
[0062] The guidance information module extracts the edge information of the portrait image with a white background through the lineart model; extracts the pose information of the portrait image with a white background through the openpose model;
[0063] The final generation module inputs the set reference portrait image into the ipadapter model to extract the portrait features; inputs the edge information, pose information, portrait image with a white background, and portrait features into the diffusion model to generate a real model product image.
[0064] In this embodiment, preferably, the preliminary generation module is specifically: input the real person image into the Redux model, and set the weight control parameter of the Redux model to 0.9. Input the clothes mask image, the first set of prompt words, the mannequin image, and the Redux model into the collector of the Flux_fill model, and then perform image generation from image through the Flux_fill model to obtain the preliminary portrait image; the parameters in the Flux_fill model are set as follows: the sampler is euler; the scheduler is simple; the redrawing amplitude is 0.9; the sampling steps are 15, and the prompt word correlation coefficient is 2; the first set of prompt words is used to describe limb details.
[0065] In this embodiment, preferably, the guidance information module is specifically: extract the edge information of the portrait image with a white background through the lineart model, and set the weight parameter of the lineart model to 0.6; extract the pose information of the portrait image with a white background through the openpose model, and set the weight parameter of the openpose model to 0.9;
[0066] The final generation module is specifically: input the set reference portrait image into the ipadapter model to extract the portrait features; input the edge information, pose information, the second set of prompt words, the portrait image with a white background, and the portrait features into the diffusion model to generate a real model product image. The parameters of the diffusion model are set as follows: the sampler is dpmpp_2m; the scheduler is Karras; the redrawing amplitude is 0.7; the sampling steps are 20, and the prompt word correlation coefficient is 7.
[0067] In this embodiment, preferably, it further includes a specific generation module. When it is necessary to generate a real model product image of a set face, it acquires the picture data of a set model and performs data cleaning to obtain cleaned data. The picture data includes a close-up face picture. If the number of pictures in the cleaned data is greater than the set threshold, and the cleaned data includes: close-up face pictures, half-body pictures, and full-body pictures, then no data augmentation processing is required. If the number of pictures in the cleaned data is less than or equal to the set threshold, the cleaned data is augmented to obtain augmented data. The augmented data includes: close-up face pictures, half-body pictures, and full-body pictures, and the quantity ratio is 2:1:1. The number of pictures in the augmented data is greater than the set threshold. Each picture in the augmented data or the cleaned data is processed to ensure that the long side of each picture is a set value to obtain training data. The pictures in the training data are formatted and tagged through a multi-modal model to form a tag file. The formatted format is: gender, skin, eye expression, hairstyle, mouth shape, expression, accessories, clothing, and background. Set the number of training rounds, learning rate, and training image pixels, and then input the training data and the tag file to perform Lora model training to obtain a dedicated model model.
[0068] Input the edge information, pose information, second prompt, and white-background portrait into the diffusion model, and be guided by the dedicated model model to generate a real model product image. The parameter settings of the diffusion model are: the sampler is dpmpp_2m; the scheduler is Karras; the redrawing amplitude is 0.7; the number of sampling steps is 20, and the prompt correlation coefficient is 7.
[0069] Since the device introduced in the second embodiment of the present invention is the device adopted to implement the method of the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and deformation of the device, so it will not be elaborated here. Any device adopted by the method of the first embodiment of the present invention belongs to the scope to be protected by the present invention.
[0070] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to the first embodiment, as detailed in the third embodiment.
[0071] Embodiment Three
[0072] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, any implementation manner in the first embodiment can be realized.
[0073] Since the electronic device introduced in this embodiment is the device used to implement the method in Embodiment 1 of this application, based on the method introduced in Embodiment 1 of this application, those skilled in the art can understand the specific implementation manners of the electronic device in this embodiment and its various variations. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of this application will not be introduced in detail here. As long as the device used by those skilled in the art to implement the method in the embodiments of this application belongs to the scope protected by this application.
[0074] Based on the same inventive concept, this application provides a storage medium corresponding to Embodiment 1, as detailed in Embodiment 4.
[0075] Embodiment 4
[0076] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any implementation manner in Embodiment 1 can be realized.
[0077] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0078] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the specified function in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0079] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the specified function in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0080] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process or multiple processes and / or one block or multiple blocks. Figure 1 one process or multiple processes and / or blocks Figure 1 steps for the functions specified in one block or multiple blocks.
[0081] Although the specific embodiments of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.
Claims
1. A method for generating a model product image based on a mannequin image, characterized in that: The steps include: Step 1, upload a mannequin image with set clothing and a real person image with set exposed skin area; obtain the clothing mask image from the mannequin image; Step 2: Input the real person image into the Redux model, input the clothing mask image, the mannequin image and the Redux model into the collector of the Flux_fill model, and then use the Flux_fill model to generate the image to obtain a preliminary portrait image; Step 3, cut out the preliminary portrait image to obtain the main portrait image, and synthesize the main portrait image with the white background image to obtain the white background portrait image; Step 4: Extract edge information of the portrait image with white background through the lineart model; extract posture information of the portrait image with white background through the openpose model; Step 5: Input the set reference portrait image into the iPad Adapter model to extract the portrait features; The edge information, posture information, white background portrait image and portrait features are input into the diffusion model to generate a real model product image.
2. The method for generating a model product image based on a mannequin image according to claim 1, characterized in that: The step 2 is specifically as follows: inputting the real person image into the Redux model, setting the weight control parameter of the Redux model to 0.9, inputting the clothing mask image, the first set prompt word, the mannequin image and the Redux model into the collector of the Flux_fill model, and then performing image generation through the Flux_fill model to obtain a preliminary portrait image; the parameters in the Flux_fill model are set as follows: the sampler is euler; the scheduler is simple; the redraw amplitude is 0.9; the sampling step number is 15, and the prompt word correlation coefficient is 2; the first set prompt word is used to describe limb details.
3. The method for generating a model product image based on a mannequin image according to claim 1, characterized in that: The step 4 specifically comprises: extracting edge information of the portrait image with white background by using a lineart model, wherein the weight parameter of the lineart model is set to 0.6; extracting posture information of the portrait image with white background by using an openpose model, wherein the weight parameter of the openpose model is set to 0.9; The step 5 specifically includes: inputting the set reference portrait image into the iPad Adapter model to extract the portrait features; The edge information, posture information, the second set prompt word, the white background portrait image and the portrait features are input into the diffusion model to generate a real model product image. The parameters of the diffusion model are set as follows: the sampler is dpmpp_2m; the scheduler is Karras; the redrawing amplitude is 0.7; the sampling step number is 20, and the prompt word correlation coefficient is 7.
4. The method for generating a model product image based on a mannequin image according to claim 1, characterized in that: The method further includes step 6, when it is necessary to generate a real model product image of a set face, obtaining image data of a set model, and performing data cleaning to obtain cleaned data, wherein the image data includes a close-up image of the face; If the number of images in the cleaned data is greater than the set threshold, and the cleaned data includes: close-up images of faces, half-body images and full-body images, data enhancement processing is not required; if the number of images in the cleaned data is less than or equal to the set threshold, the cleaned data is enhanced to obtain enhanced data, and the enhanced data includes: close-up images of faces, half-body images and full-body images, and the number ratio is 2:1:1; the number of images in the enhanced data is greater than the set threshold; each image in the enhanced data or cleaned data is processed to ensure that the long side of each image is the set value to obtain training data; the images in the training data are formatted and labeled through a multimodal model to form a label file, and the formatted format is: gender, skin, eyes, hairstyle, mouth shape, expression, accessories, clothing and background; set the number of training rounds, learning rate and training image pixels, and then input the training data and label file to perform Lora model training to obtain an exclusive model model; The edge information, posture information, the second prompt word and the white background portrait are input into the diffusion model, and guided by the exclusive model model to generate a real model product image. The parameters of the diffusion model are set as follows: the sampler is dpmpp_2m; the scheduler is Karras; the redrawing amplitude is 0.7; the sampling step number is 20, and the prompt word correlation coefficient is 7.
5. A device for generating a model product image based on a mannequin image, characterized in that: include: An upload module is used to upload a mannequin image with set clothing and a real person image with set exposed skin area; Get the clothing mask image from the mannequin image; The preliminary generation module inputs the real person image into the Redux model, inputs the clothing mask image, the mannequin image and the Redux model into the collector of the Flux_fill model, and then generates the image through the Flux_fill model to obtain a preliminary portrait image; The cutout module cuts out the preliminary portrait image to obtain the main portrait image, and synthesizes the main portrait image with the white background image to obtain the white background portrait image; The guidance information module extracts edge information of the portrait image with white background through the lineart model; and extracts posture information of the portrait image with white background through the openpose model; The final generation module inputs the set reference portrait image into the iPadapter model to extract the portrait features; The edge information, posture information, white background portrait image and portrait features are input into the diffusion model to generate a real model product image.
6. The device for generating a model product image based on a mannequin image according to claim 5, characterized in that: The preliminary generation module is specifically as follows: inputting the real-person image into the Redux model, the weight control parameter of the Redux model is set to 0.9, inputting the clothing mask image, the first set prompt word, the mannequin image and the Redux model into the collector of the Flux_fill model, and then performing image generation through the Flux_fill model to obtain a preliminary portrait image; the parameters in the Flux_fill model are set as follows: the sampler is euler; the scheduler is simple; the redrawing amplitude is 0.9; the sampling step number is 15, and the prompt word correlation coefficient is 2; the first set prompt word is used to describe limb details.
7. The device for generating a model product image based on a mannequin image according to claim 5, characterized in that: The guidance information module specifically comprises: extracting edge information of the white background portrait image through a lineart model, wherein the weight parameter of the lineart model is set to 0.6; extracting posture information of the white background portrait image through an openpose model, wherein the weight parameter of the openpose model is set to 0.9; The final generation module specifically includes: inputting the set reference portrait image into the iPad Adapter model to extract the portrait features; The edge information, posture information, the second set prompt word, the white background portrait image and the portrait features are input into the diffusion model to generate a real model product image. The parameters of the diffusion model are set as follows: the sampler is dpmpp_2m; the scheduler is Karras; the redrawing amplitude is 0.7; the sampling step number is 20, and the prompt word correlation coefficient is 7.
8. The device for generating a model product image based on a mannequin image according to claim 5, characterized in that: It also includes a specific generation module, when it is necessary to generate a real model product picture of a set face, it obtains picture data of a set model, and performs data cleaning to obtain cleaned data, wherein the picture data includes a close-up picture of the face; If the number of images in the cleaned data is greater than the set threshold, and the cleaned data includes: close-up images of faces, half-body images and full-body images, data enhancement processing is not required; if the number of images in the cleaned data is less than or equal to the set threshold, the cleaned data is enhanced to obtain enhanced data, and the enhanced data includes: close-up images of faces, half-body images and full-body images, and the number ratio is 2:1:1; the number of images in the enhanced data is greater than the set threshold; each image in the enhanced data or cleaned data is processed to ensure that the long side of each image is the set value to obtain training data; the images in the training data are formatted and labeled through a multimodal model to form a label file, and the formatted format is: gender, skin, eyes, hairstyle, mouth shape, expression, accessories, clothing and background; set the number of training rounds, learning rate and training image pixels, and then input the training data and label file to perform Lora model training to obtain an exclusive model model; The edge information, posture information, the second prompt word and the white background portrait are input into the diffusion model, and guided by the exclusive model model to generate a real model product image. The parameters of the diffusion model are set as follows: the sampler is dpmpp_2m; the scheduler is Karras; the redrawing amplitude is 0.7; the sampling step number is 20, and the prompt word correlation coefficient is 7.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Cited By
AI model chart generation method, system and device and medium
CN121685752A