This application relates to an AI
model image generation method,
system, device, and medium, belonging to the technical field of
image processing. It involves collecting
multimodal data of the target garment, including mannequin images, flat static images, and material close-up images; selecting a
reference image from a target
pose library whose
pose matches the desired
pose, and extracting the corresponding pose skeleton image as the target pose skeleton image; inputting the flat
static image and the material close-up image into a visual
language model to generate a text description; and inputting the target pose skeleton image, mannequin image, and text description into a pre-trained LoRA
diffusion model to generate a preliminary dressed image. The LoRA
diffusion model's input layer includes an independent pose condition injection channel, through which the target pose skeleton image is input into the LoRA
diffusion model. This application has the beneficial technical effect of avoiding stiff and unnatural AI model poses, thereby improving the aesthetics of the generated images.