An instruction-driven personalized fashion image editing method

By building a multi-editing task dataset and designing a unified editing framework, combining multi-task low-rank global selection mechanism, semantic control network and visual joint module, the problem of sparse datasets and confusion in intent in fashionable image editing is solved, and high-quality and personalized fashionable image editing effect is achieved.

CN119693505BActive Publication Date: 2025-05-30HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510211261.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-30
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The prior art has problems in fashion image editing with sparse data sets, compatible with different editing tasks, and confusing intentions, resulting in inaccurate editing results.

Method used

By building a multi-editing task dataset, a unified editing framework is designed, and a multi-task low-rank global selection mechanism is adopted, combining semantic control networks and visual joint modules, the unified and high-quality editing effect of multi-editing tasks is achieved.

Benefits of technology

The integration of multi-editing tasks and high-quality editing effects are realized. The generated fashionable images have realistic effects and highly personalized effects, solving the problems of dispersed and difficult integration of multi-task processing, and improving the editing implementation capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693505B_ABST
    Figure CN119693505B_ABST
Patent Text Reader

Abstract

The present invention discloses an instruction-driven personalized fashion image editing method. The present invention: 1. Defines the categories of editing tasks, and constructs a four-tuple data group of "original image-reference image-target image-text editing instruction" for different editing tasks; 2. Constructs a target semantic network to generate the target image semantic information that follows the editing instruction and the original image, and uses this as the human body semantic information of the editing model; 3. Constructs a unified editing network, including constructing a semantic control network, adding a visual joint module, and applying a low-rank fine-tuning module, to enable multiple editing tasks to obtain corresponding editing capabilities using the same framework; 4. Constructs a multi-task low-rank adjustment module, and through joint training, enables the framework to have the ability to align different editing instructions to different editing tasks. Finally, an independent and unified framework among different tasks is achieved. The present invention has conducted experiments on the constructed specific dataset and achieved good results both quantitatively and qualitatively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image editing, and particularly relates to an instruction-driven personalized fashion image editing method.

[0002] The present invention mentions a new method for instruction-driven customized fashion image editing. It mainly involves a low-rank fine-tuning technique applied to a pre-trained model to understand and implement editing intentions for different editing tasks; and a multi-task low-rank global selection mechanism is designed to provide a unified framework for the implementation of different editing tasks. Thus, it is possible to be compatible with multiple editing tasks under one framework and obtain real and effective fashion image editing results. Background Art

[0003] In recent years, the rapid development of diffusion models and the evolution of controllable generation technologies have enabled image editing technologies to make great progress in terms of authenticity, controllability, and diversity. As an important part of image editing technologies, fashion editing tasks pay more attention to the preservation of human body details and the reconstruction of fashion features, so they have always received high attention in the industrial and academic fields.

[0004] Fashion image editing tasks have certain differences in focus from traditional image editing tasks. The specific editing tasks solved by previous work mainly include two-dimensional pose transfer, fashion attribute modification, and human-centered clothing reconstruction. And the driving methods mainly include text-driven, image-driven, and multi-modal-driven. In order to cover as many editing tasks as possible and have the most compatible driving method, the invention attempts to design a unified fashion image editing framework, which has three characteristics: instruction-driven, personalized support, and task compatibility. However, the constraints of real resources and the limitations of technical conditions pose many challenges.

[0005] 1) Lack of fashion editing datasets

[0006] In order to overcome the limitations associated with the original generation framework, such as non-intentional editing and other problems, fashion editing tasks require original images, editing instructions, and high-quality edited images that follow the editing instructions as training data. Complete edited graphic data pairs are difficult to collect due to their sparsity in Internet data. Existing original data is difficult to directly use in terms of integrity and quality, and there is a large tilt in categories, which is not conducive to model training. Therefore, constructing a large number, high-quality, wide-category, and weakly tilted dataset for various editing tasks is the key to the work.

[0007] 2) Framework design compatible with different editing tasks

[0008] Different editing tasks may require different auxiliary information to enhance editing controllability. Therefore, it is necessary to determine a combination of auxiliary information that is compatible with different tasks to achieve editing guidance among different tasks.

[0009] 3) Effective fusion training strategy

[0010] There are obvious differences in editing intentions among different editing tasks. Experiments have shown that simple sample mixing training cannot enable the generative model to achieve different editing tasks while maintaining prior knowledge. Specifically, the editing results show the effects of other editing tasks, and this phenomenon is called intention confusion. Therefore, effectively overcoming intention confusion and enabling the model to obtain the ability to be compatible with different editing intentions has become an important challenge for unifying the fashion editing framework.

[0011] The present invention is committed to integrating multiple editing capabilities into a unified framework through a small amount of training based on the existing pre-trained image generation model, and generating fashion image editing results with high degrees of freedom, high authenticity, and high accuracy. Summary of the Invention

[0012] The present invention provides an instruction-driven personalized fashion image editing method. By constructing a multi-editing task dataset, designing a unified framework for multi-editing tasks, training multiple groups of task-independent editing modules, and realizing the unification of multi-editing tasks through a multi-task low-rank global selection mechanism, fashion images with realistic effects and accurate editing are generated. The invention has conducted experiments on the constructed specific dataset and achieved good results both quantitatively and qualitatively.

[0013] An instruction-driven personalized fashion image editing method, the steps of which are as follows:

[0014] Step (1), define the categories of editing tasks. For different editing tasks, use web crawler technology and existing image generation technology to construct a four-tuple data group of "original image - reference image - target image - text editing instruction".

[0015] Step (2), construct and train the target semantic network; use the text editing instruction, reference image, and original image as the input of the target semantic network, extract the semantic information of the target image as the training target, obtain the semantic information of the target image that follows the text editing instruction, and use the semantic information of the target image as the human body semantic information of the editing model.

[0016] Step (3): Construct a unified editing network; make three changes based on the pre-trained instruction-driven generation model; First, construct a semantic control network to achieve effective guidance based on the target image semantic information obtained in step (2); Second, add a visual joint module to reconstruct the reference image in the editing result; Third, apply a low-rank fine-tuning module to align the understanding and implementation of the editing intention to the low-rank fine-tuning module through a small amount of training; thus effectively achieving the corresponding editing capabilities for multiple editing tasks using the same framework;

[0017] Step (4): Construct a multi-task low-rank fine-tuning module; through joint training, enable the framework to have the ability to align different text editing instructions to different editing tasks; ultimately, achieve an independent and unified fashion editing framework between different tasks.

[0018] Advantages of the present invention:

[0019] The present invention proposes an instruction-driven personalized fashion image editing method, which realizes the integration of multiple editing tasks and high-quality editing effects by constructing a multi-editing task dataset, designing a unified editing framework, and a low-rank global selection mechanism. Through the multi-task low-rank global selection module, the independence and unity between different editing tasks are realized, and the problem of scattered multi-task processing and difficulty in integration in the existing methods is solved. By using the semantic control network and the visual joint module, the editing implementation ability is significantly improved, and the generated fashion images have realistic effects and high personalization. Through the low-rank fine-tuning technology and the unified framework, multiple tasks can be quickly adapted through a small amount of training, making the method efficient, flexible, and scalable. Experiments show that this method achieves excellent performance in both quantitative and qualitative evaluations, and is applicable to multiple scenarios such as personalized clothing matching, virtual fitting, and image editing in the e-commerce and fashion industries, and has broad practical application value. Brief Description of the Drawings

[0020] Figure 1 is a schematic diagram of the specific process of the method of the present invention.

[0021] Figure 2 is a schematic diagram of the target semantic network in the present invention.

[0022] Figure 3 is a schematic diagram of the low-rank fine-tuning module in the method of the present invention.

[0023] Figure 4 is a schematic diagram of the multi-task low-rank fine-tuning module in the method of the present invention.

[0024] Figure 5 is a schematic diagram of the unified editing network of the present invention.

[0025] Figure 6 is a schematic diagram of the data group of the present invention.

[0026] Figure 7 This is a schematic diagram of the effects of the present invention. Specific embodiments

[0027] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0028] The present invention provides an instruction-driven personalized fashion image editing method. As Figure 1 shown, this method constructs a multi-editing task dataset, designs a unified framework for multi-editing tasks, trains multiple groups of task-independent editing modules, and realizes the unification of multi-editing tasks through a multi-task low-rank global selection mechanism, generating realistic and accurately edited fashion images. The invention has conducted experiments on the constructed specific dataset and achieved good results both quantitatively and qualitatively.

[0029] An instruction-driven personalized fashion image editing method, the steps of which are as follows:

[0030] Step (1): Define the categories of editing tasks. For different editing tasks, use web crawler technology and existing image generation technology to construct a four-tuple data group of "original image - reference image - target image - text editing instruction".

[0031] Step (2): Construct and train a target semantic network; use the text editing instruction, reference image, and original image as the input of the target semantic network, extract the semantic information of the target image as the training target, obtain the semantic information of the target image following the text editing instruction, and use the semantic information of the target image as the human body semantic information of the editing model.

[0032] Step (3): Construct a unified editing network; make three changes on the basis of a pre-trained instruction-driven generation model; First, construct a semantic control network to achieve effective guidance based on the semantic information of the target image obtained in step (2); Second, add a visual joint module to reconstruct the reference image in the editing result; Third, apply a low-rank fine-tuning module to align the understanding and implementation of the editing intention to the low-rank fine-tuning module through a small amount of training; thereby effectively realizing that multiple editing tasks can obtain corresponding editing capabilities using the same framework.

[0033] Step (4): Construct a multi-task low-rank fine-tuning module; through joint training, enable the framework to have the ability to align different text editing instructions to different editing tasks; ultimately, realize an independent and unified fashion editing framework between different tasks.

[0034] The categories of the defined editing tasks in step (1) are addition, deletion, replacement, and modification. Through web crawler technology and various existing advanced image synthesis technologies, a large number of text-image data pairs with noise were obtained; after manual screening, 36,312 edited text-image data pairs were finally obtained, including 10,244 replacement tasks, 10,072 modification tasks, and 7,998 addition and deletion tasks each; in terms of categories, ADCR involves four basic editing tasks, including 8 item categories: upper wear, lower wear, full wear, shoes, glasses, belts, bags, and scarves. The data pairs for different tasks are referenced Figure 6 . The final dataset consists of the original image I src , the reference image I ref , the text editing instruction T 1 , and the target image I tar . The original image and the reference image are both obtained by the crawler, and the text editing instruction is automatically generated by the large language model based on the comparison between the original image and the target image; the acquisition of the target image has the following concentrated situations: for addition or deletion tasks, the original image containing the reference image is screened, and the reference image is removed through SD1.5-Inpaint; the target image in the replacement task is the existing virtual fitting model; the target image in the modification task is modified by the editing model and then synthesized using the virtual fitting model.

[0035] The image synthesis technology mentioned in step (1) is related to specific editing tasks. The replacement data pairs select items of the same category as the replacement target from the original image as the reference image, and the target image is obtained through advanced virtual fitting tools; the addition and deletion data pairs are essentially the same, and the reference item in the original image is removed using an advanced image-to-image inpainting model to obtain the target image; limited by the accuracy of existing editing tools, the reference image is changed in its attributes through the instruction editing model and then the target image is obtained through the virtual fitting tool, thus indirectly realizing the modification data pairs. The existing multi-modal large language model is used to analyze the original image and the target image, and then the editing instruction data with both diversity and accuracy is obtained. The final dataset consists of the original image I src , the reference image I ref , the text editing instruction T 1 , and the target image I tar . Specific examples can be referenced Figure 6 . Among them, the replacement task and the addition task contain new concepts, so there are reference images in the data pairs; while the deletion task and the modification task do not include reference images.

[0036] The human semantic information described in step (2) contains several human body region tags, and the number of human semantic tags is 18, including background, hat, hair, sunglasses, upper garment, skirt, trousers, dress, belt, left shoe, right shoe, face, left leg, right leg, left arm, right arm, bag, and scarf regions, effectively representing the structural information of the human body and related outfits; the target semantic network is essentially an image-to-image model, using the StableDiffusion model as the basic model, and using the original image in the data pair after forward noise addition as the input to the basic model; the target semantic network, as the prior of the human semantic guidance of the editing model, needs to generate a target semantic image P that follows the text editing instruction T 1 and conforms to the original human semantics tar , and the overall network structure refers to Figure 2 .

[0037] Use the image feature extractor CLIP to obtain the 24-layer feature vector of the reference image, that is, the image embedding; take the average of the first 6-layer feature vectors representing the low-frequency information of the reference image, and pass it through a mapping network composed of a group of MLPs to transform it into the text embedding space and then mix it with the text embedding to obtain the final text editing instruction T 2 , thus completing the mixed embedding

[0038] The training loss adds a regularization constraint on the feature vector on the basis of the diffusion loss; the specific formula is as follows

[0039]

[0040] Among them, E is the VAE variational autoencoder in the StableDiffusion model, z ∼ ε(x) represents the latent variable z obtained by encoding the input image x through the variational autoencoder, ε(x) represents the noise distribution extracted from the image x, t represents the time step, ∈ is the noise randomly sampled from the normal distribution, S(T, v) is a function that combines the text template T with the image embedding v to generate a mixed embedding for guiding text generation; here T 2 represents the final text editing instruction, and v is the image embedding, representing the features extracted from the image; λ visual is a hyperparameter used to adjust the strength of the regularization constraint of the feature vector; CLIP(I ref ) n represents the nth layer feature vector extracted from the reference image I ref using the image feature extractor CLIP, I ref is the reference image, which is used to provide visual information for training in combination with text, and MLP represents the multi-layer perceptron

[0041] As Figure 5 shown, the unified editing network described in step (3) is specifically implemented as follows

[0042] Input the original image I src into the unified editing network, and it is converted into the original latent representation z by the variational autoencoder ε built in the unified editing network src . Then, randomly sample noise from the normal distribution as the initial noise z T . Input the initial noise z T and the original latent representation z src along the channel dimension for concatenation to obtain the latent representation z f .

[0043] After passing through the target semantic network, the original image obtains the target semantic image P tar . The semantic control network converts the target semantic image P tar into the semantic latent representation and extracts the semantic feature vector sequence and serves as an input to the unified editing network;

[0044] After extracting the reference image features through the image feature extractor CLIP, input them into the visual joint module of the unified editing network;

[0045] The reference image features interact with the latent representation z during the editing process in the cross-attention layer of the unified editing network f to provide the visual feature information of the reference image; for each different editing task, each attention layer in the unified editing network applies a low-rank fine-tuning module. Finally, through T' steps of denoising and the decoding of the variational autoencoder, the latent representation z f obtained after T' steps of denoising in the unified editing network t0 is decoded and output as the edited image I edit .

[0046] Part 1: Semantic control network;

[0047] The described semantic control network uses the encoder part of the generative network as the basic structure, and the initial weights are copies of the corresponding generative network weights; the semantic control network uses the target semantic image P tar as the input and the final text editing instruction T 2 as the text guidance; the entire semantic control network is regarded as a semantic information feature extractor. The latent representation Z tl at each time step t and each layer l is used as the target semantic feature. After being mapped through the corresponding zero convolutional layer of each layer, it is added as a residual to the corresponding latent representation Z' tl in the decoder of the generative network to guide the generation of the target image that conforms to the text instruction.

[0048] To reduce the training burden, the training of the semantic control network will be independent of the unified editing network; a self-supervised training method is adopted for the semantic control network; specifically, the StableDiffusion model is used as the generation network, and the target image I tar is transformed into the target latent representation z through the variational autoencoder ∈ tar and after adding noise, it is used as the input z of the generation network t . The text guidance for the denoising process is the target image description T generated by the multimodal large language model desc ; the semantic control network transforms the target semantic image P tar into a semantic latent representation and extracts a sequence of semantic feature vectors and after zero convolution Zero mapping, it is added to the corresponding latent representation layer by layer in the decoder of the generation network as a residual to provide semantic guidance for the editing process; during the training process, all parameters except those of the semantic control network are frozen, and only all parameters of the semantic control network are updated; the training loss adopts the diffusion loss, and the formula is defined as follows:

[0049]

[0050] where z ∼ ε(x) represents the latent representation z obtained after the input image x is transformed by the variational autoencoder (VAE); T desc represents the description of the target image; ∈ ∼ N(0, 1) is the noise sampled from the standard normal distribution, which is added to the latent representation z and used for the denoising process of generating the image; t represents the time step in the diffusion process; P tar represents the target semantic image

[0051] Part Two: Visual Joint Module

[0052] The described visual joint module is an augmented network for the basic cross-attention layer; first, the cross-attention layer in the unified editing network is defined by the following formula:

[0053]

[0054] where Q, K t , V t respectively represent the original query feature, key feature, and value feature; the reference K t , V t are derived from the design of the text embedding. The visual joint module proposes learnable weight matrices and and maps the reference image I ref extracted by the visual encoder CLIP into the visual key feature K s and the visual value feature V s

[0055] ​Therefore, the hybrid cross-attention layer augmented by the visual association module is defined by the following formula:

[0056]

[0057] where d k represents the dimensions of the key and query features. In the attention mechanism, d k is used to scale the dot product result to prevent gradient instability caused by large dimensions. K s and V s represent the key and value features introduced by the visual association module, which are obtained by multiplying the visual features extracted from the reference image I ref through a visual encoder (such as CLIP) with the weight matrices and respectively.

[0058] Part 3: Low-rank fine-tuning module;

[0059] Apply the low-rank fine-tuning module to strip the editing capabilities of different editing tasks from the base generation network and transfer them to the corresponding low-rank adapters, so as to achieve the effect of mutual independence and non-interference between different editing capabilities, as follows:

[0060] Apply the low-rank fine-tuning module to the mapping layer, self-attention layer, and cross-attention layer in each attention layer of the unified editing network; the low-rank fine-tuning module achieves low-cost fine-tuning by adding low-rank residual networks to all linear layers and convolutional layers, as shown in Figure 3 and its definition can be abstracted as the following formula:

[0061] h(x) = W 0 x + λΔWx = W 0 x + λBAx, (6)

[0062] where h represents the forward calculation of any linear layer or convolutional layer, x is the input image of the unified editing network; λ represents the weight of the low-rank fine-tuning module, representing the activation intensity of the low-rank fine-tuning module; W 0 represents the weight of the original linear layer or convolutional layer, with dimensions a×b; ΔW represents the low-rank fine-tuning module, which is composed of B·A, where B and A are linear layers or convolutional layers with dimensions a×r and r×b respectively, and r is the rank of the low-rank fine-tuning module.

[0063] Fine-tune the corresponding unified editing network for different editing tasks to obtain a set of low-rank fine-tuning modules corresponding to the editing tasks; during training, freeze all parameters except the low-rank residual network, that is, only keep the parameter updates of the low-rank residual network; the training loss L task includes the diffusion loss and the editing localization loss, and the formula is defined as follows:

[0064] L task = L diff + L loc ,(7)

[0065]

[0066] The L represented by formula (8) diff diffusion loss, where represents the unified editing network optimized by the low-rank fine-tuning module; z ∼ ε(x) represents the latent representation z obtained by transforming the input image x through the variational autoencoder ε; T desc represents the text description of the target image; ∈ ∼ N(0,1) represents the noise sampled from the standard normal distribution; t represents the time step in diffusion, controlling the process of the generation process; z t represents the noise-added latent representation at time step t, serving as the input to the generation network; z src represents the latent representation of the original image; I ref represents the reference image, serving as a condition for the denoising process; P tar represents the target semantic image.

[0067] The editing localization loss L represented by formula (9) loc ; In the single-step denoising process, the cross-attention map of each layer is defined as That is, the attention map corresponding to the k-th token in the l-th layer, and the total number of tokens is K; A in the editing localization loss l is the average attention map of the k-th token in the l-th layer; Since the sizes of the attention maps of different layers are not all the same, therefore, according to experience, the average attention map A l is uniformly scaled to a resolution of 16*16; M in the editing localization loss represents the editing region mask, with a size of 16*16 resolution; This mask data is an augmentation of the basic dataset and is obtained through the SAM segmentation model; Through the editing localization loss, this method focuses the attention of the unified editing network's denoising process on the editing subject that follows the text editing instructions, effectively reducing the phenomenon of non-intended editing.

[0068] As Figure 4 shown, the multi-task low-rank fine-tuning module described in step (4) is essentially an MLP multi-layer perceptron for classification tasks, using the text embedding corresponding to the text editing instruction T 1 as the input, aligning the feature dimension of the text embedding with the dimension of the number of tasks, and calling the output of the multi-task low-rank fine-tuning module the low-rank weight λ; The set of low-rank fine-tuning modules trained for independent editing tasks is weighted and merged, and the mapped low-rank weight λ is used as the activation degree of the low-rank fine-tuning module corresponding to the task; Therefore, the abstract definition of formula (6) regarding the low-rank fine-tuning module is updated to the following formula:

[0069]

[0070] h(x) represents the output generated by the unified editing network; W 0 represents the initial linear transformation matrix or weight matrix, which is applied to the input image x; this matrix is fixed during network training and does not participate in low-rank fine-tuning; λ i represents the low-rank weight of the i-th low-rank fine-tuning module, which is the output result of the low-rank fine-tuning module SEL; the low-rank weight reflects the activation degree of the task-corresponding low-rank residual network in each independent editing task; ΔW i represents the weight change of the i-th low-rank fine-tuning module, that is, the deviation part generated by the low-rank fine-tuning process; B i 、A i represents the matrix decomposition part of the i-th low-rank fine-tuning module;

[0071] Among them, the low-rank weight λ = [λ 1 , λ 2 , λ 3 , λ 4 , represents 4 groups of independently trained low-rank fine-tuning modules;

[0072] Finally, a joint training strategy is designed to perform end-to-end self-supervised training on the multi-task low-rank fine-tuning module; the training samples used in the joint training are randomly sampled from the dataset of the "original image - reference image - target image - text editing instruction" four-tuple data; in order to make the activation degree of the low-rank fine-tuning module fully consider the balance between the specific editing intensity and the prior knowledge of the editing model, the present invention maintains the diffusion loss on the basis of the classification loss, and the specific selection loss L select is defined as follows:

[0073]

[0074] Among them, L″ diff is similar to formula (8), but the calculation of the combined low-rank fine-tuning module is updated to formula (10); τ θ represents the text feature extraction module identical to the unified editing network, that is, CLIP; λ * is a one-hot encoding related to the editing task of the training sample, which ensures that the correct low-rank fine-tuning module is activated while minimizing the activation degree of the wrong editing ability; SEL represents the low-rank fine-tuning module;

[0075]

[0076] L″ diff represents the diffusion loss, where Denotes the unified editing network optimized by the low-rank fine-tuning module; z ∼ ε(x) represents the latent representation z obtained by transforming the input image x through the variational autoencoder ε; T desc Represents the text description of the target image; ∈ ∼ N(0, 1) represents the noise sampled from the standard normal distribution; t represents the time step in diffusion, controlling the progress of the generation process; z t Denotes the noise-added latent representation at time step t, serving as the input to the generation network; z src Represents the latent representation of the original image; I ref Denotes the reference image, serving as a condition for the denoising process; P tar Denotes the target semantic image.

[0077] It is worth noting that the unified framework of this invention has extremely strong scalability. Facing richer fashion editing tasks, such as pose transfer tasks involving global editing, the low-rank fine-tuning network for the corresponding new task can also be trained using step (3), and compatibility with the new task can be achieved after the multi-task low-rank global selection module in the simple training step (4).

[0078] Embodiment:

[0079] In actual experiments, images can be edited more precisely through this unified editing framework. As Figure 7 shown. The first column is the original image, the second column is the reference image, the third column is the editing instruction, the fourth column is the target semantic image generated according to the editing instruction, the fifth column is the mask of the editing area, and the last column is the finally generated image. A total of 7 groups of experimental results are listed. The first group is the addition task, providing the original image, the editing instruction (holding a bag in the right hand), and the reference image (the bag), and finally a person image holding the bag can be generated; the second group is the addition of a person, providing the original image, the reference image (scarf), and the editing instruction (put on the scarf), and finally an image with the scarf can be generated; the third group is the editing task, providing the original image and the editing instruction (change the color of the pants), and finally the original image with the color of the pants changed is generated; the fourth group is the editing task, providing the original image and the editing instruction (change the belt to pure black), and finally the original image with a black belt is generated; the fifth group is the deletion task, providing the original image (a woman wearing glasses) and the editing instruction (take off the glasses), and finally a person image with the glasses taken off can be generated; the sixth group is the replacement task, providing the original image, the reference image (pants of another style), and the editing instruction (change pants), and finally the original image with the pants of the reference image changed is generated; the seventh group is the replacement task, providing the original image, the reference image (top), and the editing instruction (change the top), and finally an original image with the top changed can be generated.

Claims

1. A command-driven personalized fashion image editing method, characterized in that: The steps include: Step (1), defining the categories of editing tasks, and constructing a four-dimensional data set of "original image - reference image - target image - text editing instruction" by using web crawler technology and existing image generation technology for different editing tasks; Step (2), constructing and training a target semantic network; using the text editing instructions, the reference image and the original image as inputs of the target semantic network, extracting the semantic information of the target image as a training target, obtaining the semantic information of the target image following the text editing instructions, and using the semantic information of the target image as the human body semantic information of the editing model; Step (3), build a unified editing network; add three changes based on the pre-trained instruction-driven generation model; First, construct a semantic control network to achieve effective guidance based on the semantic information of the target image obtained in step (2); second, add a visual joint module to reconstruct the reference image in the editing result; third, apply a low-rank fine-tuning module to align the understanding and implementation of the editing intention to the low-rank fine-tuning module through a small amount of training; thereby effectively achieving the corresponding editing capabilities of multiple editing tasks using the same framework; The low-rank fine-tuning module adds a low-rank residual network to all linear layers and convolutional layers. Its definition can be abstracted into the following formula: h(x)=W0x+λΔWx=W0x+λBAx,(1) Where h represents the forward computation of any linear layer or convolutional layer, x is the input image of the unified editing network; λ represents the weight of the low-rank fine-tuning module; W0 represents the weight of the original linear layer or convolutional layer, whose dimension is a×b; ΔW represents the low-rank fine-tuning module, which is composed of B·A, where B and A are linear layers or convolutional layers with corresponding dimensions of a×r and r×b, respectively, and r is the rank of the low-rank fine-tuning module; Step (4), construct a multi-task low-rank fine-tuning module; through joint training, the framework is equipped with the ability to align different text editing instructions to different editing tasks; finally, an independent and unified fashion editing framework between different tasks is realized; the multi-task low-rank fine-tuning module is essentially an MLP multi-layer perceptron for classification tasks, using the text embedding corresponding to the text editing instruction T1 as input, aligning the feature dimension of the text embedding with the dimension of the number of tasks, and calling the output of the multi-task low-rank fine-tuning module the low-rank weight λ; weightedly merging the set of low-rank fine-tuning modules obtained by independent editing task training, and using the mapped low-rank weight λ as the activation degree of the low-rank fine-tuning module of the corresponding task; therefore, the abstract definition of the low-rank fine-tuning module in formula (1) is updated to the following formula: h(x) represents the output generated by the unified editing network; W0 represents the initial linear transformation matrix or weight matrix, which is applied to the input image x; λ i represents the low-rank weight of the i-th low-rank fine-tuning module, which is the result of the output of the low-rank fine-tuning module SEL; ΔW i represents the weight change of the i-th low-rank fine-tuning module, that is, the deviation part generated by the low-rank fine-tuning process; B i , A i represents the matrix decomposition part of the i-th low-rank fine-tuning module; Among them, the low-rank weight λ=[λ1,λ2,λ3,λ4], Represents 4 groups of independently trained low-rank fine-tuning modules.

2. The instruction-driven personalized fashion image editing method according to claim 1, characterized in that The categories of the defined editing tasks in step (1) are add, delete, replace, and modify. A large number of noisy image-text data pairs are obtained through web crawler technology and various existing advanced image synthesis technologies. After manual screening, 36,312 pairs of edited image-text data pairs are finally obtained, of which 10,244 are replacement tasks, 10,072 are modification tasks, and 7,998 are add and delete tasks respectively. From the category point of view, ADCR involves four basic editing tasks, including 8 categories of items: tops, bottoms, full clothes, shoes, glasses, belts, bags, and scarves. The final dataset consists of the original image I src , reference image I ref , text editing instruction T1, target image I tar The target image is composed of a four-tuple; the original image and the reference image are both obtained by the crawler, and the text editing instructions are generated by the large language model based on the comparison between the original image and the target image; the acquisition of the target image is concentrated in the following situations: for adding or deleting tasks, the original image containing the reference image is filtered, and the reference image is removed through SD1.5-Inpaint; the target image in the replacement task uses the existing virtual fitting model; the target image in the modification task is modified by the editing model and then synthesized using the virtual fitting model.

3. The instruction-driven personalized fashion image editing method according to claim 1, characterized in that The human body semantic information described in step (2) includes several human body region labels, and the number of human body semantic labels is 18, including background, hat, hair, sunglasses, top, skirt, pants, dress, belt, left shoe, right shoe, face, left leg, right leg, left arm, right arm, bag, and scarf regions, which effectively represent the structural information of the human body and related clothing. The target semantic network is essentially a graph-generated graph model, which uses a stable diffusion model as the basic model and uses the original image in the data pair as the basic model input after forward noise addition. The target semantic network is the human body semantic guidance prior of the editing model, and needs to generate a target semantic image P that follows the text editing instruction T1 and conforms to the original human body semantics. tar ; The image feature extractor CLIP is used to obtain the 24-layer feature vector of the reference image, i.e., the image embedding. The first 6 layers of feature vectors representing the low-frequency information of the reference image are averaged and converted to the text embedding space through a mapping network composed of a group of MLPs and then mixed with the text embedding to obtain the final text editing instruction T2 containing the image features, thus completing the mixed embedding. The training loss adds a regularization constraint on the feature vector based on the diffusion loss. The specific formula is as follows: Where E is the VAE variational autoencoder in the stable diffusion model, z~ε(x) represents the latent variable z obtained by encoding the input image x through the variational autoencoder, ε(x) represents the noise distribution extracted from the image x, t represents the time step, ∈ is the noise randomly sampled from the normal distribution, S(T,v) is a function that combines the text template T with the image embedding v to generate a hybrid embedding for guiding text generation; here T2 represents the final text editing instruction, and v is the image embedding, which represents the features extracted from the image; λ visual is a hyperparameter used to adjust the strength of the eigenvector regularization constraint; CLIP(I ref ) n Represents the use of image feature extractor CLIP from the reference image I ref The nth layer feature vector extracted from ref is a reference image used to provide visual information for training in conjunction with text, and MLP stands for multi-layer perceptron.

4. The instruction-driven personalized fashion image editing method according to claim 1, characterized in that The unified editing network described in step (3) is specifically implemented as follows: The original image I src After being input into the unified editing network, the variational autoencoder ∈ of the unified editing network is converted into the original potential representation z src , and then randomly sample noise from the normal distribution as the initial noise z T , the initial noise z T With the original latent representation z src Concatenate in the channel dimension to get the potential representation z f ; The original image passes through the target semantic network to obtain the target semantic image P tar , the semantic control network transforms the target semantic image P tar After being converted into semantic latent representation, a sequence of semantic feature vectors is extracted and serves as an input to the unified editing network; After the reference image features are extracted by the image feature extractor CLIP, they are input into the visual joint module of the unified editing network; Reference image features in the cross-attention layer of the unified editing network and the latent representation z during the editing process f For each different editing task, a low-rank fine-tuning module is applied to each attention layer in the unified editing network, and finally the potential representation z is transformed through T'-step denoising and decoding of the variational autoencoder. f The latent representation z obtained after T' steps of denoising in the unified editing network t0 , and decoded output as edited image I edit .

5. The instruction-driven personalized fashion image editing method according to claim 4, characterized in that The semantic control network uses the encoder part of the generative network as the basic structure, and the initial weights are copies of the corresponding generative network weights; The semantic control network uses the target semantic image P tar As input, the final text editing instruction T2 is used as the text guidance; the entire semantic control network is regarded as a semantic information feature extractor, and its potential representation Z of each layer l at each time step t is tl As the target semantic feature, after being mapped by the corresponding zero convolution layer of each layer, it is used as the corresponding potential representation Z' in the decoder of the residual and generative network tl Addition guides the generation of target images that conform to text instructions; Use the stable diffusion model as the generating network to transform the target image I tar Transformed into target potential representation z through variational autoencoder ∈ tar And add noise as the input z of the generating network t , the textual guidance of the denoising process is the target image description T generated by a multimodal large language model desc ; The semantic control network transforms the target semantic image P tar After being converted into semantic latent representation, a sequence of semantic feature vectors is extracted After zero convolution and Zero mapping, the residual is added to the corresponding potential representation layer by layer in the decoder of the generative network to provide semantic guidance for the editing process. The training process freezes all parameters except the semantic control network and only updates all parameters of the semantic control network. The training loss adopts diffusion loss, and the formula is defined as follows: Where z~ε(x) represents the potential representation z obtained after the input image x is transformed by the variational encoder; T desc represents the description of the target image; ∈~N(0,1) is the noise sampled from the standard normal distribution, t represents the time step in the diffusion process; P tar Represents the target semantic image.

6. The instruction-driven personalized fashion image editing method according to claim 4, characterized in that The visual joint module is an augmented network of the basic cross-attention layer; first, the cross-attention layer in the unified editing network is defined as follows: Among them, Q, K t ,V t Represent the original query features, key features, and value features respectively; refer to K t ,V t Design derived from text embedding; visual joint module proposes a learnable weight matrix and The reference image I extracted by the visual encoder CLIP ref Embedding map is the visual key feature K s and visual value feature V s ; Therefore, the hybrid cross attention layer augmented by the visual joint module is defined as follows: Among them, d k Represents the dimensions of key and query features, K s and V s Represents the key and value features introduced by the visual joint module, which are composed of the reference image I ref Visual features and weight matrix extracted by visual encoder and Multiply them together to get .

7. The instruction-driven personalized fashion image editing method according to claim 4, characterized in that: Apply the low-rank fine-tuning module to the mapping layer, self-attention layer, and cross-attention layer in each attention layer of the unified editing network; Fine-tune the corresponding unified editing network for different editing tasks to obtain a set of low-rank fine-tuning modules corresponding to the editing tasks; freeze all parameters except the low-rank residual network during training, that is, only keep the parameter update of the low-rank residual network; the training loss L task Including diffusion loss and editing positioning loss, the formula is defined as follows: L task =L diff +L loc ,(8) Formula (9) shows that L diff Diffusion loss, where represents the unified editing network optimized by the low-rank fine-tuning module; z~ε(x) represents the potential representation z obtained by transforming the input image x through the variational autoencoder ε; T desc represents the text description of the target image; ∈~N(0,1) represents the noise sampled from the standard normal distribution; t represents the time step in the diffusion, which controls the progress of the generation process; z t represents the potential representation after noise processing at time step t, which serves as the input of the generative network; z src represents the potential representation of the original image; I ref represents the reference image, which is used as a condition for the denoising process; P tar Represents the target semantic image; The edit localization loss L expressed by formula (10) loc ; Each layer of cross attention map in the single-step denoising process is defined as That is, the attention map corresponding to the kth word unit in the lth layer, the total number of words is K; A in the editing localization loss l is the average attention map of k words in the lth layer; since the sizes of attention maps in different layers are not all the same, according to experience, the average attention map A l The size is uniformly scaled to 16*16 resolution; the M in the edit localization loss represents the edit area mask, which has a size of 16*16 resolution; the mask data is an augmentation of the basic dataset and is obtained through the SAM segmentation model.

8. The instruction-driven personalized fashion image editing method according to claim 7, characterized in that In step (4), a joint training strategy is designed to perform end-to-end self-supervised training on the multi-task low-rank fine-tuning module; the training samples used in the joint training are randomly sampled from the "original image-reference image-target image-text editing instruction" four-dimensional data set; the diffusion loss is maintained on the basis of the classification loss, and the specific loss L is selected select The definition is as follows: Among them, L diff Similar to formula (9), but the merged low-rank fine-tuning module calculates the update as formula (2); τ θ represents the same text feature extraction module as the unified editing network, namely CLIP; λ * It is the unique hot encoding related to the editing task of the training sample, which ensures that the multi-task low-rank fine-tuning module activates the correct low-rank fine-tuning module while reducing the activation degree of the wrong editing ability as much as possible; SEL represents the low-rank fine-tuning module; L″ diff represents the diffusion loss, where represents the unified editing network optimized by the low-rank fine-tuning module; z~ε(x) represents the potential representation z obtained by transforming the input image x through the variational autoencoder ε; T desc represents the text description of the target image; ∈~N(0,1) represents the noise sampled from the standard normal distribution; t represents the time step in the diffusion, which controls the progress of the generation process; z t represents the potential representation after noise processing at time step t, which serves as the input of the generative network; z src represents the potential representation of the original image; I ref represents the reference image, which is used as a condition for the denoising process; P tar Represents the target semantic image.

Citation Information

Patent Citations

  • Image management method and device based on multi-task machine learning model

    CN111813532A

  • Posture and texture guided fashion costume design synthesis method

    CN113393550A