Method, device and equipment for realizing dressing of high-quality model and medium

By screening high-quality images in the clothing image library and extracting clothing feature vectors, combining deep learning and traditional algorithms, inputting them into the diffusion model for training, the problem of insufficient ability to capture and reconstruct clothing details in the existing technology is solved, and high-quality model dressing effect is achieved.

CN120070617APending Publication Date: 2025-05-30ZIXUN TECHNOLOGY (FUJIAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108213.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the field of model dressing, the existing deep learning networks have insufficient ability to capture and reconstruct clothing details, insufficient precision in clothing materials and lighting modeling, and lack of in-depth understanding of clothing structure and human body posture, resulting in the generated clothing effect not being natural enough, and the quality of the training data is uneven, which affects the generalization performance of the model.

Method used

By filtering high-quality images in the clothing image library, extracting clothing feature vectors and fusing deep learning features and features extracted by traditional algorithms, inputting them into the diffusion model for training, and generating high-quality model dressing pictures.

Benefits of technology

It realizes accurate capture and reconstruction of clothing details, improves the accuracy of clothing material and lighting modeling, makes the generated wear effect more natural, and improves the generalization performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070617A_ABST
    Figure CN120070617A_ABST
Patent Text Reader

Abstract

The invention provides a method, a device, equipment and a medium for realizing high-quality model dressing, and the method comprises the steps: screening images in a clothing picture library; removing images with set ambiguity; evaluating the uniformity of illumination, and deleting images which do not accord with a uniformity condition to obtain image data; finally, screening to obtain a set number of images as training data; obtaining a model display diagram corresponding to the training data as target data; carrying out feature vector extraction on images in the training data through a feature extractor and an existing algorithm, and carrying out fusion to obtain a clothes feature vector; inputting the clothes feature vectors and corresponding target data into a diffusion model, training the diffusion model, obtaining a trained clothes dressing model after the training is completed, inputting prompt words and required clothes pictures into the trained diffusion model, generating a model dressing picture, and displaying the model dressing picture. And the high-quality generative clothing dressing effect aiming at different models is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image generation, and in particular to a method, device, equipment and medium for implementing high-quality model dressing. Background Art

[0002] In the existing deep learning network technology solutions, due to the limitations of model architecture design and training methods, there are the following deficiencies:

[0003] First, the generative model is not able to capture and reconstruct clothing details, resulting in blurry or distorted details such as fabric texture and wrinkles;

[0004] Second, the modeling of clothing materials and lighting is not accurate enough, resulting in poor color reproduction and large color difference;

[0005] 3. Lack of in-depth understanding of clothing structure and human posture makes the generated wearing effect unnatural;

[0006] Fourth, the quality of training data is uneven and the quantity is limited, which affects the generalization performance of the model.

[0007] The above-mentioned problems seriously restrict the actual application effect of deep learning in the field of model dressing. Summary of the invention

[0008] The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for achieving high-quality model dressing, so as to achieve high-quality generative clothing dressing effects for different models.

[0009] In a first aspect, the present invention provides a method for achieving high-quality model dressing, comprising the following steps:

[0010] Step 1: In the clothing image library, filter images through the quality_filter function; remove images with set blur through blur_detection; evaluate the uniformity of lighting through lighting_analysis, delete images that do not meet the uniformity condition, and obtain image data; finally filter a set number of images as training data; and obtain the model display image corresponding to the training data as the target data;

[0011] Step 2: Extract feature vectors from images in the training data using a feature extractor based on the Swin-UNet architecture and existing algorithms, and fuse them to obtain clothing feature vectors;

[0012] Step 3: Input the clothing feature vector and the corresponding target data into the diffusion model for training. After the training is completed, a trained clothing dressing model is obtained. Input the prompt and the required clothing pictures into the trained diffusion model to generate pictures of a model wearing clothes.

[0013] In a second aspect, the present invention provides a device for achieving high-quality model dressing, including:

[0014] A training data acquisition module that, in a clothing picture library, filters images through the quality_filter function; eliminates images with a set blur degree through blur_detection; evaluates the uniformity of illumination through lighting_analysis, and deletes the images that do not meet the uniform conditions to obtain image data; finally, filters and obtains a set number of images as training data; and obtains the corresponding model display pictures of the training data as target data;

[0015] A feature vector extraction module that extracts and fuses the feature vectors of the images in the training data through a feature extractor based on the Swin-UNet architecture and an existing algorithm to obtain clothing feature vectors;

[0016] A picture generation module that inputs the clothing feature vector and the corresponding target data into the diffusion model for training. After the training is completed, a trained clothing dressing model is obtained. Input the prompt and the required clothing pictures into the trained diffusion model to generate pictures of a model wearing clothes.

[0017] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect is implemented.

[0018] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect is implemented.

[0019] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:

[0020] The present invention obtains clothing feature vectors in a specific manner. This accurate vector provides a good foundation for training the diffusion model, enabling the trained model to have strong capabilities in capturing and reconstructing clothing detail features, ensuring that the clothing does not have blurred or distorted details; being accurate enough in modeling clothing materials and illumination, ensuring color restoration and reducing color differences; making the generated wearing effects natural; and enabling the generated pictures to meet the needs of users, who can directly upload the pictures to their online stores for product display.

[0021] The above description is only an overview of the technical solution of the present invention. In order to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically given below. Brief Description of the Drawings

[0022] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.

[0023] Figure 1 It is a flowchart of the method in Embodiment 1 of the present invention;

[0024] Figure 2 It is a schematic structural diagram of the device in Embodiment 2 of the present invention. Detailed Embodiments

[0025] The technical solution in the embodiment of the present application has the following general idea:

[0026] First, based on the clothing pictures with self-owned copyright and the corresponding model display pictures, high-definition images with a resolution greater than 2048×2048 are selected through the quality_filter() function. Then, the blur_detection() is applied to detect the blur degree of the images, and unclear samples are removed; subsequently, the lighting_analysis() is used to evaluate the uniformity of the lighting to ensure the quality of the training data. Specifically:

[0027] 1. Internal process of the quality_filter() function

[0028] Read the image: Use the cv2.imread() function of OpenCV to read the image file.

[0029] Obtain the image size: Obtain the height and width of the image through the image.shape attribute.

[0030] Judge the resolution: Compare whether both the height and width of the image are greater than or equal to 2048. If so, keep the image; otherwise, remove it.

[0031] 2. Internal process of the blur_detection() function and the method for removing blur degree

[0032] Gray conversion: Convert the image to a grayscale image because the Laplacian operator is mainly used for grayscale images. Use the cv2.cvtColor() function of OpenCV to convert the image from the BGR color space to the grayscale color space.

[0033] Laplacian operator calculation: Apply the Laplacian operator to the grayscale image for convolution operation to calculate the Laplace transform result of the image. In OpenCV, the cv2.Laplacian() function can be used, setting an appropriate kernel size (such as 3) and border type (such as cv2.BORDER_DEFAULT).

[0034] Calculate the sharpness score: Calculate the variance of the Laplace transform result as the sharpness score of the image. The larger the variance, the clearer the image.

[0035] Set the threshold and remove blurred images: Set the threshold of the sharpness score to 15. If the sharpness score of the image is less than 15, the image is considered blurred and removed; otherwise, it is retained.

[0036] 3. Internal process of the lighting_analysis() function and method for evaluating uniformity

[0037] Divide the image into blocks: Divide the image into several equal-sized image blocks to calculate the brightness information of each image block separately. It can be divided according to a certain number of rows and columns. For example, divide the image into 4×4 image blocks.

[0038] Calculate the brightness of the image block: Calculate the average brightness value of each image block. First, convert the image block to a grayscale image (no conversion is required if the original image is already a grayscale image), and then calculate the average value of all pixel values in the image block, which is the brightness of the image block.

[0039] Calculate the brightness variance: Calculate the variance of the brightness values of all image blocks. The smaller the variance, the more uniform the illumination of the image.

[0040] Set the variance threshold: Set the threshold of the brightness variance to 0.15. If the brightness variance of the image is greater than 0.15, the image is considered to have uneven illumination and is removed; otherwise, it is retained.

[0041] Finally, 20,000 high-quality images are selected as training data. Next, start feature extraction on the images. The extracted features can be used for training the clothing wearing network.

[0042] To achieve more efficient and comprehensive clothing feature extraction, this solution needs to process two types of features: one is from the deep learning feature extractor based on the Swin-UNet architecture, and the other is from the feature vectors extracted by traditional algorithms. The following details the extraction, fusion of these two types of features and their application in the final network, including specific shapes and splicing details.

[0043] First, a feature extractor based on the Swin-UNet architecture was designed to obtain deep learning features in clothing images. The encoder part of this feature extractor adopted 4 layers of Swin Transformer Block, with each layer integrating the self-attention mechanism, which could effectively capture important features in the image. The size of the input image was 224×224. After being processed by the encoder, the size of the feature map gradually decreased. Assuming it was halved for each layer, the size of the feature map finally output by the encoder was 14×14. In the decoder part, deconvolution was used for upsampling to gradually restore the feature map to the original size. At the same time, high-resolution features from the encoder were fused through skip connections to ensure the retention of detailed information. At the bottleneck layer position, a channel attention module was introduced to further enhance the feature expression ability. The key parameters of the feature extractor were set as window size 7×7, number of heads 8, hidden layer dimension 96, and a drop path rate of 0.1 was adopted to prevent overfitting, with the activation function using GELU. Through the above design, the Swin-UNet feature extractor could output a deep learning feature vector with a shape of (256,224,224), where 256 was the number of channels and 224×224 was the spatial size of the feature map.

[0044] Secondly, to improve the accuracy and diversity of feature extraction, a variety of traditional image processing algorithms were combined to extract multiple feature vectors. First, the Sobel operator was used to extract the horizontal and vertical gradients of the image, generating a gradient feature map with a shape of (2,224,224). Subsequently, the Canny operator was applied for edge detection, generating an edge feature map with a shape of (1,224,224). Then, the HOG (Histogram of Oriented Gradients) descriptor was used to extract the statistical features of the local gradient direction, generating an HOG feature map with a shape of (3,224,224). In addition, SIFT (Scale-Invariant Feature Transform) was used to extract key points and their descriptors in the image. Since the number of key points was not fixed, the SIFT descriptors were usually aggregated into a feature vector with a fixed dimension through the Bag of Visual Words (BoVW) method, with a shape of (256). These preprocessing steps respectively generated different types of feature vectors, including gradient vectors, edge vectors, HOG feature vectors, and SIFT key point descriptor vectors.

[0045] Next, the above two types of feature vectors are effectively fused to construct a comprehensive clothing feature representation. First, the deep learning feature vector (with a shape of (256, 224, 224)) output by the Swin-UNet feature extractor is concatenated with various feature vectors extracted by traditional algorithms (2 channels of gradient features, 1 channel of edge features, and 3 channels of HOG features) in the channel dimension to form a comprehensive traditional feature vector with a shape of (6, 224, 224). Since the SIFT feature has a shape of (256), it needs to be mapped and extended to (256, 224, 224) through a fully connected layer to match the spatial dimensions of other features. Then, all feature vectors are concatenated in the channel dimension to obtain a comprehensive feature vector with a shape of (262, 224, 224), where 262 = 256 (deep learning features) + 6 (traditional algorithm features). Subsequently, LayerNorm is applied to normalize the concatenated comprehensive feature vector to ensure that the feature values are within a unified scale range, improving the stability and convergence speed of model training. Finally, a 1×1 convolution is used to compress the number of channels from 262 to 256, generating the final clothing feature vector with a shape of (256, 224, 224) as the data input for subsequent generation model training.

[0046] The fused clothing feature vector will be used as conditional information and input into the generation model to guide the generation process to achieve a high-quality clothing dressing effect. The specific process is as follows. First, the Euler sampler is used to perform the generation process in the Stable Diffusion model. A cross-attention layer is added to the U-Net architecture to use the clothing feature vector as conditional information to guide the generation process. The cross-attention mechanism utilizes these conditional features to guide the generation model to ensure that the generated dressing effect conforms to the input clothing features. The relevant training parameter settings include adopting a cosine strategy for the β schedule and setting the number of cross-attention heads to 8, which is consistent with the feature extractor. Combining the extraction results of the pre-trained Swin-UNet, 10 rounds of fine-tuning training are carried out to enable the generation model to fully understand and utilize the fused diverse feature information, generating a natural, high-definition, and high-quality dressing effect to obtain the final clothing dressing model.

[0047] To ensure that the generation model can achieve the expected high-definition and high-quality dressing effects, this solution designs a comprehensive evaluation mechanism that combines automated metrics with human scoring to systematically evaluate the quality of the generated images. First, the quality_assessment() function is used to automatically evaluate the dressing effects of the generated clothing, mainly calculating two key metrics: PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index). PSNR is used to measure the similarity between the generated image and the real image, with a threshold set at 30 dB to ensure that the generated image is highly consistent with the real image in terms of details and clarity. SSIM is used to evaluate the similarity between the generated image and the real image in terms of brightness, contrast, and structure, with a threshold set at 0.85 to ensure that the generated image has good visual consistency while maintaining structural information. Through the calculation of these two metrics, the quality_assessment() function can quickly screen out the generated images that meet the quality standards, providing a high-quality sample basis for subsequent human scoring.

[0048] In addition to the automated evaluation, this solution also introduces a human scoring mechanism to further ensure the visual quality and naturalness of the generated images. The specific steps include using the user_feedback() function to invite users to rate the dressing effects of the generated clothing on a scale of 1-5, where 1 indicates poor quality and 5 indicates extremely high quality, covering subjective evaluations of aspects such as image details, color restoration, and naturalness of clothing structure. A total of 10,000 pieces of human scoring data for the generated images are collected. By constructing a large-scale quality evaluation dataset, the diversity and representativeness of the scoring are ensured, avoiding the influence of single-user bias on the evaluation results. Statistical analysis is performed on the collected human scoring data to calculate the average score for each image, and the automated metrics (PSNR and SSIM) are combined with the human scoring to form a comprehensive quality evaluation system, ensuring the comprehensiveness and accuracy of the evaluation results.

[0049] Through the above-mentioned automated evaluation and human scoring, this solution comprehensively evaluates the quality of the generation model and optimizes it accordingly. The PSNR threshold is set at 30 dB, the SSIM threshold is set at 0.85, and the human scoring threshold is set at 4 points (i.e., most users give high evaluations). The proportion of generated images that meet all threshold conditions is statistically calculated, with the goal of achieving a yield rate of 70%. This means that at least 70% of all generated images can meet or exceed the set quality standards. For the generated images that fail to pass the quality assessment, analyze the existing problems, such as blurred details, color distortion, or unnatural structure, and adjust the model parameters and training strategies accordingly. Through iterative optimization, gradually improve the performance of the model in various quality metrics to ensure that the final model can stably generate high-definition and high-quality dressing effects.

[0050] Example 1

[0051] like Figure 1 As shown, this embodiment provides a method for achieving high-quality model dressing, comprising the following steps:

[0052] Step 1: In the clothing image library, filter images through the quality_filter function; remove images with set blur through blur_detection; evaluate the uniformity of lighting through lighting_analysis, delete images that do not meet the uniformity condition, and obtain image data; finally filter a set number of images as training data; and obtain the model display image corresponding to the training data as the target data;

[0053] Step 2: Extract feature vectors from images in the training data using a feature extractor based on the Swin-UNet architecture and existing algorithms, and fuse them to obtain clothing feature vectors;

[0054] Step 3: Input the clothing feature vector and the corresponding target data into the diffusion model to train the diffusion model. After the training is completed, a trained clothing dressing model is obtained. The prompt words and the required clothing pictures are input into the trained diffusion model to generate model dressing pictures.

[0055] In this embodiment, preferably, the step 1 is specifically as follows: in the clothing image library, images with a resolution greater than 2048×2048 are filtered out by the quality_filter function; high-definition images with set blur are eliminated by blur_detection; the uniformity of illumination is evaluated by lighting_analysis, and high-definition images that do not meet the uniformity condition are deleted to obtain training image data; and a model display image corresponding to the training data is obtained as the target data;

[0056] The quality_filter function filters specifically as follows:

[0057] Read image: Use OpenCV's cv2.imread function to read clothing images;

[0058] Get the image size: get the height and width of the image through the image.shape attribute;

[0059] Determine the resolution: compare whether the height and width of the image are both greater than or equal to 2048. If so, keep the image, otherwise discard it;

[0060] The blur_detection function removes the high-definition image with set blurriness as follows:

[0061] Gray-scale conversion: Convert each high-definition image into a grayscale image. Use the cv2.cvtColor function in OpenCV to convert each high-definition image from the BGR color space to the grayscale color space;

[0062] Laplacian operator calculation: Apply the Laplacian operator to the grayscale image for convolution operation to calculate the Laplace transform result of the image; In OpenCV, use the cv2.Laplacian function to set the kernel size and boundary type;

[0063] Calculate the sharpness score: Calculate the variance of the Laplace transform result as the sharpness score of the image; The larger the variance, the clearer the image;

[0064] Set the threshold and eliminate blurred images: Set the sharpness score threshold to 15. If the sharpness score of the image is less than 15, the high-definition image is blurred and will be eliminated; otherwise, it will be retained;

[0065] The specific method of evaluating the uniformity of illumination through lighting_analysis and deleting the high-definition images that do not meet the uniform conditions is as follows:

[0066] Divide the image into blocks: Divide each high-definition image into a set number of equal-sized image blocks;

[0067] Calculate the brightness of the image block: Calculate the brightness value of each image block. First, convert the image block into a grayscale image, and then calculate the average value of all pixel values in the image block, which is the brightness value of the image block;

[0068] Calculate the brightness variance: Calculate the variance of all image block brightness values. The smaller the variance, the more uniform the illumination of the image;

[0069] Set the variance threshold: Set the brightness variance threshold to 0.15. If the brightness variance of the image is greater than 0.15, the illumination of the high-definition image is not uniform and will be eliminated; otherwise, it will be retained;

[0070] Finally, 20,000 images are selected as training data.

[0071] In this embodiment, preferably, step 2 is specifically: Extract the feature vectors of the images in the training data through a feature extractor based on the Swin-UNet architecture and an existing algorithm, and perform fusion to obtain the clothing feature vectors;

[0072] First, a feature extractor based on the Swin-UNet architecture is designed to obtain deep learning features in clothing images. The feature extractor includes an encoder and a decoder. The encoder uses 4 layers of Swin Transformer Blocks, and each layer of Swin Transformer Block integrates the self-attention mechanism. The size of the input image is set to 224×224. After being processed by the encoder, the size of the image gradually decreases to obtain the encoded feature map. The decoder uses transposed convolution for upsampling, gradually restoring the encoded feature map to the original size, and at the same time fusing high-resolution features from the encoder through skip connections. A channel attention module is introduced in the bottleneck layer to enhance the feature expression ability.

[0073] The parameters of the feature extractor are set as the window size of 7×7, the number of heads of 8, and the hidden layer dimension of 96, and the droppath rate is set to 0.1. The activation function uses GELU.

[0074] Through the above design, the feature extractor outputs a deep learning feature vector with a shape of (256,224,224), where 256 is the number of channels and 224×224 is the spatial size of the feature map.

[0075] The existing algorithms are as follows:

[0076] Use the Sobel operator to extract the horizontal and vertical gradients of the images in the training data, generating a gradient feature map with a shape of (2,224,224).

[0077] Use the Canny operator for edge detection, generating an edge feature map with a shape of (1,224,224).

[0078] Use the HOG descriptor to extract the statistical features of the local gradient directions, generating a HOG feature map with a shape of (3,224,224).

[0079] Use SIFT to extract the key points and their descriptors of the images in the training data, and aggregate the SIFT descriptors into a feature vector with a fixed dimension through the Bag of Visual Words method to obtain a SIFT feature map with a shape of 256.

[0080] The existing algorithms obtain gradient vectors, edge vectors, HOG feature vectors, and SIFT key point descriptor vectors.

[0081] Concatenate the deep learning feature vector output by the Swin-UNet feature extractor with various feature vectors extracted by the existing algorithms in the channel dimension to form a comprehensive traditional feature vector with a shape of (6,224,224).

[0082] The SIFT feature map has a shape of 256 and needs to be mapped and extended to (256, 224, 224) through a fully connected layer. Then, all feature vectors are concatenated in the channel dimension to obtain a comprehensive feature vector with a shape of (262, 224, 224). Subsequently, LayerNorm is applied to normalize the concatenated comprehensive feature vector to ensure that the eigenvalues are within the set scale range. Finally, a 1×1 convolution is used to compress the number of channels from 262 to 256 to generate the final clothing feature vector with a shape of (256, 224, 224).

[0083] In this embodiment, preferably, step 3 is specifically as follows: input the clothing feature vector and the corresponding target data into the diffusion model for training the diffusion model. The diffusion model is the Stable Diffusion model, and the Euler sampler is used for sampling in the Stable Diffusion model. A cross-attention layer is added to the U-Net architecture to use the clothing feature vector as conditional information to guide the generation process of the Stable Diffusion model. The Stable Diffusion model adopts a cross-attention mechanism, and the parameter settings in the Stable Diffusion model include: the β schedule adopts a cosine strategy, and the number of cross-attention heads is set to 8. Combining the extraction results of the pre-trained Swin-UNet, 10 rounds of fine-tuning training are performed to obtain the final clothing dressing model.

[0084] After the training is completed, the trained clothing dressing model is obtained. The prompt word and the required clothing picture are input into the trained diffusion model to generate a picture of a model wearing clothes.

[0085] Based on the same inventive concept, the present application also provides an apparatus corresponding to the method in Embodiment 1, as detailed in Embodiment 2.

[0086] Embodiment 2

[0087] As Figure 2 shown, in this embodiment, an apparatus for realizing high-quality model clothing dressing is provided, including:

[0088] A training data acquisition module, in the clothing picture library, screens images through the quality_filter function; eliminates images with a set blur degree through blur_detection; evaluates the uniformity of illumination through lighting_analysis, and deletes the images that do not meet the uniform conditions to obtain image data; finally, a set number of images are screened out as training data; and the corresponding model display pictures of the training data are obtained as target data.

[0089] The feature vector extraction module extracts feature vectors from images in the training data through a feature extractor based on the Swin-UNet architecture and existing algorithms, and fuses them to obtain clothing feature vectors;

[0090] Generate a picture module, input the clothing feature vector and the corresponding target data into the diffusion model, train the diffusion model, and after the training is completed, obtain the trained clothing dressing model, input the prompt words and the required clothing pictures into the trained diffusion model to generate a model dressing picture.

[0091] In this embodiment, preferably, the module for acquiring training data is specifically as follows: in the clothing image library, images with a resolution greater than 2048×2048 are filtered out by the quality_filter function; high-definition images with set blur are eliminated by blur_detection; uniformity of illumination is evaluated by lighting_analysis, and high-definition images that do not meet the uniformity condition are deleted to obtain training image data; and a model display image corresponding to the training data is obtained as target data;

[0092] The quality_filter function filters specifically as follows:

[0093] Read image: Use OpenCV's cv2.imread function to read clothing images;

[0094] Get the image size: get the height and width of the image through the image.shape attribute;

[0095] Determine the resolution: compare whether the height and width of the image are both greater than or equal to 2048. If so, keep the image, otherwise discard it;

[0096] The blur_detection function removes the high-definition image with set blurriness as follows:

[0097] Grayscale conversion: Convert each HD image to a grayscale image. Use OpenCV's cv2.cvtColor function to convert each HD image from the BGR color space to the grayscale color space.

[0098] Laplace operator calculation: Apply the Laplace operator to the grayscale image for convolution operation and calculate the Laplace transform result of the image; in OpenCV, use the cv2.Laplacian function and set the kernel size and boundary type;

[0099] Calculate the clarity score: Calculate the variance of the Laplace transform result as the image clarity score; the larger the variance, the clearer the image;

[0100] Set a threshold and eliminate blurred images: Set the clarity score threshold to 15. If the clarity score of an image is less than 15, the high-definition image is blurred and will be eliminated; otherwise, it will be retained;

[0101] Evaluating the uniformity of illumination through lighting_analysis and deleting the high-definition images that do not meet the uniformity conditions specifically:

[0102] Divide the image into blocks: Divide each high-definition image into a set number of equal-sized image blocks;

[0103] Calculate the brightness of the image block: Calculate the brightness value of each image block. First, convert the image block into a grayscale image, and then calculate the average value of all pixel values in the image block, which is the brightness value of the image block;

[0104] Calculate the brightness variance: Calculate the variance of all image block brightness values. The smaller the variance, the more uniform the illumination of the image;

[0105] Set the variance threshold: Set the brightness variance threshold to 0.15. If the brightness variance of an image is greater than 0.15, the illumination of the high-definition image is not uniform and will be eliminated; otherwise, it will be retained;

[0106] Finally, 20,000 images are selected as training data.

[0107] In this embodiment, preferably, the feature vector extraction module is specifically: Extract and fuse the feature vectors of the images in the training data through a feature extractor based on the Swin-UNet architecture and existing algorithms to obtain clothing feature vectors;

[0108] First, design a feature extractor based on the Swin-UNet architecture to obtain deep learning features in clothing images; the feature extractor includes an encoder and a decoder. The encoder uses 4 layers of Swin TransformerBlock, and each layer of Swin TransformerBlock integrates the self-attention mechanism. Set the size of the input image to 224×224. After being processed by the encoder, the size of the image gradually decreases to obtain an encoded feature map; the decoder uses transposed convolution for upsampling to gradually restore the encoded feature map to the original size, and at the same time, fuse the high-resolution features from the encoder through skip connections; introduce a channel attention module in the bottleneck layer to enhance the feature expression ability;

[0109] The parameters of the feature extractor are set as the window size is 7×7, the number of heads is 8, and the hidden layer dimension is 96, and the droppath rate is set to 0.1, and the activation function uses GELU;

[0110] Through the above design, the feature extractor outputs a deep learning feature vector with a shape of (256, 224, 224), where 256 is the number of channels and 224×224 is the spatial size of the feature map;

[0111] The existing algorithms are as follows:

[0112] Use the Sobel operator to extract the horizontal and vertical gradients of the images in the training data, generating a gradient feature map with a shape of (2, 224, 224);

[0113] Use the Canny operator for edge detection, generating an edge feature map with a shape of (1, 224, 224);

[0114] Use the HOG descriptor to extract the statistical features of the local gradient directions, generating a HOG feature map with a shape of (3, 224, 224);

[0115] Use SIFT to extract the key points and their descriptors of the images in the training data, and aggregate the SIFT descriptors into a feature vector with a fixed dimension through the Bag of Visual Words method to obtain a SIFT feature map with a shape of 256;

[0116] The existing algorithms obtain gradient vectors, edge vectors, HOG feature vectors, and SIFT key point descriptor vectors;

[0117] Concatenate the deep learning feature vector output by the Swin-UNet feature extractor and various feature vectors extracted by the existing algorithms in the channel dimension to form a comprehensive traditional feature vector with a shape of (6, 224, 224);

[0118] The SIFT feature map has a shape of 256 and needs to be mapped through a fully connected layer and extended to (256, 224, 224); then, all feature vectors are concatenated in the channel dimension to obtain a comprehensive feature vector with a shape of (262, 224, 224); subsequently, apply LayerNorm to normalize the concatenated comprehensive feature vector to ensure that the feature values are within the set scale range; finally, use a 1×1 convolution to compress the number of channels from 262 to 256 to generate the final clothing feature vector with a shape of (256, 224, 224).

[0119] In this embodiment, preferably, the module for generating pictures specifically is: input the clothing feature vector and the corresponding target data into a diffusion model for training the diffusion model. The diffusion model is the Stable Diffusion model, and the Euler sampler is used for sampling in the Stable Diffusion model; a cross-attention layer is added to the U-Net architecture to use the clothing feature vector as conditional information to guide the generation process of the Stable Diffusion model; the Stable Diffusion model adopts a cross-attention mechanism, and the parameter settings in the Stable Diffusion model include: the β schedule adopts a cosine strategy, and the number of cross-attention heads is set to 8; combined with the extraction results of the pre-trained Swin-UNet, fine-tuning training is performed for 10 rounds to obtain the final clothing dressing model;

[0120] After the training is completed, the trained clothing dressing model is obtained. Input the prompt and the required clothing pictures into the trained diffusion model to generate pictures of a model wearing clothes.

[0121] Since the device introduced in the second embodiment of the present invention is the device adopted for implementing the method in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated here. Any device adopted for the method in the first embodiment of the present invention belongs to the scope protected by the present invention.

[0122] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to the first embodiment. For details, see the third embodiment.

[0123] Embodiment Three

[0124] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, any implementation manner in the first embodiment can be realized.

[0125] Since the electronic device introduced in this embodiment is the device adopted for implementing the method in the first embodiment of this application, based on the method introduced in the first embodiment of this application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment, so the implementation of how this electronic device realizes the method in the embodiments of this application will not be introduced in detail here. Any device adopted by those skilled in the art for implementing the method in the embodiments of this application belongs to the scope protected by this application.

[0126] Based on the same inventive concept, this application provides a storage medium corresponding to the first embodiment. For details, see the fourth embodiment.

[0127] Example 4

[0128] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any implementation manner in Embodiment 1 can be implemented.

[0129] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0130] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0131] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0132] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0133] Although the specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should all be covered by the scope protected by the claims of the present invention.

Claims

1. A method for achieving high-quality model dressing, characterized in that: The steps include: Step 1: In the clothing image library, filter images through the quality_filter function; remove images with set blur through blur_detection; evaluate the uniformity of lighting through lighting_analysis, delete images that do not meet the uniformity condition, and obtain image data; finally filter a set number of images as training data; and obtain the model display image corresponding to the training data as the target data; Step 2: Extract feature vectors from images in the training data using a feature extractor based on the Swin-UNet architecture and existing algorithms, and fuse them to obtain clothing feature vectors; Step 3: Input the clothing feature vector and the corresponding target data into the diffusion model to train the diffusion model. After the training is completed, a trained clothing dressing model is obtained. The prompt words and the required clothing pictures are input into the trained diffusion model to generate model dressing pictures.

2. A method for achieving high-quality model dressing according to claim 1, characterized in that: The step 1 is specifically as follows: in the clothing image library, images with a resolution greater than 2048×2048 are filtered out by the quality_filter function; high-definition images with set blur are eliminated by blur_detection; the uniformity of illumination is evaluated by lighting_analysis, and high-definition images that do not meet the uniformity condition are deleted to obtain training image data; and the model display image corresponding to the training data is obtained as the target data; The quality_filter function filters specifically as follows: Read image: Use OpenCV's cv2.imread function to read clothing images; Get the image size: get the height and width of the image through the image.shape attribute; Determine the resolution: compare whether the height and width of the image are both greater than or equal to 2048. If so, keep the image, otherwise discard it; The blur_detection function removes the high-definition image with set blurriness as follows: Grayscale conversion: Convert each HD image to a grayscale image. Use OpenCV's cv2.cvtColor function to convert each HD image from the BGR color space to the grayscale color space. Laplace operator calculation: Apply the Laplace operator to the grayscale image for convolution operation and calculate the Laplace transform result of the image; in OpenCV, use the cv2.Laplacian function and set the kernel size and boundary type; Calculate the clarity score: Calculate the variance of the Laplace transform result as the image clarity score; The larger the variance, the clearer the image; Set the threshold and remove blurry images: Set the clarity score threshold to 15. If the clarity score of an image is less than 15, the high-definition image is blurry and will be removed. Otherwise keep it; The method of evaluating the uniformity of illumination by lighting_analysis and deleting high-definition images that do not meet the uniformity condition is as follows: Divide image blocks: Divide each high-definition image into a set number of image blocks of equal size; Calculate the brightness of the image block: Calculate the brightness value of each image block. First, convert the image block into a grayscale image, and then calculate the average value of all pixel values ​​in the image block, which is the brightness value of the image block. Calculate brightness variance: Calculate the variance of the brightness values ​​of all image blocks. The smaller the variance, the more uniform the illumination of the image. Set variance threshold: Set the brightness variance threshold to 0.

15. If the brightness variance of an image is greater than 0.15, the illumination of the high-definition image is uneven and it will be removed. Otherwise keep it; Finally, 20,000 images were screened and used as training data.

3. A method for achieving high-quality model dressing according to claim 1, characterized in that: The step 2 specifically comprises: extracting feature vectors from images in the training data through a feature extractor based on the Swin-UNet architecture and an existing algorithm, and fusing the extracted feature vectors to obtain clothing feature vectors; Firstly, a feature extractor based on the Swin-UNet architecture is designed to obtain deep learning features in clothing images. The feature extractor includes an encoder and a decoder. The encoder adopts a 4-layer Swin TransformerBlock, and each layer of Swin TransformerBlock integrates a self-attention mechanism. The size of the input image is set to 224×224. After being processed by the encoder, the size of the image is gradually reduced to obtain an encoded feature map. The decoder uses deconvolution for upsampling to gradually restore the encoded feature map to its original size, and at the same time fuses the high-resolution features from the encoder through jump connections. Introduce the channel attention module in the bottleneck layer to enhance the feature expression capability; The parameters of the feature extractor are set to a window size of 7×7, a number of heads of 8, and a hidden layer dimension of 96, and the droppath rate is set to 0.1, and the activation function uses GELU; Through the above design, the feature extractor outputs a deep learning feature vector with a shape of (256, 224, 224), where 256 is the number of channels and 224×224 is the spatial size of the feature map; The existing algorithm is: Use the Sobel operator to extract the horizontal and vertical gradients of the image in the training data and generate a gradient feature map with a shape of (2,224,224); Use the Canny operator for edge detection to generate an edge feature map with a shape of (1,224,224); Use HOG descriptor to extract statistical features of local gradient direction and generate a HOG feature map with a shape of (3,224,224); Use SIFT to extract the key points and descriptors of the images in the training data, and aggregate the SIFT descriptors into feature vectors of fixed dimensions through the Bag of Visual Words method to obtain a SIFT feature map with a shape of 256; The existing algorithm obtains a gradient vector, an edge vector, a HOG feature vector and a SIFT key point descriptor vector; The deep learning feature vector output by the Swin-UNet feature extractor is concatenated with various feature vectors extracted by the existing algorithm in the channel dimension to form a comprehensive traditional feature vector with a shape of (6,224,224); The SIFT feature map has a shape of 256 and needs to be mapped and expanded to (256, 224, 224) through a fully connected layer. Then, all feature vectors are concatenated in the channel dimension to obtain a comprehensive feature vector with a shape of (262, 224, 224). Subsequently, LayerNorm is applied to normalize the concatenated comprehensive feature vector to ensure that the eigenvalues ​​are within the set scale range. Finally, a 1×1 convolution is used to compress the number of channels from 262 to 256 to generate the final clothing feature vector with a shape of (256, 224, 224).

4. A method for achieving high-quality model dressing according to claim 1, characterized in that: The step 3 is specifically as follows: inputting the clothing feature vector and the corresponding target data into a diffusion model to train the diffusion model, wherein the diffusion model is a Stable Diffusion model, and the Euler sampler is used for sampling in the Stable Diffusion model; adding a cross-attention layer to the U-Net architecture to guide the Stable Diffusion model generation process using the clothing feature vector as conditional information; the Stable Diffusion model adopts a cross-attention mechanism, and the parameter settings in the Stable Diffusion model include: the β scheduling adopts a cosine strategy, and the number of cross-attention heads is set to 8; combining the extraction results of the pre-trained Swin-UNet, 10 rounds of fine-tuning training are performed to obtain the final clothing dressing model; When the training is completed, a trained clothing dressing model is obtained, and the prompt words and the required clothing pictures are input into the trained diffusion model to generate a model dressing picture.

5. A device for achieving high-quality model dressing, characterized in that: include: Get the training data module and filter the images in the clothing image library using the quality_filter function; Use blur_detection to remove images with set blur; use lighting_analysis to evaluate the uniformity of lighting, delete images that do not meet the uniformity condition, and obtain image data; finally, filter out a set number of images as training data; and obtain the model display image corresponding to the training data as the target data; The feature vector extraction module extracts feature vectors from images in the training data through a feature extractor based on the Swin-UNet architecture and existing algorithms, and fuses them to obtain clothing feature vectors; Generate a picture module, input the clothing feature vector and the corresponding target data into the diffusion model, train the diffusion model, and after the training is completed, obtain the trained clothing dressing model, input the prompt words and the required clothing pictures into the trained diffusion model to generate a model dressing picture.

6. The device for achieving high-quality model dressing according to claim 5, characterized in that: The module for acquiring training data specifically comprises: in the clothing image library, images with a resolution greater than 2048×2048 are filtered out by the quality_filter function; high-definition images with set blur are eliminated by blur_detection; the uniformity of illumination is evaluated by lighting_analysis, and high-definition images that do not meet the uniformity condition are deleted to obtain training image data; and a model display image corresponding to the training data is obtained as target data; The quality_filter function filters specifically as follows: Read image: Use OpenCV's cv2.imread function to read clothing images; Get the image size: get the height and width of the image through the image.shape attribute; Determine the resolution: compare whether the height and width of the image are both greater than or equal to 2048. If so, keep the image, otherwise discard it; The blur_detection function removes the high-definition image with set blurriness as follows: Grayscale conversion: Convert each HD image to a grayscale image. Use OpenCV's cv2.cvtColor function to convert each HD image from the BGR color space to the grayscale color space. Laplace operator calculation: Apply the Laplace operator to the grayscale image for convolution operation and calculate the Laplace transform result of the image; in OpenCV, use the cv2.Laplacian function and set the kernel size and boundary type; Calculate the clarity score: Calculate the variance of the Laplace transform result as the image clarity score; The larger the variance, the clearer the image; Set the threshold and remove blurry images: Set the clarity score threshold to 15. If the clarity score of an image is less than 15, the high-definition image is blurry and will be removed. Otherwise keep it; The method of evaluating the uniformity of illumination by lighting_analysis and deleting high-definition images that do not meet the uniformity condition is as follows: Divide image blocks: Divide each high-definition image into a set number of image blocks of equal size; Calculate the brightness of the image block: Calculate the brightness value of each image block. First, convert the image block into a grayscale image, and then calculate the average value of all pixel values ​​in the image block, which is the brightness value of the image block. Calculate brightness variance: Calculate the variance of the brightness values ​​of all image blocks. The smaller the variance, the more uniform the illumination of the image. Set variance threshold: Set the brightness variance threshold to 0.

15. If the brightness variance of an image is greater than 0.15, the illumination of the high-definition image is uneven and it will be removed. Otherwise keep it; Finally, 20,000 images were screened and used as training data.

7. The device for achieving high-quality model dressing according to claim 5, characterized in that: The feature vector extraction module specifically comprises: extracting feature vectors from images in training data through a feature extractor based on the Swin-UNet architecture and an existing algorithm, and fusing them to obtain clothing feature vectors; Firstly, a feature extractor based on the Swin-UNet architecture is designed to obtain deep learning features in clothing images. The feature extractor includes an encoder and a decoder. The encoder adopts a 4-layer Swin TransformerBlock, and each layer of Swin TransformerBlock integrates a self-attention mechanism. The size of the input image is set to 224×224. After being processed by the encoder, the size of the image is gradually reduced to obtain an encoded feature map. The decoder uses deconvolution for upsampling to gradually restore the encoded feature map to its original size, and at the same time fuses the high-resolution features from the encoder through jump connections. Introduce the channel attention module in the bottleneck layer to enhance the feature expression capability; The parameters of the feature extractor are set to a window size of 7×7, a number of heads of 8, and a hidden layer dimension of 96, and the droppath rate is set to 0.1, and the activation function uses GELU; Through the above design, the feature extractor outputs a deep learning feature vector with a shape of (256, 224, 224), where 256 is the number of channels and 224×224 is the spatial size of the feature map; The existing algorithm is: Use the Sobel operator to extract the horizontal and vertical gradients of the image in the training data and generate a gradient feature map with a shape of (2,224,224); Use the Canny operator for edge detection to generate an edge feature map with a shape of (1,224,224); Use HOG descriptor to extract statistical features of local gradient direction and generate a HOG feature map with a shape of (3,224,224); Use SIFT to extract the key points and descriptors of the images in the training data, and aggregate the SIFT descriptors into feature vectors of fixed dimensions through the Bag of Visual Words method to obtain a SIFT feature map with a shape of 256; The existing algorithm obtains a gradient vector, an edge vector, a HOG feature vector and a SIFT key point descriptor vector; The deep learning feature vector output by the Swin-UNet feature extractor is concatenated with various feature vectors extracted by the existing algorithm in the channel dimension to form a comprehensive traditional feature vector with a shape of (6,224,224); The SIFT feature map has a shape of 256 and needs to be mapped and expanded to (256, 224, 224) through a fully connected layer. Then, all feature vectors are concatenated in the channel dimension to obtain a comprehensive feature vector with a shape of (262, 224, 224). Subsequently, LayerNorm is applied to normalize the concatenated comprehensive feature vector to ensure that the eigenvalues ​​are within the set scale range. Finally, a 1×1 convolution is used to compress the number of channels from 262 to 256 to generate the final clothing feature vector with a shape of (256, 224, 224).

8. The device for achieving high-quality model dressing according to claim 5, characterized in that: The image generation module is specifically as follows: inputting the clothing feature vector and the corresponding target data into a diffusion model to train the diffusion model, wherein the diffusion model is a Stable Diffusion model, and the Euler sampler is used for sampling in the Stable Diffusion model; adding a cross-attention layer to the U-Net architecture to guide the Stable Diffusion model generation process using the clothing feature vector as conditional information; the Stable Diffusion model adopts a cross-attention mechanism, and the parameter settings in the Stable Diffusion model include: using a cosine strategy for β scheduling, and setting the number of cross-attention heads to 8; combining the extraction results of the pre-trained Swin-UNet, performing 10 rounds of fine-tuning training to obtain the final clothing dressing model; When the training is completed, a trained clothing dressing model is obtained, and the prompt words and the required clothing pictures are input into the trained diffusion model to generate a model dressing picture.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.