A method and system for generating a garment mannequin image

By extracting features from clothing product images and utilizing generative adversarial networks and 3D human pose models, optimal model images are generated, solving the problems of low efficiency and large differences in clothing model image generation, and achieving efficient and accurate clothing model image generation.

CN120976404BActive Publication Date: 2026-05-08GUANGZHOU TAIDONG TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU TAIDONG TECH CO LTD
Filing Date
2025-06-17
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency in generating clothing model images and significant differences between the generated images and the original clothing.

Method used

By acquiring images of clothing products, extracting features such as color, material, size, and style, and using generative adversarial networks and 3D human pose models, combined with weighted fusion and comprehensive similarity calculation, the optimal model image is generated.

Benefits of technology

It improves the efficiency of generating clothing model images, reduces the difference between clothing models and original clothing, and makes the features of the generated clothing models closer to the features of clothing in the clothing product images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976404B_ABST
    Figure CN120976404B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image processing, and relates to a kind of clothing model image generation method and system, the method includes: obtaining clothing commodity graph, and different types of features of corresponding clothing are extracted according to clothing commodity graph;Obtain fusion feature;Generate candidate model set, and then obtain candidate model;The body shape parameters of each candidate model are adjusted using 3D human posture model, and the posture parameters of each candidate model are adjusted;The similarity of multiple types of features between each candidate model and clothing commodity graph is calculated respectively, the similarity of each type feature is weighted and summed to obtain the comprehensive similarity corresponding to each candidate model;And the model with the comprehensive similarity greater than similarity threshold is regarded as optimal model.The method of the present application can improve the generation efficiency of clothing model image and reduce the difference between the clothing of clothing model image and the clothing on original clothing commodity graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology. More specifically, this invention relates to a method and system for generating images of clothing models. Background Technology

[0002] Apparel product images are images specifically used to showcase and sell clothing products. These images highlight the key elements of the clothing product. Through high-quality images, customers can clearly see the style, color, material, and details of the clothing, thus making a purchasing decision.

[0003] When merchants list clothing products on e-commerce platforms such as Taobao and Pinduoduo, they usually need to upload product images. The traditional method for creating clothing product images is to have real models try on the clothes and take photos, but this method is relatively costly.

[0004] Alternatively, model templates from a clothing model image library can be combined with image synthesis technology to obtain clothing model images. This method requires first acquiring an image of a person in clothing; then, retrieving a model template from the model image library that has a similar pose to the person in clothing; next, determining the deformation feature parameters of the original clothing image in the person's clothing image relative to the body image, transforming the original clothing image according to the deformation feature parameters to obtain a transformed clothing image; finally, compositing the transformed clothing image onto the body image of the model template to obtain the model in clothing image. Because this method requires capturing an image of the person in clothing before compositing the clothing model image, the efficiency of clothing model image generation is low. Furthermore, this method requires transforming the original clothing image using deformation feature parameters, altering the original clothing's shape and size, resulting in significant differences between the clothing in the generated clothing model image and the original clothing. Summary of the Invention

[0005] To address the technical problems of low efficiency in generating clothing model images and significant differences between the clothing in the generated clothing model images and the clothing in the original clothing product images in existing technologies, this invention provides solutions in the following aspects.

[0006] In a first aspect, the present invention provides a method for generating images of clothing models, comprising: acquiring a clothing product image, and extracting different types of features corresponding to the clothing based on the clothing product image, wherein the extracted features include: color features, material features, size features, and style features;

[0007] The extracted features of different types are weighted and fused to obtain fused features;

[0008] The initial model set and fused features are input into the generative adversarial network model to generate a candidate model set. Then, the loss function value corresponding to each candidate model is calculated, and the candidate models whose loss function value is less than the loss threshold are selected as candidate models.

[0009] The body shape parameters of each candidate model are adjusted using a 3D human posture model to match the looseness of the clothing corresponding to the clothing product image, and the posture parameters of each candidate model are adjusted to simulate the dynamic effect of the clothing.

[0010] The similarity of each candidate model to the clothing product image is calculated for multiple types of features. The similarity of each type of feature is weighted and summed to obtain the comprehensive similarity of each candidate model. The model with a comprehensive similarity greater than the similarity threshold is selected as the optimal model.

[0011] Preferably, the expression for calculating the overall similarity is:

[0012] ;

[0013] In the formula, Similarity represents the overall similarity, and N represents the total number of feature types. This represents the weight of the feature of the k-th type. This represents the feature vector of the k-th type of a clothing product image. The feature vector representing the k-th type of feature of the candidate model. The norm of a vector.

[0014] Preferably, the multiple types of features between the candidate model and the clothing product image include color features and size features.

[0015] Preferably, if the clothing product image corresponds to a loose-fitting style, the similarity threshold is set to 0.8; otherwise, the similarity threshold is set to 0.85.

[0016] Preferably, extracting the color features of clothing includes: extracting the color distribution, primary color, secondary color, proportion of cool colors, color semantic features, and color compatibility between the clothing product image and the model's background. The color semantic features include "cool colors", "warm colors", and "contrasting color design".

[0017] Preferably, the method for extracting the material features of clothing includes: obtaining the texture features of the clothing product image using a gray-level co-occurrence matrix, obtaining the material type of the clothing product image using a VGG-16 model, and using the texture features and material type as the texture features of the clothing product image.

[0018] Preferably, extracting the size features of the clothing includes: inputting the clothing product image into the YOLOv5 model to obtain a series of predicted bounding boxes; for each predicted bounding box, calculating its intersection-union ratio with the actual clothing boundary;

[0019] From all predicted bounding boxes, the bounding box with the highest intersection-union ratio (IU) is selected as the high-confidence bounding box. The pixel-level dimensions of the high-confidence bounding box and the clothing product image are then obtained. The dimensions of the high-confidence bounding box are then standardized to obtain standard clothing dimensions. The calculation expression is as follows:

[0020] ;

[0021] In the formula, Indicates standard clothing size. This represents the pixel-level size of the high-confidence bounding box. This indicates the pixel-level dimensions of the clothing product image. Indicates a reference dimension.

[0022] Preferably, extracting the style features of the clothing includes: inputting a clothing image into a preset Mask R-CNN model to obtain the attributes of each detected clothing component, wherein the clothing component includes one or more of collar, sleeve, button and zipper, wherein the attributes of the collar include collar shape; the attributes of the sleeve include sleeve length type; and the attributes of the buttons and zippers include whether they are closed or not.

[0023] Preferably, the extracted features also include: matching information for clothing items and clothing style tags.

[0024] Preferably, the extracted features also include: applicable scenario information for the clothing, which includes applicable season and applicable occasion.

[0025] The beneficial effects of this invention are as follows: when generating clothing models, it is not necessary to take images of the person beforehand; only images of the clothing itself are needed, thereby improving the efficiency of clothing model image generation. Since the size and shape of the clothing in the original clothing image are not altered, the difference between the clothing in the clothing model image and the clothing in the original clothing product image is reduced. Furthermore, the influence of the color, material, size, and style characteristics of the clothing in the clothing product image on the generated clothing model image is fully considered, resulting in the generated clothing model displaying clothing features that are closer to those in the clothing product image. When generating clothing model images, instead of relying on a single initial model, a set of initial models (i.e., multiple initial models) is set up to match the most suitable model in terms of gender, height, and body type based on the clothing product image, thereby further improving the generation effect of clothing model images. Attached Figure Description

[0026] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of the invention are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein:

[0027] Figure 1 This is a schematic flowchart illustrating a method for generating clothing model images according to an embodiment of the present invention;

[0028] Figure 2 This is a schematic diagram illustrating the structure of a clothing model image generation system according to an embodiment of the present invention. Detailed Implementation

[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0030] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0031] Example of a method for generating images of clothing models:

[0032] like Figure 1 As shown, the present invention provides a method for generating images of clothing models, including:

[0033] S101. Obtain clothing product images.

[0034] The obtained clothing product images can be a single image or multiple different images. These images can be obtained from various e-commerce platforms or by taking photos with a camera.

[0035] S102. Extract clothing features from clothing product images. Specifically, extract different types of features corresponding to the clothing based on the clothing product images. The types of features extracted include: color features, material features, size features, and style features.

[0036] Color and material characteristics can ensure that the clothing worn by the generated model is as consistent as possible with the color and material of the clothing in the product image.

[0037] Size characteristics can determine the body type matching of the model generated in subsequent steps. For example, a long dress requires a tall model.

[0038] Style features can be used to constrain the pose of the model generated in subsequent steps. For example, if the collar shape is V-neck, the model's neck needs to be shown; the way the zipper or button closes affects the dynamic effect of the generated model.

[0039] S103. Weighted fusion of the extracted different types of features to obtain fused features.

[0040] In this embodiment, an attention mechanism can be used to weight and fuse the extracted features of different types. The attention mechanism can dynamically adjust the weights of features, making the model pay more attention to task-related features and generate a more discriminative comprehensive feature representation.

[0041] S104. Generate candidate models, specifically: set up a generative adversarial network, input the initial model set and fused features into the generative adversarial network model to generate a set of candidate models, then calculate the loss function value corresponding to each candidate model, and select the candidate models whose loss function value is less than the loss threshold as candidate models.

[0042] The initial model pool includes models of various ages, heights, postures, and body types.

[0043] The expression for calculating the loss function is:

[0044]

[0045] In the formula, This represents the total loss function of the generative adversarial network. This represents the loss term of the discriminator on the real data. This represents the loss term of the generator. Let z represent the expectation of random noise z and conditional information F. This represents the logarithm of the complement of the discriminator's predicted values ​​for the generated data G(z,F). λ is the L1 regularization term, used to constrain the differences between generated data and real data in various types of extracted features. λ is a hyperparameter that balances the importance of generative adversarial loss and L1 regularization loss. This represents the L1 distance between the generated data and the real data.

[0046] Generative adversarial networks (GANs) can be used to counteract losses, thereby driving the generator to synthesize realistic images. By introducing L1 loss into the loss function, the pixel-level similarity between the generated image and the real clothing can be constrained, reducing blur.

[0047] S105. Adjust the body shape parameters and posture parameters of each candidate model. Specifically, use a 3D human posture model to adjust the body shape parameters of each candidate model to match the clothing looseness corresponding to the clothing product image, and adjust the posture parameters of each candidate model to simulate the dynamic effect of the clothing.

[0048] The expression corresponding to the 3D human pose model is:

[0049]

[0050] In the formula, θ represents the posture parameter, β represents the body shape parameter, M(θ,β) represents the 3D human body mesh model, T(β) represents the template network, J(β) represents the joint position, and W represents the linear skin weight.

[0051] The dynamic effects of clothing include the swaying of the skirt.

[0052] S106. Obtain the comprehensive similarity of each candidate model. Specifically, calculate the similarity of multiple types of features between each candidate model and the clothing product image, and sum the similarities of each type of feature by weight to obtain the comprehensive similarity of each candidate model.

[0053] In this embodiment, the expression for calculating the overall similarity is:

[0054]

[0055] In the formula, Similarity represents the overall similarity, and N represents the total number of feature types. This represents the weight of the feature of the k-th type. This represents the feature vector of the k-th type of a clothing product image. The feature vector representing the k-th type of feature of the candidate model. The norm of a vector.

[0056] S7. Select the model with a comprehensive similarity greater than the similarity threshold as the optimal model.

[0057] In this embodiment, if the clothing type corresponding to the clothing product image is a loose fit, the similarity threshold is 0.8; otherwise, the similarity threshold is 0.85.

[0058] Since loose-fitting styles are more tolerant of size mismatches, the threshold should be appropriately lowered to enhance the diversity of the selected best models.

[0059] The clothing model image generation method of this embodiment does not require pre-shooting images of the person; only images of the clothing itself are needed, thereby improving the generation efficiency of clothing model images. Furthermore, it fully considers the impact of the color, material, size, and style characteristics of the clothing in the product image on the generated clothing model image, ensuring that the characteristics of the clothing displayed by the generated model are closer to the actual characteristics of the clothing. Instead of relying on a single initial model, the method sets up a set of initial models (i.e., multiple initial models) to match the most suitable model in terms of gender, height, and body type based on the product image, further improving the generation effect of the clothing model image.

[0060] In one embodiment, the total number of feature types is 2, namely color feature and size feature.

[0061] Since color is the primary visual element, mismatched size can render the generated results unusable. Therefore, by comprehensively considering the color and size similarities between candidate models and clothing product images, the model with the best image effect can be effectively matched.

[0062] In one embodiment, the method further includes preprocessing the clothing product image after acquisition, the preprocessing including resolution normalization and noise reduction.

[0063] By standardizing the resolution of clothing product images, all input images can have the same sharpness and level of detail. This ensures that the generative model receives input of consistent quality, helping it to learn clothing features more stably. Furthermore, consistent resolution helps the generative model capture clothing details and features, such as texture, color, and shape, more accurately. This, in turn, improves the model's ability to generate high-quality images of clothing models.

[0064] Denoising clothing images can eliminate random noise or unwanted details, such as noise caused by sensor noise or poor lighting conditions. This significantly improves image quality, makes clothing details clearer, and helps clothing modeling to extract clothing features, such as edges, contours, and textures, more accurately.

[0065] In one embodiment, extracting the color features of clothing includes: extracting the color distribution, primary color, secondary color, proportion of cool colors, color semantic features, and color compatibility between the clothing product image and the model's background. The color semantic features include "cool colors", "warm colors", and "contrasting color design".

[0066] The primary color determines the overall visual harmony of the generated model, while secondary colors are used for the model's details. The color scheme affects seasonal adaptability (cool colors are suitable for summer, and warm colors are suitable for winter). By extracting the proportions of primary, secondary, and cool colors from the clothing product image, the clothes worn by the subsequently generated model can more closely resemble the clothing in the product image.

[0067] In this embodiment, the method for extracting the color distribution of a clothing product image includes: obtaining a color histogram of the clothing product image and normalizing the color histogram; when obtaining the histogram, the expression for calculating the number of pixels with pixel value i is:

[0068]

[0069] In the formula, This represents the Kronecker function, which returns 1 when I(x,y)=i, and 0 otherwise. I(x,y) represents the number of pixels with pixel value i in the image; I(x,y) represents the pixel value of the pixel at position (x,y); i represents the pixel value.

[0070] The color histogram of a clothing product image can be used to statistically analyze the pixel distribution of each color channel (R / G / B). Normalizing the color histogram can eliminate the influence of different image resolutions on histogram comparisons.

[0071] The primary and secondary colors of clothing product images can be obtained from color histograms. Since K-means clustering can adaptively extract the dominant color, the primary and secondary colors of clothing product images can be obtained by combining the k-means clustering algorithm with color histograms.

[0072] The proportion of cool colors can be calculated using a color histogram combined with the formula for the proportion of cool colors. The formula is as follows:

[0073]

[0074] In the formula Indicates the number of cool-colored pixels. This indicates the total number of pixels in the clothing product image.

[0075] In this embodiment, the ResNet-50 classification model can be used to extract the color semantic features of clothing product images.

[0076] In this embodiment, extracting the color compatibility between the clothing product image and the model's background includes:

[0077] Both the color distribution of the clothing product image and the color distribution of the model's background are treated as probability distributions. The KL divergence metric is used to quantify the difference between these two color distributions, which is then used as a measure of the color compatibility between the clothing product image and the model's background.

[0078] The expression for calculating the KL divergence is:

[0079]

[0080] In the formula, Denotes KL divergence, This represents the probability of the i-th color level in the color distribution of the clothing product image. Let represent the probability of the i-th color level in the background color distribution of the model, and lg represent the logarithmic function with base 10.

[0081] In one embodiment, a method for extracting the material characteristics of clothing includes:

[0082] The texture features of the clothing product image are obtained using the gray-level co-occurrence matrix, and the material type of the clothing product image is obtained using the VGG-16 model; the texture features and material type are used as the texture features of the clothing product image.

[0083] When using the gray-level co-occurrence matrix to obtain the texture features of clothing product images, the expression for calculating contrast is:

[0084]

[0085] The expression for calculating correlation is:

[0086]

[0087] In the two formulas above, Contrast represents contrast. and Index representing grayscale value, This indicates that a pixel with gray value i and a pixel with gray value j in an image are at a specific distance d and direction. The probability of both appearing together; Correlation represents correlation. and Let I(x,y) and I(x+d,y+d) represent the mean gray values ​​of all pixel pairs satisfying I(x,y)=i and I(x+d,y+d)=j, respectively. Here, I(x,y) represents the pixel gray value at position (x,y) in the image, and d is the distance between pixel pairs.

[0088] In one embodiment, extracting the size features of the garment includes:

[0089] S201. Input the clothing product image into the YOLOv5 model to obtain a series of predicted bounding boxes;

[0090] For each predicted bounding box, calculate its intersection-union ratio (IU) with the actual garment boundary using the following formula:

[0091]

[0092] In the formula, IoU represents the intersection-union ratio between the predicted bounding box and the actual clothing boundary. Indicates the predicted bounding box. "Area" represents the actual boundary of the clothing.

[0093] S202. Select the bounding box with the largest intersection-union ratio (IU) from all predicted bounding boxes as the high-confidence bounding box, and obtain the pixel-level dimensions of the high-confidence bounding box and the clothing product image. Then, standardize the dimensions of the high-confidence bounding box to obtain standard clothing dimensions. The calculation expression is as follows:

[0094]

[0095] In the formula, Indicates standard clothing size. This represents the pixel-level size of the high-confidence bounding box. This indicates the pixel-level dimensions of the clothing product image. Indicates a reference dimension.

[0096] The size characteristics of clothing include garment length and looseness threshold. The formula for calculating garment length L is:

[0097]

[0098] In the formula, and These represent the maximum and minimum y-axis coordinates of the high-confidence bounding box, respectively.

[0099] The expression for calculating the leniency threshold R is:

[0100]

[0101] In the formula, and These represent the area of ​​the clothing and the area of ​​the model's body, respectively.

[0102] By setting a leniency threshold, clothing distortion on subsequent candidate models can be avoided.

[0103] In one embodiment, extracting the style features of clothing includes: inputting a clothing image into a preset Mask R-CNN model to obtain the attributes of each detected clothing component, wherein the clothing component includes one or more of collar, sleeve, button, and zipper, wherein the attributes of the collar include collar shape; the attributes of the sleeve include sleeve length type; and the attributes of the buttons and zippers include whether they are closed or not.

[0104] The training process of the Mask R-CNN model includes:

[0105] S301. Collect images of clothing items that include different collar types, different sleeve lengths, and different button closure states;

[0106] Among them, the collar type includes round neck, V-neck, square neck, etc.; the sleeve length includes: short sleeve, long sleeve, three-quarter sleeve, etc.; the button closure status includes: fully open, half open, closed, etc.

[0107] S302. Represent each garment product image at the component level, draw a precise segmentation mask for each garment component, and label its attributes to obtain the dataset; then divide the dataset into a training set and a validation set.

[0108] S303. Train the Mask R-CNN model using the training set. During the training process, optimize the model parameters to improve segmentation accuracy and attribute classification accuracy.

[0109] S304. Evaluate the model's performance using a validation set, including metrics such as segmentation accuracy and attribute classification accuracy. Adjust model parameters or training strategies based on the evaluation results to improve model performance.

[0110] In one embodiment, the extracted features may also include: matching information for clothing items and clothing style tags.

[0111] Methods for obtaining information on complete sets include:

[0112] Use object detection models to identify different parts in clothing product images;

[0113] The Cross-Union Ratio (LOU) for different component regions is calculated using the following formula:

[0114]

[0115] In the formula, A and B represent the regions of the two identified components, respectively. B represents the area of ​​the overlapping portion of the two components, A B represents the total area of ​​the two components.

[0116] Clothing product images with an intersection-union ratio greater than 0.6 are classified as sets.

[0117] The title information of clothing product images can be parsed using NLP (Natural Language Processing) models to obtain clothing style tags.

[0118] NLP can supplement the semantic information missing from visual features.

[0119] Style tags can guide the generation of backgrounds for models (e.g., formal attire requires a simple setting).

[0120] A suit typically comprises multiple complementary clothing items that coordinate in style, color, and material. Identifying whether clothing in a product image is a suit helps generative adversarial networks maintain stylistic harmony and overall coherence among these individual items when generating model images. This results in more natural and harmonious clothing combinations in the generated model images, conforming to fashion rules and aesthetic standards.

[0121] In one embodiment, the extracted features may also include: applicable scenario information for the clothing, which includes applicable season and applicable occasion.

[0122] The method for obtaining applicable scenario information is as follows: input the clothing product image into a preset ResNet-18 model to obtain the probability distribution of seasonal and occasion classifications. Select the category with the highest probability as the final result, that is, the season (Chinese New Year / Summer / Autumn / Winter) and occasion (sports / business) for which the clothing is applicable.

[0123] In this embodiment, the training process of the ResNet-18 model is as follows:

[0124] S401. Construct a dataset: Collect and label images of clothing items for different seasons (summer / winter) and occasions (sports / business). Ensure the dataset covers a variety of clothing styles, colors, materials, etc.

[0125] S401. Data Partitioning: Divide the dataset into training and validation sets, typically in a ratio of 8:2 or 7:3, to ensure a reasonable data distribution.

[0126] S403. Preprocess the training and validation data, including image enhancement and normalization.

[0127] Image enhancement: Perform image enhancement operations on the training set, such as random cropping, flipping, adjusting brightness / contrast, etc., to increase data diversity and improve the model's generalization ability.

[0128] Standardization processing: The image size is uniformly adjusted to the default size of the ResNet-18 model input (e.g., 224x224), and normalization processing is performed (e.g., scaling the pixel values ​​to the range of [0,1]).

[0129] S404, Model training, including:

[0130] 1) Loading a pre-trained model: Use a ResNet-18 model pre-trained on a large dataset (such as ImageNet) as a starting point and leverage transfer learning techniques to accelerate model convergence.

[0131] 2) Adjust the model structure: Modify the output layer of the ResNet-18 model, replacing the last fully connected layer with two independent classification heads, one for seasonal classification (summer / winter) and the other for occasion classification (sports / business). The number of neurons in each classification head is set according to the number of categories (2 for seasonal classification and 2 for occasion classification).

[0132] 3) Set hyperparameters:

[0133] Optimizer: Select Adam or AdamW optimizer. The initial learning rate can be set to 1e-4 or less, and adjusted according to the training situation.

[0134] Loss function: Use the cross-entropy loss function to calculate the loss for seasonal classification and occasion classification respectively, and then perform a weighted sum (e.g., each with a weight of 0.5).

[0135] Batch size: Set according to the video memory size, usually 32 or 64.

[0136] Training cycle: Adjusted according to the performance of the validation set, it can be set to 10-20 cycles.

[0137] 4) Model training:

[0138] The preprocessed training set is input into the model for training, and the model performance is evaluated on the validation set after each epoch. Hyperparameters, such as learning rate decay strategies and data augmentation methods, are adjusted based on the validation set results.

[0139] By obtaining the applicable season, the ambient lighting of the generated model can be constrained (e.g., warm colors are needed in winter), and by obtaining the applicable occasion, the pose of the generated model can be matched (e.g., dynamic poses are needed for sportswear).

[0140] In one embodiment, the extracted features may also include: the logo position on the garment, the pattern on the garment, and the functional label on the garment. The logo position includes two locations: the left chest and the cuff. The logo pattern includes two patterns: stripes and prints. The functional label on the garment includes waterproofing.

[0141] By obtaining the logo's location, we can ensure that the clothing logo is clearly visible in the generated model image.

[0142] Since complex patterns require high-resolution rendering, obtaining complex patterns ensures that the generated model image clearly displays the complex patterns.

[0143] Since function tags affect material rendering parameters, identifying these tags allows for adjustments to the rendering parameters of the generated model image. For example, a waterproof function requires applying a matte effect to the model image.

[0144] In this embodiment, the method for identifying the logo position is as follows:

[0145] S501. Use the Tesseract model to perform OCR recognition on clothing product images to extract text information from the clothing product images;

[0146] S501. Extract the coordinate information of the text from the recognition results of the Tesseract model to locate the logo position.

[0147] In this embodiment, a SIFT model can be used to obtain the pattern on the clothing product image, including:

[0148] 1) Install the OpenCV library.

[0149] 2) Read the clothing product image. Use OpenCV's imread function to read the clothing product image.

[0150] 3) Create a SIFT object and detect feature points: Use OpenCV's SIFT_create function to create a SIFT object, and then use the detectAndCompute method to detect keypoints and calculate descriptors:

[0151] 4) Drawing key points: Key points can be drawn on the original image using OpenCV's drawKeypoints function.

[0152] 5) Pattern matching using descriptors: Use the SIFT algorithm to extract key point descriptors from two images, and then use feature matching algorithms (such as FLANN or BFMatcher) for matching.

[0153] 6) After obtaining the pattern on the clothing product image, the entropy formula is used to quantify the pattern complexity.

[0154] Quantifying pattern complexity using the entropy formula ensures that pattern details are not lost when generating modalities.

[0155] The entropy formula is:

[0156]

[0157] In the formula, H represents the entropy value. Let represent the probability of random variable i, and log represent the logarithmic function.

[0158] Example of a clothing model image generation system:

[0159] This invention also provides a system for generating images of clothing models, such as... Figure 2As shown, the clothing model image generation system includes a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement the clothing model image generation method in the above embodiments.

[0160] The system also includes other components well known to those skilled in the art, such as communication buses and communication interfaces, the settings and functions of which are known in the art and therefore will not be described in detail here.

[0161] In the description of this specification, "multiple" means at least two, such as two, three or more, etc., unless otherwise expressly and specifically defined.

[0162] While this specification has shown and described numerous embodiments of the invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of this invention.

Claims

1. A method for generating images of clothing models, characterized in that, include: Obtain clothing product images and extract different types of features from the clothing product images. The types of features extracted include: color features, material features, size features, and style features. The extracted features of different types are weighted and fused to obtain fused features; The initial model set and fused features are input into the generative adversarial network model to generate a candidate model set. Then, the loss function value corresponding to each candidate model is calculated, and the candidate models whose loss function value is less than the loss threshold are selected as candidate models. The body shape parameters of each candidate model are adjusted using a 3D human posture model to match the looseness of the clothing corresponding to the clothing product image, and the posture parameters of each candidate model are adjusted to simulate the dynamic effect of the clothing. The similarity of each candidate model to the clothing product image is calculated for multiple types of features. The similarity of each type of feature is weighted and summed to obtain the comprehensive similarity of each candidate model. The model with a comprehensive similarity greater than the similarity threshold is selected as the optimal model. The similarity threshold is calculated as follows: if the clothing product image corresponds to a loose-fitting style, the similarity threshold is 0.8; otherwise, the similarity threshold is 0.

85. The size feature includes a looseness threshold, and the expression for calculating the looseness threshold R is: In the formula, and These represent the area of ​​the clothing and the area of ​​the model's body, respectively. The formula for calculating the overall similarity is: ; In the formula, Similarity represents the overall similarity, and N represents the total number of feature types. This represents the weight of the feature of the k-th type. This represents the feature vector of the k-th type of a clothing product image. The feature vector representing the k-th type of feature of the candidate model. The norm of a vector; Extracting the size features of clothing includes: inputting the clothing product image into the YOLOv5 model to obtain a series of predicted bounding boxes; for each predicted bounding box, calculating its intersection-union ratio with the actual clothing boundary; From all predicted bounding boxes, the bounding box with the highest intersection-union ratio (IU) is selected as the high-confidence bounding box. The pixel-level dimensions of the high-confidence bounding box and the clothing product image are then obtained. The dimensions of the high-confidence bounding box are then standardized to obtain standard clothing dimensions. The calculation expression is as follows: In the formula, Indicates standard clothing size. This represents the pixel-level size of the high-confidence bounding box. This indicates the pixel-level dimensions of the clothing product image. Indicates a reference dimension.

2. The method for generating clothing model images as described in claim 1, characterized in that, Several types of characteristics exist between candidate models and clothing product images, including color and size characteristics.

3. The method for generating clothing model images as described in claim 1, characterized in that, Extracting color features of clothing includes: extracting the color distribution, primary color, secondary color, proportion of cool colors, color semantic features, and color compatibility between the clothing product image and the model's background. The color semantic features include "cool colors", "warm colors", and "contrasting color design".

4. The method for generating clothing model images as described in claim 1, characterized in that, Methods for extracting material features of clothing include: using the gray-level co-occurrence matrix to obtain the texture features of clothing product images, using the VGG-16 model to obtain the material type of clothing product images, and using the texture features and material type as the texture features of clothing product images.

5. The method for generating clothing model images as described in claim 1, characterized in that, Extracting the style features of clothing includes: inputting the clothing product image into a preset Mask R-CNN model to obtain the attributes of each detected clothing component, which includes one or more of collar, sleeve, button and zipper, wherein the collar attribute includes collar shape; the sleeve attribute includes sleeve length type; and the button and zipper attribute includes whether they are closed or not.

6. The method for generating clothing model images as described in any one of claims 1 to 5, characterized in that, The types of features extracted also include: matching information for clothing product images and clothing style tags.

7. A clothing model image generation system, comprising a memory and a processor, wherein the memory stores computer program instructions, characterized in that, When the program instructions are executed by the processor, they implement the clothing model image generation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Virtual clothes changing method, virtual clothes changing device and electronic equipment

    CN118334285A

  • Clothing matching recommendation method and system based on multiple modes

    CN119323708A

  • Method and device for one-key generation of dummy model clothing image, and storage medium

    CN119810264A