Clothing attribute identification method and system, computer equipment and readable storage medium
Through crawler technology and visual language big models, professional clothing data sets are constructed, and a lightweight network structure based on Transformer and CNN is designed, which solves the problem of limited data sets and limited model representation capabilities in clothing attribute recognition, and achieves the multi-attribute recognition effect of clothing with high accuracy and low inference delay.
Patent Information
- Application Number
- CN202510089776.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
The clothing industry has limited open source data sets, traditional machine learning algorithms or convolutional neural network algorithms cannot meet the massive and diverse fine-grained clothing attribute recognition needs, and a single task and single branch network limits the representation ability and generalization ability of the model.
Crawl massive clothing data through crawling technology, merge part of the open source data set, and pre-notate it with the powerful content understanding ability of the visual language model to build a professional clothing attribute data set. Combining the advantages of Transformer architecture and CNN, a lightweight network structure based on local sliding frames and key point regression as auxiliary tasks are designed to solve the problem of multi-attribute recognition of clothing.
The classification accuracy of average 96.13% was achieved on the test set, which proved the feasibility and effectiveness of this method, taking into account the balance between model accuracy and inference delay.
Smart Images

Figure CN119992202A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a clothing attribute recognition method, a clothing attribute recognition system, a computer device and a readable storage medium. Background Art
[0002] In real life, clothing attribute recognition (style recognition) can help fashion recommendation systems understand users' preferences and styles more accurately, thereby providing users with personalized fashion suggestions and recommendations. Secondly, on e-commerce platforms, more accurate clothing searches and classifications can be achieved to help users quickly find the styles and styles they need. Users can also try on different styles of clothing online through virtual fitting rooms without actually putting on the clothes, providing a more convenient shopping experience. For fashion brands and retailers, clothing attribute recognition can help them understand market trends, analyze consumer preferences and purchasing behaviors, and better manage brands and formulate marketing strategies. The above application scenarios reflect the importance of clothing attribute recognition in the fashion industry, e-commerce, market research, creative design and other fields, and have a positive impact on improving user experience and promoting business development.
[0003] Earlier, machine learning algorithms were used to solve image recognition tasks. Such methods usually require manual feature extraction, which becomes difficult and time-consuming when faced with large-scale and complex clothing image data. It is also difficult to capture high-level semantic information and fine-grained features in the image, resulting in limited classification performance. With the rise of deep learning, convolutional neural networks (CNNs) have been widely used in image classification tasks. CNNs can automatically learn feature representations from raw images without the need for manual feature extraction. Secondly, the existence of inductive bias (local connections, translation invariance) makes the CNN structure very suitable for the image field. This end-to-end learning method has achieved certain results in clothing attribute recognition tasks, but because CNN was originally designed to process image data with local correlations, it has a weak ability to model global information and cannot fully capture the overall structure and semantics of the image.
[0004] Due to the diversity and variability of clothing styles, traditional image datasets are often limited and difficult to cover all possible styles. In order to solve the problem of datasets, researchers began to build larger-scale clothing image datasets, such as Deepfashion, FashionAI, etc., in order to better train and evaluate models. In addition, the emergence of large-scale visual language models (such as the language model GPT-4 released by artificial intelligence company OpenAI for the chatbot ChatGPT, the large-scale visual language model Qwen-VL launched by Alibaba Cloud, and the multimodal Chinese-English bilingual dialogue language model VisualGLM, etc.) has brought new possibilities for clothing attribute recognition. With the powerful representation and understanding capabilities of large models, users can generate detailed attributes for clothing images through question and answer (Question & Answer, Q&A) to provide more comprehensive clothing information. As the size of the dataset increases, the recognition ability of CNN is limited.
[0005] Deep Learning Model Based on Self-Attention Mechanism Transformer is a deep learning model based on self-attention mechanism, which initially achieved great success in the field of Natural Language Processing (NLP). Transformer can capture global dependencies when processing sequence data, has strong parallel computing capabilities and good representation learning capabilities. The success of this self-attention mechanism has triggered researchers to explore Transformer in the field of computer vision. Among them, Vision Transformer (VIT) is a model that applies Transformer to image classification tasks. VIT flattens images into sequences and establishes relationships between sequences through self-attention mechanisms, achieving the effect of image classification and proving the feasibility of Transformer in the field of computer vision. However, some problems of VIT have also been observed. First, VIT's operation of flattening the input image into sequence features may lead to the loss of detail information, affecting the capture of fine-grained features. Second, VIT calculates attention based on global features, and faces memory and computing power limitations when processing larger images.
[0006] However, the open source datasets in the clothing industry are limited. Although there are large-scale datasets such as Deepfashion and FashionAI, they lack professional annotations and cannot meet the needs of complex clothing applications. It is necessary to create more professional datasets for the clothing industry. In addition, traditional machine learning algorithms or convolutional neural network algorithms cannot meet the needs of massive and diverse fine-grained clothing recognition. Single-task single-branch networks limit the model's representation and generalization capabilities. In addition, the balance between model accuracy and inference latency is also an aspect that is not considered. Summary of the invention
[0007] The technical problem to be solved by the present invention is that the open source data sets in the clothing industry are limited, and traditional machine learning algorithms or convolutional neural network algorithms cannot meet the attribute recognition needs of massive and diverse fine-grained clothing. The single-task single-branch network limits the model's representation and generalization capabilities.
[0008] In view of the above-mentioned deficiencies of the prior art, the following solutions are provided:
[0009] In a first aspect, the present invention provides a method for identifying clothing attributes, comprising: obtaining a picture to be identified; obtaining a picture text database of clothing, and obtaining a clothing attribute recognition model based on the picture text database of clothing and a deep learning model; and identifying the attributes of the clothing in the picture to be identified according to the clothing attribute recognition model to obtain attribute data of the clothing in the picture to be identified. The picture text database includes multiple clothing pictures and text information corresponding to each clothing picture in the multiple clothing pictures.
[0010] Optionally, obtaining the clothing image text database includes: obtaining multiple clothing pictures; obtaining text information corresponding to each clothing picture in the multiple clothing pictures. And, obtaining the clothing image text database according to the multiple clothing pictures and the text information corresponding to each clothing picture in the multiple clothing pictures. The text information corresponding to each clothing picture in the multiple clothing pictures is related to the attributes of the clothing in each clothing picture.
[0011] Optionally, obtaining a plurality of clothing pictures and text information corresponding to each clothing picture in the plurality of clothing pictures includes: obtaining at least one clothing picture and text information corresponding to each clothing picture in at least one clothing picture obtained from at least one e-commerce platform according to crawler technology. And, obtaining at least one clothing picture and text information corresponding to each clothing picture in at least one clothing picture in the public database from a public database. Among them, the plurality of clothing pictures include at least one clothing picture obtained from at least one e-commerce platform and at least one clothing picture obtained from a public database. The text information corresponding to each clothing picture in the plurality of clothing pictures includes text information corresponding to each clothing picture in at least one clothing picture obtained from at least one e-commerce platform and text information corresponding to each clothing picture in at least one clothing picture in the public database.
[0012] Optionally, obtaining a picture text database of clothing according to a plurality of clothing pictures and text information corresponding to each clothing picture in the plurality of clothing pictures includes: using an image encoder in a contrastive learning algorithm to perform feature extraction on each clothing picture in the plurality of clothing pictures to obtain a plurality of feature-extracted data. Deduplication is performed on the plurality of feature-extracted data to obtain deduplicated feature-extracted data, and a picture database of clothing is obtained according to clothing pictures corresponding to the deduplicated feature-extracted data. Also, named entity recognition and classification are performed on the text information corresponding to each clothing picture in the plurality of clothing pictures to obtain a clothing attribute database in each clothing picture in the plurality of clothing pictures. The picture text database of clothing includes a picture database of clothing and a clothing attribute database.
[0013] Optionally, obtaining a picture-text database of clothing based on multiple clothing pictures and text information corresponding to each clothing picture in the multiple clothing pictures also includes: performing image question answering using an open source large-scale visual language model based on the multiple clothing pictures and text information corresponding to each clothing picture in the multiple clothing pictures, expanding attribute data of clothing in each clothing picture in the multiple clothing pictures to obtain an expanded attribute database of clothing, and the picture-text database of clothing includes the picture database of clothing and the expanded attribute database of clothing.
[0014] Optionally, the attributes of clothing include multiple major attribute categories, each of the multiple major attribute categories includes at least one sub-attribute category, and the multiple major attribute categories include color, sleeve length, waist type, sleeve type, type, trouser length, version, style, thickness, craftsmanship, collar type, function, popular elements, slit design, placket, trouser type, fabric, skirt type, skirt length, scene, crowd, length, pattern and season.
[0015] Optionally, obtaining a clothing attribute recognition model according to a clothing image text database and a deep learning model includes: building a deep learning model according to a block attention mechanism Swin Transformer. Making a model data set according to the clothing image text database. And, training the deep learning model built according to Swin Transformer according to the model data set to obtain a clothing attribute recognition model.
[0016] Optionally, the deep learning model built according to Swin Transformer includes: a block layer PatchPartition, a first downsampling stage, a second downsampling stage, a third downsampling stage, a fourth downsampling stage, a key point regression branch and a multi-attribute classification head, wherein PatchPartition, the first downsampling stage, the second downsampling stage, the third downsampling stage, the fourth downsampling stage and the multi-attribute classification head are connected in series in sequence. The first downsampling stage includes a linear transformation layer Linear Embedding and at least two self-attention blocks Swin Transformer blocks. The second downsampling stage, the third downsampling stage and the fourth downsampling stage respectively include a downsampling layer Patch Merging and at least two self-attention blocks Swin Transformer blocks. The key point regression branch is used to downsample and dimensionally adjust the output of the first downsampling stage of Swin Transformer to obtain the adjusted output of the first downsampling stage, downsample and dimensionally adjust the output of the second downsampling stage of Swin Transformer to obtain the adjusted output of the second downsampling stage, and perform feature fusion on the adjusted output of the first downsampling stage, the adjusted output of the second downsampling stage, and the output of the third downsampling stage of Swin Transformer.
[0017] Optionally, the first downsampling stage includes a linear transformation layer Linear Embedding and two self-attention blocks Swin Transformer block. The second downsampling stage and the fourth downsampling stage respectively include a downsampling layer Patch Merging and two self-attention blocks Swin Transformer block. The third downsampling stage includes a downsampling layer Patch Merging and 18 self-attention blocks Swin Transformer block.
[0018] Optionally, the deep learning model built according to Swin Transformer is trained according to the model data set to obtain a clothing attribute recognition model, including: setting the data enhancement method, optimizer and learning rate. Also, the deep learning model built according to Swin Transformer is trained according to the model data set, and the training parameters of the deep learning model built according to Swin Transformer are fine-tuned according to the pre-trained model to obtain a clothing attribute recognition model. Among them, the pre-trained model is obtained by training the visual model Swin-Transformer-Small using the data set ImageNet-22K. The loss function Loss during training is Loss = λL landmark +(1-λ)L ce Among them, L landmark is the key point regression branch loss function, and L ce is the multi-attribute recognition loss function, and n is the number of the current image sample, N is the total number of image samples, m is the number of the major attribute category of the current image sample, M is the total number of major attribute categories included in at least one major attribute category of the current image sample, l is the number of the sub-attribute category of the current image sample, L is the total number of sub-attribute categories in the major attribute category numbered m of the current image sample, k is the key point number of the current image sample, K is the total number of key points of the current image sample, and λ is L ce and L landmark The weight parameter v k The visibility of the current key point, the value is 0 or 1; D n (x k ,y k ) is the predicted coordinate of the key point numbered k in the image sample numbered n, is the true coordinate of the key point numbered k in the image sample numbered n; represents the true value of the image sample with image number n, major attribute category number m, and sub-attribute category number l. Represents the probability value of the model output for the image sample with image number n, major attribute category number m, and sub-attribute category number l.
[0019] In a second aspect, the present invention provides a clothing attribute recognition system, comprising a to-be-recognized picture acquisition module, a data and model acquisition module, and an attribute recognition module. The to-be-recognized picture acquisition module is configured to: acquire a to-be-recognized picture. The data and model acquisition module is configured to: acquire a picture text database of clothing, and acquire a clothing attribute recognition model based on the picture text database of clothing and a deep learning model. The attribute recognition module is configured to: recognize the attributes of clothing in the to-be-recognized picture according to the clothing attribute recognition model, and obtain attribute data of clothing in the to-be-recognized picture.
[0020] In a third aspect, the present invention provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, the processor executes the above-mentioned clothing attribute recognition method.
[0021] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor executes the above-mentioned clothing attribute recognition method.
[0022] The clothing attribute recognition method, system, computer device and readable storage medium provided by the present invention crawl massive clothing data through crawler technology, merge some open source data sets, and pre-label with the help of the powerful content understanding ability of the visual language large model to build a professional clothing attribute data set. In terms of algorithms, the advantages of the Transformer architecture and CNN are combined to design a lightweight network structure based on the self-attention mechanism of the local sliding frame and key point regression as an auxiliary task to solve the problem of clothing multi-attribute recognition, taking into account model accuracy and reasoning delay. Finally, an average classification accuracy of 96.13% was achieved on the test set collected by the present invention, proving the feasibility and effectiveness of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a flow chart of a clothing attribute recognition method in an embodiment of the present invention;
[0024] Figure 2 is a flow chart of another clothing attribute recognition method in an embodiment of the present invention;
[0025] Figure 3 A flowchart of another clothing attribute recognition method in an embodiment of the present invention;
[0026] Figure 4 is a flow chart of another clothing attribute recognition method in an embodiment of the present invention;
[0027] Figure 5 is a schematic diagram of clothing attributes in an embodiment of the present invention;
[0028] Figure 6 is a flow chart of another clothing attribute recognition method in an embodiment of the present invention;
[0029] Figure 7 A schematic diagram of the network structure of a deep learning model built according to Swin Transformer in an embodiment of the present invention;
[0030] Figure 8 is a flow chart of another clothing attribute recognition method in an embodiment of the present invention;
[0031] Fig. 9 is a flowchart of a clothing attribute recognition method in an example of an embodiment of the present invention;
[0032] Fig.10 is a structural diagram of a clothing attribute recognition system in an embodiment of the present invention;
[0033] Fig.11 is a structural diagram of another clothing attribute recognition system in an embodiment of the present invention;
[0034] Fig.12 A structural diagram of another clothing attribute recognition system in an embodiment of the present invention;
[0035] Fig.13 The figure is a structural diagram of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION
[0036] In order to enable those skilled in the art to better understand the technical solution of the present invention, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0037] It should be understood that the specific embodiments and drawings described herein are only used to explain the present invention rather than to limit the present invention.
[0038] It can be understood that, in the absence of conflict, the various embodiments of the present invention and the various features in the embodiments can be combined with each other.
[0039] It can be understood that, for the convenience of description, the drawings of the present invention only show the parts related to the present invention, while the parts irrelevant to the present invention are not shown in the drawings.
[0040] It can be understood that each unit or module involved in the embodiments of the present invention may correspond to only one physical structure, or may be composed of multiple physical structures, or multiple units or modules may be integrated into one physical structure.
[0041] It can be understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present invention may occur in an order different from that marked in the drawings.
[0042] It is understood that the flowcharts and block diagrams of the present invention illustrate the possible architectures, functions, and operations of the systems, devices, equipment, and methods according to the various embodiments of the present invention. Each box in the flowchart or block diagram may represent a unit, module, program segment, or code, which contains executable instructions for implementing the specified functions. Moreover, each box or combination of boxes in the block diagram and flowchart may be implemented by a hardware-based system that implements the specified functions, or may be implemented by a combination of hardware and computer instructions.
[0043] It can be understood that the units and modules involved in the embodiments of the present invention can be implemented by software or hardware. For example, the units and modules can be located in a processor.
[0044] The embodiment of the present invention provides a method for identifying clothing attributes. Figure 1 As shown, the clothing attribute recognition method includes steps 101 to 103.
[0045] Step 101: Obtain a picture to be identified.
[0046] It can be understood that the image information displayed by the picture to be identified includes information of the clothing picture.
[0047] Step 102: Obtain a clothing image-text database, and obtain a clothing attribute recognition model based on the clothing image-text database and the deep learning model.
[0048] In some embodiments, Figure 2 As shown, the method of obtaining the clothing image text database in step 102 includes steps 201 to 203.
[0049] Step 201: Obtain multiple clothing pictures.
[0050] Step 202: Obtain text information corresponding to each clothing picture in the plurality of clothing pictures.
[0051] In step 202, the text information corresponding to each clothing picture in the plurality of clothing pictures is related to the attribute of the clothing in each clothing picture.
[0052] In some embodiments, Figure 3 As shown, the implementation method of step 201 and step 202 includes step 2011 to step 2012.
[0053] Step 2011: Obtain at least one clothing picture and text information corresponding to each clothing picture in the at least one clothing picture obtained from the at least one e-commerce platform according to crawler technology.
[0054] Exemplarily, at least one e-commerce platform may include mainstream e-commerce platforms at home and abroad, and clothing pictures and corresponding text descriptions of the pictures may be crawled from the product homepages of the mainstream e-commerce platforms at home and abroad to solve the problem of lack of professional data sets in the clothing industry.
[0055] Step 2012: Obtain from a public database at least one clothing picture and text information corresponding to each clothing picture in the at least one clothing picture in the public database.
[0056] Exemplarily, the public database may include a public Deepfashion dataset. For example, more than 100,000 full-body clothing images are extracted from the public Deepfashion dataset.
[0057] The multiple clothing pictures in step 201 include at least one clothing picture obtained from at least one e-commerce platform in step 2011 and at least one clothing picture obtained from a public database in step 2012. The text information corresponding to each clothing picture in the multiple clothing pictures in step 202 includes text information corresponding to each clothing picture in at least one clothing picture obtained from at least one e-commerce platform in step 2011 and text information corresponding to each clothing picture in at least one clothing picture in a public database in step 2012.
[0058] Step 203: Acquire a clothing picture text database according to the plurality of clothing pictures and the text information corresponding to each clothing picture in the plurality of clothing pictures.
[0059] In some embodiments, Figure 4 As shown, the implementation method of step 203 includes steps 2031 to 2033.
[0060] Step 2031: Use the image encoder in the contrastive learning algorithm to extract features from each of the multiple clothing pictures to obtain multiple feature-extracted data.
[0061] For example, the image encoder Image Encoder of the multimodal pre-training neural network (Contrastive Language-Image Pre-Training, CLIP) can be used to extract image embedding. The image encoder Image Encoder of CLIP can extract the features of the image and save it as a low-dimensional vector representation. Among them, CLIP is also a contrastive learning model. The core idea of CLIP is to use a large amount of paired data of images and texts for pre-training to learn the alignment relationship between images and texts. The model has the ability of multimodal learning and can simultaneously understand the information of two different modalities, image and text, and establish a connection between them. CLIP consists of two parts, a text encoder TextEncoder and an image encoder ImageEncoder. Text Encoder is used to convert text into a low-dimensional vector representation, while ImageEncoder is used to convert images into similar vector representations. In the prediction stage, the CLIP model generates predictions by calculating the cosine similarity between text and image vectors. This model is particularly suitable for zero-sample learning tasks, that is, the model does not need to see new training examples of images or texts to make predictions. The CLIP model performs well in many fields, such as image text retrieval, image and text generation, etc. In addition, the CLIP model has high generalization ability and can adapt to different tasks and scenarios.
[0062] Step 2032: deduplicate the plurality of feature-extracted data to obtain deduplicated feature-extracted data, and obtain a clothing image database based on clothing images corresponding to the deduplicated feature-extracted data.
[0063] For example, the cosine similarity between image embeddings may be calculated to remove duplicates.
[0064] Exemplarily, through step 2021 and step 2022, more than 100,000 e-commerce clothing pictures can be obtained.
[0065] Step 2033: performing named entity recognition and classification on the text information corresponding to each clothing picture in the plurality of clothing pictures, and obtaining a database of attributes of the clothing in each clothing picture in the plurality of clothing pictures.
[0066] For example, the language representation model (Bidirectional Encoder Representations from Transformers, BERT) can be used for training, and the text information corresponding to each clothing image in multiple clothing images can be recognized by named entity recognition (NER) and classified into multiple clothing attribute categories. In addition, for more than 100,000 full-body clothing images extracted from the public Deepfashion dataset, the key point attribute information of this part of the Deepfashion dataset can be retained.
[0067] In some embodiments, Figure 5 As shown, the attributes of clothing include multiple major attribute categories, each of the multiple major attribute categories includes at least one sub-attribute category, and the multiple major attribute categories include color, sleeve length, waist type, sleeve type, type, trouser length, version, style, thickness, craftsmanship, collar type, function, popular elements, slit design, placket, trouser type, fabric, skirt type, skirt length, scene, crowd, length, pattern and season.
[0068] like Figure 5 As shown, three sub-attribute categories are listed under each major attribute category. For example, when the major attribute category is color, the sub-attribute categories include blue, red, and yellow.
[0069] The clothing image text database in step 203 includes a clothing image database and a clothing attribute database.
[0070] In some embodiments, the implementation method of step 203 also includes: performing image question answering using an open source large-scale visual language model based on multiple clothing pictures and text information corresponding to each clothing picture in the multiple clothing pictures, expanding the attribute data of the clothing in each clothing picture in the multiple clothing pictures to obtain an expanded clothing attribute database, and the clothing image text database includes a clothing image database and an expanded clothing attribute database.
[0071] Exemplarily, Alibaba's open source Large Vision Language Model (LVLM) Qwen-VL can be used for image question answering to obtain richer attribute data to expand the attribute data obtained in step 2023. A total of more than 200,000 pairs of clothing image-text pair data can be obtained, and each clothing image can have 24 clothing attribute data.
[0072] Understandably, for the task of text-generated images or image-generated images, in order to reduce the interference caused by other objects or image backgrounds other than the main body of the clothes, the clothing Matting model trained based on the U2Net algorithm is used to cut out the clothes separately, so as to facilitate the subsequent training of a generative model with richer clothing details. The U2Net algorithm is a deep learning model for salient object detection (SOD), which adopts a novel nested U-shaped structure design, which can deepen the network depth and maintain high resolution without significantly increasing memory and computing costs. The Matting model is an image processing technology that is mainly used to extract foreground objects from complex backgrounds.
[0073] In some embodiments, Figure 6 As shown, in step 102, the implementation method of obtaining a clothing attribute recognition model based on a clothing image text database and a deep learning model includes steps 601 to 603.
[0074] Step 601: Build a deep learning model based on Swin Transformer.
[0075] Understandably, the Transformer-based deep learning model Swin Transformer effectively captures the details and local features in the image by introducing a hierarchical window mechanism to replace the traditional image segmentation strategy. In addition, a hierarchical, cross-window self-attention mechanism is adopted to overcome the limitations of computing and memory resources, enabling the model to process larger images. Compared with VIT, Swin Transformer shows better performance in image classification tasks, especially when dealing with large images and fine-grained features, and has potential application prospects in computer vision-related tasks.
[0076] In order to avoid the problem of traditional CNN's ability degradation when processing large-scale data sets and the large amount of global attention calculation of VIT, the embodiments of the present invention draw on the excellent prior knowledge of visual signals, adopt the locality and translation invariance characteristics of CNN, and perform multi-head self-attention (Shifted Window Multi-head Self-Attention, SW-MSA) calculation in the moving window, which effectively reduces the amount of calculation. Swin Transformer generally includes 4 downsampling stages, each of which reduces the resolution of the input feature map and expands the receptive field layer by layer. In addition. Considering the close relationship between the clothing joint point regression task and the multi-attribute recognition task, the key point regression branch is introduced at the end of the third downsampling stage, and joint training is used to assist in improving the accuracy of the main branch. At the end of the fourth downsampling stage, the multi-attribute classification head is connected, and the overall network structure is as follows. Figure 7shown.
[0077] In some embodiments, Figure 7 As shown, the deep learning model built according to Swin Transformer includes: a block layer Patch Partition, a first downsampling stage, a second downsampling stage, a third downsampling stage, a fourth downsampling stage, a key point regression branch and a multi-attribute classification head, wherein Patch Partition, the first downsampling stage, the second downsampling stage, the third downsampling stage, the fourth downsampling stage and the multi-attribute classification head are connected in series in sequence. The first downsampling stage includes a linear transformation layer Linear Embedding and at least two self-attention blocks Swin Transformer blocks. The second downsampling stage, the third downsampling stage and the fourth downsampling stage respectively include a downsampling layer Patch Merging and at least two self-attention blocks Swin Transformer blocks. The key point regression branch is used to downsample and adjust the dimension of the output of the first downsampling stage to obtain the adjusted output of the first downsampling stage, downsample and adjust the dimension of the output of the second downsampling stage to obtain the adjusted output of the second downsampling stage, and feature fuse the adjusted output of the first downsampling stage, the adjusted output of the second downsampling stage and the output of the third downsampling stage. The multi-attribute classification head includes LayerNorm, AdaptiveAvgPool1d, and multiple FC layers for multi-attribute classification.
[0078] It can be understood that Patch Partition is used to perform a block operation on the input (eg, the image to be recognized).
[0079] Exemplarily, the original image (of size H×W×3, where H is the height of the original image and W is the width of the original image) is input into Patch Partition for block operation. Patch Partition can divide the original image into multiple image blocks (for example, H / 4×W / 4 image blocks of size 4×4×3).
[0080] In some embodiments, the first downsampling stage includes a linear transformation layer LinearEmbedding and two self-attention blocks Swin Transformer blocks.
[0081] Exemplarily, the output of Patch Partition is multiple image blocks, and Linear Embedding is used to perform linear transformation on each of the multiple image blocks, mapping each of the multiple image blocks into an embedding vector of dimension C (also called the number of basic channels) (after the number of channels is expanded, the final feature map size is H / 4×W / 4×C). Then, after passing through two self-attention blocks Swin Transformer Block, the output remains unchanged. Each Swin Transformer Block includes LayerNorm, Multilayer Perceptron (MLP), Window-based Multi-head Self-Attention (W-MSA) and Shifted Window Multi-head Self-Attention (SW-MSA).
[0082] Compared with MSA, W-MSA only performs self-attention calculations within the window, which greatly reduces the time complexity. MSA is shown in formula (1), and W-MSA is shown in formula (2). The two formulas compare their time complexity.
[0083] Ω(MSA)=4hwc 2 +2(hw) 2 c (1)
[0084] Ω(W-MSA)=4hwc 2 +2Win 2 hwc (2)
[0085] In formula (1) and formula (2), h and w represent the height and width of the input feature, c represents the dimension of the image embedding, and Win represents the size of each window. When the W-MSA module is used, that is, in formula (2), self-attention calculation is only performed in each window, so information cannot be transferred between windows. The SW-MSA module shifts the window from the original 4 windows to 9 windows. Some windows become smaller. The blocks that are less than Win in size are pieced together to the size of WIin. Self-attention can be performed in windows of the same size without introducing extra calculations. In order to solve the spatial discontinuity of the pieced-together Win-sized window, a mask is introduced during the calculation. Through the self-attention calculation with the mask, the spatially discontinuous pixels in the pieced-together window are cut apart. Formula (3) is the formula for calculating self-attention in a moving window. Among them, K, K, and V represent the Qurey, Key, and Value matrices in the original self-attention calculation, respectively, and B represents the position encoding matrix.
[0086]
[0087] In some embodiments, the second downsampling stage and the fourth downsampling stage include a downsampling layer Patch Merging and two self-attention blocks Swin Transformer block respectively. The third downsampling stage includes a downsampling layer Patch Merging and 18 self-attention blocks Swin Transformer block. The multi-attribute classification head includes layer normalization (Layer Normalization, LayerNorm), adaptive one-dimensional average pooling AdaptiveAvgPool1d and multiple full connection (Full Connection, FC) layers, and the multi-attribute classification head is used to perform multi-attribute classification according to the output of the fourth downsampling stage.
[0088] Exemplarily, in the second downsampling stage, the patch merging layer is first used for downsampling, and the feature scale becomes H / 8×W / 8×4C, and then a fully connected layer is passed to adjust the channel dimension to 2C. It then passes through a 2 SwinTransformer Block, and the final output dimension is H / 8×W / 8×2C. The structure of the third stage is basically similar to that of the second stage, and the hierarchical design is also maintained. The difference is that 18 Swin Transformer Blocks are used, and the feature output dimension after calculation is H / 16×W / 16×4C. The fourth stage also first passes through the patch merging layer for downsampling, and then passes through 2 Swin Transformer Blocks, and the final feature output dimension is H / 32×W / 32×8C. The multi-attribute classification head includes LayerNorm, AdaptiveAvgPool1d, and multiple FC layers for multi-attribute classification.
[0089] In some embodiments, the key point regression branch is used to downsample and dimensionally adjust the output of the first downsampling stage to obtain an adjusted output of the first downsampling stage, downsample and dimensionally adjust the output of the second downsampling stage to obtain an adjusted output of the second downsampling stage, and perform feature fusion on the adjusted output of the first downsampling stage, the adjusted output of the second downsampling stage, and the output of the third downsampling stage.
[0090] For example, the output of the first downsampling stage is downsampled twice by convolution with a stride of 2 and average pooling, and then upscaled by 1×1 convolution. Then the output of the second downsampling stage is adjusted by average pooling and 1×1 convolution. Finally, the outputs of the first two and the third downsampling stage are feature fused through the concatenation layer (concat layer), and the key point regression branch is formed through LayerNorm, AdaptiveAvgPool1d, and FC.
[0091] Step 602: Create a model data set based on the clothing image text database.
[0092] Exemplarily, the data set includes a training set, a validation set, and a test set. The training set is used to train the model to help the model learn the laws and patterns of the data. The validation set is used to adjust the model parameters and evaluate the model performance to ensure that the model performs well under different conditions. The test set is used to finally evaluate the performance of the model to ensure the predictive ability of the model.
[0093] Step 603: Train the deep learning model built according to SwinTransformer according to the model data set to obtain a clothing attribute recognition model.
[0094] For example, Figure 7 As shown in the figure, in terms of network structure, the number of basic channels C is set to 96, the window size Win is set to 7, and the dimensions of the attention heads in the four downsampling stages (the first downsampling stage to the fourth downsampling stage) are set to 3, 6, 12, and 24 respectively.
[0095] In some embodiments, Figure 8 As shown, the implementation method of step 603 includes steps 801 to 802.
[0096] Step 801: Set the data enhancement method, optimizer and learning rate.
[0097] For example, the data enhancement methods may include random cropping and interpolation RandomResizedCropAndInterpolation, random horizontal flip RandomHorizontalFlip, random vertical flip RandomVerticalFlip, random erasing RandomErasing, color jitter ColorJitter, cut mix CutMix, mixUp, normalize, etc. The optimizer may use the optimizer AdamW, with an initial learning rate of 2e-05, and training with cosine annealing and warmup WarmUp.
[0098] Step 802: train the deep learning model built according to Swin Transformer according to the model data set, and fine-tune the training parameters of the deep learning model built according to Swin Transformer according to the pre-trained model to obtain a clothing attribute recognition model.
[0099] In step 802, the pre-trained model is obtained by training the visual model Swin-Transformer-Small using the dataset ImageNet-22K.
[0100] The loss function Loss during training is as shown in formula (6), including the key point regression branch loss function L landmark (Formula (4)) and multi-attribute recognition loss function L ce (Formula (5)).
[0101]
[0102]
[0103] Loss = λLlandmark +(1-λ)L ce (6)
[0104] In formula (4) to formula (6), n is the number of the current image sample, N is the total number of image samples, m is the number of the major attribute category of the current image sample, M is the total number of major attribute categories included in at least one major attribute category of the current image sample, l is the number of the sub-attribute category of the current image sample, L is the total number of sub-attribute categories in the major attribute category numbered m of the current image sample, k is the key point number of the current image sample, K is the total number of key points of the current image sample, and λ is the key point regression branch loss function L ce And the multi-attribute recognition loss function L landmark The weight parameter v k The visibility of the current key point, the value is 0 or 1; D n (x k ,y k ) is the predicted coordinate of the key point numbered k in the image sample numbered n, is the true coordinate of the key point numbered k in the image sample numbered n; represents the true value of the image sample with image number n, major attribute category number m, and sub-attribute category number l. Represents the probability value of the model output for the image sample with image number n, major attribute category number m, and sub-attribute category number l.
[0105] Step 103: Identify the attributes of the clothing in the image to be identified according to the clothing attribute recognition model to obtain attribute data of the clothing in the image to be identified.
[0106] It can be understood that the method of identifying the attributes of the clothing in the image to be identified based on the clothing attribute recognition model in step 103 can refer to the relevant description of obtaining the clothing attribute recognition model based on the clothing image text database and deep learning model in step 102, which will not be repeated here.
[0107] In summary, in terms of data, a clothing attribute recognition method provided by an embodiment of the present invention obtains clothing pictures of an e-commerce platform through crawler technology, removes duplicates based on the CLIP model, and trains an entity discrimination model based on Bert to identify named entities in the crawled text description. In addition, a subset of the open source dataset Deepfashion is extracted, and clothing duplication and named entity recognition are also performed to retain the key point attribute information of this part of the dataset in the Deepfashion dataset. A clothing attribute recognition method provided by an embodiment of the present invention also uses the large-scale visual language model Qwen-VL and uses the designed prompt words to perform image question and answer to supplement the missing attribute data. For text-generated images or image-generated images tasks, in order to reduce the interference caused by other objects or image backgrounds other than the main body of the clothes, the U2Net algorithm is used to cut out the clothes separately, which is convenient for subsequent training to generate a generation model with richer clothing details.
[0108] In terms of algorithms, an embodiment of the present invention provides a clothing attribute recognition method, which proposes a lightweight multi-attribute classification network based on a local sliding self-attention mechanism to perform clothing multi-attribute recognition, and designs the output head of the network to adapt to multi-attribute classification. Considering the strong correlation between the clothing key point regression task and the classification task, an auxiliary key point regression branch is added to improve the accuracy of the main task, and a new loss function is designed.
[0109] The following is a specific example to illustrate a clothing attribute recognition method provided by an embodiment of the present invention.
[0110] like Fig. 9 As shown, this example includes the following aspects.
[0111] 1. Dataset establishment
[0112] (1) In order to solve the problem of lack of professional data sets in the clothing industry, crawler technology can be used to crawl clothing pictures and corresponding text descriptions from the product homepages of mainstream e-commerce platforms at home and abroad.
[0113] (2) The image part encoder of the contrastive learning model CLIP is used to extract image embeddings, and the cosine similarity between image embeddings is calculated to remove duplicates, resulting in a 100,000+ e-commerce clothing dataset.
[0114] (3) Based on Bert training, named entity recognition (NER) is performed on the text descriptions corresponding to the crawled clothing images, and they are classified into 24 major attribute categories: color, sleeve length, waist type, sleeve type, type, trouser length, version, style, thickness, craftsmanship, collar type, function, popular elements, slit design, placket, trouser type, fabric, skirt type, skirt length, scene (suitable scene), crowd (suitable crowd), clothing length, pattern and season (suitable season). Each major attribute category has different sub-attribute categories (such as Figure 5 As shown, only three sub-attribute categories are listed for each attribute for illustration purposes).
[0115] On the other hand, more than 100,000 full-body clothing images are extracted from the public Deepfashion dataset, and the clothing images are also deduplicated and the attributes (labels) are classified by named entity recognition, and the key point attribute information of this part of the Deepfashion dataset is retained.
[0116] (4) We use Alibaba's open-source large-scale visual language model Qwen-VL for image question answering to obtain richer attribute data and expand the attribute data in the pre-processing stage. We ultimately obtain a total of more than 200,000 clothing image-text pair datasets, where each clothing image can have up to 24 clothing attributes.
[0117] (5) For text-generated images or image-generated images tasks, in order to reduce the interference caused by other objects or image backgrounds other than the main body of the clothes, the clothing Matting model trained based on the U2Net algorithm is used to cut out the clothes separately, which facilitates the subsequent training of a generative model with richer clothing details.
[0118] 2. Design of network structure
[0119] In order to avoid the problem of traditional CNN's ability degradation when processing large-scale data sets and the large amount of global attention calculation of VIT, the present invention draws on the excellent prior knowledge of visual signals, adopts the characteristics of CNN's locality and translation invariance, and performs multi-head self-attention (SW-MSA, shifted window-multi head self-attention) calculation in the moving window, which effectively reduces the amount of calculation. The entire clothing attribute recognition model adopts a hierarchical design, which includes a total of 4 downsampling stages. Each stage will reduce the resolution of the input feature map and expand the receptive field layer by layer. In addition. Considering the close relationship between the clothing joint point regression task and the multi-attribute recognition task, the key point regression branch is introduced at the end of the third stage, and joint training is used to assist in improving the accuracy of the main branch. At the end of the fourth stage, the multi-attribute classification head is connected, and the overall network structure is as follows. Figure 7 shown.
[0120] Specifically, first input the original image of size H×W×3 into PatchPartition for block operation, divide the image into H / 4×W / 4 small blocks of size 4×4×3, and then send each small block to the LinearEmbedding module to map it into an embedding vector of dimension C, where dimension C can also be called the basic channel number C. After expanding the number of channels, the final feature map size is H / 4×W / 4×C. Then after passing through 2 self-attention blocks, the output remains unchanged. The self-attention block can also be called the self-attention block SwinTransformerBlock. Each self-attention block is mainly composed of LayerNorm, MLP, W-MSA and SW-MSA.
[0121] Compared with MSA, W-MSA only performs self-attention calculation within the window, which greatly reduces the time complexity. Formula (1) and (2) compare the time complexity of the two. Please refer to the relevant description of formula (1) and formula (2), which will not be repeated here.
[0122] When using the W-MSA module, self-attention calculations are only performed within each window, so information cannot be transferred between windows. The SW-MSA module offsets the windows, changing the original 4 windows to 9 windows. Some windows become smaller, and blocks that are less than Win in size are pieced together to the size of Win. Self-Attention can be performed within windows of the same size without introducing extra calculations. In order to solve the spatial discontinuity of the pieced-together Win-sized windows, a mask is introduced during the calculation, and the spatially discontinuous pixels in the pieced-together windows are cut apart through self-attention calculations with masks.
[0123] Formula (3) is the calculation formula for self-attention in the moving window. Please refer to the relevant description of formula (3) and will not be repeated here.
[0124] Then enter the second stage, first pass through the patch merging module for downsampling, the feature scale becomes H / 8×W / 8×4C, then pass through a fully connected layer to adjust the channel dimension to 2C. Then pass through a 2-self-attention block, the final output dimension is H / 8×W / 8×2C.
[0125] The structure of the third stage is basically similar to that of the second stage, and the hierarchical design is also maintained. The difference is that 18 self-attention blocks are used, and the feature output dimension after calculation is H / 16×W / 16×4C. In addition, a key point regression branch is introduced. Specifically, the output of the first stage is downsampled twice by convolution with a stride of 2 and average pooling, and then upsampled by 1×1 convolution. Then the output of the second stage is adjusted by average pooling and 1×1 convolution. Finally, the outputs of the first two and the third stage are fused through the concat layer, and the key point regression branch is formed through LayerNorm, AdaptiveAvgPool1d, and FC.
[0126] In the fourth stage, the image is downsampled by the patch merging layer and then passes through two self-attention blocks. The final feature output dimension is H / 32×W / 32×8C. Then in the fifth stage, the classification head composed of LayerNorm, AdaptiveAvgPool1d and multiple FC layers performs multi-attribute classification.
[0127] 3. Overall algorithm flow chart
[0128] After completing steps 1 and 2 above, complete model training, testing, and deployment. Fig. 9 The overall flow chart of the algorithm is shown.
[0129] 4. Specific implementation process of the algorithm
[0130] In terms of network structure, the number of basic channels C is set to 96, the window size Win is set to 7, and the dimensions of the attention heads in the four stages are set to 3, 6, 12, and 24 respectively. During training, data enhancement methods such as RandomResizedCropAndInterpolation, RandomHorizontalFlip, RandomVerticalFlip, RandomErasing, ColorJitter, CutMix, MixUp, and Normalize are used with certain probabilities. The optimizer uses AdamW, with an initial learning rate of 2e-05, combined with cosine annealing and WarmUp preheating training. Fine-tuning is performed based on the pre-trained model of Swin-Transformer-Small on ImageNet-22K. The overall loss function (Formula (6)) includes the key point regression branch loss (Formula (4)) and the multi-attribute recognition loss (Formula (5)).
[0131] Experiments show that adding a keypoint regression branch for joint training leads to a 2.9% improvement in accuracy.
[0132] During deployment, the test image is scaled geometrically with a minimum side of 256, and then a 224×224 area is taken from the center of the scaled image and normalized. The model is converted to tensorrt format for inference, the model service is deployed based on triton-server, and the interface call service is deployed based on the flask framework.
[0133] An embodiment of the present invention provides a clothing attribute recognition system. Fig.10 As shown, the clothing attribute recognition system 1000 includes a to-be-recognized picture acquisition module 1001 , a data and model acquisition module 1002 and an attribute recognition module 1003 .
[0134] The to-be-recognized picture acquisition module 1001 is configured to: acquire the to-be-recognized picture.
[0135] The data and model acquisition module 1002 is configured to: acquire a clothing image text database, and acquire a clothing attribute recognition model based on the clothing image text database and the deep learning model. The image text database includes multiple clothing images and text information corresponding to each clothing image in the multiple clothing images.
[0136] In some embodiments, Fig.11 As shown, the data and model acquisition module 1002 includes a data acquisition unit 1021. The data acquisition unit 1021 is configured to: acquire multiple clothing pictures. Acquire text information corresponding to each clothing picture in the multiple clothing pictures. And, acquire a picture text database of clothing according to the multiple clothing pictures and the text information corresponding to each clothing picture in the multiple clothing pictures. The text information corresponding to each clothing picture in the multiple clothing pictures is related to the attributes of the clothing in each clothing picture.
[0137] In some embodiments, the data acquisition unit 1021 is configured to: acquire at least one clothing picture and text information corresponding to each clothing picture in at least one clothing picture acquired from at least one e-commerce platform according to crawler technology. And, acquire at least one clothing picture and text information corresponding to each clothing picture in at least one clothing picture in the public database from a public database. Among them, the multiple clothing pictures include at least one clothing picture acquired from at least one e-commerce platform and at least one clothing picture acquired from a public database. The text information corresponding to each clothing picture in the multiple clothing pictures includes text information corresponding to each clothing picture in at least one clothing picture acquired from at least one e-commerce platform and text information corresponding to each clothing picture in at least one clothing picture in the public database.
[0138] In some embodiments, the data acquisition unit 1021 is configured to: perform feature extraction on each clothing picture in the plurality of clothing pictures using an image encoder in a contrastive learning algorithm to obtain a plurality of feature-extracted data. De-duplicate the plurality of feature-extracted data to obtain de-duplicate feature-extracted data, and obtain a clothing picture database based on the clothing pictures corresponding to the de-duplicate feature-extracted data. Also, perform named entity recognition and classification on the text information corresponding to each clothing picture in the plurality of clothing pictures to obtain a clothing attribute database in each clothing picture in the plurality of clothing pictures. The clothing picture text database includes a clothing picture database and a clothing attribute database.
[0139] In some embodiments, the data acquisition unit 1021 is configured to: perform image question answering using an open source large-scale visual language model based on multiple clothing pictures and text information corresponding to each of the multiple clothing pictures, expand the attribute data of the clothing in each of the multiple clothing pictures, and obtain an expanded clothing attribute database, wherein the clothing image text database includes a clothing image database and an expanded clothing attribute database.
[0140] In some embodiments, the attributes of clothing include multiple major attribute categories, each of the multiple major attribute categories includes at least one sub-attribute category, and the multiple major attribute categories include color, sleeve length, waist type, sleeve type, type, trouser length, version, style, thickness, craftsmanship, collar type, function, popular elements, slit design, placket, trouser type, fabric, skirt type, skirt length, scene, crowd, length, pattern and season.
[0141] In some embodiments, Fig.12 As shown, the data and model acquisition module 1002 includes a model acquisition unit 1022. The model acquisition unit 1022 is configured to: build a deep learning model according to the block attention mechanism Swin Transformer. Create a model data set according to the clothing image text database. And, train the deep learning model built according to the Swin Transformer according to the model data set to obtain a clothing attribute recognition model.
[0142] In some embodiments, the deep learning model built according to Swin Transformer includes: a block layer PatchPartition, a first downsampling stage, a second downsampling stage, a third downsampling stage, a fourth downsampling stage, a key point regression branch and a multi-attribute classification head, wherein Patch Partition, the first downsampling stage, the second downsampling stage, the third downsampling stage, the fourth downsampling stage and the multi-attribute classification head are connected in series in sequence. The first downsampling stage includes a linear transformation layer Linear Embedding and at least two self-attention blocks Swin Transformerblock. The second downsampling stage, the third downsampling stage and the fourth downsampling stage respectively include a downsampling layer Patch Merging and at least two self-attention blocks Swin Transformerblock. The key point regression branch is used to downsample and dimensionally adjust the output of the first downsampling stage of Swin Transformer to obtain the adjusted output of the first downsampling stage, downsample and dimensionally adjust the output of the second downsampling stage of Swin Transformer to obtain the adjusted output of the second downsampling stage, and feature fuse the adjusted output of the first downsampling stage, the adjusted output of the second downsampling stage and the output of the third downsampling stage of Swin Transformer.
[0143] In some embodiments, the first downsampling stage includes a linear transformation layer LinearEmbedding and two self-attention blocks Swin Transformer blocks. The second downsampling stage and the fourth downsampling stage each include a downsampling layer Patch Merging and two self-attention blocks Swin Transformer blocks. The third downsampling stage includes a downsampling layer PatchMerging and 18 self-attention blocks Swin Transformer blocks.
[0144] In some embodiments, the model acquisition unit 1022 is configured to: set a data enhancement method, an optimizer, and a learning rate. Also, train a deep learning model built according to Swin Transformer according to the model data set, and fine-tune the training parameters of the deep learning model built according to Swin Transformer according to the pre-trained model to obtain a clothing attribute recognition model. The pre-trained model is obtained by training the visual model Swin-Transformer-Small using the data set ImageNet-22K.
[0145] The loss function during training is Loss = λL landmark+(1-λ)L ce .L landmark is the key point regression branch loss function, and L ce is the multi-attribute recognition loss function, and n is the number of the current image sample, N is the total number of image samples, m is the number of the major attribute category of the current image sample, M is the total number of at least one major attribute category of the current image sample, l is the number of the sub-attribute category of the current image sample, L is the total number of sub-attribute categories in the major attribute category numbered m of the current image sample, k is the key point number of the current image sample, K is the total number of key points of the current image sample, and λ is the key point regression branch loss function L ce And the multi-attribute recognition loss function L landmark The weight parameter v k The visibility of the current key point, the value is 0 or 1; D n (x k ,y k ) is the predicted coordinate of the key point numbered k in the image sample numbered n, is the true coordinate of the key point numbered k in the image sample numbered n; represents the true value of the image sample with image number n, major attribute category number m, and sub-attribute category number l. Represents the probability value of the model output for the image sample with image number n, major attribute category number m, and sub-attribute category number l.
[0146] The attribute recognition module 1003 is configured to: recognize the attributes of the clothing in the to-be-recognized picture according to the clothing attribute recognition model, and obtain attribute data of the clothing in the to-be-recognized picture.
[0147] The specific scheme and beneficial effects of a clothing attribute recognition system provided by an embodiment of the present invention can refer to the relevant description of a clothing attribute recognition method provided by an embodiment of the present invention, which will not be repeated here.
[0148] An embodiment of the present invention provides a computer device, such as Fig.13 As shown, the computer device 1300 includes a memory 1301 and a processor 1302. The memory 1301 stores a computer program. When the processor 1302 runs the computer program stored in the memory 1301, the processor 1302 executes the above-mentioned clothing attribute recognition method.
[0149] The specific scheme and beneficial effects of a computer device provided by an embodiment of the present invention can be referred to the relevant description of a clothing attribute recognition method provided by an embodiment of the present invention, which will not be repeated here.
[0150] An embodiment of the present invention provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the processor executes the above-mentioned clothing attribute recognition method.
[0151] The specific scheme and beneficial effects of a computer-readable storage medium provided by an embodiment of the present invention can be referred to the relevant description of a clothing attribute recognition method provided by an embodiment of the present invention, which will not be repeated here.
[0152] It is to be understood that the above embodiments are merely exemplary embodiments used to illustrate the principles of the present invention, but the present invention is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also considered to be within the scope of protection of the present invention.
Claims
1. A clothing attribute recognition method, characterized in that: include: Get the image to be identified; Acquire a clothing image text database, and acquire a clothing attribute recognition model according to the clothing image text database and the deep learning model; the image text database includes a plurality of clothing images and text information corresponding to each clothing image in the plurality of clothing images; as well as The attributes of the clothing in the to-be-recognized picture are recognized according to the clothing attribute recognition model to obtain attribute data of the clothing in the to-be-recognized picture.
2. The clothing attribute recognition method according to claim 1, characterized in that: The step of obtaining a clothing image text database includes: Get multiple clothing pictures; Acquire text information corresponding to each clothing picture in a plurality of clothing pictures; the text information corresponding to each clothing picture in the plurality of clothing pictures is related to an attribute of clothing in each clothing picture; and A clothing picture text database is obtained according to a plurality of clothing pictures and text information corresponding to each clothing picture in the plurality of clothing pictures.
3. The clothing attribute recognition method according to claim 2, characterized in that: The obtaining of the plurality of clothing pictures and the text information corresponding to each of the plurality of clothing pictures comprises: Acquire at least one clothing picture and text information corresponding to each clothing picture in the at least one clothing picture acquired from the at least one e-commerce platform according to crawler technology; and Obtaining at least one clothing image and text information corresponding to each clothing image in the at least one clothing image in the public database from a public database; Among them, the multiple clothing pictures include at least one clothing picture obtained from at least one e-commerce platform and at least one clothing picture obtained from a public database; the text information corresponding to each clothing picture in the multiple clothing pictures includes text information corresponding to each clothing picture in the at least one clothing picture obtained from at least one e-commerce platform and text information corresponding to each clothing picture in the at least one clothing picture in the public database.
4. The clothing attribute recognition method according to claim 2, characterized in that: The method of acquiring a clothing picture text database according to a plurality of clothing pictures and text information corresponding to each clothing picture in the plurality of clothing pictures comprises: An image encoder in a contrastive learning algorithm is used to extract features of each clothing picture in a plurality of clothing pictures to obtain a plurality of feature-extracted data; Deduplication is performed on the plurality of feature-extracted data to obtain deduplicated feature-extracted data, and a clothing image database is obtained according to clothing images corresponding to the deduplicated feature-extracted data; and Named entity recognition and classification are performed on text information corresponding to each clothing picture in a plurality of clothing pictures to obtain a clothing attribute database in each clothing picture in the plurality of clothing pictures; the clothing picture text database includes the clothing picture database and the clothing attribute database.
5. The clothing attribute recognition method according to claim 4, characterized in that: The method of acquiring a clothing picture text database according to a plurality of clothing pictures and text information corresponding to each clothing picture in the plurality of clothing pictures further includes: An open source large-scale visual language model is used to perform image question answering based on the multiple clothing pictures and text information corresponding to each clothing picture in the multiple clothing pictures, and attribute data of clothing in each clothing picture in the multiple clothing pictures is expanded to obtain an expanded clothing attribute database, wherein the clothing image text database includes the clothing image database and the expanded clothing attribute database.
6. The clothing attribute recognition method according to claim 4 or 5, characterized in that: The attributes of clothing include multiple major attribute categories, each of the multiple major attribute categories includes at least one sub-attribute category, and the multiple major attribute categories include color, sleeve length, waist type, sleeve type, type, trouser length, version, style, thickness, craftsmanship, collar type, function, popular elements, slit design, placket, trouser type, fabric, skirt type, skirt length, scene, crowd, length, pattern and season.
7. The clothing attribute recognition method according to claim 1, characterized in that: The obtaining of a clothing attribute recognition model according to the clothing image text database and the deep learning model comprises: Build a deep learning model based on the block attention mechanism Swin Transformer; Creating a model data set based on the clothing image text database; and The deep learning model built according to Swin Transformer is trained according to the model data set to obtain a clothing attribute recognition model.
8. The clothing attribute recognition method according to claim 7, characterized in that: The deep learning model built according to Swin Transformer includes: a block layer Patch Partition, a first downsampling stage, a second downsampling stage, a third downsampling stage, a fourth downsampling stage, a key point regression branch and a multi-attribute classification head, wherein the Patch Partition, the first downsampling stage, the second downsampling stage, the third downsampling stage, the fourth downsampling stage and the multi-attribute classification head are connected in series in sequence; the first downsampling stage includes a linear transformation layer LinearEmbedding and at least two self-attention blocks Swin Transformer block; the second downsampling stage, the third downsampling stage and the fourth downsampling stage respectively include a downsampling layer Patch Merging and at least two self-attention blocks Swin Transformer block; the multi-attribute classification head includes layer normalization LayerNorm, adaptive one-dimensional average pooling AdaptiveAvgPool1d and multiple fully connected FC layers, and the multi-attribute classification head is used to perform multi-attribute classification according to the output of the fourth downsampling stage; the key point regression branch is used to downsample and dimensionally adjust the output of the first downsampling stage to obtain the adjusted output of the first downsampling stage, downsample and dimensionally adjust the output of the second downsampling stage to obtain the adjusted output of the second downsampling stage, and perform feature fusion on the adjusted output of the first downsampling stage, the adjusted output of the second downsampling stage and the output of the third downsampling stage.
9. The clothing attribute recognition method according to claim 8, characterized in that: The first downsampling stage includes a linear transformation layer Linear Embedding and two self-attention blocks Swin Transformer block; The second downsampling stage and the fourth downsampling stage respectively include a downsampling layer Patch Merging and two self-attention blocks SwinTransformer blocks; The third downsampling stage includes a downsampling layer Patch Merging and 18 self-attention blocks Swin Transformer blocks.
10. The clothing attribute recognition method according to claim 7, characterized in that: The deep learning model built according to the Swin Transformer is trained according to the model data set to obtain a clothing attribute recognition model, including: Setting data augmentation methods, optimizers, and learning rates; and The deep learning model built according to Swin Transformer is trained according to the model data set, and the training parameters of the deep learning model built according to Swin Transformer are fine-tuned according to the pre-trained model to obtain a clothing attribute recognition model; the pre-trained model is obtained by training the visual model Swin-Transformer-Small using the data set ImageNet-22K; the loss function Loss during training is: Loss=λL landmark +(1-λ)L ce Among them, L landmark is the key point regression branch loss function, and L ce is the multi-attribute recognition loss function, and n is the number of the current image sample, N is the total number of image samples, m is the number of the major attribute category of the current image sample, M is the total number of major attribute categories included in at least one major attribute category of the current image sample, l is the number of the sub-attribute category of the current image sample, L is the total number of sub-attribute categories in the major attribute category numbered m of the current image sample, k is the key point number of the current image sample, K is the total number of key points of the current image sample, and λ is the key point regression branch loss function L ce And the multi-attribute recognition loss function L landmark The weight parameter v k The visibility of the current key point, the value is 0 or 1; D n (x k ,y k ) is the predicted coordinate of the key point numbered k in the image sample numbered n, is the true coordinate of the key point numbered k in the image sample numbered n; represents the true value of the image sample with image number n, major attribute category number m, and sub-attribute category number l. Represents the probability value of the model output for the image sample with image number n, major attribute category number m, and sub-attribute category number l.
11. A clothing attribute recognition system, characterized in that: include: The module for obtaining the image to be identified is configured to: obtain the image to be identified; A data and model acquisition module, which is configured to: acquire a clothing image text database, and acquire a clothing attribute recognition model according to the clothing image text database and a deep learning model; the image text database includes a plurality of clothing images and text information corresponding to each of the plurality of clothing images; and The attribute recognition module is configured to: recognize the attributes of the clothing in the to-be-recognized picture according to the clothing attribute recognition model, and obtain attribute data of the clothing in the to-be-recognized picture.
12. A computer device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, the processor executes the clothing attribute recognition method according to any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor performs the clothing attribute recognition method according to any one of claims 1 to 10.
Citation Information
Cited By
Clothing model plane processing method based on material electrical characteristic and vision fusion
CN121500873A
Clothing attribute identification method and system based on multi-modal information hierarchy semantic modeling
CN121708402A