Clothing attribute identification method based on attribute semantics and spatial position information fusion
By using the semantic and spatial position information of the fusion attributes of lightweight MobileViT network and graph neural network in clothing image recognition, the problem of ignoring co-occurrence and spatial position relationship in existing methods is solved, and efficient clothing attribute recognition is achieved.
Patent Information
- Application Number
- CN202510516610.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-22
AI Technical Summary
Existing clothing image recognition methods ignore the co-occurrence and interdependence between clothing attributes and spatial position relationships, resulting in low recognition accuracy and long-term time.
The lightweight MobileViT network is used to extract global spatial features, combine specific attention mechanisms and spatial position information of the learning attributes of graph neural networks, and merge the semantic relationships and spatial position relationships of attributes through the multi-headed attention mechanism to directly locate key areas for identification.
It significantly improves the accuracy of clothing attribute recognition, reduces model complexity, and improves recognition efficiency.
Smart Images

Figure CN120356007A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of clothing image recognition, and particularly relates to a clothing attribute recognition method based on the fusion of attribute semantics and spatial position information. Background Art
[0002] Clothing attribute recognition is an important task in the field of computer vision and has been widely applied in multiple fields such as e-commerce recommendation, intelligent try-on, and fashion analysis. For example, when people purchase clothing on an e-commerce platform, they search for the target clothing by entering keywords of clothing tags, which requires the platform to accurately record and identify various attributes of the clothing, such as type, color, style, and pattern. Therefore, how to quickly and accurately identify a large number of clothing attributes has become a research hotspot in the field of clothing image recognition.
[0003] Traditional clothing image recognition methods mainly consist of three steps. The first step is to preprocess the clothing image, segment the image, and remove the complex background. The second step is to extract artificial features from the clothing image, including texture features, color features, and contour features. The third step is to process the extracted features and combine classification algorithms to achieve clothing attribute classification. However, using the method of manually extracting features for clothing attribute recognition has high requirements for clothing images, and clothing is a flexible object that is easily distorted in shape, resulting in the inability to extract features. Secondly, using manual feature extraction is time-consuming and cannot be applied to the recognition of a large number of clothing images.
[0004] In recent years, with the development of artificial intelligence, deep learning has gradually been applied to the field of clothing attribute recognition. Researchers have begun to use convolutional neural networks to extract clothing features and continuously improve the network model to enhance the ability to capture spatial features of clothing images, which has improved the accuracy of clothing attribute recognition to a certain extent. However, these improved methods mainly rely on detecting the positions of clothing key points and then extracting local features around the key points for recognition, ignoring the natural co-occurrence and interdependence between clothing attributes and the spatial position relationship existing between clothing attributes. For example, "shirt" usually coexists with "V-neck", and "V-neck" is usually located on the upper body. Summary of the Invention
[0005] Aiming at the problem that the existing methods ignore the co-occurrence and interdependence between clothing attributes and the spatial position relationship existing between clothing attributes, the present invention provides a clothing attribute recognition method based on the fusion of attribute semantics and spatial position information. The aim is to directly locate the key areas in the attribute recognition process by using the semantic association and spatial position relationship between attributes, avoiding the additional step of separately extracting the key areas of clothing in the traditional method, reducing the complexity of the model, and at the same time improving the accuracy of clothing attribute recognition.
[0006] The clothing attribute recognition method based on the fusion of attribute semantics and spatial location information is divided into four parts as a whole. The first part is attribute feature learning. In order to accurately extract the spatial features of each attribute in the image, a lightweight MobileViT is used as the backbone network to extract global spatial features, and then a specific attention mechanism is introduced to learn the spatial embedding representation of each attribute. The second part is the learning of attribute spatial position relationships. In order to obtain the spatial position information of each attribute in the image, an adjacency matrix containing attribute spatial direction information is constructed, and a graph neural network is used to model the spatial position relationships between attributes. The third part is the learning of attribute semantic relationships. In order to obtain the mutual relationships between attribute labels, a novel graph embedding construction method is proposed to guide the generation of attribute label embeddings. The fourth part is the fusion of spatial position and semantic relationships. In order to fully fuse the spatial position relationships and semantic relationships of attributes, a multi-head attention mechanism is adopted, using the attribute label embeddings as queries to guide the learning of the spatial position features of attributes, thereby realizing the deep fusion of attribute spatial information and semantic information. Finally, the output result is passed through a classification layer to predict the prediction probability of each clothing attribute.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] A clothing attribute recognition method based on the fusion of attribute semantics and spatial location information, comprising the following steps:
[0009] Step 1: Collect the original clothing images and perform preprocessing. The preprocessing operations mainly include: randomly cropping the pictures, horizontally flipping the pictures, centrally cropping the pictures, and picture normalization operations;
[0010] Step 2: Use a lightweight MobileViT as the backbone network to extract features from the original images and obtain rich global spatial features;
[0011] The specific process of step 2 is as follows:
[0012] Step 2.1: Remove the last 1×1 convolutional layer, pooling layer, and fully connected layer from the MobileViT network, and retain the remaining network structure as the backbone network to extract the global spatial features of the image.
[0013] Step 3: After extracting the global spatial features, introduce a specific attention mechanism to further extract fine-grained features and obtain the spatial feature representation of each attribute;
[0014] The specific process of step 3 is as follows:
[0015] The structure of the attention mechanism is three convolutional layers. The Relu function is executed after the first two convolutional layers, and the Softmax function is executed after the last convolutional layer;
[0016] Step 3.2: The input of the attention mechanism is the global spatial features extracted by the MobileViT network, and the number of output channels is equal to the number of attribute categories. Each channel corresponds to an unnormalized feature map of an attribute; the output unnormalized feature map is normalized in the channel dimension through the Softmax function to obtain the normalized feature map of each attribute. The formula is:
[0017]
[0018] where h represents the height of the feature map, w represents the width of the feature map, C represents the number of attribute categories, c' represents the current channel, and I a represents the unnormalized feature map.
[0019] Step 4: Based on the spatial feature representation of each attribute, construct an adjacency matrix containing the spatial position relationship of the relative directions of the attributes, and use it as the prior knowledge of the structure of the graph convolutional network to obtain high-level features containing the spatial position relationship of the attributes;
[0020] The specific process of the said Step 4 is:
[0021] Step 4.1: Use the graph convolutional network (GCN) to model the spatial position relationship between multi-label attributes. The propagation method between each graph convolutional network layer is:
[0022] H l+1 = F(AH l W l )(2)
[0023] where H l is the input feature of the l-th layer, H l+1 is the output feature of the (l + 1)-th layer, A is the adjacency matrix, W l is the learnable parameter matrix, and F(·) is the activation function; corresponding to the present invention, the input feature of the first layer is the output of Step 3.2, that is, the normalized feature map of each attribute, h(·) is the LeakyRelu activation function, and the adjacency matrix A is constructed according to the following steps;
[0024] Step 4.2: Perform edge detection using the 3×3 Sobel operator on each channel dimension of the feature map output in Step 3.2, and set the padding value to 0 to ensure that the size of the output feature map remains unchanged; the Sobel operator is:
[0025]
[0026] where K x is the convolutional layer in the horizontal direction, and K y is the convolutional layer in the vertical direction;
[0027] Step 4.3: Calculate the final gradient magnitude and perform min-max normalization:
[0028]
[0029]
[0030] where G x and G y are the gradient magnitudes in the horizontal and vertical directions respectively, and ρ(G) is the edge probability value of each attribute feature map. The larger the value, the more likely the position belongs to the edge region during the recognition process;
[0031] Step 4.4: Select the position with the maximum edge probability value in each attribute feature map as the spatial position of the attribute; subsequently, project this position back to the original image coordinate system through the spatial mapping relationship to determine the specific position of the attribute in the original image; finally, construct an adjacency matrix using the directional spatial relationship between the attribute position points;
[0032] Step 4.5: Let the positions of attributes i and j mapped to the original image be (x i , y i ) and (x j , y j ), and calculate the relative direction angle between attributes i and j:
[0033]
[0034]
[0035] where arctan2(x, y) is the extended arctangent function, and the calculated angle range is [-180°, 180°]. Normalize it to [0°, 360°] to obtain the final relative direction angle θ' ij ;
[0036] Step 4.6: Consider the original image as a 360-degree spatial domain and divide it into 8 sectors, each sector being 45 degrees. Each sector represents a spatial relationship and is numbered from 1 to 8 in clockwise order;
[0037] Step 4.7: Determine the sector to which it belongs according to the calculated relative direction angle θ' ij , and use the number of this sector as the value of the corresponding position in the adjacency matrix. By analogy, finally construct an adjacency matrix containing spatial position information; use this adjacency matrix as the structural prior knowledge of the graph convolutional network, and use the graph neural network model to learn the spatial direction position relationship between attributes.
[0038] Step 5: Use the graph embedding method to model the semantic relationships between attributes, construct a new adjacency matrix, and use it to guide the generation of attribute label embeddings. The generated attribute label embeddings are input into the network as additional information, thereby introducing the semantic associations between attributes;
[0039] The specific process of Step 5 is as follows:
[0040] Step 5.1: Define the graph G=(V, E), where V=(v1, v2, …, v C ) represents the set of C attribute category nodes, and E represents the edges; specifically, V is the set of all attribute labels, and E is the connection set between any two attributes; the adjacency matrix A of graph G a contains non-negative weights associated with each edge;
[0041] Step 5.2: Initialize four matrices, namely the M matrix, the P matrix, the W matrix, and the A a matrix, and set their initial values to zero. Among them, the M matrix is the co-occurrence matrix, which is used to count the number of times any two labels appear simultaneously; the P matrix is the conditional probability matrix, which is calculated based on the M matrix and represents the conditional probability of label co-occurrence; the W matrix is the label weight matrix, which is used to count the number of times each label appears and perform normalization processing on it; the A a matrix is the adjacency matrix, which is constructed based on the above three matrices and is used to describe the relationships between attribute labels. The update formula of the A a matrix is as follows:
[0042] A a [i, j]=A a [i, j]+W[j, j]*P[i, j] (8)
[0043] Step 5.3: Adopt a model of a three-layer fully connected network, which is used to learn to map the one-hot encoding of each attribute to the semantic embedding space and generate the corresponding label embeddings. Among them, batch normalization and the Relu activation function are introduced after the first two fully connected layers;
[0044] Step 5.4: Train and optimize the loss function in the semantic embedding space. The optimization goal is to make the cosine similarity cos(e i , e j ) between any two edges close to the relationship described by the adjacency matrix A of graph G a . The formula of the loss function used is as follows:
[0045]
[0046] where C is the number of attribute categories, and cos(e i , e j) is the cosine similarity between any two edges in the attribute embedding space, and A a is the adjacency matrix of graph G.
[0047] Step 6: Take the outputs of Step 4 and Step 5 as input features, and perform feature fusion through the multi-head attention mechanism, where the attribute label embedding in Step 5 is used as the query, and the output spatial position features in Step 4 are used as the key and value;
[0048] The specific process of Step 6 is as follows:
[0049] Step 6.1: Use the attribute label embedding output in Step 5 as the query of the multi-head attention mechanism, and the output spatial position features in Step 4 as the key and value of the multi-head attention mechanism. The multi-head attention mechanism is used to fuse the semantic relationship and spatial position relationship of the attributes. The calculation process of the multi-head attention mechanism:
[0050] MultiHead(Q,K,V)=Conact(head1,head2,…,head n )W O (10)
[0051] where Q represents the query, K represents the key, V represents the value, n represents the number of attention heads, and W O represents the trainable parameter matrix;
[0052] Step 6.2: The calculation formula for each attention head is as follows:
[0053]
[0054] where, is the trainable parameter matrix;
[0055] The attention score is calculated by scaled dot-product attention:
[0056]
[0057] where d k represents the scaling factor; T represents the transpose.
[0058] Step 7: Perform a linear transformation on the output of Step 6 through a fully connected layer, and calculate the prediction probability of each attribute through the Sigmoid activation function;
[0059] The specific process of Step 7 is as follows:
[0060] Use binary cross-entropy loss (BCE) as the overall loss function for training. The output result of the multi-head attention mechanism is linearly transformed through a fully connected layer, and the prediction probability of the attribute label is calculated. The formula for the overall training loss function is:
[0061]
[0062] Among them, y ic represents the true attribute value; p ic represents the predicted attribute probability value.
[0063] Compared with the prior art, the present invention has the following advantages:
[0064] The clothing attribute recognition method proposed by the present invention uses a MobileViT backbone network combined with a new attention mechanism to extract attribute space features, and uses a graph neural network to extract attribute space direction and position features on this feature region. At the same time, a graph embedding method is used to extract attribute semantic features. By inputting the attribute space position information and attribute semantic information into the multi-head attention mechanism for classification, the attribute features are greatly enriched, and the accuracy of clothing attribute recognition is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 is the overall framework diagram of the embodiment of the present invention;
[0066] Figure 2 is the module function structure diagram of the embodiment of the present invention;
[0067] Figure 3 is the flowchart of the module for extracting the attribute space position relationship in the embodiment of the present invention;
[0068] Figure 4 is the algorithm pseudocode diagram for constructing the graph embedding adjacency matrix in the embodiment of the present invention;
[0069] Figure 5 is the bar chart of the experimental effect for constructing the graph embedding adjacency matrix in the embodiment of the present invention;
[0070] Figure 6 is the image retrieval experiment diagram of the MS-COCO dataset in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] To understand the present invention in depth, we will describe it comprehensively and in detail. However, the present invention has various implementation manners and is not limited to the specific examples listed herein. The presentation of these examples is intended to deepen the comprehensive understanding of the disclosed content of the present invention.
[0072] As Figures 1-4 shown, the embodiment of the present invention provides a clothing attribute recognition method based on the fusion of attribute semantics and spatial position information. It is specifically implemented according to the following steps:
[0073] Step 1: Collect the original clothing images and perform preprocessing. The preprocessing operations mainly include: randomly cropping the pictures, horizontally flipping the pictures, centrally cropping the pictures, and normalizing the pictures;
[0074] Download the existing clothing dataset as the original images, and divide the dataset into a training set, a validation set, and a test set. Different data augmentation operations are performed on the training set and the validation set to prevent overfitting. Specifically, all images are cropped to a size of 256×256. For the training set, random cropping of the images to 256, horizontal flipping of the images, and image normalization operations are carried out. For the validation set, central cropping of the images to 256 and normalization operations are performed.
[0075] Step 2: Use the lightweight MobileViT as the backbone network to extract features from the original images and obtain rich global spatial features;
[0076] The specific process of Step 2 is as follows:
[0077] Step 2.1: Remove the last 1×1 convolutional layer, pooling layer, and fully connected layer from the MobileViT network, and retain the remaining network structure as the backbone network to extract the global spatial features of the images;
[0078] During the training process, freeze the model weights of the MobileViT backbone network structure and only update the weights of the newly added layers;
[0079] Step 3: After extracting the global spatial features, introduce a specific attention mechanism to further extract fine-grained features and obtain the spatial feature representation of each attribute;
[0080] The specific process of Step 3 is as follows:
[0081] The structure of the attention mechanism is three convolutional layers. The Relu function is executed after the first two convolutional layers, and the Softmax function is executed after the last convolutional layer;
[0082] The input of the attention mechanism is the global spatial features extracted by the MobileViT network. The number of output channels is equal to the number of attribute categories, and each channel corresponds to an unnormalized feature map of an attribute; the output unnormalized feature maps are normalized in the channel dimension through the Softmax function to obtain the normalized feature maps of each attribute. The formula is:
[0083]
[0084] where h represents the height of the feature map, w represents the width of the feature map, C represents the number of attribute categories, c' represents the current channel, and I a represents the unnormalized feature map.
[0085] Step 4: Based on the spatial feature representations of each attribute, construct an adjacency matrix that includes the spatial position relationships of the relative directions of the attributes, and use it as the prior knowledge of the structure of the graph convolutional network to obtain high-level features that contain the spatial position relationships of the attributes;
[0086] The specific process of the said Step 4 is as follows:
[0087] Step 4.1: Use a graph convolutional network (GCN) to model the spatial position relationships among multi-label attributes. The propagation method between each graph convolutional network layer is:
[0088] H l+1 = F(AH l W l ) (2)
[0089] where, H l is the input feature of the l-th layer, H l+1 is the output feature of the (l + 1)-th layer, A is the adjacency matrix, W l is the learnable parameter matrix, and F(·) is the activation function; corresponding to the present invention, the input feature of the first layer is the output of Step 3.2, that is, the normalized feature map of each attribute, h(·) is the LeakyRelu activation function, and the adjacency matrix A is constructed according to the following steps;
[0090] Step 4.2: Perform edge detection using a 3×3 Sobel operator on each channel dimension of the feature map output in Step 3.2, and set the padding value to 0 to ensure that the size of the output feature map remains unchanged; the Sobel operator is:
[0091]
[0092] where, K x is the convolutional layer in the horizontal direction, and K y is the convolutional layer in the vertical direction;
[0093] Step 4.3: Calculate the final gradient magnitude and perform min-max normalization:
[0094]
[0095]
[0096] where, G x and G y are the gradient magnitudes in the horizontal direction and the vertical direction respectively, and ρ(G) is the edge probability value of each attribute feature map. The larger its value, the more likely it is that this position belongs to the edge area during the recognition process;
[0097] Step 4.4: Select the position with the maximum marginal probability value in each attribute feature map as the spatial position of the attribute; subsequently, project this position back to the original image coordinate system through the spatial mapping relationship to determine the specific position of the attribute in the original image; finally, construct an adjacency matrix using the directional spatial relationship between the attribute position points.
[0098] Step 4.5: Let the positions of attribute i and attribute j mapped to the original image be (x i , y i ) and (x j , y j ), respectively, and calculate the relative direction angle between attribute i and attribute j:
[0099]
[0100]
[0101] where arctan2(x, y) is the extended arctangent function, and the calculated angle range is [-180°, 180°]. Normalize it to [0°, 360°] to obtain the final relative direction angle θ'. ij ;
[0102] Step 4.6: Consider the original image as a 360-degree spatial domain and divide it into 8 sectors, each sector being 45 degrees. Each sector represents a spatial relationship and is numbered from 1 to 8 in a clockwise direction.
[0103] Step 4.7: According to the calculated relative direction angle θ' ij , determine the sector to which it belongs, and use the number of this sector as the value at the corresponding position of the adjacency matrix. By analogy, finally construct an adjacency matrix containing spatial position information; use this adjacency matrix as the structural prior knowledge of the graph convolutional network, and use the graph neural network model to learn the spatial direction position relationship between attributes.
[0104] Step 5: Use the graph embedding method to model the semantic relationship between attributes, construct a new adjacency matrix, and use this to guide the generation of attribute label embeddings. The generated attribute label embeddings are input into the network as additional information, thereby introducing the semantic association between attributes;
[0105] The specific process of step 5 is as follows:
[0106] Step 5.1: Define the graph G = (V, E), where V = (v1, v2,..., v C ) represents the set of C attribute category nodes, and E represents the edges; specifically, V is the set of all attribute labels, and E is the connection set between any two attributes; the adjacency matrix A of graph G aInclude non - negative weights associated with each edge;
[0107] Step 5.2: Initialize four matrices, namely the M matrix, P matrix, W matrix, and A a matrix, and set their initial values to zero. Among them, the M matrix is a co - occurrence matrix used to count the number of times any two labels appear simultaneously; the P matrix is a conditional probability matrix calculated based on the M matrix, representing the conditional probability of label co - occurrence; the W matrix is a label weight matrix used to count the number of times each label appears and perform normalization on it; A a matrix is an adjacency matrix constructed based on the above three matrices, used to describe the relationship between attribute labels. The update formula for the A a matrix is:
[0108] A a [i, j]=A a [i, j]+W[j, j]*P[i, j] (8)
[0109] Step 5.3: Adopt a model of a three - layer fully - connected network to learn to map the one - hot encoding of each attribute to a semantic embedding space and generate corresponding label embeddings. Among them, batch normalization and Relu activation functions are introduced after the first two fully - connected layers;
[0110] Step 5.4: Train and optimize the loss function in the semantic embedding space. The optimization goal is to make the cosine similarity cos(e i , e j ) between any two edges in the semantic embedding space close to the relationship described by the adjacency matrix A a of graph G. The formula for the loss function used is as follows:
[0111]
[0112] where C is the number of attribute categories, cos(e i , e j ) is the cosine similarity between any two edges in the attribute embedding space, and A a is the adjacency matrix of graph G.
[0113] Step 6: Use the outputs of Steps 4 and 5 as input features and perform feature fusion through a multi - head attention mechanism, where the attribute label embeddings of Step 5 are used as queries, and the output spatial position features of Step 4 are used as keys and values;
[0114] The specific process of Step 6 is as follows:
[0115] Step 6.1: Use the attribute label embedding output in Step 5 as the query of the multi-head attention mechanism, and the output spatial position features in Step 4 as the key and value of the multi-head attention mechanism. The multi-head attention mechanism is used to fuse the semantic relationship and spatial position relationship of the attributes. The calculation process of the multi-head attention mechanism is as follows:
[0116] MultiHead(Q,K,V)=Conact(head1,head2,…,head n )W O (10)
[0117] where Q represents the query, K represents the key, V represents the value, n represents the number of attention heads, and W O represents the trainable parameter matrix;
[0118] Step 6.2: The calculation formula for each attention head is as follows:
[0119]
[0120] where, is the trainable parameter matrix;
[0121] The attention score is calculated by scaled dot-product attention:
[0122]
[0123] where d k represents the scaling factor; T represents the transpose.
[0124] Step 7: Linearly transform the output of Step 6 through a fully connected layer, and calculate the prediction probability of each attribute through the Sigmoid activation function;
[0125] The specific process of Step 7 is as follows:
[0126] Use binary cross-entropy loss (BCE) as the overall loss function for training. The output result of the multi-head attention mechanism is linearly transformed through a fully connected layer, and the prediction probability of the attribute label is calculated. The loss function formula for the overall training is:
[0127]
[0128] where y ic represents the true attribute value; p ic represents the predicted attribute probability value.
[0129] To test the effectiveness of the method proposed in this invention, the large-scale fashion dataset deepfashion released by LiuZiwei et al. was selected. This dataset is divided into four benchmark datasets. We chose the sub-dataset containing high-resolution images in the Category and Attribute Prediction Benchmark dataset as the experimental dataset, abbreviated as deepfashion-c-h. In addition, to demonstrate the generalization ability of our method, we also verified our invention on the benchmark dataset MS-COCO. The experiments were all conducted on the GPU of NVIDIA GeForce RTX 2070 SUPER and the PyTorch framework. The optimizer used during the training process was AdamW, with its momentum set to 0.01, the initial learning rate to 0.001, and the learning rate decayed to 0.1 times the original every 30 rounds of training. For the deepfashion-c-h dataset, the batch size was set to 16 and a total of 100 rounds of training were performed. For the MS-COCO dataset, the batch size was 32 and a total of 200 rounds of training were performed.
[0130] The results of comparing the method proposed in this invention with other state-of-the-art methods are shown in Table 1 and Table 2. Among them, Table 1 shows the experimental results of the deepfashion-c-h dataset, and Table 2 shows the results of the MS-COCO dataset. It can be seen from Table 1 that this method is superior to other methods in both category prediction and attribute recognition. Specifically, in terms of classification accuracy, this method is 1.12% higher than sRA-Net in OverallAccuracy and 0.97% higher than sRA-Net in top-1. At the same time, this method has achieved excellent results in each group of attributes. The extensive quantitative results strongly prove the good performance of the method we proposed in the aspect of clothing attribute recognition. For the MS-COCO dataset, we overall pay more attention to three evaluation metrics: mAP, OF1, and CF1. It can be seen from Table 2 that the method proposed in this invention is higher than other methods in these three metrics. Specifically, it has increased by 0.4% in mAP, 0.2% in CF1(top1), 0.2% in OF1(top1), 0.3% in CF1(top3), and 0.3% in OF1(top3).
[0131] Table 1 Comparative experimental results of the deepfashion-c-h dataset
[0132]
[0133] Table 2 Comparative experimental results of the MS-COCO dataset
[0134]
[0135] The content not described in detail in the specification of the present invention belongs to the prior art well-known to those skilled in the art. Although the illustrative specific embodiments of the present invention are described above for the understanding of those skilled in the art of the present technology, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
Claims
1. A clothing attribute recognition method based on the fusion of attribute semantics and spatial location information, characterized in that: It includes the following steps: Step 1: Collect the original clothing images and perform preprocessing. The preprocessing operations mainly include: randomly cropping the images, horizontally flipping the images, centrally cropping the images, and image normalization operations; Step 2: Use the lightweight MobileViT as the backbone network to extract features from the original images and obtain global spatial features; Step 3: After extracting the global spatial features, introduce the attention mechanism to further extract fine-grained features and obtain the spatial feature representations of each attribute; Step 4: Based on the spatial feature representations of each attribute, construct an adjacency matrix containing the spatial position relationships of the relative directions of the attributes, and use it as the structural prior knowledge of the graph convolutional network to obtain high-level features containing the spatial position relationships of the attributes; Step 5: Use the graph embedding method to model the semantic relationships between attributes, construct a new adjacency matrix, and use it to guide the generation of attribute label embeddings. The generated attribute label embeddings are input into the network as additional information, thereby introducing the semantic associations between attributes; Step 6: Take the outputs of Steps 4 and 5 as input features and perform feature fusion through the multi-head attention mechanism, where the attribute label embeddings in Step 5 are used as queries, and the output spatial position features in Step 4 are used as keys and values; Step 7: Linearly transform the output of Step 6 through a fully connected layer and calculate the prediction probability of each attribute through the Sigmoid activation function.
2. The clothing attribute recognition method based on the fusion of attribute semantics and spatial location information according to claim 1, characterized in that The specific process of Step 2 is: Step 2.1: Remove the last 1×1 convolutional layer, pooling layer, and fully connected layer from the MobileViT network, and retain the remaining network structure as the backbone network to extract the global spatial features of the images.
3. The method for identifying clothing attributes based on the fusion of attribute semantics and spatial location information according to claim 2, wherein The specific process of Step 3 is: Step 3.1: The structure of the attention mechanism is three convolutional layers. The Relu function is executed after the first two convolutional layers, and the Softmax function is executed after the last convolutional layer; Step 3.2: The input of the attention mechanism is the global spatial features extracted by the MobileViT network. The number of output channels is equal to the number of attribute categories, and each channel corresponds to an unnormalized feature map of an attribute; the output unnormalized feature maps are normalized in the channel dimension through the Softmax function to obtain the normalized feature maps of each attribute. The formula is: Among them, h represents the height of the feature map, w represents the width of the feature map, C represents the number of attribute categories, c' represents the current channel, and I a represents the unnormalized feature map.
4. The clothing attribute recognition method based on the fusion of attribute semantics and spatial location information according to claim 3, characterized in that The specific process of Step 4 is: Step 4.1: Use the graph convolutional network to model the spatial position relationships between multi-label attributes. The propagation method between each graph convolutional network layer is: H l+1 = F(AH l W l ) (2) Among them, H l is the input feature of the l-th layer, and H l+1 is the output feature of the (l + 1)-th layer. A is the adjacency matrix, and W l is the learnable parameter matrix. F(·) is the activation function. The adjacency matrix A is constructed according to the following steps; Step 4.2: Perform edge detection using the 3×3 Sobel operator on each channel dimension of the feature maps output in Step 3.2, set the padding value to 0 to ensure that the size of the output feature maps remains unchanged; the Sobel operator is: Among them, K x is a convolutional layer in the horizontal direction, and K y is a convolutional layer in the vertical direction; Step 4.3: Calculate the final gradient magnitude and perform min-max normalization: Among them, G x and G y are the gradient magnitudes in the horizontal and vertical directions respectively, and ρ(G) is the edge probability value of each attribute feature map; Step 4.4: Select the position with the maximum edge probability value in each attribute feature map as the spatial position of the attribute; subsequently, project this position back to the original image coordinate system through the spatial mapping relationship to determine the specific position of the attribute in the original image; finally, construct an adjacency matrix using the directional spatial relationships between the attribute position points; Step 4.5: Assume that the positions where attribute i and attribute j are mapped to the original image are (x i , y i ) and (x j , y j ), and calculate the relative direction angle between attribute i and attribute j: Among them, arctan2(x, y) is the extended arctangent function, and the calculated angle range is [-180°, 180°]. Normalize it to [0°, 360°] to obtain the final relative direction angle θ'. ij ; Step 4.6: Consider the original image as a 360-degree spatial domain, divide it into 8 sectors, each sector being 45 degrees, where each sector represents a spatial relationship and is sequentially numbered from 1 to 8 in a clockwise direction; Step 4.7: Determine the sector to which the calculated relative direction angle θ' belongs, and use the number of this sector as the value at the corresponding position in the adjacency matrix. By analogy, finally construct an adjacency matrix containing spatial position information; use this adjacency matrix as the structural prior knowledge of the graph convolutional network, and use the graph neural network model to learn the spatial direction and position relationship between attributes. ij 5. The clothing attribute recognition method based on the fusion of attribute semantics and spatial location information according to claim 4, wherein The specific process of the said Step 5 is as follows: Step 5.1: Define a graph G = (V, E), where V = (v1, v2, …, v C ) represents a set of C attribute category nodes, and E represents edges; specifically, V is the set of all attribute labels, and E is the connection set between any two attributes; the adjacency matrix A of graph G a contains non - negative weights associated with each edge; Step 5.2: Initialize four matrices, namely the M matrix, the P matrix, the W matrix, and the A a matrix, and set their initial values to zero. Among them, the M matrix is a co-occurrence matrix used to count the number of times any two tags appear simultaneously; the P matrix is a conditional probability matrix calculated based on the M matrix, representing the conditional probability of tag co-occurrence; the W matrix is a tag weight matrix used to count the number of times each tag appears and perform normalization on it; the A a matrix is an adjacency matrix constructed based on the above three matrices and is used to describe the relationship between attribute tags. The update formula for the A a matrix is as follows: A a [i,j] = A a [i,j] + W[j,j] * P[i,j] (8) Step 5.3: Adopt a model of a three-layer fully connected network for learning to map the one-hot encoding of each attribute to a semantic embedding space and generate corresponding label embeddings. Among them, batch normalization and the Relu activation function are introduced after the first two fully connected layers; Step 5.4: Train and optimize the loss function of the semantic embedding space. The optimization goal is to make the cosine similarity cos(e i ,e j ) between any two edges close to the relationship described by the adjacency matrix A a of the graph G. The formula of the loss function used is as follows: Among them, C is the number of attribute categories, and cos(e i , e j ) is the cosine similarity of any two edges in the attribute embedding space, and A a is the adjacency matrix of graph G.
6. The clothing attribute recognition method based on the fusion of attribute semantics and spatial location information according to claim 5, characterized in that, The specific process of the said Step 6 is as follows: Step 6.1: Use the attribute label embeddings output by Step 5 as the queries of the multi-head attention mechanism, and the output spatial position features of Step 4 as the keys and values of the multi-head attention mechanism. Adopt the multi-head attention mechanism to fuse the semantic relationship and spatial position relationship of the attributes. The calculation process of the multi-head attention mechanism: MultiHead(Q,K,V)=Conact(head1,head2,…,head n )W O (10) Among them, Q represents the query, K represents the key, V represents the value, n represents the number of attention heads, and W O represents the trainable parameter matrix; Step 6.2: The calculation formula of each attention head is as follows: Among them, is a trainable parameter matrix; The attention score is calculated by scaled dot-product attention: where d k represents a scaling factor; T represents transpose.
7. The clothing attribute recognition method based on the fusion of attribute semantics and spatial location information according to claim 6, characterized in that The specific process of the said Step 7 is as follows: Use binary cross-entropy loss as the overall loss function for training. The output result of the multi-head attention mechanism undergoes a linear transformation through a fully connected layer, and the prediction probability of the attribute label is calculated. The formula of the overall training loss function is: Among them, y ic represents the true attribute value; p ic represents the predicted attribute probability value.