Clothing style identification method based on multi-scale layered feature fusion

Through the feature extraction and adaptive fusion of MobileNetV3 and CBAM modules, combined with the dual-branch decision structure, the problem of insufficient fusion of multi-scale features in clothing style recognition is solved, and efficient and accurate clothing style recognition is achieved.

CN120259825APending Publication Date: 2025-07-04XIANGTAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510410683.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing clothing style recognition methods have problems such as insufficient fusion of multi-scale hierarchical features, low model calculation efficiency, and weak feature distinction ability.

Method used

Multi-scale features are extracted using MobileNetV3 backbone network, feature reweighting is combined with CBAM module, and hierarchical feature fusion is realized through adaptive feature pyramids, and classification decisions are used to make use of dual-branch decision structures, and category feature centers are dynamically updated.

Benefits of technology

It significantly improves the accuracy and efficiency of clothing style recognition, meets the needs of real-time interaction and cross-platform compatibility, and enhances the ability to distinguish similar style categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259825A_ABST
    Figure CN120259825A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer vision and pattern recognition, in particular to a garment style recognition method based on multi-scale layered feature fusion. According to the method, a MobileNetV3 backbone network is utilized to extract multi-scale features of a garment image, feature reweighting is performed through a CBAM module, hierarchical feature fusion is realized by adopting an adaptive feature pyramid, and style classification is completed in combination with a double-branch decision structure. According to the method, the problems of insufficient multi-scale feature fusion, low model calculation efficiency, weak feature distinguishing capability and the like in the existing garment style recognition technology are solved, and the accuracy and efficiency of garment style recognition are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and pattern recognition, and particularly to a clothing style recognition method based on multi-scale hierarchical feature fusion. Background Art

[0002] As an important application in the field of computer vision, clothing style recognition still faces multiple technical challenges. Traditional clothing image feature recognition methods mainly rely on manual annotation or basic visual feature analysis, and have certain limitations in terms of the depth of feature extraction and recognition efficiency, making it difficult to meet the recognition requirements in complex scenarios.

[0003] As one of the core technologies in the field of artificial intelligence, neural network models have made breakthrough progress in visual processing tasks such as image classification, object detection, and semantic segmentation. In particular, visual processing models based on convolutional neural networks can effectively capture multi-level semantic information in images through a hierarchical feature extraction mechanism, significantly improving the pattern recognition accuracy in complex scenarios. Through an end-to-end learning architecture, such models achieve a direct mapping from raw pixel data to high-level semantic features, providing important methodological support for clothing style recognition technology. By constructing a multi-dimensional style feature system, the limitations of traditional methods can be effectively overcome.

[0004] However, existing neural network-based clothing style recognition methods still have three technical problems: First, most methods use single-scale feature extraction, making it difficult to effectively integrate local texture details and global semantic information; second, although existing attention mechanisms can enhance the expression of key features, the computational cost is relatively high, affecting the real-time processing ability of the model; third, the traditional single-branch classification structure lacks an optimization mechanism for feature representation, resulting in insufficient discrimination between similar style categories.

[0005] Therefore, there is an urgent need for a clothing style recognition method that can effectively integrate multi-scale features, improve computational efficiency, and enhance feature discrimination ability to address the above technical challenges. Summary of the Invention

[0006] The technical problems to be solved by the present invention are the technical problems existing in existing clothing recognition methods, such as insufficient multi-scale hierarchical feature fusion, low computational efficiency of the model, and weak feature discrimination ability. The purpose of the present invention is to provide a lightweight clothing style recognition method based on multi-scale hierarchical feature fusion, which can significantly improve the style recognition accuracy while maintaining the lightweight of the model, and provide an efficient solution for real-time interaction, cross-platform compatibility, and classification applications in specific scenarios.

[0007] The present invention is achieved through the following technical solutions:

[0008] A clothing style recognition method based on multi-scale hierarchical feature fusion, characterized in that the method includes:

[0009] 1) Obtain the target image and preprocess the collected image;

[0010] 2) The preprocessed clothing image extracts the original multi-scale feature maps from the shallow, middle, and deep layers through the MobileNetV3 backbone network, performs resolution alignment and feature re-weighting on the extracted original multi-scale feature maps, and fuses the feature maps after feature re-weighting through an adaptive feature pyramid to generate unified features;

[0011] 3) Obtain the feature vector after passing the unified features through global average pooling and a fully connected layer, input the feature vector into a two-branch decision structure, and obtain the classification probability distribution and similarity probability distribution through the classification branch and the feature metric branch respectively, weighted sum them according to the learnable weights, and take the category corresponding to the maximum value as the final recognition result.

[0012] In the present invention, optionally, the feature re-weighting uses the CBAM module, and the CBAM module includes a channel attention sub-module and a spatial attention sub-module, which generate channel weights and spatial weights respectively.

[0013] In the present invention, optionally, the adaptive feature pyramid is specifically: the pyramid structure adaptively weights and fuses the deep features and the middle features through a top-down path to obtain the deep-middle fusion features, re-performs feature re-weighting on the deep-middle fusion features and adaptively weights and fuses them with the shallow features to obtain unified features; during the adaptive weighted fusion process, the normalized adaptive weights are used for weighted summation.

[0014] In the present invention, optionally, the method for obtaining the adaptive weights is specifically: generating initial weights for the multi-scale feature maps after resolution alignment through lightweight 1×1 convolution, multiplying them by the sum of the channel weights and the spatial weights to obtain corrected weights; normalizing the corrected weights through the Softmax function to obtain the normalized adaptive weights.

[0015] In the present invention, optionally, the two-branch decision structure includes a classification branch and a feature metric branch, specifically: the classification branch inputs the feature vector into a fully connected layer and obtains the classification probability distribution of the clothing style category through the Softmax function; the feature metric branch inputs the feature vector into a fully connected layer, calculates the normalized cosine similarity between the feature vector and the category feature center, and generates the similarity probability distribution after scaling by the temperature coefficient; the classification probability distribution and the similarity probability distribution are weighted and summed according to the learnable weights, and the category corresponding to the maximum value is taken as the final recognition result.

[0016] Optionally, in the present invention, the category feature center is specifically: in the training stage, the category feature centers are initialized as the mean of the feature vectors of the corresponding categories, and are updated in a moving average manner through backpropagation. In the inference stage, it is a fixed network parameter.

[0017] Optionally, in the present invention, the learnable weights are updated together with other network parameters through the backpropagation algorithm in the model training stage, and are a fixed network parameter in the inference stage.

[0018] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0019] 1. Optimization of multi-scale hierarchical feature fusion: Through an adaptive feature pyramid structure, dynamic weighted fusion of multi-level features is achieved, effectively balancing the representation capabilities of local details and global semantics, and significantly improving the style recognition accuracy in complex scenarios.

[0020] 2. Lightweight and high efficiency: Using MobileNetV3 as the backbone network, combined with a lightweight CBAM module and an adaptive weight generation mechanism, while keeping the number of model parameters small, efficient calculation of feature re-weighting and fusion is achieved, meeting the real-time requirements of mobile devices.

[0021] 3. Enhancement of discriminative power through dual-branch decision-making: The joint decision of the classification branch and the feature metric branch not only retains the direct prediction advantage of traditional classification methods but also optimizes the feature compactness through feature center measurement, significantly enhancing the discrimination ability between similar styles.

[0022] 4. Dynamic update of feature centers: Based on the category feature center update strategy of moving average, it effectively adapts to the change of training data distribution and enhances the generalization ability of the model to new samples.

[0023] Through lightweight model design, multi-scale feature fusion, and dual-branch decision-making technology, the present invention effectively solves the technical bottlenecks of traditional clothing recognition systems in terms of real-time performance, multi-scenario adaptability, and feature discriminative power, providing an efficient technical solution for various clothing image recognition application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a flowchart of a clothing style recognition method based on multi-scale hierarchical feature fusion according to the present invention;

[0025] Figure 2 is a schematic diagram of the structure of the clothing style recognition model according to the present invention;

[0026] Figure 3 is a flowchart of the image preprocessing according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0027] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0028] Embodiment

[0029] The present invention provides a clothing style recognition method based on multi-scale hierarchical feature fusion, including the following steps:

[0030] Step 1: Obtain a target image and preprocess the collected image.

[0031] Step 2: Input the preprocessed clothing image into a pre-trained clothing style recognition model to obtain a clothing style detection result.

[0032] Among them, in the clothing style recognition model, the clothing image first extracts the original multi-scale feature maps from the shallow, middle, and deep layers through the MobileNetV3 backbone network, performs resolution alignment and feature re-weighting on the extracted original multi-scale feature maps, and fuses the multi-scale feature maps after feature re-weighting through an adaptive feature pyramid to generate a unified feature.

[0033] The unified feature is passed through global average pooling and a fully connected layer to obtain a feature vector, and the feature vector is input into a two-branch decision structure. The classification probability distribution and similarity probability distribution are respectively obtained through the classification branch and the feature metric branch, weighted and summed according to the learnable weights, and the category corresponding to the maximum value is taken as the final recognition result.

[0034] As an embodiment, the specific implementation process of a clothing style recognition method based on multi-scale hierarchical feature fusion of the present disclosure is as follows:

[0035] Step 1: Construct a clothing style dataset and model training, specifically including:

[0036] S1: Construct a clothing style dataset. The clothing datasets used in this embodiment all come from Deepfasion. Screen the clothing pictures of different style types in the clothing style dataset to obtain a labeled picture dataset of 6 different clothing style types, including:

[0037] (1) Ethnic style: The main colors are saturated colors such as red, peacock blue, and emerald green, including hand embroidery, totem patterns, and tie-dyeing techniques, and are paired with decorative elements such as wide cuffs, fringes, and frog buttons;

[0038] (2) Lolita style: Using pastel colors, a large number of decorative elements such as lace, bows, and ruffles are included, combined with a fluffy skirt and puff sleeve design, and a corset and petticoat are used to construct a three-dimensional silhouette;

[0039] (3) Business style: Select steady color systems such as navy blue, charcoal gray, ivory white, etc., and adopt typical styles such as crisp suit collars, straight-leg trouser shapes, and H-shaped coats;

[0040] (4) Collegiate style: Represented by classic colors such as dark blue and wine red, use grid and stripe geometric patterns, commonly the combination of V-neck knitted vests and shirt ties, and match with A-line skirts;

[0041] (5) Minimalist style: Dominated by black, white, gray and low-saturation neutral colors, abandon decorative patterns, and the clothing is loose;

[0042] (6) Bohemian style: Adopt earthy color systems such as ochre, terracotta orange, amber yellow, etc., and present irregular hemlines, multi-layered layering and cut-out crochet designs;

[0043] Use an image enhancement method to perform image enhancement on the annotated picture dataset of different clothing style types to obtain a clothing style dataset after image enhancement;

[0044] Perform normalization and standardization processing on the clothing style dataset after image enhancement to obtain a clothing style dataset after data scaling processing;

[0045] Divide the clothing style dataset after data scaling processing into a training set, a validation set, and a test set according to the ratio of 8:1:1.

[0046] Divide the processed dataset into a training set, a validation set, and a test set according to the ratio of 8:1:1: The training set (80%) is used for model parameter learning and optimization, the validation set (10%) is used for hyperparameter tuning and model selection, and the test set (10%) is used for final model performance evaluation.

[0047] S2: Build a clothing style recognition model. The architecture of the clothing style recognition model is as Figure 2 shown.

[0048] S2.1: Original multi-scale feature extraction;

[0049] Use MobileNetV3 as the backbone network. The MobileNetV3 backbone network has two branches, Small and Large. The Small branch contains 11 residual blocks, and the Large branch contains 15 residual blocks. In this embodiment, the Large branch will be taken as an example.

[0050] The MobileNetV3 network outputs feature maps at three scales: shallow, middle, and deep in parallel. The shallow features retain rich texture details, the middle features contain local structure information, and the deep features have high-level semantic information, which are respectively represented as shallow features, middle features, and deep features.

[0051] S2.2: Feature reweighting;

[0052] First, the features at each scale are aligned in resolution through bilinear interpolation to eliminate size differences.

[0053] Subsequently, the Convolutional Block Attention Module (CBAM) is used for feature reweighting. This module includes a channel attention sub-module and a spatial attention sub-module. The channel attention sub-module generates two feature descriptors through global average pooling and max pooling respectively, which are fused after being processed by a shared multi-layer perceptron to generate channel weights. The spatial attention sub-module performs average pooling and max pooling operations on the feature map in the channel dimension, and after concatenation, generates a spatial weight map through a 7×7 convolution and Sigmoid activation.

[0054] S2.3: Adaptive Feature Fusion;

[0055] The pyramid structure adaptively weights and fuses the deep features and middle-level features through a top-down path to obtain deep-middle fusion features, and then reweights the deep-middle fusion features and adaptively weights and fuses them with the shallow features to obtain unified features. The feature fusion process uses adaptive weights for weighted summation to achieve the dynamic balance of multi-scale features.

[0056] Specifically, the adaptive weighting method is as follows: the initial weights are generated by a lightweight 1×1 convolution on the multi-scale feature maps after resolution alignment. The initial weights are multiplied by the sum of the channel weights and spatial weights to obtain corrected weights, and the corrected weights are normalized through the Softmax function to obtain the normalized adaptive weights.

[0057] The fusion process is specifically as follows:

[0058] The adaptive feature pyramid first fuses the deep and middle-level features to obtain deep-middle fusion features:

[0059]

[0060] Then, the deep-middle fusion features are reweighted and fused with the shallow features to obtain unified features:

[0061]

[0062] Among them, F deep-mid represent the shallow features, middle-level features, deep features, and deep-middle fusion features respectively, and α low 、α mid 、α high 、α deep-midThey respectively represent the normalized adaptive weights of shallow features, middle features, deep features, and deep-middle fusion features.

[0063] S2.4: Dual-branch collaborative decision-making;

[0064] The unified features obtained in S2.3 are passed through global average pooling (GAP) and a fully connected layer (FC) to obtain a feature vector. The feature vector is parallelly input into two decision branches: a classification branch and a feature metric branch.

[0065] The classification branch maps the feature vector to the category space through a fully connected layer and obtains the classification probability distribution of the clothing style category through the Softmax function.

[0066] The feature metric branch calculates the cosine similarity between the feature vector and the category feature center. First, the feature vector is projected into the metric space through a fully connected layer, then the normalized cosine similarity between the feature vector and the category feature centers of each category is calculated, and finally the similarity probability distribution after normalization is obtained, which is used to characterize the difference size of different clothing style categories.

[0067] Finally, the system fuses the two distributions through a learnable weighting coefficient λ to obtain the final decision probability:

[0068] p final = λ · p cls +(1 - λ) · p sim

[0069] Among them, p cls is the classification probability distribution, and p sim is the similarity probability distribution.

[0070] Take the category corresponding to the maximum value in p final as the final recognition result:

[0071]

[0072] Among them, j ∈ {1, 2,..., N} represents the category index.

[0073] The weight parameter λ is used as a network learnable weight, initialized to 0.5, and updated together with other network parameters through backpropagation. During the training process, the value of λ will be automatically adjusted according to the contribution of the two branches to the final classification accuracy. To prevent λ from being too large or too small resulting in one branch being completely ignored, the Sigmoid function is applied to constrain the value of λ, and its value is always between 0 and 1.

[0074] S3: Training of the clothing style recognition model.

[0075] The backbone network (MobileNetV3) is loaded with pre-trained weights, and the remaining modules (CBAM, adaptive pyramid, dual-branch) are randomly initialized. The class feature center is initialized as the mean of the feature vectors of the training samples corresponding to the class.

[0076] The training process of the model adopts a weighted combination of the cross-entropy loss function and the cosine similarity loss function:

[0077] L total = L CE + γ·L cos

[0078] where L CE is the standard cross-entropy loss, L cos is the metric learning loss based on cosine similarity, and γ is the weight coefficient that balances the two losses, which is set to 0.2 in this example.

[0079] Adam optimizer is used for training, the initial learning rate is set to 0.001, and the cosine annealing learning rate scheduling strategy is adopted. A total of 150 epochs are trained. The mini-batch training method with a batch size of 32 is used.

[0080] Dynamic feature center update strategy: In the initial stage of training, the class feature centers are initialized as the means of the feature vectors of the corresponding class samples. During the training process, they are updated by moving average after each mini-batch;

[0081] The specific formula is:

[0082]

[0083] where:

[0084] is the feature center of the k-th class at time step t;

[0085] B k is the set of samples belonging to the k-th class in the current batch;

[0086] f i is the feature vector of sample i;

[0087] α is the momentum coefficient, which is set to 0.9.

[0088] Early stopping mechanism: After each epoch ends, the performance of the model is evaluated using the validation set, and the model parameters with the best performance on the validation set are saved. When the performance on the validation set has not improved for 10 consecutive epochs, the training is terminated early, and the model with the best performance on the validation set is selected as the final model.

[0089] Final evaluation: Use the test set to evaluate the performance of the best model, and calculate the accuracy, precision, recall, and F1 score.

[0090] Step 2: Obtain the target image and preprocess the collected image. The image preprocessing process is as Figure 3 shown, including object detection and region cropping, Gaussian filtering, image enhancement, gamma transformation, size and pixel normalization;

[0091] Step 2.1 Portrait detection uses the Haar cascade classifier in OpenCV to locate the portrait region based on Haar features and cascade classification algorithm and crop to obtain the clothing image;

[0092] Step 2.2 Apply Gaussian filtering to the cropped image to denoise the image;

[0093] Step 2.3 Use the CLAHE algorithm to enhance the contrast of local details and texture features;

[0094] Step 2.4 Use gamma transformation to adjust the image brightness;

[0095] Step 2.5 Scale the image to 224×224 pixels and normalize it to the range [0,1] to reduce the differences between different images.

[0096] Step 3: Input the preprocessed clothing image into the clothing style recognition model in Step 1 to obtain the clothing style detection result;

[0097] Step 3.1: Feature extraction stage. The preprocessed image (224×224×3) is input into the MobileNetV3 backbone network to generate the shallow feature map F low (56×56×24), the middle feature map F mid (28×28×40) and the deep feature map F high (14×14×112).

[0098] Step 3.2: Feature reweighting stage. The three feature maps are first adjusted to the same resolution (28×28) through bilinear interpolation, and then respectively undergo feature reweighting through the CBAM module to generate the enhanced feature maps and

[0099] Step 3.3: Feature fusion stage. Through the adaptive feature pyramid structure, first fuse and to obtain the deep-middle fusion feature F deep-mid , F deep-mid After reweighting the features again, fuse with to obtain the unified feature F fusion .

[0100] Step 3.4: Decision stage. After the unified features are processed by global average pooling and the fully connected layer, the classification probability distribution and similarity probability distribution are calculated in parallel through the dual-branch structure, and the final prediction result is obtained by fusion according to the learned weights.

[0101] This method achieves high-precision recognition of clothing styles while maintaining the lightweight and high efficiency of the model, meeting the needs of real-time interaction on mobile terminals. This method makes full use of the complementarity of multi-scale features, effectively extracts key style features of clothing through attention mechanism and adaptive fusion strategy, and enhances the model's discriminative ability through a dual-branch decision structure, especially for the recognition of subtle differences in similar style categories.

[0102] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Technical personnel in the relevant field should understand that on the basis of the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.

Claims

1. A clothing style recognition method based on multi-scale hierarchical feature fusion, characterized in that It includes the following steps: 1) Obtain the target image and preprocess the collected image; 2) The preprocessed clothing image extracts the original multi-scale feature maps from the shallow, middle, and deep layers through the MobileNetV3 backbone network. Align the resolutions of the extracted original multi-scale feature maps, re-weight the features, and fuse the multi-scale feature maps with re-weighted features through an adaptive feature pyramid to generate unified features; 3) Obtain a feature vector after passing the unified features through global average pooling and a fully connected layer, and input the feature vector into a two-branch decision structure. Obtain the classification probability distribution and similarity probability distribution through the classification branch and the feature metric branch respectively, perform weighted summation according to learnable weights, and take the category corresponding to the maximum value as the final recognition result.

2. The clothing style recognition method based on multi-scale hierarchical feature fusion according to claim 1, wherein, The step 2) includes the following steps: The feature re-weighting uses the CBAM module. The CBAM module includes a channel attention sub-module and a spatial attention sub-module, which generate channel weights and spatial weights respectively; The pyramid structure adaptively weights and fuses the deep features and middle features through a top-down path to obtain deep-middle fusion features. After re-weighting the deep-middle fusion features, they are adaptively weighted and fused with the shallow features to obtain unified features; During the adaptive weighted fusion process, normalized adaptive weights are used for weighted summation; Among them, the acquisition of adaptive weights includes the following steps: generate initial weights for the multi-scale feature maps after resolution alignment through lightweight 1×1 convolution, and multiply them by the sum of the channel weights and spatial weights to obtain corrected weights; Normalize the corrected weights through the Softmax function to obtain normalized adaptive weights.

3. A method for clothing style recognition based on multi-scale hierarchical feature fusion according to claim 1, characterized in that, The step 3) includes the following steps: The classification branch inputs the feature vector into a fully connected layer and obtains the classification probability distribution of the clothing style category through the Softmax function; The feature metric branch inputs the feature vector into a fully connected layer, calculates the normalized cosine similarity between the feature vector and the category feature center, and generates a similarity probability distribution after scaling by a temperature coefficient; Perform weighted summation on the classification probability distribution and the similarity probability distribution according to learnable weights, and take the category corresponding to the maximum value as the final recognition result; Among them, the category feature center and the learnable weights are both fixed network parameters.

Citation Information

Cited By

  • Clothing material identification method and device

    CN121884063A

  • Clothing material identification method and device

    CN121884063B