A Method and System for Clothing Style Recognition and Database Entry Based on an Improved Convolutional Neural Network
Patent Information
- Application Number
- CN202610833848.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2046-06-10
AI Technical Summary
[0012]本发明的目的在于提供一种基于改进型卷积神经网络的服装款式识别入库方法及系统,以现有技术中服装款式识别精度有待提升的问题
[0052] 1. This invention proposes a deep feature aggregation network suitable for clothing images. It constructs a progressive feature enhancement structure of "intra-block deep aggregation - branch hierarchical completion - unified scale cross-layer feature fusion" through the design of multi-level deep aggregation layers, hierarchical feature completion branches, and cross-layer feature fusion layers, which is used to improve the feature expression capability of clothing images.
Smart Images

Figure CN122368972B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a method and system for clothing style recognition and database entry based on an improved convolutional neural network. Background Technology
[0002] With the development of e-commerce, smart retail, digitalization of the apparel supply chain, and visual artificial intelligence technology, automatic recognition of apparel images, style classification, similar style retrieval, and automatic database entry have been widely applied in apparel manufacturing, warehouse management, merchandise operation, and e-commerce database construction. In practical applications, it is usually necessary to analyze the collected apparel images and compare them with existing apparel images or style features in the database to determine whether they belong to existing styles or new styles, thereby completing information association, record updates, or new style filing.
[0003] Because clothing images possess distinctly fine-grained recognition characteristics, different garments may be quite similar in overall outline, but subtle differences exist in texture, pattern, neckline, sleeve type, pocket, splicing structure, and decorative elements. Furthermore, clothing images are easily affected by factors such as shooting angle, lighting changes, background interference, occlusion, and pose variations, causing the same style to appear significantly different under different acquisition conditions, while different styles may exhibit high similarity. Therefore, extracting stable and highly discriminative features from clothing images has become a key issue in clothing style recognition and automatic database entry.
[0004] In existing technologies, image recognition methods based on convolutional neural networks, such as VGG, ResNet, and DenseNet, are commonly used to extract features from clothing images and then identify the clothing through classification or similarity matching. Some systems also calculate the similarity between the extracted feature vectors and the features of existing clothing samples in a database to determine the closest existing style and perform database entry processing accordingly.
[0005] However, existing technologies are mostly built on general image classification or retrieval frameworks. In the context of clothing style recognition and automatic database entry, they still have shortcomings in multi-level feature expression, multi-scale information extraction, cross-level feature fusion, and differentiation of similar styles, which in turn affect the accuracy of clothing recognition and the level of intelligence of automatic database entry.
[0006] First, existing methods are still insufficient in representing the features of clothing images. In addition to the overall outline information of clothing, clothing images also contain fine-grained features such as texture, pattern, neckline, sleeve type, pocket, and splicing structure. However, traditional networks mostly use a layer-by-layer convolutional extraction method, which does not make full use of features at different levels. As a result, the obtained feature representation cannot simultaneously take into account the overall structural information and local detail information, thus affecting the recognition effect of clothing style.
[0007] Secondly, existing methods have relatively limited ability to model multi-scale features. Clothing images contain both large-scale features reflecting the overall pattern and structure, and small-scale features reflecting local embellishments and texture differences. Both types of features play an important role in clothing style recognition. In existing technologies, some feature extraction networks mainly rely on convolutional operations at a single scale or with a fixed receptive field, making it difficult to effectively and collaboratively extract information at different scales. Therefore, when faced with clothing images that are similar in style but differ in local details, the recognition accuracy is easily affected.
[0008] Furthermore, existing technologies still have room for improvement in fusing features at different levels. Generally, shallow features contain more detailed information such as edges and textures, while deep features contain stronger semantic information. For clothing style recognition tasks, relying solely on features at a single level often fails to form a complete and stable image representation. Existing methods typically employ relatively simple methods to fuse shallow and deep features, lacking an effective integration mechanism at a unified scale, thus failing to fully leverage the complementary effects of features at different levels.
[0009] Furthermore, the discriminative power of existing methods in identifying similar clothing styles still needs improvement. Because different clothing styles may have high similarity in appearance—for example, similar overall silhouettes but differences in patterns, textures, or local designs—traditional feature extraction methods often struggle to accurately represent these subtle differences. This results in insufficient matching accuracy between similar styles, hindering accurate retrieval and style classification in clothing databases.
[0010] Furthermore, existing database management methods have relatively limited intelligence in terms of automatic data entry. For newly acquired clothing images, current systems typically only provide similarity matching results. Whether the clothing image belongs to an existing style in the database or should be recorded as a new style often requires manual judgment. This not only increases database maintenance costs but also hinders the integrated automatic processing of clothing image recognition and database management.
[0011] Therefore, there is an urgent need to propose a method for clothing style recognition and automatic warehousing to improve the accuracy of style recognition and ensure the reliability of automatic warehousing. Summary of the Invention
[0012] The purpose of this invention is to provide a method and system for clothing style recognition and database entry based on an improved convolutional neural network, addressing the problem that the accuracy of clothing style recognition in existing technologies needs improvement. To this end, this invention provides the following technical solution:
[0013] On one hand, the present invention provides a method for clothing style recognition and database entry based on an improved convolutional neural network, comprising the following steps:
[0014] Image acquisition and preprocessing: Acquire clothing images and perform preprocessing;
[0015] Feature extraction involves inputting the preprocessed clothing image into a deep feature aggregation network of an improved convolutional neural network to extract a global fusion feature vector. The deep feature aggregation network has a multi-level deep aggregation layer, a hierarchical feature completion branch, and a cross-layer feature fusion layer;
[0016] The multi-level deep aggregation layer comprises multiple sequentially connected deep feature aggregation blocks; the hierarchical feature completion branch comprises multiple parallel feature completion branches, each feature completion branch being connected one-to-one with each deep feature aggregation block, and the outputs of each feature completion branch are fused through feature concatenation; the cross-layer feature fusion layer performs a unified scale transformation and cross-layer fusion on the output features of the hierarchical feature completion branches and each deep feature aggregation block to obtain a global fused feature vector used for database entry determination. ;
[0017] Feature matching will extract the global fusion feature vector. The input matching layer performs matching to obtain the style recognition and database entry results;
[0018] The style recognition model consists of a deep feature aggregation network and a matching layer. After the model is pre-trained, it identifies clothing images to be added to the database according to the above process.
[0019] Optionally, the deep feature aggregation block performs intra-block joint aggregation of output features with different convolution depths and output features of the multi-scale structural symmetry perception module.
[0020] The deep feature aggregation block contains two or three consecutively stacked convolutional layers and a multi-scale structural symmetry perception module parallel to the convolutional layers. The outputs of all convolutional layers and the output of the multi-scale structural symmetry perception module are aggregated within the block, and then... The output features of the depth feature aggregation block are obtained after convolution and pooling layers;
[0021] The multi-scale structural symmetry perception module includes multiple parallel symmetry perception convolutional branches with different kernel sizes and adaptive channel adjustment branches. Finally, the output features of each branch are concatenated and fused. Convolution is used to obtain the output features of the multi-scale structural symmetry perception module.
[0022] Optionally, the multi-scale structural symmetry sensing module has three parallel symmetry sensing convolution branches, with the following kernel sizes: , , ;
[0023] For any symmetric perceptual convolution branch, define the kernel size. Step length ,filling Position of the feature map output by the symmetry-aware convolution branch value at for:
[0024]
[0025] in, For the feature map, the first The channel corresponds to the location The value at that location, For bias terms, For output channel index, This is the maximum value of the output channel. Input channel index v is the column index within the convolution kernel. This is the weight tensor actually used for convolution after symmetric constraints; This represents the input features of the symmetry-aware convolution branch; The row and column position parameters of the feature map.
[0026] Optionally, the hierarchical feature completion branch has 5 parallel feature completion branches, specifically two 7×7, one 5×5, and two 3×3 hierarchical convolutions;
[0027] Each feature completion branch performs convolution, activation function, and adaptive pooling operations sequentially.
[0028] Optionally, the matching layer of the style recognition model introduces a bimodal joint similarity determination, which includes feature similarity and attribute similarity. The feature similarity is based on a globally fused feature vector. The similarity is represented by the attribute similarity, which is composed of a weighted average of the similarity of multiple attributes. The attribute types include all or part of the following: color, collar type, sleeve type, placket, pattern, and pattern.
[0029] The model for the bimodal joint similarity is expressed as follows:
[0030]
[0031] in, For the joint score of the two modes, For preset coefficients, The total score for attribute similarity; The total score for feature similarity; The global fusion feature vector extracted from clothing images to be imported into the database using a deep feature aggregation network. , The global fusion feature vector extracted from the inventory image of the style database using a deep feature aggregation network. .
[0032] Optionally, the bimodal joint similarity determination is based on the bimodal joint similarity and employs a dual-threshold determination, specifically as follows:
[0033]
[0034] In the formula, The upper bound of the threshold, This is the lower bound of the threshold. The central threshold represents the baseline strictness of "whether they are the same model"; The indeterminate half-width represents the degree of ambiguity in the machine's judgment, and ;k represents the clothing image identified in the kth entry into the database;
[0035] Uncertain halfwidth Driven by three dynamic variables normalized to [0,1]: image acquisition confidence Historical fluctuation range Dispersion of style feature distribution :
[0036] Uncertain halfwidth The calculation formula is:
[0037]
[0038] Where: Uncertain half-width base value , The image unreliability coefficient. Style dispersion coefficient This represents the historical volatility coefficient.
[0039] Optionally, the preprocessing includes: introducing a clothing structure-aware adaptive scaling module to normalize the size of the clothing image, that is, performing the following in sequence: grayscale conversion and edge enhancement, horizontal centering based on symmetry correction, semantic partitioning detection based on vertical projection peaks and valleys, content-aware adaptive cropping window, proportional scaling and centering fill, normalization and standardization.
[0040] The grayscale and edge enhancement operations yield an edge intensity map; the edge intensity map is then corrected by a horizontal centering operation based on symmetry correction to correct the horizontal boundaries; then, semantic partitioning detection based on vertical projection peaks and valleys identifies each garment region, namely the neckline region, the body region, and the hem region; then, adaptive cropping is performed based on the garment regions; after cropping, the window is scaled proportionally and filled in the center, and then normalized and standardized.
[0041] Secondly, the present invention also provides a system based on the above method, comprising:
[0042] The image acquisition module is used to acquire images of clothing.
[0043] The preprocessing module is used to preprocess clothing images;
[0044] The feature extraction module inputs the preprocessed clothing image into a deep feature aggregation network of an improved convolutional neural network to extract global fused feature vectors. The deep feature aggregation network has a multi-level deep aggregation layer, a hierarchical feature completion branch, and a cross-layer feature fusion layer;
[0045] The multi-level deep aggregation layer comprises multiple sequentially connected deep feature aggregation blocks; the hierarchical feature completion branch comprises multiple parallel feature completion branches, each feature completion branch being connected one-to-one with each deep feature aggregation block, and the outputs of each feature completion branch are fused through feature concatenation; the cross-layer feature fusion layer performs a unified scale transformation and cross-layer fusion on the output features of the hierarchical feature completion branches and each deep feature aggregation block to obtain a global fused feature vector used for database entry determination. ;
[0046] The feature matching module extracts the global fusion feature vector. The input matching layer performs matching to obtain the style recognition and database entry results;
[0047] The style recognition model consists of a deep feature aggregation network and a matching layer. After the model is pre-trained, it identifies clothing images to be added to the database according to the above process.
[0048] In three aspects, the present invention also provides a computer device, including: one or more processors and a memory storing a computer program;
[0049] The processor invokes a computer program to implement the steps of a method for identifying and storing clothing styles based on an improved convolutional neural network.
[0050] Fourthly, the present invention also provides a computer-readable storage medium storing a computer program that is invoked by a processor to implement the steps of a method for identifying and storing clothing styles based on an improved convolutional neural network.
[0051] Compared with the prior art, the present invention achieves the following progress and effects:
[0052] 1. This invention proposes a deep feature aggregation network suitable for clothing images. It constructs a progressive feature enhancement structure of "intra-block deep aggregation - branch hierarchical completion - unified scale cross-layer feature fusion" through the design of multi-level deep aggregation layers, hierarchical feature completion branches, and cross-layer feature fusion layers, which is used to improve the feature expression capability of clothing images.
[0053] The multi-level deep aggregation layer unifies the representation of features at different depth levels within a convolutional block, thereby enhancing the integrity and robustness of feature representation and enabling effective fusion of features at different depths within the same convolutional block. The hierarchical feature completion branch is an external hierarchical feature completion structure that works in conjunction with the multi-level deep aggregation layer. It is used to capture the top-level global pattern, intermediate-level components, and bottom-level detailed texture features of clothing according to the hierarchical rules of clothing features, thereby improving the network's ability to model targets with different receptive fields. Through this cross-level feature fusion structure, a unified representation of shallow detailed features and deep semantic features can be achieved, thereby improving the overall representation ability of clothing images. Experiments have shown that the network of this invention outperforms the traditional VGG16 network on the FashionMNIST clothing image dataset, reduces the misjudgment rate during automatic data entry, and has high application value.
[0054] 2. The present invention further optimizes the technical solution by proposing image size normalization preprocessing for clothing image characteristics. Specifically, it is implemented through a clothing structure perception adaptive scaling module, which can effectively solve or alleviate the problems in the prior art where image preprocessing only uses simple scaling or proportional filling, resulting in distortion of clothing aspect ratio, subject offset, and inability to distinguish semantic regions such as neckline and body.
[0055] 3. This invention further optimizes the technical solution by proposing a dual-modal joint similarity determination. It simultaneously considers feature similarity and attribute similarity, that is, it considers both the depth features and physical characteristics of the image, greatly improving matching accuracy. Secondly, it also proposes a dual-threshold determination rule and designs a dynamic uncertain half-width. It is determined by the confidence level of image acquisition. Historical fluctuation range Dispersion of style feature distribution The dynamic variables work together to drive the judgment conditions, enabling them to adapt dynamically to the environment and improve matching accuracy. Attached Figure Description
[0056] Figure 1 This is a flowchart illustrating the method provided in an embodiment of the present invention;
[0057] Figure 2 This is a diagram of a deep feature aggregation network architecture;
[0058] Figure 3 This is the architecture diagram of deep feature aggregation block 1;
[0059] Figure 4 This is the architecture diagram of deep feature aggregation block 2;
[0060] Figure 5 This is the architecture diagram of the multi-scale structural symmetry perception module;
[0061] Figure 6 This is a diagram showing the experimental results, specifically a diagram illustrating the accuracy of the matching results.
[0062] Figure 7 This is another experimental result diagram, specifically a schematic diagram of the training loss. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of the invention. The technical features involved in the various embodiments of the invention described below can be combined with each other as long as they do not conflict with each other.
[0064] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than those in the device or the flowchart. The terms "first," "second," etc., used in the specification, claims, and drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used herein is only for describing embodiments of this application and is not intended to limit this application.
[0065] This invention provides a method for clothing style recognition and database entry based on an improved convolutional neural network, comprising the following steps:
[0066] S1. Image acquisition: Acquire clothing images through an image acquisition device.
[0067] S2. Image preprocessing: The acquired clothing image is preprocessed, including at least one or more of image enhancement, image size normalization, pixel value normalization, and noise suppression to obtain a standardized input image. Other feasible embodiments may incorporate other image preprocessing methods, which are not specifically limited in this invention.
[0068] Input a clothing image 1 is: ,in, Indicates the image height. Indicates the image width. Indicates the number of image channels. Represents the real number field.
[0069] In this embodiment, image size normalization is preferably performed using the GarmentStructure-Aware Adaptive Rescaling Module (GSAARM). Addressing the issues in existing technologies where image preprocessing only involves simple scaling or proportional filling, leading to distortion of garment aspect ratio, subject offset, and inability to distinguish semantic regions such as the neckline and body, the GSAARM module converts the original input image (color or grayscale) of any size into a standardized 96×96 grayscale image. Specifically, it performs the following operations:
[0070] (a) Grayscale conversion and edge enhancement:
[0071] If the input clothing image is a color image, the brightness formula is used: Convert a color image to a grayscale image. R, G, and B represent the brightness values, which are the red, green, and blue channel values, respectively. If the input is a grayscale image, it is used directly. Then, the Sobel edge response (horizontal and vertical directions) of the grayscale image is calculated, and the amplitude is taken to obtain the edge intensity map (which can be achieved with existing technology) to highlight the outline and texture of clothing and suppress the smoothing of the background.
[0072] (ii) Horizontal centering based on symmetry correction:
[0073] The edge intensity map is horizontally projected (by summing the pixels in each column) to obtain the projection curve. Determine the projection curve. peak position Then, search to the left and right respectively for the position where the projected value first drops to half of the peak value. and To eliminate horizontal offset caused by unilateral decoration or shooting tilt, set the offset parameter... Correct the horizontal boundary to a symmetrical boundary: left boundary right boundary The width of the main body of the garment is: .
[0074] (III) Semantic partitioning detection based on vertical projection peaks and valleys:
[0075] The edge intensity map is vertically projected (by summing the pixels in each row) to obtain the projection curve. Utilizing the inherent vertical projection shape of clothing images (lower edge response in the neckline area, higher response in the body area, and medium response in the hem area), the following are detected sequentially:
[0076] Peak position of the garment :Pick The row corresponding to the maximum value.
[0077] Neckline end position From the peak position of the garment Search upwards for the first local minimum point (from) Search upwards to find the first point whose projected value is less than the projected values of its two adjacent positions (take this as a local minimum point), and the projected value of this point is less than 50% of the peak value. If no such point is found, then take the local minimum point. ( (Image height).
[0078] hem starting position From the peak position of the garment Search downwards for the first local minimum point (from) Search downwards to find the first point (where its projected value is less than the projected values of its two adjacent locations) or a point with a gentle gradient and a projected value less than 70% of the peak value. If not found, then take: .
[0079] This divides the image into the neckline region (top to...). ), body area ( to ), hem area ( (To the bottom).
[0080] (iv) Content-aware adaptive clipping window:
[0081] To highlight the garment area while preserving necessary context, the vertical boundary of the cropping window is defined as follows:
[0082]
[0083] in, and The vertical boundary position of the vertical clipping window; coefficient. and Both 2 are empirical values used to control the retention ratio of the neckline transition area and the hem transition area. In this embodiment, they are set to 0.15 and 0.1 respectively.
[0084] and constraints , Window height .
[0085] The horizontal boundary of the clipping window is directly adopted from the symmetrical boundary obtained in step (II). , Window width .
[0086] (v) Proportional scaling and center filling:
[0087] Let the standard target size of the network input be D×D, where D is a positive integer with a value range of [32, 256]. In this embodiment, the preferred value is D=96.
[0088] Scaling factor: The scaled width is The height is Then, the scaled image is placed in the center of the D×D canvas with a horizontal offset of . The vertical offset is ,
[0089] Uncovered areas are filled with black (pixel value 0). This transformation samples the original grayscale image using bilinear interpolation.
[0090] (vi) Normalization and standardization:
[0091] Finally, the pixel values of the clothing image are normalized to the [0,1] range, and then... Standardization is performed so that the pixel values of the output image are roughly distributed in the range of [-1,1] to match the input range of the subsequent network. x represents the pixel value normalized to the range of [0,1].
[0092] It should be understood that the GSAARM module proposed in this invention achieves adaptive scaling of garment structure perception in the preprocessing stage. Compared with the existing technology of directly stretching or simply centering and filling, this invention is better able to maintain the original length-to-width ratio of the garment, focus on the core area of the garment body and retain the appropriate context of the neckline and hem, while eliminating the interference of tilt and unilateral decoration through symmetry correction.
[0093] S3. Feature extraction: Input the preprocessed clothing image into the deep feature aggregation network of the improved convolutional neural network.
[0094] The style recognition model based on an improved convolutional neural network includes a deep feature aggregation network to extract the depth feature vector corresponding to the clothing image. For example... Figure 2 The diagram shown is a deep feature aggregation network architecture. The deep feature aggregation network mainly includes hierarchical feature completion branches, multi-level deep aggregation layers, and cross-layer feature fusion layers.
[0095] The multi-level deep aggregation layer consists of multiple depth feature aggregation blocks connected in sequence. Figure 2 The system includes five depth feature aggregation blocks for progressively extracting multi-level features from clothing images. Each depth feature aggregation block combines the output features from different convolution depths with the output features from the multi-scale structural symmetry perception module.
[0096] Let the first The output features of a depth feature aggregation block for:
[0097]
[0098] in, Indicate output features The number of channels, , These represent the output features respectively. The spatial dimensions, namely height and width.
[0099] The hierarchical feature completion branch has multiple parallel feature completion branches, which are connected one-to-one with the deep feature aggregation blocks in the multi-level deep aggregation layer. Each feature completion branch performs convolution, activation function and adaptive pooling in sequence. After the output of each feature completion branch is uniform in spatial size by adaptive average pooling, the features are spliced and fused to obtain the global representation vector, which is then input into the cross-layer feature fusion layer.
[0100] It should be noted that the hierarchical feature completion branch is an external hierarchical feature completion structure that works in conjunction with the multi-level deep aggregation layer. It is used to capture the top-level global pattern, middle-level components, and bottom-level detailed texture features of the garment according to the hierarchical rules of garment features. Its output hierarchical completion features and main features participate in cross-block aggregation.
[0101] Figure 2 In the example shown, the hierarchical feature completion branch has five parallel feature completion branches, specifically two 7×7, one 5×5, and two 3×3 hierarchical convolutions. The 7×7 convolution branch corresponds to the first and second level deep feature aggregation blocks (deep feature aggregation block 1), capturing top-level global features; the 5×5 convolution branch corresponds to the third level deep feature aggregation block of the main path (deep feature aggregation block 2), capturing intermediate-level component features; and the 3×3 convolution branch corresponds to the fourth and fifth level deep feature aggregation blocks of the main path (deep feature aggregation block 2), capturing low-level detailed features.
[0102] The computational process of completing each hierarchical convolution branch is uniformly represented as follows:
[0103]
[0104] Where m is the kernel size, which takes the values of 7, 5, and 3, corresponding to the capture of global, component, and detail-level features of the clothing, respectively; This represents an m×m convolution operation. This indicates that an m×m convolution operation is performed on the i-th level deep feature aggregation block.
[0105] By combining the features of each completed branch in the channel splicing, a global representation vector is obtained. It integrates full-level feature information of the garment, including the overall structure, components, and details. The dimensions are 3*3:
[0106] .
[0107] A cross-layer feature fusion layer is used to complete the global representation vector of the hierarchical feature branches. The output features of each deep feature aggregation block in the multi-level deep aggregation layer are subjected to a unified scale transformation and cross-layer fusion to obtain a global representation vector for database entry determination. Figure 2 In the example shown, the cross-layer feature fusion layer has four adaptive pooling layers, each corresponding to one of the first four deep feature aggregation blocks. Finally, the outputs of the four adaptive pooling layers, the output of the fifth deep feature aggregation block, and the global representation vector of the hierarchical feature completion branch are combined. The features are then concatenated and fused. The fused features are then subjected to 1*1 convolution, batch normalization, and activation function to obtain the output features.
[0108] Figure 2 In this example, let the output features of the depth feature aggregation block within the five-level block be as follows: ;
[0109] First, adaptive average pooling is performed on the output features at each level to unify the spatial size of all feature maps to a preset scale, preferably 1. ,Right now:
[0110]
[0111] Then, the features at each level after being standardized are concatenated along the channel dimension to obtain the features. :
[0112]
[0113] pass Convolution is used to reduce dimensionality, ultimately yielding a globally fused feature vector. :
[0114]
[0115] This cross-level feature fusion structure enables a unified expression of shallow detail features and deep semantic features, thereby improving the overall representation capability of clothing images.
[0116] It should be understood that in other feasible embodiments, network architectures that can achieve unified scale transformation and cross-layer fusion of fusion features of hierarchical feature completion branches and output features of each deep feature aggregation block are also considered to fall within the protection scope of this invention and meet the technical requirements of this invention.
[0117] In summary, the deep feature aggregation network proposed in this invention constitutes a progressive feature enhancement structure of "intra-block deep aggregation - branch hierarchical completion - unified scale cross-layer feature fusion", which is used to improve the feature representation capability of clothing images.
[0118] In some embodiments, to improve the network feature extraction effect, some deep feature aggregation blocks in the multi-level deep aggregation layer are deep feature aggregation blocks 1 (e.g., Figure 3 ), some deep feature aggregation blocks are deep feature aggregation blocks 2 (such as Figure 4 ). Figure 2 In the example, the first two deep feature aggregation blocks are called deep feature aggregation block 1; the last three deep feature aggregation blocks are called deep feature aggregation block 2.
[0119] like Figure 3 As shown, the deep feature aggregation block 1 contains two consecutively stacked convolutional layers and a multi-scale structural symmetry perception module parallel to the convolutional layers. The outputs of all convolutional layers and the output of the multi-scale structural symmetry perception module are aggregated within the block, and then... After convolution and pooling layers, the output features of depth feature aggregation block 1 are obtained. Each convolutional layer executes the following sequentially: Convolution—Normalization—Activation Function;
[0120] definition: Indicates the first Output features of layer convolution, Figure 3 The deep feature aggregation block 1 shown has two stacked convolutional layers, with k values of 1 and 2 respectively.
[0121] When k=1, the output features of the first convolutional layer Represented as: ;
[0122] In the formula, This can be either the initial feature input or the output feature of the previous level deep feature aggregation block. express Convolution operation, This indicates a batch normalization operation. This represents a non-linear activation function.
[0123] Figure 3 When k=2, the output features of the second convolutional layer Represented as: ;
[0124] By setting up multiple convolutional layers, edge information, texture information, and high-level semantic information in the input image can be extracted step by step.
[0125] Then, the outputs of the two convolutional layers and the output features of the multi-scale structural symmetry perception module are aggregated within a block, that is, concatenated along the channel dimension to obtain the aggregated features. :
[0126]
[0127] This represents the output features of the multi-scale structural symmetry perception module.
[0128] Aggregation features pass Convolution performs channel compression to obtain features. :
[0129]
[0130] Finally, features Use a pooling layer to reduce the output feature size:
[0131]
[0132] In the formula, the size of the pooling layer is taken as 2*2. This is the output of deep feature aggregation block 1.
[0133] Through the aforementioned deep feature aggregation block 1, features at different depth levels within the convolutional block can be uniformly expressed, thereby enhancing the integrity and robustness of the feature representation.
[0134] Different from Figure 3 The depth feature aggregation block 1 shown is as follows: Figure 4 As shown, the deep feature aggregation block 2 has three consecutively stacked convolutional layers and a multi-scale structural symmetry perception module parallel to the convolutional layers. The output of each convolutional layer is aggregated within the block with the output of the multi-scale structural symmetry perception module, and then... After convolution and pooling layers, the output features of depth feature aggregation block 2 are obtained. Each convolutional layer executes the following sequentially: Convolution—Normalization—Activation Function;
[0135] For the deep feature aggregation block 2, k takes values of 1, 2, and 3 respectively; when k > 1, the operation process of the convolutional layer can be represented as:
[0136]
[0137] In the formula, , These are the output features of the k-th and (k-1)-th convolutional layers. The processing of other network layers is the same as in depth feature aggregation block 1, and will not be repeated here.
[0138] It should be understood that by using the above network design to splice and fuse the outputs of convolutional layers of different depths within the same convolutional block with the outputs of the multi-scale structural symmetry perception module, the problem of insufficient expression of clothing image features in traditional feature extraction network applications can be effectively solved.
[0139] like Figure 5 As shown, the multi-scale structural symmetry perception module has multiple parallel symmetry perception convolutional branches and adaptive channel adjustment branches. Finally, the output features of each branch are concatenated and fused. Convolution yields the output features of the multi-scale structural symmetry-sensing module. Multiple parallel symmetry-sensing convolutional branches are used to extract image features under different receptive fields. Each symmetry-sensing convolutional branch sequentially performs symmetry-sensing convolution, normalization, and activation functions. The kernel size of the symmetry-sensing convolutions on different symmetry-sensing convolutional branches is different. Figure 5 In the example shown, there are three parallel symmetric-sensory convolution branches, as detailed below:
[0140] First symmetric perceptual convolution branch:
[0141]
[0142] Second symmetric perceptual convolution branch:
[0143]
[0144] Third symmetric perceptual convolution branch:
[0145]
[0146] In the formula, This serves as the input for the multi-scale structural symmetry sensing module. , , These are the output features of the three symmetric perceptual convolution branches, and their symmetric perceptual convolution kernel sizes are as follows: , , .
[0147] The convolution kernel used in the symmetry-sensing convolution of this invention is not an ordinary convolution kernel, but a structural symmetry convolution kernel designed based on the left-right symmetry or near-symmetry structure of most clothing. The weights of this convolution kernel are constrained to a mirror symmetry form, as shown in the following formula:
[0148]
[0149] in, The width of the convolution kernel ( Figure 5 In the example shown, the values are 5, 3, and 1). This represents the height of the convolution kernel. , Given the original convolutional kernel weight tensor without symmetric constraints, output channel index. Input channel index Convolution kernel inline index (from top to bottom) ; Convolution kernel inline column indexes (from left to right) ; This is the weight tensor actually used for convolution after symmetric constraints.
[0150] Based on the above convolution kernel design, for any symmetric perceptual convolution branch (kernel size) Step length ,filling Output feature map ( In position value at (No. The channel is:
[0151]
[0152] in, This is a bias term.
[0153] By employing this forced mirror symmetry weighting method, regardless of whether the input image has angular tilt or local deformation, this branch can produce the same response to symmetrical parts (such as left and right shoulders, left and right cuffs), thereby effectively reducing feature biases introduced by asymmetric interference factors such as shooting angle, lighting direction, or unilateral folds. Simultaneously, this symmetry constraint guides the network to pay more attention to the symmetry or structural features that truly define the garment style, such as neckline shape, collar type, and pocket position.
[0154] It should be noted that this symmetry-aware convolutional branch is only a parallel component within the multi-scale structural symmetry-aware module. For a very small number of asymmetrical clothing designs (e.g., off-shoulder styles), other standard convolutional branches in the network (such as those in the main path depth feature aggregation block) are used. Convolutional layers and residual connection paths can still freely learn asymmetric features. Therefore, the multi-scale structural symmetry perception module fully utilizes symmetric priors to improve the stability of most clothing recognition while not compromising the network's adaptability to a few exceptional styles.
[0155] Furthermore, the input feature X of the multi-scale structural symmetry perception module enters the adaptive channel adjustment branch, that is, it adjusts the feature according to whether the number of input and output channels are equal, and the feature after the feature is... :
[0156]
[0157] in, Input the number of channels. The number of output channels for the module.
[0158] Finally, the outputs of the three symmetric perceptual convolution branches and the output of the adaptive channel adjustment branch are concatenated along the channel dimension to obtain the features. :
[0159]
[0160] It should be understood that when the channel dimensions of the input features and the fused features are inconsistent, adaptive methods can be used. Convolution performs dimension matching on the input features and then performs an addition operation.
[0161] Subsequently passed Convolutional dimensionality reduction yields the output. :
[0162]
[0163] This invention enhances the network's ability to model targets at different scales through the multi-scale structural symmetry perception module, enabling it to simultaneously take into account both the overall outline information of clothing and local texture details.
[0164] S4. Feature matching: The global fusion feature vector extracted in step S3 is used to match the feature vectors obtained in step S3. The matching layer of the input style recognition model is used for matching.
[0165] Among them, the style recognition model based on the improved convolutional neural network also includes a matching layer.
[0166] The current method of determining whether clothing items are the same in the database relies solely on the similarity of a single deep feature vector, which has obvious drawbacks: the feature vectors of the same clothing items can vary greatly due to changes in shooting angle, lighting, and posture, resulting in different images of the same item being mistakenly identified as new items; and the feature vectors of similar clothing items with similar patterns and appearances are highly overlapping, resulting in similar items being mistakenly identified as the same item.
[0167] Among them, the images to be added to the database include those uploaded by users in this instance and those currently in the process of being added. Clothing images for secondary similarity determination.
[0168] Candidate styles The existing styles in the database that have the highest similarity to the image to be added to the database.
[0169] Style Database: All styles registered in the system A collection of clothing styles and their sample images.
[0170] The matching layer of this invention adopts a dual-modal joint similarity calculation architecture based on deep features and clothing-specific 6-dimensional structured attributes (color, collar type, sleeve type, placket, pattern, and pattern). Specifically, it includes feature similarity and attribute similarity, as detailed below:
[0171]
[0172] For feature similarity; The global fusion feature vector is extracted from the clothing image to be imported into the database through the deep feature aggregation network in step S3. , The global fusion feature vector is extracted from the inventory image of the style database through step S3 deep feature aggregation network. .
[0173] Attribute similarity calculation:
[0174] We select the existing pre-trained EfficientNet-B3 as the base network and add multiple classification heads on top of the base network to predict the attribute labels of the image to be added to the database (each attribute corresponds to a classification head).
[0175] For the design of each attribute header, single-selection attributes (collar type, sleeve type, placket, pattern) and multi-selection attributes (color, pattern) are proposed:
[0176] When attribute 'a' is a single-selection attribute, attribute similarity :
[0177]
[0178] When attribute 'a' is a multi-select attribute, let the set of predicted labels for the clothing images to be imported into the database be... The inventory image in the style database is as follows Attribute similarity obtained using the Jaccard similarity coefficient :
[0179]
[0180] For example, if the predicted color is {black, white} and the inventory color is {black, gray}, the intersection is {black} and the union is {black, white, gray}, then:
[0181] .
[0182] If both sides of a certain attribute are empty (no tag), define (Missing information is considered non-conflictible). For example, a white shirt without a pattern will not have a pattern tag, and both the pattern tag in the database and the one to be added to the database will be empty.
[0183] Weighted attribute similarity total score (Final attribute similarity):
[0184]
[0185] In this process, weights are assigned to each attribute. The weights are fixed weights based on the importance of the attributes. In this embodiment, the weights are set as follows: color 0.25, collar type 0.20, pattern type 0.20, sleeve type 0.10, placket 0.10, and the sum of the weights is 1.
[0186] Ultimately, the dual-modal joint score for:
[0187]
[0188] in, The coefficient is a preset value, such as 0.8 in this embodiment. In other feasible embodiments, it can be adjusted appropriately according to the scenario.
[0189] In some embodiments, a single threshold determination is adopted, that is, a single threshold determination threshold is set. When the dual-modal joint score Greater than or equal to the single threshold judgment threshold When the image of the clothing to be added to the database is determined to correspond to an existing clothing style in the database, the maximum similarity value is less than the single threshold for judgment. At that time, it is determined that the image of the clothing to be put into the warehouse corresponds to a new clothing style.
[0190] In some embodiments, a dual threshold determination is employed:
[0191]
[0192] In the formula, The upper bound of the threshold, This is the lower bound of the threshold. The central threshold represents the baseline strictness of "whether they are the same style" and is an empirical value, such as... ; The indeterminate half-width represents the degree of ambiguity in the machine's judgment, and In some embodiments, W may be an empirical value, which is optimized in some embodiments and designed as a dynamic variable driven by three dynamic variables normalized to [0,1]:
[0193] Image acquisition confidence The score is based on the current acquisition conditions, taking into account factors such as the image's clarity and illumination uniformity. The output is in the range [0,1]. A higher value indicates better image quality. The calculation formula is as follows:
[0194]
[0195] in: It is used to limit the output to between 0 and 1. : The Laplacian variance of the grayscale image to be entered into the database, used to measure sharpness. k represents the clothing image being identified in the kth entry into the database. For example, each matching identification identifies one or a batch of clothing images; this invention does not impose specific limitations on this.
[0196] (Brightness variation coefficient): Divides the image into... Non-overlapping grid blocks (such as or Calculate the average brightness for each block. (Using grayscale mean); calculate this The average of each brightness value and standard deviation :
[0197]
[0198] Brightness variation coefficient Defined as:
[0199]
[0200] To avoid When the division by zero occurs, When (extremely dark graph), directly set this item to The more uneven the lighting, the greater the difference in brightness between the different areas. The higher the score, the lower the rating. (Acceptable maximum coefficient of variation): Set according to the situation; in this embodiment... .
[0201] Dispersion of style feature distribution : For the candidate styles in the database corresponding to the current maximum similarity Assume there are candidate styles in the library. Normalized eigenvectors (Regarding candidate styles) This indicates that this style has a total of [number] instances in the database. Zhang sample images, normalized feature vectors It is a feature extracted by a deep feature aggregation network, and its center Then the intra-class dispersion Defined as:
[0202]
[0203] Retrieve all styles from the style database maximum value Normalization yields:
[0204]
[0205] If the number of style samples is less than 2, then take the neutral value. .
[0206] Historical fluctuation range : Targeting candidate styles ,recent Secondary judgment (for candidate styles) Maintain a fixed capacity ( The system uses a sliding window queue. Each time the system completes a judgment (whether automatically or after manual intervention), if the final conclusion is "belongs to style..." Then the original maximum similarity value will be used. Add the item to the end of the queue for that style; if the queue is full, remove the earliest record from the front of the queue. Confirm similarity values for the same style within this queue. Calculate its standard deviation :
[0207]
[0208] Then divide this standard deviation by the theoretical maximum value. Normalization:
[0209]
[0210] when When the standard deviation cannot be calculated effectively, the default value is used. .
[0211] Based on the above variables, the halfwidth is uncertain. The calculation formula is:
[0212]
[0213] in: Ensure an interval width of at least 0.1. ,make Roughly Between these, the corresponding width of the manual intervention interval is approximately It should be understood that the above coefficient values are for illustrative purposes.
[0214] The worse the image quality ( Small), The larger the area, the wider the gray area, and more cases are handled manually; the more dispersed the style features ( big), The larger the value, the lower the model's judgment power and the greater the uncertainty interval; the greater the fluctuation in historical similarity ( big), The larger the image, the less stable the shooting angle and lighting become.
[0215] Final dual thresholds:
[0216]
[0217] Dual threshold determination mode: using the first determination threshold ( ) and the second judgment threshold ( ) for the maximum similarity value (bimodal joint score) ) for comparison; when When the image is greater than or equal to the first threshold, it is determined that the clothing image to be added to the database corresponds to an existing clothing style; when When the value is less than or equal to the second threshold, the clothing image to be added to the database is determined to correspond to a new clothing style; when... When the image falls between the first and second judgment thresholds, the image of the garment to be put into storage is determined to enter the manual intervention judgment process.
[0218] S6. Automatic Entry into Database. When a new clothing style is identified, a unique identifier corresponding to the new style is generated, and the clothing image to be entered into the database, the corresponding feature vector, and the associated information are written into the clothing database to achieve automatic filing and entry of the new clothing style into the database. When a clothing style is identified as an existing style in the database, the clothing image to be entered into the database is associated with the existing clothing style record with the highest similarity, and the statistical information corresponding to the existing clothing style record is updated.
[0219] Example 1: Determined to be a new style
[0220] The system captured a brand new white loose shirt, whose characteristics and similarity to all of Curry's styles were below the threshold.
[0221] Automatically generate a unique ID: STYLE_003;
[0222] Store the shirt image, feature vector, and information such as acquisition time, acquisition device (CAM_005), and storage path into the database to complete the new product filing.
[0223] Example 2: Determined to be an existing style
[0224] The system captured a black hooded sweatshirt with a similarity of 0.96 to Curry's STYLE_001 (exceeding the threshold).
[0225] Link this new hoodie image to the STYLE_001 record;
[0226] Updated statistics: The frequency of the same product has increased from 12 times to 13 times, the latest data collection time has been updated, and "Store C" has been added to the data collection scenario.
[0227] This method enables the automatic updating and management of clothing image databases.
[0228] Network training:
[0229] Dataset: In this embodiment, both the training and testing datasets used are FashionMNIST, with 5 / 6 used as the training set and 1 / 6 as the testing set. The DataLoader in PyTorch is used to construct iterators for the training and testing sets, and training is executed on either a GPU or CPU device. The program first automatically selects the computing device based on the runtime environment, prioritizing CUDA-accelerated devices; when CUDA devices are unavailable, it degenerates to CPU execution. Random shuffling is set during training set loading to improve the randomness of sample traversal; the order is not shuffled during test set loading to ensure the consistency of evaluation results.
[0230] Training parameters:
[0231] Batch size: 64; Number of training epochs: 30; Initial learning rate: 0.01; Learning rate decay step size: 10; Learning rate decay factor (gamma): 0.5; Optimizer: SGD; Momentum coefficient: 0.8; Weight decay coefficient: 0.001.
[0232] Loss function:
[0233] In clothing images, the main subjects such as tops, trousers, and dresses are usually located in the vertical center of the image, while the top and bottom edges are often the background (ground, walls, clutter, etc.). The purpose of this loss function is to guide the network to "actively" focus its attention on the main clothing area in the center of the image, based on the standard cross-entropy loss function, thereby improving the robustness of style classification.
[0234] Therefore, this invention employs a joint loss function. End-to-end training consists of two parts: a standard cross-entropy loss function and a multi-level feature activation centroid centering loss.
[0235]
[0236] in, For standard cross-entropy loss, To balance the hyperparameter (0.01 in this embodiment); The centroid centering loss is activated for multi-level features and is defined as follows:
[0237] Let the first The feature map output by the deep feature aggregation block is ( ), , These represent the height and width of the feature map at this level, respectively. Let be the number of channels. For each level of feature map, calculate the weighted average centroid of the absolute values of all activations in the vertical direction. :
[0238]
[0239] in, This is the row index in the vertical direction. To prevent division by zero, 'c' corresponds to the channel, and 'w' is the row index. The mean square error is calculated by averaging the normalized centroid with the target centroid at 0.5. :
[0240]
[0241] The final multi-level center-of-gravity loss is obtained by averaging all five levels of loss. :
[0242]
[0243] This loss term simultaneously constrains the feature maps at different levels in the network, prompting the focus from shallow edge textures to deep semantic abstractions on the clothing subject area in the center of the image, suppressing background noise at the top and bottom edges, and is applicable to various types of clothing such as tops, trousers, and dresses.
[0244] Learning rate scheduling strategy:
[0245] This embodiment employs a StepLR learning rate update strategy, where the learning rate is decayed to 0.5 times its original value after every 10 epochs of training. Let the... The learning rate during round training is Then we have:
[0246]
[0247] in, This strategy can maintain a large learning rate in the early stages of training to accelerate convergence, and gradually reduce the learning rate in the later stages of training to improve the model's convergence accuracy.
[0248] In summary, the deep feature aggregation network proposed in this embodiment effectively fuses different depth features within the same convolutional block by introducing an intra-block deep aggregation mechanism into the backbone network. By introducing a multi-scale structural symmetry perception module, the network's ability to model targets with different receptive fields is improved, while suppressing interference from clothing angle tilt. The hierarchical feature completion branch accurately captures and completes all layers of clothing features, effectively compensating for the loss of deep details in the main network, further reducing the misjudgment rate of fine-grained clothing recognition and improving the accuracy of automatic data entry. Through a cross-block feature aggregation mechanism, shallow detail features and deep semantic features are further integrated. Experimental results are as follows: Figure 6 and Figure 7 As shown in the figure, experiments have verified that the network described in this invention outperforms the traditional VGG16 network on the FashionMNIST clothing image dataset, and can reduce the misjudgment rate during the automatic data entry process, thus having high application value.
[0249] In some embodiments, the present invention also provides a system based on the above method, including an image acquisition module, a preprocessing module, a feature extraction module, and a feature matching module connected in sequence or to each other.
[0250] The image acquisition module is used to acquire images of clothing; the preprocessing module is used to preprocess the images of clothing.
[0251] The feature extraction module inputs the preprocessed clothing image into a deep feature aggregation network of an improved convolutional neural network to extract global fused feature vectors. The deep feature aggregation network has a multi-level deep aggregation layer, a hierarchical feature completion branch, and a cross-layer feature fusion layer;
[0252] The multi-level deep aggregation layer comprises multiple sequentially connected deep feature aggregation blocks; the hierarchical feature completion branch comprises multiple parallel feature completion branches, each feature completion branch being connected one-to-one with each deep feature aggregation block, and the outputs of each feature completion branch are fused through feature concatenation; the cross-layer feature fusion layer performs a unified scale transformation and cross-layer fusion on the output features of the hierarchical feature completion branches and each deep feature aggregation block to obtain a global fused feature vector used for database entry determination. .
[0253] The feature matching module extracts the global fusion feature vector. The input matching layer performs matching to obtain the style recognition and database entry results; the style recognition model consists of a deep feature aggregation network and a matching layer. After model pre-training, it identifies the clothing images to be entered into the database according to the above process.
[0254] It should be understood that the specific implementation process of each module is described in the above method. This invention will not repeat the details here. The above division of functional modules is only for illustrative purposes. In some embodiments, some functional modules can be combined and some functional modules can be separated. Each functional module can be implemented in software, hardware, or a combination of software and hardware. The software and hardware devices include, but are not limited to, general-purpose computer equipment, programmable gate arrays, digital signal processors, microprocessors and their corresponding programming or burning software.
[0255] In some embodiments, the present invention also provides a computer device, including: one or more processors and a memory storing a computer program;
[0256] The processor invokes a computer program to implement the steps of a method for identifying and storing clothing styles based on an improved convolutional neural network.
[0257] Please refer to the explanation of the method above for the specific implementation process of each step.
[0258] It should be understood that, in the embodiments of the present invention, the processor may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor. The memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of the memory may also include non-volatile random access memory. For example, the memory may also store device type information.
[0259] In some embodiments, the present invention also provides a computer-readable storage medium storing a computer program that is invoked by a processor to implement the steps of a method for identifying and storing clothing styles in a database based on an improved convolutional neural network.
[0260] Please refer to the explanation of the method above for the specific implementation process of each step.
[0261] The computer-readable storage medium can be an internal storage unit of the hardware and software device in any of the foregoing embodiments (e.g., the hard disk or memory of the controller), or it can be an external storage device of the controller, such as a plug-in hard disk, smart memory card (SMC), secure digital card (SD card), flash memory card, etc., equipped on the controller. Furthermore, the storage medium can also simultaneously include both the controller's internal storage unit and external storage device.
[0262] Based on the above understanding, the core part of the technical solution of this invention that contributes to the prior art, or all or part of the content of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and contains several instructions to cause a computer device (such as a personal computer, server, or network device) to execute all or part of the steps of the methods described in the various embodiments of this invention. Available storage media include, but are not limited to: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0263] Those skilled in the art will understand that embodiments of this application can be provided in the form of a method, system, or computer program product. Therefore, this application can be implemented entirely in hardware, entirely in software, or a combination of hardware and software. Furthermore, this application can also be implemented as a computer program product containing computer-usable program code on a computer-readable storage medium (such as a disk storage device, CD-ROM, optical storage, etc.). Embodiments of this application are described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products. It should be understood that the function of each step in the flowchart and / or each block in the block diagram can be implemented by computer program instructions. These computer program instructions can be executed by a processor of a general-purpose computer, special-purpose computer, or other programmable data processing device to generate means for implementing the functions specified in the flowchart and / or block diagrams. These instructions can also be stored in a computer-readable storage medium to cause a computer or other programmable device to operate in a particular manner, thereby producing an article of manufacture containing instruction means to implement the functions specified in the flowchart and / or block diagrams. Furthermore, the instructions can be loaded onto a computer or other programmable device to form a computer-implemented processing flow by performing a series of operational steps, so that the instructions executed on the computer or other programmable device can implement the functional steps specified in the flowchart and / or block diagrams.
[0264] It should be emphasized that the examples described in this invention are illustrative rather than limiting. Therefore, this invention is not limited to the examples described in the specific embodiments. Any other embodiments derived by those skilled in the art based on the technical solutions of this invention, without departing from the spirit and scope of this invention, whether modifications or substitutions, are also within the protection scope of this invention.
Claims
1. A method for clothing style recognition and database entry based on an improved convolutional neural network, characterized in that: Includes the following steps: Image acquisition and preprocessing: Acquire clothing images and perform preprocessing; Feature extraction involves inputting the preprocessed clothing image into a deep feature aggregation network of an improved convolutional neural network to extract a global fusion feature vector. The deep feature aggregation network has a multi-level deep aggregation layer, a hierarchical feature completion branch, and a cross-layer feature fusion layer; The multi-level deep aggregation layer comprises multiple sequentially connected deep feature aggregation blocks; the hierarchical feature completion branch comprises multiple parallel feature completion branches, each feature completion branch being connected one-to-one with each deep feature aggregation block, and the outputs of each feature completion branch are fused through feature concatenation; the cross-layer feature fusion layer performs a unified scale transformation and cross-layer fusion on the output features of the hierarchical feature completion branches and each deep feature aggregation block to obtain a global fused feature vector used for database entry determination. ; The deep feature aggregation block performs intra-block joint aggregation of output features from different convolutional depths with the output features from the multi-scale structural symmetry perception module; that is, the deep feature aggregation block contains two or three consecutively stacked convolutional layers and a multi-scale structural symmetry perception module parallel to the convolutional layers. The outputs of all convolutional layers and the outputs of the multi-scale structural symmetry perception module are aggregated intra-block, and then... The output features of the depth feature aggregation block are obtained after convolution and pooling layers; The multi-scale structural symmetry perception module includes multiple parallel symmetry perception convolutional branches with different kernel sizes and adaptive channel adjustment branches. Finally, the output features of each branch are concatenated and fused. Convolution is used to obtain the output features of the multi-scale structural symmetry perception module; For any symmetric perceptual convolution branch, define the kernel size. Step length ,filling Position of the feature map output by the symmetry-aware convolution branch value at for: ; in, For the feature map, the first The channel corresponds to the location The value at that location, For bias terms, For output channel index, This is the maximum value of the output channel. v is the input channel index, and v is the column index within the convolution kernel. This is the weight tensor actually used for convolution after symmetric constraints; This represents the input features of the symmetry-aware convolution branch; Represents the row and column position parameters of the feature map; Feature matching will extract the global fusion feature vector. The input matching layer performs matching to obtain the style recognition and database entry results; The style recognition model consists of a deep feature aggregation network and a matching layer. After the model is pre-trained, it identifies clothing images to be added to the database according to the above process.
2. The method according to claim 1, characterized in that: The multi-scale structural symmetry sensing module has three parallel symmetry sensing convolution branches, with the following kernel sizes: , , .
3. The method according to claim 1, characterized in that: The hierarchical feature completion branch has 5 parallel feature completion branches, specifically two 7×7, one 5×5, and two 3×3 hierarchical convolutions; Each feature completion branch performs convolution, activation function, and adaptive pooling operations sequentially.
4. The method according to claim 1, characterized in that: The matching layer of the style recognition model introduces a dual-modal joint similarity determination, which includes feature similarity and attribute similarity. The feature similarity is based on a globally fused feature vector. The similarity is represented by the attribute similarity, which is composed of a weighted average of the similarity of multiple attributes. The attribute types include all or part of the following: color, collar type, sleeve type, placket, pattern, and pattern. The model for the bimodal joint similarity is expressed as follows: ; in, For the joint score of the two modes, For preset coefficients, The total score for attribute similarity; The total score for feature similarity; The global fusion feature vector extracted from clothing images to be imported into the database using a deep feature aggregation network. , The global fusion feature vector extracted from the inventory image of the style database using a deep feature aggregation network. .
5. The method according to claim 1, characterized in that: The bimodal joint similarity determination is based on the bimodal joint similarity and employs a dual-threshold determination method, specifically: ; In the formula, The upper bound of the threshold, This is the lower bound of the threshold. The central threshold represents the baseline strictness of "whether they are the same model"; The indeterminate half-width represents the degree of ambiguity in the machine's judgment, and ;k represents the clothing image identified in the kth entry into the database; Uncertain halfwidth Driven by three dynamic variables normalized to [0,1]: image acquisition confidence Historical fluctuation range Dispersion of style feature distribution : Uncertain halfwidth The calculation formula is: ; Where: Uncertain half-width base value , The image unreliability coefficient. Style dispersion coefficient This represents the historical volatility coefficient.
6. The method according to claim 1, characterized in that: The preprocessing includes: introducing a clothing structure-aware adaptive scaling module to normalize the size of the clothing image, that is, performing the following in sequence: grayscale conversion and edge enhancement, horizontal centering based on symmetry correction, semantic partitioning detection based on vertical projection peaks and valleys, content-aware adaptive cropping window, proportional scaling and center filling, normalization and standardization. The grayscale and edge enhancement operations yield an edge intensity map; the edge intensity map is then corrected by a horizontal centering operation based on symmetry correction to correct the horizontal boundaries; then, semantic partitioning detection based on vertical projection peaks and valleys identifies each garment region, namely the neckline region, the body region, and the hem region; then, adaptive cropping is performed based on the garment regions; after cropping, the window is scaled proportionally and filled in the center, and then normalized and standardized.
7. A system based on the method of any one of claims 1-6, characterized in that: include: The image acquisition module is used to acquire images of clothing. The preprocessing module is used to preprocess clothing images; The feature extraction module inputs the preprocessed clothing image into a deep feature aggregation network of an improved convolutional neural network to extract global fused feature vectors. The deep feature aggregation network has a multi-level deep aggregation layer, a hierarchical feature completion branch, and a cross-layer feature fusion layer; The multi-level deep aggregation layer comprises multiple sequentially connected deep feature aggregation blocks; the hierarchical feature completion branch comprises multiple parallel feature completion branches, each feature completion branch being connected one-to-one with each deep feature aggregation block, and the outputs of each feature completion branch are fused through feature concatenation; the cross-layer feature fusion layer performs a unified scale transformation and cross-layer fusion on the output features of the hierarchical feature completion branches and each deep feature aggregation block to obtain a global fused feature vector used for database entry determination. ; The feature matching module extracts the global fusion feature vector. The input matching layer performs matching to obtain the style recognition and database entry results; The style recognition model consists of a deep feature aggregation network and a matching layer. After the model is pre-trained, it identifies clothing images to be added to the database according to the above process.
8. A computer device, characterized in that: include: One or more processors; And the memory that stores computer programs; The processor invokes a computer program to achieve the following: The steps of the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that: The computer program is stored and is invoked by the processor to implement: The steps of the method according to any one of claims 1-6.