A clothing picture intelligent screening method, system, device and medium
By employing an intelligent clothing image filtering method that combines multi-dimensional feature extraction and dynamic weighted calculation, the problem of low efficiency in traditional clothing image processing is solved. This enables end-to-end intelligent filtering, improving the efficiency and accuracy of clothing image filtering and supporting business decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHIYI TECH
- Filing Date
- 2026-03-27
- Publication Date
- 2026-06-26
AI Technical Summary
Existing methods for processing clothing images are inefficient, rely on manual screening, and are greatly affected by subjective factors. They are difficult to standardize and ensure the objectivity of results, cannot accurately understand the deep semantic information of clothing, and lack an end-to-end multi-dimensional intelligent screening framework. As a result, the screening results are not comprehensive or accurate enough to support business decisions.
An intelligent screening method for clothing images is adopted, which achieves end-to-end intelligent screening by acquiring datasets, basic quality filtering, image enhancement, multi-dimensional visual feature extraction, clothing region segmentation, semantic parsing of design elements, trend matching, and innovation measurement, combined with dynamic weighted calculation to generate a comprehensive score value.
It achieves end-to-end intelligent filtering from low-level pixel processing to high-level business intelligence decision-making, improving the efficiency and accuracy of clothing image filtering, outputting structured results that can directly support business decisions, and enhancing the intelligence level of style selection, trend analysis, and design inspiration mining.
Smart Images

Figure CN122289716A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer vision and artificial intelligence, and in particular to a method, system, device and medium for intelligent screening of clothing images. Background Technology
[0002] With the rapid development of e-commerce and the widespread adoption of online shopping, the proportion of apparel in online transactions continues to rise, generating massive amounts of apparel image data daily across major e-commerce platforms, fashion brands, and social media. Simultaneously, the rapid iteration of the fashion industry means that clothing styles, design elements, and trends change in an instant, making the need for efficient management, selection, and utilization of these image resources increasingly urgent for businesses. Apparel images are no longer merely a means of product display; they have become crucial carriers of design information, market trends, and consumer preferences. Against this backdrop, the ability to quickly and accurately select images with visual appeal, market potential, and design value from a vast and heterogeneous database of apparel images has become a vital step in improving the operational efficiency and market competitiveness of apparel companies.
[0003] However, existing methods for processing clothing images still have many limitations. Traditional screening relies heavily on manual processes, which is not only inefficient but also highly susceptible to subjective factors, making it difficult to guarantee standardized and objective results. Some content-based image retrieval technologies attempt to assist screening by analyzing low-level visual features such as color, texture, and shape, but these methods often fail to accurately understand the deep semantic information of clothing (such as design elements and style details), are less robust to images with complex backgrounds, occlusions, or taken from multiple angles, and rarely comprehensively consider the clothing's fit with fashion trends and design innovation. Furthermore, although some research has attempted to combine deep learning for clothing recognition or attribute analysis, it mostly focuses on single tasks, such as image classification or retrieval, lacking an end-to-end automated screening framework that can collaboratively evaluate image quality, trend matching, and design novelty. This fragmented processing approach cannot meet the needs of multi-dimensional, intelligent, and large-scale screening of clothing images in practical applications, resulting in screening results that are often incomplete or inaccurate and cannot effectively support business decisions. Summary of the Invention
[0004] To address the aforementioned technical issues, this application provides a method, system, device, and medium for intelligent screening of clothing images.
[0005] Firstly, this application provides an intelligent screening method for clothing images, employing the following technical solution:
[0006] A smart filtering method for clothing images, the filtering method comprising:
[0007] Obtain the dataset of clothing images to be processed;
[0008] Basic quality filtering is performed on the clothing image dataset to generate a set of filtered images based on preset legality rules and image quality thresholds;
[0009] Image enhancement processing is performed on the selected image set to generate a standardized image set;
[0010] Multi-dimensional visual feature extraction is performed on the standardized image set to generate a visual feature vector and quality score for each image;
[0011] Based on the aforementioned visual feature vectors, the main body region of the clothing is identified through a pre-trained clothing region segmentation model, and a clothing region mask is generated.
[0012] By combining the clothing area mask with the visual feature vector, semantic tags of clothing design elements are extracted to generate a set of structured design points;
[0013] Obtain a pre-built clothing trend library, perform similarity matching between the structured design point set and the clothing trend library, and generate a trend matching score;
[0014] The design innovation degree is calculated based on the historical occurrence frequency of the structured design point set, and an innovation score is generated.
[0015] The quality score, trend matching score, and innovation score are dynamically weighted and calculated to generate a comprehensive score.
[0016] Based on the comprehensive score and the quality score, a grading decision is made, and the images to be added to the database and their grade labels are output.
[0017] By adopting the above technical solutions, end-to-end intelligent screening of massive amounts of clothing images, from low-level pixel processing to high-level business intelligence decision-making, has been achieved. First, standardized preprocessing and deep learning-based multi-dimensional feature extraction ensure the robustness and accuracy of the analysis process. Second, the innovative combination of clothing region segmentation, design point semantic parsing, dynamic trend matching, and innovation quantification ensures that the screening criteria simultaneously cover three key dimensions: visual aesthetics, market trends, and design originality, far exceeding simple image quality filtering or tag matching. Finally, through a dynamic weighted fusion model and two-layer hierarchical decision-making, structured results that directly support business decisions are output, significantly improving the efficiency and intelligence of applications such as clothing e-commerce selection, fashion trend analysis, and design inspiration mining. This transforms traditional processes relying on human experience into scalable and iterative data-driven processes.
[0018] Secondly, this application provides an intelligent image filtering system for clothing, employing the following technical solution:
[0019] A smart filtering method for clothing images, the filtering method comprising:
[0020] The data acquisition module is used to acquire the dataset of clothing images to be processed.
[0021] The preprocessing module is used to perform basic quality filtering operations on the clothing image dataset, and generate a set of filtered images based on preset legality rules and image quality thresholds;
[0022] The standardization processing module is used to perform image enhancement processing on the selected image set to generate a standardized image set;
[0023] The feature extraction and evaluation module is used to extract multi-dimensional visual features from the standardized image set and generate a visual feature vector and quality score for each image.
[0024] The main body segmentation module is used to identify the main body region of the clothing based on the visual feature vector using a pre-trained clothing region segmentation model, and generate a clothing region mask.
[0025] The design element parsing module is used to combine the clothing area mask with the visual feature vector to extract semantic tags of clothing design elements and generate a set of structured design points.
[0026] The trend matching module is used to obtain a pre-built clothing trend library, perform similarity matching between the structured design point set and the clothing trend library, and generate a trend matching score.
[0027] The innovation calculation module is used to calculate the design innovation based on the historical occurrence frequency of the structured design point set and generate an innovation score.
[0028] The comprehensive score calculation module is used to dynamically weight the quality score, trend matching score, and innovation score to generate a comprehensive score.
[0029] The decision and output module is used to perform a grading decision based on the comprehensive score and the quality score, and output the images to be stored and grade labels.
[0030] Thirdly, this application provides a computer-readable storage medium, which adopts the following technical solution:
[0031] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as in any of the methods in the first aspect.
[0032] In summary, this application achieves at least one of the following beneficial technical effects: It realizes intelligent evaluation and screening of apparel designs, transforming the traditional process, which heavily relies on subjective experience, into an objective, quantifiable, and fully automated data-driven system. Through computer vision and machine learning technologies, it connects the entire chain from raw images to business decisions. The system not only automatically ensures image quality and standardization but also deeply understands design semantics (identifying elements and analyzing combinations) and combines dynamic market knowledge (trend matching, timeliness and innovation calculation) for multi-dimensional comprehensive evaluation. Ultimately, it outputs structured results with comprehensive scores and grade labels, significantly improving the efficiency, consistency, and scientific rigor of large-scale apparel image screening, providing accurate and interpretable data support for business decisions such as style selection, procurement, and trend analysis. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of the first process of a smart clothing image filtering method according to one embodiment of this application.
[0034] Figure 2 This is a schematic diagram of the second process of the intelligent screening method for clothing images according to one embodiment of this application.
[0035] Figure 3 This is a schematic diagram of the third process of the intelligent screening method for clothing images according to one embodiment of this application.
[0036] Figure 4 This is a schematic diagram of the fourth process of the intelligent screening method for clothing images according to one embodiment of this application.
[0037] Figure 5 This is a schematic diagram of the fifth process of the intelligent screening method for clothing images according to one embodiment of this application.
[0038] Figure 6 This is a schematic diagram of the sixth process of the intelligent screening method for clothing images according to one embodiment of this application.
[0039] Figure 7 This is a schematic diagram of the seventh process of the intelligent screening method for clothing images according to one embodiment of this application.
[0040] Figure 8 This is a schematic diagram of the eighth process of the intelligent screening method for clothing images according to one embodiment of this application. Detailed Implementation
[0041] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figures 1-8 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.
[0042] This application discloses an intelligent filtering method for clothing images.
[0043] Reference Figure 1 A smart filtering method for clothing images, the filtering method includes:
[0044] Step S101: Obtain the dataset of clothing images to be processed;
[0045] The datasets typically come from product images on e-commerce platforms, images shared on social media, and official brand materials. These images vary greatly in resolution, lighting conditions, composition style, and background complexity.
[0046] Step S102: Perform basic quality filtering on the clothing image dataset to generate a set of filtered images based on preset legality rules and image quality thresholds;
[0047] Before conducting in-depth analysis of the content, invalid or low-value data is first eliminated based on pre-set, objective, and rigid rules to avoid wasting resources on subsequent, more complex calculations.
[0048] In this application embodiment, the legality rules may include, but are not limited to, image file integrity verification, copyright watermark detection, and content security filtering, which ensures the compliance of the data source. Image quality thresholds are typically achieved by calculating low-level visual indicators such as image sharpness (e.g., blur detection based on gradient operators), signal-to-noise ratio, brightness, and contrast distribution. This aims to filter out low-quality images that cannot be effectively extracted due to shooting errors, transmission corruption, or severe compression. The resulting set of filtered images constitutes a pool of images that meets basic technical requirements for analysis.
[0049] Step S103: Perform image enhancement processing on the selected image set to generate a standardized image set;
[0050] Due to the diverse sources of the images, even those that pass basic filtering still suffer from issues such as uneven lighting, tilted angles, and inconsistent sizes and color spaces. These problems can severely interfere with the performance and stability of subsequent deep learning-based feature extraction models. The logic of this step is to perform standardized preprocessing on the images using a series of digital image processing algorithms to improve data consistency and quality.
[0051] For example, adaptive histogram equalization can improve the visibility of details in low-light or low-contrast areas. Its principle is to adjust contrast based on the grayscale distribution of local image regions, rather than a globally uniform processing, thus avoiding excessive noise amplification. Affine transformation correction for tilted images involves detecting the principal direction of the clothing or image edges and performing rotation and translation to restore the main body of the clothing to a near-normal viewing angle. This is crucial for accurately identifying subsequent design elements such as collar and sleeve shapes.
[0052] Finally, scaling the input tensors to a fixed size (e.g., 1024×1024 pixels) and converting the color space (e.g., to RGB) is to meet the strict requirements of downstream convolutional neural network models for the input tensor format, ensuring the feasibility and computational efficiency of batch processing.
[0053] Step S104: Extract multi-dimensional visual features from the standardized image set to generate a visual feature vector and quality score for each image.
[0054] The method utilizes a dual-branch convolutional neural network to perform two tasks simultaneously: first, to quantitatively evaluate the image aesthetics and composition quality (generating quality scores); and second, to extract deep features for high-level semantic understanding (generating visual feature vectors).
[0055] Specifically, the quality scoring branch learns from a large number of manually labeled high-quality images, enabling the model to quantitatively evaluate the compositional balance (e.g., whether it conforms to the rule of thirds, whether the subject stands out), color contrast (whether the color scheme is harmonious and vivid), and texture clarity (whether the fabric texture is clearly presented). Meanwhile, another branch outputs a high-dimensional, dense visual feature vector (e.g., 512-dimensional). This vector is generated by global pooling of feature maps obtained after multiple non-linear transformations and abstractions of the input image through deep convolutional layers of the network. It densely encodes comprehensive visual information such as the style, texture, and contour of the clothing in the image, serving as the cornerstone for all subsequent semantic analysis tasks (e.g., segmentation, design point recognition).
[0056] Step S105: Based on visual feature vectors, identify the main clothing area using a pre-trained clothing area segmentation model and generate a clothing area mask.
[0057] Clothing images often contain complex backgrounds, and analyzing design elements directly on the entire image would introduce a lot of noise. The logic of this step is to focus attention and use semantic segmentation technology to accurately separate the main body of the clothing from the background, model's skin, and other areas.
[0058] In this embodiment, the pre-trained clothing region segmentation model (typically trained on a large clothing segmentation dataset such as DeepFashion2) receives the semantically rich visual feature vector obtained in the previous step as input. Through an encoder-decoder network, it predicts a category label (e.g., clothing, background) for each pixel in the image. Finally, the output is a binary clothing region mask, where pixels belonging to the clothing are labeled 1 (or white), and the background is labeled 0 (or black). This mask acts like a precise "silhouette" or "mask," ensuring that all subsequent analysis of the clothing itself only applies to the valid region, greatly improving the accuracy and robustness of the analysis.
[0059] Step S106: Combine the clothing area mask and visual feature vector to extract semantic tags of clothing design elements and generate a set of structured design points.
[0060] After focusing on the main body of the clothing, the logic of this step is to perform fine-grained visual semantic analysis, breaking down the clothing into describable and categorizable design components.
[0061] Specifically, the system uses masks to crop or weight the extracted visual feature maps, shielding them from background interference. Then, through a multi-scale semantic segmentation or attribute recognition model, it further identifies and classifies more specific design elements within the clothing area. For example, it identifies whether the collar type is "V-neck," "round neck," or "collar"; whether the sleeve type is "puff sleeve," "lantern sleeve," or "regular sleeve"; and whether the pattern is "striped," "printed," or "solid color."
[0062] In addition, each identified element is accompanied by a confidence score, representing the degree of certainty in the model's judgment. By setting a confidence threshold (e.g., 0.8), unreliable identification results are filtered out, ultimately generating a structured set, such as {collar type: V-neck, 0.95}, {sleeve type: puff sleeve, 0.92}, {pattern: solid color, 0.88}. This step achieves the transformation from overall visual features to discrete, searchable design semantic symbols.
[0063] Step S107: Obtain the pre-built clothing trend library, perform similarity matching between the structured design point set and the clothing trend library, and generate a trend matching score.
[0064] This involves measuring the design semantics of current clothing within the broader context of fashion trends to assess its trend relevance.
[0065] In this embodiment, the pre-built apparel trend library is a dynamic knowledge base. It continuously collects and analyzes cutting-edge and market data such as runway images, brand catalogs, and e-commerce bestsellers. Using unsupervised machine learning algorithms such as spectral clustering, it automatically summarizes current and near-future core trend clusters (e.g., specific design point combinations under trends like "Y2K style," "New Chinese style," and "functionalism"). During matching, the system calculates the similarity (e.g., cosine similarity) between the structured set of design points in the current image (considered a feature vector) and the representative vectors of each trend cluster in the trend library. A higher matching degree indicates that the combination of design elements in the apparel more closely matches known trends, resulting in a higher trend matching score. This gives the system a cutting-edge "fashion sense."
[0066] Step S108: Calculate the design innovation degree based on the historical occurrence frequency of the structured design point set, and generate an innovation degree score;
[0067] In contrast to following trends, this step assesses the uniqueness and novelty of the clothing design, i.e., its degree of unconventionality. The less frequently a combination of design elements appears in historical data, the higher its innovativeness. The system maintains a statistical model of the co-occurrence frequency of design points. When analyzing images of new designs, the system calculates the probability or inverse document frequency of the occurrence of its set of design points (or specific combinations) in the historical database. For example, a garment combining an "asymmetrical collar" and "detachable sleeves," if extremely rare in historical data, will receive a high innovativeness score. This metric complements the trend score, identifying potential styles that may lead the next wave of fashion trends.
[0068] Step S109: The quality score, trend matching score, and innovation score are dynamically weighted and calculated to generate a comprehensive score.
[0069] Specifically, a carefully designed multi-objective fusion function integrates three heterogeneous scores representing "visual aesthetics," "trendiness," and "design uniqueness" into a comprehensive and balanced overall score. Dynamic weighting is not a simple addition of fixed weights, but may employ non-linear strategies (for example, combining the maximum and minimum values of trend and innovation scores).
[0070] Understandably, the principle behind this design is that, for commercial selection, a garment either has the potential to become an instant bestseller due to its high degree of trend conformity (MAX trend score), or it has the potential to lead the trend and attract a specific customer group due to its unique innovation (MIN innovation score as a supplementary consideration), while ensuring basic visual presentation quality (fixed weight).
[0071] In some embodiments, the dynamic weighted calculation can be performed using the following formula: Overall score = 0.4 × Quality score + 0.5 × MAX(Trend matching score, Innovation score) + 0.1 × MIN(Trend matching score, Innovation score); where the MAX / MIN function is used to balance the conflict between trend following and innovative breakthroughs.
[0072] Step S110: Execute a grading decision based on the comprehensive score and quality score, and output the images to be stored and their grade labels.
[0073] Specifically, a hard filter based on quality scores is applied, for example, directly discarding images with quality scores below a technical threshold to ensure the visual usability of the images in the database. Secondly, images that pass the quality filter are graded according to their comprehensive score, for example, into levels such as "Excellent," "Standard," and "Observable." This strategy ensures that the final output set of "included images and grade tags" is not only technically clear and usable, but also commercially value-ranked through multiple dimensions, and can be directly used for downstream applications such as intelligent uploading, priority recommendation, or the construction of designer inspiration libraries.
[0074] The above implementation achieves end-to-end intelligent screening of massive amounts of clothing images, from low-level pixel processing to high-level business intelligence decision-making. First, standardized preprocessing and deep learning-based multi-dimensional feature extraction ensure the robustness and accuracy of the analysis process. Second, it innovatively combines clothing region segmentation, design point semantic parsing, dynamic trend matching, and innovation quantification, enabling the screening criteria to simultaneously cover three key dimensions: visual aesthetics, market trends, and design originality, far exceeding simple image quality filtering or tag matching. Finally, through a dynamic weighted fusion model and two-layer hierarchical decision-making, it outputs structured results that directly support business decisions, significantly improving the efficiency and intelligence of applications such as clothing e-commerce selection, fashion trend analysis, and design inspiration mining, transforming traditional processes reliant on human experience into scalable and iterative data-driven processes.
[0075] Reference Figure 2 As one implementation of step S102, the step of performing basic quality filtering on the clothing image dataset and generating a set of filtered images based on preset legality rules and image quality thresholds includes:
[0076] Step S201: Obtain the dataset of clothing images to be processed, as well as the preset set of legality rules and image quality thresholds;
[0077] The clothing image dataset serves as the raw input, originating from diverse and complex sources, potentially containing images of varying formats, resolutions, and image qualities. A pre-defined set of legality rules defines the compliance requirements for data format and structure, such as digital signatures for file types and color encoding modes. This essentially sets an access control framework for the system, aiming to exclude damaged, forged, or unsupported files at the source, ensuring the stability and security of the data processing flow. The image quality threshold set, on the other hand, establishes specific quantitative thresholds for the physical attributes of image content, such as minimum resolution, maximum permissible blur, and noise levels, constituting a technical standard for measuring whether an image possesses basic usability. The pre-defined sets of these two sets transform the entire filtering process from relying on empirical judgment to deterministic calculations based on objective rules.
[0078] Step S202: Perform binary data parsing on each clothing image in the clothing image dataset to generate file header feature codes and metadata structures;
[0079] In computer systems, image files are stored as binary byte streams at the underlying level. By parsing this binary stream, the system first reads a specific sequence of bytes at the beginning of the file, known as the file header signature. This signature is a unique identifier for the file format, like a "digital fingerprint" for the file; for example, the JPEG format typically begins with "FFD8".
[0080] At the same time, the parsing process decodes the metadata structure stored inside the file according to a specific standard (such as EXIF). This structure contains descriptive information embedded in the file, such as key parameters like the pixel width and height of the image, color space, and bit depth.
[0081] Step S203: Perform format verification based on the file header feature code and the set of legality rules to generate a legality verification mark;
[0082] In this process, the system precisely compares the file header signature extracted in the previous step with the pre-stored standard format signatures in the set of validity rules. For example, the rule specifies that a valid PNG file signature should be "89504E47". If the signature matches successfully, it proves that the file structure is complete, the format is correct, and it is a format supported by the system. Conversely, if the signature does not match or is corrupted, it indicates that the file may be corrupted, tampered with, or be an unknown format that the system cannot process. This verification result is converted into a binary validity check flag (such as "valid" or "invalid"). This step is crucial; it prevents the system from encountering errors or crashes due to attempting to parse illegal files, ensuring the robustness of the process.
[0083] Step S204: Extract image resolution parameters based on the metadata structure and compare them with resolution thresholds in the image quality threshold set to generate resolution detection labels;
[0084] The resolution parameter is directly obtained from the parsed metadata structure and is typically represented as a pixel value of "width × height". The system compares this value with a preset resolution threshold (e.g., a minimum requirement of 800 × 600 pixels). The logic is that resolution is the fundamental indicator of image information capacity. Images with too low a resolution, even if the content itself is clear, lack sufficient pixels to present the details of clothing (such as fabric texture, seams, and print patterns). This not only severely impacts the user's visual experience but also prevents subsequent deep learning-based feature extraction models from capturing effective information, rendering all intelligent analysis meaningless.
[0085] Step S205: Extract sharpness features from the pixel matrix of each clothing image, calculate the blur index and compare it with the sharpness threshold to generate a sharpness detection label.
[0086] Specifically, mathematical operators quantify the intensity of edges and textures in an image. Typically, the system first converts the color image into a grayscale matrix to simplify calculations. Then, an edge detection filter, such as the Laplacian operator, is applied to convolve the grayscale matrix. The Laplacian operator is highly sensitive to changes in the second derivative of the image (i.e., rapid abrupt changes in pixel intensity). Sharp images contain rich, strong edges, resulting in drastic changes in the convolution result; while blurry images have diffused edges, resulting in gradual changes in the convolution result. By calculating the variance of the convolution result matrix, a blur index can be obtained; a higher variance value indicates a sharper image. Comparing this index to a preset sharpness threshold objectively determines whether the image has lost key details due to focus failure, camera shake, or over-compression, generating a "qualified" or "unqualified" sharpness detection label.
[0087] Step S206: Calculate the noise density of the pixel matrix of each clothing image and compare it with the noise threshold to generate noise detection markers.
[0088] Noise typically originates from electronic interference or image compression artifacts in the camera sensor under low-light conditions, manifesting as randomly distributed, unrealistic color or brightness spots in the image.
[0089] In this embodiment, noise density can be calculated using local variance analysis. The system divides the image into multiple small blocks (e.g., 8×8 pixels) and calculates the variance of pixel values within each block. In smooth regions (e.g., a solid-color clothing background), the theoretical variance should be close to zero; if noise exists in the region, the variance value will be abnormally high. By statistically analyzing the proportion of blocks whose variance exceeds the noise threshold across all blocks, the noise density estimate for the entire image can be obtained. High noise density can severely interfere with subsequent algorithms; for example, noise may be misidentified as complex textures or fine patterns of fabric, leading to errors in feature extraction and design point recognition. Therefore, this step generates noise detection markers to filter out images with clean, "clean" content.
[0090] Step S207: Generate a comprehensive quality detection mark by combining the legality verification mark, resolution detection mark, sharpness detection mark, and noise detection mark;
[0091] The aforementioned detection markers represent four independent but crucial quality dimensions: file integrity, basic specifications, content clarity, and content purity. The system requires an image to meet the standards in all four dimensions simultaneously to pass the basic quality filter. Any defect in any single dimension (such as invalid file, low resolution, blurry image, or excessive noise) will result in a "fail" rating in the final overall quality detection. This design simulates a rigorous quality inspection process, ensuring the comprehensiveness and reliability of the screening results and preventing images with other serious quality issues from entering subsequent stages due to passing only one indicator.
[0092] Step S208: Based on the comprehensive quality inspection labels, select images that meet all thresholds and generate a set of filtered images.
[0093] The process involves iterating through all input images and retaining only those marked as "passed," creating a filter image set. Each image in this filter set has been verified to meet the following filtering criteria: 1) It is a complete and valid image file; 2) It has a resolution that meets the minimum requirements; 3) The image is clear and details are discernible; 4) Noise interference is at an acceptable low level.
[0094] The above implementation constructs a multi-layered, quantifiable, automated detection channel from file structure parsing to pixel content analysis, achieving efficient and reliable basic quality filtering of massive amounts of clothing images. This solution first ensures the legitimacy of the data source and the security of the processing system from the root through binary-level format verification; then, through quantitative analysis of resolution, sharpness, and noise, it filters out effective images with sufficient detail and clean visual content at both the physical specifications and visual content levels, laying a solid data foundation for downstream intelligent algorithms that rely on high-quality visual input (such as feature extraction and design point recognition).
[0095] Reference Figure 3 As one implementation of step S104, the step of extracting multi-dimensional visual features from the standardized image set to generate a visual feature vector and quality score for each image includes:
[0096] Step S301: Process the standardized image set through a pre-trained feature extraction model to generate a primary feature tensor;
[0097] Specifically, a convolutional neural network (CNN) pre-trained on a massive general-purpose image dataset (such as ImageNet) is used as a powerful general-purpose feature extractor. Taking ResNet-50 as an example, as a normalized image passes through its multiple convolutional layers, the network performs non-linear transformations and abstractions layer by layer. Shallow networks may capture basic features such as edges and corners; while deep networks can extract more complex and semantic features, such as the outline of clothing parts, the texture patterns of fabrics, and complex pattern combinations. The primary feature tensor output by this layer is a three-dimensional data structure (usually the number of channels × height × width), where each spatial location (height × width) corresponds to a region in the original image, and the activation values of multiple channels jointly describe the rich visual attributes of that region.
[0098] Step S302: Perform spatial attention weighted calculation on the primary feature tensor to generate an enhanced feature map;
[0099] Since clothing images often contain complex backgrounds, models, or other distractions, not all information in the initial feature tensor is equally important for subsequent tasks. The logic of spatial attention weighted computation is to allow the model to automatically learn and assign higher weights to features related to the main area of the clothing and key details (such as collar and cuff textures), while suppressing feature responses from irrelevant backgrounds.
[0100] Specifically, this step is typically implemented through a sub-network that analyzes the primary feature tensor and generates an attention weight matrix with the same size as the feature map space, where the weight value at each location represents the importance of the feature at that location. By performing element-wise multiplication (such as a Hadamard product) of this weight matrix with the primary feature tensor, the activation values of the background regions are weakened, while the activation values of the main clothing and detail regions are enhanced, thus obtaining an enhanced feature map. This step significantly improves the task relevance and information density of the feature representation.
[0101] Step S303: Input the enhanced feature map into the quality scoring branch network and output a multi-dimensional quality index including compositional balance, color contrast, and texture sharpness.
[0102] This step quantifies the image quality from both aesthetic and shooting technique perspectives, deconstructing the subjective "good-looking" into multiple objectively measurable professional image dimensions.
[0103] Specifically, the quality scoring branch network receives enhanced feature maps rich in semantic information and maps them to three specific quality metrics through specific network layers (usually fully connected layers). Compositional balance assesses the harmony of the layout of visual elements in the image, for example, by calculating the offset of the visual center of gravity from the classic rule of thirds. Color contrast is not simply about color vibrancy, but rather assesses the harmony and depth of color combinations, typically calculated in the LAB color space that conforms to human visual perception, by determining the color difference (ΔE) and its distribution variance between key color blocks. Texture sharpness quantifies the sharpness of details such as fabric and seams, and can be achieved by analyzing the energy proportion of high-frequency components after wavelet transform of the image; high high-frequency energy indicates well-preserved details and a sharp image. These three metrics together constitute a refined diagnosis of the visual presentation quality of an image.
[0104] Step S304: Perform linear weighted aggregation based on multidimensional quality indicators to generate quality score values;
[0105] The linear weighted aggregation method assigns different weight coefficients (α, β, γ) to the three indicators of compositional balance, color contrast, and texture clarity according to business needs or prior aesthetic knowledge. The sum of these weights is 1, ensuring a controllable range for the score. Through weighted summation, these three indicators with different dimensions and meanings are integrated into a unified quality score. This score comprehensively reflects the overall performance of the image across multiple key aesthetic dimensions and can serve as a direct basis for selecting high-quality visual materials. Its generation process is transparent and the parameters are adjustable.
[0106] Step S305: The enhanced feature map is input into the feature encoding branch network, and a fixed-dimensional visual feature vector is generated through dimensionality reduction operation;
[0107] The logic of this step is feature compression and global information aggregation. The enhanced feature map is rich in information, but has a high dimension and retains spatial structure.
[0108] Specifically, the feature encoding branch network (typically composed of global average pooling layers and fully connected layers) first averages the augmented feature map across its spatial dimensions (height and width) using a global average pooling operation. This operation collapses the two-dimensional feature map of each channel into a scalar, thereby transforming the three-dimensional C×H×W feature map into a C-dimensional vector that captures the average response intensity of the visual pattern represented by that channel across the entire image.
[0109] Subsequently, a fully connected layer is used for further linear transformation and dimensionality reduction, ultimately outputting a fixed-length (e.g., 512-dimensional) visual feature vector. This vector is the "digital fingerprint" of the entire image; it loses spatial location information but highly summarizes the essence of its visual content.
[0110] Step S306: Output the visual feature vector and quality score for each image.
[0111] The visual feature vector, as a deep abstract encoding of the image content, will be directly input into subsequent tasks such as clothing region segmentation and design point recognition, serving as the starting point for their semantic understanding. The quality score, on the other hand, acts as an independent evaluation metric, participating in the final comprehensive score calculation and hierarchical decision-making logic. These two outputs originate from the sharing and branching of the same "enhanced feature map," maximizing the reuse of computational resources and achieving the efficient goal of simultaneously acquiring content features and quality evaluation from a single forward propagation.
[0112] In the above implementation, the attention enhancement mechanism enables the model to focus on the main body of the clothing, effectively suppressing background interference, thereby improving the accuracy and robustness of all subsequent semantic analysis tasks based on this feature vector (such as design point recognition). The decoupling and quantification of multi-dimensional quality indicators transforms subjective aesthetic evaluation into objective and interpretable numerical indicators, providing a scientific and transparent decision-making basis for the visual quality screening of images. The dual-branch parallel processing architecture, based on the shared underlying deep feature calculation, synchronously outputs feature vectors for content understanding and scores for quality assessment. While ensuring feature richness and comprehensiveness of evaluation, it greatly improves the overall processing efficiency of the system, avoids redundant calculations, and provides key technical support for real-time intelligent screening and analysis of large-scale clothing images.
[0113] Reference Figure 4 As one implementation of step S106, the step of extracting semantic tags of clothing design elements and generating a set of structured design points by combining clothing region masks and visual feature vectors includes:
[0114] Step S401: Perform spatial region clipping on the visual feature vector based on the clothing region mask to generate the main feature map of the clothing.
[0115] The clothing region mask is a binary image that precisely identifies the pixel regions in the image that belong to the clothing, essentially providing the computer with accurate spatial localization of "where the target is." Meanwhile, the visual feature vector is a high-dimensional, dense abstract feature representation of the entire image extracted through a deep convolutional network, encoding comprehensive information such as the image's color, texture, and shape.
[0116] Furthermore, spatial cropping at the feature tensor level is a key preprocessing step to improve recognition accuracy. The underlying principle is to implement attention focusing and noise isolation at the feature level. The original visual feature vector (usually derived from global pooling of feature maps) or its feature map originates from the entire image, containing a mixture of information from all regions, including clothing, background, and model skin. Directly identifying design elements on this mixed feature set would result in background noise severely interfering with the model's judgment.
[0117] Therefore, the system uses a precise clothing region mask as a "digital mold" to set the activation values of those corresponding to the background region to zero or directly crop them in the spatial dimensions (height and width) of the feature map, retaining only the feature activations that fall completely within the clothing region. As a result, each effective feature point on the generated clothing main feature map corresponds purely to the visual attributes of the clothing itself, completely eliminating the interference of background clutter on the subsequent semantic recognition model and laying a pure data foundation for high-precision recognition.
[0118] Step S402: Perform multi-level feature fusion operation on the main feature map of the clothing to generate multi-scale fused features;
[0119] Among them, the identification of clothing design elements requires sensitivity to visual patterns at different scales. The logic of this step is to construct a robust feature representation that can simultaneously understand macroscopic outlines and microscopic details.
[0120] Specifically, global contextual features of clothing (such as whether the overall silhouette is A-shaped or H-shaped) can be extracted using dilated convolution. This technique can expand the receptive field without sacrificing resolution, effectively capturing a wide range of semantic information. At the same time, local detail features of clothing (such as the texture of lace, the style of buttons, and the direction of seams) need to be recovered through upsampling operations such as deconvolution (or transposed convolution) to restore the fine spatial information that may be lost due to network downsampling.
[0121] Finally, a gating mechanism (an adaptive weight learning structure) is employed to intelligently fuse global and local features. This mechanism acts like a dynamic switch, assigning appropriate weights to features at different spatial locations and through different channels. This ensures that the resulting multi-scale fused features can both capture the overall style of the clothing and clearly depict its local design essence, thus comprehensively supporting the recognition of diverse elements from collar type to patterns.
[0122] Step S403: Input the multi-scale fusion features into the pre-trained design element recognition model, and output the initial semantic label set and confidence distribution;
[0123] The specially trained design element recognition model (e.g., using a network structure with multiple parallel classification heads) receives fused multi-scale features as input, and different branches or sub-networks within the model are designed to focus on and classify design attributes in specific dimensions.
[0124] For example, one branch might focus on analyzing the structure of the upper part of the garment to output predictions about collar types (such as V-neck, round neck, stand collar); another branch might analyze the sleeve area to output sleeve structure (such as raglan sleeve, lantern sleeve); and a third branch might analyze texture and pattern to output pattern material (such as stripes, prints, lace).
[0125] For each identified potential element, the model not only provides a classification label but also outputs a confidence score, a probability value between 0 and 1 that quantifies the model's certainty about the prediction. All these predictions constitute the initial set of semantic labels and their confidence score distribution, reflecting the model's "first impression" of all possible design points in the image, but may contain some incorrect or low-certainty labels due to model uncertainty or subtle visual confusion.
[0126] Step S404: Filter the initial semantic tag set according to the preset confidence threshold to generate a valid semantic tag set;
[0127] Using all initial labels directly is not robust, as low-confidence labels are likely to correspond to incorrect identifications.
[0128] In this embodiment, an objective filtering standard is provided by setting a confidence threshold (e.g., 0.7): only those recognition results that the model is sufficiently confident in are retained. A dynamic adjustment strategy can also be adopted, setting different thresholds for different types of design elements.
[0129] For example, a higher standard threshold can be used for common, high-frequency design elements (such as a standard crew neck) to ensure the accuracy of routine identification; for designs that occur less frequently but may represent innovative elements (such as a novel asymmetrical sleeve), a more lenient threshold is used to avoid missed detections due to the model's unfamiliarity with the pattern, thus preserving the possibility of discovering potential innovative designs. After this step, ambiguous and insufficiently evidenced labels are eliminated, and the resulting set of effective semantic labels is a highly reliable and concise list of design point descriptions.
[0130] Step S405: The set of valid semantic tags is structured and encoded according to a predefined design element classification system to generate a set of structured design points.
[0131] The predefined design element classification system provides a unified "dictionary" or "outline" for all possible design points, ensuring that the labels generated by different images are consistent in semantics and format.
[0132] In this embodiment, structured encoding, for example, employs JSON-LD (a JSON-based associative data format) to transform each valid semantic label into a structured data entry. Each entry not only contains the core element_type (e.g., "collar type") and design_label (specific label, such as "V-neck"), but also encapsulates its confidence and optional position (the bounding box of its spatial location in the image). This structured set of design points is no longer just a collection of text, but a well-defined, self-describing data structure. It can be directly used to calculate similarity to a trend database (by comparing label types and names), conveniently statistically analyze historical frequency to calculate novelty, and is easily stored in a database or transmitted via a network interface.
[0133] In the above implementation, firstly, the spatial feature clipping based on the mask fundamentally isolates background interference, allowing the model to focus entirely on the garment itself, greatly improving the signal-to-noise ratio of the recognition task; secondly, the multi-level feature fusion mechanism ensures that the system can understand both the overall silhouette context of the garment and capture subtle local textures and decorative details, overcoming the limitations of single-scale feature representation; thirdly, the introduction of a dynamic confidence filtering strategy reduces the risk of missing potential innovative designs while ensuring the accuracy of conventional element recognition, making the system both reliable and exploratory; finally, through standardized structured coding, the recognition results are transformed into machine-friendly data objects, perfectly adapting to downstream advanced analysis tasks such as trend matching and innovation calculation, forming a seamless data flow from visual perception to semantic cognition to decision support.
[0134] Reference Figure 5 As one implementation of step S107, the steps of obtaining a pre-built clothing trend library, performing similarity matching between the structured design point set and the clothing trend library, and generating a trend matching score include:
[0135] Step S501: Obtain the pre-built clothing trend library and the set of structured design points for the current clothing image;
[0136] The pre-built fashion trend library is a dynamic knowledge system that has undergone continuous learning and evolution. It is not a static image set, but a structured representation formed through in-depth analysis and clustering of massive amounts of runway data, brand catalogs, and market bestsellers. Each trend cluster is abstracted as a trend feature vector and is accompanied by a generation timestamp.
[0137] On the other hand, the current set of structured design points in an image is a standardized and reliable list of descriptions of all the key design elements (such as collar type, sleeve type, and pattern) of the garment image.
[0138] Understandably, acquiring these two things means that the system has the prerequisites for meaningful comparison: one is a widely recognized trend "dictionary," and the other is a "feature description statement" of individual clothing. The two have been aligned at the semantic level (using the same design element classification system), providing a foundation for subsequent numerical similarity calculation.
[0139] Step S502: Convert the structured design point set into design point feature vectors according to a preset vector encoding rule;
[0140] The design point set is a structured data list, while efficient similarity calculation by computers requires numerical vectors of a unified dimension. The logic of this step is to perform a multimodal information fusion and numerical encapsulation. The pre-defined vector encoding rules carefully combine various encoding techniques to comprehensively preserve all dimensions of the original information.
[0141] In this embodiment, firstly, one-hot encoding is performed on design element types (such as "collar type" and "sleeve type") to generate a sparse type vector. This vector clearly identifies the category of each design point, defining clear boundaries for attributes of different categories in the vector space. Secondly, word embedding encoding is performed on specific design element tags (such as "puff sleeve" and "print"). This utilizes a language model pre-trained on apparel-related text to map text tags into a dense semantic vector. This vector captures the semantic relationships between tags (for example, "puff sleeve" is closer to "lantern sleeve" in the vector space, while it is farther from "regular sleeve"). Finally, the confidence score of each design point is concatenated with the above vector as a scalar. The resulting design point feature vector is a composite carrier that simultaneously encodes the "category attribution," "semantic connotation," and "identification confidence" of the design point, providing a comprehensive and accurate numerical basis for high-fidelity comparison with trend feature vectors.
[0142] Step S503: Extract the trend feature vector and time decay weight of each trend cluster from the clothing trend library;
[0143] Each trend cluster is assigned a trend feature vector (whose encoding rules are compatible with design point feature vectors to ensure comparability) and a timestamp upon creation. The time decay weight is dynamically calculated based on the difference between the current time and the time the trend cluster was generated (e.g., the month difference Δt) using a decay function (e.g., exponential decay). Its core principle is to simulate the natural lifecycle of fashion trends: a newly emerging trend has the greatest influence and weight; its popularity naturally decays over time. Therefore, an outdated trend, even if highly similar to current clothing in design semantics, should contribute little to the final matching score. Extracting these two elements means that when calculating similarity, the system considers not only "similarity" but also "whether it is currently trending," giving the matching score real-world temporal significance.
[0144] Step S504: Calculate the similarity value between the design point feature vector and each trend feature vector, and generate a similarity distribution sequence;
[0145] The system iterates through each trend cluster in the trend library and calculates the similarity value between the design point feature vector of the current image and the trend feature vector of each cluster.
[0146] Specifically, a similarity metric such as cosine similarity is typically used. This metric, with a value range between [-1, 1] and [0, 1], quantifies the degree of alignment between two vectors in direction. A higher value indicates that the current garment's design elements align more closely with the trend direction represented by that specific trend cluster. After performing this calculation for all trend clusters, a similarity distribution sequence is obtained, which visually demonstrates the relationship between the current garment and various trend points across the entire fashion spectrum. For example, it might show that it is most similar to "minimalism," followed by "sportswear," while having a lower similarity to "retro luxury."
[0147] Step S505: The similarity distribution sequence is weighted and corrected according to the time decay weight to generate a weighted similarity sequence;
[0148] The original similarity value is multiplied by the extracted time decay weights of the corresponding trend clusters (or fused using a more complex formula). For example, an old trend with high similarity to current clothing designs will have a significantly lower weighted score due to its low time decay weight. Conversely, a new trend with moderate similarity to current clothing designs but currently at its peak popularity may receive a considerable weighted score due to its higher time weight.
[0149] This process transforms the original similarity distribution sequence into a weighted similarity sequence. This new sequence more accurately reflects market reality: it lowers the score for "historical replicas" and raises the score for "currently trendy" styles, upgrading the matching results from academic design style comparisons to commercially relevant trend assessments.
[0150] Step S506: Perform maximum value aggregation on the weighted similarity sequence to generate a trend matching score.
[0151] Among them, finding the maximum value in the weighted similarity sequence directly reflects the matching strength of the clothing with the most suitable trend among all trends, which is defined as the trend matching score.
[0152] Furthermore, a more intelligent strategy is to consider the following: when the highest and second-highest scores are very close (the difference is less than a threshold), it indicates that the garment may incorporate elements of multiple similar trends. In this case, taking the average of the top few scores provides a more robust overall score that is less susceptible to minor fluctuations. This step ultimately outputs a concise scalar score that efficiently summarizes the garment's position within the dynamic fashion context and can be directly used for ranking, filtering, or as core input for comprehensive business scoring.
[0153] In the above implementation, a refined matching chain integrating multimodal coding, spatiotemporal joint modeling, and intelligent aggregation decision-making is constructed, realizing automated, dynamic, and high-precision quantitative evaluation of the fit between clothing design and trends. This technical solution transforms the subjective and vague "fashion sense" into an objective and calculable "trend score," providing core data-driven decision-making capabilities for intelligent selection, trend prediction, and design inspiration mining in apparel e-commerce, greatly improving the intelligence level of industry operations and market responsiveness.
[0154] Reference Figure 6 As one implementation of step S108, the step of calculating the design innovation degree based on the historical occurrence frequency of the structured design point set and generating an innovation degree score includes:
[0155] Step S601: Obtain the set of structured design points and the pre-built database of historical design elements;
[0156] The structured design point set is a precise semantic decomposition of the garment to be evaluated, specifying all its design elements (such as "V-neck," "puff sleeves," and "print"). The pre-built historical design element database is a knowledge base that has accumulated a large number of long-term design cases. It systematically records the frequency of various design elements in the past (e.g., the last five years) and the dates of their first or most recent appearance. Obtaining both means the system has mastered the "evaluation object" and the "evaluation benchmark." The innovation calculation essentially compares the feature vector of the current design with this vast historical feature distribution to determine whether it falls into the common or rare zone.
[0157] Step S602: Extract the time decay frequency weight of each design element from the historical design element database;
[0158] This step introduces a dynamic perspective on evaluating innovativeness over time. The logic behind this is to acknowledge that the "novelty" of a design is time-sensitive and naturally diminishes over time. Understandably, a design element that was highly disruptive when it first appeared five years ago (such as a particular form of deconstructivist tailoring) might no longer be considered innovative today if it has been widely imitated and used; instead, it might become a classic or common element.
[0159] In the embodiments of this application, the time decay frequency weight is designed to quantify this effect. It is typically calculated based on the time difference (Δt) between the current time and the time when the design element most recently appeared (or first appeared) in the historical database, using a function such as inverse proportional or exponential decay. The core principle is: the larger the time difference, the greater the novelty discount due to the element's "old age," and the lower its weight; conversely, elements that have only recently appeared are more "fresh," and therefore have a higher weight. This weight is used to correct its original frequency, making the innovation assessment closer to "novelty in the current market context" rather than "first appearance in absolute history."
[0160] Step S603: Perform a historical frequency query for each design element in the structured design point set to generate a basic frequency value;
[0161] The logic behind this step is to quantify the "commonness" or "scarcity" of individual design elements. The system iterates through every design element of the current garment, using its type and tag as keys to query a vast database of historical design elements. The query result is the base frequency value of the element's appearance in historical data. This can be an absolute count, a normalized probability, or a smoothed frequency estimate. This value intuitively reflects the "popularity" of the design element: a "round neck" might have an extremely high base frequency value, indicating it is extremely common and lacks novelty; while an "asymmetric magnetic snap collar" might have a base frequency value close to zero, indicating it is historically extremely rare and therefore has high potential for innovation. This step assigns an initial numerical tag to each design element, representing its individual historical status.
[0162] Step S604: Correct the base frequency value according to the time decay frequency weight to generate a weighted historical frequency value.
[0163] This step integrates statistics with timeliness, transforming static historical frequency into a dynamic, contextualized scarcity indicator. Understandably, the innovative impact of a design element depends not only on how many times it has appeared, but also on how long ago those appearances occurred. Directly using base frequency values would treat high-frequency elements from ten years ago and rare elements that only appeared twice yesterday equally, which clearly doesn't align with people's perception of "innovation."
[0164] Therefore, the system applies the extracted time-decay frequency weights to the obtained base frequency values. Through multiplication or other fusion operations, the weighted frequency value of an element that was historically high-frequency but has not appeared recently will be significantly reduced due to its low time weight; conversely, the weighted frequency value of an element with a low overall frequency but only recently observed will be relatively increased due to its high time weight. The generated weighted historical frequency value is a more refined indicator, simultaneously encoding both "how common the element is" and "whether its commonness is a thing of the past," more accurately reflecting its actual scarcity and novelty potential in the current fashion context.
[0165] Step S605: Aggregate and calculate the weighted historical frequency values of all design elements in the structured design point set to generate a combinatorial innovation index.
[0166] The logic behind this step is to move beyond evaluating individual elements in isolation and instead measure the overall innovativeness of the entire design portfolio. Innovation in clothing design often lies not only in isolated components but also in the novel and unexpected combination of multiple known elements to create a unique overall aesthetic. Since simple arithmetic averaging can obscure this synergistic effect, the system employs non-linear aggregation functions such as geometric mean to comprehensively calculate the weighted historical frequency values of all design elements.
[0167] Specifically, the geometric mean has the characteristic that when all values are low (i.e., all elements are scarce), the result will be very low, highlighting the innovation of the "completely novel combination"; conversely, as long as one value is high (i.e., it contains a very common element), it will significantly raise the overall result, reflecting the reality of "mixing classic and novel" or "the main element being common". The combination innovation index calculated in this way is a single value that integrates all elements and their combination relationships. It quantifies the overall rarity of the current combination of design points in the historical context and time sequence, and is the core intermediate representation of innovation.
[0168] Step S606: Map the combined innovation index to a preset scoring range to generate an innovation score.
[0169] The logic behind this step involves business-oriented scaling and range adjustment. The range and distribution of combined innovation indicators may not be suitable for direct use as scoring (e.g., the value range may be narrow or the distribution uneven).
[0170] In this embodiment, the preset scoring range (e.g., 0-100 points) and mapping rules (e.g., piecewise linear functions) are predefined according to business needs. The mapping operation maps the calculated indicator values to the final innovation score according to their respective ranges and corresponding transformation slopes. For example, it can be designed such that: when the indicator is extremely low (the combination is extremely rare), an extremely high score (e.g., 90-100 points) is assigned, representing a breakthrough innovation; when the indicator is moderate, the score is also moderate (e.g., 70-90 points), representing a meaningful improvement; when the indicator is high (the combination is very common), the score is lower (e.g., 30-70 points). This step ensures the consistency and interpretability of the output score range and can sensitively reflect the differences between different levels of innovation, facilitating subsequent sorting, filtering, or weighted fusion with other scores (e.g., quality score, trend score).
[0171] The above implementation achieves automated, refined, and interpretable quantitative evaluation of the innovation of clothing designs. First, by introducing a time-decay frequency weight, the system can dynamically perceive the "shelf life" of design element novelty, ensuring that the evaluation results are synchronized with the rapidly iterating fashion cycle and avoiding misjudging outdated historical innovations as current novel designs. Second, by employing aggregate calculations targeting combined innovation, it transcends the limitations of summing the frequencies of single elements, accurately identifying truly original designs composed of common elements combined in novel ways. Finally, through a non-linear mapping of preset scoring intervals, the internal calculation indicators are transformed into standardized scores that align with business intuition and decision-making needs, making the output results not only mathematically rigorous but also intuitive and operable in practical business applications.
[0172] Reference Figure 7 As a further implementation of the intelligent image filtering method for clothing, after generating an innovation score, it also includes:
[0173] Step S701: Obtain the set of structured design points and the pre-built database of historical design elements;
[0174] Step S702: Extract the first appearance timestamp and historical frequency of each design element from the historical design element database;
[0175] Two core metrics were extracted from historical databases to quantify the "historical significance" and "popularity" of design elements. The first appearance timestamp marks the "birth moment" of a design element within the system's cognitive scope, serving as the starting point for measuring its "originality" and calculating its "market lifespan." Historical frequency of appearance tracks the number of times the element has been recorded or used over a past period (e.g., 24 months), reflecting the breadth and intensity of its market acceptance, imitation, and popularization.
[0176] It should be noted that the extraction of these data requires cleaning and normalization. For example, it is necessary to handle different expressions of the same element in different data sources (such as academic literature, business reports, and patent documents), and to perform cross-source frequency aggregation through weighted merging and other methods to form a unified and comparable historical occurrence record. This ensures that the data used for subsequent calculations is reliable information that has been integrated and verified.
[0177] Step S703: Calculate the time decay factor based on the current time and the first occurrence timestamp to generate the decayed historical frequency;
[0178] This step is the core innovation of this solution, introducing a dynamic time perspective. Its logic is to apply a "fashion life cycle decay model" to correct purely historical frequency. The value of a design element, especially the impact of its "novelty," naturally diminishes as its time on the market increases.
[0179] In this embodiment of the application, the system calculates the difference (Δt) between the current time and the timestamp of the first appearance of the element, and substitutes it into a decay model based on industry experience (such as a negative exponential function) to calculate a time decay factor, which approaches 0 from 1 (just appeared) as time goes by.
[0180] The original historical frequency is then multiplied by this attenuation factor to obtain the attenuated historical frequency. The principle is that a "retro" element that appeared 100 times five years ago should contribute far less to current "innovation" than a "new" element that appeared only 10 times last year. Through this operation, historical frequency is given a time weight, allowing recently appearing design elements to receive a higher "effective scarcity" score, while elements popular in the past are automatically downweighted, thus more accurately reflecting the dynamic nature of the fashion world's "new and old" tendency.
[0181] Step S704: Match the structured design point set with the attenuated historical frequency to generate a timeliness innovation correction value;
[0182] The logic behind this step is to assess the scarcity of the current design combination within its historical context after time-weighting. The system matches each design element in the structured design point set with the historical database to query its corresponding historical frequency after attenuation. However, innovation is not only reflected in the scarcity of individual elements, but also in the novel and unconventional combinational relationships between elements.
[0183] Therefore, the system not only calculates the average decay frequency of individual elements, but also analyzes the "combination decay frequency" of combinations in the current design point set that appear simultaneously, using techniques such as association rule mining. A combination consisting of two elements that are common but have never or rarely appeared together (such as "silk" and "mecha style") has an innovation value far exceeding the sum of the novelty of the individual elements. Through a formula that integrates the scarcity of individual elements and combination relationships, the system calculates a timeliness innovation correction value. When this value is positive, it indicates that the current design (or its combination) is particularly scarce and novel after timeliness weighting, and should receive bonus points; when it is negative, it indicates that the design or its combination is already relatively common or outdated, and should receive deduction points.
[0184] In some embodiments, the timeliness innovation correction value Δ = (1-β) × element scarcity index + β × combination scarcity index; wherein, the element scarcity index is mainly used to calculate the average decay frequency of all elements in the set, and the combination scarcity index is mainly used to calculate the average combination decay frequency of all unique design element pairs in the set.
[0185] Step S705: Combine the innovation score with the timeliness innovation correction value to generate the final innovation score.
[0186] Specifically, a dynamic calibration mechanism combines the initial innovation score calculated based on the original data with a correction value reflecting spatiotemporal dynamics to arrive at a more accurate final score that better aligns with current market perceptions. The system uses this timely innovation correction value as an adjustment coefficient, applying it to the initial innovation score through weighted, proportional, or other non-linear methods.
[0187] For example, if the correction value is positive, the initial score is increased; if it is negative, the initial score is decreased. This fusion mechanism is essentially a Bayesian update process that uses newly introduced, finer-grained spatiotemporal and combinatorial information (i.e., the correction value) to calibrate and optimize the prior estimate of innovation (i.e., the initial score), which may be biased due to data lag or model limitations. The resulting final innovation score is a comprehensive and intelligent evaluation result that simultaneously considers the absolute historical status of design elements, the time decay effect, and the combinatorial innovation potential.
[0188] In some embodiments, the generated timeliness innovation correction value Δ can be directly used in fusion strategies such as final innovation score = innovation score × (1 + Δ) or final innovation score = innovation score + k × Δ (k is an adjustment coefficient).
[0189] In the above implementation, an innovation evaluation framework integrating time decay quantification, combinatorial effect analysis, and dynamic Bayesian calibration is constructed, achieving intelligent, refined, and timely quantification of the novelty of clothing designs. First, by introducing a negative exponential decay model based on the time of first appearance, the system can dynamically perceive the natural decay of the innovative value of design elements, automatically synchronizing the evaluation results with the rapidly iterating fashion cycle and effectively distinguishing between "historical classics" and "current innovations." Second, by analyzing the combinational scarcity between design points through association rule mining, the system can identify and reward overall design innovations composed of common elements combined in a groundbreaking manner, overcoming the limitation of traditional methods that only evaluate single elements while ignoring the value of combinations. Finally, by transforming the above dynamic and combinatorial information into a timely innovation correction value and intelligently fusing and calibrating the initial innovation score, the system outputs a more robust and accurate final innovation score. This score not only reflects the scarcity of the design in history but also more accurately depicts its novelty and breakthrough potential in the current market context.
[0190] Reference Figure 8 As one implementation of step S110, the step of performing a grading decision based on the comprehensive score and quality score, and outputting the imported images and grade labels, includes:
[0191] Step S801: Obtain the overall score and the corresponding quality score;
[0192] The overall score is a macro-level indicator reflecting the overall competitiveness and potential of the apparel. It dynamically weights and integrates the "trend matching score" representing market popularity, the "innovation score" representing design uniqueness, and the "quality score". The quality score, on the other hand, is a relatively independent specialized assessment that quantifies the shooting and production quality of the image itself purely from the technical and aesthetic perspectives of visual presentation (such as composition, color, and clarity).
[0193] Step S802: Generate a set of grading thresholds based on preset dynamic grading rules; wherein, the dynamic grading rules adjust the threshold range in segments according to the quality score value.
[0194] Specifically, the logic of this step is to introduce an adaptive mechanism for the grading standards to address the objective differences in basic quality between image sets from different sources.
[0195] Specifically, the dynamic grading rules dynamically adjust the comprehensive score thresholds required to classify each level (S, A, B, etc.) based on the range to which the current image's quality score falls. For example, when processing a batch of high-quality images from professional studios (with quality scores generally ≥90), the system automatically adopts a stricter grading threshold (e.g., S-level requires ≥85 points). This is because at this "high starting line" of high quality, a higher comprehensive design score is needed to stand out, preventing an overabundance of "S-level" images in the pool of high-quality images and ensuring their top-tier status. Conversely, when processing user-generated content (with quality scores potentially <80), the system uses another adjusted threshold (e.g., S-level requires ≥92 points, with a moderately relaxed threshold for B-level). This not only selects truly outstanding designs from a generally lower quality pool but also avoids burying potentially promising designs due to poor image quality. This step essentially gives the grading standards context-awareness, ensuring the comparability and fairness of rating results across datasets at different quality levels.
[0196] Step S803: Match the comprehensive score with the set of grading thresholds to determine the initial grade identifier;
[0197] After obtaining a grading standard that fits the current image quality context, this step performs preliminary quantitative classification, using the comprehensive score as the core evaluation indicator and directly comparing it with a dynamically generated set of grading thresholds.
[0198] Specifically, the system compares the overall score value sequentially with the thresholds for "S," "A," and "B" grades to determine the score range the image falls into, thus assigning the image an initial grade label (such as "S," "A," or "B"). This initial grade reflects the garment's relative position within the current batch of images after comprehensively considering factors such as trendiness, innovation, and current quality level. It is the first and most basic grade judgment based on algorithmic quantitative scoring, providing a starting point for subsequent fine-tuning.
[0199] Step S804: Perform a correction operation on the initial grade label based on the quality score to generate the final grade label;
[0200] This step is a fine-tuning and safety correction stage in the decision-making process. Its logic lies in conducting a secondary review of the preliminary results based on quality dimensions to correct extreme cases that might be masked by weighted fusion of scores. The initial grade is entirely determined by the comprehensive score, which already partially includes the quality score; however, this inclusion is a weighted average and may have a "dilution" effect. The correction operation directly examines whether the quality score matches the initial grade.
[0201] For example, the correction operation includes: if the quality score is lower than the minimum allowed value of the current initial level, the initial level is lowered by one level; if the quality score is higher than the maximum allowed value of the current initial level by 20%, the initial level is raised by one level.
[0202] For example, if a garment scores extremely high in design and trendiness, achieving an overall score of S, but its image is exceptionally blurry (very low quality score), this does not meet the high standards expected of an "S-level" premium item. In this case, the correction rules will downgrade its grade because its quality score is below the minimum acceptable standard for S-level. Conversely, if a garment's overall score is just A-level, but its photographic quality is near perfect (quality score far exceeding the average A-level), its outstanding presentation greatly enhances its commercial value, and the correction rules can upgrade its grade. This step ensures that the final grade does not deviate significantly from the hard indicator of "basic image presentation quality," thus producing a final grade label that more robustly and comprehensively reflects the overall value of the image.
[0203] Step S805: Based on preset warehousing conditions, determine the clothing images that meet the preset warehousing conditions according to the final grade label and quality score;
[0204] In some embodiments, the preset entry conditions are configured as follows: the final grade label is B or above, and the quality score value is ≥ the entry quality threshold (default 60).
[0205] Understandably, the final grade label must reach a certain level or higher (such as Grade B or above), ensuring that the clothing included in the database possesses basic commercial and design value. Secondly, the quality score must exceed an absolute quality threshold for inclusion in the database (such as 60 points), which technically prevents any images with excessively low clarity or that are unusable from being included in the resource library. Only images that meet both of these conditions can be considered qualified resources for inclusion in the database.
[0206] Step S806: Associate the clothing images that meet the preset entry conditions with the corresponding final grade tags to generate a set of entry images and grade tags.
[0207] The system binds clothing images that meet preset entry criteria to their corresponding final grade tags, generating a structured set of entry images and grade tags. This set is the final output of the entire intelligent screening system. Each record contains a unique identifier for the image and its grade obtained through dynamic, collaborative, and rigorous evaluation. It can be directly used to drive downstream business applications such as product recommendation, priority ranking, and distribution across different channels, forming a complete closed loop from raw images to manageable and graded digital assets.
[0208] The above implementation achieves automated, intelligent, and highly reliable final evaluation and screening of clothing images. First, the innovative dynamic grading rules enable the system to intelligently perceive the overall quality distribution of the input image set and adaptively adjust the grading standards, ensuring fairness and consistency of rating results across different quality batches and solving the problem of standard inaccuracy when using fixed thresholds with heterogeneous data. Second, the introduction of a grading correction operation based on quality scores establishes an effective feedback and calibration mechanism that can correct biases such as "high score, low quality" or "high quality being underestimated" that may result from solely relying on comprehensive scores, ensuring that the final grading label affirms design value while firmly maintaining the bottom line of visual presentation quality. Finally, by setting hard inclusion conditions that include both grading and quality requirements, it ensures that every image in the output set simultaneously meets the two core requirements of "valuable" and "usable." This technical solution transforms complex commercial grading decisions into a stable, transparent, and explainable automated process, greatly improving the efficiency and intelligence of managing massive clothing image resources and providing crucial technical support for building a high-quality product visual asset library.
[0209] This application also discloses an intelligent image filtering system for clothing.
[0210] A smart image filtering system for clothing, specifically including:
[0211] The data acquisition module is used to acquire the dataset of clothing images to be processed.
[0212] The preprocessing module is used to perform basic quality filtering operations on the clothing image dataset, generating a set of filtered images based on preset legality rules and image quality thresholds;
[0213] The standardization processing module is used to perform image enhancement processing on the selected image set to generate a standardized image set.
[0214] The feature extraction and evaluation module is used to extract multi-dimensional visual features from a standardized set of images, generating a visual feature vector and quality score for each image.
[0215] The main body segmentation module is used to identify the main body region of clothing based on visual feature vectors and a pre-trained clothing region segmentation model, and generate a clothing region mask.
[0216] The design element parsing module is used to combine clothing area masks and visual feature vectors to extract semantic tags of clothing design elements and generate a set of structured design points.
[0217] The trend matching module is used to obtain a pre-built clothing trend library, perform similarity matching between the structured design point set and the clothing trend library, and generate a trend matching score.
[0218] The innovation calculation module is used to calculate the design innovation based on the historical occurrence frequency of the structured design point set and generate an innovation score.
[0219] The comprehensive score calculation module is used to dynamically weight and calculate the comprehensive score by integrating the quality score, trend matching score, and innovation score.
[0220] The decision and output module is used to make grading decisions based on the comprehensive score and quality score, and output the images and grade labels to the database.
[0221] The intelligent clothing image filtering system of this application embodiment can implement any of the above methods, and the specific working process of each module in the system can refer to the corresponding process in the above method embodiments.
[0222] In the several embodiments provided in this application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for example, the division of a certain module is merely a logical functional division, and in actual implementation there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0223] This application also discloses a computer-readable storage medium.
[0224] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above in any of the intelligent image filtering methods for clothing.
[0225] The computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device; the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0226] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.
Claims
1. A method for intelligent screening of clothing pictures, characterized in that, The screening method includes: Obtain the dataset of clothing images to be processed; Basic quality filtering is performed on the clothing image dataset to generate a set of filtered images based on preset legality rules and image quality thresholds; Image enhancement processing is performed on the selected image set to generate a standardized image set; Multi-dimensional visual feature extraction is performed on the standardized image set to generate a visual feature vector and quality score for each image; Based on the aforementioned visual feature vectors, the main body region of the clothing is identified through a pre-trained clothing region segmentation model, and a clothing region mask is generated. By combining the clothing area mask with the visual feature vector, semantic tags of clothing design elements are extracted to generate a set of structured design points; Obtain a pre-built clothing trend library, perform similarity matching between the structured design point set and the clothing trend library, and generate a trend matching score; The design innovation degree is calculated based on the historical occurrence frequency of the structured design point set, and an innovation score is generated. The quality score, trend matching score, and innovation score are dynamically weighted and calculated to generate a comprehensive score. Based on the comprehensive score and the quality score, a grading decision is made, and the images to be added to the database and their grade labels are output.
2. The method of claim 1, wherein, Performing basic quality filtering on the clothing image dataset, and generating a set of filtered images based on preset validity rules and image quality thresholds, includes the following steps: Obtain the dataset of clothing images to be processed, along with a pre-defined set of legality rules and image quality thresholds; Perform binary data parsing on each clothing image in the clothing image dataset to generate file header feature codes and metadata structures; Based on the file header feature code and the set of legality rules, format verification is performed to generate a legality verification mark; Image resolution parameters are extracted from the metadata structure and compared with resolution thresholds in the image quality threshold set to generate resolution detection markers. For each clothing image, the image pixel matrix is used to extract sharpness features, calculate the blur index and compare it with the sharpness threshold to generate a sharpness detection label; For each clothing image, the noise density of the image pixel matrix is calculated and compared with the noise threshold to generate noise detection labels; A comprehensive quality detection mark is generated by combining the aforementioned legality verification mark, resolution detection mark, sharpness detection mark, and noise detection mark; Based on the comprehensive quality inspection markers, images that meet all thresholds are selected to generate a set of filtered images.
3. The method of claim 1, wherein, The steps for extracting multi-dimensional visual features from the standardized image set and generating a visual feature vector and quality score for each image include: The standardized image set is processed by a pre-trained feature extraction model to generate a primary feature tensor. Spatial attention weighting computation is performed on the primary feature tensor to generate an enhanced feature map; The enhanced feature map is input into the quality scoring branch network, and the output includes a multi-dimensional quality index including compositional balance, color contrast, and texture clarity. A quality score is generated by linear weighted aggregation based on the multidimensional quality indicators. The enhanced feature map is input into the feature encoding branch network, and a fixed-dimensional visual feature vector is generated through dimensionality reduction. Output the visual feature vector and quality score for each image.
4. The method of claim 3, wherein, The steps for extracting semantic tags of clothing design elements and generating a set of structured design points by combining the clothing region mask and the visual feature vector include: Based on the clothing area mask, spatial region clipping is performed on the visual feature vector to generate a clothing main feature map; Perform multi-level feature fusion operations on the main feature map of the clothing to generate multi-scale fused features; The multi-scale fusion features are input into a pre-trained design element recognition model, which outputs an initial set of semantic labels and a confidence distribution. The initial semantic tag set is filtered according to a preset confidence threshold to generate a valid semantic tag set; The set of valid semantic tags is structured and encoded according to a predefined design element classification system to generate a set of structured design points.
5. The method of claim 4, wherein, The steps of obtaining a pre-built clothing trend library, performing similarity matching between the structured design point set and the clothing trend library, and generating a trend matching score include: Obtain a pre-built clothing trend library and a set of structured design points for the current clothing image; The structured design point set is converted into design point feature vectors according to a preset vector encoding rule; Extract the trend feature vectors and time decay weights of each trend cluster from the clothing trend library; Calculate the similarity value between the design point feature vector and each trend feature vector to generate a similarity distribution sequence; The similarity distribution sequence is weighted and corrected according to the time decay weight to generate a weighted similarity sequence; Perform a maximum value aggregation operation on the weighted similarity sequence to generate a trend matching score.
6. The method of claim 1 to 5, wherein, The steps for calculating the design innovation degree and generating an innovation score based on the historical occurrence frequency of the structured design point set include: Obtain the set of structured design points and the pre-built database of historical design elements; Extract the time decay frequency weight of each design element from the historical design element database; Perform a historical frequency query on each design element in the structured design point set to generate a base frequency value; The base frequency value is corrected according to the time decay frequency weight to generate a weighted historical frequency value; The weighted historical frequency values of all design elements in the structured design point set are aggregated and calculated to generate a combinatorial innovation index. The combined innovation indicators are mapped to a preset scoring range to generate an innovation score.
7. The intelligent screening method for clothing images according to claim 6, further comprising, after generating the innovation score: Obtain the set of structured design points and the pre-built database of historical design elements; Extract the first occurrence timestamp and historical occurrence frequency of each design element from the historical design element database; Calculate the time decay factor based on the current time and the first occurrence timestamp, and generate the decayed historical frequency; The structured design point set is matched with the attenuated historical frequency to generate a timeliness innovation correction value; The innovation score is combined with the timeliness innovation correction value to generate the final innovation score.
8. The method of claim 1, wherein, The steps of performing a grading decision based on the comprehensive score and the quality score, and outputting the imported images and grade labels, include: Obtain the overall score and the corresponding quality score; A set of grading thresholds is generated based on preset dynamic grading rules; wherein, the dynamic grading rules adjust the threshold range in segments according to the quality score value. The comprehensive score value is matched with the set of grading thresholds to determine the initial grade identifier; Based on the quality score, the initial grade label is corrected to generate the final grade label; Based on preset entry conditions, clothing images that meet the preset entry conditions are determined according to the final grade label and quality score value; The images of clothing that meet the preset entry conditions are associated with the corresponding final grade tags to generate a set of entry images and grade tags.
9. An intelligent screening method of clothing pictures, characterized in that, The screening method includes: The data acquisition module is used to acquire the dataset of clothing images to be processed. The preprocessing module is used to perform basic quality filtering operations on the clothing image dataset, and generate a set of filtered images based on preset legality rules and image quality thresholds; The standardization processing module is used to perform image enhancement processing on the selected image set to generate a standardized image set; The feature extraction and evaluation module is used to extract multi-dimensional visual features from the standardized image set and generate a visual feature vector and quality score for each image. The main body segmentation module is used to identify the main body region of the clothing based on the visual feature vector using a pre-trained clothing region segmentation model, and generate a clothing region mask. The design element parsing module is used to combine the clothing area mask with the visual feature vector to extract semantic tags of clothing design elements and generate a set of structured design points. The trend matching module is used to obtain a pre-built clothing trend library, perform similarity matching between the structured design point set and the clothing trend library, and generate a trend matching score. The innovation calculation module is used to calculate the design innovation based on the historical occurrence frequency of the structured design point set and generate an innovation score. The comprehensive score calculation module is used to dynamically weight the quality score, trend matching score, and innovation score to generate a comprehensive score. The decision and output module is used to perform a grading decision based on the comprehensive score and the quality score, and output the images to be stored and grade labels.
10. A computer-readable storage medium, characterized in that: The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 8.