Clothing recognition method, apparatus, device, and medium

By classifying clothing images, detecting brands and defects using clothing recognition methods, and combining them with type recognition models for style and color identification, this approach solves the problems of errors and inefficiency caused by reliance on human experience in traditional laundry services, thereby improving recognition accuracy and efficiency.

CN119006874BActive Publication Date: 2025-10-17SHENZHEN HIVE BOX NETWORK TECH LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410817140.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2025-10-17
Estimated Expiration
2044-06-21

AI Technical Summary

Technical Problem

In the traditional laundry service industry, clothing sorting relies on manual experience, which is prone to errors and inefficient.

Method used

The clothing recognition method uses a detection model to classify clothing images, detect brands and defects, and combines a type recognition model to identify styles and colors, generating clothing recognition results.

Benefits of technology

It improved the accuracy of clothing brand and defect identification, enhanced clothing classification efficiency, and enabled accurate identification of clothing style and color.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119006874B_ABST
    Figure CN119006874B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of clothes recognition, and discloses a clothes recognition method, device, equipment and medium, the method comprising: acquiring at least one to-be-recognized image corresponding to clothes to be recognized, and performing image classification recognition on all to-be-recognized images through a detection model to obtain a classification detection result, a brand detection result and a defect detection result; performing clothes image recognition on all to-be-recognized images through a type recognition model corresponding to the classification detection result to obtain a style recognition result and a color recognition result; and determining a clothes recognition result corresponding to the clothes to be recognized according to the brand detection result, the style recognition result, the color recognition result and the defect detection result. In the present application, clothes images are recognized through a model, the classification of clothes is realized, the accurate recognition of clothes color and clothes style is realized, the recognition accuracy of clothes brand and clothes defects is improved, and the classification efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of clothes recognition, and in particular to a clothes recognition method, device, equipment and medium. BACKGROUND

[0002] In the traditional laundry service industry, after a user goes to a laundry shop, the user needs to place an order after communicating with the staff of the laundry shop, and then the staff classifies the clothes or sends the clothes to a laundry factory after packaging, or washes the clothes in the self-owned shop. The whole process needs manual participation, and the experience of the staff is required to be high. If the experience of the staff is insufficient, not only errors (such as incorrect classification of clothes) are prone to occur, but also the efficiency is low. SUMMARY

[0003] The present application provides a clothes recognition method, device, equipment and medium to solve the problem that the experience of the staff is required to be high in the laundry service industry in the prior art, which leads to errors and low efficiency.

[0004] A clothes recognition method comprises the following steps.

[0005] At least one to-be-recognized image corresponding to to-be-recognized clothes is acquired, and all the to-be-recognized images are subjected to image classification recognition through a detection model to obtain a classification detection result, a brand detection result and a defect detection result corresponding to the to-be-recognized clothes.

[0006] A clothes image recognition is performed on all the to-be-recognized images through a type recognition model corresponding to the classification detection result to obtain a style recognition result and a color recognition result.

[0007] A clothes recognition result corresponding to the to-be-recognized clothes is determined according to the brand detection result, the style recognition result, the color recognition result and the defect detection result.

[0008] A clothes recognition device comprises the following.

[0009] An image classification recognition module is configured to acquire at least one to-be-recognized image corresponding to to-be-recognized clothes, and to perform image classification recognition on all the to-be-recognized images through a detection model to obtain a classification detection result, a brand detection result and a defect detection result corresponding to the to-be-recognized clothes.

[0010] A clothes image recognition module is configured to perform a clothes image recognition on all the to-be-recognized images through a type recognition model corresponding to the classification detection result to obtain a style recognition result and a color recognition result.

[0011] The clothes recognition result module is configured to determine a clothes recognition result corresponding to the clothes to be recognized according to the brand detection result, the style recognition result, the color recognition result, and the defect detection result.

[0012] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor is configured to execute the clothes recognition method.

[0013] A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the clothes recognition method.

[0014] The clothes recognition method, device, equipment, and medium, the clothes recognition method of the present application classifies and recognizes all images to be recognized through a detection model, accurately recognizes the classification detection result, the brand detection result, and the defect detection result, and further improves the recognition accuracy of the clothes brand and the clothes defect and improves the classification efficiency. The clothes image recognition of all images to be recognized is performed through the type recognition model corresponding to the classification detection result, accurately recognizes the style and color of the clothes, obtains the style recognition result and the color recognition result, and further improves the recognition accuracy of the clothes style and the clothes color. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0016] Figure 1 is a flowchart of the clothes recognition method in an embodiment of the present application;

[0017] Figure 2 is a flowchart of step S10 of the clothes recognition method in an embodiment of the present application;

[0018] Figure 3 is a flowchart of step S101 of the clothes recognition method in an embodiment of the present application;

[0019] Figure 4 is a flowchart of step S102 of the clothes recognition method in an embodiment of the present application;

[0020] Figure 5 is a flowchart of step S20 of the clothes recognition method in an embodiment of the present application;

[0021] Figure 6is a flow chart of step S201 of the clothes recognition method in an embodiment of the present application;

[0022] Figure 7 is a flow chart of step S103 of the clothes recognition method in an embodiment of the present application;

[0023] Figure 8 is a principle block diagram of the clothes recognition device in an embodiment of the present application. DETAILED DESCRIPTION

[0024] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0025] In an embodiment, as shown in Figure 1 a clothes recognition method is provided, comprising the following steps:

[0026] S10: At least one image to be recognized corresponding to the clothes to be recognized is obtained, and all the images to be recognized are subjected to image classification recognition through a detection model to obtain classification detection results, brand detection results and defect detection results corresponding to the clothes to be recognized.

[0027] Understandably, the image to be recognized refers to a picture taken after the clothes. The clothes to be recognized refers to the clothes or trousers received for cleaning. The classification detection results refer to the results of identifying the type of clothes, such as shirts, trousers, skirts or shoes. The brand detection results refer to the results of identifying the brand of clothes. The defect detection results refer to the results of identifying the damage or stains on the clothes. The detection model is obtained by training the yolov8 network or the yolov8 improved network corresponding to the image data.

[0028] Specifically, after receiving the clothes to be identified, the front and back and the inside and outside of the clothes to be identified are respectively photographed by the two groups of lenses, so as to obtain at least one to-be-identified image corresponding to the clothes to be identified. Then, a plurality of trained detection models, i.e., a classification detection model, a brand detection model and a defect detection model, are obtained, and all the to-be-identified images corresponding to the clothes to be identified are input into different detection models, and all the to-be-identified images are subjected to image classification recognition by different detection models, i.e., all the to-be-identified images are subjected to classification detection by the classification detection model, so as to obtain a classification detection result, all the to-be-identified images are subjected to brand detection by the brand detection model, so as to obtain a brand detection result, and all the to-be-identified images are subjected to defect detection by the defect detection model, so as to obtain a defect detection result. In this way, the classification detection result, the brand detection result and the defect detection result of the clothes to be identified can be obtained.

[0029] The defect detection model adopts the same network structure as the brand detection model, and different training samples are used in training the model, so that the contents recognized by the model are different. The processing process of the defect detection model will not be described in detail. The defect detection model outputs a prediction probability of each region, and the region with a higher probability is determined as a defect, so as to obtain a defect detection result corresponding to the clothes to be identified.

[0030] S20: The clothes image recognition is performed on all the to-be-identified images by a type recognition model corresponding to the classification detection result, so as to obtain a style recognition result and a color recognition result.

[0031] Understandably, the type recognition model refers to a model obtained by training an EfficientNet network by corresponding type images. The style recognition result refers to the style of the clothes to be identified, for example, the style of a shirt: down jacket, windbreaker, jeans jacket, leather jacket, suit, sweater, vest, shirt, etc.; the style of shoes: leather shoes, casual shoes, canvas shoes, sports shoes, sandals, cotton shoes, slippers, etc.; the style of trousers and skirts: dress, short skirt, half-length skirt, jeans, shorts, overall, down trousers, etc.; the style of home furnishing: bed sheet, quilt, towel, cushion, etc. The color recognition result refers to the color of the clothes to be identified, for example, white, black, blue, red or combined color, etc.

[0032] Specifically, the type of clothing to be identified is determined by the classification detection results, and the type recognition model corresponding to the type is obtained based on the type, for example, a top recognition model, a culottes recognition model, a shoe recognition model, and a home recognition model. Then, the image to be identified is preprocessed, that is, the corresponding image to be identified is cropped using the detection frame corresponding to the clothing to be identified output by the classification detection model, thereby obtaining a target image that only includes clothing; then, clothing image recognition is performed on all images to be identified using the type recognition model corresponding to the classification detection results, that is, the style and color corresponding to each image to be identified are identified, and the style and color of all results corresponding to the same clothing to be identified are determined, thereby obtaining style recognition results and color recognition results. It can be understood that the detection model outputs detection frames corresponding to categories and clothing, detection frames corresponding to brands and logos, and detection frames corresponding to defects and flaws.

[0033] S30: Determine a clothing recognition result corresponding to the clothing to be recognized based on the brand detection result, the style recognition result, the color recognition result, and the defect detection result.

[0034] Understandably, the clothing recognition result is used to represent the brand, style and defect of the clothing to be recognized.

[0035] Specifically, after identifying the color and style, the information of the clothing to be identified is recorded through the brand detection results, style recognition results, color recognition results and defect detection results, thereby obtaining a document including the color, style, brand and defects of the clothing to be identified, that is, the clothing identification result corresponding to the clothing to be identified.

[0036] In the clothing recognition method of the present invention, all images to be identified are classified and identified using a detection model, achieving accurate recognition of classification detection results, brand detection results, and defect detection results, thereby improving the recognition accuracy of clothing brands and clothing defects and improving classification efficiency. Clothing image recognition is performed on all images to be identified using a type recognition model corresponding to the classification detection results, achieving accurate recognition of clothing styles and colors, and enabling the acquisition of style recognition results and color recognition results, thereby improving the recognition accuracy of clothing styles and colors.

[0037] In one embodiment, if Figure 2 As shown, in step S10, the detection model is used to perform image classification and recognition on all the images to be identified to obtain classification detection results, brand detection results, and defect detection results corresponding to the clothing to be identified. The detection model includes a classification detection model, a brand detection model, and a defect detection model; including:

[0038] S101, performing classification detection on all the to-be-recognized images through the classification detection model to obtain a classification detection result corresponding to the to-be-recognized clothes.

[0039] S102, performing brand detection on all the to-be-recognized images through the brand detection model to obtain a brand detection result corresponding to the to-be-recognized clothes.

[0040] S103, performing defect detection on all the to-be-recognized images through the defect detection model to obtain a defect detection result corresponding to the to-be-recognized clothes.

[0041] Understandably, the classification detection result is used to represent the category to which the to-be-recognized clothes belong, for example, shirts and trousers, etc. The brand detection result is used to represent the brand to which the to-be-recognized clothes belong, for example, TEBERON, Li Ning or Puma, etc. The defect detection result is used to represent whether the to-be-recognized clothes have stains or damage, etc.

[0042] Specifically, after obtaining the to-be-recognized images, the classification detection model, the brand detection model and the defect detection model are obtained, and all the to-be-recognized images are input into the classification detection model, the brand detection model and the defect detection model. Then, the classification detection model is used to perform classification detection on all the to-be-recognized images, that is, the yolov8 model is used to perform classification recognition on all the to-be-recognized images, that is, the classification ability learned by the classification detection model during training is used to detect all the to-be-recognized images, so as to obtain the detection frame corresponding to the clothes in the to-be-recognized images and the classification detection result corresponding to the clothes. Similarly, the brand detection model is used to detect all the to-be-recognized images, that is, the yolov8 model is used to perform brand recognition on all the to-be-recognized images, that is, the recognition ability learned by the brand detection model during training is used to detect the frame of the logo in all the to-be-recognized images, to obtain the detection frame corresponding to the logo on the clothes and the brand detection result. Further, the defect detection model is used to perform defect detection on all the to-be-recognized images, that is, the yolov8 model is used to perform defect recognition on all the to-be-recognized images, that is, the recognition ability learned by the defect detection model during training is used to detect the defect in all the to-be-recognized images, to obtain the detection frame corresponding to the defect on the clothes and the defect detection result.

[0043] In this embodiment, the classification detection model is used to perform classification detection on all the to-be-recognized images, so as to determine the category to which the to-be-recognized clothes belong, and obtain the classification detection result. The brand detection model is used to perform brand detection on all the to-be-recognized images, so as to determine the brand to which the to-be-recognized clothes belong, and obtain the brand detection result. The defect detection model is used to perform defect detection on all the to-be-recognized images, so as to determine the position of the defect in the to-be-recognized clothes, and obtain the defect detection result.

[0044] In one embodiment, if Figure 3 As shown, in step S101, the classification detection model is used to perform classification detection on all the images to be identified to obtain classification detection results corresponding to the clothing to be identified, including:

[0045] S1011, performing feature extraction of different scales on all the images to be identified through the first backbone network in the classification detection model to obtain first extracted features.

[0046] S1012: Perform feature fusion of different dimensions on all the first extracted features through the Neck network in the classification detection model to obtain a first fused feature.

[0047] S1013: Perform prediction processing on all the first fusion features through the first recognition network in the classification detection model to obtain a classification detection result.

[0048] It can be understood that the classification detection model refers to a model for classifying clothing types, which is trained based on the yolov8 network. The first extracted features refer to features extracted at different scales. The first fused features refer to features fused at different dimensions.

[0049] Specifically, after obtaining the image to be identified, a classification detection model is obtained, and all images to be identified are input into the classification detection model. Feature extraction is performed on the image to be identified through the first backbone network in the classification detection model, that is, the image to be identified is convolved through the two-dimensional convolution, normalization and SiLU activation function in the convolution module to reduce the dimension and enhance the nonlinear ability, thereby obtaining the first convolution feature. Then, the first convolution feature is extracted through the C2f module, that is, the number of channels is first reduced, and then the features are extracted through the first convolution, and the input and output are spliced ​​using residual connections, and finally the number of channels is restored to obtain the first extracted features. Then, the first extracted features of different scales are extracted through the convolution module and the C2f module. Among them, after the last layer of convolution module and C2f module, the first extracted features of the previous dimension are continuously subjected to multiple maximum pooling and residual connections through the SPPF module, thereby obtaining the first extracted features of different scales.

[0050] Furthermore, the Neck network in the classification detection model is used to perform feature fusion processing on the first extracted features of different scales, that is, the first extracted feature of the last dimension is upsampled and fused with the first extracted feature of the previous dimension to obtain the first feature, and the first feature is passed through the C2f layer and upsampled and fused with the first extracted feature of the previous dimension to obtain the first dimension feature. The first dimension feature is passed through the convolution layer and fused with the first feature to obtain the second feature, the second feature is passed through the C2f layer to obtain the second dimension feature, the second dimension feature is passed through the convolution layer and fused with the first extracted feature of the last dimension to obtain the third feature, the third feature is passed through the C2f layer to obtain the third dimension feature, and the three dimensional features are determined as the first fused feature.

[0051] Next, the first fused feature is input into the first recognition network of the classification detection model. The first recognition network performs image prediction on each dimension feature in the first fused feature, that is, the detect layer performs category prediction on each dimension feature, and then uses threshold filtering to eliminate low-confidence prediction results and non-maximum suppression to eliminate redundant categories, thereby obtaining the classification detection result corresponding to the image to be identified. Among them, the working principle of non-maximum suppression is: for each target category, the category with the highest confidence is selected as a reference, and only the most likely category is retained.

[0052] In this embodiment, the first backbone network extracts features of different scales from the image to be identified, thereby determining the first extracted features. The Neck network fuses features of different dimensions, thereby determining the first fused features. The first recognition network recognizes the first fused features, thereby obtaining classification detection results, thereby improving the accuracy of clothing classification.

[0053] In one embodiment, if Figure 4 As shown, in step S102, brand detection is performed on all the images to be identified by the brand detection model to obtain brand detection results corresponding to the clothing to be identified, including:

[0054] S1021: Perform feature extraction of different scales on all the images to be identified through the second backbone network in the brand detection model to obtain second extracted features.

[0055] S1022: Perform feature fusion of different dimensions on all the second extracted features through the BiFPN network in the brand detection model to obtain a second fused feature.

[0056] S1023: Perform prediction processing on all the second fusion features through the second recognition network in the brand detection model to obtain a brand detection result.

[0057] It can be understood that the brand detection model refers to a model for identifying a clothing brand, which is improved based on a yolov8 network.

[0058] Specifically, after obtaining the to-be-identified images, a brand detection model is acquired, and all the to-be-identified images are input into the brand detection model. The second backbone network in the brand detection model is used to extract features of the to-be-identified images, that is, the to-be-identified images are subjected to convolution processing through two-dimensional convolution, normalization and SiLU activation function in the convolution module, so as to reduce the dimension and enhance the nonlinear capability, thereby obtaining second convolution features. Then, the C2f module is used to extract features of the second convolution features, that is, the channel number is first reduced, then the second convolution is used to extract features, and the residual connection is used to splice the input and the output, and finally the channel number is restored, thereby obtaining second extraction features. Then, the second convolution module and the C2f module are used to extract second extraction features of different scales. After the last convolution module and the C2f module, the SPPF module is used to perform continuous maximum pooling and residual connection on the second extraction features of the last dimension, thereby obtaining second extraction features of different scales.

[0059] Further, the BiFPN network in the brand detection model is used to fuse features of different dimensions of all the second extraction features, that is, the extraction features of the second layer, the third layer, the fourth layer and the fifth layer in the second backbone network are subjected to information exchange paths from top to bottom and from bottom to top. The information exchange paths from top to bottom are subjected to information fusion in the third layer and the fourth layer, and the information of the fused upper layer features is subjected to feature fusion with the information of the layer and the lower layer from bottom to top through the skip connection, thereby obtaining second fusion features.

[0060] In a specific embodiment, the extracted features of p2, p3, p4 and p5 extracted by the second backbone network are obtained, the extracted features of p4 and p5 are fused in the first p4 fusion layer (i.e. only the middle two layers include fusion layers), to obtain the first fusion feature of p4, and the first fusion feature of p4 is transmitted to the second p4 fusion layer and the first p3 fusion layer, the first fusion feature of p4 and the extracted feature of p3 are fused in the first p3 fusion layer to obtain the first fusion feature of p3, and the first fusion feature of p3 is transmitted to the second p3 fusion layer and the second p2 fusion layer. In the second p2 fusion layer, the extracted feature of p2 and the first fusion feature of p3 are fused to obtain the second fusion feature of p2, and the second fusion feature of p2 is transmitted to the second recognition network and the second p3 fusion layer. In the second p3 fusion layer, the second fusion feature of p2, the first fusion feature of p3 and the extracted feature of p3 transmitted by jumping are fused to obtain the second fusion feature of p3, and the second fusion feature of p3 is transmitted to the second recognition network and the second p4 fusion layer. In the second p4 fusion layer, the second fusion feature of p3, the first fusion feature of p4 and the extracted feature of p4 transmitted by jumping are fused to obtain the second fusion feature of p4, and the second fusion feature of p4 is transmitted to the second recognition network and the second p5 fusion layer. In the second p5 fusion layer, the extracted feature of p5 and the second fusion feature of p4 are fused to obtain the second fusion feature of p5, and the second fusion feature of p5 is transmitted to the second recognition network.

[0061] Then, the second fusion feature is input into the second recognition network of the brand detection model, and each dimension feature in the second fusion feature is image predicted by the second recognition network, that is, each dimension feature is brand predicted by the detect layer, and low confidence prediction results are filtered by threshold, and redundant brands are eliminated by non-maximum suppression, so as to obtain the brand detection result corresponding to the image to be recognized. The working principle of non-maximum suppression is that for each target brand, the brand with the highest confidence is selected as a reference, and only one most likely brand is reserved.

[0062] In this embodiment, the second backbone network realizes the extraction of different scale features in the image to be recognized, and realizes the determination of the second extracted feature. The BiFPN network realizes the fusion of different dimension features, realizes the determination of the second fusion feature, and realizes the fusion of low dimension features. The second recognition network realizes the recognition of the second fusion feature, realizes the acquisition of the brand detection result, and further improves the accuracy of the clothing brand recognition.

[0063] In an embodiment, as Figure 5As shown, in step S20, the clothes image recognition is performed on all the to-be-recognized images by using a type recognition model corresponding to the classification detection result, to obtain a style recognition result and a color recognition result, the type recognition model including a top recognition model, a pants / skirt recognition model, a shoe recognition model, and a home recognition model; including:

[0064] S201, when the classification detection result indicates that the to-be-recognized clothes is a top, performing image recognition on all the to-be-recognized images by using the top recognition model, to obtain a top color recognition result and a top style recognition result.

[0065] S202, when the classification detection result indicates that the to-be-recognized clothes is a pants / skirt, performing image recognition on all the to-be-recognized images by using the pants / skirt recognition model, to obtain a pants / skirt color recognition result and a pants / skirt style recognition result.

[0066] S203, when the classification detection result indicates that the to-be-recognized clothes is a shoe, performing image recognition on all the to-be-recognized images by using the shoe recognition model, to obtain a shoe color recognition result and a shoe style recognition result.

[0067] S204, when the classification detection result indicates that the to-be-recognized clothes is a home, performing image recognition on all the to-be-recognized images by using the home recognition model, to obtain a home color recognition result and a home style recognition result.

[0068] Understandably, the jacket color recognition result refers to the color result recognized when the to-be-recognized clothes are jackets, for example, white, black, or a combination of colors, such as blue and white. The jacket style recognition result refers to the style result recognized when the to-be-recognized clothes are jackets, for example, down jacket, windbreaker, jeans jacket, leather jacket, suit, sweater, vest, shirt, and the like. The skirt color recognition result refers to the color result recognized when the to-be-recognized clothes are skirts, for example, white, black, or a combination of colors, such as gray and white. The skirt style recognition result refers to the style result recognized when the to-be-recognized clothes are skirts, for example, dress, short skirt, half-length skirt, jeans, shorts, jumpsuit, down pants, and the like. The shoe color recognition result refers to the color result recognized when the to-be-recognized clothes are shoes, for example, black, white, brown, or gray, or a combination of colors. The shoe style recognition result refers to the style result recognized when the to-be-recognized clothes are shoes, for example, leather shoes, casual shoes, canvas shoes, sports shoes, sandals, cotton shoes, slippers, and the like. The home color recognition result refers to the color result recognized when the to-be-recognized clothes are home, for example, black, white, brown, or gray, or a combination of colors. The home style recognition result refers to the style result recognized when the to-be-recognized clothes are home, for example, bed sheet, quilt, towel, cushion, and the like. The jacket recognition model, the skirt recognition model, the shoe recognition model, and the home recognition model are all obtained by training the EfficientNet network with different training data. The network structure of the category recognition model can be the same, or different network structures can be trained to obtain the recognition model.

[0069] Specifically, after obtaining the classification test results, if the classification test results indicate that the clothing item to be identified is a top, a top recognition model is obtained and image recognition is performed on all images to be identified using the top recognition model. Specifically, the top recognition model is used to identify the color and style of the top in the image to be identified. Specifically, the recognition capabilities learned during training are used to identify the color and style of the top. The results are then integrated and confirmed based on the recognition results of all images to be identified, thereby obtaining a top color recognition result and a top style recognition result corresponding to the clothing item to be identified. Similarly, if the classification test results indicate that the clothing item to be identified is a culottes, a culottes recognition model is obtained and image recognition is performed on all images to be identified using the culottes recognition model. Specifically, the culottes recognition model is used to identify the color and style of the culottes in the image to be identified. Specifically, the recognition capabilities learned during training are used to identify the color and style of the culottes. The results are then integrated and confirmed based on the recognition results of all images to be identified, thereby obtaining a culottes color recognition result and a culottes style recognition result corresponding to the clothing item to be identified. Furthermore, when the classification detection result indicates that the clothing to be identified is a shoe, a shoe recognition model is obtained, and image recognition is performed on all images to be identified using the shoe recognition model. That is, the color and style of the shoes in the image to be identified are recognized using the shoe recognition model, that is, the color and style of the shoes are recognized using the recognition ability learned during training, and the results are integrated and confirmed based on the recognition results of all images to be identified, thereby obtaining shoe color recognition results and shoe style recognition results corresponding to the clothing to be identified. Similarly, when the classification detection result indicates that the clothing to be identified is a home furnishing, a home furnishing recognition model is obtained, and image recognition is performed on all images to be identified using the home furnishing recognition model. That is, the color and style of the home furnishings in the image to be identified are recognized using the home furnishing recognition model, that is, the color and style of the home furnishings are recognized using the recognition ability learned during training, and the results are integrated and confirmed based on the recognition results of all images to be identified, thereby obtaining home furnishing color recognition results and home furnishing style recognition results corresponding to the clothing to be identified.

[0070] In this embodiment, by determining the category represented by the classification detection result, the recognition model corresponding to the category is obtained, and the clothing to be identified is identified by the recognition model corresponding to the category, thereby realizing the recognition of the color and style of the clothing to be identified, improving the accuracy of color recognition and style recognition, and thereby improving the efficiency of classification.

[0071] In one embodiment, if Figure 6 As shown, in step S202, the image recognition is performed on all the images to be recognized by the top recognition model to obtain the top color recognition result and the top style recognition result, including:

[0072] S2021, performing convolution processing on all the to-be-recognized images through a convolution network in the upper garment recognition model to obtain convolution features corresponding to each of the to-be-recognized images.

[0073] S2022, performing deep convolution processing on all the convolution features through an MBConv network in the upper garment recognition model to obtain output features corresponding to each of the convolution features.

[0074] S2023, performing recognition processing on all the output features through a recognition network in the upper garment recognition model to obtain upper garment color recognition results and upper garment style recognition results.

[0075] Understandably, the convolution feature refers to a specific property of an image extracted through convolution operation of the convolution network. The output feature refers to an image feature extracted through the MBConv network. The convolution network refers to a 3*3 convolution network. The MBConv network is composed of 1x1 normal convolution (dimension increasing effect, including BN and Swish activation), k*k Depthwise Conv convolution (including BN and Swish activation), k*k has two cases of 3x3 and 5x5, SE module, 1x1 normal convolution (dimension reduction effect, including BN and linear activation), and Droupout layer. The recognition network is composed of 1x1 normal convolution, average pooling layer and full connection layer.

[0076] Specifically, when the classification detection result indicates that the to-be-recognized clothes is a top, all to-be-recognized images of the to-be-recognized clothes are cropped, that is, the to-be-recognized clothes are cropped according to the target box in the classification detection result, so as to obtain target images corresponding to each to-be-recognized image. Then, all target images are input into the top recognition model, and all target images are respectively subjected to convolution processing by a convolution network in the top recognition model, that is, first subjected to convolution processing by a convolution layer with a 3*3 convolution kernel and a step of 2, and then subjected to batch normalization and activation processing on the convolution layer result by a BN layer and a Swish activation function, so as to obtain convolution features corresponding to each target image. Further, all convolution features are subjected to deep convolution processing by a plurality of stacked MBConv networks in the top recognition model, that is, subjected to dimensionality increasing processing by a 1*1 convolution layer (containing a BN layer and a Swish activation function) to obtain dimensionality increased features, then subjected to grouping convolution in the feature dimension by a plurality of deep separable convolution layers (composed of deep convolution and pointwise convolution) with a K*K convolution kernel containing a BN layer and a Swish activation function, each channel is independently subjected to deep convolution, and after all channels are aggregated by a 1x1 convolution (pointwise convolution) before output, residual processing is performed by an SE module, that is, after the feature map after the deep separable layer is transmitted into the SE module, it is divided into two branches, one branch retains the current feature map, and the other branch first performs global average pooling on the feature map, then performs dimensionality reduction operation by a first fully connected layer FC (using a Swish activation function), and then performs dimensionality increasing operation by a second fully connected layer FC (using a sigmoid activation function), and finally the feature matrices of the two branches are multiplied to obtain the residual processing result. Then, after dimensionality reduction processing by a 1x1 convolution kernel containing a BN layer, a part of neurons in the neural network is randomly discarded by a Droupout layer to prevent overfitting, so as to improve the generalization ability of the model, and the output result and the input convolution feature are connected by a shortcut, and after the connected feature passes through the plurality of MBConv networks, output features corresponding to each convolution feature are obtained. All output features are subjected to recognition processing by a recognition network in the top recognition model, that is, subjected to color recognition and style recognition by a recognition layer composed of a 1x1 convolution layer containing a BN and a Swish activation function, an average pooling layer and a fully connected layer, so as to obtain top color recognition results and top style recognition results.

[0077] In this embodiment, all the to-be-recognized images are subjected to convolution processing through the convolution network, so as to realize extraction of convolution features. All the convolution features are subjected to deep convolution processing through the MBConv network, so as to realize acquisition of output features, and then convolution is respectively performed in the width and depth directions, so as to reduce the parameter quantity and improve the recognition speed. All the output features are subjected to recognition processing through the recognition network, so as to realize acquisition of the shirt color recognition result and the shirt style recognition result, and then accurate recognition of the color and the style is realized.

[0078] In an embodiment, as shown in FIG. 1, before the image classification recognition of all the to-be-recognized images by the detection model in step S10, the method further includes: Figure 7

[0079] S103, performing slicing processing on all the to-be-recognized images through a sliced inference algorithm, to obtain at least one sliced image corresponding to each of the to-be-recognized images.

[0080] Understandably, the sliced inference algorithm is a method for object detection on large-size images. The image is divided into multiple smaller slices, and then object detection is independently performed on each slice. Finally, the results are integrated to improve the detection performance on small targets. The sliced image refers to a part of the to-be-recognized image.

[0081] Specifically, after obtaining the images of the to-be-recognized clothes, slicing processing is performed on each to-be-recognized image through the sliced inference algorithm, that is, a preset sliding window and an overlap amount (i.e., an overlap percentage between two sliced images) are obtained. The moving step of the sliding window is calculated through the size of the to-be-recognized image, the size of the sliding window, and the overlap amount, and the image is segmented in the to-be-recognized image through the sliding window with the moving step, so as to obtain at least one sliced image corresponding to each to-be-recognized image. For example, the width of the to-be-recognized image is W, the height is H, the width of the sliding window is M, the height is N, and the overlap rate is P. Then, the horizontal step S=M*(1-P), and the vertical step S=N*(1-P). The number of horizontal slices K=((W-M) / S)+1, and the number of vertical slices K=((H-N) / S)+1. The starting position of each slice is determined on the to-be-recognized image using the step S and the slice size M and N, and multiple overlapping sliced images are cut out from the to-be-recognized image using the calculated slice position and size.

[0082] In this embodiment, slicing processing is performed on all the to-be-recognized images through the sliced inference algorithm, so as to realize acquisition of the sliced images, thereby better capturing the features of small targets, and then the accuracy and robustness of the detection capability of small objects in the image are improved. ​

[0083] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0084] In an embodiment, a clothes recognition device is provided, which corresponds to the clothes recognition method in the above embodiments. As shown in the figure, the clothes recognition device includes an image classification and recognition module 10, a clothes image recognition module 20, and a clothes recognition result module 30. The functions of each module are described in detail as follows: Figure 8

[0085] The image classification and recognition module 10 is configured to obtain at least one to-be-recognized image corresponding to a to-be-recognized clothes, and perform image classification and recognition on all the to-be-recognized images through a detection model to obtain a classification detection result, a brand detection result, and a defect detection result corresponding to the to-be-recognized clothes;

[0086] The clothes image recognition module 20 is configured to perform clothes image recognition on all the to-be-recognized images through a type recognition model corresponding to the classification detection result to obtain a style recognition result and a color recognition result;

[0087] The clothes recognition result module 30 is configured to determine a clothes recognition result corresponding to the to-be-recognized clothes according to the brand detection result, the style recognition result, the color recognition result, and the defect detection result.

[0088] Optionally, the detection model includes a classification detection model, a brand detection model, and a defect detection model.

[0089] The image classification and recognition module 10 includes:

[0090] A classification detection unit is configured to perform classification detection on all the to-be-recognized images through the classification detection model to obtain a classification detection result corresponding to the to-be-recognized clothes;

[0091] A brand detection unit is configured to perform brand detection on all the to-be-recognized images through the brand detection model to obtain a brand detection result corresponding to the to-be-recognized clothes;

[0092] A defect detection unit is configured to perform defect detection on all the to-be-recognized images through the defect detection model to obtain a defect detection result corresponding to the to-be-recognized clothes.

[0093] Optionally, the classification detection unit includes:

[0094] ​a first feature extraction subunit configured to perform feature extraction of different scales on all the to-be-identified images through a first backbone network in the classification and detection model to obtain first extracted features;

[0095] a first feature fusion subunit configured to perform feature fusion of different dimensions on all the first extracted features through a Neck network in the classification and detection model to obtain first fused features;

[0096] a classification and detection result subunit configured to perform prediction processing on all the first fused features through a first recognition network in the classification and detection model to obtain a classification and detection result.

[0097] Optionally, the brand detection unit comprises:

[0098] a second feature extraction subunit configured to perform feature extraction of different scales on all the to-be-identified images through a second backbone network in the brand detection model to obtain second extracted features;

[0099] a second feature fusion subunit configured to perform feature fusion of different dimensions on all the second extracted features through a BiFPN network in the brand detection model to obtain second fused features;

[0100] a brand detection result subunit configured to perform prediction processing on all the second fused features through a second recognition network in the brand detection model to obtain a brand detection result.

[0101] Optionally, the type recognition model comprises a top recognition model, a trousers and skirts recognition model, a shoe recognition model and a home recognition model.

[0102] The clothing image recognition module 20 comprises:

[0103] a top recognition unit configured to perform image recognition on all the to-be-identified images through the top recognition model when the classification and detection result indicates that the to-be-identified clothing is a top to obtain a top color recognition result and a top style recognition result;

[0104] a trousers and skirts recognition unit configured to perform image recognition on all the to-be-identified images through the trousers and skirts recognition model when the classification and detection result indicates that the to-be-identified clothing is trousers and skirts to obtain a trousers and skirts color recognition result and a trousers and skirts style recognition result;

[0105] a shoe recognition unit configured to perform image recognition on all the to-be-identified images through the shoe recognition model when the classification and detection result indicates that the to-be-identified clothing is a shoe to obtain a shoe color recognition result and a shoe style recognition result;

[0106] The home recognition unit is configured to, when the classification detection result indicates that the clothes to be recognized are home clothes, perform image recognition on all the to-be-recognized images by using the home recognition model to obtain home color recognition results and home style recognition results.

[0107] Optionally, the upper garment recognition unit comprises:

[0108] The convolution processing subunit is configured to perform convolution processing on all the to-be-recognized images by using a convolution network in the upper garment recognition model to obtain convolution features corresponding to each of the to-be-recognized images.

[0109] The deep convolution subunit is configured to perform deep convolution processing on all the convolution features by using an MBConv network in the upper garment recognition model to obtain output features corresponding to each of the convolution features.

[0110] The recognition processing subunit is configured to perform recognition processing on all the output features by using a recognition network in the upper garment recognition model to obtain upper garment color recognition results and upper garment style recognition results.

[0111] Optionally, the image classification recognition module 10 further comprises:

[0112] The segmentation processing unit is configured to perform segmentation processing on all the to-be-recognized images by using a slice-assisted super-inference algorithm to obtain at least one segmented image corresponding to each of the to-be-recognized images.

[0113] A computer device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor is configured to execute the clothes recognition method.

[0114] The specific definitions of the computer device, the processor, and the units and modules thereof can be found in the above definitions of the clothes recognition method, which will not be repeated here. Each module in the processor can be implemented by software, hardware, or a combination thereof, in whole or in part. Understandably, the processor comprises a processor, a memory, a network interface, and a database connected by a device bus. Each module of the processor can be embedded in the processor in hardware form or independent of the processor, or stored in the memory in software form to be called and executed by the processor to perform the operations corresponding to each module. The processor is configured to provide computing and control capabilities. The memory comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores operating devices, computer programs, and databases. The internal memory provides an environment for the operation of the operating devices and computer programs in the non-volatile storage medium. The database is configured to store data used by the clothes recognition method in the above embodiments. The network interface is configured to communicate with external terminals through network connection. The computer program is executed by the processor to implement a clothes recognition method.

[0115] In one embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the above-mentioned clothes recognition method.

[0116] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the processes of the above-mentioned embodiments can be included. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM) and the like.

[0117] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is exemplified. In actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the above-mentioned functions.

[0118] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, and not to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalent ones. These modifications or replacements do not change the essence of the corresponding technical solutions, and should be included in the protection scope of the present application.

Claims

1. A clothing identification method, characterized in that: include: Acquire at least one image to be identified corresponding to the clothing to be identified, and perform image classification and recognition on all the images to be identified using a detection model to obtain a classification detection result, a brand detection result, and a defect detection result corresponding to the clothing to be identified; Perform clothing image recognition on all the images to be recognized using a type recognition model corresponding to the classification detection result to obtain a style recognition result and a color recognition result; Determining a clothing recognition result corresponding to the clothing to be recognized based on the brand detection result, the style recognition result, the color recognition result, and the defect detection result; The detection models include a classification detection model, a brand detection model and a defect detection model; The image classification and recognition is performed on all the images to be recognized by the detection model to obtain classification detection results, brand detection results, and defect detection results corresponding to the clothing to be recognized, including: Performing classification detection on all the images to be identified using the classification detection model to obtain classification detection results corresponding to the clothing to be identified; Performing brand detection on all the images to be identified using the brand detection model to obtain brand detection results corresponding to the clothing to be identified; Performing defect detection on all the images to be identified using the defect detection model to obtain defect detection results corresponding to the clothing to be identified; The step of performing brand detection on all the images to be identified by using the brand detection model to obtain brand detection results corresponding to the clothing to be identified includes: Performing feature extraction of different scales on all the images to be identified using the second backbone network in the brand detection model to obtain second extracted features; Performing feature fusion of different dimensions on all the second extracted features through the BiFPN network in the brand detection model to obtain a second fused feature; All the second fusion features are predicted and processed by the second recognition network in the brand detection model to obtain a brand detection result.

2. The clothing identification method according to claim 1, wherein: The step of performing classification detection on all the images to be identified by using the classification detection model to obtain classification detection results corresponding to the clothing to be identified includes: Performing feature extraction of different scales on all the images to be identified by using the first backbone network in the classification detection model to obtain first extracted features; Performing feature fusion of different dimensions on all the first extracted features through the Neck network in the classification detection model to obtain a first fused feature; All the first fusion features are predicted and processed by the first recognition network in the classification detection model to obtain a classification detection result.

3. The clothing identification method according to claim 1, wherein: The type recognition model includes a top recognition model, a trouser skirt recognition model, a shoe recognition model and a home recognition model; The method of performing clothing image recognition on all the images to be recognized by using a type recognition model corresponding to the classification detection result to obtain a style recognition result and a color recognition result includes: When the classification detection result indicates that the clothing to be identified is a top, performing image recognition on all the images to be identified using the top recognition model to obtain a top color recognition result and a top style recognition result; When the classification detection result indicates that the clothing to be identified is a culottes, performing image recognition on all the images to be identified using the culottes recognition model to obtain a culottes color recognition result and a culottes style recognition result; When the classification detection result indicates that the clothing to be identified is a shoe, performing image recognition on all the images to be identified using the shoe recognition model to obtain a shoe color recognition result and a shoe style recognition result; When the classification detection result indicates that the clothing to be identified is home furnishings, image recognition is performed on all the images to be identified using the home furnishings recognition model to obtain home furnishings color recognition results and home furnishings style recognition results.

4. The clothing identification method according to claim 3, wherein: The method of performing image recognition on all the images to be recognized by the top recognition model to obtain a top color recognition result and a top style recognition result includes: Performing convolution processing on all the images to be recognized through the convolution network in the top recognition model to obtain convolution features corresponding to each of the images to be recognized; Performing deep convolution processing on all the convolution features through the MBConv network in the top recognition model to obtain output features corresponding to each convolution feature; All the output features are processed by the recognition network in the top recognition model to obtain a top color recognition result and a top style recognition result.

5. The clothing identification method according to claim 1, wherein: Before performing image classification and recognition on all the images to be recognized by the detection model, the method further includes: All the images to be identified are segmented using a slice-assisted hyper-inference algorithm to obtain at least one segmented image corresponding to each of the images to be identified.

6. A clothing identification device, characterized in that: include: An image classification and recognition module is configured to obtain at least one image to be recognized corresponding to the clothing to be recognized, and perform image classification and recognition on all of the images to be recognized using a detection model to obtain classification detection results, brand detection results, and defect detection results corresponding to the clothing to be recognized; a clothing image recognition module, configured to perform clothing image recognition on all the images to be recognized using a type recognition model corresponding to the classification detection result, and obtain style recognition results and color recognition results; a clothing recognition result module, configured to determine a clothing recognition result corresponding to the clothing to be recognized based on the brand detection result, the style recognition result, the color recognition result, and the defect detection result; The image classification and recognition module includes: a classification detection unit, configured to perform classification detection on all the images to be identified using a classification detection model, and obtain classification detection results corresponding to the clothing to be identified; a brand detection unit, configured to perform brand detection on all the images to be identified using a brand detection model, and obtain brand detection results corresponding to the clothing to be identified; a defect detection unit, configured to perform defect detection on all the images to be identified using a defect detection model, and obtain defect detection results corresponding to the clothing to be identified; The brand detection unit includes: A second feature extraction subunit is configured to extract features of different scales from all the images to be identified using the second backbone network in the brand detection model to obtain second extracted features; A second fusion feature subunit is configured to perform feature fusion of different dimensions on all the second extracted features through the BiFPN network in the brand detection model to obtain a second fused feature; The brand detection result subunit is used to perform prediction processing on all the second fusion features through the second recognition network in the brand detection model to obtain a brand detection result.

7. A computer device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to execute the clothing identification method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the clothing identification method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Machine control method and system based on object recognition

    CN114466954A