Clothing Recognition Method, Device, Equipment and Medium

Through the clothing recognition model and type recognition model, combined with the local sensitive hashing algorithm and anchor box algorithm, the problem of inaccurate pedestrian attribute recognition is solved and the accuracy of pedestrian recognition is improved.

CN113486855BActive Publication Date: 2025-07-04ZHEJIANG DAHUA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110868406.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-30
Publication Date
2025-07-04
Estimated Expiration
2041-07-30

AI Technical Summary

Technical Problem

In the prior art, pedestrian attribute identification is inaccurate, resulting in inaccurate pedestrian identification.

Method used

The style and position information of the clothing in the image is determined through the clothing recognition model, and the type recognition model is used to determine the clothing category and logo position, combining the local sensitive hashing algorithm and anchor box algorithm to improve the accuracy of attribute recognition.

Benefits of technology

The identification accuracy of the style and target type identification of different clothing categories in the image is improved, thereby improving the accuracy of pedestrian identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113486855B_ABST
    Figure CN113486855B_ABST
Patent Text Reader

Abstract

The present invention provides a clothing recognition method, device, equipment and medium, which are used to solve the problem that the attribute recognition in the prior art is inaccurate, resulting in inaccurate pedestrian recognition. Since in the embodiments of the present invention, through the clothing style recognition model, the style of the clothing in the image can be determined, and through the type recognition model, the position information of the logo, the position information of the clothing and the clothing category corresponding to the clothing in the image can be determined, and further the target type identifier corresponding to different clothing categories can be determined. In addition, since the area occupied by the pedestrian's clothing in the image is large and relatively clear, the style and target type identifier corresponding to different clothing categories in the image can be accurately recognized, improving the accuracy of attribute recognition, and thus improving the accuracy of pedestrian recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a clothing recognition method, device, equipment and medium. Background Art

[0002] In recent years, with the rapid development of video surveillance technology, a vast amount of video image data is generated every moment. Among the vast amount of video image data, quickly retrieving specific pedestrians is one of the most important tasks. Among them, when performing pedestrian retrieval, it is based on a pedestrian database constructed by video structured description, and combined with computer vision algorithms such as image optimization, attribute recognition, and target tracking for retrieval. Therefore, it is crucial to perform attribute recognition on pedestrians.

[0003] In the prior art, the method of attribute recognition is to recognize features such as the appearance features or face features of pedestrians. However, there are many appearance features and face features, and the areas they occupy in the images collected by image acquisition devices are small and relatively blurred, and it is very easy to have inaccurate recognition problems when performing recognition, resulting in inaccurate pedestrian recognition. Summary of the Invention

[0004] The present invention provides a clothing recognition method, device, equipment and medium, which are used to solve the problem that inaccurate attribute recognition in the prior art leads to inaccurate pedestrian recognition.

[0005] In a first aspect, an embodiment of the present invention provides a clothing recognition method, and the method includes:

[0006] Receiving an image containing clothing; obtaining the style of the clothing in the image through a pre-trained clothing style recognition model;

[0007] Obtaining the position information of the logo, the position information of the clothing, and the clothing category corresponding to the clothing in the image through a pre-trained trademark type recognition model;

[0008] For each clothing category, determining the target type logo corresponding to the clothing category according to the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to different types of logos pre-saved.

[0009] Further, the clothing style recognition model is trained in the following manner:

[0010] Obtaining any sample image in the first sample set, and the sample style logo of each piece of clothing included in the sample image;

[0011] Inputting the sample image into the original recognition model to obtain the recognized style logo of each piece of clothing included in the sample image;

[0012] Train the original recognition model according to the sample style identifier and the recognized style identifier.

[0013] Further, the method further includes:

[0014] Perform offline distillation learning training on a pre-configured clothing style recognition model according to the trained original recognition model to obtain a trained clothing style recognition model.

[0015] Further, the determining the target type identifier corresponding to the clothing category according to the position information of the identifier of the clothing corresponding to the clothing category and the images corresponding to different types of identifiers pre-saved includes:

[0016] Use a preset algorithm to determine the similarity between the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to each type of identifier pre-saved; wherein, the preset algorithm includes: local sensitive hashing algorithm;

[0017] Determine the type identifier corresponding to the image with the highest similarity and exceeding the preset threshold as the target type identifier corresponding to the clothing category.

[0018] Further, the method further includes:

[0019] For the position information of each logo obtained, determine the overlap degree between the position information of the logo and the position information of the clothing obtained; determine the position information of the clothing with the largest overlap degree as the position information of the target clothing; according to the corresponding relationship between the position information of the clothing and the clothing category, determine the target clothing category corresponding to the position information of the target clothing as the clothing category corresponding to the logo.

[0020] Further, the obtaining the position information of the logo of the clothing in the image through a pre-trained type recognition model includes:

[0021] Extract features from the image through the first network layer in the pre-trained type recognition model to obtain the feature maps output by each sub-network layer in the first network layer;

[0022] Obtain each position information of the logo in the feature map through the detection layer in the type recognition model; wherein, each position information of the logo is determined according to the regions containing the logo framed in the feature map by a preset number of anchor boxes with different ratios, and the preset number is not less than four;

[0023] Obtain the position information of the logo through the second network layer in the type recognition model based on each position information of the logo.

[0024] Further, the ratio of the anchor is obtained in the following manner:

[0025] According to the sizes of the pre-annotated logos and the sizes of the clothing items in the second sample set, obtain the aspect ratio of each logo and the aspect ratio of each clothing item in the second sample set;

[0026] Cluster the aspect ratios of each logo and the aspect ratios of each clothing item in the second sample set through a clustering algorithm and the preset quantity;

[0027] According to the aspect ratios of the logos and the aspect ratios of the clothing items in each category of the determined clustering result, determine the anchor ratio corresponding to each category.

[0028] In a second aspect, an embodiment of the present invention provides a clothing recognition device, and the device includes:

[0029] A receiving and obtaining module, configured to receive an image including clothing; obtain the style of the clothing in the image through a pre-trained clothing style recognition model; obtain the position information of the logo, the position information of the clothing, and the clothing category corresponding to the clothing in the image through a pre-trained trademark type recognition model;

[0030] A processing module, configured to, for each clothing category, determine the target type identifier corresponding to the clothing category according to the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to different types of identifiers pre-stored.

[0031] Further, the processing module is further configured to obtain any sample image in the first sample set and the sample style identifier of each piece of clothing included in the sample image; input the sample image into the original recognition model to obtain the recognized style identifier of each piece of clothing included in the sample image; and train the original recognition model according to the sample style identifier and the recognized style identifier.

[0032] Further, the processing module is further configured to perform offline distillation learning training on a pre-configured clothing style recognition model according to the trained original recognition model to obtain a trained clothing style recognition model.

[0033] Further, the processing module is specifically configured to use a preset algorithm to determine the similarity between the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to each type identifier pre-stored; wherein, the preset algorithm includes: a locality-sensitive hashing algorithm; and determine the type identifier corresponding to the image with the highest similarity and exceeding the preset threshold as the target type identifier corresponding to the clothing category.

[0034] Further, the processing module is further configured to determine, for the position information of each logo obtained, the degree of overlap between the position information of the logo and the position information of the clothing obtained; determine the position information of the clothing with the largest degree of overlap as the position information of the target clothing; and determine, according to the corresponding relationship between the position information of the clothing and the clothing category, that the target clothing category corresponding to the position information of the target clothing is the clothing category corresponding to the logo.

[0035] Further, the receiving and obtaining module is specifically configured to extract features of the image through the first network layer in a pre-trained type recognition model, and obtain the feature maps output by each sub-network layer in the first network layer; obtain each position information of the logo in the feature map through the detection layer in the type recognition model; wherein, each position information of the logo is determined according to the regions containing the logo boxed in the feature map by a preset number of anchor boxes with different ratios, and the preset number is not less than four; and obtain the position information of the logo based on each position information of the logo through the second network layer in the type recognition model.

[0036] Further, the receiving and obtaining module is further configured to obtain the aspect ratio of each logo and the aspect ratio of each piece of clothing in the second sample set according to the size of the logo and the size of the clothing pre-annotated in the second sample set; perform clustering on the aspect ratio of each logo and the aspect ratio of each piece of clothing in the second sample set through a clustering algorithm and the preset number; and determine the anchor ratio corresponding to each category according to the aspect ratio of the logo and the aspect ratio of the clothing in each category in the determined clustering result.

[0037] In a third aspect, an embodiment of the present invention provides an electronic device, which at least includes a processor and a memory. The processor is configured to execute the steps of the clothing recognition method according to any one of the above claims when executing a computer program stored in the memory.

[0038] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. The computer program is configured to execute the steps of the clothing recognition method according to any one of the above claims when executed by a processor.

[0039] Since in the embodiment of the present invention, through the clothing style recognition model, the style of the clothing in the image can be determined, and through the type recognition model, the position information of the logo, the position information of the clothing, and the clothing category corresponding to the clothing in the image can be determined, and further the target type identifier corresponding to different clothing categories can be determined. In addition, since the area occupied by the clothing of the pedestrian in the image is large and relatively clear, the style and target type identifier corresponding to different clothing categories in the image can be accurately recognized, improving the accuracy of attribute recognition, and thus improving the accuracy of pedestrian recognition. Brief Description of the Drawings

[0040] To more clearly illustrate the technical solutions of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0041] Figure 1 Schematic diagram of a clothing recognition process provided by an embodiment of the present invention;

[0042] Figure 2 Schematic diagram of the position information of the marks output by the type recognition model provided by an embodiment of the present invention;

[0043] Figure 3 Detailed schematic diagram of a clothing recognition process provided by an embodiment of the present invention;

[0044] Figure 4 Schematic diagram of the structure of a clothing recognition device provided by an embodiment of the present invention;

[0045] Figure 5 Schematic diagram of the result of an electronic device provided by an embodiment of the present invention. Detailed Embodiments

[0046] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present invention.

[0047] Embodiment 1:

[0048] Figure 1 Schematic diagram of a clothing recognition process provided by an embodiment of the present invention, and this process includes the following steps:

[0049] S101: Receive an image containing clothing; through a pre-trained clothing style recognition model, obtain the style of the clothing in the image.

[0050] The clothing recognition method provided by the embodiment of the present invention is applied to an electronic device, and this electronic device can be an intelligent device such as an image acquisition device, a PC, or a server.

[0051] The electronic device first receives an image containing a pedestrian and the clothing worn by the pedestrian. Among them, the image is an image that has been processed to contain only one pedestrian. After receiving the image containing the clothing, the electronic device can first input the received image into a pre-trained clothing style recognition model to obtain the output of the clothing style recognition model. Among them, the output of the clothing style recognition model contains the style of the clothing contained in the image.

[0052] The styles of clothing include styles of upper garments in clothing such as T-shirts, suits, sweaters, jackets, coats, dresses, shirts, etc.; styles of lower garments such as shorts, trousers, skirts, etc.; styles of shoes such as sports shoes, leather shoes, boots, sandals, no shoes, etc.; styles of accessories such as backpacks, handbags, shoulder bags, suitcases, etc.

[0053] S102: Obtain the position information of the logo, the position information of the clothing, and the clothing category corresponding to the clothing in the image through a pre-trained trademark type recognition model.

[0054] The electronic device inputs the image into the trademark type recognition model to obtain the output of the type recognition model. Among them, the output contains the position information of the logo in the image, the position information of the clothing, and the clothing category corresponding to the clothing. Among them, the position information of the clothing in the output corresponds to the clothing category of the clothing corresponding to this position information. That is to say, the position information of the clothing corresponds to whether the clothing at this position information is an upper garment, a lower garment, shoes, or an accessory. And the position information of the logo and the position information of the clothing can be the position information outlined by a rectangular frame. Among them, the logo can be a trademark (Logo).

[0055] S103: For each clothing category, determine the target type identifier corresponding to the clothing category according to the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to different types of identifiers pre-saved.

[0056] In order to determine the type identifier corresponding to each clothing category, images corresponding to different types of identifiers are pre-saved in the electronic device. For each clothing category, the electronic device determines the target type identifier corresponding to the clothing category according to the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to different types of identifiers. Specifically, the process for the electronic device to determine the target type identifier corresponding to the clothing category is as follows: intercept the position information of the logo corresponding to the clothing category in the image to obtain a sub-image containing the logo corresponding to the clothing category after interception. According to the sub-image and the images corresponding to different types of identifiers pre-saved, determine the image corresponding to the type identifier that matches the sub-image, obtain the image corresponding to the type identifier determined to match the sub-image, and determine the type identifier corresponding to the obtained image corresponding to the type identifier as the target type identifier corresponding to the clothing category. Among them, the type identifier includes a trademark identifier.

[0057] In the embodiments of the present invention, through the clothing style recognition model, the style of the clothing in the image can be determined, and through the type recognition model, the position information of the logo, the position information of the clothing, and the clothing category corresponding to the clothing in the image can be determined. Furthermore, the target type identifier corresponding to different clothing categories can be determined. Additionally, since the area occupied by the pedestrian's clothing in the image is large and relatively clear, the styles and target type identifiers corresponding to different clothing categories in the image can be accurately recognized, improving the accuracy of attribute recognition and thus improving the accuracy of pedestrian recognition.

[0058] Embodiment 2:

[0059] In order to accurately obtain the clothing style recognition model, based on the above embodiments, in the embodiments of the present invention, the clothing style recognition model is trained in the following manner:

[0060] Obtain any sample image in the first sample set and the sample style identifier of each piece of clothing included in the sample image;

[0061] Input the sample image into the original recognition model to obtain the recognized style identifier of each piece of clothing included in the sample image;

[0062] Train the original recognition model according to the sample style identifier and the recognized style identifier.

[0063] In order to enable the clothing style recognition model with a relatively small network depth to achieve the recognition effect of a model with a relatively large network depth and accurately recognize the style of the clothing in the image, in the embodiments of the present invention, the original recognition model with a relatively large network depth is first trained. Specifically, in order to implement the training of the original recognition model, in the embodiments of the present invention, a first sample set for training the original recognition model is stored. This first sample set contains a large number of sample images, where the sample images are images of pedestrians wearing clothing of different styles. For the convenience of training the original recognition model, for each sample image in this first sample set, the sample style identifier of each piece of clothing included in the sample image in the sample image is also stored.

[0064] Among them, the sample style identifier can include the identifier of the clothing category and the identifier of the style, or can include only the identifier of the style. And the styles of clothing include styles of upper garments such as T-shirts, suits, sweaters, jackets, coats, dresses, shirts, etc.; styles of lower garments such as shorts, trousers, skirts, etc.; styles of shoes such as sports shoes, leather shoes, boots, sandals, no shoes, etc.; styles of accessories such as backpacks, handbags, shoulder bags, suitcases, etc. Taking the case where the sample style identifier includes the identifier of the clothing category and the identifier of the style as an example, for example, the identifier of upper garments in the clothing category can be 00, the identifier of lower garments can be 01, the identifier of shoes can be 02, and the identifier of accessories can be 03; and among the styles of upper garments, the identifier of T-shirts can be a, the identifier of suits can be b, among the styles of lower garments, the identifier of shorts can be a, and the identifier of trousers can be b. Then the sample style identifier of T-shirts is 00a, and the sample style identifier of shorts is 01a. Taking the case where the sample style identifier includes only the identifier of the style as an example, for example, the identifier of T-shirts is a, the identifier of suits can be b, the identifier of shorts can be c, and the identifier of trousers can be d. Then the sample style identifier of T-shirts is a, and the sample style identifier of shorts is c.

[0065] In an embodiment of the present invention, after obtaining any sample image in the first sample set and the sample style identifier of the clothing included in the sample image, the sample image is input into the original recognition model, and the output of the original recognition model is obtained, where the output includes the recognized style identifier of the clothing included in the sample image.

[0066] After obtaining the output of the original recognition model, the electronic device determines that the output of the original recognition model is the recognized style identifier of the clothing included in the sample image. The electronic device adjusts the parameters of the original recognition model according to the sample style identifier in the sample image and the recognized style identifier output by the original recognition model, so as to train the original recognition model.

[0067] When the original recognition model is trained in the above manner and meets the preset conditions, a trained original recognition model is obtained. Among them, the preset conditions can be that the number of sample images whose recognized style identifiers obtained after training the sample images in the first sample set through the original recognition model are consistent with the sample style identifiers is greater than the set number; or the number of iterations for training the original recognition model reaches the set maximum number of iterations, etc. Specifically, the embodiments of the present invention do not limit this.

[0068] In an embodiment of the present invention, the network depth of the original recognition model is relatively large, and the original recognition model can be a Resnet-50 model.

[0069] In order to accurately obtain a clothing style recognition model, based on the above embodiments, in an embodiment of the present invention, the method further includes:

[0070] Perform offline distillation learning training on the pre-configured clothing style recognition model according to the originally trained recognition model, and obtain the trained clothing style recognition model.

[0071] In the embodiment of the present invention, in order to enable the pre-configured clothing style recognition model with less network depth to achieve the recognition effect of the original recognition model with more network depth, the clothing style recognition model is trained by the originally trained recognition model. Specifically, the process of training the clothing style recognition model by the originally trained recognition model is as follows: Use the originally trained recognition model as the teacher model to perform offline distillation learning training on the clothing style recognition model, and obtain the trained clothing style recognition model.

[0072] Among them, the original recognition model has a relatively large network depth, and the original recognition model can be a Resnet-50 model, and the pre-configured clothing style recognition model can be a Resnet-18 model. That is to say, in the embodiment of the present invention, the trained Resnet-50 model can be used as the teacher model for offline distillation learning training, so that the small model Resnet-18 can learn the recognition ability of the large model Resnet-50, thereby improving the performance of the Resnet-18 model without increasing the algorithm complexity and model parameters.

[0073] Embodiment 3:

[0074] In order to accurately determine the Logo category corresponding to the clothing style, on the basis of the above embodiments, in the embodiment of the present invention, the determining the target type logo corresponding to the clothing category according to the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to different types of logos pre-stored includes:

[0075] Adopt a preset algorithm to determine the similarity between the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to each type of logo pre-stored; wherein, the preset algorithm includes: local sensitive hashing algorithm;

[0076] Determine the type logo corresponding to the image with the highest similarity and exceeding the preset threshold as the target type logo corresponding to the clothing category.

[0077] In an embodiment of the present invention, in order to determine the type identifier corresponding to the clothing category in the image, different type identifiers corresponding to images are pre-stored in the electronic device, and a sub-image of the area containing the logo corresponding to the clothing category in the image is intercepted according to the position information of the logo of the clothing corresponding to the clothing category in the image. After intercepting the sub-image of the area containing the logo corresponding to the clothing category, a preset algorithm is used to determine the similarity between the sub-image and each image corresponding to the pre-stored type identifier, so as to determine the type identifier of the logo contained in the sub-image, and the preset algorithm can be the locality-sensitive hashing algorithm. Specifically, how to use the locality-sensitive hashing algorithm to determine the similarity between the sub-image and the images corresponding to different type identifiers is the prior art and will not be elaborated here.

[0078] After determining the similarity between the sub-image and each image corresponding to the type identifier, the image corresponding to the type identifier with the highest determined similarity is obtained, and it is determined whether the similarity of the image corresponding to the type identifier with the highest similarity exceeds a preset threshold. If the similarity of the image corresponding to the type identifier with the highest similarity exceeds the preset threshold, the type identifier of the image corresponding to the type identifier with the highest similarity is determined as the target type identifier corresponding to the clothing category. If the similarity of the image corresponding to the type identifier with the highest similarity does not exceed the preset threshold, the type identifier corresponding to the clothing category is determined to be unknown.

[0079] In order to accurately determine the position information of the logos contained in the clothing of different clothing categories, based on the above embodiments, in an embodiment of the present invention, the method further includes:

[0080] For the position information of each obtained logo, the coincidence degree between the position information of the logo and the position information of the obtained clothing is determined; the position information of the clothing with the largest coincidence degree is determined as the position information of the target clothing; according to the corresponding relationship between the position information of the clothing and the clothing category, the target clothing category corresponding to the position information of the target clothing is determined as the clothing category corresponding to the logo.

[0081] In an embodiment of the present invention, an image is input into a type recognition model. The output of the type recognition model includes the position information of each logo in the image, the position information of each piece of clothing, and the clothing category corresponding to the clothing. In order to determine the clothing category corresponding to each logo, after the electronic device obtains the output of the type recognition model, it determines the clothing category corresponding to each logo in the image according to the output of the type recognition model. Specifically, the process of determining the clothing category corresponding to any logo in the image is as follows: according to the position information of the logo and the position information of each piece of clothing obtained, determine the degree of overlap between the position information of the logo and the position information of each piece of clothing, obtain the position information of the clothing with the highest degree of overlap with the position information of the logo. If the degree of overlap between the position information of the logo and the position information of a certain piece of clothing is the highest, it means that the logo on the clothing at the determined position information of the clothing is the logo at the position information of the logo. Then the clothing category corresponding to the position information of the clothing is the clothing category corresponding to the logo. Therefore, after determining the position information of the clothing with the highest degree of overlap with the position information of the logo, determine the clothing category corresponding to the position information of the clothing, which is the clothing category corresponding to the logo at the position information of the logo.

[0082] The process of determining the degree of overlap between the position information of the logo and the position information of a certain piece of clothing can be as follows: if the position information of the logo in the image is within the position information of the clothing, obtain the ratio of the area of the logo to the area of the clothing, and determine this ratio as the degree of overlap between the position information of the logo and the position information of the clothing; if the position information of the logo in the image is not within the position information of the clothing, determine the degree of overlap between the position information of the logo and the position information of the clothing as 0. Of course, other methods can also be used to determine the degree of overlap between the position information of the logo and the position information of the clothing. Specifically, how to determine it is not limited here.

[0083] And in an embodiment of the present invention, a training data set for training the type recognition model is pre-constructed to train the position information of the upper body, lower body, shoes, accessories of a pedestrian and the position information of each logo through the training data set.

[0084] Figure 2 It is a schematic diagram of the position information of the logo output by the type recognition model provided by the embodiment of the present invention.

[0085] As Figure 2 can be seen, the position information of the logo output by the type recognition model can be framed by a rectangular box.

[0086] Embodiment 4:

[0087] To improve the accuracy of the type recognition model, based on the above embodiments, in the embodiments of the present invention, obtaining the position information of the logo of the clothing in the image by the pre-trained type recognition model includes:

[0088] Performing feature extraction on the image through the first network layer in the pre-trained type recognition model to obtain the feature maps output by each sub-network layer in the first network layer;

[0089] Obtaining each position information of the logo in the feature map through the detection layer in the type recognition model; wherein, each position information of the logo is determined by the regions containing the logo framed in the feature map according to a preset number of anchor boxes with different ratios, and the preset number is not less than four;

[0090] Obtaining the position information of the logo based on each position information of the logo through the second network layer in the type recognition model.

[0091] To improve the accuracy of the type recognition model, in the embodiments of the present invention, the image is input into the type recognition model, and each sub-network layer in the first network layer of the type recognition model performs feature extraction on the image. After each sub-network layer in the first network layer performs feature extraction on the image, the feature maps output by each sub-network layer in the first network layer are obtained.

[0092] Among them, the feature maps output by each sub-network layer in the deep learning model include feature maps with resolutions of 28*28, 14*14, and 7*7. Since the position information of the clothing and the position information of the logo in the image are relatively large, and the feature map with a resolution of 7*7 is suitable for features with a smaller area, and the feature map output by the last sub-network layer in the deep learning model is the feature map with a resolution of 7*7, therefore, in the embodiments of the present invention, in order to reduce the parameters and time consumption of the type recognition model and improve the recognition efficiency of the type recognition model, the sub-network layer that outputs the feature map with a resolution of 7*7 in the deep learning model may not be included in the first network layer of the configured type recognition model, and if the sub-network layer that outputs the feature map with a resolution of 7*7 in the deep learning model is not included in the first network layer, it will not affect the effect of the type recognition model. Additionally, the sub-network layers included in the first network layer can be set according to the resolution size of the feature maps output by the first network layer. For example, the feature maps output by each sub-network layer included in the first network layer can be feature maps with a resolution of 28*28 and feature maps with a resolution of 14*14.

[0093] After obtaining the feature maps output by each sub-network layer, the detection layer in the type recognition model processes the feature maps output by each sub-network layer to obtain the output of the detection layer in the type recognition model, where the output of the detection layer in the type recognition model includes the position information of each position of each logo in the feature map. And in the embodiment of the present invention, in order to more accurately determine the position information of the logo and the position information of the clothing, when processing in the detection layer of the type recognition model, for any logo in the feature map, a preset number of position information corresponding to the logo is output, where the preset number of position information is framed in the feature map according to a preset number of anchor boxes with different ratios, and the areas framed by the preset number of anchor boxes with different ratios in the feature map contain the logo. In the embodiment of the present invention, in order to more accurately determine the position information of the logo, the preset number is not less than four.

[0094] After obtaining the position information of each position of each logo in the feature map, the second network layer in the type recognition model processes the position information of each position of each logo in the feature map to obtain the output of the second network layer in the type recognition model, where the output includes the position information of each logo. That is to say, for each logo, the feature map output by the detection layer in the type recognition model contains the position information of each position of the logo, and the feature map output by the second network layer in the type recognition model contains the position information of the logo.

[0095] In the embodiment of the present invention, the type recognition model can be an improved Yolov5s object detection model. Compared with the original Yolov5s object detection model, this model removes the upsampling layer in the network, thereby improving the forward propulsion speed of the type recognition model; removes the sub-network layer with an output resolution of 7*7 in the network layer, reducing the model parameters and time consumption, and the network performance drops less; the number of anchors used in each detection layer increases, making it more accurate to determine the position information of the logo and the position information of the clothing in the image.

[0096] In order to improve the accuracy of determining the position information of the logo and the position information of the clothing, on the basis of the above embodiments, in the embodiment of the present invention, the ratio of the anchor is obtained in the following manner:

[0097] According to the sizes of the logos and the sizes of the clothing pre-annotated in the second sample set, obtain the aspect ratio of the length and width of each logo and the aspect ratio of the length and width of each clothing in the second sample set;

[0098] Through the clustering algorithm and the preset number, cluster the aspect ratio of the length and width of each logo and the aspect ratio of the length and width of each clothing in the second sample set;

[0099] According to the aspect ratios of the logos and the aspect ratios of the clothing in each category of the determined clustering results, determine the anchor ratio corresponding to each category.

[0100] In the embodiments of the present invention, by increasing the number of anchors in the detection layer of the type recognition model, the accuracy of determining the position information of the logo and the position information of the clothing is improved. Among them, the ratio of the anchor can be determined by a clustering algorithm and the number of categories to be classified. In order to accurately determine the ratio of the anchor, a second sample set for determining the ratio of the anchor is pre-stored in the electronic device, and the second sample set contains the sizes of different pre-labeled logos and the sizes of different clothing. According to the sizes of different logos and different clothing in the second sample set, determine the aspect ratios of different logos and different clothing in the second sample set.

[0101] And perform clustering on the determined aspect ratios of different logos and different clothing through a clustering algorithm and a preset number. According to the aspect ratios of the logos and the aspect ratios of the clothing in each category of the determined clustering results, determine the anchor ratio corresponding to each category. That is, perform clustering on the aspect ratios of different logos and different clothing through a clustering algorithm and the number of categories to be classified, cluster the clustering results of a preset number of aspect ratios, and determine that the preset number of aspect ratios obtained by clustering are respectively the ratios of the anchors. Among them, in the embodiments of the present invention, the clustering algorithm can be the Kmeans clustering algorithm.

[0102] Embodiment 5:

[0103] Figure 3 It is a detailed schematic diagram of a clothing recognition process provided by an embodiment of the present invention.

[0104] The electronic device receives an image including a pedestrian and the pedestrian's clothing, and inputs the image into a clothing style recognition model to obtain the style of the clothing in the image. Among them, the determined styles of the clothing in the image are that the style of the upper garment is a jacket, the style of the lower garment is trousers, the style of the shoes is sports shoes, and the style of the accessory is a backpack.

[0105] And input the image including the clothing into the type recognition model to obtain the position information of the logo, the position information of the clothing, and the clothing category corresponding to the clothing in the image. And by determining the coincidence degree between the position information of the logo and the position information of each clothing, determine the clothing category corresponding to the position information of the logo. By comparing with the images corresponding to different type identifiers pre-stored, determine the target type identifier corresponding to each clothing category. Among them, the determined target type identifier corresponding to each clothing category can be that the accessory is Nike, the lower garment is unknown, and no logo is detected in other clothing categories.

[0106] Integrate the determined styles of the clothing and the target type identifiers of different clothing categories. Among them, the determined result based on the received image containing clothing is that the style of the upper garment is a jacket and the type identifier of the upper garment is none; the style of the lower garment is trousers and the type identifier of the lower garment is unknown; the style of the shoes is sports shoes and the type identifier of the shoes is none; the style of the accessory is a backpack and the type identifier of the accessory is Nike.

[0107] Embodiment 6:

[0108] Figure 4 A schematic structural diagram of a clothing recognition device provided by an embodiment of the present invention. The device includes:

[0109] A receiving and obtaining module 401, configured to receive an image containing clothing; obtain the style of the clothing in the image through a pre-trained clothing style recognition model; obtain the position information of the logo in the image, the position information of the clothing, and the clothing category corresponding to the clothing through a pre-trained trademark type recognition model.

[0110] A processing module 402, configured to, for each clothing category, determine the target type identifier corresponding to the clothing category according to the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to different type identifiers pre-stored.

[0111] In a possible implementation manner, the processing module 402 is further configured to obtain any sample image in the first sample set and the sample style identifier of each piece of clothing included in the sample image; input the sample image into the original recognition model to obtain the recognized style identifier of each piece of clothing included in the sample image; and train the original recognition model according to the sample style identifier and the recognized style identifier.

[0112] In a possible implementation manner, the processing module 402 is further configured to perform offline distillation learning training on a pre-configured clothing style recognition model according to the trained original recognition model to obtain a trained clothing style recognition model.

[0113] In a possible implementation manner, the processing module 402 is specifically configured to use a preset algorithm to determine the similarity between the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to each type identifier pre-stored; wherein the preset algorithm includes: a locality-sensitive hashing algorithm; and determine the type identifier corresponding to the image with the highest similarity and exceeding a preset threshold as the target type identifier corresponding to the clothing category.

[0114] In a possible implementation, the processing module 402 is further configured to determine, for the position information of each obtained logo, the degree of overlap between the position information of the logo and the position information of the obtained clothing; determine the position information of the clothing with the largest degree of overlap as the position information of the target clothing; and determine, according to the corresponding relationship between the position information of the clothing and the clothing category, that the target clothing category corresponding to the position information of the target clothing is the clothing category corresponding to the logo.

[0115] In a possible implementation, the receiving and obtaining module 401 is specifically configured to extract features of the image through the first network layer in a pre-trained type recognition model, and obtain the feature maps output by each sub-network layer in the first network layer; obtain each position information of the logo in the feature map through the detection layer in the type recognition model; wherein, each position information of the logo is determined according to the regions containing the logo boxed in the feature map by a preset number of anchor boxes with different ratios, and the preset number is not less than four; and obtain the position information of the logo based on each position information of the logo through the second network layer in the type recognition model.

[0116] In a possible implementation, the receiving and obtaining module 401 is further configured to obtain the aspect ratio of each logo and the aspect ratio of each piece of clothing in the second sample set according to the pre-annotated size of the logo and the size of the clothing in the second sample set; perform clustering on the aspect ratio of each logo and the aspect ratio of each piece of clothing in the second sample set through a clustering algorithm and the preset number; and determine the anchor ratio corresponding to each category according to the aspect ratio of the logo and the aspect ratio of the clothing in each category in the determined clustering result.

[0117] Embodiment 7:

[0118] Based on the above embodiments, Figure 5 A schematic diagram of an electronic device provided by an embodiment of the present invention is shown in Figure 5 As shown, it includes: a processor 501, a communication interface 502, a memory 503, and a communication bus 504. Among them, the processor 501, the communication interface 502, and the memory 503 complete communication with each other through the communication bus 504.

[0119] A computer program is stored in the memory 503. When the program is executed by the processor 501, the processor 501 is caused to execute the following steps:

[0120] Receive an image containing clothing; obtain the style of the clothing in the image through a pre-trained clothing style recognition model.

[0121] Obtain the position information of the logo, the position information of the clothing, and the clothing category corresponding to the clothing in the image through a pre-trained trademark type recognition model;

[0122] For each clothing category, determine the target type logo corresponding to the clothing category according to the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to different types of logos pre-saved.

[0123] In a possible implementation manner, the clothing style recognition model is trained in the following way:

[0124] Obtain any sample image in the first sample set and the sample style logo of each piece of clothing included in the sample image;

[0125] Input the sample image into the original recognition model to obtain the recognized style logo of each piece of clothing included in the sample image;

[0126] Train the original recognition model according to the sample style logo and the recognized style logo.

[0127] In a possible implementation manner, the method further includes:

[0128] Perform offline distillation learning training on the pre-configured clothing style recognition model according to the trained original recognition model to obtain the trained clothing style recognition model.

[0129] In a possible implementation manner, the determining the target type logo corresponding to the clothing category according to the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to different types of logos pre-saved includes:

[0130] Adopt a preset algorithm to determine the similarity between the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to each type of logo pre-saved; wherein, the preset algorithm includes: local sensitive hashing algorithm;

[0131] Determine the type logo corresponding to the image with the highest similarity and exceeding the preset threshold as the target type logo corresponding to the clothing category.

[0132] In a possible implementation manner, the method further includes:

[0133] For the position information of each logo obtained, determine the overlap degree between the position information of the logo and the position information of the clothing obtained; determine the position information of the clothing with the largest overlap degree as the position information of the target clothing; according to the corresponding relationship between the position information of the clothing and the clothing category, determine the target clothing category corresponding to the position information of the target clothing as the clothing category corresponding to the logo.

[0134] In a possible implementation manner, obtaining the position information of the logo of the clothing in the image through the pre-trained type recognition model includes:

[0135] Performing feature extraction on the image through the first network layer in the pre-trained type recognition model to obtain the feature maps output by each sub-network layer in the first network layer;

[0136] Obtaining each position information of the logo in the feature map through the detection layer in the type recognition model; wherein, each position information of the logo is determined according to the regions containing the logo framed in the feature map by a preset number of anchor boxes with different ratios, and the preset number is not less than four;

[0137] Obtaining the position information of the logo through the second network layer in the type recognition model based on each position information of the logo.

[0138] In a possible implementation manner, the ratio of the anchor box is obtained through the following method:

[0139] According to the sizes of the logos pre-annotated in the second sample set and the sizes of the clothing, obtaining the aspect ratios of the length and width of each logo and the aspect ratios of the length and width of each piece of clothing in the second sample set;

[0140] Performing clustering on the aspect ratios of the length and width of each logo and the aspect ratios of the length and width of each piece of clothing in the second sample set through a clustering algorithm and the preset number;

[0141] Determining the anchor box ratio corresponding to each category according to the aspect ratios of the length and width of the logos and the aspect ratios of the length and width of the clothing in each category in the determined clustering result.

[0142] The communication bus mentioned in the above server may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0143] The communication interface 502 is used for communication between the above electronic device and other devices.

[0144] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0145] The aforementioned processor may be a general-purpose processor, including a central processing unit, a Network Processor (NP), etc.; it may also be a Digital Signal Processing (DSP), an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0146] Example 8:

[0147] Based on the above embodiments, an embodiment of the present invention further provides a computer-readable storage medium, in which a computer program executable by an electronic device is stored. When the program runs on the electronic device, the electronic device is caused to execute the following steps when executing:

[0148] A computer program is stored in the memory. When the program is executed by the processor, the processor is caused to execute the following steps:

[0149] Receive an image containing clothing; obtain the style of the clothing in the image through a pre-trained clothing style recognition model;

[0150] Obtain the position information of the logo in the image, the position information of the clothing, and the clothing category corresponding to the clothing through a pre-trained trademark type recognition model;

[0151] For each clothing category, determine the target type identifier corresponding to the clothing category according to the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to different types of identifiers pre-stored.

[0152] In a possible implementation manner, the clothing style recognition model is trained in the following manner:

[0153] Obtain any sample image in the first sample set, and the sample style identifier of each piece of clothing included in the sample image;

[0154] Input the sample image into the original recognition model to obtain the recognized style identifier of each piece of clothing included in the sample image;

[0155] Train the original recognition model according to the sample style identifier and the recognized style identifier.

[0156] In one possible implementation, the method further includes:

[0157] Performing offline distillation learning training on a pre-configured clothing style recognition model according to the originally trained recognition model that has been completed, to obtain a trained clothing style recognition model.

[0158] In one possible implementation, the determining the target type identifier corresponding to the clothing category according to the position information of the identifier of the clothing corresponding to the clothing category and the images corresponding to different types of identifiers pre-stored includes:

[0159] Using a preset algorithm to determine the similarity between the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to each type of identifier pre-stored; wherein, the preset algorithm includes: local sensitive hashing algorithm;

[0160] Determining the type identifier corresponding to the image with the highest similarity and exceeding the preset threshold as the target type identifier corresponding to the clothing category.

[0161] In one possible implementation, the method further includes:

[0162] For the position information of each logo obtained, determining the overlap degree between the position information of the logo and the position information of the clothing obtained; determining the position information of the clothing with the largest overlap degree as the position information of the target clothing; according to the corresponding relationship between the position information of the clothing and the clothing category, determining the target clothing category corresponding to the position information of the target clothing as the clothing category corresponding to the logo.

[0163] In one possible implementation, the obtaining the position information of the logo of the clothing in the image through a pre-trained type recognition model includes:

[0164] Performing feature extraction on the image through the first network layer in the pre-trained type recognition model that has been completed, to obtain the feature maps output by each sub-network layer in the first network layer;

[0165] Obtaining each position information of the logo in the feature map through the detection layer in the type recognition model; wherein, each position information of the logo is determined according to the regions containing the logo boxed out in the feature map by a preset number of anchor boxes with different ratios, and the preset number is not less than four;

[0166] Obtaining the position information of the logo through the second network layer in the type recognition model based on each position information of the logo.

[0167] In one possible implementation, the ratio of the anchor box is obtained by the following method:

[0168] According to the sizes of the pre-annotated signs and the sizes of the clothing items in the second sample set, obtain the aspect ratios of each sign and the aspect ratios of each clothing item in the second sample set;

[0169] Cluster the aspect ratios of each sign and the aspect ratios of each clothing item in the second sample set through a clustering algorithm and the preset quantity;

[0170] According to the aspect ratios of the signs and the aspect ratios of the clothing items in each category in the determined clustering result, determine the anchor ratio corresponding to each category.

[0171] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0172] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in the process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0173] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in the process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0174] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide means for implementing the specified functions in the process Figure 1One or more processes and / or blocks Figure 1 Steps of functions specified in one or more blocks.

[0175] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A clothing recognition method, characterized in that, The method includes: Receiving an image containing clothing; inputting the image into a pre-trained clothing style recognition model to obtain the clothing style in the image output by the clothing style recognition model; Inputting the image into a pre-trained type recognition model to obtain the position information of the logo, the position information of the clothing, and the clothing category corresponding to the clothing in the image output by the type recognition model; wherein, the clothing style recognition model and the type recognition model are deep learning models; For each clothing category, using a preset algorithm to determine the similarity between the position information of the logo of the clothing corresponding to the clothing category and the images corresponding to each type identifier pre-stored; wherein, the preset algorithm includes: locality-sensitive hashing algorithm; determining the type identifier corresponding to the image with the highest similarity and exceeding the preset threshold as the target type identifier corresponding to the clothing category; wherein, the type identifier includes a trademark identifier.

2. The method according to claim 1, characterized in that, The clothing style recognition model is trained in the following manner: Obtaining any sample image in the first sample set, and the sample style identifier of each piece of clothing included in the sample image; Inputting the sample image into the original recognition model to obtain the recognized style identifier of each piece of clothing included in the sample image; Training the original recognition model according to the sample style identifier and the recognized style identifier.

3. The method according to claim 2, wherein The method further includes: Performing offline distillation learning training on the pre-configured clothing style recognition model according to the trained original recognition model to obtain the trained clothing style recognition model.

4. The method according to claim 1, wherein The method further includes: For the position information of each obtained logo, determining the overlap degree between the position information of the logo and the position information of the obtained clothing; determining the position information of the clothing with the largest overlap degree as the position information of the target clothing; according to the corresponding relationship between the position information of the clothing and the clothing category, determining the target clothing category corresponding to the position information of the target clothing as the clothing category corresponding to the logo.

5. The method according to claim 1, wherein Obtaining the position information of the logo of the clothing in the image through the pre-trained type recognition model includes: Performing feature extraction on the image through the first network layer in the pre-trained type recognition model to obtain the feature maps output by each sub-network layer in the first network layer; Obtaining each position information of the logo in the feature map through the detection layer in the type recognition model; wherein, each position information of the logo is determined by the regions containing the logo boxed in the feature map according to a preset number of anchor boxes with different ratios, and the preset number is not less than four; Obtaining the position information of the logo through the second network layer in the type recognition model based on each position information of the logo.

6. The method according to claim 5, characterized in that, The ratio of the anchor box is obtained in the following manner: According to the sizes of the logos and the sizes of the clothing pre-annotated in the second sample set, obtaining the aspect ratio of the length and width of each logo and the aspect ratio of the length and width of each piece of clothing in the second sample set; Performing clustering on the aspect ratio of the length and width of each logo and the aspect ratio of the length and width of each piece of clothing in the second sample set through a clustering algorithm and the preset number; Determine the anchor ratio corresponding to each category according to the aspect ratio of the logo and the aspect ratio of the clothing in each category in the determined clustering results.

7. A clothing recognition device, characterized in that, The device includes: A receiving and obtaining module, configured to receive an image including clothing; input the image into a pre-trained clothing style recognition model to obtain the clothing style in the image output by the clothing style recognition model; input the image into a pre-trained type recognition model to obtain the position information of the logo, the position information of the clothing, and the clothing category corresponding to the clothing in the image output by the type recognition model; wherein, the clothing style recognition model and the type recognition model are deep learning models; A processing module, configured to, for each clothing category, use a preset algorithm to determine the similarity between the position information of the logo of the clothing corresponding to the clothing category and the image corresponding to each type identifier pre-stored; wherein, the preset algorithm includes: a locality-sensitive hashing algorithm; determine the type identifier corresponding to the image with the highest similarity and exceeding the preset threshold as the target type identifier corresponding to the clothing category; wherein, the type identifier includes a trademark identifier.

8. An electronic device, characterized in that, The electronic device includes at least a processor and a memory, and the processor is configured to execute the steps of the clothing recognition method according to any one of claims 1-6 when executing a computer program stored in the memory.

9. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by a processor, it executes the steps of the clothing recognition method according to any one of claims 1-6.

Citation Information

Patent Citations

  • A method and equipment for carrying out target detection based on an SSD model

    CN109949359A

  • Clothes attribute identification method and device and electronic equipment

    CN111079757A

  • Personnel dressing monitoring method, device and equipment and storage medium

    CN111401301A