System and method for accurately segmenting shoes and clothes and identifying attributes based on end-to-end

Through the end-to-end precise segmentation and attribute recognition system of shoes and clothing, multi-scale feature extraction and information fusion, the efficiency and accuracy of instance segmentation and attribute recognition in the prior art are solved, real-time response and efficient multi-attribute recognition are achieved.

CN120356040APending Publication Date: 2025-07-22HANGZHOU DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410131025.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-31
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art cannot quickly and conveniently realize end-to-end real-time instance segmentation and attribute recognition of shoe and clothing pictures, and cannot integrate prior knowledge to improve the accuracy of instance segmentation, and cannot accurately classify ultra-fine-grained multi-label recognition.

Method used

The end-to-end precision segmentation and attribute recognition system for shoes and clothing, including picture encoder and feature decoder, uses the backbone network to extract multi-scale features, generate high-dimensional features through feature fusion modules and position encoders, and fuse information through object query modules and attribute query modules, and finally output mask segmentation and attribute recognition results through multi-layer rendering modules.

Benefits of technology

Real-time accurate instance segmentation and fine-grained attribute recognition of shoe and clothing pictures are realized, improving the accuracy of instance segmentation and the accuracy and generalization of attribute recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356040A_ABST
    Figure CN120356040A_ABST
Patent Text Reader

Abstract

The invention discloses an end-to-end based shoe and clothing precise segmentation and attribute recognition system and recognition method, and the system comprises a picture encoder which is used for receiving an input picture and then generating high-quality / high-dimension picture features; and the feature decoder is connected with the picture encoder, receives the high-quality / high-dimension picture features generated by the picture encoder, decodes the high-quality / high-dimension picture features into semantic information, and outputs a mask segmentation result and an attribute prediction result. According to the end-to-end based shoe and clothes precise segmentation and attribute identification system, the end-to-end shoe and clothes picture shoe and clothes fine instance segmentation and fine-grained multi-attribute identification system and the identification method, the requirements of a user on precise instance segmentation and fine-grained attribute identification of a plurality of shoe and clothes instances in a single picture can be met; and outputting the shoe and clothing vector of the instance level so as to facilitate subsequent similar shoe and clothing search and establishment of a shoe and clothing knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a shoe and clothing recognition system, and more specifically to a system and method for end-to-end accurate segmentation and attribute recognition of shoes and clothing. Background Art

[0002] Currently, the solutions for accurate segmentation and recognition analysis of shoes and clothing in the prior art can only meet the instance segmentation or attribute recognition capabilities. 1) It is impossible to quickly and conveniently complete instance segmentation and attribute recognition in real time end-to-end. 2) It is impossible to fuse the prior knowledge in instance segmentation and attribute recognition to improve the accuracy of instance segmentation. 3) It is impossible to accurately classify the multi-label recognition of hundreds of ultra-fine-grained levels. Summary of the Invention

[0003] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide an end-to-end system for fine instance segmentation and fine-grained multi-attribute recognition of shoe and clothing pictures, which can meet the user's needs for accurate instance segmentation and fine-grained attribute recognition of multiple shoe and clothing instances in a single picture, and output instance-level shoe and clothing vectors for subsequent similar shoe and clothing search and the establishment of a shoe and clothing knowledge graph.

[0004] To achieve the above purpose, the present invention provides the following technical solution: A system for end-to-end accurate segmentation and attribute recognition of shoes and clothing, characterized in that it includes:

[0005] A picture encoder, which is used to receive the input picture and then generate high-quality / high-dimensional picture features;

[0006] A feature decoder, connected to the picture encoder, receiving the high-quality / high-dimensional picture features generated by the picture encoder, and decoding the high-quality / high-dimensional picture features into semantic information, and outputting a mask segmentation result and an attribute prediction result.

[0007] As a further improvement of the present invention, the picture encoder includes:

[0008] A backbone network, which is used to receive the input picture and extract multi-scale features from the input picture;

[0009] A feature fusion module, connected to the backbone network, which is used to fuse the features extracted by the backbone network to generate high-quality / high-dimensional picture features;

[0010] A position encoder, connected to the feature fusion module, which is used to record the position information of the picture blocks.

[0011] As a further improvement of the present invention, the feature decoder includes:

[0012] An object query module and an attribute query module are used to receive high-quality / high-dimensional image features, and then perform instance segmentation to obtain object query information and attribute prediction information. The object query module and the attribute query module are connected through mask prediction to fuse the object query information and the attribute prediction information through the potential connection between the object query module and the attribute query module;

[0013] A multi-layer rendering module is connected to the feature fusion module to receive the high-quality / high-dimensional image features generated by the feature fusion module and perform more refined feature learning and fusion on the high-quality / high-dimensional image features; A query representation module is connected to the object query module and the attribute query module to decouple the fused object query information and attribute prediction information, and finally output a mask segmentation result and an attribute prediction result.

[0014] As a further improvement of the present invention, the backbone network is composed of a pre-trained swim transformer. On the other hand, the present invention provides an identification method, including the following steps:

[0015] Step 1, scale and normalize the input image and then input it into the image encoder;

[0016] Step 2, use the image encoder to receive the image and then generate high-quality / high-dimensional image features;

[0017] Step 3, use the feature decoder to receive the high-quality / high-dimensional image features, decode the high-dimensional features into semantic information, and output a mask segmentation result and an attribute recognition result.

[0018] As a further improvement of the identification method, the specific steps for generating high-quality / high-dimensional image features in Step 2 are as follows:

[0019] Step 2-1, use the backbone network to extract features from the input image;

[0020] Step 2-2, use the features extracted in Step 2 to generate multi-scale features;

[0021] Step 2-3, add information position encoding to the segmented image through the position encoder;

[0022] Step 2-4, fuse the multi-scale features generated in Step 3 through the feature fusion module and splice the position encoding added in Step 4 to output high-dimensional features.

[0023] As a further improvement of the identification method, the specific steps for decoding the high-dimensional features into semantic information and outputting a mask segmentation result and an attribute recognition result in Step 3 are as follows:

[0024] Step 3-1: Fuse the object query information and attribute prediction information output by the object query module and the attribute query module;

[0025] Step 3-2: Input the fused object query information and attribute prediction information into the multi-layer rendering module, and then use two independent dynamic convolutions and forward networks in the multi-layer rendering module to output the mask segmentation result and the attribute recognition result.

[0026] As a further improvement of the recognition method, the specific method of fusing the object query information and the attribute prediction information in Step 3-1 is as follows: First, use the high-quality / high-dimensional image features combined with the mask segmentation result and the attribute recognition result output by the multi-layer rendering module to enable the object query module to learn mask prediction and the attribute query module to learn attribute label prediction. Then, splice the object query information and the attribute prediction information, and use the multi-layer fully connected network in the multi-layer rendering module to fuse the object query information and the attribute prediction information.

[0027] As a further improvement of the recognition method, the specific method of outputting the mask segmentation result and the attribute recognition result in Step 3-2 is:

[0028] Use two two-layer forward networks to output the mask category prediction and the mask prediction result respectively;

[0029] And, use a multi-label classification head to output the attribute label prediction result;

[0030] Among them, the multi-label classification head and the two two-layer forward networks are both built in the query representation module.

[0031] The beneficial effects of the present invention are as follows: By fusing instance segmentation and attribute recognition end-to-end, real-time response is achieved; based on the attention mechanism, the prior knowledge of attribute annotation is used to improve the accuracy of instance segmentation; using the attention mechanism, a new classification head is developed, which significantly improves the accuracy and generalization of attribute recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a schematic flowchart of the recognition method for end-to-end accurate segmentation and attribute recognition of shoes and clothing according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The following will further elaborate on the present invention in combination with the embodiments given in the drawings.

[0034] A system for end-to-end accurate segmentation and attribute recognition of shoes and clothing according to this embodiment includes:

[0035] An image encoder, which is used to receive the input image and then generate high-quality / high-dimensional image features;

[0036] A feature decoder, connected to the image encoder, receives the high-quality / high-dimensional image features generated by the image encoder, decodes the high-quality / high-dimensional image features into semantic information, and outputs a mask segmentation result and an attribute prediction result. Through the settings of the above-mentioned image encoder and feature decoder, the features of the image can be simply and effectively extracted, and then high-quality / high-dimensional image features can be generated. After that, by decoding the high-quality / high-dimensional image features, the mask segmentation result and the attribute prediction result can be simply and effectively output.

[0037] As a specific implementation of the improvement, the image encoder includes:

[0038] A backbone network, which is used to receive the input image and extract multi-scale features from the input image;

[0039] A feature fusion module, connected to the backbone network, which is used to fuse the features extracted by the backbone network to generate high-quality / high-dimensional image features;

[0040] A position encoder, connected to the feature fusion module, which is used to record the position information of the image patches. Through the settings of the above-mentioned network modules, the effect of simply and effectively extracting features and then generating high-quality / high-dimensional image features can be achieved.

[0041] As a specific implementation of the improvement, the feature decoder includes:

[0042] An object query module and an attribute query module, which are used to receive the high-quality / high-dimensional image features, and then perform instance segmentation to obtain object query information and attribute prediction information. The object query module and the attribute query module are connected through mask prediction, so that the object query information and the attribute prediction information can be fused through the potential connection between the object query module and the attribute query module;

[0043] A multi-layer rendering module, connected to the feature fusion module, to receive the high-quality / high-dimensional image features generated by the feature fusion module, and perform more refined feature learning and fusion on the high-quality / high-dimensional image features; A query representation module, connected to the object query module and the attribute query module, to decouple the fused object query information and attribute prediction information, and finally output a mask segmentation result and an attribute prediction result. Through the settings of the above-mentioned modules, the use of high-quality / high-dimensional image features and then the output of the mask segmentation result and the attribute prediction result can be simply and effectively achieved. The object query module and the attribute query module in this embodiment both have learnable parameters. Through the connection method of mask prediction, the two can learn the potential connection between them and help each other optimize the performance.

[0044] As a specific implementation of the improvement, the backbone network is composed of a pre-trained swim transformer. In this way, it is possible to use the pre-trained swim transformer as the backbone network and the feature pyramid network, which can be used to generate multi-scale feature vectors and effectively reduce the construction cost of the backbone network. On the other hand, this embodiment provides an identification method based on the above system, as Figure 1 shown, including the following steps:

[0045] Step 1, scale and normalize the input image, and then input it into the image encoder;

[0046] Step 2, use the image encoder to receive the image and then generate high-quality / high-dimensional image features;

[0047] Step 3, use the feature decoder to receive the high-quality / high-dimensional image features, decode the high-dimensional features into semantic information, and output the mask segmentation result and the attribute recognition result. Through the settings of the above three steps, the image encoder and the feature decoder of the system can be effectively used to process the image, and finally the mask segmentation result and the attribute recognition result are output.

[0048] The specific methods of the above three steps are as follows:

[0049] The specific steps of generating high-quality / high-dimensional image features in Step 2 are as follows:

[0050] Step 2-1, use the backbone network to extract features from the input image;

[0051] Step 2-2, use the features extracted in Step 2 to generate multi-scale features;

[0052] Step 2-3, add information position encoding to the segmented image through the position encoder;

[0053] Step 2-4, fuse the multi-scale features generated in Step 3 through the feature fusion module and splice the position encoding added in Step 4 to output high-dimensional features;

[0054] The specific steps of decoding the high-dimensional features into semantic information and outputting the mask segmentation result and the attribute recognition result in Step 3 are as follows:

[0055] Step 3-1, fuse the object query information and the attribute prediction information output by the object query module and the attribute query module;

[0056] Step 3-2, input the fused object query information and attribute prediction information into the multi-layer rendering module, and then use the two independent dynamic convolutions and the forward network in the multi-layer rendering module to output the mask segmentation result and the attribute recognition result.

[0057] The specific method for fusing the object query information and the attribute prediction information in Step 31 is as follows: First, the high-quality / high-dimensional image features are combined with the mask segmentation result and the attribute recognition result output by the multi-layer rendering module, so that the object query module learns mask prediction and the attribute query module learns attribute label prediction. Then, the object query information and the attribute prediction information are concatenated, and the multi-layer fully connected network in the multi-layer rendering module is used to fuse the object query information and the attribute prediction information.

[0058] The specific method for outputting the mask segmentation result and the attribute recognition result in Step 32 is:

[0059] Two two-layer forward networks are used to output the mask category prediction and the mask prediction result respectively. Here, the mask category and the mask prediction result are combined into the mask segmentation result;

[0060] And, a multi-label classification head is used to output the attribute label prediction result, and the attribute label prediction result is used as the attribute recognition result;

[0061] Among them, the multi-label classification head and the two two-layer forward networks are both built in the query representation module;

[0062] Thus, it can be seen that through the above three steps, the image encoder and the feature decoder in the above system can be effectively used to process the input image, and then the mask segmentation result and the attribute recognition result are output.

[0063] In summary, the system and recognition method for end-to-end accurate segmentation and attribute recognition of shoes and clothing in this embodiment fuse instance segmentation and attribute recognition end-to-end to achieve real-time response; and based on the attention mechanism, use the prior knowledge of attribute annotation to improve the accuracy of instance segmentation; use the attention mechanism to develop a new classification head, which significantly improves the accuracy and generalization of attribute recognition.

[0064] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and retouches should also be regarded as the protection scope of the present invention.

Claims

1. An end-to-end system for precise segmentation and attribute recognition of shoes and clothing, characterized in that: Comprising: An image encoder for receiving an input image and then generating high-quality / high-dimensional image features; Feature A decoder, connected to the image encoder, receiving the high-quality / high-dimensional image features generated by the image encoder, and decoding the high-quality / high-dimensional image features into semantic information, outputting a mask segmentation result and an attribute prediction result.

2. The system for accurate segmentation and attribute recognition of shoes and clothing based on end-to-end according to claim 1, characterized in that: The image encoder includes: A backbone network for receiving the input image and extracting multi-scale features from the input image; A feature fusion module, connected to the backbone network, for fusing the features extracted by the backbone network to generate high-quality / high-dimensional image features; A position encoder, connected to the feature fusion module, for recording the position information of the image patches.

3. The system for precise segmentation and attribute recognition of shoes and clothing based on end-to-end according to claim 2, characterized in that: The feature decoder includes: An object query module and an attribute query module for receiving the high-quality / high-dimensional image features, and then performing instance segmentation to obtain object query information and attribute prediction information. The object query module and the attribute query module are connected through mask prediction to enable the fusion of the object query information and the attribute prediction information through the potential connection between the object query module and the attribute query module; A multi-layer rendering module, connected to the feature fusion module, to receive the high-quality / high-dimensional image features generated by the feature fusion module, and perform more refined feature learning and fusion on the high-quality / high-dimensional image features; A query representation module, connected to the object query module and the attribute query module, to decouple the fused object query information and attribute prediction information, and finally output a mask segmentation result and an attribute prediction result.

4. The system for end-to-end accurate segmentation and attribute recognition of footwear and clothing according to claim 3, characterized in that: The backbone network is composed of a pre-trained swim transformer.

5. A recognition method using the system according to any one of claims 1 to 3, characterized in that: Including the following steps: Step 1, scale and normalize the input image and then input it into the image encoder; Step 2, use the image encoder to receive the image and then generate high-quality / high-dimensional image features; Step 3, use the feature decoder to receive the high-quality / high-dimensional image features, decode the high-dimensional features into semantic information, and output a mask segmentation result and an attribute recognition result.

6. The recognition method according to claim 5, characterized in that: The specific steps for generating the high-quality / high-dimensional image features in Step 2 are as follows: Step 2-1, use the backbone network to extract features from the input image; Step 2-2, use the features extracted in Step 2 to generate multi-scale features; Step 2-3, add information position encoding to the segmented image through the position encoder; Step 2-4, fuse the multi-scale features generated in Step 3 through the feature fusion module and splice the position encoding added in Step 4 to output high-dimensional features.

7. The recognition method according to claim 6, characterized in that: The specific steps for decoding the high-dimensional features into semantic information and outputting a mask segmentation result and an attribute recognition result in Step 3 are as follows: Step 3-1, fuse the object query information and the attribute prediction information output by the object query module and the attribute query module; Step 3-2, input the fused object query information and attribute prediction information into the multi-layer rendering module, and then use two independent dynamic convolutions and forward networks in the multi-layer rendering module to output a mask segmentation result and an attribute recognition result.

8. The recognition method according to claim 7, wherein: The specific method for fusing the object query information and the attribute prediction information in Step 31 is as follows: First, use the high-quality / high-dimensional image features, combined with the mask segmentation result and the attribute recognition result output by the multi-layer rendering module, to enable the object query module to learn mask prediction and the attribute query module to learn attribute label prediction. Then, splice the object query information and the attribute prediction information, and use the multi-layer fully connected network in the multi-layer rendering module to fuse the object query information and the attribute prediction information.

9. The recognition method according to claim 8, characterized in that: The specific method for outputting the mask segmentation result and the attribute recognition result in Step 32 is: Use two two-layer forward networks to respectively output the mask category prediction and the mask prediction result; And use the multi-label classification head to output the attribute label prediction result; Among them, the multi-label classification head and the two two-layer forward networks are both built in the query representation module.