Clothing Recognition Method, Device, Equipment and Storage Medium

By extracting keyframes in the video and performing clothing area detection and feature fusion methods, the problem of low clothing recognition accuracy in the prior art is solved, and higher recognition accuracy and efficiency are achieved, reducing the ambiguity of cross-time styles.

CN113887426BActive Publication Date: 2025-07-29BEIJING IQIYI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202111164025.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2025-07-29
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

The existing clothing recognition technology has low recognition accuracy in videos, making it difficult to meet the needs of actual application scenarios. In particular, there is a problem of ambiguity in cross-time styles, and it is easy to identify the same clothing as different styles or different styles as the same style.

Method used

By extracting multiple key frames of different moments from the video to be identified, clothing area detection and feature extraction are performed, target clothing areas with similar features and high time domain consistency are divided into one category based on clustering technology, and image features of various target clothing areas are fused into one target image feature for identification.

Benefits of technology

It improves the accuracy of clothing recognition, reduces the ambiguity problem of identifying the same clothing as different models or different models as the same models at different moments, improves the time-domain consistency of the recognition results and improves the recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113887426B_ABST
    Figure CN113887426B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a clothing recognition method, device, equipment and storage medium. The method includes: extracting multiple key frames at different times from a video to be recognized; detecting clothing regions in the key frames; for the detected target clothing regions, extracting features of the target clothing regions, where the features of the target clothing regions at least include image features; based on the features of each target clothing region, clustering each target clothing region to obtain at least one class of target clothing regions; for each class of target clothing regions, fusing the image features of each target clothing region in the class into a target image feature, and based on the target image feature, performing clothing recognition to obtain the recognition result of the class of target clothing regions. In this way, the accuracy of clothing recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a clothing recognition method, device, equipment and storage medium. Background Art

[0002] Clothing is an essential item in people's daily lives. With the emergence of online shopping platforms, people often purchase clothing on online shopping platforms, which greatly facilitates people's lives. In related technologies, in order to further provide convenience for people, for the clothing that appears in the video, clothing recognition technology can be used to identify it and provide people with information about relevant clothing. The existing clothing recognition technology is still in the development stage, and the recognition accuracy is still relatively low, making it difficult to meet the recognition requirements of actual application scenarios. For example, currently, the maximum accuracy of most clothing recognition technologies based on open-source datasets is lower than 70%, making it difficult to meet the recognition requirement of an accuracy higher than 90% required by actual application scenarios. In particular, the following recognition problems are likely to occur: Since clothing is a non-rigid object, differences in human postures are likely to cause situations such as occlusion and incomplete exposure of the clothing area, which may result in the same piece of clothing appearing at two different times, such as t1 (for example, the moment when the front of the clothing is shown) and t2 (for example, the moment when the back of the clothing is shown), being recognized as different pieces of clothing, or two different pieces of clothing appearing at two different times, such as t3 (for example, the back of a pure white T-shirt) and t4 (for example, the back of a T-shirt with a large pattern on the front chest but a pure white back), being recognized as the same piece of clothing. This is a severe challenge in clothing recognition in videos. Summary of the Invention

[0003] The purpose of the embodiments of the present invention is to provide a clothing recognition method, device, equipment and storage medium to improve the accuracy of clothing recognition. The specific technical solutions are as follows:

[0004] In the first aspect implemented by the present invention, first, a clothing recognition method is provided, including:

[0005] Extracting multiple key frames at different times from the video to be recognized;

[0006] Performing clothing area detection on the key frames;

[0007] For the detected target clothing area, extracting the features of the target clothing area, where the features of the target clothing area at least include image features;

[0008] Based on the features of each target clothing area, clustering each target clothing area to obtain at least one category of target clothing area;

[0009] For each type of target clothing region, the image features of each target clothing region in the type of target clothing region are fused into a target image feature, and clothing recognition is performed based on the target image feature to obtain the recognition result of the type of target clothing region.

[0010] In a second aspect of the implementation of the present invention, there is also provided a clothing recognition device, including:

[0011] A key frame extraction module, configured to extract multiple key frames at different times from the video to be recognized;

[0012] A region detection module, configured to perform clothing region detection on the key frames;

[0013] A feature extraction module, configured to extract the features of the detected target clothing region, and the features of the target clothing region at least include image features;

[0014] A region clustering module, configured to cluster each target clothing region based on the features of each target clothing region to obtain at least one type of target clothing region;

[0015] A clothing recognition module, configured to, for each type of target clothing region, fuse the image features of each target clothing region in the type of target clothing region into a target image feature, and perform clothing recognition based on the target image feature to obtain the recognition result of the type of target clothing region.

[0016] In yet another aspect of the implementation of the present invention, there is also provided an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus;

[0017] The memory is used to store a computer program;

[0018] The processor is configured to, when executing the program stored on the memory, implement the steps of the clothing recognition method described in any one of the above.

[0019] In yet another aspect of the implementation of the present invention, there is also provided a computer-readable storage medium, in which instructions are stored, and when it runs on a computer, it causes the computer to execute the clothing recognition method described in any one of the above.

[0020] In yet another aspect of the implementation of the present invention, there is also provided a computer program product containing instructions, and when it runs on a computer, it causes the computer to execute the clothing recognition method described in any one of the above.

[0021] The clothing recognition method, device, equipment, and storage medium provided by the embodiments of the present invention perform clustering on each target clothing area corresponding to multiple key frames at different times in the video to be recognized, based on the characteristics of each target clothing area. Target clothing areas with similar characteristics and high temporal consistency can be classified into one category. Then, clothing recognition is performed based on one category of target clothing areas. The image features of each target clothing area in this category of target clothing areas are fused into a target image feature, and based on this target image feature, clothing recognition is performed to obtain the recognition result of this category of target clothing areas. In this way, considering the factor of temporal consistency, the image features of each target clothing area with similar features and high temporal consistency are fused, that is, image feature fusion is performed in the time domain. Clothing recognition is performed based on the target image feature that fuses the image features of each target clothing area. Compared with clothing recognition based on the image features of the target clothing area at a single moment, the recognition result is more accurate. In this way, the influence of the target clothing area at a certain single moment is reduced, the temporal consistency of the clothing recognition result in the video is improved, and the cross-time ambiguity problems in the related art, such as the same clothing at different times being recognized as different styles of clothing, or different styles of clothing at different times being recognized as the same style of clothing, are reduced, thereby improving the clothing recognition accuracy. And because the image features of one category of target clothing areas are fused to obtain one target clothing area, only one clothing recognition needs to be performed on one target clothing area. Compared with performing one clothing recognition on each target clothing area, the clothing recognition efficiency is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.

[0023] Figure 1 It is an exemplary system architecture diagram in the embodiments of the present invention.

[0024] Figure 2 It is a flowchart of an exemplary clothing recognition method in the embodiments of the present invention.

[0025] Figure 3 It is a flowchart of an exemplary clothing recognition method in the embodiments of the present invention.

[0026] Figure 4 It is a flowchart of an exemplary clothing recognition method in the embodiments of the present invention.

[0027] Figure 5 It is a schematic structural diagram of an exemplary clothing recognition device in the embodiments of the present invention.

[0028] Figure 6It is a schematic structural diagram of an exemplary clothing recognition device in an embodiment of the present invention.

[0029] Figure 7 It is a schematic structural diagram of an exemplary electronic device for implementing a clothing recognition method in an embodiment of the present invention. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present invention will be described with reference to the accompanying drawings in the embodiments of the present invention.

[0031] Clothing is an essential item in people's daily lives. With the emergence of online shopping platforms, people often purchase clothing on online shopping platforms, which greatly facilitates people's lives. In related technologies, in order to further provide convenience for people, for the clothing appearing in the video, clothing recognition technology can be used for recognition to provide people with information about relevant clothing. The existing clothing recognition technology is still in the development stage, and the recognition accuracy is still relatively low, making it difficult to meet the recognition requirements of actual application scenarios. For example, currently, the maximum accuracy of most clothing recognition technologies based on open-source datasets is lower than 70%, making it difficult to meet the recognition requirement of an accuracy higher than 90% required by actual application scenarios. In particular, the following recognition problems are likely to occur: Since clothing is a non-rigid object, differences in human postures are likely to cause situations such as occlusion and incomplete exposure of the clothing area, which may result in the same piece of clothing appearing at two different times, t1 (for example, the moment when the front of the clothing is shown) and t2 (for example, the moment when the back of the clothing is shown), being recognized as different pieces of clothing (i.e., one-to-many), or two different pieces of clothing appearing at two different times, t3 (for example, the back of a pure white T-shirt) and t4 (for example, the back of a T-shirt with a large pattern on the front chest but a pure white back), being recognized as the same piece of clothing (i.e., many-to-one) ambiguity problems, which are severe challenges in clothing recognition in videos.

[0032] Regarding the cross-time style ambiguity problem existing in clothing recognition of videos, the embodiments of the present invention provide a clothing recognition method. This method combines clustering technology and image feature fusion technology in the time domain to effectively reduce the one-to-many and many-to-one mappings between the query clothing and the styles in the database, improve the temporal consistency of the clothing recognition results in the video, improve the rationality of the clothing same-style or similar-style recognition results at the video level, and thus improve the user experience. The clothing recognition method provided by the embodiments of the present invention can be executed by a server or a user terminal, etc. Taking the execution by the server as an example, Figure 1This is a system architecture diagram in an embodiment of the present invention. The system includes a server and user terminals. In practical applications, the server can execute the clothing recognition method provided in the embodiment of the present invention, and apply the results to processing scenarios such as clothing recognition or clothing recommendation, and send the processing results to the user terminals, so that the users of the user terminals can understand the information of the clothing. The solution provided in the embodiment of the present invention will be described in more detail below.

[0033] Figure 2 This is a flowchart of an exemplary clothing recognition method provided in an embodiment of the present invention. As Figure 2 shown, the clothing recognition method provided in this embodiment may at least include the following steps:

[0034] Step 201: Extract multiple key frames at different times from the video to be recognized.

[0035] Step 202: Detect the clothing regions in the key frames.

[0036] Step 203: Extract the features of the target clothing regions for the detected target clothing regions, and the features of the target clothing regions at least include image features.

[0037] Step 204: Cluster the target clothing regions based on the features of each target clothing region to obtain at least one class of target clothing regions.

[0038] Step 205: For each class of target clothing regions, fuse the image features of the target clothing regions in this class of target clothing regions into a target image feature, and perform clothing recognition based on the target image feature to obtain the recognition result of this class of target clothing regions.

[0039] Among them, the video to be recognized is the video to be subjected to clothing recognition. After extracting the key frames from the video to be recognized, detect the clothing regions in each key frame, and the detected regions for clothing recognition are used as the target clothing regions. The image features of the target clothing regions refer to the features in the spatial domain of the key frames. Here, the spatial domain refers to the space where the pixels of the key frames are located.

[0040] Clustering is the process of dividing a set of physical or abstract objects into multiple classes composed of similar objects. Then, clustering the target clothing regions based on the features of each target clothing region can divide the target clothing regions with similar features into one class. In the time domain, since multiple key frames come from multiple different times, the target clothing regions corresponding to the multiple key frames also come from multiple different times. It can be understood that one class of target clothing regions also comes from different times, and most of the times when the target clothing regions in one class with similar features are located are close, that is, the time domain consistency is relatively high.

[0041] After clustering, one type of target clothing region or multiple types of target clothing regions can be obtained. For example, if there is actually only one type of clothing in the video to be recognized, then the target clothing regions corresponding to the key frames at different times are all of the same type of clothing. After clustering, one type of target clothing region corresponding to this type of clothing may be obtained, or multiple types of target clothing regions corresponding to this type of clothing may be obtained. If there are multiple types of clothing in the video to be recognized, then after clustering, multiple types of target clothing regions are obtained. For example, for the same type of clothing, one type of target clothing region corresponding to this type of clothing may be obtained, or multiple types of target clothing regions corresponding to this type of clothing may be obtained. Assume that there are two types of clothing in the video to be recognized. After clustering, three types of target clothing regions are obtained, where two of them are the target clothing regions of one type of clothing, and one is the target clothing region of the other type of clothing.

[0042] Based on this, in this embodiment, for each target clothing region corresponding to multiple key frames at different times in the video to be recognized, clustering is performed based on the characteristics of each target clothing region. Target clothing regions with similar characteristics and high temporal consistency can be classified into one type. Then, clothing recognition is performed based on one type of target clothing region. The image features of each target clothing region in this type of target clothing region are fused into a target image feature, and based on this target image feature, clothing recognition is performed to obtain the recognition result of this type of target clothing region. In this way, considering the factor of temporal consistency, the image features of each target clothing region with similar characteristics and high temporal consistency are fused, that is, image feature fusion is performed in the time domain. Clothing recognition is performed based on the target image feature that has fused the image features of each target clothing region. Compared with clothing recognition based on the image features of the target clothing region at a single time, the recognition result is more accurate. In this way, the influence of the target clothing region at a single time is reduced, the temporal consistency of the clothing recognition result in the video is improved, and the cross-time ambiguity problems in the related art, such as the same type of clothing at different times being recognized as different types of clothing, or different types of clothing at different times being recognized as the same type of clothing, are reduced, thereby improving the clothing recognition accuracy. And because the image features of one type of target clothing region are fused to obtain one target clothing region, only one clothing recognition needs to be performed on one target clothing region. Compared with performing one clothing recognition on each target clothing region, the clothing recognition efficiency is greatly improved.

[0043] In one embodiment, the features of the target clothing area may further include time domain features. The time domain features of the target clothing area represent the moment when the key frame where the target clothing area is located appears in the video to be identified, that is, the moment when the target clothing area appears in the video to be identified. In this way, when clustering is performed, clustering can be performed based on the image features and time domain features of each target clothing area. Clustering based on image features can implicitly classify areas with higher time domain consistency into one category. In this embodiment, further, the time domain features are used as a display feature, in parallel with the image features, as the basis for clustering, thereby further improving the consistency of each type of target clothing area in the time domain, further avoiding the situation in the related art where the same clothing at different times is identified as different clothing or different clothing at different times is identified as the same clothing, thereby improving the accuracy of clothing recognition.

[0044] Based on this, each target clothing area is clustered based on its characteristics to obtain at least one type of target clothing area, such as Figure 3 As shown, its specific implementation may include:

[0045] Step 301: For each target clothing area, determine an associated clothing area corresponding to the target clothing area, and determine association information corresponding to the target clothing area, wherein the associated clothing area is from a set of candidate clothing areas, the candidate clothing areas included in the set of candidate clothing areas are target clothing areas characterized by time domain features earlier than the target clothing area, the associated clothing area is a candidate clothing area in the set of candidate clothing areas having features most similar to those of the target clothing area, and the association information includes an identifier combination formed by an identifier of the target clothing area and an identifier of the corresponding associated clothing area.

[0046] Step 302: Classify the identification combinations corresponding to the target clothing areas to obtain at least one type of identification combination, wherein the identification combinations of the same type can be spliced in sequence from beginning to end to form an identification string, and the identification of the spliced parts of the two spliced identification combinations is the same.

[0047] Step 303: For each type of logo combination, each target clothing area corresponding to the logo combination is regarded as a type.

[0048] It should be noted that if the target clothing region has no corresponding candidate clothing region, no processing is performed. For example, the target clothing region with the earliest time sequence has no corresponding candidate clothing region and no processing is performed.

[0049] For example, key frames 1, 2, and 3 are sequentially extracted from the video to be recognized according to the time sequence. Among them, two target clothing regions are detected in key frame 1, and the identifiers of these two target clothing regions are 1 and 2 respectively. Two target clothing regions are detected in key frame 2, and the identifiers of these two target clothing regions are 3 and 4 respectively. Two target clothing regions are detected in key frame 3, and the identifiers of these two target clothing regions are 5 and 6 respectively.

[0050] Taking the target clothing region 4 in key frame 2 as an example, the associated clothing region of this target clothing region 2 comes from the candidate clothing region set formed by the two target clothing regions in key frame 1 with an earlier time sequence. Assuming that the feature of the target clothing region 2 in key frame 1 is the most similar to the feature of the target clothing region 4, then the associated clothing region of the target clothing region 4 is the target clothing region 2. Establish the association information between the target clothing region 4 and the target clothing region 2. This association information includes the identifier combination formed by the identifier of the target clothing region 4 and the identifier of the target clothing region 2. For example: 4, 2. Accordingly, the same processing is performed on the target clothing region 3, the target clothing region 5, and the target clothing region 6. Since the target clothing region 1 and the target clothing region 2 are the earliest in time sequence, no processing is required.

[0051] Assume that the identifier combinations: 3, 1, the identifier combination: 5, 4, and the identifier combination: 6, 3 are also obtained. Based on this, the identifier combination: 6, 3 and the identifier combination: 3, 1 can be classified into one category because these two identifier combinations can be spliced to form the identifier string 6, 3, 1, and the spliced part of the two identifier combinations has the identifier "3". Then, the target clothing regions 6, 3, and 1 corresponding to this category of identifier combinations are classified into one category; and the identifier combination: 5, 4 and the identifier combination: 4, 2 can be classified into one category because these two identifier combinations can be spliced to form the identifier string 5, 4, 2, and the spliced part of the two identifier combinations has the identifier "4". The target clothing regions 5, 4, and 2 corresponding to this category of identifier combinations are classified into one category.

[0052] For a target clothing region, the time moment represented by the time domain feature of the associated clothing region is earlier than the time moment represented by the time domain feature of this target clothing region. That is, the time sequence of the associated clothing region is before the target clothing region. In this embodiment, a target clothing region is associated with the most similar target clothing region with an earlier time sequence, and the association information is established through the identifier combination, that is, the forward mapping relationship from the target clothing region to the associated clothing region is established. Then, through the identifier string formed by connecting the heads and tails of the identifier combinations, the target clothing regions with the most similar features and the same time domain are associated, and the associated target clothing regions are classified into one category, so that the target clothing regions with the most similar features and the same time domain are classified into one category.

[0053] In one implementation, after classifying each type of identification combination and treating each target clothing area corresponding to the type of identification combination as one category, before fusing the image features of each target clothing area in each type of target clothing area into a target image feature for each type of target clothing area and performing clothing recognition based on the target image feature to obtain the recognition result of the type of target clothing area, the method may further include: using at least one type of classified target clothing area as at least one clustering object for reclustering to obtain at least one type of target clothing area corresponding to the reclustering.

[0054] Correspondingly, for each type of target clothing area, fusing the image features of each target clothing area in the type of target clothing area into a target image feature and performing clothing recognition based on the target image feature to obtain the recognition result of the type of target clothing area may specifically include: for each type of target clothing area in at least one type of target clothing area corresponding to the reclustering, fusing the image features of each target clothing area in the type of target clothing area into a target image feature and performing clothing recognition based on the target image feature to obtain the recognition result of the type of target clothing area.

[0055] In this embodiment, after the first clustering is performed by associating target clothing areas with the most similar features and consistent time domain, among at least one type of target clothing area, there may be two or more types of target clothing areas that belong to the target clothing areas of the same clothing style. Therefore, a second clustering is performed here, that is, a hierarchical clustering is performed. In this way, target clothing areas with the most similar features and consistent time domain can be further divided into one category, which not only makes the classification more accurate, but also makes the information of the fused target image feature richer, thereby further improving the accuracy of clothing recognition.

[0056] In one implementation, for each type of target clothing area, fusing the image features of each target clothing area in the type of target clothing area into a target image feature and performing clothing recognition based on the target image feature to obtain the recognition result of the type of target clothing area may be specifically implemented as: for each type of target clothing area in at least one type of classified target clothing area, fusing the image features of each target clothing area in the type of target clothing area into a target image feature and performing clothing recognition based on the target image feature to obtain the recognition result of the type of target clothing area.

[0057] Then, the above-mentioned clothing recognition method may further include: using at least one type of classified target clothing area as at least one clustering object for re-clustering to obtain at least one type of target clothing area corresponding to the re-clustering; for each type of target clothing area among the at least one type of target clothing area corresponding to the re-clustering, selecting at least one candidate style from the recognition results of the clustering objects in this type of target clothing area as the recognition result of this type of target clothing area.

[0058] In this way, after the first clustering, based on the results of fine-grained clustering, clothing recognition is performed. Since the number of categories is large and the categories are divided in a more detailed manner at this time, for a type of clothing area, there are fewer interfering clothing areas, and the recognition result is more accurate. After the second clustering, the recognition results of the clustering objects in a type of clothing area are then fused, and at least one candidate style is selected from them as the recognition result of this type of clothing area, and this recognition result is also more accurate.

[0059] Specifically, selecting at least one candidate style from the recognition results of the clustering objects in this type of target clothing area may specifically include: forming a candidate style set based on the candidate styles in the recognition results of the clustering objects in this type of target clothing area; calculating the score of each candidate style in the candidate style set; and selecting at least one candidate style from the candidate style set based on the scores of the candidate styles.

[0060] For example, a type of target clothing area L1 includes target clothing areas a, b, and c. The recognition result of target clothing area a includes candidate styles h1, h2, h3, and h4. The recognition result of target clothing area b includes candidate styles h1, h2, h3, and h5. The recognition result of target clothing area c includes candidate styles h1, h2, h4, and h6. Then, the candidate styles in the recognition result of target clothing area a, the candidate styles in the recognition result of target clothing area b, and the candidate styles in the recognition result of target clothing area c form a candidate style set P1, and this candidate style set P1 includes candidate styles h1, h2, h3, h4, h5, and h6. For each candidate style in this candidate style set P1, calculate the score of this candidate style, and select at least one candidate style from the candidate style set P1 based on the scores of the candidate styles. For example, candidate styles h1 and h2 are selected.

[0061] Among them, the score of the candidate style can characterize the rationality of the candidate style as the clothing recognition result. In this embodiment, by calculating the score of the candidate style, a selection basis is provided for the selection of the candidate style, and a reasonable candidate style can be accurately selected, further improving the accuracy of clothing recognition.

[0062] If the score of the candidate style is higher, the rationality is higher. Then, when selecting at least one candidate style, at least one candidate style with the highest score can be selected. In this way, the accuracy of clothing recognition is further improved.

[0063] In one implementation manner, calculate the score of each candidate style in the candidate style set, and its specific implementation manner may include: for each candidate style in the candidate style set, based on the similarity between each feature of the candidate style and the same feature of the corresponding target clothing area, and / or, the first hit count, and / or, the second hit count, calculate the score of the candidate style; where the first hit count is the total hit count of the candidate style in the recognition results of all target clothing areas forming the candidate style set; the second hit count is the total hit count of the candidate style in the recognition results of all target clothing areas in the video to be recognized.

[0064] The similarity of features is an important factor in measuring the rationality of the candidate style. The more similar the features are, the more likely it is to be the same style. Therefore, in this embodiment, the similarity of features is used as the calculation basis for the score of the candidate style.

[0065] Among them, hitting for the first time, that is, the number of occurrences.

[0066] Continuing with the above example of the candidate style set P1, the candidate style set P1 includes candidate style h1, candidate style h2, candidate style h3, candidate style h4, candidate style h5, and candidate style h6.

[0067] The first hit count is the total hit count of the candidate style in the recognition results of the target clothing area a, the target clothing area b, and the target clothing area c that form the candidate style set P1. Taking the candidate style h5 as an example, the candidate style h5 hits 0 times in the recognition result of the target clothing area a, the candidate style h1 hits 1 time in the recognition result of the target clothing area b, and the candidate style h1 hits 0 times in the recognition result of the target clothing area c. At this time, the total hit count is statistically 1 time, that is, the first hit count of the candidate style h5 is 1.

[0068] The first hit count can reflect the hit frequency of a candidate style in the recognition results of a category of target clothing regions. In a category of target clothing regions, since the characteristics of each target clothing region are similar and the time-domain consistency is relatively high, it is more likely to be the same style. If the hit frequency of a candidate style in the recognition results of a category of target clothing regions is higher, that is, the recognition results of more target clothing regions hit the same candidate style, then the candidate style is more likely to be accurate, and it is more reasonable to use the candidate style as the recognition result. Therefore, the first hit count is an important factor for measuring the rationality of a candidate style and can be used as the calculation basis for scoring the candidate style.

[0069] Suppose that after clustering all the target clothing regions corresponding to the video to be recognized, two categories of target clothing regions are obtained. In addition to the above-mentioned category of target clothing regions L1, it also includes another category of target clothing regions L2. The other category of target clothing regions L2 includes target clothing regions d, e, and f. The recognition results of target clothing region d include candidate styles h7, h8, h9, and h5. The recognition results of target clothing region e include candidate styles h7, h8, h9, and h11. The recognition results of target clothing region f include candidate styles h7, h8, h10, and h12.

[0070] The second hit count is the total hit count of a candidate style in the recognition results of each target clothing region in a category of target clothing regions L1 and the recognition results of each target clothing region in another category of target clothing regions L2. Taking candidate style h5 as an example, candidate style h5 hits 0 times in the recognition result of target clothing region a, candidate style h1 hits 1 time in the recognition result of target clothing region b, candidate style h1 hits 0 times in the recognition result of target clothing region c, candidate style h5 hits 1 time in the recognition result of target clothing region d, candidate style h5 hits 0 times in the recognition result of target clothing region e, and candidate style h5 hits 0 times in the recognition result of target clothing region f. At this time, the total hit count is statistically 2 times, that is, the second hit count of candidate style h5 is 2.

[0071] The second hit count can reflect the hit frequency of the candidate style in the recognition results of each target clothing area of the entire video to be recognized. From the perspective of the entire video to be recognized, the greater the difference between the recognition results of different types of target clothing areas is better, because the features of different types of target clothing areas are not similar and the temporal consistency is poor, so it is more likely that they are different styles. Therefore, what we expect more is that the candidate style is only hit by one type of target clothing area. If the candidate style is hit by the recognition results of at least one type of target clothing area, the distinguishability for different styles is relatively poor, and it is unreasonable to use this candidate style as the recognition result. Taking an extreme case as an example, if there are actually multiple different styles in the entire video to be recognized, and a candidate style appears in the recognition results of each type of target clothing area of the entire video to be recognized, then there is no distinguishability at all and it is not suitable to be output as the recognition result. Therefore, the second hit count is also an important factor in measuring the rationality of the candidate style and can also be used as the calculation basis for the scoring of the candidate style.

[0072] When calculating the score of the candidate style h5, based on the similarity between the texture feature of the candidate style h5 and the texture feature of the corresponding target clothing area, the similarity between the color feature of the candidate style h5 and the color feature of the corresponding target clothing area, the first hit count, and the second hit count, calculate the score of the candidate style h5.

[0073] In this embodiment, when calculating the score of the candidate style, in addition to considering the important factor of feature similarity, the hit frequency of the candidate style in the recognition results of one type of target clothing area and the hit frequency in the recognition results of each target clothing area of the entire video to be recognized are also considered. The factors considered are very comprehensive, and a more accurate score of the candidate style can be calculated to select a more reasonable candidate style.

[0074] In one implementation manner, based on the similarity between each feature of the candidate style and the same feature of the corresponding target clothing area, and / or the first hit count, and / or the second hit count, calculate the score of the candidate style. The specific implementation method may include: weighted summation of the similarity, the first hit count, and the ratio of the first hit count to the second hit count to obtain the score of the candidate style. In this embodiment, by the method of weighted summation of the similarity, the first hit count, and the ratio of the first hit count to the second hit count, the score of the candidate style can be calculated more accurately.

[0075] In addition, since we expect that the higher the score of the candidate style, the higher the rationality of the candidate style as the clothing result. As mentioned before, the higher the first hit count, the higher the rationality, and the higher the second hit count, the lower the rationality. Therefore, here the ratio of the first hit count to the second hit count is used as one item to participate in the weighted summation. Based on this, the higher the first hit count, the higher the score of the candidate style, and the higher the second hit count, the lower the score of the candidate style, and the rationality of the candidate style can be accurately evaluated.

[0076] Alternatively, a preset constant can also be directly used to replace the first hit count, and the ratio of the preset constant to the second hit count is used as one item to participate in the weighted summation.

[0077] In one implementation, the weight of the above similarity is greater than the weight of the first hit count, and / or the weight of the first hit count is greater than the weight of the ratio.

[0078] In practical applications, since clothing recognition needs to be mainly based on features, the weight of features is the largest, the weight of the first hit count is the second, and finally the weight of the ratio of the first hit count to the second hit count is the smallest. In this way, the output of unreasonable candidate styles can be minimized.

[0079] It should be noted that the above is only an example of a method for calculating the score of a candidate style, and the score of a candidate style can also be calculated by other methods. For example, the similarities corresponding to at least one feature of the candidate style are weighted and summed. Or, the similarities corresponding to at least one feature of the candidate style and the first hit count are weighted and summed. Or, the similarities corresponding to at least one feature of the candidate style and the second hit count are weighted and summed. Or, the first hit count and the second hit count are weighted and summed, and so on.

[0080] It should be noted that the above is only an example of an implementation method for selecting at least one candidate style, and at least one candidate style can also be selected by other methods. For example, the similarities or the first hit counts corresponding to the above candidate styles are directly sorted, and at least one candidate style is selected based on the sorting result. For example, at least one candidate style with the highest similarity or the most first hit counts is selected.

[0081] The above is only an exemplary way to cluster each target clothing area. It can be understood that clustering can also be performed in other ways. For example, clustering algorithms such as K-Means can be used to cluster each target clothing area, and so on.

[0082] In one implementation, for each target clothing region, the associated clothing region corresponding to the target clothing region is determined. The specific implementation method may include: for each target clothing region, the feature distance between the features of the target clothing region and the features of each candidate clothing region in the candidate clothing region set is determined respectively, and a candidate clothing region with the smallest feature distance is selected as the associated clothing region. The feature distance is used to measure the similarity between two objects. The smaller the feature distance, the more similar. Therefore, in this embodiment, by determining the feature distance between the features of the candidate clothing region and the features of the target clothing region, the most similar candidate clothing region is accurately selected as the associated clothing region.

[0083] In one implementation, the feature distance between the features of the target clothing region and the features of each candidate clothing region in the candidate clothing region set is determined respectively. The specific implementation method may include: calculating the distance between the image features of the target clothing region and the image features of the candidate clothing region as the first distance; calculating the distance between the time-domain features of the target clothing region and the time-domain features of the candidate clothing region as the second distance; determining the weight coefficient of the first distance based on the second distance; and determining the feature distance between the features of the target clothing region and the features of the candidate clothing region based on the first distance and the weight coefficient. Wherein, the larger the second distance, the larger the weight coefficient of the first distance, and the larger the feature distance.

[0084] The above-mentioned first distance, that is, the image feature distance, is used to measure the similarity of the image features of two clothing regions. The smaller the image feature distance, the more similar. The larger the image feature distance, the less similar. The above-mentioned second distance, that is, the time-domain interval distance, is used to measure the time-domain consistency of two clothing regions. The smaller the time-domain interval distance, the higher the time-domain consistency. The larger the time-domain interval distance, the lower the time-domain consistency.

[0085] In practical applications, although the image feature distance between the target clothing region and the candidate clothing region is very small, that is, the two clothing regions are very similar, but the time-domain interval distance between them is very large. This is very likely to be from two different clothing items. Therefore, the feature distance brought by the combination of the image feature distance and the time-domain interval distance is relatively large.

[0086] For example, at a certain moment, the back of a pure white T-shirt is shown, and at another moment, the front chest with a large area of patterns is shown, but the back of the T-shirt with a pure white back. These are two different clothing items that appear at two relatively far times. From the perspective of image features, the backs are all pure white. However, due to the relatively large time-domain interval, it is considered that the feature distance is relatively large when combined.

[0087] Therefore, in this embodiment, based on the second distance, the weight coefficient of the first distance is determined, and based on the first distance and the weight coefficient, the characteristic distance is obtained. The larger the second distance, the larger the weight coefficient applied to the first distance, so that the characteristic distance obtained from the first distance and its weight coefficient is larger. In this way, the characteristic distance can be accurately obtained.

[0088] Exemplarily, the characteristic distance can be calculated according to the following formula:

[0089] D = gamma + D1 * scale_ratio(1)

[0090] where D represents the characteristic distance, D1 represents the first distance, scale_ratio represents the weight coefficient of the first distance, and gamma is a preset constant.

[0091] In one implementation, the distance between the time-domain feature of the target clothing region and the time-domain feature of the candidate clothing region is calculated as the second distance. The specific implementation manner may include: calculating the time interval between the time represented by the time-domain feature of the target clothing region and the time represented by the time-domain feature of the candidate clothing region; if the time interval is less than or equal to the first preset interval, determining the second distance as the first preset interval; if the time interval is greater than the first preset interval and less than the second preset interval, determining the second distance as the time interval; if the time interval is greater than the first preset interval and greater than or equal to the second preset interval, determining the second distance as the second preset interval; where the second preset interval is greater than the first preset interval.

[0092] In practical applications, if the time interval corresponding to the actual target clothing area and the candidate clothing area is relatively small, within a certain range, it can be considered that both belong to the range with a relatively high time-domain consistency, and the effects on the above-mentioned feature distances are similar and can be not distinguished. If the actual time interval is relatively large and exceeds a certain range, both belong to the range with a relatively low time-domain consistency, and the effects on the above-mentioned feature distances are similar and can also be not distinguished. Based on this, in this embodiment, in order to simplify the processing, two time intervals are preset, namely the first preset interval and the second preset interval. Among them, one value is larger and the other value is smaller, and the second preset interval can be set to be greater than the first preset interval. Based on this, the time interval corresponding to the actual target clothing area and the candidate clothing area is compared with the first preset interval and the second preset interval respectively. If the time interval is less than or equal to the first preset interval, it can be considered that both belong to the range with a relatively high time-domain consistency, and the second distance is uniformly determined to be the first preset interval; if the time interval is greater than the first preset interval and less than the second preset interval, the actual time-domain consistency degree can be adopted, and the second distance is determined to be the time interval; if the time interval is greater than the first preset interval and greater than or equal to the second preset interval, it can be considered that both belong to the range with a relatively low time-domain consistency, and the second distance is uniformly determined to be the second preset interval, thus simplifying the processing.

[0093] Exemplarily, the second distance can be calculated by the following formula:

[0094] tm_coef = min(max(0, tm1 – tm2 - min_tm_gap), max_tm_gap) * PI / max_tm_gap (2)

[0095] Among them, tm_coef represents the second distance, tm1 represents the moment represented by the time-domain feature of the target clothing area, tm2 represents the moment represented by the time-domain feature of the candidate clothing area, min() represents the minimum value function, max() represents the maximum value function, min_tm_gap represents the first preset interval, max_tm_gap represents the second preset interval, and PI is a preset constant.

[0096] The specific values of the first preset interval and the second preset interval can be set according to the actual situation.

[0097] Exemplarily, based on the second distance, the weight coefficient of the first distance is determined. Specifically, it can be implemented by the following formula:

[0098] scale_ratio = (alpha + cosine(tm_coef)) * beta (3)

[0099] Among them, cosine() represents the value of the cosine function, and alpha and beta are both preset constants.

[0100] In the above - related embodiments, only the method of calculating the feature distance is illustrated by way of example. The feature distance can also be calculated by other methods. For example, the feature distance can be calculated by weighting the first distance and the second distance.

[0101] In one implementation, the image features of each target clothing region in this type of target clothing region are fused into a target image feature. The specific implementation method may include: weighting and averaging the image features of each target clothing region in this type of target clothing region to obtain the target image feature. In this embodiment, the image features of each target clothing region are fused by weighting and averaging, which is simple to implement and the obtained target image feature is very accurate.

[0102] In addition, the image features of each target clothing region in this type of target clothing region can also be fused into a target image feature by other methods. For example, the image features of each target clothing region in this type of target clothing region can be clustered, and weights are assigned according to the number of categories where the image features are located. The larger the number, the greater the weight. Then, based on the assigned weights, the image features of this type of target clothing region are weighted and summed.

[0103] In one implementation, after detecting the clothing region of the key frame to obtain at least one target clothing region corresponding to the key frame, and before extracting at least one feature of the target clothing region, the above - mentioned clothing recognition method may further include: detecting the human body region of the key frame; in response to detecting at least one clothing region and at least one human body region, comparing the clothing region with each human body region, and based on the comparison result, determining whether the clothing region is a clothing region to be removed, where the overlapping ratio of the clothing region to be removed and the human body region reaches a first threshold and the pose similarity does not exceed a second threshold; removing the clothing region to be removed in the key frame.

[0104] If the overlapping ratio of the target clothing region and the human body region reaches the first threshold, it can be considered that the clothing included in the target clothing region may be worn on the human body included in the human body region. Further, if the pose similarity between the target clothing region and the human body region does not exceed the second threshold, it can be considered that the pose of the clothing included in the target clothing region may be different from the pose of the human body included in the human body region. At this time, the possibility that the clothing included in the target clothing region is worn on the human body included in the human body region is excluded, and it is considered that the target clothing region is not worn on the human body included in the human body region. This situation may be caused by factors such as the target clothing region being blocked by a passing human body, which may affect the clothing recognition result. It is determined as a clothing region to be removed, and then the clothing region to be removed in the image to be processed is removed. The specific values of the first threshold and the second threshold can be set according to the actual situation and are not specifically limited here.

[0105] Based on this, in this embodiment, after detecting the target clothing area of the image to be processed, the target clothing area is compared with each human body area. Based on this, it is determined whether the target clothing area is a clothing area to be removed, and the overlapping ratio of the clothing area to be removed and the human body area reaches a first threshold and the pose similarity does not exceed a second threshold. In this way, by combining at least one human body area in the image, that is, combining the context information, some target clothing areas that affect recognition caused by factors such as being blocked by passing human bodies are removed, and the target clothing areas outside the clothing areas to be removed in the image to be processed are determined as valid target clothing areas. In this way, the low-quality target clothing areas that affect recognition are removed. Generally speaking, the number of target clothing areas participating in subsequent clothing recognition is reduced, laying a good foundation for subsequent clothing recognition, thereby improving the accuracy of clothing recognition.

[0106] The pose can be a wearing state or a non-wearing state (such as a flat state). For clothing in the wearing state, it may be blocked by other passing human bodies. For clothing in the non-wearing state, it may also be blocked by other passing human bodies. In this case, the target clothing areas will affect the clothing recognition result and need to be removed.

[0107] The above-mentioned comparison of the target clothing area with each human body area and determining whether the target clothing area is a clothing area to be removed based on the comparison result can be specifically implemented as follows: calculating the intersection-over-union ratio of the target clothing area and each human body area; based on the calculation result, determining that the target clothing area is a clothing area to be removed.

[0108] In practical applications, the human body area of the image to be processed can be detected in advance. The detected clothing area and human body area are framed by a rectangular box.

[0109] The intersection-over-union ratio is the ratio of the intersection of two boxes to their union. The intersection-over-union ratio can measure the overlapping ratio of two boxes. Here, the intersection-over-union ratio of the target clothing area and the human body area can reflect the overlapping ratio of the target clothing area and the human body area, that is, the relative size of the overlap. From this, it can be reflected whether the clothing included in the target clothing area may be worn on the human body included in the human body area. Based on this, some rationality judgments can be made to determine whether the target clothing area is a clothing area to be removed, so as to remove some detected unreasonable target clothing areas.

[0110] In one implementation manner, the above-mentioned determination that the target clothing area is a clothing area to be removed based on the calculation result can be specifically implemented as follows: in response to the maximum value among the intersection-over-union ratios being greater than or equal to the first threshold and the pose similarity corresponding to the maximum value not exceeding the second threshold, determining that the target clothing area is a clothing area to be removed.

[0111] In this embodiment, if the maximum value among the above-mentioned intersection over union ratios is greater than or equal to the first threshold, and the pose similarity corresponding to the maximum value does not exceed the second threshold, it indicates that at least one human body has blocked the target clothing area. Therefore, the target clothing area can be removed to improve the accuracy of subsequent clothing recognition.

[0112] Exemplarily, the value of the first threshold can be less than or equal to 0.5. It can be understood that the first threshold can be set to other values according to the actual situation.

[0113] In one implementation manner, the image to be processed contains two or more human body areas. Based on the calculation results, the target clothing area is determined as the clothing area to be removed. The specific implementation manner can include: in response to two or more intersection over union ratios among the intersection over union ratios being greater than or equal to the first threshold, determining the target clothing area as the clothing area to be removed.

[0114] In practical applications, for clothing in a worn state, if two or more of the above-mentioned intersection over union ratios are greater than or equal to the first threshold, it indicates that the clothing included in the detected target clothing area may be worn on the human bodies included in multiple human body areas. However, in reality, a piece of clothing can only be worn on one human body, which does not conform to the actual situation, indicating that the detected target clothing area is problematic and is very likely to be an overlapping area caused by occlusion between multiple human bodies. For clothing in a non-worn state, if two or more of the above-mentioned intersection over union ratios are greater than or equal to the first threshold, it indicates that at least two human bodies have simultaneously blocked the target clothing area. Such a target clothing area is ambiguous. Therefore, the target clothing area can be removed to improve the accuracy of subsequent clothing recognition.

[0115] Further, in response to two or more intersection over union ratios among the intersection over union ratios being greater than or equal to the first threshold and at least one of the two or more intersection over union ratios corresponding to a pose similarity that does not exceed the second threshold, determining the target clothing area as the clothing area to be removed. For clothing in a worn state, the pose of the clothing included in the target clothing area is different from the pose of the human body included in at least one of the above-mentioned possible worn human body areas, indicating that at least one human body has blocked it. In this embodiment, by combining the factor of pose similarity, the accuracy of the determined clothing area to be removed is further improved.

[0116] Further, in response to more than two intersection-over-union ratios among the intersection-over-union ratios being greater than or equal to the first threshold and the pose similarities corresponding to the more than two intersection-over-union ratios not exceeding the second threshold, it is determined that the target clothing region is the clothing region to be removed. For clothing in a non-worn state, if at least two people block it simultaneously, since the clothing is not worn on the human body, then the pose of the clothing included in the target clothing region is different from the pose of the human body included in each human body region. Considering the factor of pose similarity further improves the accuracy of the determined clothing region to be removed.

[0117] In one implementation, the human body region includes at least one key human body pose node; based on the calculation result, it is determined that the target clothing region is the clothing region to be removed, and its specific implementation manner may include: determining the human body region corresponding to the maximum value among the intersection-over-union ratios as the target human body region; obtaining the style category of the identified target clothing region; obtaining at least one key human body pose node included in the wearing part of the pre-set identified style category; counting the number of the same key human body pose nodes included in the wearing part of the identified style category and the target human body region; in response to the counted number not exceeding the second threshold, determining the target clothing region as the clothing region to be removed.

[0118] In practical applications, the detected human body region is a rectangular frame including at least one key human body pose node. The key human body pose nodes may be the left eye, right eye, left ear, right ear, nose, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left foot, right foot, etc. Specifically, it can be detected by a pre-trained human body region detection model, and relevant technologies can be referred to, which will not be elaborated here. By selecting different key human body pose nodes, different parts of the human body can be formed. For example, the upper body part of the human body can be formed by the neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, and right hip.

[0119] Taking the human body region corresponding to the maximum value among the intersection-over-union ratios as the target human body region means that it is considered that the clothing included in the target clothing region is worn on the human body included in the target human body region.

[0120] In addition, the style category of the target clothing region can also be recognized in advance. The style category may include categories such as short-sleeved tops, long-sleeved tops, vests, shorts, long pants, knee-length skirts, and one-piece dresses. Specifically, a pre-trained category classifier can be used to recognize the style category of the target clothing region.

[0121] In this embodiment, the number of the same key body posture nodes between the worn parts of the recognized style categories and the target human body area does not exceed the second threshold, indicating that the worn parts of the clothing included in the clothing area do not match the parts of the human body formed by the key body posture nodes included in the target human body area, that is, there is a conflict. For example, the key body posture nodes included in the target human body area belong to the upper body part of the human body, while the recognized style category is trousers, which are actually worn on the lower body of the human body. Then, there is a conflict between the two. The recognition result of such a clothing area may be ambiguous and needs to be removed to improve the accuracy of subsequent clothing recognition.

[0122] In an exemplary embodiment, for the image to be recognized, target clothing area and human body area detection are performed. The specific implementation manner may include: performing target clothing area detection on the image to be recognized; in response to detecting at least one target clothing area, performing posture category recognition on each target clothing area, where the posture category includes a worn state and an unworn state; in response to the posture category of the target clothing area being the worn state, performing human body area detection on the image to be recognized. In this way, for the target clothing area in the worn state, the clothing area that needs to be removed can be determined.

[0123] Correspondingly, based on the calculation result, determining that the target clothing area is the clothing area that needs to be removed. The specific implementation manner may further include: in response to the maximum value among the intersection-over-union ratios being less than the first threshold, determining that the target clothing area is the clothing area that needs to be removed. In practical applications, if the maximum value among the above intersection-over-union ratios is less than the first threshold, it means that the intersection-over-union ratios are relatively small. Then, the clothing of the detected target clothing area may not be worn on any human body in any human body area, which does not match the previously recognized posture category of the worn state, indicating that the detected target clothing area is problematic and is very likely to be an invalid area without any clothing actually. Therefore, this target clothing area can be filtered out to improve the accuracy of subsequent clothing recognition.

[0124] Then, when extracting the features of the target clothing area, specifically, for each non-removed target clothing area, the features of the target clothing area are extracted. It should be noted that if all the target clothing areas are finally removed, the subsequent clothing recognition steps may not be performed. A prompt message indicating that no clothing is recognized can be sent.

[0125] In one implementation manner, the above image features may include texture features and may also include color features. In this embodiment, color features are added on the basis of texture features. Although a certain amount of color features are included in the texture features, they are not obvious. Here, the color features and texture features are listed as display features together, which can increase the proportion of color features, eliminate the errors caused by color, and improve the recognition accuracy.

[0126] Taking a specific application scenario as an example, a clothing recognition method provided by an embodiment of the present invention will be described in more detail below.

[0127] In this embodiment, aiming at the problem of cross-time style ambiguity in clothing recognition of videos, a clothing clustering method combining image features and time-domain interval metrics is proposed, and then a video clothing style recognition method combining clothing clustering, time-domain feature fusion, and image clothing recognition technology is proposed, improving the rationality of the recognition results of the same or similar styles of clothing at the video level.

[0128] As Figure 4 shown, the video clothing recognition method proposed in this embodiment mainly includes several main steps: video key frame extraction, clothing region detection, clothing region clustering, image feature aggregation (i.e., fusion) of clothing regions, clothing recognition, and two-way tracking of recognition results. The detailed process is as follows:

[0129] The first step: Video key frame extraction.

[0130] In this step, K key frames at different times are extracted from the video to be recognized, obtaining key frame (or frame) 1, ……, key frame K.

[0131] The second step: Clothing region detection.

[0132] In this step, clothing region detection (i.e., clothing region detection) is performed on the key frames. In addition, the effectiveness of the target clothing region can be further judged. Specifically, it can be: for each detected target clothing region, clothing category recognition is performed, including style category and pose category, and the pose category includes the wearing state and the non-wearing state; if the pose category of the target clothing region is the wearing state, human region detection is performed on the key frame where the target clothing region is located. Based on at least one human region detected in the key frame where the target clothing region is located, the target clothing region is compared with each human region. Based on the comparison result, the target clothing region to be removed is determined, and each target clothing region other than the target clothing region to be removed in the key frame is determined as the effective clothing region (i.e., the effectiveness judgment of the clothing region), obtaining the set of effective clothing regions in the video. The specific implementation method can refer to the description of the above relevant embodiments and will not be elaborated here.

[0133] Among them, the extraction method of key frames can refer to related technologies and will not be elaborated here.

[0134] In addition, a clothing category classifier can be pre-trained, and this clothing category classifier is used to recognize clothing categories and output the pose category and style category of the clothing.

[0135] The third step: Clothing region clustering.

[0136] In this step, a hierarchical clustering method for clothing regions that combines image and time-domain features is adopted to cluster clothing regions. Specifically as follows:

[0137] First, extract the image features and time-domain features of the target clothing region. Among them, the image features can be texture features and color features, and the specific extraction methods can refer to related technologies and will not be elaborated here. The time-domain feature is the time when the key frame where the target clothing region is located appears in the video to be recognized.

[0138] Secondly, combine the image feature distance and the time-domain interval distance to construct a distance metric matrix between all valid target clothing regions.

[0139] Among them, the distance metric matrix D is an N*N matrix, where N represents the number of valid target clothing regions in the video obtained in the second step. D(i,j) is an element in the distance metric matrix D, and i and j are the identifiers of the target clothing regions among the N target clothing regions. j represents the jth target clothing region, and D(i,j) represents the feature distance between the features of the ith target clothing region and the jth target clothing region. Among them, the distance metric matrix D is constructed using a label mapping algorithm, and the "i,j" included in D(i,j) is the forward mapping label from target clothing region i to target clothing region j.

[0140] Then, assign values to each element D(i,j) in the distance metric matrix D. Specifically, it can be:

[0141] Calculate the distance between the image feature i_feat of the ith target clothing region and the image feature j_feat of the jth target clothing region to obtain the image feature distance. Specifically, the similarity simi_val(i,j) can be calculated, and the similarity can be calculated through a pre-constructed similarity function simi(). The calculation formula is as follows:

[0142] simi_val(i,j) = simi(i_feat,j_feat) (4)

[0143] Since the greater the similarity, the smaller the distance, -simi_val(i,j) can be taken as the image feature distance.

[0144] Calculate the distance between the time-domain feature i_tm of the ith target clothing region and the image feature j_tm of the jth target clothing region to obtain the time-domain interval distance. The calculation formula is as follows:

[0145] tm_coef = min(max(0,i_tm – j_tm - min_tm_gap),max_tm_gap)*PI / max_tm_gap(5)

[0146] Calculate the weight coefficient of the image feature distance. The calculation formula is shown in Formula (3).

[0147] Calculate D(i, j) based on the image feature distance and the weight coefficient. The calculation formula is as follows:

[0148] D(i, j) = gamma - simi_val(i, j) * scale_ratio (6)

[0149] Exemplarily, min_tm_gap = 30 and max_tm_gap = 300 can be set.

[0150] Since it is necessary to determine the j-th target clothing area with the smallest feature distance (i.e., the most similar feature) in terms of time sequence for the i-th target clothing area as the associated clothing area, to meet this condition, the following settings can be adopted to increase the forward mapping label constraint:

[0151] (1) If i = j, it means that the j-th target clothing area and the i-th target clothing area are the same clothing area, which does not meet the requirements. At this time, set D(i, j) = val1, where val1 is the first preset feature distance. Exemplarily, it is set to 2.0.

[0152] (2) If i_tm = j_tm, it means that the j-th target clothing area and the i-th target clothing area are at the same moment, which does not meet the requirements. At this time, set D(i, j) = val2, where val2 is the second preset feature distance. Exemplarily, it is set to 3.0.

[0153] (3) If i_tm < j_tm, it means that in terms of time sequence, the j-th target clothing area is later than the i-th target clothing area, which does not meet the requirements. At this time, set D(i, j) = val3, where val3 is the third preset feature distance. Exemplarily, it is set to 4.0.

[0154] The above first preset feature distance, second preset feature distance, and third preset feature distance are all greater than the calculated feature distance and will not be selected when choosing the smallest feature distance.

[0155] Based on the above distance metric matrix D, select an element D(i, j) with the smallest feature distance corresponding to the i-th target clothing area. Based on the forward mapping label included in the selected D(i, j), obtain the forward mapping relationship i, j (i.e., the association information) corresponding to the i-th target clothing area.

[0156] Cluster the obtained forward mapping relationships to obtain a fine-grained clustering result. The forward mapping relationships in the same class can be concatenated in the order of head and tail to form an identification string. In two concatenated forward mapping relationships, the identifications of the concatenated parts are the same. Perform secondary clustering on the obtained fine-grained clustering result, which can be achieved by using the general KMeans clustering method, to obtain a hierarchical clustering result. Finally, obtain C subclasses, namely subclass 1, ……, subclass C. The obtained subclass 1 includes the target clothing area 1 in key frame 1, the target clothing area 1 in key frame 2, ……; ……; subclass C includes the target clothing area 1 in key frame 1, the target clothing area 2 in key frame K, ……

[0157] Step 4: Aggregate the image features of the clothing area.

[0158] In this step, weight and average the image features of each target clothing area in each class of target clothing areas to obtain the target image features. The specific formula is as follows:

[0159] feat_cluster_p = ∑ f∈{R(p)} w f feat f (7)

[0160] where R(p) represents the set of all target clothing areas included in subclass p, f represents the target clothing area in R(p), num({R(p)}) represents the number of target clothing areas included in subclass p, feat f represents the vector of image features, w f represents the weighting coefficient.

[0161] Step 4: Clothing recognition.

[0162] In this step, perform clothing recognition on the target image features corresponding to each subclass to obtain the recognition results of each subclass.

[0163] In implementation, the texture features and color features of candidate styles can be extracted in advance and saved in the database. When performing clothing recognition, based on the similarity between the texture features and color features of the target clothing area and the texture features and color features of the candidate styles, a set of candidate styles can be obtained.

[0164] By determining the time point of the key frame where the target clothing area is located, and thus based on the recognition result of the target clothing area, the recognition result of the key frame where the target clothing area is located is obtained. In this way, the recognition result of the target clothing area is extended to the time points of all key frames included in the class where the target clothing area is located, improving the clothing recall duration at the video level.

[0165] Step 5: Bidirectional tracking of the video sequence.

[0166] In this step, region detection can be performed on the video frames before and after the key frame in the video to be identified. The video frame in which the target clothing region appears in the video frame sequence of the video to be identified is determined. The time point at which the target clothing region appears is obtained, and the spatial location information of the target clothing region is determined. This allows the time point and location information of clothing with the same or similar styles appearing in the video to be identified to be obtained. Based on the determined time point and location information, the display of the clothing recognition results is controlled.

[0167] In this embodiment, the clustering method is optimized, and a hierarchical clustering method combining image and time domain features is used to effectively improve the video-level clothing clustering effect. By combining image clothing recognition, clustering technology, and time domain feature fusion, a method for identifying the same or similar clothing styles with time domain consistency is provided. This effectively reduces the one-to-many and many-to-one mapping between the query clothing and the styles in the database, improves the time domain consistency of the video recognition results, and effectively improves the rationality of clothing recognition results at the video level. It can be used in scenarios such as clothing recommendations of the same or similar styles (such as the same or similar styles worn by celebrities) in videos and clothing recognition, thereby improving the user experience. Based on this, this solution can adopt a more general implementation method in the clothing area detection and clothing recognition steps, reducing computational complexity.

[0168] Figure 5 This is a schematic diagram of the structure of an exemplary clothing recognition device provided by an embodiment of the present invention. Figure 5 As shown, the device 500 includes:

[0169] A key frame extraction module 501 is used to extract key frames at different moments from the video to be identified;

[0170] The region detection module 502 is used to perform clothing region detection on the key frame;

[0171] A feature extraction module 503 is used to extract features of the detected target clothing area, where the features of the target clothing area include at least image features;

[0172] A region clustering module 504 is configured to cluster the target clothing regions based on the characteristics of the target clothing regions to obtain at least one type of target clothing regions;

[0173] The clothing recognition module 505 is used to fuse the image features of each target clothing area in each target clothing area into a target image feature, and perform clothing recognition based on the target image feature to obtain a recognition result of the target clothing area.

[0174] In one embodiment, the characteristics of the target clothing area further include time domain characteristics.

[0175] In one implementation, the region clustering module 504 is specifically configured to:

[0176] For each target clothing region, determine the associated clothing region corresponding to the target clothing region, and determine the associated information corresponding to the target clothing region, where the associated clothing region is from the set of candidate clothing regions, and the candidate clothing regions included in the set of candidate clothing regions are target clothing regions whose time domain features represent a time earlier than that of the target clothing region, the associated clothing region is the candidate clothing region in the set of candidate clothing regions with the most similar features to the target clothing region, and the associated information includes the identification combination formed by the identification of the target clothing region and the identification of the corresponding associated clothing region;

[0177] Classify the identification combinations corresponding to each target clothing region to obtain at least one class of identification combinations, where the identification combinations of the same class can be spliced end to end to form an identification string, and in two spliced identification combinations, the identifications of the spliced parts are the same;

[0178] For each class of identification combinations, regard the target clothing regions corresponding to the class of identification combinations as one class.

[0179] In one implementation, the region clustering module 504 is specifically configured to:

[0180] After, for each class of identification combinations, regarding the target clothing regions corresponding to the class of identification combinations as one class, and before, for each class of target clothing regions, fusing the image features of the target clothing regions in the class of target clothing regions into a target image feature, and performing clothing recognition based on the target image feature to obtain the recognition result of the class of target clothing regions, use the at least one class of classified target clothing regions as at least one clustering object for reclustering to obtain at least one class of target clothing regions corresponding to the reclustering;

[0181] The clothing recognition module is specifically configured to:

[0182] For each class of target clothing regions in the at least one class of target clothing regions corresponding to the reclustering, fuse the image features of the target clothing regions in the class of target clothing regions into a target image feature, and perform clothing recognition based on the target image feature to obtain the recognition result of the class of target clothing regions.

[0183] In one implementation, the clothing recognition module is specifically configured to:

[0184] For each type of target clothing area among the at least one type of classified target clothing areas, fuse the image features of each target clothing area in this type of target clothing area into a target image feature, and perform clothing recognition based on the target image feature to obtain the recognition result of this type of target clothing area;

[0185] The area clustering module is further configured to: use the at least one type of classified target clothing areas as at least one clustering object, and perform re-clustering to obtain at least one type of target clothing areas corresponding to the re-clustering;

[0186] The clothing recognition module is further configured to: for each type of target clothing area among the at least one type of target clothing areas corresponding to the re-clustering, select at least one candidate style from the recognition results of each clustering object in this type of target clothing area as the recognition result of this type of target clothing area.

[0187] In one implementation, the area clustering module 504 is specifically configured to:

[0188] For each target clothing area, respectively determine the feature distance between the feature of this target clothing area and the feature of each candidate clothing area in the candidate clothing area set, and select a candidate clothing area with the smallest feature distance as the associated clothing area.

[0189] In one implementation, the area clustering module 504 is specifically configured to:

[0190] Calculate the distance between the image feature of this target clothing area and the image feature of the candidate clothing area as the first distance;

[0191] Calculate the distance between the time domain feature of this target clothing area and the time domain feature of the candidate clothing area as the second distance;

[0192] Based on the second distance, determine the weight coefficient of the first distance;

[0193] Based on the first distance and the weight coefficient, determine the feature distance between the feature of this target clothing area and the feature of the candidate clothing area;

[0194] Wherein, the larger the second distance, the larger the weight coefficient of the first distance, and the larger the feature distance.

[0195] In one implementation, the area clustering module 504 is specifically configured to:

[0196] Calculate the time interval between the moment represented by the time domain feature of this target clothing area and the moment represented by the time domain feature of the candidate clothing area;

[0197] If the time interval is less than or equal to the first preset interval, determine the second distance as the first preset interval;

[0198] If the time interval is greater than the first preset interval and less than the second preset interval, determine the second distance as the time interval;

[0199] If the time interval is greater than the first preset interval and greater than or equal to the second preset interval, determine the second distance as the second preset interval;

[0200] Wherein, the second preset interval is greater than the first preset interval.

[0201] In one implementation, the clothing recognition module is specifically configured to:

[0202] Weight and average the image features of each target clothing area in the target clothing areas of this type to obtain the target image features.

[0203] In one implementation, as Figure 6 shown, it further includes:

[0204] The area removal module 506 is used to detect the human body area in the key frame;

[0205] In response to detecting at least one target clothing area and at least one human body area, compare the target clothing area with each human body area, and based on the comparison result, determine whether the target clothing area is a clothing area to be removed, where the overlap ratio between the clothing area to be removed and the human body area reaches the first threshold and the pose similarity does not exceed the second threshold;

[0206] Remove the target clothing area to be removed in the key frame.

[0207] In one implementation, the area removal module is specifically configured to:

[0208] Calculate the intersection-over-union ratio of the target clothing area and each human body area;

[0209] Based on the calculation result, determine that the target clothing area is a clothing area to be removed.

[0210] In one implementation, in one implementation, the area removal module is specifically configured to:

[0211] In response to the maximum value among the intersection-over-union ratios being greater than or equal to the first threshold and the pose similarity corresponding to the maximum value not exceeding the second threshold, determine that the target clothing area is a clothing area to be removed.

[0212] In one implementation, in one implementation, the area removal module is specifically configured to:

[0213] In response to two or more intersection-over-union ratios among the intersection-over-union ratios being greater than or equal to the first threshold, determine that the clothing area is a clothing area to be removed.

[0214] In one embodiment, the human body area includes at least one key node of human body posture;

[0215] In one embodiment, the area removal module is specifically configured to:

[0216] Determine the human body area corresponding to the maximum value in each intersection over union as the target human body area;

[0217] Obtain the style category of the identified target clothing area;

[0218] Obtain at least one key node of human body posture included in the wearing part of the identified style category that is preset;

[0219] Count the number of the same key nodes of human body posture included in the wearing part of the identified style category and the target human body area;

[0220] In response to the counted number not exceeding the second threshold, determine the target clothing area as the clothing area to be removed.

[0221] In one embodiment, the image features include texture features and color features.

[0222] For the functions of the modules in each device provided in the embodiments of the present invention, reference may be made to the corresponding descriptions in the embodiments of the above clothing recognition method, which will not be elaborated here.

[0223] The embodiments of the present invention further provide an electronic device, as Figure 7 shown, including a processor 701, a communication interface 702, a memory 703, and a communication bus 704. Among them, the processor 701, the communication interface 702, and the memory 703 complete mutual communication through the communication bus 704,

[0224] The memory 703 is used to store a computer program;

[0225] The processor 701, when executing the program stored on the memory 703, implements the following steps:

[0226] Extract key frames at multiple different moments from the video to be recognized;

[0227] Perform clothing area detection on the key frames;

[0228] For the detected target clothing area, extract the features of the target clothing area, and the features of the target clothing area at least include image features;

[0229] Cluster each target clothing area based on the features of each target clothing area to obtain at least one category of target clothing areas;

[0230] For each type of target clothing region, the image features of each target clothing region in the type of target clothing region are fused into a target image feature, and based on the target image feature, clothing recognition is performed to obtain the recognition result of the type of target clothing region.

[0231] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0232] The communication interface is used for communication between the above terminal and other devices.

[0233] The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0234] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0235] In another embodiment provided by the present invention, a computer-readable storage medium is also provided. Instructions are stored in the computer-readable storage medium, and when it runs on a computer, it causes the computer to execute the clothing recognition method described in any one of the above embodiments.

[0236] In another embodiment provided by the present invention, a computer program product containing instructions is also provided. When it runs on a computer, it causes the computer to execute the clothing recognition method described in any one of the above embodiments.

[0237] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0238] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.

[0239] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.

[0240] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A clothing recognition method, characterized in that, Including: Extracting multiple key frames at different moments from the video to be recognized; Performing clothing area detection on the key frames; For the detected target clothing areas, extracting the features of the target clothing areas, where the features of the target clothing areas at least include image features; Based on the features of each of the target clothing areas, clustering each of the target clothing areas to obtain at least one class of target clothing areas, where the at least one class of target clothing areas at least includes a first class of target clothing areas of a first type of clothing and a second class of target clothing areas of a second type of clothing, and the first type of clothing is different from the second type of clothing; For each class of target clothing areas, fusing the image features of each of the target clothing areas in this class of target clothing areas into a target image feature, and based on the target image feature, performing clothing recognition to obtain the recognition result of this class of target clothing areas; Wherein, after performing clothing area detection on the key frames to obtain at least one target clothing area corresponding to the key frames, and before extracting the features of the detected target clothing areas, the method further includes: Performing human body area detection on the key frames; In response to detecting at least one of the target clothing areas and at least one of the human body areas, comparing the target clothing area with each of the human body areas, and based on the comparison result, determining whether the target clothing area is a clothing area to be removed, where the overlapping ratio of the clothing area to be removed and the human body area reaches a first threshold and the pose similarity does not exceed a second threshold; Removing the target clothing areas to be removed in the key frames.

2. The method according to claim 1, characterized in that, The features of the target clothing areas further include time domain features.

3. The method according to claim 2, characterized in that, The clustering of each of the target clothing areas based on the features of each of the target clothing areas to obtain at least one class of target clothing areas includes: For each of the target clothing areas, determining the associated clothing area corresponding to this target clothing area, and determining the associated information corresponding to this target clothing area, where the associated clothing area is from a set of candidate clothing areas, and the candidate clothing areas included in the set of candidate clothing areas are target clothing areas whose time domain features represent moments earlier than this target clothing area, the associated clothing area is the candidate clothing area in the set of candidate clothing areas with the most similar features to this target clothing area, and the associated information includes an identification combination formed by the identification of this target clothing area and the identification of the corresponding associated clothing area; Classifying the identification combinations corresponding to each of the target clothing areas to obtain at least one class of identification combinations, where each of the identification combinations in the same class can be sequentially spliced end to end to form an identification string, and in two spliced identification combinations, the identifications of the spliced parts are the same; For each class of identification combinations, taking each of the target clothing areas corresponding to this class of identification combinations as one class.

4. The method according to claim 3, wherein After classifying each combination of identifiers and treating each of the target clothing regions corresponding to the combination of identifiers as one category, before fusing the image features of each of the target clothing regions in each category of the target clothing regions into a target image feature and performing clothing recognition based on the target image feature to obtain the recognition result of each category of the target clothing regions, the following steps are also included: Using at least one category of the classified target clothing regions as at least one clustering object and performing re-clustering to obtain at least one category of target clothing regions corresponding to the re-clustering; The step of fusing the image features of each of the target clothing regions in each category of the target clothing regions into a target image feature and performing clothing recognition based on the target image feature to obtain the recognition result of each category of the target clothing regions includes: For each category of the target clothing regions corresponding to the re-clustering, fusing the image features of each of the target clothing regions in the category of the target clothing regions into a target image feature and performing clothing recognition based on the target image feature to obtain the recognition result of the category of the target clothing regions.

5. The method according to claim 3, wherein The step of fusing the image features of each of the target clothing regions in each category of the target clothing regions into a target image feature and performing clothing recognition based on the target image feature to obtain the recognition result of each category of the target clothing regions includes: For each category of the target clothing regions of at least one category of the classified target clothing regions, fusing the image features of each of the target clothing regions in the category of the target clothing regions into a target image feature and performing clothing recognition based on the target image feature to obtain the recognition result of the category of the target clothing regions; The method also includes: Using at least one category of the classified target clothing regions as at least one clustering object and performing re-clustering to obtain at least one category of target clothing regions corresponding to the re-clustering; For each category of the target clothing regions corresponding to the re-clustering, selecting at least one candidate style from the recognition results of each clustering object in the category of the target clothing regions as the recognition result of the category of the target clothing regions.

6. The method according to claim 3, wherein The step of determining the associated clothing region corresponding to each of the target clothing regions includes: For each of the target clothing regions, respectively determining the feature distance between the feature of the target clothing region and the feature of each of the candidate clothing regions in the set of candidate clothing regions, and selecting the candidate clothing region with the smallest feature distance as the associated clothing region.

7. The method according to claim 6, characterized in that, The step of respectively determining the feature distance between the feature of the target clothing region and the feature of each of the candidate clothing regions in the set of candidate clothing regions includes: Calculating the distance between the image feature of the target clothing region and the image feature of the candidate clothing region as the first distance; Calculating the distance between the time-domain feature of the target clothing region and the time-domain feature of the candidate clothing region as the second distance; Determine the weight coefficient of the first distance based on the second distance; Determine the feature distance between the features of the target clothing region and the features of the candidate clothing region based on the first distance and the weight coefficient; Wherein, the larger the second distance, the larger the weight coefficient of the first distance, and the larger the feature distance.

8. The method according to claim 7, wherein The calculating the distance between the temporal features of the target clothing region and the temporal features of the candidate clothing region as the second distance includes: Calculate the time interval between the moment represented by the temporal features of the target clothing region and the moment represented by the temporal features of the candidate clothing region; If the time interval is less than or equal to the first preset interval, determine that the second distance is the first preset interval; If the time interval is greater than the first preset interval and less than the second preset interval, determine that the second distance is the time interval; If the time interval is greater than the first preset interval and greater than or equal to the second preset interval, determine that the second distance is the second preset interval; Wherein, the second preset interval is greater than the first preset interval.

9. The method according to claim 1, wherein The fusing the image features of each of the target clothing regions in this type of target clothing region into one target image feature includes: Perform weighted averaging on the image features of each of the target clothing regions in this type of target clothing region to obtain the target image feature.

10. The method according to claim 1, wherein The comparing the target clothing region with each of the human body regions, and determining whether the target clothing region is a clothing region to be removed based on the comparison result includes: Calculate the intersection over union of the target clothing region and each of the human body regions; Based on the calculation result, determine that the target clothing region is a clothing region to be removed.

11. The method according to claim 10, wherein The determining that the target clothing region is a clothing region to be removed based on the calculation result includes: In response to the maximum value among the intersection over unions being greater than or equal to the first threshold and the pose similarity corresponding to the maximum value not exceeding the second threshold, determine that the target clothing region is a clothing region to be removed.

12. The method according to claim 10, wherein The key frame contains more than two human body regions. The determining that the target clothing region is a clothing region to be removed based on the calculation result includes: In response to more than two of the intersection over unions being greater than or equal to the first threshold, determine that the clothing region is a clothing region to be removed.

13. The method according to claim 10, wherein The human body region contains at least one human pose key node; The determining that the target clothing region is a clothing region to be removed based on the calculation result includes: Determine the human body region corresponding to the maximum value among the intersection over unions as the target human body region; Obtain the style category of the identified target clothing region; Obtain at least one human pose key node included in the wearing part of the identified style category that is preset; Count the number of the same human pose key nodes included in the wearing part of the identified style category and the target human body region; In response to the counted number not exceeding the second threshold, determine the target clothing region as a clothing region to be removed.

14. The method according to claim 1, wherein The image features include texture features and color features.

15. A clothing recognition device, characterized in that, Comprising: A key frame extraction module, configured to extract multiple key frames at different moments from the video to be recognized; A region detection module, configured to perform clothing region detection on the key frames; A feature extraction module, configured to extract features of the detected target clothing region, where the features of the target clothing region at least include image features; A region clustering module, configured to cluster each of the target clothing regions based on the features of each of the target clothing regions, so as to obtain at least one type of target clothing region, where the at least one type of target clothing region at least includes a first type of target clothing region of a first piece of clothing and a second type of target clothing region of a second piece of clothing, and the first piece of clothing is different from the second piece of clothing; A clothing recognition module, configured to, for each type of target clothing region, fuse the image features of each of the target clothing regions in the type of target clothing region into a target image feature, and perform clothing recognition based on the target image feature, so as to obtain a recognition result of the type of target clothing region; Wherein, the apparatus further includes: A region removal module, configured to perform human body region detection on the key frames after performing clothing region detection on the key frames to obtain at least one target clothing region corresponding to the key frames, and before extracting the features of the detected target clothing region; in response to detecting at least one of the target clothing regions and at least one of the human body regions, compare the target clothing region with each of the human body regions, and based on the comparison result, determine whether the target clothing region is a clothing region to be removed, where the overlapping ratio of the clothing region to be removed and the human body region reaches a first threshold and the pose similarity does not exceed a second threshold; remove the target clothing region to be removed in the key frames.

16. An electronic device, characterized in that, Including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used for storing a computer program; The processor is configured to implement the method steps described in any one of claims 1-14 when executing the program stored on the memory.

17. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1-14.

Citation Information

Patent Citations

  • Face identification method and apparatus

    CN104408404A

  • Face identification method, device and system

    CN105956518A

  • Information acquisition method and device, storage medium and electronic device

    CN111126179A

  • Facial information acquisition method and device

    CN112101197A

  • Costume searching method and device, electronic equipment and medium

    CN112905889A