Clothing recognition method, device and equipment and storage medium
A clustering method that extracts keyframes from videos and performs clothing region detection and feature extraction solves the problem of cross-temporal ambiguity in clothing recognition technology, improves the accuracy of clothing recognition, and meets the needs of practical applications.
Patent Information
- Application Number
- CN202111164023.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2041-09-30
AI Technical Summary
Existing clothing recognition technologies have low accuracy in video recognition, making it difficult to meet the needs of real-world applications. In particular, when faced with situations where clothing is obscured or partially exposed, ambiguity issues can easily arise across time, leading to the same garment being identified as different styles or different garments being identified as the same style.
By extracting keyframes from multiple different moments in the video to be identified, clothing area detection and feature extraction are performed. Clustering is then performed based on feature similarity and temporal consistency. Reasonable candidate styles are selected as the recognition results. By combining texture and color features, the influence of individual moments is reduced, and the recognition accuracy is improved.
It improves the temporal consistency of clothing recognition, reduces cross-time ambiguity, enhances clothing recognition accuracy, and meets the needs of practical application scenarios.
Smart Images

Figure CN113869247B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a clothing recognition method, device, equipment and storage medium. BACKGROUND
[0002] Clothing is an essential item in people's daily life. With the emergence of online shopping platforms, people often buy clothing on online shopping platforms, which greatly facilitates people's life. In related technologies, in order to further provide convenience for people, the clothing appearing in the video can be recognized by clothing recognition technology to provide information of related clothing. The existing clothing recognition technology is still in the development stage, and the recognition accuracy is still low, which is difficult to meet the recognition needs of actual application scenarios. For example, the maximum accuracy of most current clothing recognition technologies based on open source data sets is less than 70%, which is difficult to meet the recognition needs of actual application scenarios with an accuracy higher than 90%. Especially, the following recognition problems are prone to occur: due to the non-rigid nature of clothing, the difference in human posture can cause the clothing area to be blocked and not fully exposed, which may lead to the same clothes appearing at two different time points (e.g., t1 time point (e.g., time point showing the front of the clothing) and t2 time point (e.g., time point showing the back of the clothing)) being recognized as different clothing, or two different clothes appearing at two different time points (e.g., t3 time point (e.g., showing the back of a pure white T-shirt) and t4 time point (e.g., showing a T-shirt with a large area of pattern on the front, but the back is pure white)) being recognized as the same clothes, which is a serious problem in clothing recognition in videos. SUMMARY
[0003] The purpose of the embodiments of the present application is to provide a clothing recognition method, device, equipment and storage medium to improve the clothing recognition accuracy. The specific technical solutions are as follows:
[0004] In the first aspect of the present application, a clothing recognition method is first provided, comprising:
[0005] extracting a plurality of key frames at different time points from a to-be-recognized video;
[0006] detecting clothing regions of the key frames;
[0007] extracting at least one feature of the target clothing region from the detected target clothing region;
[0008] clustering the target clothing regions based on the at least one feature of each target clothing region to obtain at least one type of target clothing region;
[0009] For each type of target clothing area, clothing recognition is performed on each target clothing area in the type of target clothing area to obtain a recognition result of each target clothing area, the recognition result of the target clothing area including at least one candidate style, and from the recognition results of the target clothing areas in the type of target clothing area, at least one candidate style is selected as the recognition result of the type of target clothing area.
[0010] In a second aspect of the embodiment of the present application, a clothing recognition device is further provided, comprising:
[0011] a key frame extraction module configured to extract key frames at different time points from the video to be recognized;
[0012] a region detection module configured to perform clothing region detection on the key frames;
[0013] a feature extraction module configured to extract at least one feature of the target clothing area from the detected target clothing area;
[0014] a region clustering module configured to cluster the target clothing areas based on the at least one feature of the target clothing areas to obtain at least one type of target clothing area;
[0015] a clothing recognition module configured to, for each type of target clothing area, perform clothing recognition on each target clothing area in the type of target clothing area to obtain a recognition result of each target clothing area, the recognition result of the target clothing area including at least one candidate style, and from the recognition results of the target clothing areas in the type of target clothing area, at least one candidate style is selected as the recognition result of the type of target clothing area.
[0016] In another aspect of the embodiment of the present application, an electronic device is further provided, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus;
[0017] the memory is configured to store a computer program;
[0018] the processor is configured to execute the program stored on the memory to implement the steps of the clothing recognition method described above.
[0019] In another aspect of the embodiment of the present application, a computer readable storage medium is further provided, the computer readable storage medium storing instructions, when the instructions are run on a computer, causing the computer to execute the clothing recognition method described above.
[0020] In another aspect of the embodiment of the present application, a computer program product containing instructions is further provided, when the instructions are run on a computer, causing the computer to execute the clothing recognition method described above.
[0021] The clothing recognition method, device, equipment and storage medium provided by the embodiment of the application are used for clustering each target clothing region corresponding to a plurality of key frames at different moments in a to-be-recognized video based on at least one feature of each target clothing region, and the target clothing regions with similar features and high time domain consistency can be classified into a category. Then, clothing recognition is performed based on the target clothing regions in the category, and at least one candidate style is selected from the recognition results of the target clothing regions in the category as the recognition result of the target clothing regions in the category. In this way, the time domain consistency is considered, the clothing recognition is performed based on the target clothing regions with similar features and high time domain consistency in a category, the influence of the target clothing region at a single moment is reduced, the time domain consistency of the clothing recognition result in the video is improved, and the ambiguity problem across time, such as that the same clothing at different moments is recognized as different styles or that different styles at different moments are recognized as the same style, is reduced, thereby improving the clothing recognition accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced.
[0023] Figure 1 An exemplary system architecture diagram in the embodiment of the application.
[0024] Figure 2 An exemplary clothing recognition method flowchart in the embodiment of the application.
[0025] Figure 3 An exemplary clothing recognition method flowchart in the embodiment of the application.
[0026] Figure 4 An exemplary clothing recognition device structure schematic diagram in the embodiment of the application.
[0027] Figure 5 An exemplary clothing recognition device structure schematic diagram in the embodiment of the application.
[0028] Figure 6 An exemplary electronic device structure schematic diagram for implementing the clothing recognition method in the embodiment of the application. DETAILED DESCRIPTION
[0029] The technical solutions in the embodiments of the application will be described below with reference to the drawings in the embodiments of the application.
[0030] Clothing is an essential item in people's daily life. With the emergence of online shopping platforms, people often buy clothing on online shopping platforms, which greatly facilitates people's life. In related technologies, in order to further provide convenience for people, clothing appearing in a video can be identified through clothing identification technology to provide people with information about related clothing. The existing clothing identification technology is still in the development stage, and the identification accuracy is still low, which is difficult to meet the identification needs of actual application scenarios. For example, the maximum accuracy of most current clothing identification technologies based on open source data sets is less than 70%, which is difficult to meet the identification needs of actual application scenarios with an accuracy higher than 90%. Especially, the following identification problems are prone to occur: due to the fact that clothing is a non-rigid object, the difference in human posture can easily cause the clothing area to be blocked and not completely exposed, which can cause the same clothes appearing at two different time points (for example, t1 time point (for example, time point of showing the front of the clothing) and t2 time point (for example, time point of showing the back of the clothing)) to be identified as different clothing (that is, one-to-many), or two different clothes appearing at two different time points (for example, t3 time point (for example, time point of showing the back of a pure white T-shirt) and t4 time point (for example, time point of showing a T-shirt with a large area of pattern on the front, but the back of the T-shirt is pure white)) to be identified as the same clothes (that is, many-to-one), which is a serious challenge in clothing identification in videos.
[0031] In view of the cross-time style ambiguity problem existing in clothing identification in a video, an embodiment of the present application provides a clothing identification method, which combines clustering technology to effectively reduce one-to-many and many-to-one mapping of query clothing and styles in a database, improve the time domain consistency of clothing identification results in a video, improve the rationality of clothing same style or similar style identification results in a video level, and thus improve user experience. The clothing identification method provided by the embodiment of the present application can be executed by a server or a user terminal. Taking execution of the server as an example, Figure 1 is a system architecture diagram in an embodiment of the present application. The system includes a server and a user terminal. In actual application, the server can execute the clothing identification method provided by the embodiment of the present application, and apply the result to a clothing identification or clothing recommendation processing scenario, and send the processing result to the user terminal, so that the user of the user terminal can understand the information of the clothing. The scheme provided by the embodiment of the present application is described in more detail below.
[0032] Figure 2 is a flowchart of an exemplary clothing identification method provided by an embodiment of the present application. As shown in Figure 2 , the clothing identification method provided by the embodiment can at least include the following steps:
[0033] Step 201, extracting a plurality of key frames at different time points from a to-be-identified video.
[0034] Step 202, detecting a clothing region of the key frame.
[0035] Step 203, extracting at least one feature of the detected target clothing region.
[0036] Step 204, clustering the target clothing regions based on the at least one feature of each target clothing region to obtain at least one type of target clothing region.
[0037] Step 205, for each type of target clothing region, performing clothing recognition on each target clothing region in the type of target clothing region to obtain a recognition result of each target clothing region, the recognition result of the target clothing region including at least one candidate style, and selecting at least one candidate style from the recognition results of the target clothing regions in the type of target clothing region as the recognition result of the type of target clothing region.
[0038] The to-be-recognized video is a video to be recognized. After extracting the key frames from the to-be-recognized video, the clothing region detection is performed on each key frame, and the detected region for clothing recognition is taken as a target clothing region. The at least one feature extracted from the target clothing region can be an image feature of the key frame.
[0039] Clustering is a process of dividing a collection of physical or abstract objects into multiple classes composed of similar objects. Therefore, clustering the target clothing regions based on the at least one feature of each target clothing region can divide the target clothing regions with similar features into a class. In the time domain, since the multiple key frames come from multiple time points, the target clothing regions corresponding to the multiple key frames also come from multiple time points, and accordingly, a type of target clothing region also comes from different time points, and the time points at which the target clothing regions with similar features are located are mostly close, that is, the time domain consistency is high.
[0040] After clustering, one type of target clothing region can be obtained, or multiple types of target clothing regions can be obtained. For example, if there is only one clothing in the to-be-recognized video, the target clothing regions of the key frames at different time points correspond to the same clothing, and after clustering, one type of target clothing region corresponding to the clothing can be obtained, or multiple types of target clothing regions corresponding to the clothing can be obtained. If there are multiple clothes in the to-be-recognized video, after clustering, multiple types of target clothing regions are obtained, for example, for the same clothing, one type of target clothing region corresponding to the clothing can be obtained, or multiple types of target clothing regions corresponding to the clothing can be obtained. Assuming that there are two clothes in the to-be-recognized video, after clustering, three types of target clothing regions are obtained, two types of target clothing regions are for one clothing, and one type of target clothing region is for another clothing.
[0041] Based on this, in the embodiment, for each target clothing region corresponding to multiple key frames at different time moments in the to-be-identified video, clustering is performed on each target clothing region based on at least one feature of each target clothing region, which can divide target clothing regions with similar features and high time consistency into a category, and then clothing identification is performed based on the target clothing regions in the category, and at least one candidate style is selected from the identification results of each target clothing region in the category as the identification result of the target clothing regions in the category. In this way, the time consistency is considered, clothing identification is performed based on target clothing regions with similar features and high time consistency, the influence of target clothing regions at a single time moment is reduced, the time consistency of the clothing identification result in the video is improved, the ambiguity problem across time, such as that the same clothing at different time moments is identified as different styles or different styles at different time moments are identified as the same style, is reduced, and therefore the clothing identification accuracy is improved.
[0042] In an implementation manner, the at least one feature can include a texture feature and can also include a color feature. In the embodiment, the color feature is added on the basis of the texture feature. Although the texture feature contains a certain amount of color features, the color features are not obvious. Here, the color feature is listed as a display feature together with the texture feature, which can increase the proportion of the color feature, eliminate the error caused by the color feature, and improve the identification accuracy.
[0043] In an implementation manner, in step 205, the specific implementation manner of selecting at least one candidate style from the identification results of each target clothing region in the category can include the following steps: forming a candidate style set based on the candidate styles in the identification results of each target clothing region in the category; calculating the score of each candidate style in the candidate style set; and selecting at least one candidate style from the candidate style set based on the scores of the candidate styles.
[0044] For example, a target clothing region L1 includes a target clothing region a, a target clothing region b, and a target clothing region c. The identification result of the target clothing region a includes a candidate style h1, a candidate style h2, a candidate style h3, and a candidate style h4. The identification result of the target clothing region b includes the candidate style h1, the candidate style h2, the candidate style h3, and a candidate style h5. The identification result of the target clothing region c includes the candidate style h1, the candidate style h2, the candidate style h4, and a candidate style h6. Then, the candidate styles in the identification result of the target clothing region a, the candidate styles in the identification result of the target clothing region b, and the candidate styles in the identification result of the target clothing region c form a candidate style set P1, which includes the candidate style h1, the candidate style h2, the candidate style h3, the candidate style h4, the candidate style h5, and the candidate style h6. For each candidate style in the candidate style set P1, a score of the candidate style is calculated, and at least one candidate style is selected from the candidate style set P1 based on the scores of the candidate styles. For example, the candidate style h1 and the candidate style h2 are selected.
[0045] The score of the candidate style can represent the rationality of the candidate style as the clothing identification result. In this embodiment, the score of the candidate style is calculated to provide a selection basis for the selection of the candidate style, so that a reasonable candidate style can be accurately selected, and the accuracy of the clothing identification is further improved.
[0046] If the score of the candidate style is higher, the rationality is higher. Therefore, when at least one candidate style is selected, at least one candidate style with the highest score can be selected. In this way, the accuracy of the clothing identification is further improved.
[0047] In an implementation, the score of each candidate style in the candidate style set is calculated, and the specific implementation can include: for each candidate style in the candidate style set, the score of the candidate style is calculated based on the similarity of each feature of the candidate style and the same feature of the corresponding target clothing region, and / or the first hit number, and / or the second hit number; the first hit number is the total hit number of the candidate style in the identification result of all target clothing regions forming the candidate style set; and the second hit number is the total hit number of the candidate style in the identification result of all target clothing regions of the video to be identified.
[0048] The similarity of the features is an important factor for measuring the rationality of the candidate style. The more similar the features are, the more likely it is the same style. Therefore, in this embodiment, the similarity of the features is used as the calculation basis of the score of the candidate style.
[0049] The hit number is the initial number, i.e., the number of occurrences.
[0050] Continuing with the above example candidate style set P1, the candidate style set P1 includes candidate style h1, candidate style h2, candidate style h3, candidate style h4, candidate style h5, and candidate style h6.
[0051] The first hit number is the total hit number of the candidate style in the identification result of the target clothing region a, the identification result of the target clothing region b, and the identification result of the target clothing region c forming the candidate style set P1. Taking the candidate style h5 as an example, the candidate style h5 hits 0 times in the identification result of the target clothing region a, the candidate style h1 hits 1 time in the identification result of the target clothing region b, and the candidate style h1 hits 0 times in the identification result of the target clothing region c. At this time, the total hit number is obtained as 1 time, that is, the first hit number of the candidate style h5 is 1.
[0052] The first hit number can reflect the hit frequency of the candidate style in the identification result of a type of target clothing region. In a type of target clothing region, since the characteristics of each target clothing region are similar and the time domain consistency is high, it is more likely to be the same style. If the hit frequency of the candidate style in the identification result of a type of target clothing region is higher, that is, more target clothing region identification results hit the same candidate style, the candidate style is more likely to be accurate, and it is more reasonable to take the candidate style as the identification result. Therefore, the first hit number is an important factor for measuring the rationality of the candidate style and can be used as the basis for calculating the score of the candidate style.
[0053] Suppose that after clustering all target clothing regions corresponding to the video to be recognized, two types of target clothing regions are obtained. In addition to the above-mentioned first type of target clothing region L1, another type of target clothing region L2 is included. The other type of target clothing region L2 includes target clothing region d, target clothing region e, and target clothing region f. The identification result of the target clothing region d includes candidate style h7, candidate style h8, candidate style h9, and candidate style h5. The identification result of the target clothing region e includes candidate style h7, candidate style h8, candidate style h9, and candidate style h11. The identification result of the target clothing region f includes candidate style h7, candidate style h8, candidate style h10, and candidate style h12.
[0054] The second hit number is the total hit number of the candidate style in the identification result of each target clothing area of one type of target clothing area L1 and the identification result of each target clothing area of another type of target clothing area L2. Taking the candidate style h5 as an example, the candidate style h5 hits 0 times in the identification result of the target clothing area a, the candidate style h1 hits 1 time in the identification result of the target clothing area b, the candidate style h1 hits 0 times in the identification result of the target clothing area c, the candidate style h5 hits 1 time in the identification result of the target clothing area d, the candidate style h5 hits 0 times in the identification result of the target clothing area e, and the candidate style h5 hits 0 times in the identification result of the target clothing area f. At this time, the total hit number is 2, that is, the second hit number of the candidate style h5 is 2.
[0055] The second hit number can reflect the hit frequency of the candidate style in the identification result of each target clothing area of the entire to-be-identified video. From the perspective of the entire to-be-identified video, the greater the difference between the identification results of different types of target clothing areas is, the better. Because the features of different types of target clothing areas are not similar and the time domain consistency is poor, it is more likely that they are not the same style. Therefore, we expect that the candidate style is hit only by one type of target clothing area. If the candidate style is hit by the identification result of at least one type of target clothing area, the distinguishability of different types of styles is poor, and it is unreasonable to take the candidate style as the identification result. Taking an extreme case as an example, if there are actually multiple different styles in the entire to-be-identified video, and a candidate style appears in the identification result of each type of target clothing area of the entire to-be-identified video, then it has no distinguishability and is not suitable for being output as the identification result. Therefore, the second hit number is also an important factor for measuring the rationality of the candidate style and can be used as a basis for calculating the score of the candidate style.
[0056] When calculating the score of the candidate style h5, the score of the candidate style h5 is calculated based on the similarity between the texture feature of the candidate style h5 and the texture feature of the corresponding target clothing area, the similarity between the color feature of the candidate style h5 and the color feature of the corresponding target clothing area, the first hit number, and the second hit number.
[0057] In this embodiment, when calculating the score of the candidate style, in addition to considering the important factor of feature similarity, the hit frequency of the candidate style in the identification result of one type of target clothing area and the hit frequency of the candidate style in the identification result of each target clothing area of the entire to-be-identified video are also considered. The considered factors are very comprehensive, and a more accurate score of the candidate style can be calculated to select a more reasonable candidate style.
[0058] In an embodiment, the score of the candidate style is calculated based on the similarity of each feature of the candidate style to the corresponding feature of the target clothing region, and / or the first hit number, and / or the second hit number. The specific implementation can include: weighting and summing the similarity, the first hit number, and the ratio of the first hit number to the second hit number to obtain the score of the candidate style. In this embodiment, by weighting and summing the similarity, the first hit number, and the ratio of the first hit number to the second hit number, the score of the candidate style can be more accurately calculated.
[0059] In addition, because we expect that the higher the score of the candidate style, the higher the rationality of the candidate style as the clothing result, as described above, the higher the first hit number, the higher the rationality, and the higher the second hit number, the lower the rationality, therefore, the ratio of the first hit number to the second hit number is used as an item participating in the weighted sum, based on which, the higher the first hit number, the higher the score of the candidate style, and the higher the second hit number, the lower the score of the candidate style, which can accurately evaluate the rationality of the candidate style.
[0060] Alternatively, a preset constant can be directly used to replace the first hit number, and the ratio of the preset constant to the second hit number is used as an item participating in the weighted sum.
[0061] In an embodiment, the weight of the similarity is greater than the weight of the first hit number, and / or the weight of the first hit number is greater than the weight of the ratio.
[0062] In actual application, because the feature is mainly used for clothing identification, the weight of the feature is the largest, the weight of the first hit number is the second, and the weight of the ratio of the first hit number to the second hit number is the smallest, so that the output of the unreasonable candidate style can be maximally reduced.
[0063] It should be noted that the above only illustrates one way of calculating the score of the candidate style, and the score of the candidate style can also be calculated by other ways. For example, weighting and summing the similarity corresponding to at least one feature of the candidate style. Or, weighting and summing the similarity and the first hit number corresponding to at least one feature of the candidate style. Or, weighting and summing the similarity and the second hit number corresponding to at least one feature of the candidate style. Or, weighting and summing the first hit number and the second hit number, and the like.
[0064] It should be noted that the above only exemplarily illustrates an implementation manner of selecting at least one candidate style, and at least one candidate style can also be selected in other manners, for example, the similarities or the first hit times corresponding to the respective candidate styles are directly sorted, and at least one candidate style is selected based on a sorting result. For example, at least one candidate style with the highest similarity or the most first hit times is selected.
[0065] In an implementation manner, after the costume region detection on the key frame is performed to obtain at least one target costume region corresponding to the key frame, before the at least one feature of the target costume region is extracted, the costume recognition method can further include:
[0066] The human body region detection is performed on the key frame; in response to the detection of at least one costume region and at least one human body region, the costume region is compared with each human body region, and based on a comparison result, it is determined whether the costume region is a costume region that needs to be removed, the overlap ratio of the costume region and the human body region reaches a first threshold value and the posture similarity does not exceed a second threshold value; and the costume region that needs to be removed in the key frame is removed.
[0067] If the overlap ratio of the target costume region and the human body region reaches the first threshold value, it can be considered that the costume included in the target costume region is possibly worn on the human body included in the human body region. Further, if the posture similarity of the target costume region and the human body region does not exceed the second threshold value, it can be considered that the posture of the costume included in the target costume region is possibly different from the posture of the human body included in the human body region. At this time, the possibility that the costume included in the target costume region is worn on the human body included in the human body region is excluded, and it is considered that the target costume region is not worn on the human body included in the human body region. This situation can be caused by factors such as the target costume region being blocked by the human body passing by, which can affect the costume recognition result, and the costume region that needs to be removed is determined. The costume region that needs to be removed in the image to be processed is removed. The specific values of the first threshold value and the second threshold value can be set according to actual conditions, which are not limited here.
[0068] Based on this, in the embodiment, after the target clothing region of the to-be-processed image is detected, the target clothing region is compared with each human body region. Based on this, it is determined whether the target clothing region is a clothing region that needs to be removed, which has an overlap ratio with the human body region reaching a first threshold and a pose similarity not exceeding a second threshold. In this way, in combination with at least one human body region in the image, that is, in combination with the context information, some target clothing regions that are affected by the recognition due to factors such as occlusion by a passing human body are removed. The target clothing regions in the to-be-processed image that do not need to be removed are determined as effective target clothing regions. In this way, the low-quality target clothing regions that affect the recognition are removed. Overall, the base number of target clothing regions participating in subsequent clothing recognition is reduced, which lays a good foundation for subsequent clothing recognition, thereby improving the accuracy of clothing recognition.
[0069] The pose can be a wearing state or a non-wearing state (such as a flat state). For clothing in the wearing state, it is possible to be occluded by other passing human bodies. For clothing in the non-wearing state, it is also possible to be occluded by other passing human bodies. In this case, the target clothing region will affect the clothing recognition result and needs to be removed.
[0070] The above comparison of the target clothing region with each human body region determines whether the target clothing region is a clothing region that needs to be removed based on the comparison result. The specific implementation manner can include: calculating the intersection-over-union of the target clothing region and each human body region; and determining the target clothing region as a clothing region that needs to be removed based on the calculation result.
[0071] In actual application, the human body region of the to-be-processed image can be detected in advance. The detected clothing region and human body region are framed by a rectangular box.
[0072] The intersection-over-union is the ratio of the intersection to the union of two boxes. The intersection-over-union can measure the overlap ratio of two boxes. Here, the intersection-over-union of the target clothing region and the human body region can reflect the overlap ratio of the target clothing region and the human body region, that is, the relative size of the overlap. From this, it can be reflected whether the clothing contained in the target clothing region can be worn on the human body contained in the human body region. Based on this, some rationality judgment can be performed to determine whether the target clothing region is a clothing region that needs to be removed, thereby removing some detected unreasonable target clothing regions.
[0073] In an implementation manner, the above determination of the target clothing region as a clothing region that needs to be removed based on the calculation result can include: in response to a maximum value in the intersection-over-union being greater than or equal to a first threshold and a pose similarity corresponding to the maximum value not exceeding a second threshold, determining the target clothing region as a clothing region that needs to be removed.
[0074] In the embodiment, if the maximum value of each of the above intersection-over-union ratios is greater than or equal to the first threshold value, and the pose similarity corresponding to the maximum value does not exceed the second threshold value, it is indicated that at least one human body causes occlusion to the target clothing region, and thus the target clothing region can be removed to improve the accuracy of subsequent clothing recognition.
[0075] For example, the first threshold value can be less than or equal to 0.5. It can be understood that the first threshold value can be set to other values according to actual conditions.
[0076] In an implementation, the image to be processed contains two or more human body regions, and based on the calculation result, the target clothing region is determined as a clothing region that needs to be removed, and a specific implementation manner can include: in response to two or more intersection-over-union ratios being greater than or equal to the first threshold value, the target clothing region is determined as a clothing region that needs to be removed.
[0077] In actual application, for clothing in a wearing state, if two or more intersection-over-union ratios are greater than or equal to the first threshold value, it is indicated that the clothing contained in the detected target clothing region can be worn on multiple human bodies, however, actually, a piece of clothing can only be worn on one human body, which is inconsistent with the actual situation, indicating that the detected target clothing region is problematic, and it is very likely that the target clothing region is actually an overlapping region caused by occlusion between multiple human bodies. For clothing in a non-wearing state, if two or more intersection-over-union ratios are greater than or equal to the first threshold value, it is indicated that at least two human bodies simultaneously occlude the target clothing region, and such a target clothing region is ambiguous, and thus the target clothing region can be removed to improve the accuracy of subsequent clothing recognition.
[0078] Further, in response to two or more intersection-over-union ratios being greater than or equal to the first threshold value and at least one of the two or more intersection-over-union ratios corresponding to a pose similarity not exceeding the second threshold value, the target clothing region is determined as a clothing region that needs to be removed. For clothing in a wearing state, the pose of the clothing contained in the target clothing region is different from the pose of the human body contained in at least one of the multiple human body regions that can be worn, and thus at least one human body causes occlusion. In the embodiment, the pose similarity is combined as a factor to further improve the accuracy of the determined clothing region that needs to be removed.
[0079] Further, in response to two or more of the intersection-over-union ratios being greater than or equal to the first threshold and the corresponding pose similarities of the two or more intersection-over-union ratios all not exceeding the second threshold, the target clothing region is determined as a clothing region that needs to be removed. For the clothing in a non-wearing state, if at least two people are simultaneously shielding, since the clothing is not worn on the human body, the pose of the clothing contained in the target clothing region is different from the pose of the human body contained in each human body region. In combination with the factor of the pose similarity, the accuracy of determining the clothing region that needs to be removed is further improved.
[0080] In an implementation manner, the human body region contains at least one human body pose key node; based on the calculation result, the target clothing region is determined as a clothing region that needs to be removed, and a specific implementation manner can include: determining a human body region corresponding to a maximum value in each intersection-over-union ratio as a target human body region; obtaining a style category of the recognized target clothing region; obtaining at least one human body pose key node contained in a wearing part of the pre-set recognized style category; counting a number of the same human body pose key nodes contained in the wearing part of the recognized style category and the target human body region; and in response to the counted number not exceeding the second threshold, determining the target clothing region as the clothing region that needs to be removed.
[0081] In actual application, the detected human body region is a rectangular frame containing at least one human body pose key node. The human body pose key node can be left eye, right eye, left ear, right ear, nose, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left foot, right foot, etc. The human body region detection model can be detected by a pre-trained human body region detection model, and reference can be made to related technologies, which will not be described herein. Different parts of the human body can be formed by selecting different human body pose key nodes. For example, the upper body part of the human body can be formed by the neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip and right hip.
[0082] The human body region corresponding to the maximum value in each intersection-over-union ratio is determined as the target human body region, that is, it is considered that the clothing contained in the target clothing region is worn on the human body contained in the target human body region.
[0083] In addition, the style category of the target clothing region can also be identified in advance. The style category can include short-sleeved shirt, long-sleeved shirt, vest, short pants, long pants, half skirt, one-piece skirt, etc. The style category of the target clothing region can be identified by a pre-trained category classifier.
[0084] In this embodiment, the number of the same human posture key nodes contained in the target human region and the recognized style category of the wearing part does not exceed the second threshold, which indicates that the wearing part of the clothing contained in the clothing region is inconsistent with the part of the human body formed by the human posture key nodes contained in the target human region, i.e., there is a conflict. For example, the human posture key nodes contained in the target human region belong to the upper body part of the human body, and the recognized style category is trousers, which are actually worn on the lower body of the human body. Therefore, there is a conflict between them, and the recognition result of the clothing region may be ambiguous and needs to be removed to improve the accuracy of subsequent clothing recognition.
[0085] In an example embodiment, the target clothing region and human region of the to-be-recognized image are detected, and the specific implementation manner can include: detecting the target clothing region of the to-be-recognized image; in response to detecting at least one target clothing region, performing posture category recognition on each target clothing region, the posture category including a wearing state and an unwearing state; and in response to the posture category of the target clothing region being the wearing state, detecting the human region of the to-be-recognized image. In this way, the clothing region that needs to be removed can be determined for the target clothing region in the wearing state.
[0086] Correspondingly, based on the calculation result, the target clothing region is determined as the clothing region that needs to be removed, and the specific implementation manner can further include: in response to the maximum value in each intersection-over-union being less than the first threshold, determining the target clothing region as the clothing region that needs to be removed. In actual application, if the maximum value in each intersection-over-union is less than the first threshold, it indicates that each intersection-over-union is relatively small, and then the clothing of the detected target clothing region can not be worn on any human body, which is inconsistent with the previously recognized posture category in the wearing state, indicating that the detected target clothing region is problematic and is likely to be an invalid region without any clothing. Therefore, the target clothing region can be filtered out to improve the accuracy of subsequent clothing recognition.
[0087] Then, when extracting at least one feature of the target clothing region, specifically, at least one feature of each target clothing region that is not removed is extracted. It should be noted that if all target clothing regions are finally removed, the subsequent clothing recognition steps can not be performed. A prompt information that no clothing is recognized can be issued.
[0088] In one implementation manner, based on at least one feature of each target clothing region, each target clothing region is clustered to obtain at least one type of target clothing region, and the specific implementation manner can include the following two kinds:
[0089] The first clustering manner is: performing one clustering on each target clothing region based on at least one feature of each target clothing region to obtain at least one type of target clothing region corresponding to the one clustering.
[0090] The second clustering manner is: performing one clustering on each target clothing region based on at least one feature of each target clothing region to obtain at least one type of target clothing region corresponding to the one clustering; and performing again clustering on the at least one type of target clothing region corresponding to the one clustering as at least one clustering object to obtain at least one type of target clothing region corresponding to the again clustering.
[0091] Compared with the first clustering manner, the second clustering manner is two times of clustering, which is called coarse-grained clustering, while the first clustering manner is only one time of clustering, which is called fine-grained clustering.
[0092] After the first clustering (i.e., fine-grained clustering), there can be more than two types of target clothing regions belonging to the target clothing region of the same clothing, so here the second clustering (i.e., coarse-grained clustering) is performed, that is, hierarchical clustering is performed, so that each target clothing region with the most similar features and consistent time domain can be further classified into a type, which not only makes the classification more accurate, but also makes the information of the fused target image features more rich, thereby further improving the accuracy of clothing recognition.
[0093] In the first clustering manner, for each type of target clothing region, clothing recognition is performed on each target clothing region in the type of target clothing region to obtain an identification result of each target clothing region, the identification result of the target clothing region includes at least one candidate style, and at least one candidate style is selected from the identification results of the target clothing regions in the type of target clothing region as the identification result of the type of target clothing region. Specifically, for each type of target clothing region in the at least one type of target clothing region corresponding to the one clustering, clothing recognition is performed on each target clothing region in the type of target clothing region to obtain an identification result of each target clothing region, the identification result of the target clothing region includes at least one candidate style, and at least one candidate style is selected from the identification results of the target clothing regions in the type of target clothing region as the identification result of the type of target clothing region.
[0094] In the second clustering manner, for each type of target clothing region, clothing recognition is performed on each target clothing region in the type of target clothing region to obtain an identification result of each target clothing region, the identification result of the target clothing region includes at least one candidate style, and at least one candidate style is selected from the identification results of the target clothing regions in the type of target clothing region as the identification result of the type of target clothing region. Specifically, for each type of target clothing region in the at least one type of target clothing region corresponding to the one clustering, clothing recognition is performed on each target clothing region in the type of target clothing region to obtain an identification result of each target clothing region, the identification result of the target clothing region includes at least one candidate style, and at least one candidate style is selected from the identification results of the target clothing regions in the type of target clothing region as the identification result of the type of target clothing region.
[0095] For each target clothing area in the at least one type of target clothing area corresponding to the second clustering, clothing recognition is performed on each target clothing area in the type of target clothing area to obtain an identification result of each target clothing area, the identification result of the target clothing area including at least one candidate style, and at least one candidate style is selected from the identification results of the target clothing areas in the type of target clothing area as the identification result of the type of target clothing area. In this way, after two times of clustering, clothing recognition is performed, and the recognition speed is fast.
[0096] Alternatively, for each target clothing area in the at least one type of target clothing area corresponding to the first clustering, clothing recognition is performed on each target clothing area in the type of target clothing area to obtain an identification result of each target clothing area, the identification result of the target clothing area including at least one candidate style, and at least one candidate style is selected from the identification results of the target clothing areas in the type of target clothing area as the identification result of the type of target clothing area; and for each target clothing area in the at least one type of target clothing area corresponding to the second clustering, at least one candidate style is selected from the identification results of the clustering objects in the type of target clothing area as the identification result of the type of target clothing area.
[0097] In this way, after the first clustering, clothing recognition is performed based on the result of the fine-grained clustering. Since the number of categories is large at this time, the categories are divided in a relatively fine manner, the interference target clothing areas are less for a type of target clothing area, and the identification result is more accurate. After the second clustering, the identification results of the clustering objects in the type of target clothing area are fused, and at least one candidate style is selected therefrom as the identification result of the type of target clothing area. The identification result is also more accurate.
[0098] Specifically, selecting at least one candidate style from the identification results of the clustering objects in the type of target clothing area can include: forming a candidate style set based on the candidate styles in the identification results of the clustering objects in the type of target clothing area; calculating a score of each candidate style in the candidate style set; and selecting at least one candidate style from the candidate style set based on the scores of the candidate styles. The specific implementation of calculating the score of each candidate style in the candidate style set and selecting at least one candidate style from the candidate style set based on the scores of the candidate styles can refer to the related embodiments described above, and will not be described here.
[0099] In an implementation, each target clothing region is clustered once based on at least one feature of each target clothing region to obtain at least one type of target clothing region corresponding to the once clustering, and a specific implementation can include: for each target clothing region, determining an associated clothing region corresponding to the target clothing region, and determining associated information corresponding to the target clothing region, wherein the associated clothing region is from a candidate clothing region set, the candidate clothing region set contains a target clothing region earlier in time than the target clothing region represented by the time domain feature, the associated clothing region is the candidate clothing region in the candidate clothing region set that is most similar to the target clothing region in feature, and the associated information contains an identification combination formed by an identification of the target clothing region and an identification of the corresponding associated clothing region; classifying the identification combinations corresponding to each target clothing region to obtain at least one type of identification combination, wherein each identification combination of the same type can be sequentially concatenated to form an identification string, and the identifications of the concatenated parts in the two identification combinations to be concatenated are the same; for each type of identification combination, each target clothing region corresponding to the type of identification combination is a type.
[0100] The at least one feature includes a time domain feature and an image feature such as a color feature and a texture feature. The image feature of the target clothing region refers to a feature in the spatial domain of the key frame. The spatial domain here refers to the space where the pixels of the key frame are located.
[0101] It should be noted that if the target clothing region does not have a corresponding candidate clothing region, it is not processed. For example, the earliest target clothing region in time does not have a corresponding candidate clothing region, and can not be processed.
[0102] For example, key frame 1, key frame 2 and key frame 3 are sequentially extracted in time from the video to be recognized, wherein two target clothing regions are detected in key frame 1, the identifications of the two target clothing regions are 1 and 2 respectively, two target clothing regions are detected in key frame 2, the identifications of the two target clothing regions are 3 and 4 respectively, and two target clothing regions are detected in key frame 3, the identifications of the two target clothing regions are 5 and 6 respectively.
[0103] Taking the target clothing region 4 in the key frame 2 as an example, the associated clothing region of the target clothing region 2 is from a candidate clothing region set formed by two target clothing regions in the key frame 1 in time sequence in front, assuming that the feature of the target clothing region 2 in the key frame 1 is the most similar to the feature of the target clothing region 4, then the associated clothing region of the target clothing region 4 is the target clothing region 2. The association information of the target clothing region 4 and the target clothing region 2 is established, the association information includes an identification combination formed by the identification of the target clothing region 4 and the identification of the target clothing region 2, for example: 4, 2. In this way, the same processing is performed on the target clothing region 3, the target clothing region 5 and the target clothing region 6. Since the target clothing region 1 and the target clothing region 2 are the earliest in time sequence, no processing is performed.
[0104] Assuming that the identification combination: 3, 1, the identification combination: 5, 4 and the identification combination: 6, 3 are also obtained, based on this, the identification combination: 6, 3 and the identification combination: 3, 1 can be classified into a category, because the two identification combinations can be spliced to form an identification string 6, 3, 1, and the spliced part of the two identification combinations has the identification "3", then the target clothing region 6, the target clothing region 3 and the target clothing region 1 corresponding to this category of identification combinations are classified into a category; and the identification combination: 5, 4 and the identification combination: 4, 2 can be classified into a category, because the two identification combinations can be spliced to form an identification string 5, 4, 2, and the spliced part of the two identification combinations has the identification "4", then the target clothing region 5, the target clothing region 4 and the target clothing region 2 corresponding to this category of identification combinations are classified into a category.
[0105] For the target clothing region, the time point represented by the time domain feature of the associated clothing region is earlier than the time point represented by the time domain feature of the target clothing region, that is, the time sequence of the associated clothing region is in front of the target clothing region. In this embodiment, a target clothing region is associated with the most similar target clothing region in time sequence in front, and the association information is established through the identification combination, that is, the forward mapping relationship from the target clothing region to the associated clothing region is established, then the target clothing regions with the most similar features and consistent time domains are associated through the identification string formed by the head and tail of the identification combination, and the target clothing regions associated are classified into a category, so that the target clothing regions with the most similar features and consistent time domains are classified into a category.
[0106] The above is only an example of a way of clustering the target clothing regions, and it can be understood that the target clothing regions can be clustered in other ways, for example, the target clothing regions can be clustered through a KMeans clustering algorithm, etc.
[0107] In an embodiment, for each target clothing region, a corresponding associated clothing region is determined, and the implementation can include: for each target clothing region, respectively determining the feature distance between the feature of the target clothing region and the feature of each candidate clothing region in the candidate clothing region set, and selecting the candidate clothing region with the smallest feature distance as the associated clothing region. The feature distance is used to measure the similarity between two objects, and the smaller the feature distance, the more similar the two objects. Therefore, in this embodiment, by determining the feature distance between the feature of the candidate clothing region and the feature of the target clothing region, the most similar candidate clothing region is accurately selected as the associated clothing region.
[0108] In an embodiment, the feature distance between the feature of the target clothing region and the feature of each candidate clothing region in the candidate clothing region set is respectively determined, and the implementation can include: calculating the distance between the image feature of the target clothing region and the image feature of the candidate clothing region as a first distance; calculating the distance between the time domain feature of the target clothing region and the time domain feature of the candidate clothing region as a second distance; determining a weight coefficient of the first distance based on the second distance; and determining the feature distance between the feature of the target clothing region and the feature of the candidate clothing region based on the first distance and the weight coefficient, wherein the greater the second distance, the greater the weight coefficient of the first distance, and the greater the feature distance.
[0109] The first distance, i.e. the image feature distance, is used to measure the similarity of the image features of two clothing regions, and the smaller the image feature distance, the more similar the two clothing regions, and the greater the image feature distance, the less similar the two clothing regions. The second distance, i.e. the time domain interval distance, is used to measure the time domain consistency of two clothing regions, and the smaller the time domain interval distance, the higher the time domain consistency, and the greater the time domain interval distance, the lower the time domain consistency.
[0110] In practical applications, although the image feature distance between the target clothing region and the candidate clothing region is very small, i.e. the two clothing regions are very similar, the time domain interval distance between the two clothing regions is very large, which is likely to come from two different clothing, therefore, the feature distance caused by the image feature distance and the time domain interval distance is relatively large.
[0111] For example, a T-shirt with a white back at a certain time and a T-shirt with a large area of pattern on the front at another time, but the back of the T-shirt is white, which is two different clothing at different times, and from the perspective of image features, the back of the T-shirt is white, but because the time domain interval is far apart, the feature distance is considered to be large in combination.
[0112] Therefore, in the embodiment, the weight coefficient of the first distance is determined based on the second distance, and the feature distance is obtained based on the first distance and the weight coefficient. The greater the second distance is, the greater the weight coefficient acting on the first distance is, so that the feature distance obtained by the first distance and the weight coefficient is greater. In this way, the feature distance can be accurately obtained.
[0113] For example, the feature distance can be calculated according to the following formula:
[0114] D = gamma + D1*scale_ratio (1)
[0115] wherein, D represents the feature distance, D1 represents the first distance, scale_ratio represents the weight coefficient of the first distance, and gamma is a preset constant.
[0116] In an embodiment, the distance between the time domain feature of the target clothing region and the time domain feature of the candidate clothing region is calculated as the second distance. The specific implementation manner can include: calculating the time interval between the time point represented by the time domain feature of the target clothing region and the time point represented by the time domain feature of the candidate clothing region; if the time interval is less than or equal to a first preset interval, determining the second distance as the first preset interval; if the time interval is greater than the first preset interval and less than a second preset interval, determining the second distance as the time interval; if the time interval is greater than the first preset interval and greater than or equal to the second preset interval, determining the second distance as the second preset interval; wherein the second preset interval is greater than the first preset interval.
[0117] In actual application, if the time intervals corresponding to the actual target clothing region and the candidate clothing region are relatively small, it can be considered that they all belong to a range with high time domain consistency within a certain range, and the effects of the above feature distances are similar, and they can be distinguished. If the actual time intervals are relatively large and exceed a certain range, they all belong to a range with low time domain consistency, and the effects of the above feature distances are similar, and they can not be distinguished. Based on this, in the embodiment, in order to simplify the processing, two time intervals are preset, that is, a first preset interval and a second preset interval, one of which is larger and the other of which is smaller, and the second preset interval can be set to be greater than the first preset interval. Based on this, the time interval corresponding to the actual target clothing region and the candidate clothing region is compared with the first preset interval and the second preset interval respectively, if the time interval is less than or equal to the first preset interval, it can be considered that it belongs to a range with high time domain consistency, and the second distance is uniformly determined as the first preset interval; if the time interval is greater than the first preset interval and less than the second preset interval, the actual time domain consistency degree can be used to determine the second distance as the time interval; if the time interval is greater than the first preset interval and greater than or equal to the second preset interval, it can be considered that it belongs to a range with low time domain consistency, and the second distance is uniformly determined as the second preset interval, thereby simplifying the processing.
[0118] Exemplarily, the second distance can be calculated by the following formula:
[0119] tm_coef = min(max(0, tm1-tm2-min_tm_gap), max_tm_gap)*PI / max_tm_gap (2)
[0120] Wherein, tm_coef represents the second distance, tm1 represents the time point represented by the time domain feature of the target clothing region, tm2 represents the time point represented by the time domain feature of the candidate clothing region, min() represents the minimum value function, max() represents the maximum value function, min_tm_gap represents the first preset interval, max_tm_gap represents the second preset interval, and PI is a preset constant.
[0121] The specific values of the first preset interval and the second preset interval can be set according to actual conditions.
[0122] Exemplarily, based on the second distance, the weight coefficient of the first distance is determined, which can be realized by the following formula:
[0123] scale_ratio = (alpha+cosine(tm_coef))*beta (3)
[0124] Wherein, cosine() represents the cosine function value, and alpha and beta are both preset constants.
[0125] The above related embodiments only exemplarily illustrate the way of calculating the feature distance, and the feature distance can also be calculated in other ways, for example, the feature distance can be calculated by weighting the first distance and the second distance.
[0126] In actual application, when fine-grained clustering, the image feature distance and the time domain interval distance can be combined to construct the distance measurement matrix between all effective target clothing regions.
[0127] The distance measurement matrix D is an N*N matrix, and N represents the number of effective target clothing regions in the video obtained in the second step. D(i,j) is an element in the distance measurement matrix D, i and j are identifiers of target clothing regions in the N target clothing regions, j represents the jth target clothing region, and D(i,j) represents the feature distance between the features of the ith target clothing region and the features of the jth target clothing region. The label mapping algorithm is used to construct the distance measurement matrix D, and "i,j" contained in D(i,j) is the forward mapping label of the target clothing region i to the target clothing region j.
[0128] Then, each element D(i,j) in the distance measurement matrix D is assigned a value, which can be:
[0129] The distance between the image features i_feat of the ith target clothing region and the image features j_feat of the jth target clothing region is calculated to obtain the image feature distance, which can be calculated by calculating the similarity simi_val(i,j). The similarity can be calculated by a pre-constructed similarity function simi(), and the calculation formula is as follows:
[0130] simi_val(i,j)=simi(i_feat,j_feat) (4)
[0131] Since the greater the similarity, the smaller the distance, -simi_val(i,j) can be taken as the image feature distance.
[0132] The distance between the time domain features i_tm of the ith target clothing region and the image features j_tm of the jth target clothing region is calculated to obtain the time domain interval distance, and the calculation formula is as follows:
[0133] tm_coef=min(max(0,i_tm–j_tm-min_tm_gap),max_tm_gap)*PI / max_tm_gap(5)
[0134] The weight coefficient of the image feature distance is calculated, and the calculation formula is referred to formula (3).
[0135] D(i,j) is calculated based on image feature distance and weight coefficient, and the calculation formula is as follows:
[0136] D(i,j) = gamma-simi_val(i,j)*scale_ratio (6)
[0137] Exemplarily, min_tm_gap = 30 and max_tm_gap = 300 can be set.
[0138] Since it is required to determine the jth target clothing region with the minimum feature distance in time sequence before the ith target clothing region (i.e., the most similar features), in order to meet this condition, the following setting can be used to increase the forward mapping label constraint:
[0139] (1) If i = j, it indicates that the jth target clothing region and the ith target clothing region are the same clothing region, which does not meet the requirement, at this time, D(i,j) = val1 is set, val1 is a first preset feature distance, and exemplarily, 2.0 is set.
[0140] (2) If i_tm = j_tm, it indicates that the jth target clothing region and the ith target clothing region are at the same time, which does not meet the requirement, at this time, D(i,j) = val2 is set, val2 is a second preset feature distance, and exemplarily, 3.0 is set.
[0141] (3) If i_tm < j_tm, it indicates that the jth target clothing region is later than the ith target clothing region in time sequence, which does not meet the requirement, at this time, D(i,j) = val3 is set, val3 is a third preset feature distance, and exemplarily, 4.0 is set.
[0142] The first preset feature distance, the second preset feature distance, and the third preset feature distance are all greater than the calculated feature distance, and will not be selected when the minimum feature distance is selected.
[0143] Based on the distance measurement matrix D, the minimum element D(i,j) corresponding to the ith target clothing region is selected, and the forward mapping relationship i,j (i.e., the association information) corresponding to the ith target clothing region is obtained based on the forward mapping label contained in the selected D(i,j).
[0144] The obtained forward mapping relationship is clustered to obtain a fine-grained clustering result. The forward mapping relationships in the same class can be sequentially spliced to form an identification string, and the identification of the spliced part in the two forward mapping relationships to be spliced is the same.
[0145] The clothing recognition method provided by the embodiment of the application will be described in more detail below with reference to a specific application scenario.
[0146] The embodiment proposes a video clothing style recognition method with time consistency combining clothing clustering and image clothing recognition technology, which improves the rationality of clothing same or similar style recognition results at the video level.
[0147] As shown in Figure 3 The clothing recognition method of the embodiment mainly includes video key frame extraction, clothing area detection and effectiveness confirmation, clothing area clustering, clothing recognition, recognition result aggregation, and bidirectional tracking of recognition results. The detailed process is as follows:
[0148] First step, video key frame extraction.
[0149] In this step, K key frames at different times are extracted from the video to be recognized, obtaining key frame (or frame) 1, …, key frame K. The extraction method of the key frame can refer to related technologies, which will not be described here.
[0150] Second step, clothing area detection and effectiveness determination.
[0151] In this step, the key frame is subjected to clothing area detection (i.e., clothing area detection). Each target clothing area detected is subjected to clothing category recognition, including style category and posture category, which includes wearing state and non-wearing state; if the posture category of the target clothing area is wearing state, the human body area in the key frame where the target clothing area is located is detected, at least one human body area in the key frame where the target clothing area is located is detected, the target clothing area is compared with each human body area, and based on the comparison result, the clothing area to be removed is determined, the clothing area outside the clothing area to be removed in the key frame is determined as the effective clothing area (i.e., the effectiveness determination of the clothing area), and the effective clothing area set in the video is obtained. The specific implementation mode can refer to the description of the above related embodiments, which will not be described here.
[0152] In addition, the clothing category classifier can be pre-trained, which is used for recognizing the clothing category, and outputs the posture category and style category of the clothing.
[0153] Third step, clothing area clustering.
[0154] In this step, texture features and color features of the target clothing region are extracted; based on the texture features and color features of each target clothing region, each target clothing region is clustered to obtain at least one type of target clothing region, for example, including sub-class 1, …, sub-class C, the obtained sub-class 1 includes target clothing region 1 in key frame 1, target clothing region 1 in key frame 2, …; …; sub-class C includes target clothing region 1 in key frame 1, target clothing region 2 in key frame K, … Clustering can use mature clustering techniques such as KMeans or DBSCAN (Density-Based Spatial Clustering of Applications with Noise) to achieve simplicity.
[0155] In this step, texture features and color features of the target clothing region are extracted; based on the texture features and color features of each target clothing region, each target clothing region is clustered to obtain at least one type of target clothing region, for example, including sub-class 1, …, sub-class C, the obtained sub-class 1 includes target clothing region 1 in key frame 1, target clothing region 1 in key frame 2, …; …; sub-class C includes target clothing region 1 in key frame 1, target clothing region 2 in key frame K, … Clustering can use mature clustering techniques such as KMeans or DBSCAN (Density-Based Spatial Clustering of Applications with Noise) to achieve simplicity.
[0156] Fourth step, clothing recognition.
[0157] In this step, clothing recognition is performed on each target clothing region in each type of target clothing region to obtain the recognition result of each target clothing region. For example, in sub-class 1, the recognition result of target clothing region 1 in key frame 1 is candidate style set 1-1, the recognition result of target clothing region 1 in key frame 2 is candidate style set 2-1, …; in sub-class C, the recognition result of target clothing region 1 in key frame 1 is candidate style set 1-2, the recognition result of target clothing region 2 in key frame K is candidate style K-2, ….
[0158] In implementation, texture features and color features of candidate styles can be extracted in advance and saved to a database. During clothing recognition, candidate style sets can be obtained based on the similarity of the texture features and color features of the target clothing region and the texture features and color features of the candidate styles.
[0159] Fifth step, aggregation of recognition results.
[0160] In this step, for each sub-class, all candidate styles of the sub-class are voted (i.e., the score is calculated), and the candidate style with the highest score is taken as the recognition result of the sub-class.
[0161] The voting score is calculated based on the similarity s1 of the texture features of the candidate style and the texture features of the target clothing region, the similarity s2 of the color features of the candidate style and the color features of the target clothing region, the number of intra-class style hits (i.e., the first hit number) N1, and the number of overall video style hits (i.e., the second hit number) N2.
[0162] score=w1*s1+w2*s2+w3*N1+w4(N1 / N2)(1)
[0163] Wherein, score represents the voting score. w1, w2, w3, w4 are all weights. For example, the values of the weights can be set as w1=0.65, w2=0.2, w3=0.08, w3=0.07, and the sum of w1, w2, w3 and w4 is 1.
[0164] By determining the time point of the key frame where the target clothing region is located, the identification result of the key frame where the target clothing region is located is obtained based on the identification result of the target clothing region. In this way, the identification result of the target clothing region is extended to all time points of the key frame contained in the target clothing region category, and the clothing recall time length at the video level is improved.
[0165] Step 6, bidirectional tracking of the video sequence.
[0166] In this step, the video frames before and after the key frame in the to-be-identified video can be subjected to region detection to determine the video frame where the target clothing region appears in the video frame sequence of the to-be-identified video, obtain the time point where the target clothing region appears, and determine the position information of the target clothing region in the spatial domain, so as to obtain the time point and position information of the clothing with the same or similar style appearing in the to-be-identified video. Based on the determined time point and position information, the display of the clothing identification result is controlled.
[0167] In this embodiment, a clothing identification method with time domain consistency is proposed by combining clothing identification and clustering technology, which effectively reduces the one-to-many and many-to-one mapping of the query clothing and the styles in the database, improves the time domain consistency of the clothing identification result in the video, effectively improves the rationality of the clothing identification result at the video level, and can be used in scenes such as clothing recommendation and clothing identification of the same or similar style (such as star same or similar style) in the video, thereby improving the user experience.
[0168] Figure 4 An exemplary clothing identification device structure diagram is provided for the embodiments of the present application. As shown in the figure, Figure 4 The device 400 comprises:
[0169] The key frame extraction module 401 is configured to extract a plurality of key frames at different time points from the to-be-identified video.
[0170] The region detection module 402 is configured to perform clothing region detection on the key frame.
[0171] The feature extraction module 403 is configured to extract at least one feature of the target clothing region.
[0172] The region clustering module 404 is configured to cluster the target clothing regions based on at least one feature of the target clothing regions, to obtain at least one category of target clothing regions.
[0173] The clothing recognition module 405 is configured to perform clothing recognition on each target clothing region in each category of target clothing regions, to obtain a recognition result of each target clothing region, wherein the recognition result of the target clothing region includes at least one candidate style, and the at least one candidate style is selected from the recognition results of the target clothing regions in the category of target clothing regions as the recognition result of the category of target clothing regions.
[0174] In an embodiment, the clothing recognition module 405 is specifically configured to:
[0175] form a candidate style set based on the candidate styles in the recognition results of the target clothing regions in the category of target clothing regions;
[0176] calculate a score of each candidate style in the candidate style set;
[0177] select the at least one candidate style from the candidate style set based on the scores of the candidate styles.
[0178] In an embodiment, the clothing recognition module 405 is specifically configured to:
[0179] calculate the score of each candidate style in the candidate style set based on the similarity between each feature of the candidate style and the same feature of the corresponding target clothing region, and / or the first hit number, and / or the second hit number;
[0180] wherein the first hit number is the total hit number of the candidate style in the recognition results of all target clothing regions in which the candidate style is formed into the candidate style set;
[0181] the second hit number is the total hit number of the candidate style in the recognition results of all target clothing regions in the video to be recognized.
[0182] In an embodiment, the clothing recognition module 405 is specifically configured to:
[0183] weight-sum the similarity, the first hit number, and the ratio of the first hit number to the second hit number, to obtain the score of the candidate style.
[0184] In an embodiment, the weight of the similarity is greater than the weight of the first hit number, and / or the weight of the first hit number is greater than the weight of the ratio.
[0185] In an embodiment, the at least one feature includes a texture feature and a color feature.
[0186] In an embodiment, as shown in Figure 5 Further comprising:
[0187] The region removing module 406 is configured to perform human region detection on the key frame.
[0188] In response to detecting the at least one target clothing region and the at least one human region, the target clothing region is compared with each human region, and based on a comparison result, it is determined whether the target clothing region is a clothing region that needs to be removed, the clothing region that needs to be removed has an overlap ratio with the human region reaching a first threshold and a pose similarity not exceeding a second threshold.
[0189] The target clothing region in the key frame that needs to be removed is removed.
[0190] In an embodiment, the region removing module is specifically configured to:
[0191] Calculate an intersection-over-union of the target clothing region and each human region.
[0192] Based on the calculation result, it is determined that the target clothing region is a clothing region that needs to be removed.
[0193] In an embodiment, in an embodiment, the region removing module is specifically configured to:
[0194] In response to a maximum value in the intersection-overs-union being greater than or equal to a first threshold and a pose similarity corresponding to the maximum value not exceeding a second threshold, it is determined that the target clothing region is a clothing region that needs to be removed.
[0195] In an embodiment, in an embodiment, the region removing module is specifically configured to:
[0196] In response to more than two intersection-overs-union in the intersection-overs-union being greater than or equal to the first threshold, it is determined that the clothing region is a clothing region that needs to be removed.
[0197] In an embodiment, the human region contains at least one human pose key node.
[0198] In an embodiment, the region removing module is specifically configured to:
[0199] Determine a human region corresponding to a maximum value in the intersection-overs-union as a target human region.
[0200] Obtain a style category of the recognized target clothing region.
[0201] Obtain at least one human pose key node contained in a wearing part of the recognized style category that is preset.
[0202] The number of the same human posture key nodes contained in the target human body region and the wearing part of the style category recognized by the statistics;
[0203] In response to the number of the statistics not exceeding the second threshold value, the target clothing region is determined as a clothing region that needs to be removed.
[0204] In an embodiment, the region clustering module is specifically configured to:
[0205] Based on at least one feature of each target clothing region, the target clothing regions are clustered once to obtain at least one type of target clothing region corresponding to the once clustering;
[0206] The at least one type of target clothing region corresponding to the once clustering is taken as at least one clustering object, and the at least one clustering object is clustered again to obtain at least one type of target clothing region corresponding to the twice clustering.
[0207] In an embodiment, the clothing recognition module is specifically configured to:
[0208] For each target clothing region in each type of target clothing region corresponding to the twice clustering, clothing recognition is performed on each target clothing region in the type of target clothing region to obtain an identification result of each target clothing region, the identification result of the target clothing region includes at least one candidate style, and at least one candidate style is selected from the identification results of the target clothing regions in the type of target clothing region as the identification result of the type of target clothing region.
[0209] Alternatively, for each target clothing region in each type of target clothing region corresponding to the once clustering, clothing recognition is performed on each target clothing region in the type of target clothing region to obtain an identification result of each target clothing region, the identification result of the target clothing region includes at least one candidate style, and at least one candidate style is selected from the identification results of the target clothing regions in the type of target clothing region as the identification result of the type of target clothing region; and for each type of target clothing region corresponding to the twice clustering, at least one candidate style is selected from the identification results of the clustering objects in the type of target clothing region as the identification result of the type of target clothing region.
[0210] In an embodiment, the region clustering module is specifically configured to:
[0211] For each target clothing region, the associated clothing region corresponding to the target clothing region is determined, as well as the associated information corresponding to the target clothing region. The associated clothing region comes from the candidate clothing region set. The candidate clothing region set contains target clothing regions whose time is earlier than the target clothing region as represented by the time domain features. The associated clothing region is the candidate clothing region in the candidate clothing region set whose features are most similar to the target clothing region. The associated information includes the identifier of the target clothing region and the identifier of the corresponding associated clothing region.
[0212] Classify the logo combinations corresponding to each target clothing area to obtain at least one type of logo combination. Among them, the logo combinations of the same type can be spliced together in the order of their beginning and end to form a logo string. In the two spliced logo combinations, the logos of the spliced parts are the same.
[0213] For each type of logo combination, the target clothing areas corresponding to that type of logo combination are treated as a category.
[0214] The functions of each module in the devices provided in the embodiments of the present invention can be found in the corresponding descriptions in the above embodiments of the clothing recognition method, and will not be repeated here.
[0215] This invention also provides an electronic device, such as... Figure 6 As shown, it includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604.
[0216] Memory 603 is used to store computer programs;
[0217] When processor 601 executes a program stored in memory 603, it performs the following steps:
[0218] Extract keyframes from multiple different moments in the video to be identified;
[0219] Perform clothing region detection on keyframes;
[0220] For the detected target clothing area, extract at least one feature of the target clothing area;
[0221] Based on at least one feature of each target clothing region, the target clothing regions are clustered to obtain at least one class of target clothing regions.
[0222] For each type of target clothing area, clothing recognition is performed on each target clothing area in the type of target clothing area to obtain a recognition result of each target clothing area, the recognition result of the target clothing area including at least one candidate style, and at least one candidate style is selected from the recognition results of the target clothing areas in the type of target clothing area as the recognition result of the type of target clothing area.
[0223] The communication bus mentioned above can be a Perheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0224] The communication interface is used for communication between the terminal and other devices.
[0225] The memory can include a Random Access Memory (RAM) and can also include a non-volatile memory, such as at least one disk memory. Optionally, the memory can also be at least one storage device located away from the processor.
[0226] The processor mentioned above can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; can also be a Digital Signal Processing (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0227] In another embodiment provided by the application, a computer readable storage medium is also provided, and the computer readable storage medium stores instructions, when the instructions are run on a computer, the computer is caused to execute the clothing recognition method in any of the above embodiments.
[0228] In yet another embodiment of the present disclosure, there is provided a computer program product comprising instructions which, when executed on a computer, cause the computer to carry out the method of identifying the clothing according to any one of the above embodiments.
[0229] In the above embodiments, the whole or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, the whole or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the whole or part of the processes or functions according to the embodiments of the present disclosure are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (for example, floppy disk, hard disk, magnetic tape), optical media (for example, DVD), or semiconductor media (for example, solid state disk (SSD)) and the like.
[0230] It should be noted that, in this document, the terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the processes, methods, articles or devices including a series of elements not only include those elements, but also include other elements not explicitly listed or inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0231] The various embodiments in the specification are described in a related manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.
[0232] The above only describes the preferred embodiments of the present application, and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for identifying clothing, characterized in that, include: Extract keyframes from multiple different moments in the video to be identified; Perform clothing region detection on the keyframes to obtain at least one target clothing region; Human region detection is performed on the keyframes to obtain at least one human region; In response to detecting at least one target clothing region and at least one human body region, the target clothing region is compared with each human body region. Based on the comparison results, it is determined whether the target clothing region is a clothing region that needs to be removed. The overlap ratio between the clothing region that needs to be removed and the human body region reaches a first threshold and the pose similarity does not exceed a second threshold. Remove the target clothing area that needs to be removed from the keyframe; For the removed target clothing area, extract at least one feature of the target clothing area; Based on at least one feature of each target clothing region, the target clothing regions are clustered to obtain at least one class of target clothing regions. The temporal consistency of target clothing regions of the same class is higher than that of target clothing regions of different classes, and the feature similarity of target clothing regions of the same class is higher than that of target clothing regions of different classes. The temporal consistency is inversely proportional to the temporal interval distance. For each target clothing region, the temporal interval distance is the distance between the temporal features of the target clothing region and the temporal features of the candidate clothing region. The candidate clothing region is the target clothing region whose time represented by the temporal features is earlier than that of the target clothing region. For each type of target clothing area, clothing recognition is performed on each of the target clothing areas in that type of target clothing area to obtain the recognition result of each target clothing area. The recognition result of the target clothing area includes at least one candidate style. At least one candidate style is selected from the recognition results of each target clothing area in that type of target clothing area to be used as the recognition result of that type of target clothing area.
2. The method according to claim 1, characterized in that, The step of selecting at least one candidate style from the identification results of each of the target clothing regions in this type of target clothing region includes: Based on the candidate styles in the identification results of each of the target clothing regions in this type of target clothing region, a candidate style set is formed; Calculate a score for each of the candidate styles in the candidate style set; Based on the scores of each candidate style, at least one candidate style is selected from the set of candidate styles.
3. The method according to claim 2, characterized in that, The calculation of the score for each candidate style in the candidate style set includes: For each candidate style in the candidate style set, a score is calculated based on the similarity of each feature of the candidate style to the same feature of the corresponding target clothing area, and / or, the first hit count, and / or, the second hit count. Wherein, the first hit count is the total hit count of the candidate style in the recognition results of all the target clothing areas that form the candidate style set; The second hit count is the total number of times the candidate style is hit in the recognition results of all the target clothing areas in the video to be recognized.
4. The method according to claim 3, characterized in that, The calculation of the score for the candidate style based on the similarity between each feature of the candidate style and the corresponding feature of the target clothing region, and / or, the first hit count, and / or, the second hit count, includes: The similarity, the first hit count, and the ratio of the first hit count to the second hit count are weighted and summed to obtain the score of the candidate style.
5. The method according to claim 4, characterized in that, The weight of the similarity is greater than the weight of the first hit count, and / or the weight of the first hit count is greater than the weight of the ratio.
6. The method according to claim 1, characterized in that, The at least one feature includes texture features and color features.
7. The method according to claim 1, characterized in that, The step of comparing the target clothing area with each of the human body areas, and determining whether the target clothing area is a clothing area that needs to be removed based on the comparison results, includes: Calculate the intersection-union ratio (IUU) between the target clothing area and each of the human body areas; Based on the calculation results, the target clothing area is determined to be the clothing area that needs to be removed.
8. The method according to claim 7, characterized in that, The step of determining the target clothing area as the clothing area to be removed based on the calculation results includes: In response to the fact that the maximum value of each of the intersection-union ratios is greater than or equal to the first threshold and the pose similarity corresponding to the maximum value does not exceed the second threshold, the target clothing region is determined to be a clothing region that needs to be removed.
9. The method according to claim 7, characterized in that, The keyframe contains two or more human body regions, and determining the target clothing region as the clothing region to be removed based on the calculation results includes: In response to two or more of the cross-union ratios being greater than or equal to a first threshold, the clothing area is determined to be a clothing area that needs to be removed.
10. The method according to claim 7, characterized in that, The human body region contains at least one key node for human posture. The step of determining the target clothing area as the clothing area to be removed based on the calculation results includes: The human body region corresponding to the maximum value among the intersection-union ratios is determined as the target human body region; Obtain the style category of the identified target clothing area; Obtain at least one key human posture node of the wearing part of the identified style category as preset; The number of identical human posture key nodes contained in the target human body region for the identified style category of the wearing part is counted. If the number of statistics does not exceed the second threshold, the target clothing area is determined as the clothing area that needs to be removed.
11. The method according to claim 1, characterized in that, The step of clustering the target clothing regions based on at least one feature of each target clothing region to obtain at least one class of target clothing regions includes: Based on at least one feature of each of the target clothing regions, each of the target clothing regions is clustered once to obtain at least one type of target clothing region corresponding to the first clustering. The at least one type of target clothing region corresponding to the first clustering is used as at least one clustering object, and then clustered again to obtain the at least one type of target clothing region corresponding to the second clustering.
12. The method according to claim 11, characterized in that, For each type of target clothing area, clothing recognition is performed on each target clothing area within that type of target clothing area to obtain a recognition result for each target clothing area. The recognition result for each target clothing area includes at least one candidate style. From the recognition results of each target clothing area within that type of target clothing area, at least one candidate style is selected as the recognition result for that type of target clothing area. This includes: For each of the at least one type of target clothing regions corresponding to the re-clustering, clothing identification is performed on each of the target clothing regions in the type of target clothing regions to obtain the identification result of each target clothing region. The identification result of the target clothing region includes at least one candidate style. At least one candidate style is selected from the identification results of each of the target clothing regions in the type of target clothing regions to be used as the identification result of the type of target clothing region. Alternatively, for each of the at least one type of target clothing regions corresponding to the first clustering, clothing identification is performed on each of the target clothing regions in that type of target clothing region to obtain an identification result for each target clothing region. The identification result of the target clothing region includes at least one candidate style. At least one candidate style is selected from the identification results of each target clothing region in that type of target clothing region to serve as the identification result for that type of target clothing region. For each of the at least one type of target clothing regions corresponding to the second clustering, at least one candidate style is selected from the identification results of each cluster object in that type of target clothing region to serve as the identification result for that type of target clothing region.
13. The method according to claim 11, characterized in that, The step of clustering each of the target clothing regions based on at least one feature of each target clothing region to obtain at least one class of target clothing regions corresponding to the first clustering includes: For each target clothing region, the associated clothing region corresponding to the target clothing region is determined, and the associated information corresponding to the target clothing region is determined. The associated clothing region comes from a set of candidate clothing regions. The candidate clothing regions included in the set of candidate clothing regions are target clothing regions whose time is earlier than the target clothing region as represented by the time domain features. The associated clothing region is the candidate clothing region in the set of candidate clothing regions whose features are most similar to the target clothing region. The associated information includes an identifier combination formed by the identifier of the target clothing region and the identifier of the corresponding associated clothing region. The identification combinations corresponding to each target clothing area are classified to obtain at least one type of identification combination. The identification combinations of the same type can be spliced together in the order of their beginning and end to form an identification string. In two spliced identification combinations, the identifications of the spliced parts are the same. For each type of identifier combination, the target clothing areas corresponding to that type of identifier combination are treated as a category.
14. A clothing recognition device, characterized in that, include: The keyframe extraction module is used to extract keyframes from multiple different moments in the video to be identified. The region detection module is used to perform clothing region detection on the keyframe to obtain at least one target clothing region. A region removal module is used to perform human region detection on the keyframe to obtain at least one human region. In response to detecting at least one target clothing region and at least one human body region, the target clothing region is compared with each human body region. Based on the comparison results, it is determined whether the target clothing region is a clothing region that needs to be removed. The overlap ratio between the clothing region that needs to be removed and the human body region reaches a first threshold and the pose similarity does not exceed a second threshold. Remove the target clothing area that needs to be removed from the keyframe; The feature extraction module is used to extract at least one feature of the target clothing area after it has been removed. The region clustering module is used to cluster the target clothing regions based on at least one feature of each target clothing region to obtain at least one class of target clothing regions. The temporal consistency of target clothing regions of the same class is higher than that of target clothing regions of different classes, and the feature similarity of target clothing regions of the same class is higher than that of target clothing regions of different classes. The temporal consistency is inversely proportional to the temporal interval distance. For each target clothing region, the temporal interval distance is the distance between the temporal features of the target clothing region and the temporal features of the candidate clothing region. The candidate clothing region is a target clothing region whose time represented by the temporal features is earlier than that of the target clothing region. The clothing recognition module is used to perform clothing recognition on each target clothing area in each category of target clothing areas to obtain a recognition result for each target clothing area. The recognition result of the target clothing area includes at least one candidate style. At least one candidate style is selected from the recognition results of each target clothing area in the category of target clothing areas to be used as the recognition result of the category of target clothing areas.
15. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-13.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-13.
Citation Information
Patent Citations
Face identification method and apparatus
CN104408404A
Face identification method, device and system
CN105956518A
Information acquisition method and device, storage medium and electronic device
CN111126179A
Facial information acquisition method and device
CN112101197A
Costume searching method and device, electronic equipment and medium
CN112905889A