Training method of costume category identification model and costume category identification method and device
By training multimodal pre-training models and combining adversarial learning methods, the problem of low recognition accuracy in clothing category recognition is solved, and the recognition accuracy of clothing categories and the ability to distinguish similar categories are improved.
Patent Information
- Application Number
- CN202510231613.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
AI Technical Summary
When identifying clothing categories based on deep learning models, the recognition accuracy of clothing in the video frame is low due to reasons such as occlusion, few exposed areas, and large angle changes.
By obtaining sample images and clothing category description text information, training multimodal pre-trained models, extracting image features, and generating pseudo-sample image features, and further training classifiers in combination with adversarial learning to improve the recognition accuracy of clothing categories.
The recognition accuracy of clothing categories is improved, and the contribution of image features to category determination is enhanced through the combination of multimodal pre-training model and adversarial learning is enhanced, and the ability to distinguish similar clothing categories is improved.
Smart Images

Figure CN120071013A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer application technologies, and in particular, to a method for training a clothing category recognition model, a method for recognizing a clothing category, and a device therefor. Background Art
[0002] The recognition of clothing categories is widely applied in video scenarios. For example, it can be used for the collection of clothing materials in videos, the summary and analysis of clothing categories, etc.
[0003] In related technologies, the recognition of clothing categories is directly based on a deep learning model, that is, a large number of sample images labeled with clothing categories are pre-obtained to train the deep learning model, and the trained deep learning model can recognize the clothing categories included in the input image.
[0004] However, in the above-mentioned method for recognizing clothing categories based on a deep learning model, when recognizing clothing in a video, the clothing in the video frame may have a low recognition accuracy of the deep learning model due to reasons such as occlusion, small exposed areas, and large angle changes. Summary of the Invention
[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method for training a clothing category recognition model, a method for recognizing a clothing category, and a device therefor, which can solve the technical problem of low recognition accuracy when recognizing clothing categories based on a deep learning model.
[0006] An embodiment of the present disclosure provides a method for training a clothing category recognition model. The method includes: obtaining a sample image, obtaining a plurality of pre-generated clothing category description text information, and training a pre-constructed initial multi-modal pre-training model according to the sample image and the plurality of clothing category description text information to obtain a trained multi-modal pre-training model, wherein the trained multi-modal pre-training model learns to output image features associated with clothing categories according to the input image; extracting a plurality of second sample image features of the sample image according to the multi-modal pre-training model, and generating a plurality of pseudo-sample image features corresponding to the plurality of second sample image features one by one, wherein the feature similarity between each pseudo-sample image feature and the sample image is less than the feature similarity between the second sample image feature corresponding to each pseudo-sample image feature and the sample image; respectively inputting the plurality of second sample image features and the plurality of pseudo-sample image features into a pre-constructed initial classifier, obtaining a second clothing category corresponding to the plurality of second sample image features and a third clothing category corresponding to the plurality of pseudo-sample image features, determining a second category loss value between the second clothing category and a preset standard clothing category corresponding to the sample image, and determining a third category loss value between the third clothing category and the preset standard clothing category, and obtaining a trained classifier according to the second category loss value and the third category loss value, wherein the trained classifier learns to output a clothing category according to the input image features.
[0007] An embodiment of the present disclosure further provides a method for recognizing clothing categories, including: extracting at least one video frame from a target video stream; extracting a plurality of target sample image features corresponding to each video frame through a pre-trained multi-modal pre-training model, wherein the pre-trained multi-modal pre-training model is trained by the training method of the clothing category recognition model as described above; inputting the plurality of target sample image features into a pre-trained classifier, obtaining a target clothing category output by the classifier, wherein the pre-trained classifier is trained by the training method of the clothing category recognition model as described above; and taking all the target clothing categories corresponding to the at least one video frame as the clothing category recognition result of the target video stream.
[0008] The embodiment of the present disclosure also provides a training device for a clothing category recognition model. The device includes: an acquisition module, configured to acquire sample images and acquire a plurality of pre-generated clothing category description text information; a first training module, configured to train a pre-constructed initial multi-modal pre-training model according to the sample images and the plurality of clothing category description text information to obtain a trained multi-modal pre-training model, wherein the trained multi-modal pre-training model learns to output image features associated with clothing categories according to the input images; a first feature extraction module, configured to extract a plurality of second sample image features of the sample images according to the multi-modal pre-training model, and generate a plurality of pseudo-sample image features corresponding to the plurality of second sample image features one by one, wherein the feature similarity between each pseudo-sample image feature and the sample image is less than the feature similarity between the second sample image feature corresponding to each pseudo-sample image feature and the sample image; a second training module, configured to respectively input the plurality of second sample image features and the plurality of pseudo-sample image features into a pre-constructed initial classifier, obtain a second clothing category corresponding to the plurality of second sample image features and a third clothing category corresponding to the plurality of pseudo-sample image features, determine a second category loss value between the second clothing category and a preset standard clothing category corresponding to the sample image, and determine a third category loss value between the third clothing category and the preset standard clothing category, and obtain a trained classifier according to the second category loss value and the third category loss value, wherein the trained classifier learns to output clothing categories according to the input image features.
[0009] The embodiment of the present disclosure also provides a recognition device for clothing categories. The device includes: an extraction module, configured to extract at least one video frame from a target video stream; a second feature extraction module, configured to extract a plurality of target sample image features corresponding to each frame of the video frame through a pre-trained multi-modal pre-training model, wherein the pre-trained multi-modal pre-training model is trained by the training method of the clothing category recognition model as described above; a second clothing category recognition module, configured to input the plurality of target sample image features into a pre-trained classifier, and obtain a target clothing category output by the classifier, wherein the pre-trained classifier is trained by the training method of the clothing category recognition model as described above; a recognition result determination module, configured to use all the target clothing categories corresponding to the at least one video frame as the clothing category recognition result of the target video stream.
[0010] The embodiment of the present disclosure also provides an electronic device. The electronic device includes: a processor; a memory for storing executable instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the training method of the clothing category recognition model provided by the embodiment of the present disclosure, or the recognition method of clothing categories.
[0011] The embodiments of the present disclosure also provide a computer-readable storage medium storing a computer program for executing the training method of the clothing category recognition model or the recognition method of clothing categories provided by the embodiments of the present disclosure.
[0012] The technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art:
[0013] In the clothing category recognition solution provided by the embodiments of the present disclosure, a sample image is obtained, multiple pre-generated clothing category description text information is obtained, and an initially constructed initial multi-modal pre-trained model is trained according to the sample image and the multiple clothing category description text information to obtain a trained multi-modal pre-trained model. Among them, the trained multi-modal pre-trained model learns to output image features associated with clothing categories according to the input image. Multiple second sample image features of the sample image are extracted according to the multi-modal pre-trained model, and multiple pseudo-sample image features corresponding to the multiple second sample image features are generated. Among them, the feature similarity between each pseudo-sample image feature and the sample image is less than the feature similarity between the second sample image feature corresponding to each pseudo-sample image feature and the sample image. Furthermore, the multiple second sample image features and the multiple pseudo-sample image features are respectively input into an initially constructed initial classifier to obtain a second clothing category corresponding to the multiple second sample image features and a third clothing category corresponding to the multiple pseudo-sample image features. A second category loss value between the second clothing category and the preset standard clothing category corresponding to the sample image is determined, and a third category loss value between the third clothing category and the preset standard clothing category is determined. A trained classifier is obtained according to the second category loss value and the third category loss value. Among them, the trained classifier learns to output clothing categories according to the input image features. In this technical solution, the multi-modal pre-trained model is trained based on the sample image and the clothing category description text information. When the clothing category is recognized based on the multi-modal pre-trained model, the contribution degree of the image features extracted by the multi-modal pre-trained model to the determination of the image category helps to improve the determination accuracy of the clothing category. And on the basis of the trained multi-modal pre-trained model, the classifier is further trained, and the classifier is trained in the form of adversarial learning, which ensures the recognition accuracy of the clothing category. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In combination with the drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale.
[0015] Figure 1Flow chart of a method for training a clothing category recognition model provided by an embodiment of the present disclosure;
[0016] Figure 2 Scene diagram of a multi-modal pre-training model provided by an embodiment of the present disclosure;
[0017] Figure 3 Scene diagram of generating pseudo-sample image features provided by an embodiment of the present disclosure;
[0018] Figure 4 Flow chart of a method for recognizing clothing categories provided by an embodiment of the present disclosure;
[0019] Figure 5 Structural diagram of a training device for a clothing category recognition model provided by an embodiment of the present disclosure;
[0020] Figure 6 Structural diagram of a recognition device for clothing categories provided by an embodiment of the present disclosure;
[0021] Figure 7 Structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0022] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not used to limit the protection scope of the present disclosure.
[0023] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0024] As used herein, the term "including" and its variants are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0025] It should be noted that the concepts such as "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0026] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0027] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0028] To solve the above problems, an embodiment of this disclosure provides a method for identifying clothing categories. In this method, a multi-modal pre-training model (Contrastive Language-Image Pre-training, CLIP) and the like are applied.
[0029] Therefore, before explaining the method for identifying clothing categories, the training method of the clothing category recognition model is first explained.
[0030] Figure 1 As shown in the flowchart of the training method of the clothing category recognition model according to an embodiment of this disclosure, this method can be executed by a training device of the clothing category recognition model, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device. Figure 1 As shown, this method includes:
[0031] Step 101, obtain sample images, obtain multiple pre-generated clothing category description text information, and train a pre-constructed initial multi-modal pre-training model according to the sample images and the multiple clothing category description text information to obtain a trained multi-modal pre-training model. Among them, the trained multi-modal pre-training model learns to output image features associated with clothing categories according to the input images.
[0032] Among them, the sample images can be sourced from a video stream. In the embodiments of this disclosure, the sample images can be one or multiple. When there are multiple sample images, the multiple sample images can be jointly used for training the multi-modal pre-training model. The method of training the multi-modal pre-training model with multiple sample images is similar to the method of training the multi-modal pre-training model with one sample image. The difference is that when calculating the loss value, the average loss value of multiple sample images, or the number of sample images that meet the preset convergence condition, can be used to determine whether to complete the training of the multi-modal pre-training model.
[0033] In one embodiment of the present disclosure, a sample image is obtained, and a plurality of pre-generated clothing category description text information is obtained. The clothing category description text information can be any text describing clothing categories, including but not limited to clothing category texts, clothing category description short sentences, etc. In some possible embodiments, the clothing category description text information may include "chef's uniform" and "doctor's uniform", etc.
[0034] In this embodiment, with reference to Figure 2 , the sample image and the plurality of clothing category description text information can be input into a pre-constructed initial multi-modal pre-training model.
[0035] Among them, considering that the sample image may contain some content unrelated to clothing, therefore, in one embodiment of the present disclosure, before inputting the sample image and the plurality of clothing category description text information into the pre-constructed initial multi-modal pre-training model, the sample image region containing clothing in the sample image can be identified. For example, the sample image region can be detected by an open-source deep learning detector (such as the YOLO series). Furthermore, the first region coordinate position information of the sample image region in the sample image is determined. The first region coordinate position information may include ('cloth_id', 'top_left_x', 'top_left_y', 'bottom_right_x', 'bottom_right_y'), which respectively represent the clothing box number, the upper left coordinate x, the upper left coordinate y, the lower right coordinate x, the lower right coordinate y, etc. Furthermore, the initial multi-modal pre-training model extracts a plurality of first sample image features corresponding to the sample image region according to the first region coordinate position information, that is, only the relevant features of the image region containing clothing are extracted.
[0036] In one embodiment of the present disclosure, in order to further ensure the training effect of the multi-modal pre-training model, the confidence level that the sample image region contains clothing can also be extracted. When the confidence level is greater than the preset confidence threshold, it will be input into the initial multi-modal pre-training model. Otherwise, the sample image is replaced.
[0037] In this embodiment, different methods can be adopted to train the pre-constructed initial multi-modal pre-training model according to different application scenarios. In some possible embodiments, a plurality of first sample image features corresponding to the sample image are extracted through the initial multi-modal pre-training model, and a plurality of clothing category text features corresponding one-to-one to the plurality of clothing category description text information, that is, with reference to Figure 2, the initial multi-modal pre-training model includes an image feature extraction module and a text feature extraction module. Based on the image feature extraction module, multiple first sample image features I1, I2, I3... IN corresponding to the sample image are extracted, and multiple clothing category text features corresponding to multiple clothing category description text information are extracted through the text feature extraction module, namely multiple first sample image features T1, T2, T3... TN.
[0038] Furthermore, the feature similarity between each first sample image feature and each clothing category text feature can be calculated through the initial multi-modal pre-training model, that is, continue to refer to Figure 2 , the feature similarities I1T1, I1T2, I1T3... I1TN between I1 and T1, T2, T3... TN can be calculated, the feature similarities I2T1, I2T2, I2T3... I2TN between I2 and T1, T2, T3... TN can be calculated... the feature similarities INT1, INT2, INT3... INTN between IN and T1, T2, T3... TN can be calculated.
[0039] In this embodiment, after calculating the feature similarity, the first clothing category is determined and output through the initial multi-modal pre-training model. For example, the maximum feature similarity among all feature similarities can be determined through the initial multi-modal pre-training model, and the first clothing category corresponding to the clothing category text feature corresponding to the maximum feature similarity is determined, and the first clothing category is output through the initial multi-modal pre-training model.
[0040] After determining the first clothing category, calculate the first category loss value between the first clothing category and the preset standard clothing category corresponding to the sample image, that is, determine whether the accuracy of the initial multi-modal pre-training model is appropriate. Determine whether the first category loss value is less than or equal to the first preset classification loss threshold, where the first preset classification loss threshold can be set according to the scenario requirements.
[0041] In this embodiment, when the first category loss value is greater than the first preset classification loss threshold, the model weight parameters of the initial multi-modal pre-training model are modified until the first category loss value is less than or equal to the first preset classification loss threshold to obtain the trained multi-modal pre-training model.
[0042] Thus, in an embodiment of the present disclosure, a CLIP model with strong generalization ability is trained to facilitate the subsequent CLIP model to identify multiple clothing categories and ensure the recognition accuracy of clothing categories.
[0043] Step 102: Extract multiple second sample image features of the sample image according to the multi-modal pre-trained model, and generate multiple pseudo-sample image features corresponding one-to-one to the multiple second sample image features. Among them, the feature similarity between each pseudo-sample image feature and the sample image is less than the feature similarity between the second sample image feature corresponding to each pseudo-sample image feature and the sample image.
[0044] Step 103: Input the multiple second sample image features and the multiple pseudo-sample image features into the pre-constructed initial classifier respectively, obtain the second clothing category corresponding to the multiple second sample image features and the third clothing category corresponding to the multiple pseudo-sample image features, determine the second category loss value between the second clothing category and the preset standard clothing category corresponding to the sample image, and determine the third category loss value between the third clothing category and the preset standard clothing category. Obtain the trained classifier according to the second category loss value and the third category loss value. Among them, the trained classifier is trained and learned to output the clothing category according to the input image features.
[0045] In an embodiment of the present disclosure, considering that some clothing categories are relatively similar. For example, doctor's uniforms and chef's uniforms are both white, and misjudgments may occur sometimes. Therefore, in order to further improve the recognition accuracy of clothing categories, a classifier can also be trained. This classifier pre-learns different features of similar clothing categories and can accurately perform image classification.
[0046] In the embodiment of the present disclosure, multiple second sample image features of the sample image are extracted through the trained multi-modal pre-trained model. Among them, the multiple second sample image features are features that contribute greatly to the recognition of clothing categories.
[0047] In the embodiment of the present disclosure, multiple pseudo-sample image features corresponding one-to-one to the multiple second sample image features are generated through a generative adversarial network. Among them, the multiple pseudo-sample image features are as similar as possible to the second sample image features, but there are also differences, so that the multiple pseudo-sample image features can be used as image features of clothing categories similar to the clothing categories corresponding to the second sample image features.
[0048] That is, referring to Figure 3 , in the embodiment of the present disclosure, the multiple second sample image features are input into the generator of the adversarial network to obtain multiple pseudo-sample image features.
[0049] In an embodiment of the present disclosure, in order to ensure the similarity between the multiple pseudo-sample image features and the multiple second sample image features, in this embodiment, the classification loss value and the adversarial loss value between the multiple pseudo-sample image features and the multiple second sample image features can be calculated. Among them, the calculation function of the classification loss function can be any function such as the cross-entropy loss function, and the calculation function of the adversarial loss value can be any function that can calculate the follow-up loss value, etc.
[0050] Determine the image feature loss value according to the classification loss value and the adversarial loss value. For example, the first weight corresponding to the classification loss value can be determined, the second weight corresponding to the adversarial loss value can be determined, the first product value of the classification loss value and the corresponding first weight can be calculated, the second product value of the adversarial loss value and the corresponding second weight can be calculated, and the sum of the first product value and the second product value can be calculated to obtain the image feature loss value.
[0051] In this embodiment, determine whether the image feature loss value is less than the preset feature loss threshold. If it is less than or equal to the preset feature loss threshold, the subsequent step 303 can be executed. When the image feature loss value is not less than the preset feature loss threshold, generate a plurality of updated pseudo-sample image features corresponding one-to-one to the plurality of second sample image features through the generative adversarial network until it is determined that the image feature loss value is less than the preset feature loss threshold, that is, ensure that the plurality of pseudo-sample image features can perturb the images of similar clothing categories of the plurality of second sample image features. Enhance the feature expression through adversarial training to ensure that the trained classifier maintains a greater feature dispersion between similar clothing categories.
[0052] In this embodiment, input the plurality of second sample image features and the plurality of pseudo-sample image features into the pre-constructed initial classifier respectively, obtain the second clothing category corresponding to the plurality of second sample image features and the third clothing category corresponding to the plurality of pseudo-sample image features, determine the second category loss value between the second clothing category and the preset standard clothing category corresponding to the sample image, and determine the third category loss value between the third clothing category and the preset standard clothing category.
[0053] In the embodiments of the present disclosure, determine the second category loss value and the third category loss value between the second clothing category and the third clothing category and the preset standard clothing category respectively. When the second category loss value is greater than the second preset classification loss threshold, and / or when the third category loss value is less than or equal to the third preset classification loss threshold, modify the model weight parameters of the initial classifier until the second category loss value is not greater than the second preset classification loss threshold and the third category loss value is greater than the third preset classification loss threshold, then obtain the trained classifier, that is, ensure that the classifier can distinguish the image features with similar image features.
[0054] In the embodiments of the present disclosure, if the classifier identifies a plurality of clothing categories, output the one with a high confidence level as the final output clothing category.
[0055] In summary, for the training method of the clothing category recognition model according to the embodiments of the present disclosure, sample images are obtained, and multiple pre-generated clothing category description text information is obtained. Then, the pre-constructed initial multi-modal pre-training model is trained based on the sample images and the multiple clothing category description text information to obtain a trained multi-modal pre-training model. Among them, the trained multi-modal pre-training model learns to output image features associated with clothing categories according to the input images. Multiple second sample image features of the sample images are extracted according to the multi-modal pre-training model, and multiple pseudo-sample image features corresponding to the multiple second sample image features are generated. Among them, the feature similarity between each pseudo-sample image feature and the sample image is less than the feature similarity between the second sample image feature corresponding to each pseudo-sample image feature and the sample image. Furthermore, the multiple second sample image features and the multiple pseudo-sample image features are respectively input into the pre-constructed initial classifier to obtain a second clothing category corresponding to the multiple second sample image features and a third clothing category corresponding to the multiple pseudo-sample image features. A second category loss value between the second clothing category and the preset standard clothing category corresponding to the sample image is determined, and a third category loss value between the third clothing category and the preset standard clothing category is determined. A trained classifier is obtained according to the second category loss value and the third category loss value. Among them, the trained classifier learns to output clothing categories according to the input image features. In this technical solution, the multi-modal pre-training model is trained based on the sample images and the clothing category description text information. When the multi-modal pre-training model is used to identify clothing categories, the contribution degree of the image features extracted by the multi-modal pre-training model to the determination of the image category helps to improve the determination accuracy of clothing categories. And on the basis of the trained multi-modal pre-training model, the classifier is further trained, and the classifier is trained in combination with the form of adversarial learning, which ensures the recognition accuracy of clothing categories.
[0056] Next, the clothing category recognition method according to the embodiments of the present disclosure will be described with reference to the embodiments. This method can be executed by a clothing category recognition device, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 4 shown, this method includes:
[0057] Step 401, extract at least one video frame from the target video stream.
[0058] In an embodiment of the present disclosure, at least one video frame in the video stream can be extracted through tools such as OpenCV. Among them, in some possible examples, the number of frames extracted can be preset, and at least one video frame is extracted from the target video stream according to the number of frames extracted; in some possible instances, the frame extraction frequency can also be set, for example, the number of frames extracted per second, etc., and at least one video frame is extracted from the target video stream according to the frame extraction frequency.
[0059] Step 402: Extract multiple target sample image features corresponding to each video frame through a pre-trained multi-modal pre-training model.
[0060] Among them, the pre-trained multi-modal pre-training model is trained by the training method of the above-mentioned clothing category recognition model.
[0061] In an embodiment of the present disclosure, multiple target sample image features corresponding to each video frame are extracted through a pre-trained multi-modal pre-training model, where the multiple target sample image features are image features related to clothing categories.
[0062] In an embodiment of the present disclosure, in order to avoid the influence of other non-clothing image contents, the target image region of the clothing included in each video frame can also be recognized. Furthermore, the second region coordinate position information of the target image region in the corresponding video frame can be determined. For example, the clothing region box coordinate detection can be performed through an open-source deep learning detector (such as the YOLO series). The detected second region coordinate position information can include ('cloth_id', 'top_left_x', 'top_left_y', 'bottom_right_x', 'bottom_right_y'), where the above five data respectively identify the clothing box number, the upper left coordinate x, the upper left coordinate y, the lower right coordinate x, and the lower right coordinate y.
[0063] Furthermore, in this embodiment, through the pre-trained multi-modal pre-training model, multiple target sample image features corresponding to each video frame are extracted according to the second region coordinate position information, that is, only the multiple target sample image features in the target image region containing clothing are extracted.
[0064] In an embodiment of the present disclosure, in order to avoid the problem of misrecognition, the confidence level of the clothing included in the target image region can also be recognized. Among them, the confidence level can also be recognized through the open-source deep learning detector in the above embodiment, etc. Furthermore, it is determined whether the confidence level of the target image region corresponding to each video frame is greater than a preset confidence threshold. The preset confidence threshold can be set according to the scenario requirements. For example, the preset confidence threshold can be 0.25, etc. In this embodiment, the video frames not greater than the preset confidence threshold are deleted, that is, when the preset confidence threshold is not greater than the preset confidence threshold, it is considered that the corresponding video frame may not contain clothing, and thus, the corresponding video frame is directly deleted.
[0065] Step 403: Input the multiple target sample image features into a pre-trained classifier to obtain the target clothing category output by the classifier.
[0066] Among them, the pre-trained classifier is trained by the training method of the above-mentioned clothing category recognition model.
[0067] In an embodiment of the present disclosure, multiple target sample image features are input into a pre-trained classifier to obtain the target clothing categories output by the classifier. The pre-trained classifier is trained based on adversarial learning and can better distinguish clothing categories with similar image features.
[0068] Step 404: Use all the target clothing categories corresponding to at least one video frame as the clothing category recognition result of the target video stream.
[0069] In this embodiment, the classifier outputs the corresponding target clothing categories for each video frame. In this embodiment, all the target clothing categories corresponding to at least one video frame can be used as the clothing category recognition result of the target video stream. For example, if all the target clothing categories corresponding to at least one video frame include "nurse uniform", "academic dress", and "Lolita", then "nurse uniform", "academic dress", and "Lolita" can be used as the clothing category recognition result.
[0070] Of course, considering that the classifier can recognize the category confidence of the target clothing categories, in an embodiment of the present disclosure, all the target clothing categories with a category confidence greater than a preset category confidence threshold can also be used as the clothing category recognition result.
[0071] In the actual execution process, after obtaining the clothing category recognition result, the clothing category recognition result can be provided in the form of a clothing category label, that is, the clothing category recognition result is output in the form of a clothing category label, or the clothing category recognition result can be output in the way of annotating the target clothing category in the video frame, etc.
[0072] In an embodiment of the present disclosure, after obtaining the clothing category recognition result, the face features corresponding to each target clothing category can also be recognized, and all the video frames corresponding to each target clothing category under the corresponding face features in the target video stream can be extracted (which can be extracted based on the target sample image features corresponding to the target clothing category combined with the face features). Thus, a video clip of a specific person wearing the same type of clothing can be generated based on all the video frames.
[0073] Of course, in other possible embodiments, after obtaining the clothing category recognition result, it can also be applied to other scenarios, which will not be listed one by one here.
[0074] In summary, for the clothing category recognition method in the embodiments of the present disclosure, at least one video frame in the target video stream is extracted, and multiple target sample image features corresponding to each video frame are extracted through a pre-trained multi-modal pre-training model. Furthermore, the multiple target sample image features are input into a pre-trained classifier to obtain the target clothing categories output by the classifier. In this technical solution, the recognition accuracy of clothing categories is improved.
[0075] To implement the above embodiments, the present disclosure also proposes a training device for a clothing category recognition model.
[0076] Figure 5 As shown in the structure diagram of a training device for a clothing category recognition model provided by an embodiment of the present disclosure, the device can be implemented by software and / or hardware and is generally integrated in an electronic device. Figure 5 As shown, the device includes: an acquisition module 510, a first training module 520, a first feature extraction module 530, and a second training module 540, where
[0077] The acquisition module 510 is configured to acquire sample images and acquire a plurality of pre-generated clothing category description text information.
[0078] The first training module 520 is configured to train a pre-constructed initial multi-modal pre-training model according to the sample images and the plurality of clothing category description text information to obtain a trained multi-modal pre-training model. Among them, the trained multi-modal pre-training model learns to output image features associated with clothing categories according to the input images.
[0079] The first feature extraction module 530 is configured to extract a plurality of second sample image features of the sample images according to the multi-modal pre-training model and generate a plurality of pseudo-sample image features corresponding one-to-one to the plurality of second sample image features. Among them, the feature similarity between each pseudo-sample image feature and the sample image is less than the feature similarity between the second sample image feature corresponding to each pseudo-sample image feature and the sample image.
[0080] The second training module 540 is configured to input the plurality of second sample image features and the plurality of pseudo-sample image features into a pre-constructed initial classifier respectively, obtain a second clothing category corresponding to the plurality of second sample image features and a third clothing category corresponding to the plurality of pseudo-sample image features, determine a second category loss value between the second clothing category and a preset standard clothing category corresponding to the sample image, and determine a third category loss value between the third clothing category and the preset standard clothing category. A trained classifier is obtained according to the second category loss value and the third category loss value. Among them, the trained classifier learns to output clothing categories according to the input image features. In an embodiment of the present disclosure, the first training module 520 is configured to:
[0081] Extract a plurality of first sample image features corresponding to the sample images and a plurality of clothing category text features corresponding one-to-one to the plurality of clothing category description text information through the initial multi-modal pre-training model;
[0082] Calculate the feature similarity between each first sample image feature and each clothing category text feature through the initial multi-modal pre-training model, and determine and output the first clothing category according to all the feature similarities through the initial multi-modal pre-training model;
[0083] Calculate the first category loss value between the first clothing category and the preset standard clothing category corresponding to the sample image, and determine whether the first category loss value is less than or equal to the first preset classification loss threshold;
[0084] When the first category loss value is greater than the first preset classification loss threshold, modify the model weight parameters of the initial multi-modal pre-training model until the first category loss value is less than or equal to the first preset classification loss threshold to obtain the trained multi-modal pre-training model.
[0085] In an embodiment of the present disclosure, the first training module 520 is configured to:
[0086] Determine the maximum feature similarity among all the feature similarities through the initial multi-modal pre-training model, and determine the first clothing category corresponding to the clothing category text feature corresponding to the maximum feature similarity;
[0087] Output the first clothing category through the initial multi-modal pre-training model.
[0088] In an embodiment of the present disclosure, the device further includes: a region determination module, configured to:
[0089] Identify the sample image region containing clothing in the sample image;
[0090] Determine the first region coordinate position information of the sample image region in the sample image;
[0091] The first training module 520 is configured to:
[0092] Extract a plurality of first sample image features corresponding to the sample image region according to the first region coordinate position information through the initial multi-modal pre-training model.
[0093] In an embodiment of the present disclosure, the first feature extraction module 530 is configured to:
[0094] Generate a plurality of pseudo-sample image features corresponding one-to-one with a plurality of second sample image features through a generative adversarial network.
[0095] In an embodiment of the present disclosure, the second training module 540 is configured to:
[0096] When the loss value of the second category is greater than the second preset classification loss threshold, and / or when the loss value of the third category is less than or equal to the third preset classification loss threshold, modify the model weight parameters of the initial classifier until the loss value of the second category is not greater than the second preset classification loss threshold and the loss value of the third category is greater than the third preset classification loss threshold, and then obtain the trained classifier.
[0097] In an embodiment of the present disclosure, it further includes: a pseudo-sample image feature generation module, configured to:
[0098] Calculate the classification loss value and the adversarial loss value of multiple pseudo-sample image features and multiple second-sample image features;
[0099] Determine the image feature loss value according to the classification loss value and the adversarial loss value;
[0100] Determine whether the image feature loss value is less than the preset feature loss threshold;
[0101] When the image feature loss value is not less than the preset feature loss threshold, generate multiple updated pseudo-sample image features corresponding one-to-one to the multiple second-sample image features through a generative adversarial network until it is determined that the image feature loss value is less than the preset feature loss threshold.
[0102] The training device for the clothing category recognition model provided by the embodiments of the present disclosure can execute the training method for the clothing category recognition model provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0103] To implement the above embodiments, the present disclosure also proposes a clothing category recognition device. This device can be implemented by software and / or hardware, and is generally integrated in an electronic device. As Figure 6 shown, this device includes: an extraction module 610, a second feature extraction module 620, an identification module 630, and a determination module 640, where
[0104] The extraction module 610 is configured to extract at least one video frame from the target video stream;
[0105] The second feature extraction module 620 is configured to extract multiple target sample image features corresponding to each video frame through a pre-trained multi-modal pre-trained model, where the pre-trained multi-modal pre-trained model is trained by the training method of the clothing category recognition model according to any one of claims 1-7;
[0106] The identification module 630 is configured to input the multiple target sample image features into a pre-trained classifier to obtain the target clothing category output by the classifier, where the pre-trained classifier is trained by the training method of the clothing category recognition model according to any one of claims 1-7;
[0107] A determination module 640 is configured to use all target clothing categories corresponding to at least one video frame as the clothing category recognition result of the target video stream.
[0108] In an embodiment of the present disclosure, the second feature extraction module 620 is specifically configured to:
[0109] Identify a target image region containing clothing in each video frame;
[0110] Determine the second region coordinate position information of the target image region in the corresponding video frame;
[0111] Extract multiple target sample image features corresponding to each video frame through a pre-trained multi-modal pre-training model according to the second region coordinate position information.
[0112] In an embodiment of the present disclosure, it further includes: a video frame screening module, configured to:
[0113] Identify the confidence level of the clothing contained in the target image region;
[0114] Determine whether the confidence level of the target image region corresponding to each video frame is greater than a preset confidence threshold;
[0115] Delete the video frames not greater than the preset confidence threshold.
[0116] The clothing category recognition device provided by the embodiments of the present disclosure can execute the clothing category recognition method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.
[0117] To implement the above embodiments, the present disclosure also proposes a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the training method of the clothing category recognition model and the clothing category recognition method in the above embodiments are implemented.
[0118] To implement the above embodiments, the present disclosure also proposes a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and the computer program is used to execute the training method of the clothing category recognition model and the clothing category recognition method.
[0119] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure.
[0120] Specifically refer to the following Figure 7, which shows a schematic structural diagram of an electronic device 700 suitable for implementing the embodiments of the present disclosure. The electronic device 700 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0121] As Figure 7 shown, the electronic device 700 may include a processor (such as a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the memory 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0122] Generally, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a memory 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 can allow the electronic device 700 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 the electronic device 700 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.
[0123] Specifically, according to the embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the memory 708, or installed from the ROM 702. When the computer program is executed by the processor 701, it executes the above-mentioned functions defined in the training method of the clothing category recognition model or the recognition method of clothing categories in the embodiments of the present disclosure.
[0124] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0125] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0126] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.
[0127] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the training method of the clothing category recognition model or the recognition method of clothing categories.
[0128] An electronic device can write computer program code for performing the operations of the present disclosure in one or more programming languages or combinations thereof. The above-mentioned programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and the module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0130] The units involved in the embodiments described in the present disclosure can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases.
[0131] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), and so on.
[0132] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0133] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0134] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0135] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A training method for a clothing category recognition model, characterized in that: include: Acquire a sample image, acquire a plurality of pre-generated clothing category description text information, and train a pre-constructed initial multimodal pre-trained model according to the sample image and the plurality of clothing category description text information to acquire a trained multimodal pre-trained model, wherein the trained multimodal pre-trained model learns to output image features associated with clothing categories according to the input image; Extracting a plurality of second sample image features of the sample image according to the multimodal pre-trained model, and generating a plurality of pseudo sample image features corresponding to the plurality of second sample image features one by one, wherein a feature similarity between each of the pseudo sample image features and the sample image is less than a feature similarity between a second sample image feature corresponding to each of the pseudo sample image features and the sample image; The multiple second sample image features and the multiple pseudo-sample image features are respectively input into a pre-constructed initial classifier, a second clothing category corresponding to the multiple second sample image features and a third clothing category corresponding to the multiple pseudo-sample image features are obtained, a second category loss value between the second clothing category and the preset standard clothing category corresponding to the sample image is determined, and a third category loss value between the third clothing category and the preset standard clothing category is determined, and a trained classifier is obtained according to the second category loss value and the third category loss value, wherein the trained classifier is trained and learned to output clothing categories according to the input image features.
2. The method according to claim 1, characterized in that The step of training a pre-built initial multimodal pre-training model according to the sample image and the plurality of clothing category description text information to obtain a trained multimodal pre-training model includes: Extracting a plurality of first sample image features corresponding to the sample image and a plurality of clothing category text features corresponding one-to-one to the plurality of clothing category description text information through the initial multimodal pre-training model; Calculating the feature similarity between each of the first sample image features and each of the clothing category text features respectively through the initial multimodal pre-trained model, and determining and outputting the first clothing category according to all of the feature similarities through the initial multimodal pre-trained model; Calculating a first category loss value of a preset standard clothing category corresponding to the first clothing category and the sample image, and determining whether the first category loss value is less than or equal to a first preset classification loss threshold; When the first category loss value is greater than the first preset classification loss threshold, the model weight parameters of the initial multimodal pre-trained model are modified until the trained multimodal pre-trained model is obtained when the first category loss value is less than or equal to the first preset classification loss threshold.
3. The method according to claim 2, characterized in that The determining and outputting a first clothing category according to all the feature similarities through the initial multimodal pre-training model includes: Determine the maximum feature similarity among all the feature similarities through the initial multimodal pre-training model, and determine the first clothing category corresponding to the clothing category text feature corresponding to the maximum feature similarity; The first clothing category is outputted through the initial multimodal pre-trained model.
4. The method according to claim 2, characterized in that Before extracting a plurality of first sample image features corresponding to the sample image through the initial multimodal pre-training model, the method further includes: Identifying a sample image region containing clothing in the sample image; Determine first region coordinate position information of the sample image region in the sample image; The extracting a plurality of first sample image features corresponding to the sample image through the initial multimodal pre-training model includes: The multiple first sample image features corresponding to the sample image area are extracted according to the first area coordinate position information through the initial multimodal pre-trained model.
5. The method according to claim 1, characterized in that The generating of a plurality of pseudo sample image features corresponding one-to-one to the plurality of second sample image features comprises: The plurality of pseudo sample image features corresponding one-to-one to the plurality of second sample image features are generated by a generative adversarial network.
6. The method according to claim 1 or 5, characterized in that The step of obtaining a trained classifier according to the second category loss value and the third category loss value includes: When the second category loss value is greater than the second preset classification loss threshold, and / or when the third category loss value is less than or equal to the third preset classification loss threshold, modify the model weight parameters of the initial classifier until the second category loss value is no greater than the second preset classification loss threshold and the third category loss value is greater than the third preset classification loss threshold, thereby obtaining the trained classifier.
7. The method according to claim 6, characterized in that Before inputting the plurality of pseudo sample image features into the initial classifier, the method further comprises: Calculating classification loss values and adversarial loss values of the plurality of pseudo sample image features and the plurality of second sample image features; Determine an image feature loss value according to the classification loss value and the adversarial loss value; Determining whether the image feature loss value is less than a preset feature loss threshold; When the image feature loss value is not less than the preset feature loss threshold, a plurality of updated pseudo sample image features corresponding one-to-one to the plurality of second sample image features are generated by the generative adversarial network until it is determined that the image feature loss value is less than the preset feature loss threshold.
8. A method for identifying clothing categories, characterized in that: include: Extracting at least one video frame from a target video stream; Extracting multiple target sample image features corresponding to each video frame through a pre-trained multimodal pre-training model, wherein the pre-trained multimodal pre-training model is trained by the training method of the clothing category recognition model according to any one of claims 1 to 7; Inputting the plurality of target sample image features into a pre-trained classifier to obtain a target clothing category output by the classifier, wherein the pre-trained classifier is trained by the training method of the clothing category recognition model according to any one of claims 1 to 7; All target clothing categories corresponding to the at least one video frame are used as clothing category recognition results of the target video stream.
9. The method according to claim 8, characterized in that The extracting a plurality of target sample image features corresponding to each video frame by a pre-trained multimodal pre-training model includes: Identifying a target image region containing clothing in each video frame; Determine second region coordinate position information of the target image region in the corresponding video frame; A plurality of target sample image features corresponding to each video frame are extracted based on the coordinate position information of the second region through a pre-trained multimodal pre-training model.
10. The method according to claim 9, characterized in that Before extracting a plurality of target sample image features corresponding to each video frame according to the second region coordinate position information through the pre-trained multimodal pre-trained model, the method further includes: Identifying the confidence that clothing is contained in the target image area; Determine whether the confidence of the target image area corresponding to each video frame is greater than a preset confidence threshold; The video frames whose confidence level is not greater than the preset confidence threshold are deleted.
11. A training device for a clothing category recognition model, characterized in that: include: An acquisition module, used to acquire a sample image and obtain pre-generated multiple clothing category description text information; A first training module is used to train a pre-built initial multimodal pre-training model according to the sample image and the multiple clothing category description text information to obtain a trained multimodal pre-training model, wherein the trained multimodal pre-training model learns to output image features associated with clothing categories according to the input image; a first feature extraction module, configured to extract a plurality of second sample image features of the sample image according to the multimodal pre-trained model, and generate a plurality of pseudo sample image features corresponding to the plurality of second sample image features one by one, wherein a feature similarity between each of the pseudo sample image features and the sample image is less than a feature similarity between a second sample image feature corresponding to each of the pseudo sample image features and the sample image; The second training module is used to input the multiple second sample image features and the multiple pseudo-sample image features into a pre-constructed initial classifier respectively, obtain a second clothing category corresponding to the multiple second sample image features and a third clothing category corresponding to the multiple pseudo-sample image features, determine a second category loss value between the second clothing category and a preset standard clothing category corresponding to the sample image, and determine a third category loss value between the third clothing category and the preset standard clothing category, and obtain a trained classifier according to the second category loss value and the third category loss value, wherein the trained classifier is trained and learned to output clothing categories according to the input image features.
12. A device for identifying clothing categories, characterized in that: include: An extraction module, used for extracting at least one video frame from a target video stream; A second feature extraction module, configured to extract a plurality of target sample image features corresponding to each video frame by using a pre-trained multimodal pre-training model, wherein the pre-trained multimodal pre-training model is trained by using the training method of the clothing category recognition model according to any one of claims 1 to 7; A recognition module, used for inputting the plurality of target sample image features into a pre-trained classifier to obtain a target clothing category output by the classifier, wherein the pre-trained classifier is trained by the training method of the clothing category recognition model according to any one of claims 1 to 7; The determination module is used to use all target clothing categories corresponding to the at least one video frame as clothing category recognition results of the target video stream.
13. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the executable instructions to implement the training method of the clothing category recognition model described in any one of claims 1-7 above, or the clothing category recognition method described in any one of claims 8-10.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is used to execute the training method of the clothing category recognition model described in any one of claims 1-7 above, or the clothing category recognition method described in any one of claims 8-10.