Clothing recognition method and device based on deep learning

Through the deep learning-based clothing recognition method, human body and upper body detection, combined with clothing recognition and classification models, the problem of interference in the lower body background in reflective clothing recognition is solved, and the recognition accuracy and accuracy are improved.

CN119206796BActive Publication Date: 2025-05-06UNIVERSAL UBIQUITOUS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411712859.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-05-06
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

When identifying reflective clothing, the prior art is susceptible to interference from the background information in the lower body area, resulting in insufficient recognition accuracy and accuracy.

Method used

The clothing recognition method based on deep learning is adopted to determine the upper body area of ​​the human body through human body detection and upper body detection. Combined with the clothing recognition model and classification model, feature vectors are extracted and wear state classification is performed, and finally the clothing state is determined through weighted sum.

Benefits of technology

It improves the accuracy and accuracy of clothing recognition, reduces interference from background information, and realizes real-time monitoring and accurate clothing status judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206796B_ABST
    Figure CN119206796B_ABST
Patent Text Reader

Abstract

This application provides a deep learning-based clothing recognition method and apparatus. The method includes: inputting a scene image into a preset human body detection model to perform human body detection, determining the corresponding human body region; expanding the human body region according to a set expansion ratio to determine the corresponding expanded human body region; inputting the expanded human body region into a preset upper body detection model to perform upper body detection, determining the corresponding upper body region; inputting the upper body region into a preset clothing recognition model to determine the corresponding similarity score; inputting the upper body region into a preset clothing classification model to determine the corresponding wearing status score; performing a weighted summation operation on the similarity score and the wearing status score to determine the corresponding comprehensive score; and determining the clothing status in the scene image based on the comprehensive score and a preset threshold. This application can improve the accuracy and precision of clothing recognition based on multi-model fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and specifically to a clothing recognition method and device based on deep learning. Background Art

[0002] In environments such as construction sites, road maintenance, police, firefighters, and rescue workers, employees often need to wear specific clothing to improve their visibility, especially reflective clothing. By identifying reflective clothing, we can understand the location of workers in real time and prevent possible safety risks. In an emergency, by identifying reflective clothing, we can improve emergency response efficiency.

[0003] At present, in order to avoid the situation where staff do not wear clothing in the work area, manual inspections are usually adopted. This method requires manpower costs and cannot achieve real-time monitoring. In recent years, methods based on surveillance cameras using artificial intelligence algorithms have also emerged. For example, a neural network detection model is used to detect pedestrians, and then a classification model is used to classify whether pedestrians are wearing reflective clothing. This method is low-cost and can achieve real-time monitoring and real-time alarms when there are problems.

[0004] However, most clothing, such as reflective clothing, is different from ordinary clothing mainly because of the difference in clothing on the upper body of the person, that is, the upper body area is the most important feature. Therefore, in the field of image recognition, only the upper body area needs to be considered. The previous existing solutions took the entire area of ​​the person as the research object, including the lower body area. This area does not have the characteristics of reflective clothing, but is easily mixed with background information, which will interfere with the results. Therefore, there is an urgent need for a clothing recognition method based on deep learning, which can improve the precision and accuracy of clothing recognition by specifically identifying upper body clothing. Summary of the invention

[0005] In response to the problems in the prior art, the present application provides a clothing recognition method and device based on deep learning, which can improve the precision and accuracy of clothing recognition based on multi-model fusion.

[0006] In order to solve at least one of the above problems, the present application provides the following technical solutions:

[0007] In a first aspect, the present application provides a clothing recognition method based on deep learning, comprising:

[0008] Acquire a live image captured by a camera in the working area in real time, input the live image into a preset human body detection model to perform a human body detection operation, determine a corresponding human body area, perform an outward expansion operation on the human body area according to a set outward expansion ratio, determine a corresponding expanded human body area, input the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determine a corresponding upper body area;

[0009] Input the upper body region into a set clothing recognition model to perform a feature vector extraction operation to determine a corresponding feature vector, compare the feature vector with a preset feature base library to determine a corresponding similarity score; input the upper body region into a set clothing classification model to perform a wearing state classification operation to determine a corresponding wearing state score, wherein the wearing state score includes a first wearing state score, a second wearing state score, and a third wearing state score, wherein the first wearing state score indicates a score for not being a piece of clothing, the second wearing state score indicates a score for being a piece of clothing but not being worn correctly, and the third wearing state score indicates a score for being a piece of clothing and being worn correctly;

[0010] A weighted sum operation is performed on the similarity score and the wearing status score to determine a corresponding comprehensive score, and the clothing status in the scene image is determined based on the comprehensive score and a preset threshold.

[0011] Further, performing an expansion operation on the human body region according to the set expansion ratio to determine a corresponding expanded human body region includes:

[0012] Performing an aspect ratio calculation operation on the human body region, and determining a corresponding outward expansion pixel range according to the aspect ratio of the human body region obtained after the calculation operation, the width value of the human body region, and a preset outward expansion ratio formula;

[0013] An expansion operation is performed on the human body region according to the expanded pixel range to determine a corresponding expanded human body region.

[0014] Furthermore, before inputting the upper body region into a set clothing recognition model to perform a feature vector extraction operation and determining a corresponding feature vector, the method includes:

[0015] Acquire upper body image data in different scenes, input the upper body image data into a RepVGG or ResNet backbone network model for model training, and determine corresponding training results, wherein the upper body image data is data consisting of multiple categories that classify clothes of the same style but different colors into one category;

[0016] The learning rate parameter is adjusted according to the optimal training result to determine the corresponding set clothing recognition model.

[0017] Furthermore, the comparing the feature vector with a preset feature base database to determine a corresponding similarity score includes:

[0018] Perform feature vector calculation operations on all samples in the preset feature base library to determine the corresponding feature vector samples;

[0019] A similarity calculation operation is performed on the feature vector and the feature vector sample, and a descending sorting operation is performed on the similarities obtained after the similarity calculation operation, and a corresponding similarity score is determined according to the highest similarity.

[0020] Furthermore, before inputting the upper body region into a set clothing classification model to perform a wearing state classification operation and determining a corresponding wearing state score, the method includes:

[0021] Performing a labeling operation on the upper body image data to determine a corresponding image data label, wherein the image data label includes at least one of not being a garment, being a garment but not being worn correctly, and being a garment and being worn correctly;

[0022] The image data with the image data label is input into the ResNet18 backbone network model for model training to obtain a set clothing classification model after the model training.

[0023] Furthermore, performing a weighted sum operation on the similarity score and the wearing status score to determine a corresponding comprehensive score includes:

[0024] determining whether a third wearing state score in the wearing state score is greater than the second wearing state score, and if so, performing a weighted sum operation according to the third wearing state score and the similarity score to determine a corresponding first comprehensive score;

[0025] If not, a weighted sum operation is performed according to the second wearing state score and the similarity score to determine a corresponding second comprehensive score.

[0026] Furthermore, determining the clothing status in the scene image according to the comprehensive score and a preset threshold value includes:

[0027] Determine whether the first comprehensive score is higher than a first threshold, if so, clothing exists in the scene image and is worn correctly, if not, clothing does not exist in the scene image;

[0028] It is determined whether the second comprehensive score is higher than a second threshold value. If so, clothing exists in the scene image but is not worn correctly. If not, clothing does not exist in the scene image.

[0029] In a second aspect, the present application provides a clothing recognition device based on deep learning, comprising:

[0030] An upper body detection module is used to obtain a real-time live image captured by a camera in the working area, input the live image into a preset human body detection model to perform a human body detection operation, determine a corresponding human body area, perform an outward expansion operation on the human body area according to a set outward expansion ratio, determine a corresponding expanded human body area, input the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determine a corresponding upper body area;

[0031] a clothing recognition module, for inputting the upper body region into a set clothing recognition model to perform a feature vector extraction operation, determining a corresponding feature vector, and performing a comparison operation on the feature vector and a preset feature base library to determine a corresponding similarity score; inputting the upper body region into a set clothing classification model to perform a wearing state classification operation, and determining a corresponding wearing state score, wherein the wearing state score includes a first wearing state score, a second wearing state score, and a third wearing state score, wherein the first wearing state score indicates a score for not being a clothing, the second wearing state score indicates a score for being a clothing but not being worn correctly, and the third wearing state score indicates a score for being a clothing and being worn correctly;

[0032] The clothing judgment module is used to perform a weighted sum operation on the similarity score and the wearing status score to determine the corresponding comprehensive score, and determine the clothing status in the scene image according to the comprehensive score and a preset threshold.

[0033] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the deep learning-based clothing recognition method when executing the program.

[0034] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the deep learning-based clothing recognition method.

[0035] In a fifth aspect, the present application provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the deep learning-based clothing recognition method.

[0036] It can be seen from the above technical scheme that the present application provides a clothing recognition method and device based on deep learning, which inputs a scene image into a preset human body detection model to perform a human body detection operation, determines the corresponding human body area, performs an outward expansion operation on the human body area according to a set outward expansion ratio, determines the corresponding expanded human body area, inputs the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determines the corresponding upper body area; inputs the upper body area into a set clothing recognition model to determine the corresponding similarity score; inputs the upper body area into a set clothing classification model to determine the corresponding wearing state score, performs a weighted sum operation on the similarity score and the wearing state score to determine the corresponding comprehensive score, and determines the clothing state in the scene image according to the comprehensive score and a preset threshold, thereby improving the precision and accuracy of clothing recognition based on multi-model fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0038] Figure 1 This is one of the flowcharts of the clothing recognition method based on deep learning in the embodiment of the present application;

[0039] Figure 2 This is a second flow chart of a clothing recognition method based on deep learning in an embodiment of the present application;

[0040] Figure 3 The third flowchart of the clothing recognition method based on deep learning in the embodiment of the present application;

[0041] Figure 4 This is a fourth flow chart of a clothing recognition method based on deep learning in an embodiment of the present application;

[0042] Figure 5 This is a fifth flow chart of a clothing recognition method based on deep learning in an embodiment of the present application;

[0043] Figure 6 This is a sixth flow chart of a clothing recognition method based on deep learning in an embodiment of the present application;

[0044] Figure 7 FIG7 is a flowchart of a clothing recognition method based on deep learning in an embodiment of the present application;

[0045] Figure 8is a structural diagram of a clothing recognition device based on deep learning in an embodiment of the present application;

[0046] Fig. 9 It is a schematic diagram of the structure of an electronic device in an embodiment of the present application.

[0047] Reference numerals:

[0048] Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0050] The acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.

[0051] Considering that the current workwear recognition scheme takes the entire human body as the research object, it is easy to be mixed with background information, which will interfere with the results. The present application provides a clothing recognition method and device based on deep learning, which inputs the scene image into a preset human body detection model to perform a human body detection operation, determines the corresponding human body area, performs an external expansion operation on the human body area according to a set external expansion ratio, determines the corresponding expanded human body area, inputs the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determines the corresponding upper body area; inputs the upper body area into a set clothing recognition model to determine the corresponding similarity score; inputs the upper body area into a set clothing classification model to determine the corresponding wearing state score, performs a weighted sum operation on the similarity score and the wearing state score to determine the corresponding comprehensive score, and determines the clothing state in the scene image according to the comprehensive score and a preset threshold, thereby improving the precision and accuracy of clothing recognition based on multi-model fusion.

[0052] In order to improve the precision and accuracy of clothing recognition based on multi-model fusion, the present application provides an embodiment of a clothing recognition method based on deep learning, see Figure 1, the clothing recognition method based on deep learning specifically includes the following contents:

[0053] Step S101: obtaining a real-time live image captured by a camera in the working area, inputting the live image into a preset human body detection model to perform a human body detection operation, determining a corresponding human body area, performing an outward expansion operation on the human body area according to a set outward expansion ratio, determining a corresponding expanded human body area, inputting the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determining a corresponding upper body area;

[0054] Optionally, in this embodiment, a human body detection model is preset, and a YOLO series model is used to detect human body areas in an image.

[0055] Human body detection model settings: The input image size is (1, 3, w1, h1), the loss function is the loss function corresponding to the YOLO series model, such as the CIOU function, etc., and the detection category is set to: human body.

[0056] Human body detection model training: Collect human body images in different scenes, angles, and brightness, annotate the human body areas in the images, and divide the large amount of annotated data into training sets, validation sets, and test sets. Adjust parameters such as the learning rate to train the model.

[0057] Optional, input image size (1, 3, w1, h1) for human detection model D1:

[0058] Batch size (1): Indicates that one image is processed at a time, which is conducive to real-time processing and suitable for scenarios that require fast response.

[0059] Color channel (3): corresponds to the three RGB color channels, retaining complete color information, which helps to improve detection accuracy.

[0060] Width and height (w1, h1): determine the size of the receptive field of the model. Larger w1 and h1 can provide more contextual information, but also increase computational complexity.

[0061] The uniform input size (w1, h1) helps with batch processing, improving the generalization ability of the model, and simplifying the network structure design.

[0062] Optionally, the CIOU (Complete Intersection over Union) function is a metric used to evaluate the overlap between the predicted bounding box and the true bounding box in object detection tasks. By considering the center point distance, CIOU can more accurately evaluate the position of the predicted box, making the predicted box closer to the shape of the true target. At the same time, CIOU loss can provide the network with more fine-grained gradient information, which helps the model converge faster.

[0063] Optionally, in this embodiment, an expansion operation is performed on the human body region according to a set expansion ratio to determine a corresponding expanded human body region.

[0064] Specifically, the human body region is adaptively expanded: the height is kept unchanged and the width is expanded by a certain ratio. Considering that the human body region is a rectangle, the aspect ratio of the training image is avoided to be too large, which affects the effect of the algorithm and improves the robustness of the algorithm.

[0065] For example, assuming that the width and height of the human body area are w and h respectively, the aspect ratio a=h / w, and the external expansion b satisfies:

[0066] 1. When a≧4, b=w / 3.5

[0067] 2. When 4>a≧3, b=w / 5

[0068] 3. When 3>a≧2.5, b=w / 6

[0069] 4. When a<2.5, b=w / 20

[0070] First, the aspect ratio of the human body area is calculated to obtain the range of a. Secondly, the calculation formula for the external expansion b is selected according to the range of a. Finally, the width value w is substituted into the calculation formula to obtain the b value.

[0071] It can be understood that the expansion process is centered on the human body area and expands b pixels to the left and right. After the expansion, the width and height of the human body area are w+2b and h respectively.

[0072] Optionally, in this embodiment, the expanded human body area is input into a preset upper body detection model to perform an upper body detection operation to determine a corresponding upper body area.

[0073] Specifically, an upper body detection model is preset, and the YOLO series model is selected to detect the upper body area in the human body area.

[0074] Upper body detection model settings: The input image size is (1, 3, w2, h2), the loss function is the loss function corresponding to the YOLO series model, such as the CIOU function, etc., and the detection category is set to upper body.

[0075] Upper body detection model training: The human body region in the human body detection model dataset is adaptively expanded and used as the dataset for the upper body detection model. The upper body region in the dataset is then labeled, that is, the image input to the upper body detection model is an image with a certain ratio of human body expanded outward. Adjust parameters such as the learning rate to train the model.

[0076] Step S102: input the upper body region into a set clothing recognition model to perform a feature vector extraction operation, determine a corresponding feature vector, compare the feature vector with a preset feature base library, and determine a corresponding similarity score; input the upper body region into a set clothing classification model to perform a wearing state classification operation, and determine a corresponding wearing state score, wherein the wearing state score includes a first wearing state score, a second wearing state score, and a third wearing state score, wherein the first wearing state score indicates a score for not being a clothing, the second wearing state score indicates a score for being a clothing but not being worn correctly, and the third wearing state score indicates a score for being a clothing and being worn correctly;

[0077] Optionally, in this embodiment, the upper body area obtained in step S101 is respectively input into a set clothing recognition model and a set clothing classification model, wherein the set clothing recognition model is used for binary classification to determine whether work clothes are worn, and the set clothing classification model is used for multi-classification to determine whether work clothes are worn correctly. Then, the scores output by the above two models can be obtained, laying the foundation for step S103.

[0078] Optionally, in this embodiment, the upper body area is input into a set clothing recognition model to perform a feature vector extraction operation to determine a corresponding feature vector.

[0079] Specifically, a clothing recognition model is set to identify whether workwear reflective clothing is worn, and the output is a binary classification result.

[0080] Clothing recognition model configuration: select RepVGG or ResNet as the backbone network model, select the recognition model loss function, such as ArcFace Loss, etc., the input image size is (1, 3, w3, h3), and the output feature vector dimension is (1,512).

[0081] Clothing recognition model training: The data set is upper body image data in different scenes. People wearing the same style (including different colors) of clothes are classified into one category. A large amount of data is obtained through web crawling and other methods and divided into training sets, validation sets, and test sets. Adjust parameters such as learning rate to train the model.

[0082] Optionally, backbone refers to the basic network structure used for feature extraction. Choosing RepVGG or ResNet as backbone means using these two different network architectures as the main feature extractors of the model.

[0083] RepVGG, a simple VGG-like structure, but uses reparameterization technology. It contains a multi-branch structure during training and will be reparameterized into a single convolution operation during inference. It has a large number of parameters during training and a small number of parameters during inference, and has high computational efficiency during inference.

[0084] ResNet, a deep network using residual connections, is relatively efficient in both training and reasoning, solves the gradient vanishing problem of deep networks, and is widely used in various computer vision tasks.

[0085] During training, consider the training task requirements, computing resources, accuracy requirements, and inference speed, and choose a suitable backbone network. If you pursue inference speed, RepVGG has more advantages; if you have higher accuracy requirements, ResNet can provide higher accuracy.

[0086] ArcFace Loss is optional. Its main function is to enhance the discriminative ability of features and improve recognition accuracy. It adds edge penalties in the angle space and operates in the angle space instead of the Euclidean space to increase the distance between classes, reduce the distance within classes, and improve the accuracy of various recognition tasks. For clothing categories with fewer samples, ArcFace Loss can help learn more discriminative features.

[0087] It can be understood that in the multi-model fusion clothing recognition system, using ArcFace Loss to train the recognition model can further improve the performance of the system and work together with other models to achieve more accurate clothing recognition and wearing judgment.

[0088] Optional, feature vector extraction, the input image size is (1, 3, w3, h3), and the output feature vector dimension is (1,512). This means that the model compresses the input image into a 512-dimensional feature representation. This reduces the computational complexity of subsequent processing.

[0089] Optionally, in this embodiment, the feature vector is compared with a preset feature base database to determine a corresponding similarity score.

[0090] Specifically, a feature base database is preset, and in order to improve the accuracy of recognition, the base database images collect image data of the front, back and side surfaces.

[0091] More specifically, first, the feature vectors of all samples in the base library are pre-calculated and stored. Secondly, the similarity between the input feature vector and each feature vector in the base library is calculated. The similarity calculation usually uses cosine similarity. Then, the calculated similarities are sorted in descending order. Finally, the highest similarity is selected as the final score s1.

[0092] Preferably, in order to improve efficiency, the following optimization strategies can be adopted:

[0093] Feature vector normalization: Normalizing all feature vectors in advance can simplify the calculation of cosine similarity.

[0094] Index structure: Use index structures such as KD-tree or Locality Sensitive Hashing (LSH) to speed up nearest neighbor searches.

[0095] Batch processing: Process multiple queries simultaneously, leveraging the parallel computing power of GPUs.

[0096] Finally, the obtained similarity score s1 can be used in the subsequent decision-making process.

[0097] Optionally, in this embodiment, the upper body area is input into a set clothing classification model to perform a wearing state classification operation to determine a corresponding wearing state score, wherein the wearing state score includes a first wearing state score, a second wearing state score, and a third wearing state score, wherein the first wearing state score indicates a score for not being a clothing item, the second wearing state score indicates a score for being a clothing item but not being worn correctly, and the third wearing state score indicates a score for being a clothing item and being worn correctly.

[0098] Specifically, a clothing classification model is set to determine whether the workwear reflective clothing is worn correctly, and the output is a multi-classification result.

[0099] Set the clothing classification model configuration: select ResNet18 as the network model backbone, select the classification model loss function as loss, such as the cross entropy loss function, the input image size is (1, 3, w3, h3), the labels are 0 (not clothing), 1 (clothing but not worn correctly) and 2 (clothing and worn correctly), and the output dimension is (1,3).

[0100] Set up clothing classification model training: The data set is upper body image data in different scenes. The same as the data set for clothing recognition model, the labeling categories are three categories: not clothing, clothing but not worn correctly, clothing and worn correctly, and the data is divided into training set, validation set and test set. Adjust the learning rate and other parameters to train the model.

[0101] Optionally, after training the above-mentioned set clothing classification model, the upper body area is input into the set clothing classification model, and a softmax operation is performed on the wearing state result output by the set clothing classification model. The original output is converted into a probability distribution to obtain a wearing state score [s20, s21, s22], where s20, s21, and s22 are respectively the scores of the three categories: not clothing, clothing but not worn correctly, and clothing and worn correctly.

[0102] Step S103: performing a weighted sum operation on the similarity score and the wearing status score to determine a corresponding comprehensive score, and determining the clothing status in the scene image according to the comprehensive score and a preset threshold.

[0103] Optionally, in this embodiment, the similarity score s1 output by the clothing recognition model set in step S102 and the wearing status scores s20, s21, s22 output by the clothing classification model set have been obtained, wherein s21 is the score for the "incorrect wearing" category and s22 is the score for the "correct wearing" category.

[0104] Optionally, in this embodiment, a weighted sum operation is performed on the similarity score and the wearing status score to determine a corresponding comprehensive score.

[0105] Specifically, the weight of the clothing recognition model is set to w, and the weight of the clothing classification model is set to 1-w.

[0106] When s22>s21, the probability of "it is clothing and worn correctly" is greater than the probability of "it is clothing but not worn correctly". The first comprehensive score S1=w s1+(1-w) s22.

[0107] When s21>s22, the probability of "it is clothing but not worn correctly" is greater than the probability of "it is clothing and worn correctly". The second comprehensive score is S2=w s1+(1-w) s21.

[0108] Optionally, in this embodiment, the clothing status in the scene image is determined based on the comprehensive score and a preset threshold.

[0109] If the first comprehensive score S1 is greater than the threshold a, the reflective clothing is worn correctly; if the first comprehensive score S1 is less than the threshold a, the reflective clothing is not worn;

[0110] If the second comprehensive score S2 is greater than the threshold value b, the reflective clothing is not worn correctly; if the second comprehensive score S2 is less than the threshold value b, the reflective clothing is not worn.

[0111] Optionally, the weight w, threshold a, and threshold b are all traversed on [0,1] at intervals of 0.05, and the corresponding value when the test accuracy is the highest is the actual usage value.

[0112] It can be understood that S1 and S2 are obtained by weighting the scores of the recognition model and the classification model, providing complementary information, reflecting the possibility of correct wearing and incorrect wearing, respectively.

[0113] This example demonstrates how this embodiment uses two algorithms in series to accurately locate the upper body area of ​​the human body, and then uses a combination of a recognition model and a classification model based on the upper body area to determine whether the clothing is worn correctly.

[0114] Optionally, the complete solution process of this embodiment is as follows:

[0115] 1. The camera in the working area captures the on-site images in real time;

[0116] 2. Input the above image into the trained human body detection model D1. If no human body area is detected, the process ends. If a human body area is detected, the process proceeds to the next step.

[0117] 3. Expand the human body area detected in the previous step on the original image according to the following formula;

[0118] When a≧4, b=w / 3.5

[0119] When 4>a≧3, b=w / 5

[0120] When 3>a≧2.5, b=w / 6

[0121] When a<2.5, b=w / 20

[0122] 4. Input the human body data obtained after the previous step of expansion into the human upper body detection model D2. If the upper body area is not detected, the process ends. If the upper body area is detected, the process proceeds to the next step.

[0123] 5. Input the upper body area detected in the previous step into the clothing recognition model R, extract the feature vector, and compare the feature vector with the feature vector of the base database to obtain the top1 score s1;

[0124] 6. In the previous step, in order to improve the accuracy of recognition, the base library images should collect image data from the front, back and sides.

[0125] 7. Input the upper body data detected in step 4 into the clothing classification model C. Model C has three categories of outputs, namely 0 (not clothing), 1 (clothing but not worn correctly), and 2 (work clothes and worn correctly);

[0126] 8. Perform softmax operation on the result of model C to obtain [s20, s21, s22], where s20, s21, and s22 are the scores of the three categories respectively;

[0127] 9. The score weight of the recognition model R is w, the score weight of the classification model is 1-w, and the comprehensive score is:

[0128] When s22>s21, S1=w s1+(1-w) s22

[0129] When s21>s22, S2=w s1+(1-w) s21

[0130] If S1 is greater than the threshold a, the reflective clothing is worn correctly;

[0131] If S1 is less than the threshold a, the reflective clothing is not worn;

[0132] If S2 is greater than the threshold value b, the reflective clothing is not worn correctly;

[0133] If S2 is less than the threshold value b, the person is not wearing reflective clothing.

[0134] The weight w, threshold a and threshold b are all traversed on [0,1] with an interval of 0.05, and the corresponding value when the test accuracy is the highest is the actual usage value.

[0135] From the above description, it can be seen that the deep learning-based clothing recognition method provided in the embodiment of the present application can determine the corresponding human body area by inputting the scene image into a preset human body detection model for human body detection operation, and perform an outward expansion operation on the human body area according to the set outward expansion ratio to determine the corresponding expanded human body area, and input the expanded human body area into a preset upper body detection model for upper body detection operation to determine the corresponding upper body area; input the upper body area into the set clothing recognition model to determine the corresponding similarity score; input the upper body area into the set clothing classification model to determine the corresponding wearing state score, perform a weighted sum operation on the similarity score and the wearing state score to determine the corresponding comprehensive score, and determine the clothing state in the scene image according to the comprehensive score and the preset threshold, thereby improving the precision and accuracy of clothing recognition based on multi-model fusion.

[0136] In one embodiment of the clothing recognition method based on deep learning of the present application, see Figure 2 , and can also include the following:

[0137] Step S201: performing an aspect ratio calculation operation on the human body region, and determining a corresponding outward expansion pixel range according to the aspect ratio of the human body region obtained after the calculation operation, the width value of the human body region, and a preset outward expansion ratio formula;

[0138] Step S202: performing an expansion operation on the human body region according to the expansion pixel range to determine a corresponding expanded human body region.

[0139] Optionally, in this embodiment, the human body region is adaptively expanded: the height is kept unchanged and the width is expanded by a certain ratio. Considering that the human body region is a rectangle, the aspect ratio of the training image is avoided to be too large, which affects the effect of the algorithm and improves the robustness of the algorithm.

[0140] For example, assuming that the width and height of the human body area are w and h respectively, the aspect ratio a=h / w, and the external expansion b satisfies:

[0141] 1. When a≧4, b=w / 3.5

[0142] 2. When 4>a≧3, b=w / 5

[0143] 3. When 3>a≧2.5, b=w / 6

[0144] 4. When a<2.5, b=w / 20

[0145] First, the aspect ratio of the human body area is calculated to obtain the range of a. Secondly, the calculation formula for the external expansion b is selected according to the range of a. Finally, the width value w is substituted into the calculation formula to obtain the b value.

[0146] It can be understood that the expansion process is centered on the human body area and expands b pixels to the left and right. After the expansion, the width and height of the human body area are w+2b and h respectively.

[0147] Through step S202, this embodiment successfully adaptively expands the human body recognition area, laying a foundation for subsequent accurate clothing recognition and classification.

[0148] In one embodiment of the clothing recognition method based on deep learning of the present application, see Figure 3 , and can also include the following:

[0149] Step S301: obtaining upper body image data in different scenes, inputting the upper body image data into a RepVGG or ResNet backbone network model for model training, and determining corresponding training results, wherein the upper body image data is data consisting of multiple categories that classify clothes of the same style but different colors into one category;

[0150] Step S302: adjusting the learning rate parameter according to the optimal training result to determine the corresponding set clothing recognition model.

[0151] Optionally, in this embodiment, the clothing recognition model is configured as follows: the backbone network model (backbone) selects RepVGG or ResNet, the loss selects the recognition model loss function, such as ArcFace Loss, etc., the input image size is (1, 3, w3, h3), and the output feature vector dimension is (1,512).

[0152] Clothing recognition model training: The data set is upper body image data in different scenes. People wearing the same style (including different colors) of clothes are classified into one category. A large amount of data is obtained through web crawling and other methods and divided into training sets, validation sets, and test sets. Adjust parameters such as learning rate to train the model.

[0153] Optionally, backbone refers to the basic network structure used for feature extraction. Choosing RepVGG or ResNet as backbone means using these two different network architectures as the main feature extractors of the model.

[0154] RepVGG, a simple VGG-like structure, but uses reparameterization technology. It contains a multi-branch structure during training and will be reparameterized into a single convolution operation during inference. It has a large number of parameters during training and a small number of parameters during inference, and has high computational efficiency during inference.

[0155] ResNet, a deep network using residual connections, is relatively efficient in both training and reasoning, solves the gradient vanishing problem of deep networks, and is widely used in various computer vision tasks.

[0156] During training, consider the training task requirements, computing resources, accuracy requirements, and inference speed, and choose a suitable backbone network. If you pursue inference speed, RepVGG has more advantages; if you have higher accuracy requirements, ResNet can provide higher accuracy.

[0157] ArcFace Loss is optional. Its main function is to enhance the discriminative ability of features and improve recognition accuracy. It adds edge penalties in the angle space and operates in the angle space instead of the Euclidean space to increase the distance between classes, reduce the distance within classes, and improve the accuracy of various recognition tasks. For clothing categories with fewer samples, ArcFace Loss can help learn more discriminative features.

[0158] It can be understood that in the multi-model fusion clothing recognition system, using ArcFace Loss to train the recognition model can further improve the performance of the system and work together with other models to achieve more accurate clothing recognition and wearing judgment.

[0159] Through step S302, this embodiment trains the clothing recognition model, laying a foundation for subsequent joint identification with the clothing classification model.

[0160] In one embodiment of the clothing recognition method based on deep learning of the present application, see Figure 4 , and can also include the following:

[0161] Step S401: performing feature vector calculation operations on all samples in a preset feature base database to determine corresponding feature vector samples;

[0162] Step S402: performing a similarity calculation operation on the feature vector and the feature vector sample, performing a descending sorting operation on the similarities obtained after the similarity calculation operation, and determining a corresponding similarity score according to the highest similarity.

[0163] Optionally, in this embodiment, a feature base library is preset, and in order to improve the accuracy of recognition, the base library images collect image data of the front, back and sides.

[0164] Specifically, first, the feature vectors of all samples in the base database are pre-calculated and stored. Secondly, the similarity between the input feature vector and each feature vector in the base database is calculated. The similarity calculation usually uses cosine similarity. Then, the calculated similarities are sorted in descending order. Finally, the highest similarity is selected as the final score s1.

[0165] Preferably, in order to improve efficiency, the following optimization strategies can be adopted:

[0166] Feature vector normalization: Normalizing all feature vectors in advance can simplify the calculation of cosine similarity.

[0167] Index structure: Use index structures such as KD-tree or Locality Sensitive Hashing (LSH) to speed up nearest neighbor searches.

[0168] Batch processing: Process multiple queries simultaneously, leveraging the parallel computing power of GPUs.

[0169] Finally, the obtained similarity score s1 can be used in the subsequent decision-making process.

[0170] Through step S402, this embodiment successfully performs feature extraction and feature vector similarity calculation through the clothing recognition model, laying a foundation for subsequent clothing discrimination.

[0171] In one embodiment of the clothing recognition method based on deep learning of the present application, see Figure 5 , and can also include the following:

[0172] Step S501: performing a labeling operation on the upper body image data to determine a corresponding image data label, wherein the image data label includes at least one of not being a garment, being a garment but not being worn correctly, and being a garment and being worn correctly;

[0173] Step S502: input the image data with the image data label into the ResNet18 backbone network model for model training, and obtain a set clothing classification model after the model training.

[0174] Optionally, in this embodiment, the clothing classification model configuration is set: the network model backbone selects ResNet18, the loss selects the classification model loss function, such as the cross entropy loss function, the input image size is (1, 3, w3, h3), the labels are 0 (not clothing), 1 (clothing but not worn correctly) and 2 (clothing and worn correctly), and the output dimension is (1,3).

[0175] Set up clothing classification model training: The data set is upper body image data in different scenes. The same as the data set for clothing recognition model, the labeling categories are three categories: not clothing, clothing but not worn correctly, clothing and worn correctly, and the data is divided into training set, validation set and test set. Adjust the learning rate and other parameters to train the model.

[0176] It can be understood that after the above-mentioned set clothing classification model is trained, the upper body area is input into the set clothing classification model, and the wearing state result output by the set clothing classification model is subjected to a softmax operation. The original output is converted into a probability distribution to obtain a wearing state score [s20, s21, s22], where s20, s21, and s22 are the scores of the three categories: not clothing, clothing but not worn correctly, and clothing and worn correctly.

[0177] Through step S502, this embodiment successfully trains the set clothing classification model, laying a foundation for the subsequent joint clothing judgment of the two models.

[0178] In one embodiment of the clothing recognition method based on deep learning of the present application, see Figure 6 , and can also include the following:

[0179] Step S601: determining whether the third wearing state score in the wearing state score is greater than the second wearing state score, and if so, performing a weighted sum operation based on the third wearing state score and the similarity score to determine a corresponding first comprehensive score;

[0180] Step S602: If not, a weighted sum operation is performed according to the second wearing state score and the similarity score to determine a corresponding second comprehensive score.

[0181] Optionally, in this embodiment, the first wearing state score s20, the second wearing state score s21, and the third wearing state score s22 output by the clothing classification model are set, wherein s21 is the score for the "incorrect wearing" category, and s22 is the score for the "correct wearing" category.

[0182] Specifically, the weight of the clothing recognition model is set to w, and the weight of the clothing classification model is set to 1-w.

[0183] When s22>s21, the probability of "it is clothing and worn correctly" is greater than the probability of "it is clothing but not worn correctly". The first comprehensive score S1=w s1+(1-w) s22.

[0184] When s21>s22, the probability of "it is clothing but not worn correctly" is greater than the probability of "it is clothing and worn correctly". The second comprehensive score is S2=w s1+(1-w) s21.

[0185] It can be understood that in the above formula, the first comprehensive score S1 and the second comprehensive score S2 are obtained by weighting the scores of the recognition model and the classification model, providing complementary information, reflecting the possibility of correct wearing and incorrect wearing respectively.

[0186] Through step S602, this embodiment successfully obtains the comprehensive score of the two models, which lays a foundation for subsequent judgment of whether the clothing is worn and whether it is worn correctly.

[0187] In one embodiment of the clothing recognition method based on deep learning of the present application, see Figure 7 , and can also include the following:

[0188] Step S701: determining whether the first comprehensive score is higher than a first threshold, if so, clothing exists in the scene image and is worn correctly, if not, clothing does not exist in the scene image;

[0189] Step S702: Determine whether the second comprehensive score is higher than a second threshold value. If so, clothing exists in the scene image but is not worn correctly. If not, clothing does not exist in the scene image.

[0190] Optional, the judgment rules are as follows:

[0191] If the first comprehensive score S1 is greater than the threshold a, the reflective clothing is worn correctly; if the first comprehensive score S1 is less than the threshold a, the reflective clothing is not worn;

[0192] If the second comprehensive score S2 is greater than the threshold value b, the reflective clothing is not worn correctly; if the second comprehensive score S2 is less than the threshold value b, the reflective clothing is not worn.

[0193] Optionally, the weight w, threshold a, and threshold b are all traversed on [0,1] at intervals of 0.05, and the corresponding value when the test accuracy is the highest is the actual usage value.

[0194] It can be understood that S1 and S2 are obtained by weighting the scores of the recognition model and the classification model, providing complementary information, reflecting the possibility of correct wearing and incorrect wearing, respectively.

[0195] Through step S702, this embodiment successfully makes a judgment based on the comprehensive score, obtains a clothing judgment result, and improves the accuracy and precision of clothing recognition.

[0196] In order to improve the precision and accuracy of clothing recognition based on multi-model fusion, the present application provides an embodiment of a clothing recognition device based on deep learning for implementing all or part of the content of the clothing recognition method based on deep learning, see Figure 8 , the clothing recognition device based on deep learning specifically includes the following contents:

[0197] The form feature extraction module 10 is used to obtain a real-time live image captured by a camera in the working area, input the live image into a preset human body detection model to perform a human body detection operation, determine a corresponding human body area, perform an outward expansion operation on the human body area according to a set outward expansion ratio, determine a corresponding expanded human body area, input the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determine a corresponding upper body area;

[0198] The model training module 20 is used to input the upper body area into a set clothing recognition model to perform a feature vector extraction operation, determine a corresponding feature vector, compare the feature vector with a preset feature base library, and determine a corresponding similarity score; input the upper body area into a set clothing classification model to perform a wearing state classification operation, and determine a corresponding wearing state score, wherein the wearing state score includes a first wearing state score, a second wearing state score, and a third wearing state score, wherein the first wearing state score indicates a score for not being a clothing item, the second wearing state score indicates a score for being a clothing item but not being worn correctly, and the third wearing state score indicates a score for being a clothing item and being worn correctly;

[0199] The prediction and maintenance optimization module 30 is used to perform a weighted sum operation on the similarity score and the wearing status score to determine the corresponding comprehensive score, and determine the clothing status in the scene image based on the comprehensive score and a preset threshold.

[0200] From the above description, it can be seen that the deep learning-based clothing recognition device provided in the embodiment of the present application can input the scene image into a preset human body detection model to perform a human body detection operation, determine the corresponding human body area, expand the human body area according to the set expansion ratio, determine the corresponding expanded human body area, input the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determine the corresponding upper body area; input the upper body area into the set clothing recognition model to determine the corresponding similarity score; input the upper body area into the set clothing classification model to determine the corresponding wearing state score, perform a weighted sum operation on the similarity score and the wearing state score to determine the corresponding comprehensive score, and determine the clothing state in the scene image according to the comprehensive score and the preset threshold, thereby improving the precision and accuracy of clothing recognition based on multi-model fusion.

[0201] From the hardware level, in order to improve the precision and accuracy of clothing recognition based on multi-model fusion, the present application provides an embodiment of an electronic device for implementing all or part of the content of the clothing recognition method based on deep learning, and the electronic device specifically includes the following content:

[0202] Processor, memory, communication interface and bus; wherein the processor, memory and communication interface communicate with each other through the bus; the communication interface is used to realize the information transmission between the clothing recognition method based on deep learning and the core business system, user terminal and related database and other related devices; the logic controller can be a desktop computer, a tablet computer and a mobile terminal, etc., but the present embodiment is not limited thereto. In the present embodiment, the logic controller can be implemented with reference to the embodiment of the clothing recognition method based on deep learning and the embodiment of the clothing recognition method based on deep learning in the embodiment, and the contents thereof are incorporated herein, and the repeated parts are not repeated.

[0203] It is understandable that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.

[0204] In practical applications, part of the clothing recognition method based on deep learning can be executed on the electronic device side as described above, or all operations can be completed in the client device. The specific selection can be based on the processing capability of the client device and the limitations of the user's usage scenario. This application does not limit this. If all operations are completed in the client device, the client device may also include a processor.

[0205] The client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and other implementation scenarios may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster consisting of multiple servers, or a server structure of a distributed device.

[0206] Fig. 9 FIG. 9 is a schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. Fig. 9As shown, the electronic device 9600 may include a central processor 9100 and a memory 9140; the memory 9140 is coupled to the central processor 9100. It is worth noting that Fig. 9 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.

[0207] In one embodiment, the clothing recognition method function based on deep learning can be integrated into the central processing unit 9100. The central processing unit 9100 can be configured to perform the following control:

[0208] Step S101: obtaining a real-time live image captured by a camera in the working area, inputting the live image into a preset human body detection model to perform a human body detection operation, determining a corresponding human body area, performing an outward expansion operation on the human body area according to a set outward expansion ratio, determining a corresponding expanded human body area, inputting the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determining a corresponding upper body area;

[0209] Step S102: input the upper body region into a set clothing recognition model to perform a feature vector extraction operation, determine a corresponding feature vector, compare the feature vector with a preset feature base library, and determine a corresponding similarity score; input the upper body region into a set clothing classification model to perform a wearing state classification operation, and determine a corresponding wearing state score, wherein the wearing state score includes a first wearing state score, a second wearing state score, and a third wearing state score, wherein the first wearing state score indicates a score for not being a clothing, the second wearing state score indicates a score for being a clothing but not being worn correctly, and the third wearing state score indicates a score for being a clothing and being worn correctly;

[0210] Step S103: performing a weighted sum operation on the similarity score and the wearing status score to determine a corresponding comprehensive score, and determining the clothing status in the scene image according to the comprehensive score and a preset threshold.

[0211] From the above description, it can be seen that the electronic device provided in the embodiment of the present application inputs the scene image into a preset human body detection model to perform a human body detection operation, determines the corresponding human body area, performs an outward expansion operation on the human body area according to a set outward expansion ratio, determines the corresponding expanded human body area, inputs the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determines the corresponding upper body area; inputs the upper body area into a set clothing recognition model to determine the corresponding similarity score; inputs the upper body area into a set clothing classification model to determine the corresponding wearing state score, performs a weighted sum operation on the similarity score and the wearing state score to determine the corresponding comprehensive score, and determines the clothing state in the scene image according to the comprehensive score and a preset threshold, thereby improving the precision and accuracy of clothing recognition based on multi-model fusion.

[0212] In another embodiment, the clothing recognition method based on deep learning can be configured separately from the central processing unit 9100. For example, the clothing recognition method based on deep learning can be configured as a chip connected to the central processing unit 9100, and the function of the clothing recognition method based on deep learning can be realized through the control of the central processing unit.

[0213] like Fig. 9 As shown, the electronic device 9600 may also include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily have to include Fig. 9 In addition, the electronic device 9600 may also include Fig. 9 For components not shown, reference may be made to the prior art.

[0214] like Fig. 9 As shown, the central processing unit 9100 is sometimes also referred to as a controller or an operation control, and may include a microprocessor or other processor device and / or logic device, which receives input and controls the operation of various components of the electronic device 9600.

[0215] The memory 9140 may be, for example, one or more of a cache, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory or other suitable devices. The above-mentioned information related to the failure may be stored, and a program for executing the relevant information may also be stored. The CPU 9100 may execute the program stored in the memory 9140 to implement information storage or processing, etc.

[0216] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to provide power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display may be, for example, an LCD display, but is not limited thereto.

[0217] The memory 9140 may be a solid-state memory, such as a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It may also be a memory that saves information even when the power is off, can be selectively erased, and is provided with more data, examples of which are sometimes referred to as EPROMs, etc. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142, which is used to store application programs and function programs or processes for executing the operation of the electronic device 9600 through the central processor 9100.

[0218] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).

[0219] The communication module 9110 is a transmitter / receiver that sends and receives signals via the antenna 9111. The communication module 9110 is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as the case of a conventional mobile communication terminal.

[0220] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module and / or a wireless local area network module, etc. The communication module 9110 is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, thereby realizing a common telecommunication function. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor 9130 is also coupled to the central processor 9100, so that recording can be performed on the machine through the microphone 9132, and the sound stored on the machine can be played through the speaker 9131.

[0221] The embodiments of the present application also provide a computer-readable storage medium capable of implementing all the steps of the clothing recognition method based on deep learning in the above embodiments, where the execution subject is a server or a client. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, all the steps of the clothing recognition method based on deep learning in the above embodiments are implemented. For example, when the processor executes the computer program, the following steps are implemented:

[0222] Step S101: obtaining a real-time live image captured by a camera in the working area, inputting the live image into a preset human body detection model to perform a human body detection operation, determining a corresponding human body area, performing an outward expansion operation on the human body area according to a set outward expansion ratio, determining a corresponding expanded human body area, inputting the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determining a corresponding upper body area;

[0223] Step S102: input the upper body region into a set clothing recognition model to perform a feature vector extraction operation, determine a corresponding feature vector, compare the feature vector with a preset feature base library, and determine a corresponding similarity score; input the upper body region into a set clothing classification model to perform a wearing state classification operation, and determine a corresponding wearing state score, wherein the wearing state score includes a first wearing state score, a second wearing state score, and a third wearing state score, wherein the first wearing state score indicates a score for not being a clothing, the second wearing state score indicates a score for being a clothing but not being worn correctly, and the third wearing state score indicates a score for being a clothing and being worn correctly;

[0224] Step S103: performing a weighted sum operation on the similarity score and the wearing status score to determine a corresponding comprehensive score, and determining the clothing status in the scene image according to the comprehensive score and a preset threshold.

[0225] From the above description, it can be seen that the computer-readable storage medium provided in the embodiment of the present application inputs the scene image into a preset human body detection model to perform a human body detection operation, determines the corresponding human body area, performs an outward expansion operation on the human body area according to a set outward expansion ratio, determines the corresponding expanded human body area, inputs the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determines the corresponding upper body area; inputs the upper body area into a set clothing recognition model to determine the corresponding similarity score; inputs the upper body area into a set clothing classification model to determine the corresponding wearing state score, performs a weighted sum operation on the similarity score and the wearing state score to determine the corresponding comprehensive score, and determines the clothing state in the scene image according to the comprehensive score and a preset threshold, thereby improving the precision and accuracy of clothing recognition based on multi-model fusion.

[0226] The embodiments of the present application also provide a computer program product capable of implementing all the steps of the clothing recognition method based on deep learning in the above embodiments, where the execution subject is a server or a client. When the computer program / instruction is executed by a processor, the steps of the clothing recognition method based on deep learning are implemented. For example, the computer program / instruction implements the following steps:

[0227] Step S101: obtaining a real-time live image captured by a camera in the working area, inputting the live image into a preset human body detection model to perform a human body detection operation, determining a corresponding human body area, performing an outward expansion operation on the human body area according to a set outward expansion ratio, determining a corresponding expanded human body area, inputting the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determining a corresponding upper body area;

[0228] Step S102: input the upper body region into a set clothing recognition model to perform a feature vector extraction operation, determine a corresponding feature vector, compare the feature vector with a preset feature base library, and determine a corresponding similarity score; input the upper body region into a set clothing classification model to perform a wearing state classification operation, and determine a corresponding wearing state score, wherein the wearing state score includes a first wearing state score, a second wearing state score, and a third wearing state score, wherein the first wearing state score indicates a score for not being a clothing, the second wearing state score indicates a score for being a clothing but not being worn correctly, and the third wearing state score indicates a score for being a clothing and being worn correctly;

[0229] Step S103: performing a weighted sum operation on the similarity score and the wearing status score to determine a corresponding comprehensive score, and determining the clothing status in the scene image according to the comprehensive score and a preset threshold.

[0230] From the above description, it can be seen that the computer program product provided in the embodiment of the present application inputs the scene image into a preset human body detection model to perform a human body detection operation, determines the corresponding human body area, performs an outward expansion operation on the human body area according to a set outward expansion ratio, determines the corresponding expanded human body area, inputs the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determines the corresponding upper body area; inputs the upper body area into a set clothing recognition model to determine the corresponding similarity score; inputs the upper body area into a set clothing classification model to determine the corresponding wearing state score, performs a weighted sum operation on the similarity score and the wearing state score to determine the corresponding comprehensive score, and determines the clothing state in the scene image according to the comprehensive score and a preset threshold, thereby improving the precision and accuracy of clothing recognition based on multi-model fusion.

[0231] It should be understood by those skilled in the art that embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0232] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0233] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0234] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0235] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A tool identification method based on deep learning, characterized in that: The method comprises: Acquire a live image captured by a camera in the working area in real time, input the live image into a preset human body detection model to perform a human body detection operation, determine a corresponding human body area, perform an outward expansion operation on the human body area according to a set outward expansion ratio, determine a corresponding expanded human body area, input the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determine a corresponding upper body area; Input the upper body region into a set clothing recognition model to perform a feature vector extraction operation to determine a corresponding feature vector, compare the feature vector with a preset feature base library to determine a corresponding similarity score; input the upper body region into a set clothing classification model to perform a wearing state classification operation to determine a corresponding wearing state score, wherein the wearing state score includes a first wearing state score, a second wearing state score, and a third wearing state score, wherein the first wearing state score indicates a score for not being a piece of clothing, the second wearing state score indicates a score for being a piece of clothing but not being worn correctly, and the third wearing state score indicates a score for being a piece of clothing and being worn correctly; A weighted sum operation is performed on the similarity score and the wearing status score to determine a corresponding comprehensive score, and the clothing status in the scene image is determined based on the comprehensive score and a preset threshold.

2. The tooling identification method based on deep learning according to claim 1, characterized in that: The step of performing an expansion operation on the human body region according to the set expansion ratio to determine a corresponding expanded human body region includes: Performing an aspect ratio calculation operation on the human body region, and determining a corresponding outward expansion pixel range according to the aspect ratio of the human body region obtained after the calculation operation, the width value of the human body region, and a preset outward expansion ratio formula; An expansion operation is performed on the human body region according to the expanded pixel range to determine a corresponding expanded human body region.

3. The tooling identification method based on deep learning according to claim 1, characterized in that: Before inputting the upper body region into a set clothing recognition model to perform a feature vector extraction operation and determining a corresponding feature vector, the method includes: Acquire upper body image data in different scenes, input the upper body image data into a RepVGG or ResNet backbone network model for model training, and determine corresponding training results, wherein the upper body image data is data consisting of multiple categories that classify clothes of the same style but different colors into one category; The learning rate parameter is adjusted according to the optimal training result to determine the corresponding set clothing recognition model.

4. The tooling identification method based on deep learning according to claim 1, characterized in that: The comparing the feature vector with a preset feature base database to determine a corresponding similarity score includes: Perform feature vector calculation operations on all samples in the preset feature base library to determine the corresponding feature vector samples; A similarity calculation operation is performed on the feature vector and the feature vector sample, and a descending sorting operation is performed on the similarities obtained after the similarity calculation operation, and a corresponding similarity score is determined according to the highest similarity.

5. The tooling identification method based on deep learning according to claim 3, characterized in that: Before inputting the upper body region into a set clothing classification model to perform a wearing state classification operation and determining a corresponding wearing state score, the method includes: Performing a labeling operation on the upper body image data to determine a corresponding image data label, wherein the image data label includes at least one of not being a garment, being a garment but not being worn correctly, and being a garment and being worn correctly; The image data with the image data label is input into the ResNet18 backbone network model for model training to obtain a set clothing classification model after the model training.

6. The tooling identification method based on deep learning according to claim 1, characterized in that: The performing a weighted sum operation on the similarity score and the wearing status score to determine a corresponding comprehensive score includes: determining whether a third wearing state score in the wearing state score is greater than the second wearing state score, and if so, performing a weighted sum operation according to the third wearing state score and the similarity score to determine a corresponding first comprehensive score; If not, a weighted sum operation is performed according to the second wearing state score and the similarity score to determine a corresponding second comprehensive score.

7. The tooling identification method based on deep learning according to claim 6, characterized in that: The step of determining the clothing state in the scene image according to the comprehensive score and a preset threshold value includes: Determine whether the first comprehensive score is higher than a first threshold, if so, clothing exists in the scene image and is worn correctly, if not, clothing does not exist in the scene image; It is determined whether the second comprehensive score is higher than a second threshold value. If so, clothing exists in the scene image but is not worn correctly. If not, clothing does not exist in the scene image.

8. A tool identification device based on deep learning, characterized in that: The device comprises: An upper body detection module is used to obtain a live image captured by a camera in a working area in real time, input the live image into a preset human body detection model to perform a human body detection operation, determine a corresponding human body area, perform an outward expansion operation on the human body area according to a set outward expansion ratio, determine a corresponding expanded human body area, input the expanded human body area into a preset upper body detection model to perform an upper body detection operation, and determine a corresponding upper body area; a clothing recognition module, for inputting the upper body region into a set clothing recognition model to perform a feature vector extraction operation, determining a corresponding feature vector, and performing a comparison operation on the feature vector and a preset feature base library to determine a corresponding similarity score; inputting the upper body region into a set clothing classification model to perform a wearing state classification operation, and determining a corresponding wearing state score, wherein the wearing state score includes a first wearing state score, a second wearing state score, and a third wearing state score, wherein the first wearing state score indicates a score for not being a clothing, the second wearing state score indicates a score for being a clothing but not being worn correctly, and the third wearing state score indicates a score for being a clothing and being worn correctly; The clothing judgment module is used to perform a weighted sum operation on the similarity score and the wearing status score to determine the corresponding comprehensive score, and determine the clothing status in the scene image according to the comprehensive score and a preset threshold.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the tool recognition method based on deep learning described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the tool recognition method based on deep learning described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Wearable specification identification method and device, medium and electronic equipment

    CN115984897A

  • Non-standard wearing detection method and device based on deep learning and computer equipment

    CN117197580A