A method for video recognition of multi-person clothing features suitable for complex scenarios

By using the improved SE-InceptionV4 network and CPN/ResNet50 network cascade method in complex scenarios, the problem of low accuracy of clothing feature recognition in multi-person scenarios is solved, and a higher accuracy of pedestrian detection and clothing feature recognition is achieved.

CN114821477BActive Publication Date: 2025-06-27NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210481470.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-05
Publication Date
2025-06-27
Estimated Expiration
2042-05-05

AI Technical Summary

Technical Problem

In complex scenarios, the existing clothing feature recognition network cannot accurately identify multi-person clothing features, especially in multi-person scenarios. The accuracy of the existing technology is low, resulting in low manual search efficiency and high cost.

Method used

The InceptionV4 network, which incorporates the improved SE module, is used as the backbone network of the SSD, to build a pedestrian detection network, and combines the CPN network and the ResNet50 network to build a key point detection network and a clothing feature recognition network, and realizes multi-person clothing feature video recognition through cascading networks.

Benefits of technology

It improves the accuracy of pedestrian detection, simplifies the key point detection network, and enhances the accuracy of clothing feature recognition. Compared with the network directly classifying clothing, the accuracy rate is increased by 10%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821477B_ABST
    Figure CN114821477B_ABST
Patent Text Reader

Abstract

A method for video recognition of multi-person clothing features suitable for complex scenarios, including a cascaded pedestrian detection network, a key point detection network, and a clothing feature recognition network. The pedestrian detection network outputs the coordinates of the pedestrian detection box. After the key point detection network reads the coordinates of the pedestrian detection box, it outputs the key point coordinates. After the clothing feature recognition network reads the key point coordinates, it outputs the length and color of the clothing. The present invention uses the SE-InceptionV4 network as the backbone network of SSD, improving the accuracy of pedestrian detection. An improved SE module is proposed, which can extract more representative features in each channel feature map. The SE-InceptionV4 network can simultaneously take into account the extraction of more effective features in space and channels. The present invention specifically intercepts and identifies the clothing features from the pictures at the human key points related to the recognition task, avoiding the influence of complex clothing types on feature extraction and improving the accuracy compared with the network directly for classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent video detection, and relates to the recognition and positioning of pedestrians in videos and the recognition of pedestrian clothing features, and is a video recognition method for pedestrian clothing features suitable for complex scenarios. Background Art

[0002] Clothing is the most obvious and representative feature of each pedestrian. Using computer image processing to recognize the clothing features of the target person can filter out most of the useless information in the surveillance videos, and has great value in many practical application scenarios. Recognizing the clothing features of pedestrians in videos based on deep learning can effectively improve the search efficiency. Most of the existing clothing feature recognition networks directly adopt classification networks and directly put the pictures containing pedestrians into the classification networks for training. However, due to the rich variety of clothing types and the complex and diverse clothing combinations on pedestrians, such methods cannot accurately recognize the clothing features, especially in the multi-person scenario, and the existing classification networks cannot adapt to the recognition and marking of multiple targets. Summary of the Invention

[0003] The problems to be solved by the present invention are: when searching for the target person in various complex scenarios, manual search takes a long time and costs a lot; common clothing feature recognition networks cannot be applied to complex multi-person scenarios and have low accuracy.

[0004] The technical solution of the present invention is: a video recognition method for multi-person clothing features suitable for complex scenarios, including the following steps:

[0005] step1: Construct a pedestrian data set, and label the pedestrian bounding box, pedestrian key points, and pedestrian clothing features, including the lengths and colors of the upper and lower clothes of the pedestrian;

[0006] step2: Use InceptionV4 incorporating an improved SE module as the backbone network of SSD to build a pedestrian detection network;

[0007] step3: Use the pedestrian data set to train the pedestrian detection network in step2.

[0008] step4: Use the CPN network without the RefinelNet part to build a key point detection network;

[0009] step5: Use the pedestrian data set to train the key point detection network;

[0010] step6: Build a clothing feature recognition network based on the ResNet50 network, and judge the color recognition according to the values in the HSV space of the picture;

[0011] Step 7: Read the pedestrian clothing feature dataset and train the clothing feature recognition network;

[0012] Step 8: After training, cascade each network to obtain a multi-person clothing feature video recognition and detection network. For the input video or image, the pedestrian detection network outputs the coordinates of the pedestrian detection box. After the key point detection network reads the coordinates of the pedestrian detection box, it outputs the key point coordinates. After the clothing feature recognition network reads the key point coordinates, it outputs the length and color of the clothing.

[0013] Further, the detection network in step step2 is specifically:

[0014] Based on the InceptionV4 network, construct an InceptionV4 network integrated with an improved SE module, called the SE-InceptionV4 network. The InceptionV4 network includes a stem module, an Inception-A module group, an Inception-B module group, an Inception-C module group, a Reduction-A module, and a Reduction-B module. The feature map output by the Inception-A module group is numbered A1, the feature map output by the Inception-B module group is numbered B1, and the feature map output by the Inception-C module group is numbered C1. The improved SE module is integrated after the Inception-A module group, the Inception-B module group, and the Inception-C module group. The improved SE module includes a Max poling layer, a Global poling layer, a fully connected layer, a ReLu activation layer, a fully connected layer, and a Sigmoid activation layer in sequence. The size selected by the Max poling layer varies according to the position where the channel attention module is added. Specifically as follows:

[0015] For the Inception-A module group, add an improved SE module branch A, specifically add a 3*3 Max poling layer, a Global poling layer, a 1*1*24 fully connected layer, a ReLu activation layer, a 1*1*384 fully connected layer, and a Sigmoid activation layer in sequence. The feature map numbered A1 passes through branch A to obtain a 1*1*384 feature map, numbered A2. After multiplying the feature values of each channel of the feature map numbered A1 by the feature values of the corresponding channels of the feature map numbered A2, it is then sent to the subsequent convolutional layer of the Inception-A module group;

[0016] For the Inception-B module group, an improved SE module branch B is added. Specifically, a 2*2 Maxpoling layer, a Global poling layer, a fully connected layer of 1*1*64, a ReLu activation layer, a fully connected layer of 1*1*1024, and a Sigmoid activation layer are added in sequence. The feature map numbered B1 passes through branch B to obtain a feature map of 1*1*1024, numbered B2. After multiplying the feature values of each channel of the feature map numbered B1 by the corresponding channel feature values of the feature map numbered B2, it is then fed into the subsequent convolutional layer of the Inception-B module group;

[0017] For the Inception-C module group, an SE module branch is added. Specifically, a Global poling layer, a fully connected layer of 1*1*96, a ReLu activation layer, a fully connected layer of 1*1*1536, and a Sigmoid activation layer are added in sequence. The feature map numbered C1 passes through the SE module branch to obtain a feature map of 1*1*1536, numbered C2. After multiplying the feature values of each channel of the feature map numbered C1 by the corresponding channel feature values of the feature map numbered C2, it is then fed into the subsequent convolutional layer of the Inception-C module group;

[0018] The SE-InceptionV4 network is used as the feature extraction network and as the backbone network of SSD; the feature maps A1×A2, B1×B2, C1×C2 after fusing the SE module, and the features generated by the conv9, conv10, conv11 of the SSD network Figure 1 are output to the prediction network of SSD to output the prediction results and obtain the coordinates of the detection box.

[0019] Furthermore, step step4 is specifically as follows:

[0020] step4.1: Build a key point detection network based on the CPN network, extract features using the ResNet network as the backbone network, then detect key points by GlobalNet, remove the original RefineNet part of the CPN network, and output the key point coordinates;

[0021] step4.2: During training, only take the sum of the losses of the key points required by the subsequent clothing feature recognition network for gradient backpropagation.

[0022] Furthermore, step step6 is specifically as follows:

[0023] Step 6.1: Build a clothing feature recognition network based on the ResNet50 network. Remove the last softmax layer of the ResNet50 network, change the output dimension of the fully connected layer to 512, denoted as FC1. After the fully connected layer FC1, add a set of parallel fully connected layers FC2. Each fully connected layer classifies one of the attributes of the clothing. According to the key point coordinates output by the key point detection network, intercept the pictures at the key points, splice the intercepted pictures, and then send them to the ResNet50 network to recognize the clothing features; Step 6.2 Color recognition, convert the intercepted pictures from the RGB space to the HSV space, judge the color of each pixel point according to the H, S, and V values of the color, and select the color with the most pixel points as the color of the clothing in the picture.

[0024] There are also some clothing detection methods in the prior art. Some detection networks first use the object detection network to detect the clothing area, and then classify the clothing in the detected area. However, the overall performance of this solution has a great relationship with the dataset annotation carried out in the early stage. How to annotate a suitable and correct dataset is time-consuming and laborious, and it is still affected by the types of clothing.

[0025] Compared with the prior art, the present invention has the following advantages:

[0026] First, the present invention uses the SE-InceptionV4 network as the backbone network of SSD, which improves the accuracy of pedestrian detection. The improved SE module of the present invention can extract more representative features in each channel feature map. After adding the SE module to each Inception module of the InceptionV4 network, the network can simultaneously take into account extracting more effective features in space and channels.

[0027] Second, the present invention simplifies the human key point detection network. In view of the fact that the subsequent network only needs to intercept the surrounding pictures based on the key points and the accuracy requirement is not high, the subsequent RefineNet is removed, and the loss of some key points is specifically selected for gradient backpropagation.

[0028] Third, the present invention specifically intercepts the pictures at the human key points related to the recognition task, splices them and then recognizes the clothing features, avoiding the influence of complex clothing types on feature extraction. Using the FashiobAI dataset for testing, the accuracy of the present invention is improved by 10% compared with the network directly classifying the clothing. Brief Description of the Drawings

[0029] Figure 1 It is a flow schematic diagram of the present invention.

[0030] Figure 2 It is a structural diagram of the detection network of the present invention.

[0031] Figure 3 is the structural diagram of each module in the InceptionV4 network. (a) is the Inception-A module, (b) is the Inception-B module, (c) is the Inception-C module, (d) is the Reduction-A module, and (e) is the Reduction-B module.

[0032] Figure 4 This is the structural diagram of the clothing feature recognition network of the present invention.

[0033] Figure 5 This is the schematic diagram of the recognition accuracy result of skirt length in the embodiment of the present invention.

[0034] Figure 6 This is the schematic diagram of the recognition accuracy result of pants length in the embodiment of the present invention.

[0035] Figure 7 This is the schematic diagram of the recognition accuracy result of sleeve length in the embodiment of the present invention. Detailed implementation manners

[0036] The present invention proposes a method for video recognition of multi-person clothing features suitable for complex scenarios. The present invention solves the problems of low efficiency and high cost in finding target persons by manually watching videos in various complex scenarios; the present invention has strong practicability and is applicable to a variety of monitoring scenarios.

[0037] As Figure 1 shown, the specific implementation process of the present invention is as follows:

[0038] step1: Construct a pedestrian dataset, and label pedestrian boxes, pedestrian key points, and pedestrian clothing features: the lengths and colors of the upper and lower clothes of pedestrians.

[0039] step2: Use InceptionV4 integrated with an improved SE module as the backbone network of SSD to build a pedestrian detection network.

[0040] The present invention improves the SE (Squeeze-and-Excitation) module and integrates the improved SE module into the existing InceptionV4 network to construct an SE-InceptionV4 network as the backbone network of the SSD network, realizing pedestrian detection in complex scenarios. The contributions of the features of each channel to the extraction of key information in the next layer vary. The improved SE module can effectively learn the relationship between channels, assign different weights to different feature channels, and thus improve the network's feature extraction ability. The InceptionV4 network includes a stem module, an Inception-A module group, an Inception-B module group, an Inception-C module group, a Reduction-A module, and a Reduction-B module. After passing through 4 Inception-A modules, a feature map with a size of 35*35 and 384 channels is obtained, numbered A1; after passing through 7 Inception-B modules, a feature map with a size of 17*17 and 1024 channels is obtained, numbered B1; after passing through 3 Inception-C modules, a feature map with a size of 8*8 and 1536 channels is obtained, numbered C1. The structural diagrams of each module in the InceptionV4 network are shown in Figure 3. The InceptionV4 network is prior art, and the stem module, Inception-A module, Inception-B module, Inception-C module, Reduction-A module, and Reduction-B module in it, as well as the data relationship between them, will not be described in detail.

[0041] The present invention improves the InceptionV4 network, such as Figure 2As shown in the figure, an improved SE module is incorporated into the InceptionV4 network. The improved SE module sequentially includes a Max poling layer, a Global poling layer, a fully connected layer, a ReLu activation layer, a fully connected layer, and a Sigmoid activation layer. The size selected by the Max poling layer varies according to the position where the channel attention module is placed. The improved SE module first extracts more representative features in each feature channel through the Max poling layer, and then compresses each two-dimensional feature channel into a real number through the Global poling layer. This real number represents the global receptive field to some extent. Then, it first reduces the dimension through a fully connected layer to limit the model complexity and reduce the number of network parameters, and then increases the non-linearity of the module through the ReLu activation layer. Finally, a one-dimensional vector with the same number of channels as the original is obtained through a fully connected layer, and the Sigmoid function is used to limit the vector value to 0-1 as the weight for each channel. The weights output by the SE module are used as the learned channel-level weight relationship, and the original feature map is weighted channel by channel through multiplication to complete the change of channel weights. The specific steps are as follows.

[0042] For the Inception-A module group, an improved SE module branch A is added. Branch A specifically consists of a sequentially connected 3*3 Max poling layer, a Global poling layer, a 1*1*24 fully connected layer, a ReLu activation layer, a 1*1*384 fully connected layer, and a Sigmoid activation layer. The feature map numbered A1 passes through branch A to obtain a 1*1*384 feature map, numbered A2. After multiplying the feature values of each channel of the feature map numbered A1 by the corresponding channel feature values of the feature map numbered A2, it is then sent to the subsequent convolutional layer of the Inception-A module group. For example, Figure 2 the scale in [reference] represents the fusion operation of A1 and A2. The fused feature map is sent to the Reduction module and conv9 of the original SSD network as shown by the arrow.

[0043] For the Inception-B module group, an improved SE module branch B is added. Branch B specifically consists of a sequentially connected 2*2 Max poling layer, a Global poling layer, a 1*1*64 fully connected layer, a ReLu activation layer, a 1*1*1024 fully connected layer, and a Sigmoid activation layer. The feature map numbered B1 passes through branch B to obtain a 1*1*1024 feature map, numbered B2. After multiplying the feature values of each channel of the feature map numbered B1 by the corresponding channel feature values of the feature map numbered B2, it is then sent to the subsequent convolutional layer of the Inception-B module group, as Figure 2 shown.

[0044] For the Inception-C module group, an SE module branch is added. Specifically, a Global pooling layer, a fully connected layer of 1*1*96, a ReLu activation layer, a fully connected layer of 1*1*1536, and a Sigmoid activation layer are added in sequence. The feature map numbered C1 passes through this branch to obtain a feature map of 1*1*1536, numbered C2. After multiplying the feature values of each channel of the feature map numbered C1 by the corresponding channel feature values of the feature map numbered C2, it is then sent to the subsequent convolutional layer of the Inception-C module group, as Figure 2 shown.

[0045] Essentially, the SE module redefines the relationship between each feature channel through learning, performs feature recalibration on different feature channels, strengthens useful information and weakens irrelevant information. In the SE module, global average pooling operation is used to assign the same weight to each position of each channel feature map, which strengthens unimportant information and suppresses important information to some extent.

[0046] In the present invention, the SE module is improved. In order to extract important feature information of each feature channel, first, a Max pooling layer is used to extract better features with stronger semantic information for each feature channel, and then global pooling is performed on the feature map obtained by the Max pooling layer. However, considering that the Inception-A, Inception-B, and Inception-C modules in the InceptionV4 network have different depths and feature map sizes in the entire network, the receptive field of the feature map in the shallow layer of the network is smaller, and the receptive field of the feature map in the deep layer is larger. Therefore, Max pooling layers with different strides are selected for the three Inception modules. The Inception-A module is located in the shallow layer of the network, and the size of the feature map is 35*35. At this time, the receptive field of the feature map is not large, so a Max pooling layer of 3*3 size is selected. The Inception-B module is located in the middle layer of the network and uses a 2*2 Max pooling layer. The Inception-C module is located in the deep layer of the network, and the size of the feature map is 7*7. At this time, the feature map already has more global and higher-level semantic information and is no longer suitable for using the Max pooling layer. Therefore, an SE module is selected here instead of the improved SE module.

[0047] The present invention uses an improved SE-InceptionV4 network to replace the VGG16 network in the existing SSD, and removes conv6, conv7, and conv8 in the SSD. And the feature maps A1×A2, B1×B2, C1×C2 after fusing the attention mechanism are combined with the features generated by conv9, conv10, and conv11 in the SSD network Figure 1Output to the original prediction network of the SSD, output the prediction result, and obtain the coordinates of the detection box.

[0048] Step 3: Use the pedestrian dataset to train the pedestrian detection network.

[0049] Step 4: Build a human keypoint detection network based on the CPN network. The CPN body consists of two parts: GlobalNet and RefineNet. In the present invention, the RefineNet part is removed, only the GlobalNet part is used, and the loss function is modified as follows:

[0050] Step 4.1: Build a keypoint detection network based on the CPN network, use the ResNet network as the backbone network to extract features, and GlobalNet detects keypoints, removing the original RefineNet part.

[0051] Step 4.2: During training, only take the sum of the losses of the keypoints required by the subsequent clothing feature recognition network for gradient backpropagation.

[0052] Step 5: Use the pedestrian dataset to train the keypoint detection network.

[0053] Step 6: Build a clothing feature recognition network based on the ResNet50 network. Color recognition is judged according to the values in the HSV space of the picture. Specifically:

[0054] Step 6.1: The present invention also improves the ResNet50 network. Build a network based on the ResNet50 network, remove the last softmax layer, change the output dimension of the fully connected layer to 512, denoted as FC1. After the fully connected layer FC1, add a set of parallel fully connected layers FC2. Each fully connected layer FC2 corresponds to one of the attributes of the clothing for classification. When classifying clothing, according to the keypoint coordinates output by the keypoint detection network, intercept the picture at the keypoints, splice the intercepted pictures, and then send them to the improved ResNet50 network to recognize the clothing features. For example, the sleeve length attribute is classified using a fully connected layer FC2: long sleeves, medium sleeves, short sleeves, etc. The specific structure is as Figure 4 , and the purpose of designing the parallel fully connected layer group is to obtain the result by classifying all the attributes of the clothing at one time, without classifying each attribute separately, reducing the calculation cost and accelerating the speed of the entire network.

[0055] Step 6.2: In addition to clothing classification, color recognition is also performed. Convert the intercepted picture from the RGB space to the HSV space, and judge the color of each pixel point according to the H, S, V values of various colors, such as black, white, gray, red, orange, yellow, green, blue, purple, etc. Select the color with the most pixel points as the color of the clothing in the picture.

[0056] Step 7: Read the pedestrian clothing feature dataset, intercept the pictures at the key points of the left and right elbows and knees, splice them and use them as the input of the clothing feature recognition network to train the clothing feature recognition network.

[0057] Step 8: After the training is completed, cascade each network. Specifically:

[0058] Step 8.1: Pedestrian detection network: Input: video or picture, Output: pedestrian detection box coordinates;

[0059] Step 8.2: Key point detection network: Input: pedestrian detection box coordinates, Output: key point coordinates;

[0060] Step 8.3: Clothing feature recognition network: Input: key point coordinates, Output: clothing length and color.

[0061] The cascading of the three networks realizes the video recognition of the clothing features of multiple people in complex scenarios.

[0062] The present invention conducts simulation experiments. The dataset includes the collected pedestrian images and some public datasets, such as the datasets for detecting pedestrians: caltech, INRIA; the key point datasets: COCO, MPII, and the clothing classification dataset: fashionAI. Figure 5 、 Figure 6 and Figure 7 They are the accuracy results for skirt length, trouser length, and sleeve length respectively. The ordinate is the accuracy, and the abscissa is the number of training times for the entire dataset. It can be seen that as the number of training times increases, the accuracy of the present invention can reach 97-98%. On this basis, combined with the clothing color, the present invention can effectively achieve the accurate recognition of the diverse clothing features of people in pictures or videos of complex multi-person scenarios.

Claims

1. A method for video recognition of multi-person clothing features suitable for complex scenarios, characterized in that It includes the following steps: Step 1: Construct a pedestrian dataset, and annotate pedestrian bounding boxes, pedestrian key points, and pedestrian clothing features, including the lengths and colors of the upper and lower clothes of pedestrians; Step 2: Use InceptionV4 incorporated with an improved SE module as the backbone network of SSD to build a pedestrian detection network. The specific detection network is as follows: Based on the InceptionV4 network, construct an InceptionV4 network incorporated with an improved SE module, called the SE-InceptionV4 network. The InceptionV4 network includes a stem module, an Inception-A module group, an Inception-B module group, an Inception-C module group, a Reduction-A module, and a Reduction-B module. Number the feature maps output by the Inception-A module group as A1, the feature maps output by the Inception-B module group as B1, and the feature maps output by the Inception-C module group as C1. The improved SE module is incorporated after the Inception-A module group, the Inception-B module group, and the Inception-C module group in the SE-InceptionV4 network. The improved SE module sequentially includes a Max poling layer, a Global poling layer, a fully connected layer, a ReLu activation layer, a fully connected layer, and a Sigmoid activation layer. The size selected by the Max poling layer varies according to the position where the channel attention module is added. Specifically as follows: For the Inception-A module group, add an improved SE module branch A, specifically add a 3*3 Max poling layer, a Global poling layer, a 1*1*24 fully connected layer, a ReLu activation layer, a 1*1*384 fully connected layer, and a Sigmoid activation layer in sequence. The feature map numbered A1 passes through branch A to obtain a 1*1*384 feature map, numbered A2. After multiplying the feature values of each channel of the feature map numbered A1 by the feature values of the corresponding channels of the feature map numbered A2, it is then sent to the subsequent convolutional layer of the Inception-A module group; For the Inception-B module group, add an improved SE module branch B, specifically add a 2*2 Max poling layer, a Global poling layer, a 1*1*64 fully connected layer, a ReLu activation layer, a 1*1*1024 fully connected layer, and a Sigmoid activation layer in sequence. The feature map numbered B1 passes through branch B to obtain a 1*1*1024 feature map, numbered B2. After multiplying the feature values of each channel of the feature map numbered B1 by the feature values of the corresponding channels of the feature map numbered B2, it is then sent to the subsequent convolutional layer of the Inception-B module group; For the Inception-C module group, add the SE module branch, specifically add the Global poling layer, 1*1*96 fully connected layer, ReLu activation layer, 1*1*1536 fully connected layer and Sigmoid activation layer in sequence. The feature map numbered C1 passes through the SE module branch to obtain a 1*1*1536 feature map numbered C2. The feature values ​​of each channel of the feature map numbered C1 are multiplied by the feature values ​​of the corresponding channels of the feature map numbered C2, and then sent to the subsequent convolutional layers of the Inception-C module group; The SE-InceptionV4 network is used as the feature extraction network and the backbone network of SSD. The feature maps A1×A2, B1×B2, C1×C2 after the fusion of the SE module are output to the prediction network of SSD together with the feature maps generated by the SSD network conv9, conv10, and conv11. The prediction results are output to obtain the detection box coordinates. Step 3: Use the pedestrian dataset to train the pedestrian detection network in step 2; Step 4: Use the CPN network without the RefinelNet part to build a key point detection network; Step 5: Use the pedestrian dataset to train the key point detection network; Step 6: Build a clothing feature recognition network based on the ResNet50 network, where color recognition is determined based on the value of the HSV space of the image; Step 7: Read the pedestrian clothing feature dataset and train the clothing feature recognition network; Step 8: After training, cascade each network to obtain a multi-person clothing feature video recognition and detection network. For the input video or picture, the pedestrian detection network outputs the coordinates of the pedestrian detection frame, the key point detection network reads the coordinates of the pedestrian detection frame and outputs the key point coordinates, and the clothing feature recognition network reads the key point coordinates and outputs the length and color of the clothing.

2. The method for identifying video features of multi-person clothing suitable for complex scenarios according to claim 1, wherein Step 4 is as follows: Step 4.1: Build a key point detection network based on the CPN network, use the ResNet network as the backbone network to extract features, then use GlobalNet to detect key points, remove the original RefineNet part of the CPN network, and output the key point coordinates; Step 4.2: During training, only the sum of the losses of the key points required by the subsequent clothing feature recognition network is taken for gradient backpropagation.

3. The method for identifying multi-person clothing feature videos suitable for complex scenarios according to claim 1, characterized in that Step 6 is as follows: Step 6.1: Build a clothing feature recognition network based on the ResNet50 network, remove the last softmax layer of the ResNet50 network, change the output dimension of the fully connected layer to 512, denoted as FC1, and add a set of parallel fully connected layers FC2 after the fully connected layer FC1. Each fully connected layer classifies one of the attributes of the clothing. According to the key point coordinates output by the key point detection network, the images at the key points are intercepted, the intercepted images are spliced, and then sent to the ResNet50 network to identify clothing features; step6.2: Color recognition. Convert the captured image from the RGB color space to the HSV color space. Based on the H, S, and V values of the colors, determine the color to which each pixel belongs, and select the color with the most pixels as the color of the clothing in the image.

Citation Information

Patent Citations

  • Clothes size measuring method and system based on neural network and electronic device

    CN110349201A

  • Target retrieval device and method and electronic equipment

    CN112417205A