Clothing style identification method and device based on few-sample metric learning
Through the clothing style discrimination method based on the few-sample metric learning, the feature extraction network and the target part border generation network are used to solve the problem that a large amount of data is required for detecting new clothing styles in the prior art, and the rapid and accurate judgment of new clothing styles is achieved.
Patent Information
- Application Number
- CN202111572462.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-12-21
AI Technical Summary
The prior art requires a large amount of data to be retrained when detecting new clothing styles, and the scope of application is small, making it difficult to effectively detect whether the clothing style meets the requirements.
The clothing style discrimination method based on the learning of few-sample metrics is adopted. By constructing a sample data set and a clothing style discrimination model, a feature extraction network and a target part border are used to generate a network and feature measurement module, offline training and online detection are carried out to achieve rapid discrimination of new clothing styles.
It realizes the rapid identification of new clothing styles, reduces the need to retrain new clothing styles, and completes detection without a large amount of sample data, which expands the scope of application.
Smart Images

Figure CN114359793B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a clothing style discrimination method and device based on few-sample metric learning. Background Art
[0002] During the construction and maintenance of power facilities, workers are required to perform on-site operations. These workers include both team members and members of the construction unit. The Electric Power Safety Production Management Measures have clear requirements for work clothes styles. At present, the main method for detecting clothing that does not meet the requirements is to perform target detection through deep neural networks, but this method can generally only detect several existing clothing styles. For new clothing styles, retraining with a large amount of data is required, and the scope of application is limited. Summary of the invention
[0003] In view of the above-mentioned defects, the embodiments of the present invention disclose a method and device for distinguishing clothing styles based on few-sample metric learning, which can distinguish whether new clothing with a small number of samples meets the requirements.
[0004] A first aspect of an embodiment of the present invention discloses a clothing style discrimination method based on few-sample metric learning, the method comprising an offline training step and an online detection step, wherein:
[0005] The offline training step includes:
[0006] Step 1, constructing a sample data set and a clothing style discrimination model, wherein the sample data set includes multiple sample images, each of which includes a person, and the sample images include sample images of various types of work clothes styles and sample images of various general clothing styles; for each work clothes sample, a number of sample images are selected to form a support set, and the positions of all the person's torsos are marked with a labeling box for the sample images in the support set. Among the sample images in the support set, a labeling box is selected, and only the images of 8 pixels inside and around the labeling box are retained, and the pixels of the remaining parts are set to zero; the clothing style discrimination model includes a feature extraction network, a target part border generation network and a feature measurement module, and the feature measurement module includes a pooling layer and a full convolutional network;
[0007] The clothing style discrimination model also includes a first input end and a second input end, the input image of the first input end is passed through a feature extraction network to form a first feature map, and the first feature map is subjected to position-sensitive pooling through a pooling layer using artificial annotation information to obtain a first feature vector; the input image of the second input end is passed through a feature extraction network to form a second feature map, and the character torso frame obtained by the target part frame generation network is subjected to position-sensitive pooling through a pooling layer to obtain a second feature vector, and the first feature vector and the second feature vector are sent to a fully convolutional network to obtain a judgment result;
[0008] Step 2, randomly select two types of work clothes from the sample images, denoted as A and B respectively, randomly select multiple A-type sample images from the support set as positive support set images, randomly select multiple B-type sample images from the support set as negative support set images, and randomly select multiple A-type sample images, B-type sample images and general clothing style images from the sample data set to form training sample images; initialize the clothing style discrimination model; send the positive support set images and negative support set images to the first input end respectively, send the training sample images to the second input end respectively, train the clothing style discrimination model, and obtain the trained clothing style discrimination model;
[0009] The online detection step comprises the following steps:
[0010] Step 3, taking multiple standard work clothes style images from different angles, using a marking frame to mark the positions of all the standard work clothes style images, only retaining the standard work clothes style images of 8 pixels inside and around the marking frame, and setting the other parts to zero, to form a reference set image;
[0011] Step 4, obtaining a real-time monitoring image, and sending the reference set image and the monitoring image to the first input end and the second input end of the trained clothing style discrimination model respectively to determine whether the clothing meets the requirements.
[0012] As a preferred embodiment, in the first aspect of the embodiment of the present invention, training the clothing style discrimination model includes:
[0013] Step 21, initializing the clothing style discrimination model;
[0014] Step 22, input the positive support set image into the first input end, obtain the feature map of each positive support set image through the feature extraction network, and record it as the third feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding third feature map through the pooling layer to obtain the feature vector of each positive support set image, which is recorded as the third feature vector, and average the third feature vectors of all positive support set images to obtain the positive support set feature vector f As;
[0015] Step 23, input the negative support set image into the first input end, obtain the feature map of each negative support set image through the feature extraction network, and record it as the fourth feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding fourth feature map through the pooling layer to obtain the feature vector of each negative support set image, which is recorded as the fourth feature vector, and average the fourth feature vectors of all negative support set images to obtain the negative support set feature vector f Bs ;
[0016] Step 24, sending the i-th training sample image to the second input terminal, obtaining its feature map through the feature extraction network, recorded as the fifth feature map, and using the target part border generation network to obtain all the human torso borders in the i-th training sample image, 1≤i≤M, M is the total number of training sample images;
[0017] Step 25, using all the character trunk frames to perform position-sensitive pooling on the fifth feature map through a pooling layer to obtain a training sample feature vector f cij , 1≤j≤N, N is the total number of human torso bounding boxes in the i-th training sample image;
[0018] Step 26: Use the full convolutional network to calculate the feature vector f of each training sample cij and the positive support set eigenvector f As The first correlation metric score P(f cij ,f As ), and each training sample feature vector f cij and the negative support eigenvector f Bs The second correlation metric score P(f cij ,f Bs );
[0019] Step 27, calculate the score loss of the i-th training sample image:
[0020] L i =L ci +L bi +L pi
[0021] Where L ci is the classification loss, L bi is the border loss, L pi To measure the score loss:
[0022]
[0023] Among them, p i is the indicator parameter, when the sample feature vector f cijWhen the corresponding clothing style is consistent with Class A work clothes, p i =1, otherwise, p i =0;
[0024] Step 28, calculate the average score loss of all training sample images:
[0025]
[0026] Step 29, obtain the gradient through the back propagation of the average score loss, use the Adam optimization algorithm to update all network parameters of the clothing style discrimination model, and obtain the trained clothing style discrimination model.
[0027] As a preferred embodiment, in the first aspect of the embodiment of the present invention, a real-time monitoring image is obtained, and the reference set image and the monitoring image are respectively sent to the first input end and the second input end of the trained clothing style discrimination model to judge whether the clothing meets the requirements, including:
[0028] Step 41, input each reference set image into the first input end of the trained clothing style discrimination model, obtain a feature map of each reference set image through a feature extraction network, and record it as the sixth feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding sixth feature map through a pooling layer to obtain a feature vector of each reference set image, which is recorded as the sixth feature vector, and average the sixth feature vectors of all reference set images to obtain a positive reference set feature vector f s ;
[0029] Step 42, sending the monitoring image to the second input end of the trained clothing style discrimination model, obtaining its feature map through the feature extraction network, recorded as the seventh feature map, and using the target part border generation network to obtain the torso borders of all the characters in the monitoring image;
[0030] Step 43, using all the person trunk frames in the monitoring image to perform position-sensitive pooling on the seventh feature map through a pooling layer, to obtain a monitoring image feature vector f cj , 1≤j≤N, N is the total number of human torso bounding boxes in the surveillance image;
[0031] Step 44, use the full convolutional network to calculate each monitoring image feature vector f cj and the positive reference set eigenvector f s The correlation metric score P(f cj ,f s ),
[0032] Step 45, when the correlation metric score P(f cj ,f s) is greater than the threshold value T, then the clothing of the person in the torso frame of the j-th person in the monitoring image meets the requirements; otherwise, the clothing of the person in the torso frame of the j-th person in the monitoring image does not meet the requirements.
[0033] The second aspect of the embodiment of the present invention discloses a clothing style discrimination device based on few-sample metric learning, which includes an offline training module and an online detection module, wherein:
[0034] The offline training module includes:
[0035] A construction unit is used to construct a sample data set and a clothing style discrimination model, wherein the sample data set includes multiple sample images, each of which includes a person, and the sample images include sample images of various types of work clothes styles and sample images of various general clothing styles; for each work clothes sample, a number of sample images are selected to form a support set, and the positions of all the person's torsos are marked with a labeling box for the sample images in the support set; among the sample images in the support set, a labeling box is selected, and only the images of 8 pixels inside and around the labeling box are retained, and the pixels of the remaining parts are set to zero; the clothing style discrimination model includes a feature extraction network, a target part border generation network and a feature measurement module, and the feature measurement module includes a pooling layer and a full convolutional network;
[0036] The clothing style discrimination model also includes a first input end and a second input end, the input image of the first input end is passed through a feature extraction network to form a first feature map, and the first feature map is subjected to position-sensitive pooling through a pooling layer using artificial annotation information to obtain a first feature vector; the input image of the second input end is passed through a feature extraction network to form a second feature map, and the character torso frame obtained by the target part frame generation network is subjected to position-sensitive pooling through a pooling layer to obtain a second feature vector, and the first feature vector and the second feature vector are sent to a fully convolutional network to obtain a judgment result;
[0037] A training unit is used to randomly select two types of work clothes from the sample images, which are respectively recorded as A and B, randomly select multiple A-type sample images from the support set as positive support set images, randomly select multiple B-type sample images from the support set as negative support set images, and randomly select multiple A-type sample images, B-type sample images and general clothing style images from the sample data set to form training sample images; initialize the clothing style discrimination model; send the positive support set images and negative support set images to the first input end respectively, send the training sample images to the second input end respectively, train the clothing style discrimination model, and obtain a trained clothing style discrimination model;
[0038] The online detection module comprises:
[0039] A shooting unit is used to shoot a plurality of standard work clothes style images from different angles, mark the positions of all the standard work clothes style images with a marking frame, and only retain the standard work clothes style images of 8 pixels inside and around the marking frame, and set the other parts to zero to form a reference set image;
[0040] The discrimination unit is used to obtain real-time monitoring images, and respectively send the reference set images and the monitoring images to the first input end and the second input end of the trained clothing style discrimination model to judge whether the clothing meets the requirements.
[0041] As a preferred embodiment, in the second aspect of the embodiment of the present invention, the training unit includes:
[0042] An initialization subunit, used for initializing the clothing style discrimination model;
[0043] The first input subunit is used to input the positive support set image into the first input end, obtain the feature map of each positive support set image through the feature extraction network, and record it as the third feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding third feature map through the pooling layer to obtain the feature vector of each positive support set image, which is recorded as the third feature vector, and average the third feature vectors of all positive support set images to obtain the positive support set feature vector f As ;
[0044] The second input subunit is used to input the negative support set image into the first input end, obtain the feature map of each negative support set image through the feature extraction network, and record it as the fourth feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding fourth feature map through the pooling layer to obtain the feature vector of each negative support set image, which is recorded as the fourth feature vector, and average the fourth feature vectors of all negative support set images to obtain the negative support set feature vector f Bs ;
[0045] The third input subunit is used to send the i-th training sample image to the second input end, obtain its feature map through the feature extraction network, recorded as the fifth feature map, and use the target part border generation network to obtain all the human torso borders in the i-th training sample image, 1≤i≤M, M is the total number of training sample images;
[0046] The first pooling subunit is used to use the trunk frames of all characters to perform position-sensitive pooling on the fifth feature map through a pooling layer to obtain a training sample feature vector f cij , 1≤j≤N, N is the total number of human torso bounding boxes in the i-th training sample image;
[0047] The first subunit is used to calculate the feature vector f of each training sample using the full convolutional network. cij and the positive support set eigenvector f As The first correlation metric score P(f cij ,f As ), and each training sample feature vector f cij and the negative support eigenvector f Bs The second correlation metric score P(f cij ,f Bs );
[0048] The first calculation subunit is used to calculate the score loss of the i-th training sample image:
[0049] L i =L ci +L bi +L pi
[0050] Where L ci is the classification loss, L bi is the border loss, L pi To measure the score loss:
[0051]
[0052] Among them, p i is the indicator parameter, when the sample feature vector f cij When the corresponding clothing style is consistent with Class A work clothes, p i =1, otherwise, p i =0;
[0053] The second calculation subunit is used to calculate the average score loss of all training sample images:
[0054]
[0055] The updating subunit is used to obtain the gradient through the back propagation of the average score loss, and use the Adam optimization algorithm to update all network parameters of the clothing style discrimination model to obtain the trained clothing style discrimination model.
[0056] As a preferred embodiment, in the second aspect of the embodiment of the present invention, the discrimination unit includes:
[0057] The fourth input subunit is used to input each reference set image into the first input end of the trained clothing style discrimination model, obtain a feature map of each reference set image through a feature extraction network, and record it as a sixth feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding sixth feature map through a pooling layer to obtain a feature vector of each reference set image, which is recorded as the sixth feature vector; average the sixth feature vectors of all reference set images to obtain a positive reference set feature vector f s ;
[0058] a fifth input subunit, configured to input the monitoring image into the second input end of the trained clothing style discrimination model, obtain a feature map thereof through the feature extraction network, recorded as the seventh feature map, and obtain the torso frames of all the characters in the monitoring image using the target part frame generation network;
[0059] The first pooling subunit is used to use all the trunk frames of the characters in the monitoring image to perform position-sensitive pooling on the seventh feature map through the pooling layer to obtain the monitoring image feature vector f cj , 1≤j≤N, N is the total number of human torso bounding boxes in the surveillance image;
[0060] The second subunit is used to calculate the feature vector f of each monitoring image using the full convolutional network. cj and the positive reference set eigenvector f s The correlation metric score P(f cj ,f s ),
[0061] The discriminant subunit is used to determine the correlation metric score P(f cj ,f s ) is greater than the threshold value T, then the clothing of the person in the torso frame of the j-th person in the monitoring image meets the requirements; otherwise, the clothing of the person in the torso frame of the j-th person in the monitoring image does not meet the requirements.
[0062] A third aspect of an embodiment of the present invention discloses an electronic device, comprising: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute a clothing style discrimination method based on few-sample metric learning disclosed in the first aspect of an embodiment of the present invention.
[0063] A fourth aspect of an embodiment of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute a clothing style discrimination method based on few-sample metric learning disclosed in the first aspect of an embodiment of the present invention.
[0064] A fifth aspect of an embodiment of the present invention discloses a computer program product. When the computer program product runs on a computer, the computer executes a clothing style discrimination method based on few-sample metric learning disclosed in the first aspect of an embodiment of the present invention.
[0065] A sixth aspect of an embodiment of the present invention discloses an application publishing platform, which is used to publish a computer program product. When the computer program product runs on a computer, the computer executes a clothing style discrimination method based on few-sample metric learning disclosed in the first aspect of an embodiment of the present invention.
[0066] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0067] The embodiment of the present invention introduces a few-sample metric learning method into the target detection model, and judges whether the target clothing meets the requirements by comparing the features of a small number of supporting samples with the features of the detection target. After the network training is completed, only a small number of clothing samples are needed to complete the detection, and there is no need to retrain new clothing styles. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0069] Figure 1 It is a flow chart of a method for distinguishing clothing styles based on few-sample metric learning disclosed in an embodiment of the present invention;
[0070] Figure 2 is a structural schematic diagram of a clothing style discrimination model disclosed in an embodiment of the present invention;
[0071] Figure 3 It is a schematic diagram of the flow of clothing style discrimination model training disclosed in an embodiment of the present invention;
[0072] Figure 4 It is a schematic diagram of a specific process of clothing style identification disclosed in an embodiment of the present invention;
[0073] Figure 5 It is a structural schematic diagram of a clothing style discrimination device based on few-sample metric learning disclosed in an embodiment of the present invention;
[0074] Figure 6 It is a structural schematic diagram of an electronic device disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0075] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0076] It should be noted that the terms "first", "second", "third", "fourth", etc. in the specification and claims of the present invention are used to distinguish different objects rather than to describe a specific order. The terms "including" and "having" in the embodiments of the present invention and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device including a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0077] The embodiment of the present invention discloses a clothing style discrimination method and device based on few-sample metric learning, which introduces the few-sample metric learning method into the target detection model, and judges whether the target clothing meets the requirements by comparing the features of a small number of supporting samples with the features of the detection target. After the network training is completed, only a small number of clothing samples are needed to complete the detection, and there is no need to retrain the new clothing style. The following is a detailed description with reference to the accompanying drawings.
[0078] Embodiment 1
[0079] See also Figure 1 , Figure 1 is a flow chart of a method for distinguishing clothing styles based on few-sample metric learning disclosed in an embodiment of the present invention. Figure 1 As shown, the clothing style discrimination method based on few-sample metric learning mainly includes an offline training step S110 and an online detection step S120, wherein:
[0080] The offline training step S110 includes:
[0081] S111, construct a sample data set.
[0082] The sample data set includes a plurality of sample images, each of which includes a person, and the sample images include sample images of various types of work clothes styles and sample images of various general clothing styles.
[0083] In a preferred embodiment of the present invention, at least 10,000 sample images containing people and of uniform size are collected, including at least 3,000 general clothing styles and 100 work clothes styles, and each work clothes style includes at least 20 images.
[0084] If the sample images are not the same size, they can be scaled to a uniform size of w×h by scaling, where w is the width of the image and h is the height of the image. It should be noted that scaling includes two situations: enlargement or reduction, so that all sample images have the same size and the same number of pixels. During the enlargement process, the pixel points can be increased by interpolation, etc., and during reduction, the equal-space sampling method or the local mean method can be used.
[0085] For each work clothing sample, several sample images (for example, 10 images for each work clothing style) are selected to form a support set. The sample images in the support set are manually labeled, and the positions of all the human torsos are marked with a labeling box. Among the sample images in the support set, a labeling box is selected, and only the image of 8 pixels inside and around the labeling box is retained, and the pixels of the remaining parts are set to zero.
[0086] The annotation box is set as needed, which can be a rectangular box or other shapes. There is no specific limitation here. The annotation box marks the position of all the torsos of the characters in each sample image of the support set (when there are multiple characters, there may be multiple annotation boxes). The annotation box needs to cover the torso boundary more accurately.
[0087] Select any annotation box. Of course, there must be a corresponding work clothes style in this annotation box. You can also select multiple annotation boxes to obtain multiple sample images in the support set. In order to avoid the influence of background and other environmental factors on the training process, in a preferred embodiment of the present invention, only the image of 8 pixels inside and around the annotation box (the four corners of the rectangular box and the midpoint of each line) is retained, and the pixels of the remaining parts are set to zero (black). Of course, in some other scenarios, images in other positions can also be removed by cropping.
[0088] S112, constructing a clothing style discrimination model.
[0089] In a preferred embodiment of the present invention, please refer to Figure 2 As shown, the clothing style discrimination model includes a feature extraction network, a target part border generation network and a feature measurement module, and the feature measurement module includes a pooling layer and a full convolutional network.
[0090] The feature extraction network uses a deep residual network (ResNet50), removes its final average pooling layer and fully connected layer, and only uses the convolution layer to calculate the feature map. In addition, a 1×1×1024 full convolution layer is added to reduce the output of ResNet50 from 2048 dimensions to 1024 dimensions. After the support set image and the sample image pass through the network, the corresponding feature maps can be obtained respectively.
[0091] The target part border generation network uses the last level network of the R-FCN network, which has three branches: the first branch is the Region Proposal Networks (RPN), which takes the output of conv4 in ResNet50 as input and aims to generate the Region-of-Interest (ROI); the second branch is the k- 2 The convolutional layer of (C+1) channels is input to the final output of the feature extraction module. The purpose is to generate k for the background and each type of target. 2 Position sensitive score maps, where C is the type of human body part to be detected. In this embodiment, only the torso is detected, C=1, k=3; the third branch has 4k 2 The convolutional layer of the channel has the same input as the second branch, and its purpose is to predict the four regression parameters (t x ,t y ,t w ,t h ) each produces k 2 Finally, for each ROI generated by RPN, position-sensitive pooling is performed on the score maps obtained in the second and third branches to obtain its classification results and bounding box regression parameters, thereby obtaining the bounding box of the target part.
[0092] Both the feature extraction network and the target part bounding box generation network use pre-trained models.
[0093] The feature measurement module first uses the manually annotated box information to perform position-sensitive pooling on the support set feature map to obtain a 1024-dimensional feature vector f s Then, each target part border obtained by the target part border generation network is used to perform position-sensitive pooling on the sample image feature map to obtain a 1024-dimensional feature vector f c , and finally use a three-layer fully convolutional network F to output f s With f c The correlation metric score P(f s ,f c ). The score is between 0 and 1, and the larger the score, the higher the relevance. The fully convolutional network is randomly initialized using a Gaussian distribution.
[0094] In some other embodiments, the feature extraction network may also use other feature extraction methods, and the full convolutional network may also be replaced by other types of correlation algorithms. Of course, it may also be implemented directly using a correlation calculation function.
[0095] The clothing style discrimination model also includes a first input end and a second input end. The input image of the first input end is passed through a feature extraction network to form a first feature map, and the first feature map is subjected to position-sensitive pooling through a pooling layer using manually labeled information to obtain a first feature vector; the input image of the second input end is passed through a feature extraction network to form a second feature map, and the character torso bounding box obtained by a target part bounding box generation network is subjected to position-sensitive pooling through a pooling layer to obtain a second feature vector. The first feature vector and the second feature vector are sent to a fully convolutional network to obtain a judgment result.
[0096] S113, randomly select two types of work clothes from the sample images, denoted as A and B respectively, randomly select multiple A-type sample images (for example, 5) from the support set as positive support set images, randomly select multiple B-type sample images (for example, 5) from the support set as negative support set images, and randomly select multiple A-type sample images (for example, 10), B-type sample images (for example, 2) and general clothing style images (for example, 8) from the sample data set to form training sample images.
[0097] Initialize the clothing style discrimination model, send the positive support set images and the negative support set images to the first input end respectively, send the training sample images to the second input end respectively, train the clothing style discrimination model, and obtain the trained clothing style discrimination model.
[0098] For details, please refer to Figure 3 As shown, the training process includes the following steps:
[0099] S1131, initializing the clothing style discrimination model.
[0100] S1132, input the positive support set images into the first input end respectively, obtain the feature map of each positive support set image through the feature extraction network, and record it as the third feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding third feature map through the pooling layer to obtain the feature vector of each positive support set image, which is recorded as the third feature vector, and average the third feature vectors of all positive support set images to obtain the positive support set feature vector f As .
[0101] S1133, in the same manner, input the negative support set image into the first input end, obtain the feature map of each negative support set image through the feature extraction network, and record it as the fourth feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding fourth feature map through the pooling layer to obtain the feature vector of each negative support set image, which is recorded as the fourth feature vector, and average the fourth feature vectors of all negative support set images to obtain the negative support set feature vector f Bs .
[0102] S1134, sending the i-th training sample image to the second input terminal, obtaining its feature map through the feature extraction network, recorded as the fifth feature map, and using the target part border generation network to obtain all the human torso borders in the i-th training sample image, 1≤i≤M, M is the total number of training sample images.
[0103] S1135, using all the character trunk frames to perform position-sensitive pooling on the fifth feature map through a pooling layer, to obtain a training sample feature vector f cij , 1≤j≤N, N is the total number of human torso bounding boxes in the i-th training sample image.
[0104] S1136, using the full convolutional network to calculate the feature vector f of each training sample cij and the positive support set eigenvector f As The first correlation metric score P(f cij ,f As ), and each training sample feature vector f cij and the negative support eigenvector f Bs The second correlation metric score P(f cij ,f Bs ).
[0105] S1137, calculate the score loss L of the i-th training sample image i :
[0106] L i =L ci +L bi +L pi
[0107] Where L ci is the classification loss, L bi is the border loss, which is defined in the same way as the R-FCN network. pi To measure the score loss, it is calculated as:
[0108]
[0109] Among them, p i is the indicator parameter, when the sample feature vector fcij When the corresponding clothing style is consistent with Class A work clothes, p i =1, otherwise, p i =0;
[0110] The above formula can be used to calculate the score loss L of the i-th training sample image: i , after the score losses of all training sample images are calculated, the average score loss L of all training sample images can be calculated through S1138.
[0111] S1138, calculate the average score loss of all training sample images:
[0112]
[0113] S1139, obtaining the gradient through back propagation of the average score loss, and using the Adam optimization algorithm to update all network parameters of the clothing style discrimination model to obtain the trained clothing style discrimination model.
[0114] The online detection step 120 includes the following steps:
[0115] S121, taking multiple standard work clothes style images from different angles, using a labeling box to mark the positions of all the standard work clothes style images, only retaining the standard work clothes style images of 8 pixels inside and around the labeling box, and setting the other parts to zero, to form a reference set image.
[0116] The present invention can be applied to the identification of new styles of clothing. For the work clothes styles mentioned in the sample, the work clothes styles and monitoring images supporting the type can be directly input into the first input terminal and the second input terminal respectively to determine whether they meet the requirements.
[0117] For new styles of work clothes, multiple photos of new styles of work clothes (e.g., 5 photos, called standard work clothes styles) can be taken from different angles and then compared with the surveillance images, so that a small number of samples can be used to determine whether the clothing style meets the requirements. For new styles of work clothes that have been determined, their reference set images can be saved in the support set for subsequent use.
[0118] S122, acquiring a real-time monitoring image, and sending the reference set image and the monitoring image to the first input terminal and the second input terminal of the trained clothing style discrimination model respectively, to judge whether the clothing meets the requirements.
[0119] For details, please refer to Figure 4 As shown, it may include the following steps:
[0120] S1221, input each reference set image into the first input end of the trained clothing style discrimination model, obtain a feature map of each reference set image through a feature extraction network, and record it as the sixth feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding sixth feature map through a pooling layer to obtain a feature vector of each reference set image, which is recorded as the sixth feature vector, and average the sixth feature vectors of all reference set images to obtain a positive reference set feature vector f s .
[0121] S1222, sending the monitoring image to the second input end of the trained clothing style discrimination model, obtaining its feature map through the feature extraction network, recorded as the seventh feature map, and using the target part border generation network to obtain the torso borders of all the characters in the monitoring image.
[0122] S1223, using all the person trunk frames in the monitoring image to perform position-sensitive pooling on the seventh feature map through a pooling layer, to obtain a monitoring image feature vector f cj , 1≤j≤N, N is the total number of human torso bounding boxes in the surveillance image.
[0123] S1224, using the full convolutional network to calculate each monitoring image feature vector f cj and the positive reference set eigenvector f s The correlation metric score P(f cj ,f s ).
[0124] Among them, the processes of S1221 and S1132 are similar, and the processes of S1222-S1224 and S1134-S1136 are similar.
[0125] S1225, when the correlation metric score P(f cj ,f s ) is greater than the threshold value T, then the clothing of the person in the torso frame of the j-th person in the monitoring image meets the requirements; otherwise, the clothing of the person in the torso frame of the j-th person in the monitoring image does not meet the requirements.
[0126] Embodiment 2
[0127] See also Figure 5 , Figure 5 is a schematic diagram of the structure of a clothing style discrimination device based on few-sample metric learning disclosed in an embodiment of the present invention. Figure 5 As shown, the clothing style discrimination device based on few-sample metric learning may include an offline training module 210 and an online detection module 220, wherein:
[0128] The offline training module 210 includes:
[0129] A construction unit 211 is used to construct a sample data set and a clothing style discrimination model, wherein the sample data set includes multiple sample images, each of which includes a person, and the sample images include sample images of various types of work clothes styles and sample images of various general clothing styles; for each work clothes sample, a number of sample images are selected to form a support set, and the positions of all the person's torsos are annotated using a labeling box for the sample images in the support set. Among the sample images in the support set, a labeling box is selected, and only the images of 8 pixels inside and around the labeling box are retained, and the pixels of the remaining parts are set to zero; the clothing style discrimination model includes a feature extraction network, a target part border generation network, and a feature measurement module, and the feature measurement module includes a pooling layer and a full convolutional network;
[0130] The clothing style discrimination model also includes a first input end and a second input end, the input image of the first input end is passed through a feature extraction network to form a first feature map, and the first feature map is subjected to position-sensitive pooling through a pooling layer using artificial annotation information to obtain a first feature vector; the input image of the second input end is passed through a feature extraction network to form a second feature map, and the character torso frame obtained by the target part frame generation network is subjected to position-sensitive pooling through a pooling layer to obtain a second feature vector, and the first feature vector and the second feature vector are sent to a fully convolutional network to obtain a judgment result;
[0131] The training unit 212 is used to randomly select two types of work clothes from the sample images, which are respectively denoted as A and B, randomly select multiple A-type sample images from the support set as positive support set images, randomly select multiple B-type sample images from the support set as negative support set images, and randomly select multiple A-type sample images, B-type sample images and general clothing style images from the sample data set to form training sample images; initialize the clothing style discrimination model; send the positive support set images and negative support set images to the first input terminal respectively, send the training sample images to the second input terminal respectively, train the clothing style discrimination model, and obtain a trained clothing style discrimination model;
[0132] The online detection module 220 includes:
[0133] The shooting unit 221 is used to shoot a plurality of standard work clothes style images from different angles, mark the positions of all the standard work clothes style images with a marking frame, and only retain the standard work clothes style images of 8 pixels inside and around the marking frame, and set the other parts to zero to form a reference set image;
[0134] The discrimination unit 222 is used to obtain real-time monitoring images, and send the reference set image and the monitoring image to the first input terminal and the second input terminal of the trained clothing style discrimination model respectively to judge whether the clothing meets the requirements.
[0135] Preferably, the training unit 212 may include:
[0136] An initialization subunit, used for initializing the clothing style discrimination model;
[0137] The first input subunit is used to input the positive support set image into the first input end, obtain the feature map of each positive support set image through the feature extraction network, and record it as the third feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding third feature map through the pooling layer to obtain the feature vector of each positive support set image, which is recorded as the third feature vector, and average the third feature vectors of all positive support set images to obtain the positive support set feature vector f As ;
[0138] The second input subunit is used to input the negative support set image into the first input end, obtain the feature map of each negative support set image through the feature extraction network, and record it as the fourth feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding fourth feature map through the pooling layer to obtain the feature vector of each negative support set image, which is recorded as the fourth feature vector, and average the fourth feature vectors of all negative support set images to obtain the negative support set feature vector f Bs ;
[0139] The third input subunit is used to send the i-th training sample image to the second input end, obtain its feature map through the feature extraction network, recorded as the fifth feature map, and use the target part border generation network to obtain all the human torso borders in the i-th training sample image, 1≤i≤M, M is the total number of training sample images;
[0140] The first pooling subunit is used to use the trunk frames of all characters to perform position-sensitive pooling on the fifth feature map through a pooling layer to obtain a training sample feature vector f cij , 1≤j≤N, N is the total number of human torso bounding boxes in the i-th training sample image;
[0141] The first subunit is used to calculate the feature vector f of each training sample using the full convolutional network. cij and the positive support set eigenvector f As The first correlation metric score P(f cij ,f As ), and each training sample feature vector f cij and the negative support eigenvector f Bs The second correlation metric score P(fcij ,f Bs );
[0142] The first calculation subunit is used to calculate the score loss of the i-th training sample image:
[0143] L i =L ci +L bi +L pi
[0144] Where L ci is the classification loss, L bi is the border loss, L pi To measure the score loss:
[0145]
[0146] Among them, p i is the indicator parameter, when the sample feature vector f cij When the corresponding clothing style is consistent with Class A work clothes, p i =1, otherwise, p i =0;
[0147] The second calculation subunit is used to calculate the average score loss of all training sample images:
[0148]
[0149] The updating subunit is used to obtain the gradient through the back propagation of the average score loss, and use the Adam optimization algorithm to update all network parameters of the clothing style discrimination model to obtain the trained clothing style discrimination model.
[0150] Preferably, the determination unit 222 includes:
[0151] The fourth input subunit is used to input each reference set image into the first input end of the trained clothing style discrimination model, obtain a feature map of each reference set image through a feature extraction network, and record it as a sixth feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding sixth feature map through a pooling layer to obtain a feature vector of each reference set image, which is recorded as the sixth feature vector; average the sixth feature vectors of all reference set images to obtain a positive reference set feature vector f s ;
[0152] a fifth input subunit, configured to input the monitoring image into the second input end of the trained clothing style discrimination model, obtain a feature map thereof through the feature extraction network, recorded as the seventh feature map, and obtain the torso frames of all the characters in the monitoring image using the target part frame generation network;
[0153] The first pooling subunit is used to use all the trunk frames of the characters in the monitoring image to perform position-sensitive pooling on the seventh feature map through the pooling layer to obtain the monitoring image feature vector f cj , 1≤j≤N, N is the total number of human torso bounding boxes in the surveillance image;
[0154] The second subunit is used to calculate the feature vector f of each monitoring image using the full convolutional network. cj and the positive reference set eigenvector f s The correlation metric score P(f cj ,f s ),
[0155] The discriminant subunit is used to determine the correlation metric score P(f cj ,f s ) is greater than the threshold value T, then the clothing of the person in the torso frame of the j-th person in the monitoring image meets the requirements; otherwise, the clothing of the person in the torso frame of the j-th person in the monitoring image does not meet the requirements.
[0156] Embodiment 3
[0157] See also Figure 6 , Figure 6 Schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention. Figure 6 As shown, the electronic device may include:
[0158] A memory 310 storing executable program codes;
[0159] a processor 320 coupled to the memory 310;
[0160] The processor 320 calls the executable program code stored in the memory 310 to execute part or all of the steps in the clothing style identification method based on few-sample metric learning in the first embodiment.
[0161] An embodiment of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute part or all of the steps in a clothing style discrimination method based on few-sample metric learning in embodiment one.
[0162] An embodiment of the present invention further discloses a computer program product, wherein when the computer program product runs on a computer, the computer executes part or all of the steps in a clothing style discrimination method based on few-sample metric learning in embodiment one.
[0163] An embodiment of the present invention further discloses an application publishing platform, wherein the application publishing platform is used to publish a computer program product, wherein when the computer program product runs on a computer, the computer executes part or all of the steps in a clothing style discrimination method based on few-sample metric learning in embodiment one.
[0164] In various embodiments of the present invention, it should be understood that the size of the serial numbers of the processes does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0165] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed over multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0166] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0167] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-accessible memory. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a memory and includes several requests for a computer device (which can be a personal computer, a server or a network device, etc., specifically a processor in a computer device) to perform some or all of the steps of the method described in each embodiment of the present invention.
[0168] In the embodiments provided by the present invention, it should be understood that "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.
[0169] A person of ordinary skill in the art can understand that some or all of the steps in the various methods of the embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0170] The above is a detailed introduction to a clothing style discrimination method and device based on few-sample metric learning disclosed in an embodiment of the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A clothing style identification method based on few-sample metric learning, characterized in that: It includes an offline training step and an online detection step, where: The offline training step includes: Step 1, constructing a sample data set and a clothing style discrimination model, wherein the sample data set includes multiple sample images, each of which includes a person, and the sample images include sample images of various types of work clothes styles and sample images of various general clothing styles; for each work clothes sample, a number of sample images are selected to form a support set, and the positions of all the person's torsos are marked with a labeling box for the sample images in the support set. Among the sample images in the support set, a labeling box is selected, and only the images of 8 pixels inside and around the labeling box are retained, and the pixels of the remaining parts are set to zero; the clothing style discrimination model includes a feature extraction network, a target part border generation network and a feature measurement module, and the feature measurement module includes a pooling layer and a full convolutional network; The clothing style discrimination model also includes a first input end and a second input end, the input image of the first input end is passed through a feature extraction network to form a first feature map, and the first feature map is subjected to position-sensitive pooling through a pooling layer using artificial annotation information to obtain a first feature vector; the input image of the second input end is passed through a feature extraction network to form a second feature map, and the character torso frame obtained by the target part frame generation network is subjected to position-sensitive pooling through a pooling layer to obtain a second feature vector, and the first feature vector and the second feature vector are sent to a fully convolutional network to obtain a judgment result; Step 2, randomly select two types of work clothes from the sample images, denoted as A and B respectively, randomly select multiple A-type sample images from the support set as positive support set images, randomly select multiple B-type sample images from the support set as negative support set images, and randomly select multiple A-type sample images, B-type sample images and general clothing style images from the sample data set to form training sample images; initialize the clothing style discrimination model; send the positive support set images and negative support set images to the first input end respectively, send the training sample images to the second input end respectively, train the clothing style discrimination model, and obtain the trained clothing style discrimination model; The online detection step comprises the following steps: Step 3, taking multiple standard work clothes style images from different angles, using a marking frame to mark the positions of all the standard work clothes style images, only retaining the standard work clothes style images of 8 pixels inside and around the marking frame, and setting the other parts to zero, to form a reference set image; Step 4, obtaining a real-time monitoring image, and sending the reference set image and the monitoring image to the first input terminal and the second input terminal of the trained clothing style discrimination model respectively to judge whether the clothing meets the requirements; Training the clothing style discrimination model includes: Step 21, initializing the clothing style discrimination model; Step 22, input the positive support set image into the first input end, obtain the feature map of each positive support set image through the feature extraction network, and record it as the third feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding third feature map through the pooling layer to obtain the feature vector of each positive support set image, which is recorded as the third feature vector, and average the third feature vectors of all positive support set images to obtain the positive support set feature vector f As ; Step 23, input the negative support set image into the first input end, obtain the feature map of each negative support set image through the feature extraction network, and record it as the fourth feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding fourth feature map through the pooling layer to obtain the feature vector of each negative support set image, which is recorded as the fourth feature vector, and average the fourth feature vectors of all negative support set images to obtain the negative support set feature vector f Bs ; Step 24, sending the i-th training sample image to the second input terminal, obtaining its feature map through the feature extraction network, recorded as the fifth feature map, and using the target part border generation network to obtain all the human torso borders in the i-th training sample image, 1≤i≤M, M is the total number of training sample images; Step 25, using all the character trunk frames to perform position-sensitive pooling on the fifth feature map through a pooling layer to obtain a training sample feature vector f cij , 1≤j≤N, N is the total number of human torso bounding boxes in the i-th training sample image; Step 26: Use the full convolutional network to calculate the feature vector f of each training sample cij and the positive support set eigenvector f As The first correlation metric score P(f cij ,f As ), and each training sample feature vector f cij and the negative support eigenvector f Bs The second correlation metric score P(f cij ,f Bs ); Step 27, calculate the score loss of the i-th training sample image: L i =L ci +L bi +L pi Where L ci is the classification loss, L bi is the border loss, L pi To measure the score loss: Among them, p i is the indicator parameter, when the sample feature vector f cij When the corresponding clothing style is consistent with Class A work clothes, p i =1, otherwise, p i =0; Step 28, calculate the average score loss of all training sample images: Step 29, obtain the gradient through the back propagation of the average score loss, use the Adam optimization algorithm to update all network parameters of the clothing style discrimination model, and obtain the trained clothing style discrimination model.
2. The clothing style identification method based on few-sample metric learning according to claim 1, characterized in that: Acquire a real-time monitoring image, and send the reference set image and the monitoring image to the first input terminal and the second input terminal of the trained clothing style discrimination model respectively to judge whether the clothing meets the requirements, including: Step 41, input each reference set image into the first input end of the trained clothing style discrimination model, obtain a feature map of each reference set image through a feature extraction network, and record it as the sixth feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding sixth feature map through a pooling layer to obtain a feature vector of each reference set image, which is recorded as the sixth feature vector, and average the sixth feature vectors of all reference set images to obtain a positive reference set feature vector f s ; Step 42, sending the monitoring image to the second input end of the trained clothing style discrimination model, obtaining its feature map through the feature extraction network, recorded as the seventh feature map, and using the target part border generation network to obtain the torso borders of all the characters in the monitoring image; Step 43, using all the person trunk frames in the monitoring image to perform position-sensitive pooling on the seventh feature map through a pooling layer, to obtain a monitoring image feature vector f cj , 1≤j≤N, N is the total number of human torso bounding boxes in the surveillance image; Step 44, use the full convolutional network to calculate each monitoring image feature vector f cj and the positive reference set eigenvector f s The correlation metric score P(f cj ,f s ), Step 45, when the correlation metric score P(f cj ,f s ) is greater than the threshold value T, then the clothing of the person in the torso frame of the j-th person in the monitoring image meets the requirements; otherwise, the clothing of the person in the torso frame of the j-th person in the monitoring image does not meet the requirements.
3. A clothing style identification device based on few-sample metric learning, characterized in that: It includes an offline training module and an online detection module, where: The offline training module includes: A construction unit is used to construct a sample data set and a clothing style discrimination model, wherein the sample data set includes multiple sample images, each of which includes a person, and the sample images include sample images of various types of work clothes styles and sample images of various general clothing styles; for each work clothes sample, a number of sample images are selected to form a support set, and the positions of all the person's torsos are marked with a labeling box for the sample images in the support set; among the sample images in the support set, a labeling box is selected, and only the images of 8 pixels inside and around the labeling box are retained, and the pixels of the remaining parts are set to zero; the clothing style discrimination model includes a feature extraction network, a target part border generation network and a feature measurement module, and the feature measurement module includes a pooling layer and a full convolutional network; The clothing style discrimination model also includes a first input end and a second input end, the input image of the first input end is passed through a feature extraction network to form a first feature map, and the first feature map is subjected to position-sensitive pooling through a pooling layer using artificial annotation information to obtain a first feature vector; the input image of the second input end is passed through a feature extraction network to form a second feature map, and the character torso frame obtained by the target part frame generation network is subjected to position-sensitive pooling through a pooling layer to obtain a second feature vector, and the first feature vector and the second feature vector are sent to a fully convolutional network to obtain a judgment result; A training unit is used to randomly select two types of work clothes from the sample images, which are respectively recorded as A and B, randomly select multiple A-type sample images from the support set as positive support set images, randomly select multiple B-type sample images from the support set as negative support set images, and randomly select multiple A-type sample images, B-type sample images and general clothing style images from the sample data set to form training sample images; initialize the clothing style discrimination model; send the positive support set images and negative support set images to the first input end respectively, send the training sample images to the second input end respectively, train the clothing style discrimination model, and obtain a trained clothing style discrimination model; The online detection module comprises: A shooting unit is used to shoot a plurality of standard work clothes style images from different angles, mark the positions of all the standard work clothes style images with a marking frame, and only retain the standard work clothes style images of 8 pixels inside and around the marking frame, and set the other parts to zero to form a reference set image; A discrimination unit, used for acquiring a real-time monitoring image, and sending the reference set image and the monitoring image to a first input terminal and a second input terminal of a trained clothing style discrimination model, respectively, to discriminate whether the clothing meets the requirements; The training unit comprises: An initialization subunit, used for initializing the clothing style discrimination model; The first input subunit is used to input the positive support set image into the first input end, obtain the feature map of each positive support set image through the feature extraction network, and record it as the third feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding third feature map through the pooling layer to obtain the feature vector of each positive support set image, which is recorded as the third feature vector, and average the third feature vectors of all positive support set images to obtain the positive support set feature vector f As ; The second input subunit is used to input the negative support set image into the first input end, obtain the feature map of each negative support set image through the feature extraction network, and record it as the fourth feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding fourth feature map through the pooling layer to obtain the feature vector of each negative support set image, which is recorded as the fourth feature vector, and average the fourth feature vectors of all negative support set images to obtain the negative support set feature vector f Bs ; The third input subunit is used to send the i-th training sample image to the second input end, obtain its feature map through the feature extraction network, recorded as the fifth feature map, and use the target part border generation network to obtain all the human torso borders in the i-th training sample image, 1≤i≤M, M is the total number of training sample images; The first pooling subunit is used to use the trunk frames of all characters to perform position-sensitive pooling on the fifth feature map through a pooling layer to obtain a training sample feature vector f cij , 1≤j≤N, N is the total number of human torso bounding boxes in the i-th training sample image; The first subunit is used to calculate the feature vector f of each training sample using the full convolutional network. cij and the positive support set eigenvector f As The first correlation metric score P(f cij ,f As ), and each training sample feature vector f cij and the negative support eigenvector f Bs The second correlation metric score P(f cij ,f Bs ); The first calculation subunit is used to calculate the score loss of the i-th training sample image: L i =L ci +L bi +L pi Where L ci is the classification loss, L bi is the border loss, L pi To measure the score loss: Among them, p i is the indicator parameter, when the sample feature vector f cij When the corresponding clothing style is consistent with Class A work clothes, p i =1, otherwise, p i =0; The second calculation subunit is used to calculate the average score loss of all training sample images: The updating subunit is used to obtain the gradient through the back propagation of the average score loss, and use the Adam optimization algorithm to update all network parameters of the clothing style discrimination model to obtain the trained clothing style discrimination model.
4. The clothing style identification device based on few-sample metric learning according to claim 3, characterized in that: The determination unit comprises: The fourth input subunit is used to input each reference set image into the first input end of the trained clothing style discrimination model, obtain a feature map of each reference set image through a feature extraction network, and record it as a sixth feature map; use the manual annotation information to perform position-sensitive pooling on the corresponding sixth feature map through a pooling layer to obtain a feature vector of each reference set image, which is recorded as the sixth feature vector; average the sixth feature vectors of all reference set images to obtain a positive reference set feature vector f s ; a fifth input subunit, configured to input the monitoring image into the second input end of the trained clothing style discrimination model, obtain a feature map thereof through the feature extraction network, recorded as the seventh feature map, and obtain the torso frames of all the characters in the monitoring image using the target part frame generation network; The first pooling subunit is used to use all the trunk frames of the characters in the monitoring image to perform position-sensitive pooling on the seventh feature map through the pooling layer to obtain the monitoring image feature vector f cj , 1≤j≤N, N is the total number of human torso bounding boxes in the surveillance image; The second subunit is used to calculate the feature vector f of each monitoring image using the full convolutional network. cj and the positive reference set eigenvector f s The correlation metric score P(f cj ,f s ), The discriminant subunit is used to determine the correlation metric score P(f cj ,f s ) is greater than the threshold value T, then the clothing of the person in the torso frame of the j-th person in the monitoring image meets the requirements; otherwise, the clothing of the person in the torso frame of the j-th person in the monitoring image does not meet the requirements.
5. An electronic device, characterized in that: include: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the clothing style discrimination method based on few-sample metric learning as described in any one of claims 1 to 2.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program enables a computer to execute a clothing style discrimination method based on few-sample metric learning as described in any one of claims 1 to 2.
Citation Information
Patent Citations
Clothes attribute identification method and device and electronic equipment
CN111079757A
End-to-end 3D-CapsNet flame detection method and end-to-end 3D-CapsNet flame detection device
CN111353412A
Image processing method and device based on artificial intelligence, and storage medium
CN113392866A