Dress code discrimination method, person re-identification model training method, and apparatus

GB202508530D0Pending Publication Date: 2025-07-16BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
GB2025008530
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-11-30
Publication Date
2025-07-16

AI Technical Summary

Technical Problem

Existing pedestrian re-identification technology cannot accurately identify the dress code of each part of the human body.

Method used

By performing human body detection and area division on the image to be recognized, the pedestrian re-identification model is used to extract features of the target human body area image, and compared with the feature vectors in the dress code sample image to determine whether the dress code is standard.

Benefits of technology

It achieves accurate identification of clothing for each part of the human body and improves identification accuracy.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Provided in the present invention are a dress code discrimination method, a person re-identification model training method, and an apparatus. The dress code discrimination method comprises: performing human body detection on an image to be identified, so as to obtain a first human body detection image of a target person in the image to be identified; performing human body area division on the first human body detection image to obtain a target human body area image of the target person; using a person re-identification model to perform feature extraction on the target human body area image of the target person to obtain a first feature vector; comparing the first feature vector with a second feature vector of the target human body area in a dress code example image to obtain a comparison result; and, according to the comparison result, determining whether the dress of the target human body area of the target person complies with the code.
Need to check novelty before this filing date? Find Prior Art

Description

Dress code identification method, pedestrian re-identification model training method and device Technical Field

[0001] The embodiments of the present invention relate to the field of image processing technology, and in particular to a dress code identification method, a pedestrian re-identification model training method and a device. Background Art

[0002] With the development of science and technology, computer vision technology has been increasingly applied in daily life and production. In places like factories and bank lobbies, employees need to adhere to dress codes. Person re-identification technology within computer vision can compare the similarity between a real-life image of a person and a given dress code sample, helping to determine whether the employee's attire meets the requirements.

[0003] However, current person re-identification technology cannot accurately identify the clothing specifications of each part of the human body.

[0004] Summary of the Invention

[0005] The embodiments of the present invention provide a dress code identification method, a pedestrian re-identification model training method and a device, which are used to solve the problem that pedestrian re-identification technology in the prior art cannot accurately identify the dress code of each part of the human body.

[0006] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:

[0007] In a first aspect, an embodiment of the present invention provides a dress code determination method, comprising:

[0008] Performing human body detection on the image to be identified to obtain a first human body detection image of a target pedestrian in the image to be identified;

[0009] Performing human body region segmentation on the first human body detection image to obtain a target human body region image of the target pedestrian;

[0010] Using a person re-identification model to perform feature extraction on the target human body region image of the target pedestrian to obtain a first feature vector;

[0011] Comparing the first feature vector with a second feature vector of a target human body region in the dress code sample image to obtain a comparison result;

[0012] Determine whether the clothing of the target human body area of ​​the target pedestrian is standard according to the comparison result.

[0013] Optionally, dividing the first human body detection image into human body regions includes:

[0014] Extracting human key points from the first human detection image using a preprocessing model to obtain first human key points;

[0015] The first human body detection image is divided into human body areas according to the first human body key points.

[0016] Optionally, before comparing the first feature vector with the second feature vector of the target human body region in the dress code sample image, the method further includes:

[0017] Performing human body detection on the dress code sample image to obtain a second human body detection image;

[0018] performing human body region segmentation on the second human body detection image to obtain a target human body region image in the dress code sample image;

[0019] The pedestrian re-identification model is used to extract features of the target human body area image in the dress code sample image to obtain the second feature vector.

[0020] Optionally, comparing the first feature vector with a second feature vector of a target human body region in the dress code sample image to obtain a comparison result includes:

[0021] The cosine similarity between the first feature vector and the second feature vector of the target human body region in the dress code sample image is calculated to obtain similarity information as the comparison result.

[0022] Optionally, determining whether the target body area of ​​the target pedestrian is properly dressed according to the comparison result includes:

[0023] For a target human body region of the target pedestrian, if the comparison results of N frames of images to be identified containing the target pedestrian indicate that similarity information between a first feature vector of the target human body region in at least M frames of the images to be identified and a second feature vector of the target human body region in the dress code sample image does not reach a preset threshold, it is determined that the dress code of the target human body region is not standardized;

[0024] Wherein, N is a positive integer greater than or equal to 1, and M is a positive integer greater than or equal to 1 and less than N.

[0025] In a second aspect, an embodiment of the present invention further provides a method for training a person re-identification model, comprising:

[0026] Determining a plurality of training image pairs, each of the training image pairs including at least two training images;

[0027] Performing human body detection on the training image in the training image pair to obtain a third human body detection image in the training image;

[0028] Performing human body region division on the third human body detection image to obtain a target human body region image;

[0029] Using the pedestrian re-identification model to be trained to perform feature extraction on the target human body area image to obtain a third feature vector;

[0030] Comparing the third eigenvectors of the target human body region of each training image in the training image pair to obtain a comparison result;

[0031] The pedestrian re-identification model to be trained is optimized according to the comparison result to obtain a trained pedestrian re-identification model.

[0032] Optionally, dividing the third human body detection image into human body regions includes:

[0033] Extracting human key points from the third human body detection image using a preprocessing model to obtain third human body key points;

[0034] The third human body detection image is divided into human body areas according to the third human body key points.

[0035] Optionally, determining a plurality of training image pairs includes:

[0036] Extracting human attribute information from the candidate image using a preprocessing model to obtain the human attribute information of the candidate image;

[0037] A training image is selected from the candidate images according to the human body attribute information to form the training image pair.

[0038] Optionally, the human body attribute information includes human body orientation, and selecting training images from the candidate images according to the human body attribute information to form the training image pair includes: selecting training images of the same person with the same and / or different orientations from the candidate images as the training images in the training image pair;

[0039] and / or

[0040] The human body attribute information includes human body orientation and clothing color. Selecting training images from the candidate images according to the human body attribute information to form the training image pair includes: selecting training images of different people with the same orientation and the same color clothing from the candidate images as the training images in the training image pair.

[0041] Optionally, selecting training images of the same person in the same and / or different orientations from the candidate images as training images in the training image pair includes:

[0042] For a training image, from multiple candidate images containing the same person, select a first image of a first difficulty with a first probability, select a second image of a second difficulty with a second probability, and select a third image of a third difficulty with a third probability as the training images in the training image pair;

[0043] The first difficulty level refers to: the person in one of the training image and the first image is facing forward, and the person in the other image is facing backward; or the person in one of the training image and the first image is facing left, and the person in the other image is facing right;

[0044] The second difficulty level refers to: the person in one of the training image and the second image is facing forward, and the person in the other image is facing left or right; or the person in one of the training image and the first image is facing backward, and the person in the other image is facing left or right;

[0045] The third difficulty level means that the characters in the training image and the third image are oriented in the same direction.

[0046] Optionally, selecting training images of different persons wearing the same clothing in the same orientation and the same color from the candidate images as training images in the training image pair includes:

[0047] For a training image, selecting a candidate image containing a different person than the training image;

[0048] Calculating similarity information between the training image and the candidate image, wherein the similarity information is determined by at least one of the following: clothing color, hat wearing, and orientation of the person in the training image and the candidate image;

[0049] Dividing the candidate images into multiple sets according to the similarity information of the candidate images;

[0050] Candidate images are selected from different sets as training images in the training image pairs.

[0051] Optionally, determining a plurality of training image pairs includes:

[0052] For a training image, selecting a candidate image containing a different person than the training image;

[0053] Performing human body detection on the candidate image to obtain a fourth human body detection image in the candidate image;

[0054] extracting human key points from the fourth human body detection image using a preprocessing model to obtain second human body key points;

[0055] dividing the fourth human body detection image into human body regions according to the second human body key points to obtain a target human body region image in the fourth human body detection image;

[0056] Determine the mean and variance of the target human body region image in the fourth human body detection image on the RGB three channels;

[0057] Convert the target human body region image in the fourth human body detection image into HSV space, and calculate the mean and variance of the HSV three channels after conversion into HSV space;

[0058] Obtaining the image features after dimensionality reduction according to the mean and variance of the RGB three channels and the mean and variance of the HSV three channels;

[0059] Clustering the image features after dimensionality reduction of the plurality of candidate images, and dividing the plurality of candidate images into different clusters;

[0060] Candidate images are selected from different clusters as training images in the training image pairs.

[0061] In a third aspect, an embodiment of the present invention further provides a dress code determination device, comprising:

[0062] A first human detection module is used to perform human detection on the image to be identified, and obtain a first human detection image of a target pedestrian in the image to be identified;

[0063] a first human body region division module, configured to perform human body region division on the first human body detection image to obtain a target human body region image of the target pedestrian;

[0064] A first feature extraction module is used to extract features from the target human body region image of the target pedestrian using a pedestrian re-identification model to obtain a first feature vector;

[0065] a first comparison module, configured to compare the first feature vector with a second feature vector of a target human body region in a dress code sample image to obtain a comparison result;

[0066] The discrimination module is used to determine whether the clothing of the target human body area of ​​the target pedestrian is standard according to the comparison result.

[0067] In a fourth aspect, an embodiment of the present invention further provides a person re-identification model training device, comprising:

[0068] a determination module, configured to determine a plurality of training image pairs, each of the training image pairs including at least two training images;

[0069] a third human body detection module, configured to perform human body detection on the training image in the training image pair to obtain a third human body detection image in the training image;

[0070] A third human body region division module is used to divide the third human body detection image into human body regions to obtain a target human body region image;

[0071] a third feature extraction module, configured to extract features from the target human body region image using a person re-identification model to be trained, to obtain a third feature vector;

[0072] a second comparison module, configured to compare the third eigenvectors of the target human body region of each training image in the training image pair to obtain a comparison result;

[0073] The optimization module is used to optimize the pedestrian re-identification model to be trained according to the comparison result to obtain a trained pedestrian re-identification model.

[0074] In a fifth aspect, an embodiment of the present invention further provides an electronic device comprising: a processor, a memory, and a program stored on the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the dress code determination method as described in the first or second aspect above are implemented.

[0075] In a sixth aspect, an embodiment of the present invention further provides a non-volatile computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the dress code determination method as described in the first or second aspect above are implemented.

[0076] In an embodiment of the present invention, before using the pedestrian re-identification module to identify the clothing of the target pedestrian in the image to be identified, the human body detection image in the image to be identified is divided into human body regions to obtain the human body region image of the target pedestrian. The pedestrian re-identification module is used to identify each human body region image separately, so as to accurately identify the standard situation of the clothing of each part of the human body and improve the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0078] FIG1 is a flow chart of a dress code identification method according to an embodiment of the present invention;

[0079] FIG2 is a flow chart of a method for training a person re-identification model according to an embodiment of the present invention;

[0080] FIG3 is a schematic diagram of a training method for a person re-identification model according to an embodiment of the present invention;

[0081] FIG4 is a schematic structural diagram of a preprocessing model according to an embodiment of the present invention;

[0082] FIG5 is a schematic structural diagram of a dress code identification device according to an embodiment of the present invention;

[0083] FIG6 is a schematic structural diagram of a pedestrian re-identification model training device according to an embodiment of the present invention;

[0084] FIG7 is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0085] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0086] Referring to FIG1 , an embodiment of the present invention provides a method for determining a dress code, including:

[0087] Step 11: performing human body detection on the image to be identified to obtain a first human body detection image of the target pedestrian in the image to be identified;

[0088] In an embodiment of the present invention, a variety of algorithms can be used to perform human detection on the image to be identified. For example, a target detection algorithm such as the yolo-v5 algorithm is used to detect the target pedestrian, and a target tracking algorithm such as the sort algorithm is used to track the target.

[0089] In the embodiment of the present invention, the image to be recognized may be an image in a surveillance video stream captured by a camera device in a preset location (such as a factory, a bank lobby, etc.).

[0090] In the embodiment of the present invention, the number of target pedestrians in the image to be identified may be one or more, and each target pedestrian may be assigned a unique tracking ID.

[0091] In an embodiment of the present invention, it is optional to detect specific pedestrians in the image to be identified. For example, for an image in a surveillance video stream of a bank lobby, only bank staff members appearing in the image can be detected. In this case, it is necessary to pre-establish a target pedestrian database, which stores facial images of one or more target pedestrians. When performing human body detection on the image to be identified, facial recognition is first performed on the detected pedestrian based on the target pedestrian database. If the pedestrian in the image to be identified is identified as a pedestrian in the target pedestrian database, subsequent steps are executed. If the pedestrian in the image to be identified is not identified as a pedestrian in the target pedestrian database, the process ends.

[0092] In the embodiment of the present invention, optionally, all pedestrians in the image to be identified may be taken as target pedestrians for subsequent dress code identification, for example, in a factory or other place where outsiders are generally not allowed to enter.

[0093] Step 12: performing human body region division on the first human body detection image to obtain a target human body region image of the target pedestrian;

[0094] In an embodiment of the present invention, the target human body region may be one or more, for example, including a head region, an upper body region, and a lower body region, etc. That is, a first human body detection image may be divided into one or more target human body region images.

[0095] Step 13: Using a person re-identification model to perform feature extraction on the target human body region image of the target pedestrian to obtain a first feature vector;

[0096] In an embodiment of the present invention, optionally, when the first human detection image corresponds to multiple target human region images, the multiple target human region images may be first stitched together to obtain a stitched image and input into the person re-identification model. Alternatively, stitching may be omitted and the multiple target human region images may be input into the person re-identification model separately.

[0097] Step 14: Compare the first feature vector with the second feature vector of the target human body region in the dress code sample image to obtain a comparison result;

[0098] In an embodiment of the present invention, if the first human detection image corresponds to multiple target human region images, each target human region image of the first human detection image may be compared with the corresponding target human region image in the dress code sample image.

[0099] For example, assuming that the target human body region image includes: a head region image, an upper body region image, and a lower body region image, then the first feature vector of the head region image in the first human body detection image can be compared with the second feature vector of the head region in the dress code sample image, the first feature vector of the upper body region image in the first human body detection image can be compared with the second feature vector of the upper body region in the dress code sample image, and the first feature vector of the lower body region image in the first human body detection image can be compared with the second feature vector of the lower body region in the dress code sample image, to obtain three comparison results.

[0100] Step 15: Determine whether the target pedestrian's target body area is dressed in a standard manner based on the comparison result.

[0101] In an embodiment of the present invention, before using the pedestrian re-identification module to identify the clothing of the target pedestrian in the image to be identified, the human body detection image in the image to be identified is divided into human body regions to obtain the human body region image of the target pedestrian. The pedestrian re-identification module is used to identify each human body region image separately, so as to accurately identify the standard situation of the clothing of each part of the human body and improve the recognition accuracy.

[0102] In the embodiment of the present invention, optionally, dividing the first human body detection image into human body regions includes:

[0103] Step 121: extracting human key points from the first human detection image using a preprocessing model to obtain first human key points;

[0104] In the embodiment of the present invention, the key points of the human body may include, for example, the top of the head, the neck, the limbs and other key points of the human body.

[0105] Step 122: Divide the first human body detection image into human body areas according to the first human body key points.

[0106] In the embodiment of the present invention, optionally, before comparing the first feature vector with the second feature vector of the target human body region in the dress code sample image, the method further includes:

[0107] Step 01: Perform human body detection on the dress code sample image to obtain a second human body detection image;

[0108] Step 02: performing human body region segmentation on the second human body detection image to obtain a target human body region image in the dress code sample image;

[0109] Step 03: Use the pedestrian re-identification model to extract features of the target human body area image in the dress code sample image to obtain the second feature vector.

[0110] In this embodiment of the present invention, before using a person re-identification model to determine the dress code of a target pedestrian in an image to be identified, dress code sample images are identified. Different dress code sample images are pre-identified for different locations. When the application location changes, there is no need to re-collect training images to train the person re-identification model; only the dress code sample images need to be replaced.

[0111] In an embodiment of the present invention, optionally, comparing the first feature vector with the second feature vector of the target human body area in the dress code sample image to obtain a comparison result includes: calculating the cosine similarity between the first feature vector and the second feature vector of the target human body area in the dress code sample image, and obtaining similarity information as the comparison result.

[0112] Of course, in some other embodiments of the present invention, the calculation of similarity is not limited to using cosine similarity, and other methods can also be used to calculate the similarity.

[0113] In an embodiment of the present invention, optionally, determining whether the target body area of ​​the target pedestrian is dressed in standard manner based on the comparison result includes: for a target body area of ​​the target pedestrian, if the comparison result of N frames containing the image to be identified of the target pedestrian indicates: there are at least M frames of the first feature vector of the target body area in the image to be identified and the second feature vector of the target body area in the dress standard sample image whose similarity information does not reach a preset threshold, determining that the dress of the target body area is not standard; wherein N is a positive integer greater than or equal to 1, and M is a positive integer greater than or equal to 1 and less than N.

[0114] For example, 5 (i.e., N) frames of images containing the target pedestrian to be identified can be identified, and each target human body region (such as the head region, upper body region, and lower body region) in each frame of the image to be identified can be compared with the corresponding target human body region in the dress code sample image. Assuming that the similarity information of 3 (i.e., M) frames of the head region image does not reach a preset threshold (such as 0.45), it is determined that the dress of the target pedestrian's head region is not standard.

[0115] In the embodiment of the present invention, optionally, after determining whether the target pedestrian's target body region is dressed properly based on the comparison result, the method further includes: outputting a warning message if the comparison result indicates that the target body region is dressed improperly. For example, the output warning message may be: Zhang San's hat is not worn properly.

[0116] The following describes the training method of the person re-identification model in the above embodiment.

[0117] Referring to FIG. 2 , an embodiment of the present invention further provides a person re-identification model training method, including:

[0118] Step 21: determining a plurality of training image pairs, each of the training image pairs including at least two training images;

[0119] Step 22: performing human body detection on the training image in the training image pair to obtain a third human body detection image in the training image;

[0120] Step 23: performing human body region division on the third human body detection image to obtain a target human body region image;

[0121] Step 24: using the pedestrian re-identification model to be trained to perform feature extraction on the target human body region image to obtain a third feature vector;

[0122] Step 25: comparing the third eigenvectors of the target human body region of each training image in the training image pair to obtain a comparison result;

[0123] Step 26: Optimize the person re-identification model to be trained according to the comparison result to obtain a trained person re-identification model.

[0124] Please refer to Figure 3, which is a schematic diagram of a training method for a person re-identification model according to an embodiment of the present invention. The input of the person re-identification model is a spliced ​​image of the target human region images corresponding to the training image (i.e., a full-body image). In this embodiment of the present invention, three target human region images are included: a head region image, an upper-body region image, and a lower-body region image. Feature 1, feature 2, and feature 3 are feature vectors extracted from the head region image, the upper-body region image, and the lower-body region image, respectively, i.e., the head, upper-body, and lower-body feature vectors. FC-Total refers to a fully connected network layer that outputs the full-body feature vector extracted from the full-body image.

[0125] In an embodiment of the present invention, a person re-identification model can optionally be trained using a combination of triplet loss and softmax loss. This combination involves calculating a weighted average of multiple loss methods, taking the derivative of the weighted average with respect to the input, and then using gradient descent to update network parameters for training.

[0126] In the embodiment of the present invention, optionally, dividing the third human body detection image into human body regions includes:

[0127] Step 231: extracting human key points from the third human detection image using a preprocessing model to obtain third human key points;

[0128] In the embodiment of the present invention, the key points of the human body may include, for example, the top of the head, the neck, the limbs, etc. The number of key points of the human body may be set as needed, for example, 21 key points of the human body.

[0129] Step 232: Divide the third human body detection image into human body areas according to the third human body key points.

[0130] For example, the human body detection image can be divided into three regions: head, upper body, and lower body according to the key points of the human body.

[0131] Taking the head as an example: select the left and right ear keypoints as reference points and crop the [center_x–d:center_x+d, center_y–d:center_y+d] region from the entire image. Center_x and center_y are the horizontal and vertical coordinates of the midpoint of the line connecting the two keypoints, respectively, and d is the distance between the two keypoints. Rotate the cropped image so that the line connecting the left and right ear keypoints is horizontal. Resize the rotated image to a predetermined size (e.g., 128*128) as the head region image.

[0132] The same method can be used to process the upper body and lower body regions to obtain three images of predetermined sizes. The three images of predetermined sizes are spliced ​​to obtain a full-body image (e.g., 384*128 size) as the input of the pedestrian re-identification model.

[0133] In an embodiment of the present invention, before training a pedestrian re-identification model, human body key points are extracted through a preprocessing model, the human body detection image is divided into regions according to the human body key points, and the divided human body region images are used to train the pedestrian re-identification model, thereby obtaining a pedestrian re-identification model that can target each human body region, improving the accuracy of the pedestrian re-identification model, and at the same time obtaining clothing information of different parts of the human body, thereby enhancing the pedestrian re-identification model's ability to recognize detailed information.

[0134] In the embodiment of the present invention, optionally, determining a plurality of training image pairs includes:

[0135] Step 211: extracting human attribute information from the candidate image using a preprocessing model to obtain the human attribute information of the candidate image;

[0136] In the embodiment of the present invention, optionally, the human attribute information may include, for example, at least one of the following: top color, bottom color, human orientation, whether wearing a hat, whether being blocked, etc.

[0137] Step 212: selecting training images from the candidate images according to the human body attribute information to form the training image pairs.

[0138] In an embodiment of the present invention, before training the pedestrian re-identification model, the human attribute information of the candidate image is extracted through a preprocessing model, and training images are selected based on the human attribute information of the candidate image, so that the required training images can be obtained according to different needs, such as performing difficult sample mining to improve the accuracy of the pedestrian re-identification model.

[0139] It should be noted that the preprocessing model for extracting key points of the human body and the preprocessing model for extracting human attribute information in the embodiment of the present invention may be different models or may be integrated into the same model.

[0140] Please refer to Figure 4, which is a structural diagram of the preprocessing model of an embodiment of the present invention. The preprocessing model includes conv (convolutional neural network), which is the backbone of the preprocessing model and is used to extract feature vectors from the input image. Conv can choose structures such as resnet50 or mobilenetv2; the two network branches in the dotted box are respectively a human key point extraction network and a human attribute information extraction network. The human key point extraction network extracts human key points from the feature vector extracted by conv. Optionally, the last layer of the human key point extraction network can be an N1 (for example, 4 2) A fully connected layer with N1 / 2 dimensions (for example, 21) outputs the horizontal and vertical coordinates of N1 / 2 (for example, 21) key points of the human body; the human attribute information extraction network extracts the human attribute information from the feature vector extracted by conv. Optionally, the last layer of the human attribute information extraction network can be a fully connected layer with N2 dimensions, each dimension being the binary classification result of a certain human attribute, such as whether the top is red, whether the top is green, whether the human body is facing forward, whether the human body is facing backward, etc. N2 is 24, for example, and the binary classification results may include: 8 top colors, 8 bottom colors, 4 human body orientations, whether wearing a hat, whether the head is blocked, whether the upper body is blocked, and whether the lower body is blocked.

[0141] The advantage of integrating the human key point extraction network and the human attribute information extraction network into one model is that the human key point extraction network and the human attribute information extraction network can share the same conv.

[0142] The following describes in detail the method for selecting training images based on the human attribute information of candidate images.

[0143] In the embodiment of the present invention, the training image pair can optionally be a triple, for example, [img, img+, img-], where img+ is a different image containing the same person as img, and img- is a different image containing a different person than img. Of course, in other embodiments of the present invention, the training image pair is not limited to a triple.

[0144] In an embodiment of the present invention, optionally, the human body attribute information includes human body orientation, and selecting training images from the candidate images according to the human body attribute information to form the training image pair includes: selecting training images of the same person with the same and / or different orientations from the candidate images as training images in the training image pair.

[0145] Generally, image pairs of the same person with different body orientations can be used as difficult samples. In an embodiment of the present invention, the two most difficult directions are front and back, and left and right; the four second most difficult directions are front and left, front and right, back and left, and back and right; and the image pairs with the same orientation have the lowest difficulty. For a given training image img, images are extracted from the directions with the highest difficulty, the second highest difficulty, and the lowest difficulty to form positive image pairs with img for training.

[0146] That is, selecting training images of the same person in the same and / or different orientations from the candidate images as training images in the training image pair includes:

[0147] For a training image, from multiple candidate images containing the same person, select a first image of a first difficulty with a first probability, select a second image of a second difficulty with a second probability, and select a third image of a third difficulty with a third probability as the training images in the training image pair;

[0148] Optionally, the first probability>the second probability>the third probability; for example, the first probability is 50%, the second probability is 30%, and the third probability is 20%;

[0149] The first difficulty level refers to: the person in one of the training image and the first image is facing forward, and the person in the other image is facing backward; or the person in one of the training image and the first image is facing left, and the person in the other image is facing right;

[0150] The second difficulty level refers to: the person in one of the training image and the second image is facing forward, and the person in the other image is facing left or right; or the person in one of the training image and the first image is facing backward, and the person in the other image is facing left or right;

[0151] The third difficulty level means that the characters in the training image and the third image are oriented in the same direction.

[0152] In some other embodiments of the present invention, the first probability, the second probability and the third probability may be the same, or partially the same.

[0153] Generally, a pair of pedestrian images of different people wearing the same color clothing and facing the same direction can be used as difficult samples, and color has a greater impact on recognition difficulty than direction. In this embodiment of the present invention, optionally, the human attribute information includes human direction and clothing color, and selecting training images from the candidate images based on the human attribute information to form the training image pairs includes: selecting training images of different people wearing the same clothing and facing the same direction from the candidate images as the training images in the training image pairs.

[0154] In the embodiment of the present invention, optionally, selecting training images of different persons wearing clothing of the same color and facing the same direction from the candidate images as training images in the training image pair includes:

[0155] For a training image, selecting a candidate image containing a different person than the training image;

[0156] Calculating similarity information between the training image and the candidate image, wherein the similarity information is determined by at least one of the following: clothing color, hat wearing, and orientation of the person in the training image and the candidate image;

[0157] Dividing the candidate images into multiple sets according to the similarity information of the candidate images;

[0158] Candidate images are selected from different sets as training images in the training image pairs.

[0159] Optionally, candidate images are selected from different sets with different probabilities as training images in the training image pairs.

[0160] Optionally, the following formula can be used to calculate the similarity information between the training image img containing different people and the candidate image:

[0161] score=3.5*same upper +3.5*same down +2*same hat +same direction

[0162] Among them, same upper Indicates whether the colors of the tops are the same (0 for different, 1 for the same), same downIndicates whether the bottom color is the same (0 for different, 1 for the same), same hat Indicates whether the hats are the same (0 for different, 1 for the same), same direction Indicates whether the pedestrians are facing the same direction (0 for different, 1 for the same).

[0163] The candidate images containing different characters can be divided into several sets according to the similarity information, and the candidate images are extracted from different sets with a probability of score / 10 to form negative image pairs with img for training.

[0164] In this embodiment of the present invention, it is also possible to mine difficult samples without using a preprocessing model. As an alternative to mining difficult samples containing different people, since clothing color has the most significant impact on sample difficulty, this embodiment considers mining difficult samples containing different people using only color information to improve training speed.

[0165] That is, optionally, determining a plurality of training image pairs includes:

[0166] For a training image, selecting a candidate image containing a different person than the training image;

[0167] Performing human body detection on the candidate image to obtain a fourth human body detection image in the candidate image;

[0168] extracting human key points from the fourth human body detection image using a preprocessing model to obtain second human body key points;

[0169] dividing the fourth human body detection image into human body regions according to the second human body key points to obtain a target human body region image in the fourth human body detection image;

[0170] Determine the mean and variance of the target human body region image in the fourth human body detection image on the RGB three channels;

[0171] Convert the target human body region image in the fourth human body detection image into HSV space, and calculate the mean and variance of the HSV three channels after conversion to HSV space; that is, obtain a total of 12-dimensional features;

[0172] According to the mean and variance of the three RGB channels and the mean and variance of the three HSV channels, the image features after dimensionality reduction are obtained; optionally, the t-SNE method can be used to reduce the dimensionality of the 12-dimensional features to 2 dimensions;

[0173] Clustering the image features after dimensionality reduction of the plurality of candidate images, and dividing the plurality of candidate images into different clusters; optionally, a DBSCAN method may be used to cluster the image features after dimensionality reduction;

[0174] Candidate images are selected from different clusters as training images in the training image pairs.

[0175] Optionally, candidate images are selected from different clusters with different probabilities as training images in the training image pairs. For example, images belonging to different clusters and the same cluster as the training image are selected with probabilities of 80% and 20% to form negative sample pairs for training.

[0176] In the embodiment of the present invention, optionally, the method may further include: training a preprocessing model.

[0177] In an embodiment of the present invention, when the human key point extraction network and the human attribute information extraction network are different preprocessing models, the two preprocessing models are trained separately. When the human key point extraction network and the human attribute information extraction network are integrated into one preprocessing model, the human key point extraction network and the human attribute information extraction network are trained separately to obtain the losses of the two. The losses of the two are then combined (added or weighted added, etc.) to calculate the gradient and update the conv parameters.

[0178] In an embodiment of the present invention, optionally, wing loss is used to train the human body key point extraction network in the preprocessing model.

[0179] In the embodiment of the present invention, optionally, BCE Loss is used to train the human attribute information extraction network in the preprocessing model.

[0180] In the above embodiments of the present invention, in addition to using human key points to divide the human body image into regions, a segmentation model can also be used to divide the human body image into regions.

[0181] Referring to FIG5 , an embodiment of the present invention further provides a dress code determination device 50, comprising:

[0182] A first human detection module 51 is configured to perform human detection on the image to be identified, and obtain a first human detection image of a target pedestrian in the image to be identified;

[0183] A first human body region division module 52 is configured to perform human body region division on the first human body detection image to obtain a target human body region image of the target pedestrian;

[0184] A first feature extraction module 53 is configured to extract features from the target human body region image of the target pedestrian using a pedestrian re-identification model to obtain a first feature vector;

[0185] A first comparison module 54 is configured to compare the first feature vector with a second feature vector of the target human body region in the dress code sample image to obtain a comparison result;

[0186] The judging module 55 is configured to determine whether the target pedestrian's target body area is dressed in a standard manner based on the comparison result.

[0187] Optionally, the first human body region division module 52 is used to extract human body key points from the first human body detection image using a preprocessing model to obtain first human body key points; and divide the first human body detection image into human body regions according to the first human body key points.

[0188] Optionally, the dress code determination device 50 further includes:

[0189] A second human body detection module is used to perform human body detection on the dress code sample image to obtain a second human body detection image;

[0190] a second human body region segmentation module, configured to perform human body region segmentation on the second human body detection image to obtain a target human body region image in the dress code sample image;

[0191] The second feature extraction module is used to extract features of the target human body area image in the dress code sample image using the pedestrian re-identification model to obtain the second feature vector.

[0192] Optionally, the first comparison module 54 is configured to calculate the cosine similarity between the first feature vector and a second feature vector of the target human body region in the dress code sample image, and obtain similarity information as the comparison result.

[0193] Optionally, the discrimination module 55 is used to determine that the dress code of a target human body region of the target pedestrian is not standard if the comparison result of N frames of images to be identified containing the target pedestrian indicates that: the similarity information between the first eigenvector of the target human body region in at least M frames of the images to be identified and the second eigenvector of the target human body region in the dress code sample image does not reach a preset threshold; wherein N is a positive integer greater than or equal to 1, and M is a positive integer greater than or equal to 1 and less than N.

[0194] Referring to FIG. 6 , an embodiment of the present invention further provides a person re-identification model training device 60, comprising:

[0195] A determination module 61 is configured to determine a plurality of training image pairs, each of which includes at least two training images;

[0196] A third human body detection module 62 is configured to perform human body detection on the training image in the training image pair to obtain a third human body detection image in the training image;

[0197] A third human body region division module 63 is configured to divide the third human body detection image into human body regions to obtain a target human body region image;

[0198] a third feature extraction module 64 for extracting features from the target human body region image using a person re-identification model to be trained to obtain a third feature vector;

[0199] A second comparison module 65 is configured to compare the third eigenvectors of the target human body region of each training image in the training image pair to obtain a comparison result;

[0200] The optimization module 66 is configured to optimize the person re-identification model to be trained according to the comparison result to obtain a trained person re-identification model.

[0201] Optionally, the third human body region division module 63 is used to extract human body key points from the third human body detection image using a preprocessing model to obtain third human body key points; and divide the third human body detection image into human body regions according to the third human body key points.

[0202] Optionally, the determination module 61 is configured to extract human attribute information from the candidate image using a preprocessing model to obtain the human attribute information of the candidate image; and select training images from the candidate image according to the human attribute information to form the training image pair.

[0203] Optionally, the human attribute information includes human orientation, and the determining module 61 is configured to select training images of the same person with the same and / or different orientations from the candidate images as training images in the training image pair.

[0204] Optionally, the human attribute information includes human orientation and clothing color, and the determination module 61 is configured to select training images of different persons with the same orientation and clothing color from the candidate images as training images in the training image pair.

[0205] Optionally, the determination module 61 is used to select, for a training image, a first image of a first difficulty with a first probability, a second image of a second difficulty with a second probability, and a third image of a third difficulty with a third probability from multiple candidate images containing the same person, as training images in the training image pair; the first difficulty refers to: the person in one of the training image and the first image is facing forward, and the person in the other is facing backward, or the person in one of the training image and the first image is facing left, and the person in the other is facing right; the second difficulty refers to: the person in one of the training image and the second image is facing forward, and the person in the other is facing left or right, or the person in one of the training image and the first image is facing backward, and the person in the other is facing left or right; the third difficulty refers to: the person in the training image and the third image has the same orientation.

[0206] Optionally, the determination module 61 is used to select, for a training image, a candidate image containing a different person from the training image; calculate similarity information between the training image and the candidate image, the similarity information being determined by at least one of the following: clothing color, hat wearing, and person orientation of the person in the training image and the candidate image; divide the candidate images into multiple sets based on the similarity information of the candidate images; and select candidate images from different sets as training images in the training image pair.

[0207] Optionally, the determination module 61 is used to select, for a training image, a candidate image that contains a different person from the training image; perform human body detection on the candidate image to obtain a fourth human body detection image in the candidate image; use a preprocessing model to extract human body key points on the fourth human body detection image to obtain a second human body key point; divide the fourth human body detection image into human body areas based on the second human body key point to obtain a target human body area image in the fourth human body detection image; determine the mean and variance of the target human body area image in the fourth human body detection image on the RGB three channels; convert the target human body area image in the fourth human body detection image to the HSV space, and calculate the mean and variance on the HSV three channels after conversion to the HSV space; obtain the image features after dimensionality reduction based on the mean and variance on the RGB three channels and the mean and variance on the HSV three channels; cluster the image features after dimensionality reduction of multiple candidate images, and divide the multiple candidate images into different cluster clusters; select candidate images from different cluster clusters as training images in the training image pair.

[0208] Please refer to Figure 7. An embodiment of the present invention further provides an electronic device 70, including a processor 71, a memory 72, and a computer program stored in the memory 72 and executable on the processor 71. When the computer program is executed by the processor 71, the various processes of the above-mentioned dress code identification method or pedestrian re-identification model training method embodiment are implemented, and the same technical effects can be achieved. To avoid repetition, they will not be described here.

[0209] The present invention also provides a non-transitory computer-readable storage medium that stores a computer program. When the computer program is executed by a processor, the computer program implements the various processes of the above-mentioned dress code discrimination method or pedestrian re-identification model training method embodiments, and can achieve the same technical effects. To avoid repetition, the above-mentioned non-transitory computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0210] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0211] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0212] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.

Claims

1. A dress code identification method, characterized in that: include: Performing human body detection on the image to be identified to obtain a first human body detection image of a target pedestrian in the image to be identified; Performing human body region segmentation on the first human body detection image to obtain a target human body region image of the target pedestrian; Using a person re-identification model to perform feature extraction on the target human body region image of the target pedestrian to obtain a first feature vector; Comparing the first feature vector with a second feature vector of a target human body region in the dress code sample image to obtain a comparison result; Determine whether the clothing of the target human body area of ​​the target pedestrian is standard according to the comparison result.

2. The method according to claim 1, characterized in that The dividing the first human body detection image into human body regions includes: Extracting human key points from the first human detection image using a preprocessing model to obtain first human key points; The first human body detection image is divided into human body areas according to the first human body key points.

3. The method according to claim 1, characterized in that Before comparing the first feature vector with the second feature vector of the target human body region in the dress code sample image, the method further includes: Performing human body detection on the dress code sample image to obtain a second human body detection image; performing human body region segmentation on the second human body detection image to obtain a target human body region image in the dress code sample image; The pedestrian re-identification model is used to extract features of the target human body area image in the dress code sample image to obtain the second feature vector.

4. The method according to claim 1, wherein The comparing the first feature vector with the second feature vector of the target human body region in the dress code sample image to obtain a comparison result includes: The cosine similarity between the first feature vector and the second feature vector of the target human body region in the dress code sample image is calculated to obtain similarity information as the comparison result.

5. The method according to claim 1, wherein The determining, based on the comparison result, whether the clothing of the target human body area of ​​the target pedestrian is standard includes: For a target human body region of the target pedestrian, if the comparison results of N frames of images to be identified containing the target pedestrian indicate that similarity information between a first feature vector of the target human body region in at least M frames of the images to be identified and a second feature vector of the target human body region in the dress code sample image does not reach a preset threshold, it is determined that the dress code of the target human body region is not standardized; Wherein, N is a positive integer greater than or equal to 1, and M is a positive integer greater than or equal to 1 and less than N.

6. A person re-identification model training method, characterized in that: include: Determining a plurality of training image pairs, each of the training image pairs including at least two training images; Performing human body detection on the training image in the training image pair to obtain a third human body detection image in the training image; Performing human body region division on the third human body detection image to obtain a target human body region image; Using the pedestrian re-identification model to be trained to perform feature extraction on the target human body area image to obtain a third feature vector; Comparing the third eigenvectors of the target human body region of each training image in the training image pair to obtain a comparison result; The pedestrian re-identification model to be trained is optimized according to the comparison result to obtain a trained pedestrian re-identification model.

7. The method according to claim 6, characterized in that The dividing the third human body detection image into human body regions comprises: Extracting human key points from the third human body detection image using a preprocessing model to obtain third human body key points; The third human body detection image is divided into human body areas according to the third human body key points.

8. The method according to claim 6, characterized in that Determining a plurality of training image pairs comprises: Extracting human attribute information from the candidate image using a preprocessing model to obtain the human attribute information of the candidate image; A training image is selected from the candidate images according to the human body attribute information to form the training image pair.

9. The method according to claim 8, characterized in that The human body attribute information includes human body orientation, and selecting training images from the candidate images according to the human body attribute information to form the training image pair includes: selecting training images of the same person with the same and / or different orientations from the candidate images as the training images in the training image pair; and / or The human body attribute information includes human body orientation and clothing color. Selecting training images from the candidate images according to the human body attribute information to form the training image pair includes: selecting training images of different people with the same orientation and the same color clothing from the candidate images as the training images in the training image pair.

10. The method according to claim 9, characterized in that The selecting, from the candidate images, training images of the same person in the same and / or different orientations as training images in the training image pair includes: For a training image, from multiple candidate images containing the same person, select a first image of a first difficulty with a first probability, select a second image of a second difficulty with a second probability, and select a third image of a third difficulty with a third probability as the training images in the training image pair; The first difficulty level refers to: the person in one of the training image and the first image is facing forward, and the person in the other image is facing backward; or the person in one of the training image and the first image is facing left, and the person in the other image is facing right; The second difficulty level refers to: the person in one of the training image and the second image is facing forward, and the person in the other image is facing left or right; or the person in one of the training image and the first image is facing backward, and the person in the other image is facing left or right; The third difficulty level means that the characters in the training image and the third image are oriented in the same direction.

11. The method according to claim 9, characterized in that Selecting training images of different people with the same orientation and the same color clothing from the candidate images as training images in the training image pair includes: For a training image, selecting a candidate image containing a different person than the training image; Calculating similarity information between the training image and the candidate image, wherein the similarity information is determined by at least one of the following: clothing color, hat wearing, and orientation of the person in the training image and the candidate image; Dividing the candidate images into multiple sets according to the similarity information of the candidate images; Candidate images are selected from different sets as training images in the training image pairs.

12. The method according to claim 6, characterized in that Determining a plurality of training image pairs comprises: For a training image, selecting a candidate image containing a different person than the training image; Performing human body detection on the candidate image to obtain a fourth human body detection image in the candidate image; extracting human key points from the fourth human body detection image using a preprocessing model to obtain second human body key points; dividing the fourth human body detection image into human body regions according to the second human body key points to obtain a target human body region image in the fourth human body detection image; Determine the mean and variance of the target human body region image in the fourth human body detection image on the RGB three channels; Convert the target human body region image in the fourth human body detection image into HSV space, and calculate the mean and variance of the HSV three channels after conversion into HSV space; Obtaining the image features after dimensionality reduction according to the mean and variance of the RGB three channels and the mean and variance of the HSV three channels; Clustering the image features after dimensionality reduction of the plurality of candidate images, and dividing the plurality of candidate images into different clusters; Candidate images are selected from different clusters as training images in the training image pairs.

13. A dress code identification device, characterized in that: include: A first human detection module is used to perform human detection on the image to be identified, and obtain a first human detection image of a target pedestrian in the image to be identified; a first human body region division module, configured to perform human body region division on the first human body detection image to obtain a target human body region image of the target pedestrian; A first feature extraction module is used to extract features from the target human body region image of the target pedestrian using a pedestrian re-identification model to obtain a first feature vector; a first comparison module, configured to compare the first feature vector with a second feature vector of a target human body region in a dress code sample image to obtain a comparison result; The discrimination module is used to determine whether the clothing of the target human body area of ​​the target pedestrian is standard according to the comparison result.

14. A pedestrian re-identification model training device, characterized in that: include: a determination module, configured to determine a plurality of training image pairs, each of the training image pairs including at least two training images; a third human body detection module, configured to perform human body detection on the training image in the training image pair to obtain a third human body detection image in the training image; A third human body region division module is used to divide the third human body detection image into human body regions to obtain a target human body region image; a third feature extraction module, configured to extract features from the target human body region image using a person re-identification model to be trained, to obtain a third feature vector; a second comparison module, configured to compare the third eigenvectors of the target human body region of each training image in the training image pair to obtain a comparison result; The optimization module is used to optimize the pedestrian re-identification model to be trained according to the comparison result to obtain a trained pedestrian re-identification model.

15. An electronic device, characterized in that: include: A processor, a memory, and a program stored on the memory and executable on the processor, wherein when the program is executed by the processor, the program implements the steps of the dress code identification method according to any one of claims 1 to 5, or the program implements the steps of the pedestrian re-identification model training method according to any one of claims 6 to 12.

16. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the steps of the dress code discrimination method according to any one of claims 1 to 5; or, when the computer program is executed by the processor, it implements the steps of the pedestrian re-identification model training method according to any one of claims 6 to 12.