A method, device and equipment for pedestrian re-identification and a storage medium

By using a pre-trained pedestrian feature recognition model that combines biometric and clothing features for feature fusion, the problem of poor recognition performance in pedestrian re-identification has been solved, and accurate recognition has been achieved in intelligent security scenarios.

CN114627310BActive Publication Date: 2026-02-27SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210257785.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-16
Publication Date
2026-02-27
Estimated Expiration
2042-03-16

AI Technical Summary

Technical Problem

Existing pedestrian re-identification technologies struggle to accurately identify individuals after they have changed clothes, primarily due to changes in body posture and shooting angles, resulting in unstable extracted body features and poor recognition performance.

Method used

A pre-trained pedestrian feature recognition model for changing clothes is used, which combines biometric features and clothing features. The target person is identified in the base pedestrian image set through a feature fusion sub-model, including a biometric feature recognition sub-model, a clothing feature recognition sub-model, and a feature fusion sub-model.

Benefits of technology

It enables accurate identification of target individuals even when they are changing clothes, reduces labor costs, and improves the accuracy of identification, with particularly significant results in intelligent security scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627310B_ABST
    Figure CN114627310B_ABST
Patent Text Reader

Abstract

The application discloses a clothes-changing pedestrian re-identification method and device, equipment and storage medium. The method comprises the following steps: obtaining a to-be-inquired human body picture, a target clothes template picture and a base library pedestrian picture set; using a pre-trained clothes-changing pedestrian feature recognition model to obtain target pedestrian fusion features corresponding to the to-be-inquired human body picture and the target clothes template picture, and base library pedestrian fusion features corresponding to each base library pedestrian picture in the base library pedestrian picture set, wherein the clothes-changing pedestrian feature recognition model comprises a biological feature recognition sub-model, a clothes feature recognition sub-model and a feature fusion sub-model; determining target pedestrian similarities corresponding to each base library pedestrian picture according to each base library pedestrian fusion feature and the target pedestrian fusion feature; and determining a base library pedestrian picture with a target pedestrian similarity greater than or equal to a preset similarity threshold value as a target pedestrian picture. The application realizes accurate identification of clothes-changing pedestrians and reduces manual searching costs.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image recognition, and in particular to a method and device for recognizing a person who has changed clothes, a device, and a storage medium. BACKGROUND

[0002] The person who has changed clothes refers to finding a picture of a target person who has changed clothes from a set of millions of pictures in a database given a picture of the target person as a query condition. The person re-identification technology plays a huge role in the intelligent security monitoring scene, and it can assist in tasks such as finding lost children and tracking personnel.

[0003] The ordinary person re-identification technology mainly focuses on the features of the clothes of the target person and uses the clothes to "find people". The difficulty of the person who has changed clothes lies in the fact that the target person has changed clothes, and it is difficult to extract features from the human body picture that are not related to the clothes, because the proportion of biological features such as human faces in the human body picture is small and the human face may not be visible. The existing person who has changed clothes re-identification method is mainly focused on how to extract the edge and body shape information of the person, and the body shape is used to "find people", but the body contour information is abstract and difficult to extract, and is easily affected by the shooting angle, the interference of the shielding object, and the change of the human body posture, resulting in poor recognition effect. SUMMARY

[0004] The present application provides a method and device for recognizing a person who has changed clothes, a device, and a storage medium to achieve accurate identification of the person who has changed clothes.

[0005] According to an aspect of the present application, a method for recognizing a person who has changed clothes is provided, the method comprising:

[0006] obtaining a to-be-queried human body picture, a target clothes template picture, and a set of database person pictures;

[0007] using a pre-trained clothes-changing person feature recognition model to obtain a target person fusion feature corresponding to the to-be-queried human body picture and the target clothes template picture, and a database person fusion feature corresponding to each database person picture in the set of database person pictures, wherein the clothes-changing person feature recognition model comprises a biological feature recognition sub-model, a clothes feature recognition sub-model, and a feature fusion sub-model;

[0008] determining a target person similarity corresponding to each database person picture according to each database person fusion feature and the target person fusion feature;

[0009] determining a database person picture with a target person similarity greater than or equal to a preset similarity threshold as a target person picture.

[0010] According to another aspect of the present application, there is provided a clothes-changing pedestrian re-identification device, comprising:

[0011] a data acquisition module configured to acquire a query human body picture, a target clothes template picture and a set of base library pedestrian pictures;

[0012] a feature extraction module configured to acquire target pedestrian fusion features corresponding to the query human body picture and the target clothes template picture, and base library pedestrian fusion features corresponding to each base library pedestrian picture in the set of base library pedestrian pictures, by using a pre-trained clothes-changing pedestrian feature recognition model, wherein the clothes-changing pedestrian feature recognition model comprises a biological feature recognition sub-model, a clothes feature recognition sub-model and a feature fusion sub-model;

[0013] a similarity determination module configured to determine target pedestrian similarities corresponding to each base library pedestrian picture according to each base library pedestrian fusion feature and the target pedestrian fusion features;

[0014] a target determination module configured to determine a base library pedestrian picture as a target pedestrian picture if the target pedestrian similarity corresponding to the base library pedestrian picture is greater than or equal to a preset similarity threshold.

[0015] According to another aspect of the present application, there is provided an electronic device, comprising:

[0016] at least one processor; and

[0017] a memory connected to the at least one processor in communication; wherein

[0018] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the clothes-changing pedestrian re-identification method according to any one of the embodiments of the present application.

[0019] According to another aspect of the present application, there is provided a computer readable storage medium storing computer instructions for enabling a processor to perform the clothes-changing pedestrian re-identification method according to any one of the embodiments of the present application when executed by the processor.

[0020] The technical scheme of the embodiment of the present application comprises the following steps: obtaining a to-be-queried human body picture and a target clothes template picture, and a base library pedestrian picture set; taking the to-be-queried human body picture and the target clothes template picture as input data, performing pedestrian similarity recognition on the base library pedestrian pictures in the base library pedestrian picture set through a pre-trained clothes-changing pedestrian re-identification model, obtaining target pedestrian similarities of the base library pedestrian pictures, wherein the clothes-changing pedestrian re-identification model comprises a feature recognition model and a similarity recognition model; and determining the base library pedestrian picture with a target pedestrian similarity greater than or equal to a preset similarity threshold as a target pedestrian picture. The present application solves the problem that the prior art is affected by the shooting angle and the human body posture change when recognizing the clothes-changing pedestrian by extracting the human body features, and the features are difficult to extract and the recognition effect is poor. The present application realizes accurate recognition of the clothes-changing pedestrian by pre-training the clothes-changing pedestrian re-identification model, extracting features from the input to-be-queried human body picture and the target clothes template picture respectively, and performing clothes-changing pedestrian re-identification on the base library pedestrian picture set according to the fusion features, thereby reducing the labor cost of finding the target pedestrian and having a great effect in intelligent security and the like.

[0021] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0023] Figure 1a is a flowchart of a clothes-changing pedestrian re-identification method according to the first embodiment of the present application;

[0024] Figure 1b is a model training schematic diagram of a clothes-changing pedestrian re-identification method according to the first embodiment of the present application;

[0025] Figure 1c is an application effect diagram of a clothes-changing pedestrian re-identification method according to the first embodiment of the present application;

[0026] Figure 2 is a flowchart of a clothes-changing pedestrian re-identification method according to the second embodiment of the present application;

[0027] Figure 3 is a structural schematic diagram of a clothes-changing pedestrian re-identification device according to the third embodiment of the present application;

[0028] Figure 4 Figure 1 is a structural schematic diagram of an electronic device implementing a method for recognizing a person in a changed clothes according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to make the technical personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0030] It should be noted that the terms "first", "second", "target" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0031] Embodiment one

[0032] Figure 1a A flowchart of a method for recognizing a person in a changed clothes is provided for the first embodiment of the present application. The present embodiment can be applied to the case of recognizing and positioning a person in a changed clothes. The method can be executed by a person in a changed clothes recognition device, which can be realized in the form of hardware and / or software, and can be configured in a computer device.

[0033] The application scenarios of the present embodiment can be finding a lost child, an old person, and positioning a person through monitoring video. For example, in the scenario of finding a lost child, the child's family can provide an arbitrary picture of the child and a template picture of the clothes worn by the child when the child was lost. By using the changed clothes feature recognition model provided by the present embodiment, the picture of the child wearing the specific clothes can be found in the city monitoring picture database. In the scenario of tracking a person, the police can provide an arbitrary picture of the person and a template picture of the clothes described by the latest eyewitness, and quickly locate the image of the person in the monitoring video.

[0034] As Figure 1aAs shown, the method comprises:

[0035] S110, obtaining a to-be-queried human body picture, a target clothes template picture, and a bottom library pedestrian picture set.

[0036] The to-be-queried human body picture can be understood as an arbitrary human body picture of a target person to be queried. The target clothes template picture can be understood as a photo or image of target clothes worn by the target person to be queried. The bottom library pedestrian picture set can be understood as a set containing all material pictures that can be queried. That is, the clothes-changing pedestrian re-identification method of the embodiment is to find a picture taken when the target person in the to-be-queried human body picture wears the target clothes in the target clothes template picture in the bottom library pedestrian picture set.

[0037] S120, using a pre-trained clothes-changing pedestrian feature recognition model to obtain a target pedestrian fusion feature corresponding to the to-be-queried human body picture and the target clothes template picture, and a bottom library pedestrian fusion feature corresponding to each bottom library pedestrian picture in the bottom library pedestrian picture set. The clothes-changing pedestrian feature recognition model comprises a biological feature recognition sub-model, a clothes feature recognition sub-model, and a feature fusion sub-model.

[0038] The clothes-changing pedestrian feature recognition model in the embodiment can be obtained by training a large amount of training data in advance. The clothes-changing pedestrian feature recognition model can comprise a biological feature recognition sub-model, a clothes feature recognition sub-model, and a feature fusion sub-model. The biological feature recognition sub-model can be used to obtain biological features such as the appearance contour of a human body in a picture. The clothes feature recognition sub-model can be used to obtain clothes features of clothes worn by a pedestrian in a picture. The feature fusion sub-model can be used to fuse the biological features and the clothes features to obtain a fusion feature containing more information of the pedestrian in the picture.

[0039] The target pedestrian fusion feature can be understood as a feature value containing biological features of a human body and target clothes features worn by a target person. The bottom library pedestrian fusion feature can be understood as a feature value containing biological features of a human body and clothes features worn by a pedestrian in a bottom library pedestrian picture.

[0040] Specifically, the to-be-queried human body picture, the target clothes template picture, and all bottom library pedestrian pictures in the bottom library pedestrian picture set can be input as input data into the clothes-changing pedestrian feature recognition model. The clothes-changing pedestrian feature recognition model outputs the target pedestrian fusion feature corresponding to the to-be-queried human body picture and the target clothes template picture, and the bottom library pedestrian fusion feature corresponding to each bottom library pedestrian picture.

[0041] Optionally, the training process of the clothes-changing pedestrian feature recognition model in the embodiment can comprise:

[0042] A1, obtain training data containing query human body training pictures, clothes template training pictures and a set of base library pedestrian training pictures, and perform pedestrian recognition annotation on the base library pedestrian training pictures in the set of base library pedestrian training pictures in the same group of training data according to the query human body training pictures and the clothes template training pictures, to obtain the standard recognition result corresponding to each base library pedestrian training picture.

[0043] The training data used for training the model in the embodiment contains query human body training pictures, clothes template training pictures and a set of base library pedestrian training pictures. For a model training process, it is necessary to identify the base library pedestrian training picture in which the pedestrian in the human body training picture in the group of training data wears the clothes in the clothes template training picture in the group of training data, that is, for a group of training data, the query human body training picture contains the pedestrian to be searched this time, the clothes template training picture contains the clothes image worn by the pedestrian to be searched this time, and the set of base library pedestrian training pictures contains the collection of material pictures that can be queried. In order to measure the effect of model training, the base library pedestrian training pictures in the set of base library pedestrian training pictures can be pre-annotated for pedestrian recognition. If the pedestrian in a base library pedestrian training picture is the pedestrian in the query human body training picture wearing the clothes in the clothes template training picture, the base library pedestrian training picture and the query human body training picture are annotated to contain the same pedestrian.

[0044] In the embodiment, a data set containing a large amount of real scene clothes-changing pedestrian data can be selected, and after data preprocessing and other operations, training data containing query human body training pictures, clothes template training pictures and a set of base library pedestrian training pictures are obtained, which are used to train the clothes-changing pedestrian feature recognition model to be trained. In the case where the amount of real data is insufficient, a large amount of simulated clothes-changing data can be made as training data through game simulation means. The simulation data can contain different background conditions, shooting angle conditions, weather conditions, shooting distance conditions, etc. The simulation data has low production cost and strong controllability, and a large amount of clothes-changing pedestrian data with very rich conditions can be generated in a short time, so that more and richer learning samples participate in the training process of the model. The combination of real data and simulation data can make the model exhibit higher generalization in real application scenarios.

[0045] A2, input the training data into the clothes-changing pedestrian feature recognition model to be trained, and obtain the output query biological feature, query fusion feature, base library biological feature and base library fusion feature corresponding to each base library pedestrian training picture, wherein the clothes-changing pedestrian feature recognition model to be trained includes a biological feature recognition sub-model to be trained, a clothes feature recognition sub-model to be trained and a feature fusion sub-model to be trained.

[0046] The to-be-trained clothing-changing pedestrian feature recognition model in this embodiment can include a to-be-trained biological feature recognition sub-model, a to-be-trained clothing feature recognition sub-model, and a to-be-trained feature fusion sub-model. The to-be-trained biological feature recognition sub-model, the to-be-trained clothing feature recognition sub-model, and the to-be-trained feature fusion sub-model can be built using a network structure such as a transformer or a convolutional neural network. The number of layers of the network can be set according to the amount of training data. The larger the amount of data, the deeper the number of layers can be.

[0047] Specifically, the training data can be used as input data, and the to-be-trained clothing-changing pedestrian feature recognition model can be used for recognition, to output the query biological feature, the query fusion feature, the gallery biological feature, and the gallery fusion feature corresponding to each group of training data.

[0048] Further, step A2 can be implemented through the following specific steps:

[0049] A21, input the query human training picture into the to-be-trained biological feature recognition sub-model to output the query biological feature intermediate value and the query biological feature, and input each gallery pedestrian training picture into the to-be-trained biological feature recognition sub-model to output the corresponding gallery biological feature training intermediate value and the gallery biological feature.

[0050] The query biological feature and the gallery biological feature can be understood as the model output result of the to-be-trained biological feature recognition sub-model; the query biological feature intermediate value and the gallery biological feature training intermediate value can be understood as the output result of the intermediate layer of the to-be-trained biological feature recognition sub-model.

[0051] A22, input the clothing template training picture into the to-be-trained clothing feature recognition sub-model to output the query clothing feature intermediate value, perform clothing detection on each gallery pedestrian training picture to obtain the corresponding gallery clothing training picture, and input each gallery clothing training picture into the to-be-trained clothing feature recognition sub-model to output the corresponding gallery clothing feature training intermediate value.

[0052] The query clothing feature intermediate value and the gallery clothing feature training intermediate value can be understood as the output result of the intermediate layer of the to-be-trained clothing feature recognition sub-model.

[0053] A23, input the query biological feature intermediate value and the query clothing feature intermediate value into the to-be-trained feature fusion sub-model to output the query fusion feature, and input each gallery biological feature training intermediate value and the corresponding gallery clothing feature training intermediate value into the to-be-trained feature fusion sub-model to output the corresponding gallery fusion feature.

[0054] In the embodiment, the output results of the model intermediate layers of the to-be-trained biological feature recognition sub-model and the output results of the model intermediate layers of the to-be-trained clothing feature recognition sub-model can be fused to obtain the fused features. Therefore, the query human training picture can be input into the to-be-trained biological feature recognition sub-model, the clothing template training picture can be input into the to-be-trained clothing feature recognition sub-model, and then the query biological feature intermediate value output by the to-be-trained biological feature recognition sub-model and the query clothing feature intermediate value output by the to-be-trained clothing feature recognition sub-model can be input into the to-be-trained feature fusion sub-model to obtain the query fused features. For any one base library pedestrian training picture, the base library pedestrian training picture can be input into the to-be-trained biological feature recognition sub-model, the clothing detection technology can be used to perform clothing detection on the base library pedestrian training picture, the obtained base library clothing training picture can be input into the to-be-trained clothing feature recognition sub-model, and then the base library biological feature training intermediate value output by the to-be-trained biological feature recognition sub-model and the base library clothing feature training intermediate value output by the to-be-trained clothing feature recognition sub-model can be input into the to-be-trained feature fusion sub-model to obtain the base library fused features.

[0055] In addition, the query biological features and the base library biological features output by the to-be-trained biological feature recognition sub-model are used to provide data support for detecting the recognition effect of the to-be-trained biological feature recognition sub-model.

[0056] A3, by performing feature classification on the query biological features, the query fused features, the base library biological features and the base library fused features, the training biological feature recognition results and the training fused feature recognition results corresponding to the base library pedestrian training pictures are obtained.

[0057] Further, step A3 can be implemented through the following specific steps:

[0058] A31, using a first classifier to perform feature classification on the query biological features and the base library biological features, to obtain the query biological feature classification result and the base library biological feature classification result corresponding to each base library biological feature, and determining the training biological feature recognition result corresponding to each base library biological feature according to the base library biological feature classification result and the query biological feature classification result.

[0059] A32, using a second classifier to perform feature classification on the query fused features and the base library fused features, to obtain the query fused feature classification result and the base library fused feature classification result corresponding to each base library fused feature, and determining the training fused feature recognition result corresponding to each base library fused feature according to the base library fused feature classification result and the query fused feature classification result.

[0060] The first classifier and the second classifier can be selected according to specific scenes.

[0061] Specifically, the feature classifier can be used to classify the query biological features, the query fusion features, the biological features of each database, and the fusion features of each database. If a biological feature of a database and a query biological feature are classified into the same class, the training biological feature recognition result corresponding to the training image of the pedestrian in the database can be marked as being in the same class as the query biological feature. Similarly, if a fusion feature of a database and a query fusion feature are classified into the same class, the training fusion feature recognition result corresponding to the training image of the pedestrian in the database can be marked as being in the same class as the query fusion feature.

[0062] A4, the standard recognition result, the training biological feature recognition result, and the training fusion feature recognition result are substituted into the given at least two loss function expressions, respectively, to obtain corresponding loss functions.

[0063] Specifically, by comparing the standard recognition result and the training biological feature recognition result, the feature recognition effect of the to-be-trained biological feature recognition sub-model can be reflected. A suitable classification loss function expression can be preselected to obtain a corresponding loss function. By comparing the standard recognition result and the training fusion feature recognition result, the feature recognition effect of the to-be-trained biological feature recognition sub-model and the to-be-trained clothes feature recognition sub-model, and the feature fusion effect of the to-be-trained feature fusion sub-model can be reflected. A suitable classification loss function expression or a triple loss function expression can be preselected to obtain a corresponding loss function.

[0064] A5, the to-be-trained clothes-changing pedestrian feature recognition model is trained by using the loss functions, to obtain a clothes-changing pedestrian feature recognition model.

[0065] Specifically, after obtaining the loss functions, the to-be-trained clothes-changing pedestrian feature recognition model can be trained by using the loss functions, and the model parameters are continuously adjusted, to finally obtain a clothes-changing pedestrian feature recognition model.

[0066] An exemplary clothes-changing pedestrian re-identification method is shown in FIG. 1. As shown in FIG. 1, the clothes-changing pedestrian re-identification method includes the following steps. Figure 1b An exemplary clothes-changing pedestrian re-identification method is shown in FIG. 1. As shown in FIG. 1, the clothes-changing pedestrian re-identification method includes the following steps. Figure 1bAs shown, the transformer network can be selected to build the to-be-trained biological feature recognition sub-model, the to-be-trained clothes feature recognition sub-model, and the to-be-trained feature fusion sub-model. On the one hand, the query human training picture is input into the to-be-trained biological feature recognition sub-model to obtain the query biological feature; the base library pedestrian training picture is input into the to-be-trained biological feature recognition sub-model to obtain the base library biological feature; the query biological feature and the base library biological feature are combined with the preselected classification loss function expression 1 to obtain the classification loss function 1. On the other hand, the query human training picture is input into the to-be-trained biological feature recognition sub-model, the clothes template training picture is input into the to-be-trained clothes feature recognition sub-model, the query biological feature intermediate value output by the to-be-trained biological feature recognition sub-model and the query clothes feature intermediate value output by the to-be-trained clothes feature recognition sub-model are input into the to-be-trained feature fusion sub-model to obtain the output query fusion feature; the base library pedestrian training picture is input into the to-be-trained biological feature recognition sub-model, the base library clothes training picture is obtained by performing clothes detection on each base library pedestrian training picture using the clothes detector, the base library clothes training picture is input into the to-be-trained clothes feature recognition sub-model, the base library biological feature training intermediate value output by the to-be-trained biological feature recognition sub-model and the base library clothes feature training intermediate value output by the to-be-trained clothes feature recognition sub-model are input into the to-be-trained feature fusion sub-model to obtain the base library fusion feature; the query fusion feature and the base library fusion feature are combined with the preselected classification loss function expression 2 and the triplet loss function expression to obtain the classification loss function 2 and the triplet loss function. The classification loss function 1 can be used for back propagation of the to-be-trained biological feature recognition sub-model, and the classification loss function 2 and the triplet loss function can be used for back propagation of the to-be-trained biological feature recognition sub-model, the to-be-trained clothes feature recognition sub-model, and the to-be-trained feature fusion sub-model, so as to obtain the clothes-changing pedestrian feature recognition model.

[0067] In S130, the target pedestrian similarity corresponding to each base library pedestrian picture is determined according to each base library pedestrian fusion feature and the target pedestrian fusion feature.

[0068] The target pedestrian similarity can be understood as the similarity between the pedestrian in the base library pedestrian picture and the pedestrian in the to-be-queried human picture. In this embodiment, it can be considered that the target pedestrian similarity reflects the possibility that the pedestrian in the base library pedestrian picture and the pedestrian in the to-be-queried human picture are the same person.

[0069] Specifically, the feature similarity of each base library pedestrian fusion feature and the target pedestrian fusion feature can be calculated respectively, and the calculated similarity can be taken as the target pedestrian similarity of the base library pedestrian fusion feature. The calculation method of the feature similarity can be selected as any similarity calculation method, and this embodiment is not limited.

[0070] S140, determine the gallery pedestrian picture corresponding to the target pedestrian similarity greater than or equal to the preset similarity threshold as the target pedestrian picture.

[0071] Specifically, a similarity threshold can be preset, and each target pedestrian similarity is compared with the preset similarity threshold, and if the target pedestrian similarity is greater than or equal to the preset similarity threshold, the corresponding gallery pedestrian picture is determined as the target pedestrian picture.

[0072] The preset similarity threshold can reflect the strictness of feature comparison of the clothes-changing pedestrian re-identification, and the threshold can be set according to a specific application scenario. When the set threshold value is high, it can be considered that the feature comparison is strict, and only when the gallery pedestrian fusion feature and the target pedestrian fusion feature have a high similarity, it is considered that the pedestrian in the gallery pedestrian picture corresponding to the gallery pedestrian fusion feature and the pedestrian in the to-be-queried human body picture corresponding to the target pedestrian fusion feature are the same person; when the set threshold value is low, it can be considered that the feature comparison is loose, and as long as the gallery pedestrian fusion feature and the target pedestrian fusion feature have a certain similarity, it is considered that the pedestrian in the gallery pedestrian picture corresponding to the gallery pedestrian fusion feature and the pedestrian in the to-be-queried human body picture corresponding to the target pedestrian fusion feature are the same person.

[0073] Exemplarily, Figure 1c is an application effect diagram of a clothes-changing pedestrian re-identification method according to Embodiment One of the present application. As shown in Figure 1c the left side is a to-be-queried human body picture and a target clothes template picture provided by a user, and through the clothes-changing pedestrian feature recognition model provided in the present embodiment, the picture taken after the pedestrian in the to-be-queried human body picture changes the clothes in the target clothes template picture can be found in the gallery pedestrian picture set on the right side, that is, the target pedestrian picture shown in the figure.

[0074] Generally, after clothes-changing pedestrian re-identification is performed using a certain method or model, the recognition effect of the clothes-changing pedestrian re-identification method or model can be measured by statistics of two indexes of mAP and Rank1, and Table 1 is a recognition index statistics table of several clothes-changing pedestrian re-identification methods and models.

[0075] Table 1 Recognition index statistics table of various clothes-changing pedestrian re-identification methods and models

[0076] Method and model mAP (%) Rank1 (%) ResNet50ibn 1.2 1.0 Vit 5.5 17.5 Pixel sampling 2.1 11.6 The clothes-changing pedestrian re-identification method provided in the embodiment 25 41.3

[0077] As shown in the table above, the ResNet50ibn and Vit methods suffer from low mAP and Rank1 scores after pedestrian re-identification due to their difficulty in learning and capturing features unrelated to clothing. While the pixel sampling method uses human body analysis as an auxiliary means, guiding the network to focus on body contour features unrelated to clothing by changing the color of clothing in the human image, its performance is also poor because body contour features are relatively abstract and occupy a small proportion of the image, making feature extraction difficult and resulting in overall poor performance. In contrast, the pedestrian re-identification method proposed in this embodiment significantly improves the performance under clothing-changing conditions, achieving mAP and Rank1 scores of 25% and 41.3% respectively in real-world testing scenarios. This is because the method in this embodiment simultaneously inputs the image of the human body and the target clothing template image during feature extraction, extracting the human body's biological features and the target clothing's features through dual channels, and then fusing these features. This results in the fused features containing both the human body's biological features and the target person's clothing features, making it easier to find the target person wearing the target clothing in the base pedestrian image set.

[0078] This invention, in its embodiments, acquires a query image of a human body and a target clothing template image, as well as a database of pedestrian images. Using the query image of the human body and the target clothing template image as input data, a pre-trained pedestrian re-identification model for changing clothes is used to perform pedestrian similarity recognition on the images in the database, obtaining the target pedestrian similarity for each image. The pedestrian re-identification model for changing clothes includes a feature recognition model and a similarity recognition model. Images in the database with a target pedestrian similarity greater than or equal to a preset similarity threshold are identified as target pedestrian images. This invention solves the problem that existing technologies rely on extracting human body shape features to identify pedestrians changing clothes, which is easily affected by changes in shooting angle and human posture, leading to difficulty in feature extraction and poor recognition results. This invention, through a pre-trained pedestrian re-identification model for changing clothes, extracts features from the input query image of the human body and the target clothing template image respectively, and performs pedestrian re-identification for changing clothes in the database of pedestrian images based on the fused features. This achieves accurate identification of pedestrians changing clothes, reduces the manual cost of finding target pedestrians, and has a significant role in intelligent security and other fields.

[0079] Example 2

[0080] Figure 2 This is a flowchart of a pedestrian re-identification method for changing clothes provided in Embodiment 2 of the present invention. This embodiment further optimizes the above-mentioned pedestrian re-identification method for changing clothes based on the previous embodiment. Figure 2 As shown, the method includes:

[0081] S210, acquire a to-be-queried human body picture, a target clothes template picture and a bottom library pedestrian picture set.

[0082] S220, take the to-be-queried human body picture and the target clothes template picture as input data, perform feature extraction and feature fusion through the pre-trained clothes-changing pedestrian feature recognition model, and output target pedestrian fusion features.

[0083] In the embodiment, biological features such as a pedestrian posture and body shape in the to-be-queried human body picture and clothes features in the target clothes template picture can be extracted respectively, the extracted features are fused, and target pedestrian fusion features are obtained. The target pedestrian fusion features contain both biological features of the pedestrian and clothes features after the pedestrian changes clothes, and the accuracy of pedestrian re-identification can be improved.

[0084] Optionally, S220 can be implemented through the following specific steps.

[0085] S2201, biological feature extraction is performed on the to-be-queried human body picture by using a biological feature recognition sub-model to obtain target biological feature intermediate values.

[0086] Specifically, the biological feature recognition sub-model in the clothes-changing pedestrian feature recognition model is mainly used for recognizing biological features of the pedestrian, and therefore, the to-be-queried human body picture can be taken as input data, feature extraction is performed through the biological feature recognition sub-model, and target biological feature intermediate values are output from the network intermediate layer of the biological feature recognition sub-model.

[0087] S2202, clothes feature extraction is performed on the target clothes template picture by using a clothes feature recognition sub-model to obtain target clothes feature intermediate values.

[0088] Specifically, the clothes feature recognition sub-model in the clothes-changing pedestrian feature recognition model is mainly used for recognizing clothes features after the pedestrian changes clothes, and therefore, the target clothes template picture can be taken as input data, feature extraction is performed through the clothes feature recognition sub-model, and target clothes feature intermediate values are output from the network intermediate layer of the clothes feature recognition sub-model.

[0089] S2203, feature fusion is performed on the target biological feature intermediate values and the target clothes feature intermediate values by using a feature fusion sub-model to obtain target pedestrian fusion features.

[0090] Specifically, the feature fusion sub-model in the clothes-changing pedestrian feature recognition model is mainly used for feature fusion of the biological features of the pedestrian and the clothes features of the pedestrian after changing clothes, and therefore the target biological feature intermediate value output by the biological feature recognition sub-model and the target clothes feature intermediate value output by the clothes feature recognition sub-model can be taken as input data and input into the feature fusion sub-model. The feature fusion sub-model performs feature fusion on the target biological feature intermediate value and the target clothes feature intermediate value, and outputs a target pedestrian fusion feature.

[0091] S230, clothes detection is performed on each base library pedestrian picture in the base library pedestrian picture set to obtain a corresponding base library clothes template picture. The base library pedestrian picture and the corresponding base library clothes template picture are taken as input data, feature extraction and feature fusion are performed through the clothes-changing pedestrian feature recognition model, and a base library pedestrian fusion feature corresponding to each base library pedestrian picture is output.

[0092] In this embodiment, for each base library pedestrian picture, since the clothes worn by the pedestrian in the base library pedestrian picture are the clothes to be recognized, the clothes area of the pedestrian can be detected and the clothes picture can be obtained through clothes detection technology to obtain the base library clothes template picture. When the base library pedestrian picture and the corresponding base library clothes template picture are obtained, the biological features such as the posture and body shape of the pedestrian in the base library pedestrian picture and the clothes features in the base library clothes template picture can be extracted respectively, the extracted features are fused, and a base library pedestrian fusion feature is obtained. The base library pedestrian fusion feature contains both the biological features of the pedestrian in the base library pedestrian picture and the clothes features of the pedestrian, and the accuracy of pedestrian re-identification can be improved.

[0093] Optionally, for each base library pedestrian picture, the base library pedestrian fusion feature can be extracted through the following specific steps:

[0094] S2301, biological feature extraction is performed on the base library pedestrian picture by using the biological feature recognition sub-model to obtain a base library biological feature intermediate value.

[0095] Specifically, similar to the extraction of the biological features in the to-be-queried human body picture, the base library pedestrian picture can be taken as input data, feature extraction is performed through the biological feature recognition sub-model, and a base library biological feature intermediate value is output by the network intermediate layer of the biological feature recognition sub-model.

[0096] S2302, clothes feature extraction is performed on the base library clothes template picture corresponding to the base library pedestrian picture by using the clothes feature recognition sub-model to obtain a base library clothes feature intermediate value.

[0097] Specifically, similar to the extraction of the clothes features in the target clothes template picture, the base library clothes template picture can be taken as input data, feature extraction is performed through the clothes feature recognition sub-model, and a base library clothes feature intermediate value is output by the network intermediate layer of the clothes feature recognition sub-model.

[0098] S2303, adopt the feature fusion sub-model to perform feature fusion on the base library biological feature intermediate value and the base library clothes feature intermediate value, and obtain the base library pedestrian fusion feature.

[0099] Specifically, similar to the fusion of the target biological feature intermediate value and the target clothes feature intermediate value, the base library biological feature intermediate value output by the biological feature recognition sub-model and the base library clothes feature intermediate value output by the clothes feature recognition sub-model can be taken as input data and input into the feature fusion sub-model, and the feature fusion sub-model performs feature fusion on the base library biological feature intermediate value and the base library clothes feature intermediate value, and outputs the base library pedestrian fusion feature.

[0100] S240, according to the base library pedestrian fusion feature and the target pedestrian fusion feature, determine the target pedestrian similarity corresponding to each base library pedestrian picture.

[0101] S250, determine the base library pedestrian picture with a target pedestrian similarity greater than or equal to a preset similarity threshold as the target pedestrian picture.

[0102] In the embodiment of the application, the set of base library pedestrian pictures, the target clothes template picture and the to-be-queried human body picture are obtained, the to-be-queried human body picture and the target clothes template picture are taken as input data, feature extraction and feature fusion are performed through the pre-trained clothes-changing pedestrian feature recognition model, and the target pedestrian fusion feature is output; clothes detection is performed on each base library pedestrian picture in the set of base library pedestrian pictures, and the corresponding base library clothes template picture is obtained, each base library pedestrian picture and the corresponding base library clothes template picture are taken as input data, feature extraction and feature fusion are performed through the clothes-changing pedestrian feature recognition model, and the base library pedestrian fusion feature corresponding to each base library pedestrian picture is output; according to the base library pedestrian fusion feature and the target pedestrian fusion feature, the target pedestrian similarity corresponding to each base library pedestrian picture is determined, and the base library pedestrian picture with a target pedestrian similarity greater than or equal to a preset similarity threshold is determined as the target pedestrian picture. The method can obtain biological features such as pedestrian posture and body shape, and clothes features of the pedestrian, fuse the two to obtain a fusion feature, and perform pedestrian clothes-changing recognition according to the fusion feature, thereby improving the recognition accuracy of the clothes-changing pedestrian.

[0103] Embodiment three

[0104] Figure 3 A structure schematic diagram of a clothes-changing pedestrian re-identification device provided for the embodiment three of the application is shown in FIG. 3. Figure 3 As shown in the figure, the device comprises:

[0105] The data acquisition module 310 is configured to acquire a to-be-queried human body picture, a target clothes template picture and a set of base library pedestrian pictures.

[0106] The feature extraction module 320 is configured to obtain target pedestrian fusion features corresponding to the to-be-queried human body picture and the target clothes template picture, and base library pedestrian fusion features corresponding to each base library pedestrian picture in the base library pedestrian picture set by using a pre-trained clothes-changing pedestrian feature recognition model. The clothes-changing pedestrian feature recognition model comprises a biological feature recognition sub-model, a clothes feature recognition sub-model, and a feature fusion sub-model.

[0107] The similarity determination module 330 is configured to determine target pedestrian similarities corresponding to each base library pedestrian picture according to the base library pedestrian fusion features and the target pedestrian fusion features.

[0108] The target determination module 340 is configured to determine, as a target pedestrian picture, a base library pedestrian picture with a target pedestrian similarity greater than or equal to a preset similarity threshold.

[0109] Optionally, the feature extraction module 320 comprises:

[0110] The target pedestrian fusion feature extraction unit is configured to take the to-be-queried human body picture and the target clothes template picture as input data, perform feature extraction and feature fusion by using the pre-trained clothes-changing pedestrian feature recognition model, and output target pedestrian fusion features.

[0111] The base library pedestrian fusion feature extraction unit is configured to perform clothes detection on each base library pedestrian picture in the base library pedestrian picture set to obtain a corresponding base library clothes template picture, take each base library pedestrian picture and the corresponding base library clothes template picture as input data, perform feature extraction and feature fusion by using the clothes-changing pedestrian feature recognition model, and output base library pedestrian fusion features corresponding to each base library pedestrian picture.

[0112] Optionally, the target pedestrian fusion feature extraction unit is specifically configured to:

[0113] The biological feature recognition sub-model is used to perform biological feature extraction on the to-be-queried human body picture to obtain target biological feature intermediate values.

[0114] The clothes feature recognition sub-model is used to perform clothes feature extraction on the target clothes template picture to obtain target clothes feature intermediate values.

[0115] The feature fusion sub-model is used to perform feature fusion on the target biological feature intermediate values and the target clothes feature intermediate values to obtain target pedestrian fusion features.

[0116] Optionally, the base library pedestrian fusion feature extraction unit is specifically configured to:

[0117] For each base library pedestrian picture, a biological feature recognition sub-model is used to perform biological feature extraction on the base library pedestrian picture to obtain a base library biological feature intermediate value;

[0118] A clothing feature recognition sub-model is used to perform clothing feature extraction on a base library clothing template picture corresponding to the base library pedestrian picture to obtain a base library clothing feature intermediate value.

[0119] A feature fusion sub-model is used to perform feature fusion on the base library biological feature intermediate value and the base library clothing feature intermediate value to obtain a base library pedestrian fusion feature.

[0120] Optionally, the training process of the clothes-changing pedestrian feature recognition model comprises:

[0121] Training data containing query human training pictures, clothing template training pictures, and a base library pedestrian training picture set are obtained, base library pedestrian training pictures in the base library pedestrian training picture set in the same group of training data are annotated for pedestrian recognition according to the query human training pictures and the clothing template training pictures, and standard recognition results corresponding to each base library pedestrian training picture are obtained.

[0122] The training data are input into a to-be-trained clothes-changing pedestrian feature recognition model to obtain output query biological features, query fusion features, base library biological features corresponding to each base library pedestrian training picture, and base library fusion features, wherein the to-be-trained clothes-changing pedestrian feature recognition model comprises a to-be-trained biological feature recognition sub-model, a to-be-trained clothing feature recognition sub-model, and a to-be-trained feature fusion sub-model.

[0123] Feature classification is performed on the query biological features, the query fusion features, each base library biological feature, and each base library fusion feature to obtain training biological feature recognition results and training fusion feature recognition results corresponding to each base library pedestrian training picture.

[0124] The standard recognition results, the training biological feature recognition results, and the training fusion feature recognition results are respectively substituted into at least two given loss function expressions to obtain corresponding loss functions.

[0125] The to-be-trained clothes-changing pedestrian feature recognition model is subjected to back propagation through the loss functions to obtain the clothes-changing pedestrian feature recognition model.

[0126] Optionally, the inputting of the training data into the to-be-trained clothes-changing pedestrian feature recognition model to obtain the output query biological features, the query fusion features, the base library biological features corresponding to each base library pedestrian training picture, and the base library fusion features comprises:

[0127] input the query human training picture into the to-be-trained biometric feature recognition sub-model, output a query biometric feature intermediate value and the query biometric feature, input each of the base library pedestrian training pictures into the to-be-trained biometric feature recognition sub-model, and output a corresponding base library biometric feature training intermediate value and a base library biometric feature;

[0128] input the clothes template training picture into the to-be-trained clothes feature recognition sub-model, output a query clothes feature intermediate value, perform clothes detection on each of the base library pedestrian training pictures to obtain a corresponding base library clothes training picture, input each of the base library clothes training pictures into the to-be-trained clothes feature recognition sub-model, and output a corresponding base library clothes feature training intermediate value;

[0129] input the query biometric feature intermediate value and the query clothes feature intermediate value into the to-be-trained feature fusion sub-model, output a query fusion feature, input each of the base library biometric feature training intermediate value and the corresponding base library clothes feature training intermediate value into the to-be-trained feature fusion sub-model, and output a corresponding base library fusion feature.

[0130] Optionally, the training biometric feature recognition result and the training fusion feature recognition result corresponding to each of the base library pedestrian training pictures are obtained by performing feature classification on the query biometric feature, the query fusion feature, each of the base library biometric features, and each of the base library fusion features, and the method comprises the following steps.

[0131] a first classifier is used to perform feature classification on the query biometric feature and each of the base library biometric features, to obtain a query biometric feature classification result and a base library biometric feature classification result corresponding to each of the base library biometric features, and each of the base library biometric feature classification results is used to determine a training biometric feature recognition result corresponding to each of the base library biometric features according to the query biometric feature classification result;

[0132] a second classifier is used to perform feature classification on the query fusion feature and each of the base library fusion features, to obtain a query fusion feature classification result and a base library fusion feature classification result corresponding to each of the base library fusion features, and each of the base library fusion feature classification results is used to determine a training fusion feature recognition result corresponding to each of the base library fusion features according to the query fusion feature classification result.

[0133] The clothes-changing pedestrian re-identification device provided in the embodiment of the present application can perform the clothes-changing pedestrian re-identification method provided in any embodiment of the present application, has the function modules and beneficial effects corresponding to the execution method.

[0134] Embodiment Four

[0135] Figure 4A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0136] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0137] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0138] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the person re-identification method for people changing clothes.

[0139] In some embodiments, the method of dressed pedestrian re-identification can be implemented as a computer program tangibly embodied in a computer readable storage medium, e.g., storage unit 18. In some embodiments, parts or all of the computer program can be loaded and / or installed onto electronic device 10 via, e.g., ROM 12 and / or communication unit 19. When the computer program is loaded onto RAM 13 and executed by processor 11, one or more steps of the method of dressed pedestrian re-identification described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the method of dressed pedestrian re-identification by other means, e.g., via firmware.

[0140] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, specially designed application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0141] Computer programs used to implement the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor of the machine, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0142] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0143] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0144] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0145] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0146] It should be understood that the various forms of flow shown above can be reordered, added to, or have steps deleted. For example, the steps described in the present application can be performed in parallel, in series, or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, and this is not limited herein.

[0147] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for pedestrian re-identification of a changing clothes pedestrian, characterized in that, The method comprises the following steps: acquiring a to-be-inquired human body picture, a target clothes template picture and a base library pedestrian picture set; using a pre-trained clothes-changing pedestrian feature recognition model to obtain target pedestrian fusion features corresponding to the to-be-inquired human body picture and the target clothes template picture, and base library pedestrian fusion features corresponding to each base library pedestrian picture in the base library pedestrian picture set, wherein the clothes-changing pedestrian feature recognition model comprises a biological feature recognition sub-model, a clothes feature recognition sub-model and a feature fusion sub-model; determining target pedestrian similarities corresponding to each base library pedestrian picture according to each base library pedestrian fusion feature and the target pedestrian fusion feature; determining a base library pedestrian picture with a target pedestrian similarity greater than or equal to a preset similarity threshold as a target pedestrian picture; the step of using the pre-trained clothes-changing pedestrian feature recognition model to obtain the target pedestrian fusion features corresponding to the to-be-inquired human body picture and the target clothes template picture, and the base library pedestrian fusion features corresponding to each base library pedestrian picture in the base library pedestrian picture set comprises: using the to-be-inquired human body picture and the target clothes template picture as input data, performing feature extraction and feature fusion through the pre-trained clothes-changing pedestrian feature recognition model, and outputting target pedestrian fusion features; performing clothes detection on each base library pedestrian picture in the base library pedestrian picture set to obtain a corresponding base library clothes template picture, using each base library pedestrian picture and the corresponding base library clothes template picture as input data, performing feature extraction and feature fusion through the clothes-changing pedestrian feature recognition model, and outputting base library pedestrian fusion features corresponding to each base library pedestrian picture; the biological feature recognition sub-model is used to acquire biological features of an appearance contour of a human body in a picture, the clothes feature recognition sub-model is used to acquire clothes features of clothes worn by a pedestrian in a picture, and the feature fusion sub-model is used to fuse the biological features and the clothes features to obtain fusion features.

2. The method of claim 1, wherein, the step of using the to-be-inquired human body picture and the target clothes template picture as input data, performing feature extraction and feature fusion through the pre-trained clothes-changing pedestrian feature recognition model, and outputting target pedestrian fusion features comprises: using the biological feature recognition sub-model to perform biological feature extraction on the to-be-inquired human body picture to obtain target biological feature intermediate values; using the clothes feature recognition sub-model to perform clothes feature extraction on the target clothes template picture to obtain target clothes feature intermediate values; using the feature fusion sub-model to perform feature fusion on the target biological feature intermediate values and the target clothes feature intermediate values to obtain target pedestrian fusion features.

3. The method of claim 1, wherein, the step of using each base library pedestrian picture and the corresponding base library clothes template picture as input data, performing feature extraction and feature fusion through the clothes-changing pedestrian feature recognition model, and outputting base library pedestrian fusion features corresponding to each base library pedestrian picture comprises: for each base library pedestrian picture, using the biological feature recognition sub-model to perform biological feature extraction on the base library pedestrian picture to obtain base library biological feature intermediate values; The clothes feature recognition sub-model is used to extract clothes features from the base clothes template pictures corresponding to the base pedestrian pictures, to obtain base clothes feature intermediate values; The feature fusion sub-model is used to fuse the base biological feature intermediate values and the base clothes feature intermediate values, to obtain base pedestrian fusion features.

4. The method of claim 1, wherein, The training process of the clothes-changing pedestrian feature recognition model includes: Training data including query human training pictures, clothes template training pictures, and a base pedestrian training picture set are obtained, base pedestrian training pictures in the base pedestrian training picture set in the same group of training data are annotated for pedestrian recognition according to the query human training pictures and the clothes template training pictures, to obtain standard recognition results corresponding to each base pedestrian training picture; The training data are input into a to-be-trained clothes-changing pedestrian feature recognition model, to obtain output query biological features, query fusion features, base biological features corresponding to each base pedestrian training picture, and base fusion features, wherein the to-be-trained clothes-changing pedestrian feature recognition model includes a to-be-trained biological feature recognition sub-model, a to-be-trained clothes feature recognition sub-model, and a to-be-trained feature fusion sub-model; Feature classification is performed on the query biological features, the query fusion features, the base biological features, and the base fusion features, to obtain training biological feature recognition results and training fusion feature recognition results corresponding to each base pedestrian training picture; The standard recognition results, the training biological feature recognition results, and the training fusion feature recognition results are substituted into at least two given loss function expressions respectively, to obtain corresponding loss functions; The to-be-trained clothes-changing pedestrian feature recognition model is subjected to back propagation through the loss functions, to obtain the clothes-changing pedestrian feature recognition model.

5. The method of claim 4, wherein, The training data are input into a to-be-trained clothes-changing pedestrian feature recognition model, to obtain output query biological features, query fusion features, base biological features corresponding to each base pedestrian training picture, and base fusion features, wherein the to-be-trained clothes-changing pedestrian feature recognition model includes a to-be-trained biological feature recognition sub-model, a to-be-trained clothes feature recognition sub-model, and a to-be-trained feature fusion sub-model; The query human training pictures are input into the to-be-trained biological feature recognition sub-model, to output query biological feature intermediate values and query biological features, and each base pedestrian training picture is input into the to-be-trained biological feature recognition sub-model, to output corresponding base biological feature training intermediate values and base biological features; The clothes template training pictures are input into the to-be-trained clothes feature recognition sub-model, to output query clothes feature intermediate values, clothes detection is performed on each base pedestrian training picture, to obtain corresponding base clothes training pictures, and each base clothes training picture is input into the to-be-trained clothes feature recognition sub-model, to output corresponding base clothes feature training intermediate values; The query biological feature intermediate values and the query clothes feature intermediate values are input into the to-be-trained feature fusion sub-model, to output query fusion features, and each base biological feature training intermediate value and corresponding base clothes feature training intermediate value are input into the to-be-trained feature fusion sub-model, to output corresponding base fusion features.

6. The method of claim 4, wherein, The training biological feature recognition result and the training fusion feature recognition result corresponding to each of the library pedestrian training picture are obtained by performing feature classification on the query biological feature, the query fusion feature, each of the library biological features, and each of the library fusion features, including: The first classifier is used for feature classification on the query biological feature and each of the library biological features, to obtain a query biological feature classification result and a library biological feature classification result corresponding to each of the library biological features, and the training biological feature recognition result corresponding to each of the library biological features is determined according to each of the library biological feature classification results and the query biological feature classification result; The second classifier is used for feature classification on the query fusion feature and each of the library fusion features, to obtain a query fusion feature classification result and a library fusion feature classification result corresponding to each of the library fusion features, and the training fusion feature recognition result corresponding to each of the library fusion features is determined according to each of the library fusion feature classification results and the query fusion feature classification result.

7. A device for pedestrian re-identification of a changing clothes, characterized in that, It includes: The data acquisition module is used for acquiring a to-be-queried human body picture, a target clothes template picture, and a library pedestrian picture set; The feature extraction module is used for obtaining target pedestrian fusion features corresponding to the to-be-queried human body picture and the target clothes template picture, and library pedestrian fusion features corresponding to each of the library pedestrian pictures in the library pedestrian picture set by using a pre-trained clothes-changing pedestrian feature recognition model, wherein the clothes-changing pedestrian feature recognition model includes a biological feature recognition sub-model, a clothes feature recognition sub-model, and a feature fusion sub-model; The similarity determination module is used for determining target pedestrian similarities corresponding to each of the library pedestrian pictures according to each of the library pedestrian fusion features and the target pedestrian fusion features; The target determination module is used for determining a library pedestrian picture with a target pedestrian similarity greater than or equal to a preset similarity threshold as a target pedestrian picture; The feature extraction module includes a target pedestrian fusion feature extraction unit configured to input the to-be-queried human body picture and the target clothes template picture as input data, perform feature extraction and feature fusion on the input data by using the pre-trained clothes-changing pedestrian feature recognition model, and output target pedestrian fusion features; A library pedestrian fusion feature extraction unit is configured to perform clothes detection on each of the library pedestrian pictures in the library pedestrian picture set to obtain a corresponding library clothes template picture, input each of the library pedestrian pictures and the corresponding library clothes template picture as input data, perform feature extraction and feature fusion on the input data by using the clothes-changing pedestrian feature recognition model, and output library pedestrian fusion features corresponding to each of the library pedestrian pictures; The biological feature recognition sub-model is used for obtaining biological features of an appearance contour of a human body in a picture, the clothes feature recognition sub-model is used for obtaining clothes features of clothes worn by a pedestrian in a picture, and the feature fusion sub-model is used for fusing the biological features and the clothes features to obtain fusion features.

8. An electronic device, comprising: The electronic device includes: at least one processor; and a memory connected with the at least one processor in communication; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the clothes-changing pedestrian re-identification method in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the clothes-changing pedestrian re-identification method in any one of claims 1-6 when executed.

Citation Information

Patent Citations

  • Clothes changing pedestrian re-identification method and system based on auto-encoding network

    CN110321801A

  • Cross-domain pedestrian re-identification method and system based on multi-feature mixed learning

    CN113221770A