Pedestrian re-identification method, device, equipment and readable storage medium
Through the feature extraction model based on RGB and infrared image training, feature extraction and matching of the query image is solved, and the problem of information loss in infrared image feature extraction is improved.
Patent Information
- Application Number
- CN202111679104.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-31
AI Technical Summary
When processing infrared images, existing feature extraction models will cause features to lose important information and reduce the accuracy of pedestrian re-identification results.
A pedestrian re-identification method is proposed. By using the trained feature extraction model based on an image containing RGB images and infrared images, the feature extraction of the query image is extracted and matched with the features of each pedestrian image in the query database to obtain the re-identification result.
By using feature extraction models based on RGB and infrared images, the features of infrared images can be effectively extracted, and the accuracy of pedestrian re-identification results can be improved.
Smart Images

Figure CN114445859B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and more specifically, to a pedestrian re-identification method, related equipment and a readable storage medium. Background Art
[0002] Person re-identification (ReID) aims to find images of the same person as the query image from a large image library, which contains multiple images with known identity information of pedestrians. When performing person re-identification on the query image, the query image needs to be input into a feature extraction model to obtain the features of the query image, and then the features of the query image are matched with the features of each image in the large image library to determine the person re-identification result of the query image.
[0003] However, the current feature extraction model is trained based on a large number of RGB modal images, while in some scenarios, the queried image may be an infrared (IR) image (for example, an image captured by an infrared camera at night). There is a huge difference between infrared images and RGB images (for example, infrared images lose rich color information). Therefore, when the current feature extraction model is used to extract features from infrared images, the obtained features will lose some important information, resulting in a decrease in the accuracy of pedestrian re-identification results.
[0004] Therefore, how to optimize the pedestrian re-identification scheme to improve the accuracy of pedestrian re-identification results has become a technical problem that needs to be urgently solved by technical personnel in this field. Summary of the invention
[0005] In view of the above problems, this application proposes a pedestrian re-identification method, related equipment and readable storage medium. The specific scheme is as follows:
[0006] A pedestrian re-identification method, the method comprising:
[0007] Get the image to be queried;
[0008] Inputting the image to be queried into a feature extraction model, the feature extraction model outputting features of the image to be queried; the feature extraction model is trained based on an image pair including an RGB image and an infrared image;
[0009] The features of the image to be queried are matched with the features of each pedestrian image in the query database to obtain a pedestrian re-identification result corresponding to the image to be queried.
[0010] Optionally, the feature extraction model is trained as follows:
[0011] Obtaining a pre-built feature extraction model training network, wherein the feature extraction model training network includes a feature extraction module, an attribute classification module, a feature fusion module, a feature alignment module, and an identity prediction module;
[0012] Acquire a training image pair, wherein the training image pair includes an RGB image and an infrared image; each training image pair is annotated with an identity label and an attribute label;
[0013] The training image pairs are used as training samples, the identity labels and attribute labels annotated by the training image pairs are used as sample labels, and the feature extraction model training network is trained with the joint loss of the attribute classification module, the identity prediction module and the feature alignment module. When the feature extraction model training network converges, the feature extraction module is the feature extraction model.
[0014] Optionally, during the training process of the feature extraction model training network, the feature extraction module performs feature extraction on the RGB image and the infrared image in the training image pair to obtain global features of the RGB image, global features of the infrared image, and overall local features of the image pair;
[0015] The attribute classification module generates new overall local features of the image pair based on the overall local features of the image pair, and predicts the attributes of the image pair based on the overall local features of the image pair to obtain an attribute prediction result, wherein the difference between the attribute prediction result and the attribute label of the training image pair is the loss of the attribute classification module;
[0016] The feature fusion module fuses the new overall local features of the image pair with the global features of the RGB image to obtain new global features of the RGB image, and fuses the new overall local features of the image pair with the global features of the infrared image to obtain new global features of the infrared image;
[0017] The feature alignment module obtains a synthesized infrared modal feature based on the new global feature of the RGB image and the attribute label of the image pair; obtains a real infrared modal feature based on the new global feature of the infrared image and the attribute label of the image pair; generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature, and the difference between the intermediate modal feature and the synthesized infrared modal feature and the difference between the intermediate modal feature and the real infrared modal feature are the losses of the feature alignment module;
[0018] The identity prediction module predicts the identity of the image pair based on the new global features of the RGB image and the new global features of the infrared image to obtain an identity prediction result, and the difference between the identity prediction result and the identity label of the training image pair is the loss of the identity prediction module.
[0019] Optionally, the feature extraction module includes an input layer, an intermediate layer and a final layer;
[0020] The feature extraction module extracts features from the RGB image and the infrared image in the training image pair to obtain global features of the RGB image, global features of the infrared image, and overall local features of the image pair, including:
[0021] The input layer generates an input sequence of RGB images in the training image pair, and an input sequence of infrared images in the training image pair;
[0022] The intermediate layer encodes the input sequence of the RGB image to obtain local features of the RGB image, and encodes the input sequence of the infrared image to obtain local features of the infrared image;
[0023] The last layer encodes the local features of the RGB image to obtain the global features of the RGB image, and encodes the local features of the infrared image to obtain the global features of the infrared image.
[0024] Optionally, the input layer generates an input sequence of RGB images in the training image pair, comprising:
[0025] Obtaining a feature sequence and a position sequence of each block corresponding to the RGB image in the training image pair;
[0026] Combining the feature sequence of each block corresponding to the preset mark and the RGB image, and the position sequence into an input sequence of the RGB image;
[0027] The input layer generates an input sequence of infrared images in the training image pair, including:
[0028] Obtaining a feature sequence and a position sequence of each block corresponding to the infrared image in the training image pair;
[0029] The input sequence of the infrared image is composed of the feature sequence of each block corresponding to the preset mark and the position sequence of the infrared image.
[0030] Optionally, the attribute classification module includes: an attention layer, an attribute prediction layer, the attribute classification module generates new overall local features of the image pair based on the overall local features of the image pair, and predicts the attributes of the image pair based on the overall local features of the image pair to obtain attribute prediction results, including:
[0031] The attention layer generates attribute features of the image pair based on the overall local features of the image pair, and the attribute features of the image pair are combined to obtain new overall local features of the image pair;
[0032] The attribute prediction layer predicts the attributes of the image pair based on the attribute features of the image pair to obtain an attribute prediction result.
[0033] Optionally, the feature alignment module includes:
[0034] Infrared code generator, attribute label fusion layer, adversarial network;
[0035] The feature alignment module obtains a synthesized infrared modal feature based on the new global feature of the RGB image and the attribute label of the image pair; obtains a real infrared modal feature based on the new global feature of the infrared image and the attribute label of the image pair; and generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature, including:
[0036] The infrared code generator encodes the new global features of the RGB image to obtain the encoded infrared modal features;
[0037] The attribute label fusion layer fuses the encoded infrared modal features and the attribute labels of the image pair to obtain a synthesized infrared modal feature; and fuses the new global features of the infrared image and the attribute labels of the image pair to obtain a real infrared modal feature;
[0038] The adversarial network generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature.
[0039] A pedestrian re-identification device, the device comprising:
[0040] An acquisition unit, used for acquiring an image to be queried;
[0041] A feature extraction unit, used for inputting the image to be queried into a feature extraction model, wherein the feature extraction model outputs features of the image to be queried; the feature extraction model is obtained by training based on an image pair including an RGB image and an infrared image;
[0042] The matching unit is used to match the features of the image to be queried with the features of each pedestrian image in the query database to obtain a pedestrian re-identification result corresponding to the image to be queried.
[0043] Optionally, the feature extraction model is trained as follows:
[0044] Obtaining a pre-built feature extraction model training network, wherein the feature extraction model training network includes a feature extraction module, an attribute classification module, a feature fusion module, a feature alignment module, and an identity prediction module;
[0045] Acquire a training image pair, wherein the training image pair includes an RGB image and an infrared image; each training image pair is annotated with an identity label and an attribute label;
[0046] The training image pairs are used as training samples, the identity labels and attribute labels annotated by the training image pairs are used as sample labels, and the feature extraction model training network is trained with the joint loss of the attribute classification module, the identity prediction module and the feature alignment module. When the feature extraction model training network converges, the feature extraction module is the feature extraction model.
[0047] Optionally, during the training process of the feature extraction model training network, the feature extraction module performs feature extraction on the RGB image and the infrared image in the training image pair to obtain global features of the RGB image, global features of the infrared image, and overall local features of the image pair;
[0048] The attribute classification module generates new overall local features of the image pair based on the overall local features of the image pair, and predicts the attributes of the image pair based on the overall local features of the image pair to obtain an attribute prediction result, wherein the difference between the attribute prediction result and the attribute label of the training image pair is the loss of the attribute classification module;
[0049] The feature fusion module fuses the new overall local features of the image pair with the global features of the RGB image to obtain new global features of the RGB image, and fuses the new overall local features of the image pair with the global features of the infrared image to obtain new global features of the infrared image;
[0050] The feature alignment module obtains a synthesized infrared modal feature based on the new global feature of the RGB image and the attribute label of the image pair; obtains a real infrared modal feature based on the new global feature of the infrared image and the attribute label of the image pair; generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature, and the difference between the intermediate modal feature and the synthesized infrared modal feature and the difference between the intermediate modal feature and the real infrared modal feature are the losses of the feature alignment module;
[0051] The identity prediction module predicts the identity of the image pair based on the new global features of the RGB image and the new global features of the infrared image to obtain an identity prediction result, and the difference between the identity prediction result and the identity label of the training image pair is the loss of the identity prediction module.
[0052] Optionally, the feature extraction module includes an input layer, an intermediate layer and a final layer;
[0053] The feature extraction module extracts features from the RGB image and the infrared image in the training image pair to obtain global features of the RGB image, global features of the infrared image, and overall local features of the image pair, including:
[0054] The input layer generates an input sequence of RGB images in the training image pair, and an input sequence of infrared images in the training image pair;
[0055] The intermediate layer encodes the input sequence of the RGB image to obtain local features of the RGB image, and encodes the input sequence of the infrared image to obtain local features of the infrared image;
[0056] The last layer encodes the local features of the RGB image to obtain the global features of the RGB image, and encodes the local features of the infrared image to obtain the global features of the infrared image.
[0057] Optionally, the input layer generates an input sequence of RGB images in the training image pair, comprising:
[0058] Obtaining a feature sequence and a position sequence of each block corresponding to the RGB image in the training image pair;
[0059] Combining the feature sequence of each block corresponding to the preset mark and the RGB image, and the position sequence into an input sequence of the RGB image;
[0060] The input layer generates an input sequence of infrared images in the training image pair, including:
[0061] Obtaining a feature sequence and a position sequence of each block corresponding to the infrared image in the training image pair;
[0062] The input sequence of the infrared image is composed of the feature sequence of each block corresponding to the preset mark and the position sequence of the infrared image.
[0063] Optionally, the attribute classification module includes: an attention layer, an attribute prediction layer, the attribute classification module generates new overall local features of the image pair based on the overall local features of the image pair, and predicts the attributes of the image pair based on the overall local features of the image pair to obtain attribute prediction results, including:
[0064] The attention layer generates attribute features of the image pair based on the overall local features of the image pair, and the attribute features of the image pair are combined to obtain new overall local features of the image pair;
[0065] The attribute prediction layer predicts the attributes of the image pair based on the attribute features of the image pair to obtain an attribute prediction result.
[0066] Optionally, the feature alignment module includes:
[0067] Infrared code generator, attribute label fusion layer, adversarial network;
[0068] The feature alignment module obtains a synthesized infrared modal feature based on the new global feature of the RGB image and the attribute label of the image pair; obtains a real infrared modal feature based on the new global feature of the infrared image and the attribute label of the image pair; and generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature, including:
[0069] The infrared code generator encodes the new global features of the RGB image to obtain the encoded infrared modal features;
[0070] The attribute label fusion layer fuses the encoded infrared modal features and the attribute labels of the image pair to obtain a synthesized infrared modal feature; and fuses the new global features of the infrared image and the attribute labels of the image pair to obtain a real infrared modal feature;
[0071] The adversarial network generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature.
[0072] A pedestrian re-identification device, comprising a memory and a processor;
[0073] The memory is used to store programs;
[0074] The processor is used to execute the program to implement the various steps of the pedestrian re-identification method as described above.
[0075] A readable storage medium stores a computer program, and when the computer program is executed by a processor, each step of the pedestrian re-identification method described above is implemented.
[0076] By means of the above technical scheme, the present application discloses a pedestrian re-identification method, related equipment and readable storage medium. A feature extraction model is first obtained based on training of an image pair including an RGB image and an infrared image. After obtaining the image to be queried, the image to be queried is input into the feature extraction model. The feature extraction model outputs the features of the image to be queried. By matching the features of the image to be queried with the features of each pedestrian image in the query database, the pedestrian re-identification result corresponding to the image to be queried can be obtained. In the present application, the feature extraction model is obtained by training based on RGB images and infrared images. Whether feature extraction is performed on RGB images or infrared images, the effectiveness of the extracted features can be guaranteed. Therefore, the accuracy of the pedestrian re-identification results can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0078] Figure 1 A flowchart of a pedestrian re-identification method disclosed in an embodiment of the present application;
[0079] Figure 2 A schematic diagram of the structure of a feature extraction model training network disclosed in an embodiment of the present application;
[0080] Figure 3 A structural diagram of a feature extraction module disclosed in an embodiment of the present application;
[0081] Figure 4 A schematic diagram of the structure of an attribute classification module disclosed in an embodiment of the present application;
[0082] Figure 5 A schematic diagram of a structural example of an attribute classification module disclosed in an embodiment of the present application;
[0083] Figure 6 A structural schematic diagram of a feature alignment module disclosed in an embodiment of the present application;
[0084] Figure 7A schematic diagram of the structure of a pedestrian re-identification device disclosed in an embodiment of the present application;
[0085] Figure 8 This is a hardware structure block diagram of a pedestrian re-identification device disclosed in an embodiment of the present application. DETAILED DESCRIPTION
[0086] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0087] Next, the pedestrian re-identification method provided by the present application is introduced through the following embodiments.
[0088] Reference Figure 1 , Figure 1 This is a flow chart of a pedestrian re-identification method disclosed in an embodiment of the present application. The method may include:
[0089] Step S101: Acquire an image to be queried.
[0090] In the present application, the image to be queried may be an RGB image or an infrared image.
[0091] Step S102: inputting the image to be queried into a feature extraction model, and the feature extraction model outputs the features of the image to be queried; the feature extraction model is trained based on an image pair including an RGB image and an infrared image.
[0092] In this application, the feature extraction model is trained based on RGB images and infrared images. Whether feature extraction is performed on RGB images or infrared images, the effectiveness of the extracted features can be guaranteed, thereby ensuring the accuracy of subsequent pedestrian re-identification results.
[0093] Step S103: matching the features of the image to be queried with the features of each pedestrian image in a preset query database to obtain a pedestrian re-identification result corresponding to the image to be queried.
[0094] In the present application, the multiple pedestrian images with known identities included in the query database may also be RGB images or infrared images, and the features of each pedestrian image in the query database are also extracted based on the feature extraction model. In order to improve the efficiency of pedestrian re-identification, after the training based on the image pairs including the RGB image and the infrared image is completed, each pedestrian image in the query database can be input into the feature extraction model to obtain the features of each pedestrian image in the query database and store them.
[0095] In the present application, as an implementable method, the similarity between the features of the image to be queried and the features of each pedestrian image in the query database can be calculated, and the identities corresponding to the features of a preset number of pedestrian images in the query database whose similarity is greater than a preset threshold and whose similarity is ranked high can be determined as the pedestrian re-identification results corresponding to the image to be queried.
[0096] The present embodiment discloses a pedestrian re-identification method, which first obtains a feature extraction model based on an image pair including an RGB image and an infrared image. After obtaining the image to be queried, the image to be queried is input into the feature extraction model, and the feature extraction model outputs the features of the image to be queried. By matching the features of the image to be queried with the features of each pedestrian image in the query database, the pedestrian re-identification result corresponding to the image to be queried can be obtained. In the present application, the feature extraction model is obtained by training based on RGB images and infrared images. Whether feature extraction is performed on RGB images or infrared images, the effectiveness of the extracted features can be guaranteed, and therefore, the accuracy of the pedestrian re-identification results can be improved.
[0097] In another embodiment of the present application, a training method of the feature extraction model is described, and the method may include the following steps:
[0098] Step S201: obtaining a pre-constructed feature extraction model training network, wherein the feature extraction model training network includes a feature extraction module, an attribute classification module, a feature fusion module, a feature alignment module and an identity prediction module.
[0099] See also Figure 2 , Figure 2 This is a structural schematic diagram of a feature extraction model training network disclosed in an embodiment of the present application, wherein the feature extraction model training network includes a feature extraction module, an attribute classification module, a feature fusion module, a feature alignment module, and an identity prediction module.
[0100] refer to Figure 2 ,During the training process, the specific functions of each module are implemented as follows:
[0101] The feature extraction module extracts features from the RGB image and the infrared image in the training image pair to obtain global features of the RGB image, global features of the infrared image, and overall local features of the image pair;
[0102] The attribute classification module generates new overall local features of the image pair based on the overall local features of the image pair, and predicts the attributes of the image pair based on the overall local features of the image pair to obtain an attribute prediction result, wherein the difference between the attribute prediction result and the attribute label of the training image pair is the loss of the attribute classification module;
[0103] The feature fusion module fuses the new overall local features of the image pair with the global features of the RGB image to obtain new global features of the RGB image, and fuses the new overall local features of the image pair with the global features of the infrared image to obtain new global features of the infrared image;
[0104] The feature alignment module obtains a synthesized infrared modal feature based on the new global feature of the RGB image and the attribute label of the image pair; obtains a real infrared modal feature based on the new global feature of the infrared image and the attribute label of the image pair; generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature, and the difference between the intermediate modal feature and the synthesized infrared modal feature and the difference between the intermediate modal feature and the real infrared modal feature are the losses of the feature alignment module;
[0105] The identity prediction module predicts the identity of the image pair based on the new global features of the RGB image and the new global features of the infrared image to obtain an identity prediction result, and the difference between the identity prediction result and the identity label of the training image pair is the loss of the identity prediction module.
[0106] Step S202: Acquire a pair of training images, wherein the pair of training images includes an RGB image and an infrared image; each pair of training images is annotated with an identity label and an attribute label.
[0107] In this application, multiple identities can be preset to indicate different pedestrian identities. For example, an identity can be represented by a name. In addition, in this application, multiple attributes can be preset to represent the characteristics of pedestrians. It should be noted that since infrared images lack color information, in this application, the preset attributes are attributes that are not related to color. As an implementable method, the attributes can include eight attributes: gender, whether wearing glasses, long hair or short hair, long pants or shorts, long sleeves or short sleeves, whether making a phone call, whether carrying a bag, and whether carrying a backpack. In this application, in order to annotate the training image pairs, an identity label can be set for each identity, and an attribute label can be set for the attribute. As an implementable method, the identity label can be used to indicate the identity of the pedestrian. The identity label can be represented by different identifiers, and the attribute label is used to indicate the characteristics of the pedestrian. For each attribute, the positive value is 1 and the negative value is 0.
[0108] As an implementable embodiment, the attribute labels include any one or more of a label for indicating the gender of the pedestrian, a label for indicating whether the pedestrian wears glasses, a label for indicating whether the pedestrian has long hair or short hair, a label for indicating whether the pedestrian wears long pants or shorts, a label for indicating whether the pedestrian has long sleeves or short sleeves, a label for indicating whether the pedestrian is making a phone call, a label for indicating whether the pedestrian is carrying a bag, and a label for indicating whether the pedestrian is carrying a backpack.
[0109] Step S203: Using the training image pairs as training samples, using the identity labels and attribute labels annotated by the training image pairs as sample labels, and training the feature extraction model training network with the joint loss of the attribute classification module, the identity prediction module and the feature alignment module; when the feature extraction model training network converges, the feature extraction module is the feature extraction model.
[0110] It should be noted that in the present application, the feature extraction module can adopt any network structure, such as CNN structure, transformer structure, etc., but the CNN structure generally uses downsampling (such as pooling and striped convolution) to reduce the spatial resolution of the output features, which greatly affects the ability to distinguish objects with similar appearances. For example, the details of some attributes (whether to wear glasses) will be lost in the output features of the feature extraction module using the CNN structure. The use of the transformer structure can focus on more diverse parts of the human body, and there is no need for downsampling, so that more detailed information is retained in the output features. For example, the differences in the features around the eyes can be observed, so as to help the model distinguish between two people more easily. Therefore, in the present application, the feature extraction module preferably adopts the transformer structure to make the extracted features more robust.
[0111] In another embodiment of the present application, the specific implementation of the feature extraction module is described, see Figure 3 , Figure 3 A structural diagram of a feature extraction module disclosed in an embodiment of the present application is shown in FIG. Figure 3 As shown, the feature extraction module includes an input layer, an intermediate layer and a final layer. The specific functions of each layer are implemented as follows:
[0112] The input layer generates an input sequence of RGB images in the training image pair, and an input sequence of infrared images in the training image pair;
[0113] As an implementation method, the input layer generates an input sequence of the RGB image in the training image pair, including: obtaining a feature sequence and a position sequence of each block corresponding to the RGB image in the training image pair; combining a preset mark with the feature sequence and the position sequence of each block corresponding to the RGB image to form the input sequence of the RGB image;
[0114] As an implementable embodiment, the input layer generates an input sequence of the infrared image in the training image pair, including: obtaining a feature sequence and a position sequence of each block corresponding to the infrared image in the training image pair; and combining preset marks with the feature sequence and the position sequence of each block corresponding to the infrared image to form the input sequence of the infrared image.
[0115] For ease of understanding, RGB images and infrared images are collectively referred to as images. In this application, each image can be divided into N patches of fixed size: i (i=1, 2, 3,…, N).
[0116] Then, use the linear classification block (LinearProjection) to map the N blocks to the D-dimensional linear space, and obtain the sequence F(x1); F(x2); ···; F(x N );
[0117] Then, an additional marker x0 is added before the sequence corresponding to the N blocks, and the input sequence corresponding to the image is obtained as S = [x0; F(x1); F(x2); ...; F(x N )]+P, where S represents sequence embedding and P represents position embedding, which is represented by the position sequence of each block.
[0118] The intermediate layer encodes the input sequence of the RGB image to obtain the local features of the RGB image, and encodes the input sequence of the infrared image to obtain the local features of the infrared image.
[0119] It should be noted that in this application, the reason for selecting the output of the middle layer as the local feature is: it can avoid the loss of too much information in the process of deep feature generation, and the pedestrian identity feature is high-level semantics, while the attribute classification result is low-level semantics, and low-level semantics is more suitable for comparing the appearance similarity between different pedestrians.
[0120] The last layer encodes the local features of the RGB image to obtain the global features of the RGB image, and encodes the local features of the infrared image to obtain the global features of the infrared image.
[0121] In another embodiment of the present application, the specific implementation of the attribute classification module is described, see Figure 4 , Figure 4 This is a schematic diagram of the structure of an attribute classification module disclosed in an embodiment of the present application, such as Figure 4 As shown, the attribute classification module includes: an attention layer and an attribute prediction layer.
[0122] The attention layer generates attribute features of the image pair based on the overall local features of the image pair, and the attribute features of the image pair are combined to obtain new overall local features of the image pair; the attribute prediction layer predicts the attributes of the image pair based on the attribute features of the image pair to obtain attribute prediction results.
[0123] As an implementation option, see Figure 5 , Figure 5 FIG. 1 is a schematic diagram of a structural example of an attribute classification module disclosed in an embodiment of the present application. In this example, there are 8 attributes of an image pair, such as Figure 5 As shown, the attribute classification module uses the attention module ( Figure 5 The AttentionBlock shown in ) is used to calculate the overall local features of the image pair ( Figure 5 LocFeat) shown in the figure obtains the concentrated areas of different attributes and generates different attribute features ( Figure 5 The different attribute features are then fused to generate new overall local features (Loc-Feat1 to Loc-Feat8). Figure 5 Loc Feat' shown in ); different attribute features use simplified gating ( Figure 5 The Gating shown in the figure determines the importance of the attribute. w is the weight. The sigmoid function is used because its function value range is [0,1]. The function value represents the importance of the attribute. Then, the features are reduced in dimension using 1×1 convolution and input into the fully connected layer ( Figure 5Attribute classification can be achieved by using linear classifiers, and eight attributes are predicted by eight linear classifiers to obtain attribute prediction results ( Figure 5 Attribution1 to Attribution8 shown in ).
[0124] In another embodiment of the present application, the specific implementation of the feature alignment module is described, see Figure 6 , Figure 6 This is a schematic diagram of the structure of a feature alignment module disclosed in an embodiment of the present application, such as Figure 6 As shown, the feature alignment module includes: an infrared code generator, an attribute label fusion layer, and an adversarial network.
[0125] The feature alignment module obtains a synthesized infrared modal feature based on the new global feature of the RGB image and the attribute label of the image pair; obtains a real infrared modal feature based on the new global feature of the infrared image and the attribute label of the image pair; and generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature, including:
[0126] The infrared code generator encodes the new global features of the RGB image to obtain the encoded infrared modal features;
[0127] The attribute label fusion layer fuses the encoded infrared modal features and the attribute labels of the image pair to obtain a synthesized infrared modal feature; and fuses the new global features of the infrared image and the attribute labels of the image pair to obtain a real infrared modal feature;
[0128] The adversarial network generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature.
[0129] The pedestrian re-identification device disclosed in the embodiment of the present application is described below. The pedestrian re-identification device described below and the pedestrian re-identification method described above can be referenced to each other.
[0130] Reference Figure 7 , Figure 7 Schematic diagram of the structure of a pedestrian re-identification device disclosed in an embodiment of the present application. Figure 7 As shown, the pedestrian re-identification device may include:
[0131] An acquisition unit 11, used for acquiring an image to be queried;
[0132] A feature extraction unit 12, configured to input the image to be queried into a feature extraction model, and the feature extraction model outputs the features of the image to be queried; the feature extraction model is obtained by training based on an image pair including an RGB image and an infrared image;
[0133] The matching unit 13 is used to match the features of the image to be queried with the features of each pedestrian image in the query database to obtain a pedestrian re-identification result corresponding to the image to be queried.
[0134] As an implementable method, the feature extraction model is trained as follows:
[0135] Obtaining a pre-built feature extraction model training network, wherein the feature extraction model training network includes a feature extraction module, an attribute classification module, a feature fusion module, a feature alignment module, and an identity prediction module;
[0136] Acquire a training image pair, wherein the training image pair includes an RGB image and an infrared image; each training image pair is annotated with an identity label and an attribute label;
[0137] The training image pairs are used as training samples, the identity labels and attribute labels annotated by the training image pairs are used as sample labels, and the feature extraction model training network is trained with the joint loss of the attribute classification module, the identity prediction module and the feature alignment module. When the feature extraction model training network converges, the feature extraction module is the feature extraction model.
[0138] As an implementation method, during the training process of the feature extraction model training network, the feature extraction module extracts features from the RGB image and the infrared image in the training image pair to obtain global features of the RGB image, global features of the infrared image, and overall local features of the image pair;
[0139] The attribute classification module generates new overall local features of the image pair based on the overall local features of the image pair, and predicts the attributes of the image pair based on the overall local features of the image pair to obtain an attribute prediction result, wherein the difference between the attribute prediction result and the attribute label of the training image pair is the loss of the attribute classification module;
[0140] The feature fusion module fuses the new overall local features of the image pair with the global features of the RGB image to obtain new global features of the RGB image, and fuses the new overall local features of the image pair with the global features of the infrared image to obtain new global features of the infrared image;
[0141] The feature alignment module obtains a synthesized infrared modal feature based on the new global feature of the RGB image and the attribute label of the image pair; obtains a real infrared modal feature based on the new global feature of the infrared image and the attribute label of the image pair; generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature, and the difference between the intermediate modal feature and the synthesized infrared modal feature and the difference between the intermediate modal feature and the real infrared modal feature are the losses of the feature alignment module;
[0142] The identity prediction module predicts the identity of the image pair based on the new global features of the RGB image and the new global features of the infrared image to obtain an identity prediction result, and the difference between the identity prediction result and the identity label of the training image pair is the loss of the identity prediction module.
[0143] As an implementable embodiment, the feature extraction module includes an input layer, an intermediate layer and a final layer;
[0144] The feature extraction module extracts features from the RGB image and the infrared image in the training image pair to obtain global features of the RGB image, global features of the infrared image, and overall local features of the image pair, including:
[0145] The input layer generates an input sequence of RGB images in the training image pair, and an input sequence of infrared images in the training image pair;
[0146] The intermediate layer encodes the input sequence of the RGB image to obtain local features of the RGB image, and encodes the input sequence of the infrared image to obtain local features of the infrared image;
[0147] The last layer encodes the local features of the RGB image to obtain the global features of the RGB image, and encodes the local features of the infrared image to obtain the global features of the infrared image.
[0148] As an implementation method, the input layer generates an input sequence of RGB images in the training image pair, including:
[0149] Obtaining a feature sequence and a position sequence of each block corresponding to the RGB image in the training image pair;
[0150] Combining the feature sequence of each block corresponding to the preset mark and the RGB image, and the position sequence into an input sequence of the RGB image;
[0151] The input layer generates an input sequence of infrared images in the training image pair, including:
[0152] Obtaining a feature sequence and a position sequence of each block corresponding to the infrared image in the training image pair;
[0153] The input sequence of the infrared image is composed of the feature sequence of each block corresponding to the preset mark and the position sequence of the infrared image.
[0154] As an implementable embodiment, the attribute classification module includes: an attention layer, an attribute prediction layer, the attribute classification module generates new overall local features of the image pair based on the overall local features of the image pair, and predicts the attributes of the image pair based on the overall local features of the image pair to obtain attribute prediction results, including:
[0155] The attention layer generates attribute features of the image pair based on the overall local features of the image pair, and the attribute features of the image pair are combined to obtain new overall local features of the image pair;
[0156] The attribute prediction layer predicts the attributes of the image pair based on the attribute features of the image pair to obtain an attribute prediction result.
[0157] As an implementable embodiment, the feature alignment module includes:
[0158] Infrared code generator, attribute label fusion layer, adversarial network;
[0159] The feature alignment module obtains a synthesized infrared modal feature based on the new global feature of the RGB image and the attribute label of the image pair; obtains a real infrared modal feature based on the new global feature of the infrared image and the attribute label of the image pair; and generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature, including:
[0160] The infrared code generator encodes the new global features of the RGB image to obtain the encoded infrared modal features;
[0161] The attribute label fusion layer fuses the encoded infrared modal features and the attribute labels of the image pair to obtain a synthesized infrared modal feature; and fuses the new global features of the infrared image and the attribute labels of the image pair to obtain a real infrared modal feature;
[0162] The adversarial network generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature.
[0163] Reference Figure 8 , Figure 8The hardware structure block diagram of the pedestrian re-identification device provided in the embodiment of the present application is shown in FIG. Figure 8 ,The hardware structure of the pedestrian re-identification device may include: at least one processor 1, at least one communication interface 2, at least one memory 3 and at least one communication bus 4;
[0164] In the embodiment of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 communicate with each other through the communication bus 4;
[0165] The processor 1 may be a central processing unit CPU, or an application-specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;
[0166] The memory 3 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory;
[0167] The memory stores a program, and the processor can call the program stored in the memory, wherein the program is used to:
[0168] Get the image to be queried;
[0169] Inputting the image to be queried into a feature extraction model, the feature extraction model outputting features of the image to be queried; the feature extraction model is trained based on an image pair including an RGB image and an infrared image;
[0170] The features of the image to be queried are matched with the features of each pedestrian image in the query database to obtain a pedestrian re-identification result corresponding to the image to be queried.
[0171] Optionally, the detailed functions and extended functions of the program may refer to the above description.
[0172] The embodiment of the present application further provides a readable storage medium, which may store a program suitable for execution by a processor, wherein the program is used to:
[0173] Get the image to be queried;
[0174] Inputting the image to be queried into a feature extraction model, the feature extraction model outputting features of the image to be queried; the feature extraction model is trained based on an image pair including an RGB image and an infrared image;
[0175] The features of the image to be queried are matched with the features of each pedestrian image in the query database to obtain a pedestrian re-identification result corresponding to the image to be queried.
[0176] Optionally, the detailed functions and extended functions of the program may refer to the above description.
[0177] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0178] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0179] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A pedestrian re-identification method, characterized in that: The method comprises: Get the image to be queried; Inputting the image to be queried into a feature extraction model, the feature extraction model outputting features of the image to be queried; the feature extraction model is trained based on an image pair including an RGB image and an infrared image; Matching the features of the image to be queried with the features of each pedestrian image in the query database to obtain a pedestrian re-identification result corresponding to the image to be queried; The feature extraction model is trained based on a pre-built feature extraction model training network, wherein the feature extraction model training network includes a feature extraction module, an attribute classification module, a feature fusion module, a feature alignment module, and an identity prediction module; During the training process of the feature extraction model training network, the feature extraction module performs feature extraction on the RGB image and the infrared image in the training image pair to obtain the global features of the RGB image, the global features of the infrared image, and the overall local features of the image pair; The attribute classification module generates new overall local features of the image pair based on the overall local features of the image pair, and predicts the attributes of the image pair based on the overall local features of the image pair to obtain an attribute prediction result, wherein the difference between the attribute prediction result and the attribute label of the training image pair is the loss of the attribute classification module; The feature fusion module fuses the new overall local features of the image pair with the global features of the RGB image to obtain new global features of the RGB image, and fuses the new overall local features of the image pair with the global features of the infrared image to obtain new global features of the infrared image; The feature alignment module obtains a synthesized infrared modal feature based on the new global feature of the RGB image and the attribute label of the image pair; obtains a real infrared modal feature based on the new global feature of the infrared image and the attribute label of the image pair; generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature, and the difference between the intermediate modal feature and the synthesized infrared modal feature and the difference between the intermediate modal feature and the real infrared modal feature are the losses of the feature alignment module; The identity prediction module predicts the identity of the image pair based on the new global features of the RGB image and the new global features of the infrared image to obtain an identity prediction result, and the difference between the identity prediction result and the identity label of the training image pair is the loss of the identity prediction module.
2. The method according to claim 1, characterized in that The feature extraction model is trained as follows: Acquire a training image pair, wherein the training image pair includes an RGB image and an infrared image; each training image pair is annotated with an identity label and an attribute label; The training image pairs are used as training samples, the identity labels and attribute labels annotated by the training image pairs are used as sample labels, and the feature extraction model training network is trained with the joint loss of the attribute classification module, the identity prediction module and the feature alignment module. When the feature extraction model training network converges, the feature extraction module is the feature extraction model.
3. The method according to claim 1, characterized in that The feature extraction module includes an input layer, an intermediate layer and a final layer; The feature extraction module extracts features from the RGB image and the infrared image in the training image pair to obtain global features of the RGB image, global features of the infrared image, and overall local features of the image pair, including: The input layer generates an input sequence of RGB images in the training image pair, and an input sequence of infrared images in the training image pair; The intermediate layer encodes the input sequence of the RGB image to obtain local features of the RGB image, and encodes the input sequence of the infrared image to obtain local features of the infrared image; The last layer encodes the local features of the RGB image to obtain the global features of the RGB image, and encodes the local features of the infrared image to obtain the global features of the infrared image.
4. The method according to claim 3, characterized in that The input layer generates an input sequence of RGB images in the training image pair, including: Obtaining a feature sequence and a position sequence of each block corresponding to the RGB image in the training image pair; Combining the feature sequence of each block corresponding to the preset mark and the RGB image, and the position sequence into an input sequence of the RGB image; The input layer generates an input sequence of infrared images in the training image pair, including: Obtaining a feature sequence and a position sequence of each block corresponding to the infrared image in the training image pair; The input sequence of the infrared image is composed of the feature sequence of each block corresponding to the preset mark and the position sequence of the infrared image.
5. The method according to claim 1, characterized in that: The attribute classification module includes: an attention layer, an attribute prediction layer, the attribute classification module generates new overall local features of the image pair based on the overall local features of the image pair, and predicts the attributes of the image pair based on the overall local features of the image pair to obtain attribute prediction results, including: The attention layer generates attribute features of the image pair based on the overall local features of the image pair, and the attribute features of the image pair are combined to obtain new overall local features of the image pair; The attribute prediction layer predicts the attributes of the image pair based on the attribute features of the image pair to obtain an attribute prediction result.
6. The method according to claim 1, characterized in that The feature alignment module comprises: Infrared code generator, attribute label fusion layer, adversarial network; The feature alignment module obtains a synthesized infrared modal feature based on the new global feature of the RGB image and the attribute label of the image pair; obtains a real infrared modal feature based on the new global feature of the infrared image and the attribute label of the image pair; and generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature, including: The infrared code generator encodes the new global features of the RGB image to obtain the encoded infrared modal features; The attribute label fusion layer fuses the encoded infrared modal features and the attribute labels of the image pair to obtain a synthesized infrared modal feature; and fuses the new global features of the infrared image and the attribute labels of the image pair to obtain a real infrared modal feature; The adversarial network generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature.
7. A pedestrian re-identification device, characterized in that: The device comprises: An acquisition unit, used for acquiring an image to be queried; A feature extraction unit, used for inputting the image to be queried into a feature extraction model, wherein the feature extraction model outputs features of the image to be queried; the feature extraction model is obtained by training based on an image pair including an RGB image and an infrared image; A matching unit, used for matching the features of the image to be queried with the features of each pedestrian image in the query database to obtain a pedestrian re-identification result corresponding to the image to be queried; The feature extraction model is trained based on a pre-built feature extraction model training network, wherein the feature extraction model training network includes a feature extraction module, an attribute classification module, a feature fusion module, a feature alignment module, and an identity prediction module; During the training process of the feature extraction model training network, the feature extraction module performs feature extraction on the RGB image and the infrared image in the training image pair to obtain the global features of the RGB image, the global features of the infrared image, and the overall local features of the image pair; The attribute classification module generates new overall local features of the image pair based on the overall local features of the image pair, and predicts the attributes of the image pair based on the overall local features of the image pair to obtain an attribute prediction result, wherein the difference between the attribute prediction result and the attribute label of the training image pair is the loss of the attribute classification module; The feature fusion module fuses the new overall local features of the image pair with the global features of the RGB image to obtain new global features of the RGB image, and fuses the new overall local features of the image pair with the global features of the infrared image to obtain new global features of the infrared image; The feature alignment module obtains a synthesized infrared modal feature based on the new global feature of the RGB image and the attribute label of the image pair; obtains a real infrared modal feature based on the new global feature of the infrared image and the attribute label of the image pair; generates an intermediate modal feature based on the synthesized infrared modal feature and the real infrared modal feature, and the difference between the intermediate modal feature and the synthesized infrared modal feature and the difference between the intermediate modal feature and the real infrared modal feature are the losses of the feature alignment module; The identity prediction module predicts the identity of the image pair based on the new global features of the RGB image and the new global features of the infrared image to obtain an identity prediction result, and the difference between the identity prediction result and the identity label of the training image pair is the loss of the identity prediction module.
8. A pedestrian re-identification device, characterized in that: including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the pedestrian re-identification method according to any one of claims 1 to 6.
9. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the pedestrian re-identification method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Pedestrian re-identification method and device, electronic equipment and storage medium
CN111967314A