A pedestrian retrieval method based on hybrid component transformation network

By introducing a hybrid component transformation network in pedestrian retrieval, the component interaction information of pedestrian images is learned, and feature discrimination is improved through sequence block screening steps, the problem of limited pedestrian retrieval performance in the prior art is solved, and higher accuracy and feature representation ability are achieved.

CN116246305BActive Publication Date: 2025-05-23TIANJIN NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310081039.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-31
Publication Date
2025-05-23
Estimated Expiration
2043-01-31

AI Technical Summary

Technical Problem

In the process of processing the query of the same pedestrian under multiple different cameras, it is difficult to effectively utilize the component interaction information in the pedestrian image, resulting in limited retrieval performance.

Method used

A pedestrian search method based on a hybrid component transformation network is proposed. The complete component interaction information of pedestrian images is learned through the cascaded transform network model and the component global transform network model, and the sequence block with more information is retained through the sequence block screening step, thereby improving the discriminantity of pedestrian characteristics.

Benefits of technology

By learning complete component interaction information, the accuracy of pedestrian retrieval is significantly improved and the feature representation ability of pedestrian images is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246305B_ABST
    Figure CN116246305B_ABST
Patent Text Reader

Abstract

The present invention discloses a pedestrian retrieval method based on a hybrid component transformation network. The method comprises: constructing a pedestrian retrieval model; segmenting a training pedestrian image to obtain stripe components and sequence blocks of the training pedestrian image; inputting the sequence blocks into the pedestrian retrieval model to obtain stripe component features and complete features of the training pedestrian image; calculating component masks of the stripe component features, screening, and retaining some sequence blocks; calculating loss values ​​using the stripe component features and complete features corresponding to the retained sequence blocks, and optimizing the pedestrian retrieval model; extracting the final features of the query image and the pedestrian library image using the optimal pedestrian retrieval model, and obtaining the pedestrian retrieval result by means of the similarity between the query image and the pedestrian library image. The present invention makes full use of the advantages of the hybrid component transformation network, learns the complete component information of the pedestrian image, and further improves the accuracy of pedestrian retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the fields of computer vision, pattern recognition and artificial intelligence, and in particular relates to a pedestrian retrieval method based on a hybrid component transformation network. Background Art

[0002] In recent years, pedestrian retrieval has been widely used in the fields of human behavior analysis and multi-target recognition, and has therefore attracted extensive attention from academia and industry. Pedestrian retrieval mainly studies methods for querying the same pedestrian under multiple different cameras. However, there are many difficult factors in pedestrian images obtained in real scenes, such as posture, clothing, lighting, and camera angle, which makes pedestrian retrieval technology face huge challenges.

[0003] In recent years, the component information of pedestrian images has been proven to be effective in pedestrian retrieval. However, the interaction between components is often ignored when using transformer networks to learn long-range dependencies. He et al. proposed a pure transformer network to learn discriminative features by using auxiliary information embedding and patch block token reordering modules. In addition, some researchers combined convolutional neural networks and transformer networks for pedestrian retrieval. Liao et al. designed an encoder-decoder transformer network by combining convolutional neural networks to consider the attention between pedestrian images. Zhang et al. proposed a hierarchical aggregation transformer network to learn multi-scale pedestrian features by embedding a transformer network into a convolutional neural network. Li et al. proposed a component-aware transformer network to learn robust component diversity features from the perspective of semantic information by combining convolutional neural networks. Wang et al. first used convolutional neural networks to learn pedestrian posture information and then decoupled the semantic information of pedestrian images through a transformer network. Wang et al. proposed a neighborhood transformer network to explicitly model the interaction between images to improve the performance of pedestrian retrieval.

[0004] Different from the above methods, the present invention proposes a transformation network model and a component global transformation network model to learn a hybrid component transformation network with complete component interaction for pedestrian retrieval. In addition, the present invention also proposes a sequence block screening step to improve the discriminability of pedestrian features by retaining sequence blocks with more information. Summary of the invention

[0005] The purpose of the present invention is to design a transformation network suitable for learning the complete component interaction of pedestrian images. To this end, the present invention provides a pedestrian retrieval method based on a hybrid component transformation network.

[0006] In order to achieve the above object, the present invention proposes a pedestrian retrieval method based on a hybrid component transformation network, comprising the following steps:

[0007] Step S1, constructing a pedestrian retrieval model using a pre-trained deep learning model, wherein the pedestrian retrieval model includes a cascaded transformation network model and a component global transformation network model;

[0008] Step S2, segmenting the training pedestrian image to obtain stripe components of the training pedestrian image and sequence blocks of the stripe components;

[0009] Step S3, inputting the sequence blocks of the stripe components of the training pedestrian image into the pedestrian retrieval model to obtain the stripe component features of the training pedestrian image and the complete features of the training pedestrian image;

[0010] Step S4, calculating the component mask of the stripe component feature by using the attention weight of the affinity matrix of the component global transformation layer in the component global transformation network model and a preset threshold, and screening the sequence blocks according to the component mask to retain some sequence blocks;

[0011] Step S5, constructing a loss calculation module, inputting the retained sequence blocks into the transformation network model to obtain the stripe component features and the complete features of the training pedestrian image into the loss calculation module, and optimizing the pedestrian retrieval model using the obtained loss value to obtain the optimal pedestrian retrieval model;

[0012] Step S6, in the test phase, the optimal pedestrian retrieval model is used to extract the final features of the query image and the pedestrian library image, and the similarity between the query image and the pedestrian library image is calculated based on the final features to obtain the pedestrian retrieval result.

[0013] Optionally, the step S1 includes the following steps:

[0014] Step S11, determining a pre-trained deep learning model, and using the pre-trained deep learning model to construct a transformation network model and a component global transformation network model to obtain a pedestrian retrieval model;

[0015] Step S12, initializing parameters of the transformation network model and the component global transformation network model.

[0016] Optionally, step S2 includes the following steps:

[0017] Step S21, preprocessing N training pedestrian images in the training set;

[0018] Step S22, performing horizontal segmentation on the preprocessed training pedestrian image to obtain stripe components of the training pedestrian image;

[0019] Step S23, serializing the stripe component to obtain a plurality of sequence blocks of the stripe component.

[0020] Optionally, step S3 includes the following steps:

[0021] Step S31, inputting a sequence block of stripe components of a single training pedestrian image into the pedestrian retrieval model, and the output of the last transformation layer of the transformation network model is the stripe component feature of the training pedestrian image;

[0022] Step S32, performing maximum pooling aggregation on the output of the last component global transformation layer of the component global transformation network model to obtain the complete features of the training pedestrian image.

[0023] Optionally, in step S31, a class token is added during the learning process of each stripe component sequence block in the transformation network model. Multi-head self-attention learning is performed, where the class token is a feature vector used to learn the features of the stripe components.

[0024] Optionally, step S4 includes the following steps:

[0025] Step S41, calculating the attention weight of the sequence block in each stripe component in the training pedestrian image based on the affinity matrix of the component global transformation layer in the component global transformation network model;

[0026] Step S42, using the obtained attention weight of the sequence block in each stripe component in the training pedestrian image and a preset threshold, calculate the component mask of the stripe component feature;

[0027] Step S43: retain the sequence blocks whose component mask value is 1.

[0028] Optionally, the loss calculation module includes a cross entropy loss calculation module and a triplet loss calculation module.

[0029] Optionally, step S5 includes the following steps:

[0030] Step S51, constructing a loss calculation module, and using the loss calculation module to calculate the cross entropy loss and triplet loss of the stripe component features obtained by inputting the retained sequence blocks into the transformation network model and the complete features of the training pedestrian image;

[0031] Step S52, adding up the calculated losses to obtain a total loss value, and using the total loss value to optimize the parameters of the pedestrian retrieval model to obtain an optimal pedestrian retrieval model.

[0032] Optionally, in step S6, the final feature is a feature obtained by connecting the complete feature of the pedestrian image and the stripe component feature corresponding to the retained sequence block in series.

[0033] Optionally, in step S6, the similarity between the query image and the pedestrian library image is calculated using cosine distance.

[0034] The beneficial effects of the present invention are as follows: the present invention proposes to use a hybrid component transformation network designed by the present invention to learn complete component interactions for pedestrian retrieval, making full use of the advantages of the hybrid component transformation network to learn complete component information of pedestrian images. In addition, the present invention also designs a sequence block screening step to improve the discriminability of pedestrian features by retaining sequence blocks with more information, so that the present invention scheme effectively improves the accuracy of pedestrian retrieval.

[0035] It should be noted that this invention was funded by the National Natural Science Foundation of China Project No.62171321, the Tianjin Natural Science Foundation Key Project No.20JCZDJC00180, the Tianjin Municipal Education Commission Scientific Research Plan Project No.2022KJ011, the Tianjin Applied Basic Research Project (Research on Stereo Image Stitching Technology Based on Depth Preservation) and the Tianjin Normal University Postgraduate Scientific Research Innovation Key Project No.2022KYCX032Z. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a flowchart of a pedestrian retrieval method based on a hybrid component transformation network according to an embodiment of the present invention. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessary confusion of the concept of the present invention.

[0038] Figure 1 is a flowchart of a pedestrian retrieval method based on a hybrid component transformation network according to an embodiment of the present invention. Figure 1 Take some specific implementation processes of the present invention as an example. Figure 1 As shown, the pedestrian retrieval method based on the hybrid component transformation network includes the following steps:

[0039] Step S1, constructing a pedestrian retrieval model using a pre-trained deep learning model, wherein the pedestrian retrieval model includes a cascaded transformation network model and a component global transformation network model;

[0040] Furthermore, the step S1 includes the following steps:

[0041] Step S11, determining a pre-trained deep learning model, and using the pre-trained deep learning model to construct a transformation network model and a component global transformation network model to obtain a pedestrian retrieval model;

[0042] In one embodiment of the present invention, the pre-trained deep learning model may use ViT pre-trained on the dataset ImageNet-21K and fine-tuned on the dataset ImageNet-1K, or DeiT pre-trained on the dataset ImageNet-1K.

[0043] In one embodiment of the present invention, the structure of the transformation network model is the same as that of the pre-trained deep learning model, that is, the transformation network model includes L transformation layers. The component global transformation network model also includes L component global transformation layers. The transformation network model and the component global transformation network model are cascaded to obtain the pedestrian retrieval model, wherein the L transformation layers of the transformation network model are cascaded in sequence, and the L component global transformation layers of the component global transformation network model are cascaded in sequence. In addition, the L transformation layers of the transformation network model are also correspondingly connected to the L component global transformation layers of the component global transformation network model, that is, except for the first layer of the L transformation layers of the transformation network model, the input of each transformation layer is the output of the previous transformation layer, and among the L component global transformation layers of the component global transformation network model, except for the input of the first layer of the component global transformation layer which is the output of the first layer of the transformation layer of the transformation network model, the input of each subsequent component global transformation layer includes not only the output of the previous component global transformation layer, but also the output of the corresponding transformation layer in the transformation network model.

[0044] Step S12, initializing parameters of the transformation network model and the component global transformation network model.

[0045] In one embodiment of the present invention, the parameters of the pre-trained deep learning model can be used to initialize the parameters of the transformation network model and the component global transformation network model.

[0046] Step S2, segmenting the training pedestrian image to obtain stripe components of the training pedestrian image and sequence blocks of the stripe components;

[0047] Furthermore, the step S2 comprises the following steps:

[0048] Step S21, preprocessing N training pedestrian images in the training set;

[0049] In one embodiment of the present invention, preprocessing the training pedestrian image includes: cropping the size of the training pedestrian image to a preset size, such as 256×128, and proportionally reducing all pixel values ​​of the training pedestrian image to a preset range, such as between 0 and 1, and then subtracting the pixel average of the corresponding training pedestrian image from each pixel value in the training pedestrian image, and then dividing it by the pixel variance of the training pedestrian image.

[0050] Step S22, horizontally segmenting the preprocessed training pedestrian image to obtain stripe components of the training pedestrian image, wherein each training pedestrian image can obtain S stripe components, and thus N training pedestrian images can obtain (N×S) stripe components;

[0051] In one embodiment of the present invention, for each training pedestrian image Perform horizontal segmentation, that is, segmentation in the height direction, where H, W, and C are the height, width, and number of channels of the training pedestrian image, respectively, and the sub-image obtained by segmentation is the stripe component of the training pedestrian image.

[0052] Step S23, serializing the stripe components to obtain a plurality of sequence blocks of the stripe components, wherein the i-th sequence block of the p-th stripe component can be expressed as: Among them, K×K is the size of the sequence block, M is the number of sequence blocks in each stripe component, and S is the number of stripe components in each training pedestrian image.

[0053] In one embodiment of the present invention, H=128, W=256, C=3, S=2, N=64, and K=16.

[0054] Step S3, inputting the sequence blocks of the stripe components of the training pedestrian image into the pedestrian retrieval model to obtain the stripe component features of the training pedestrian image and the complete features of the training pedestrian image;

[0055] Furthermore, step S3 includes the following steps:

[0056] Step S31, inputting a sequence block of stripe components of a single training pedestrian image into the pedestrian retrieval model, and the output of the last transformation layer of the transformation network model, i.e., the Lth transformation layer, is the stripe component feature of the training pedestrian image;

[0057] Furthermore, a class token can be added to the learning process of each stripe component sequence block in the transformation network model. Multi-head self-attention learning is performed, where the class token is a feature vector used to learn the stripe component features, so that the stripe component features can be expressed as: Wherein, D is the size of the class token.

[0058] Step S32, performing maximum pooling aggregation on the output of the last component global transformation layer of the component global transformation network model, that is, the Lth component global transformation layer, to obtain the complete features of the training pedestrian image, which can be expressed as:

[0059] In one embodiment of the present invention, the output of the lth component global transformation layer of the component global transformation network model It can be calculated using the following formula:

[0060]

[0061]

[0062] in, It is a function based on multi-head cross attention (MCA), multi-layer perceptron (MLP) and layer normalization (LN). It is obtained by aggregating the output of the l-1th transformation layer of the transformation network model, where Q is the number of all sequence blocks in a single training pedestrian image, It is a function based on multi-head cross attention (MCA) and layer normalization (LN), stripe component features Connect in series

[0063] In addition, the multi-head cross attention value (MCA) of the lth component global transformation layer of the component global transformation network model can be calculated using the following formula:

[0064]

[0065]

[0066]

[0067] Where a represents T l-1 or Y l-1 , b represents G l-1 or C l-1 , cat2 means row concatenation, represents a linear projection, and They represent three linear projection parameters respectively, H is the number of heads in the multi-head self-attention mechanism, h represents the h-th head in H heads, d = D / H, Represents the affinity matrix in the component global transformation network model.

[0068] In one embodiment of the present invention, B=128, D=768, and H=12.

[0069] Step S4, calculating the component mask of the stripe component feature by using the attention weight of the affinity matrix of the component global transformation layer in the component global transformation network model and a preset threshold, and screening the sequence blocks according to the component mask to retain some sequence blocks;

[0070] Furthermore, the step S4 comprises the following steps:

[0071] Step S41, calculating the attention weight of the sequence block in each stripe component in the training pedestrian image based on the affinity matrix of the component global transformation layer in the component global transformation network model

[0072] In one embodiment of the present invention, the attention weight of the sequence block in each stripe component in the training pedestrian image is calculated using the following formula:

[0073]

[0074] Where H is the number of heads in the multi-head self-attention mechanism, represents the attention weight of the p-th row of the affinity matrix of the h-th head in the global transformation layer of the j-th component, M is the number of all sequence blocks in each stripe component, p = 1, 2, ..., S, l = 2, ..., L, ((p-1)·M+1): p·M represents from (p-1)·M+1) to p·M.

[0075] Step S42, using the obtained attention weights of the sequence blocks in each stripe component in the training pedestrian image and a preset threshold value, and calculate the component mask of the stripe component feature

[0076] In one embodiment of the present invention, the component mask of the stripe component feature can be calculated using the following formula:

[0077]

[0078] Where i = 1, 2, ..., M, and τ is a preset threshold to retain the sequence blocks with more information in the features of each stripe component.

[0079] In one embodiment of the present invention, τ=0.3.

[0080] Step S43: retain the sequence blocks whose component mask value is 1.

[0081] Step S5, constructing a loss calculation module, inputting the retained sequence blocks into the transformation network model to obtain the stripe component features and the complete features of the training pedestrian image into the loss calculation module, and optimizing the pedestrian retrieval model using the obtained loss value to obtain the optimal pedestrian retrieval model;

[0082] Furthermore, the step S5 comprises the following steps:

[0083] Step S51, constructing a loss calculation module, and using the loss calculation module to calculate the stripe component features obtained by inputting the retained sequence blocks into the transformation network model and the complete features of the training pedestrian image Cross entropy loss and triplet loss;

[0084] The loss calculation module includes a cross entropy loss calculation module and a triplet loss calculation module. j and the predicted value p j , the cross entropy loss calculation module can calculate the cross entropy loss using the following formula:

[0085]

[0086] Where N is the maximum value of j.

[0087] Given a triple set {a, p, n}, the triple loss calculation module may calculate the triple loss using the following formula:

[0088]

[0089] Among them, f a represents the input sample, i.e., the stripe component feature or the complete feature, f p represents the positive sample of the input sample, f n Represents the negative sample of the input sample.

[0090] Step S52, adding up the calculated losses to obtain a total loss value Loss, and using the total loss value to optimize the parameters of the pedestrian retrieval model to obtain an optimal pedestrian retrieval model.

[0091] In one embodiment of the present invention, the total loss function Loss can be expressed as:

[0092]

[0093] in, and represent the cross entropy loss and triplet loss of the complete features of the training pedestrian image, respectively, and They respectively represent the cross entropy loss and triplet loss of the sequence block of the p-th stripe component of the training pedestrian image.

[0094] In one embodiment of the present invention, in step S52, the parameter update calculation process of the pedestrian retrieval model can be expressed as:

[0095]

[0096] Among them, θ s : is the updated model parameter of the pedestrian retrieval model, θ s are the model parameters before the pedestrian retrieval model is updated, and σ is the learning rate.

[0097] In one embodiment of the present invention, an optimizer based on stochastic gradient descent (SGD) and cosine decay strategy may be used to optimize the pedestrian retrieval model, with a learning rate σ=0.01.

[0098] Step S6, in the test phase, the optimal pedestrian retrieval model is used to extract the final features of the query image and the pedestrian library image, wherein the final features are the features obtained by concatenating the complete features of the pedestrian image and the stripe component features corresponding to the retained sequence blocks, and the similarity between the query image and the pedestrian library image is calculated based on the final features to obtain the pedestrian retrieval result.

[0099] In one embodiment of the present invention, based on the final feature, the similarity between the query image and the pedestrian library image is calculated using the cosine distance, wherein the pedestrian library image refers to an image with a known pedestrian recognition result.

[0100] The similarity between the query image and the pedestrian library image can be expressed as:

[0101]

[0102] Among them, C qg Refers to the final feature I of the query image q And the final feature I of the pedestrian library image g The cosine similarity between .

[0103] It should be understood that the above specific embodiments of the present invention are only used to illustrate or explain the principles of the present invention, and do not constitute a limitation of the present invention. Therefore, any modifications, equivalent substitutions, improvements, etc. made without departing from the spirit and scope of the present invention should be included in the protection scope of the present invention. In addition, the appended claims of the present invention are intended to cover all changes and modifications that fall within the scope and boundaries of the appended claims, or the equivalent forms of such scope and boundaries.

Claims

1. A pedestrian retrieval method based on hybrid component transformation network, It is characterized in that The method comprises the following steps: Step S1, constructing a pedestrian retrieval model using a pre-trained deep learning model, wherein the pedestrian retrieval model includes a cascaded transformation network model and a component global transformation network model; Step S2, segmenting the training pedestrian image to obtain stripe components of the training pedestrian image and sequence blocks of the stripe components; Step S3, inputting the sequence blocks of the stripe components of the training pedestrian image into the pedestrian retrieval model to obtain the stripe component features of the training pedestrian image and the complete features of the training pedestrian image; Step S4, calculating the component mask of the stripe component feature by using the attention weight of the affinity matrix of the component global transformation layer in the component global transformation network model and a preset threshold, and screening the sequence blocks according to the component mask to retain some sequence blocks; Step S5, constructing a loss calculation module, inputting the retained sequence blocks into the transformation network model to obtain the stripe component features and the complete features of the training pedestrian image into the loss calculation module, and optimizing the pedestrian retrieval model using the obtained loss value to obtain the optimal pedestrian retrieval model; Step S6, in the test phase, the optimal pedestrian retrieval model is used to extract the final features of the query image and the pedestrian library image, and the similarity between the query image and the pedestrian library image is calculated based on the final features to obtain the pedestrian retrieval result.

2. The method according to claim 1, It is characterized in that The step S1 comprises the following steps: Step S11, determining a pre-trained deep learning model, and using the pre-trained deep learning model to construct a transformation network model and a component global transformation network model to obtain a pedestrian retrieval model; Step S12, initializing parameters of the transformation network model and the component global transformation network model.

3. The method according to claim 2, It is characterized in that The step S2 comprises the following steps: Step S21, preprocessing N training pedestrian images in the training set; Step S22, performing horizontal segmentation on the preprocessed training pedestrian image to obtain stripe components of the training pedestrian image; Step S23, serializing the stripe component to obtain a plurality of sequence blocks of the stripe component.

4. The method according to claim 1, It is characterized in that The step S3 comprises the following steps: Step S31, inputting a sequence block of stripe components of a single training pedestrian image into the pedestrian retrieval model, and the output of the last transformation layer of the transformation network model is the stripe component feature of the training pedestrian image; Step S32, performing maximum pooling aggregation on the output of the last component global transformation layer of the component global transformation network model to obtain the complete features of the training pedestrian image.

5. The method according to claim 4, It is characterized in that In step S31, a class token is added during the learning process of each stripe component sequence block in the transformation network model. Multi-head self-attention learning is performed, where the class token is a feature vector used to learn the features of the stripe components.

6. The method according to claim 1, It is characterized in that The step S4 comprises the following steps: Step S41, calculating the attention weight of the sequence block in each stripe component in the training pedestrian image based on the affinity matrix of the component global transformation layer in the component global transformation network model; Step S42, using the obtained attention weight of the sequence block in each stripe component in the training pedestrian image and a preset threshold, calculate the component mask of the stripe component feature; Step S43: retain the sequence blocks whose component mask value is 1.

7. The method according to claim 1, It is characterized in that The loss calculation module includes a cross entropy loss calculation module and a triplet loss calculation module.

8. The method according to claim 7, It is characterized in that The step S5 comprises the following steps: Step S51, constructing a loss calculation module, and using the loss calculation module to calculate the cross entropy loss and triplet loss of the stripe component features obtained by inputting the retained sequence blocks into the transformation network model and the complete features of the training pedestrian image; Step S52, adding up the calculated losses to obtain a total loss value, and using the total loss value to optimize the parameters of the pedestrian retrieval model to obtain an optimal pedestrian retrieval model.

9. The method according to claim 1, It is characterized in that In step S6, the final feature is a feature obtained by connecting the complete feature of the pedestrian image and the stripe component feature corresponding to the retained sequence block in series.

10. The method according to claim 1, It is characterized in that In step S6, the similarity between the query image and the pedestrian library image is calculated using the cosine distance.

Citation Information

Patent Citations

  • Mask pooling model training and pedestrian re-identification method for pedestrian re-identification

    CN109977798A

  • Global and local sensing pedestrian re-identification method fusing segmentation information

    CN113657355A