Cross-resolution pedestrian re-identification method based on double-flow input feature reconstruction

By adopting the dual-stream input feature reconstruction method in cross-resolution pedestrian recognition, the problems of low-resolution image information loss and feature mismatch are solved, and a higher recognition accuracy is achieved.

CN120071380AActive Publication Date: 2025-05-30SICHUAN UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202311603648.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30
Estimated Expiration
2043-11-28

AI Technical Summary

Technical Problem

The existing cross-resolution pedestrian recognition method can easily lead to the loss of key discriminant information when processing low-resolution images, and there is a mismatch problem in pedestrian feature extraction at different resolutions, making it difficult to effectively deal with the problem of resolution mismatch in practical applications.

Method used

Using a method based on dual-stream input feature reconstruction, a cross-resolution pedestrian image pair is constructed, and a dual-stream reconstruction network and feature enhancement module are used to combine feature degradation losses and reconstruction losses to learn feature reconstruction capabilities, and acquire deep invariant features for cross-resolution pedestrian matching.

Benefits of technology

The image information is effectively restored, the pedestrian feature distance at different resolutions is reduced, and the accuracy of pedestrian re-identification across resolutions is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004574882980000052
    Figure BDA0004574882980000052
  • Figure BDA0004574882980000068
    Figure BDA0004574882980000068
  • Figure BDA0004574882980000069
    Figure BDA0004574882980000069
Patent Text Reader

Abstract

The invention discloses a cross-resolution pedestrian re-identification method based on double-flow input feature reconstruction, and relates to the field of computer vision and artificial intelligence. The method comprises the following steps: firstly, constructing a double-flow input and feature reconstruction module, adaptively reconstructing features of a low-resolution image, and learning feature reconstruction capability by using feature degradation loss and reconstruction loss in combination with a lightweight residual decoder; secondly, constructing a weight-shared twin network structure, and obtaining depth invariant features of pedestrians in combination with a feature enhancement module; and finally, pooling the depth invariant features by using maximum average fusion pooling operation to obtain vectorized representation, and constraining distribution of pedestrian features by using cross-resolution triple loss, cross-resolution center loss and identity loss to improve cross-resolution matching precision. The method is mainly applied to a video monitoring intelligent analysis system, and has a wide application prospect in the fields of intelligent security and protection and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a cross-resolution person re-identification method based on dual-stream input feature reconstruction, and to the problem of cross-resolution person re-identification in the field of intelligent video surveillance, belonging to the fields of computer vision and security surveillance. Background Art

[0002] Person Re-Identification (ReID) is a task of great practical significance and challenge, whose goal is to identify the same target pedestrian under the perspectives of different surveillance cameras. In recent years, due to its practicality in the fields of pedestrian tracking, surveillance, etc., ReID has become a research hotspot in the field of computer vision. Although the person re-identification methods based on deep learning have made remarkable progress in aspects such as human body pose changes, background clutter, and partial occlusion, and have shown performance close to or even exceeding the human level in many public benchmark tests, most of these methods assume that the resolution of the query image is comparable to that of the image gallery. However, in practical applications, due to different camera performances and changes in the distance between the camera and the target pedestrian, the image resolution results in images of the same pedestrian at different resolution levels, which makes it difficult for current research to effectively address the resolution mismatch problem existing in pedestrian images in real application scenarios. To solve this problem, Cross-resolution person re-identification (CRReID) has emerged as a new research branch.

[0003] The challenges of CRReID are mainly reflected in two aspects: First, the blurred appearance of low-resolution pedestrian images often leads to the loss of key discriminative information. Second, due to the different resolutions of the collected images, pedestrians may present different fine-grained and detailed textures, so that for the same pedestrian image at different resolutions, the extracted features may have obvious mismatches. To solve this problem, many studies have proposed different cross-resolution pedestrian re-identification algorithms in recent years, which can be generally divided into two strategies: a) image resolution reconstruction and b) resolution-invariant feature extraction. Strategy a) uses super-resolution reconstruction technology to reconstruct low-resolution (LR) images, solves the resolution mismatch problem in cross-resolution pedestrian re-identification, and uses the super-resolution image as a bridge to match with the real high-resolution (HR) image. However, general super-resolution methods aim to improve the visual fidelity of images, and there may be compatibility problems when directly combined with the ReID network. Strategy b) notices the compatibility problem between super-resolution technology and pedestrian re-identification, and tries to learn pedestrian feature representations that are invariant or robust to resolution to achieve cross-resolution pedestrian comparison. Although this framework notices the recovery of information, this method of forcibly using the same ReID network to extract invariant representations at different resolutions is difficult to ensure the effective recovery of information, and there is still a large risk of losing detailed information. In summary, existing studies either have compatibility problems or cannot effectively recover the lost information. Therefore, it is necessary to develop a method that can effectively recover image information while solving the compatibility problem of super-resolution technology and organically combining super-resolution technology into the field of CRReID. Summary of the Invention

[0004] To solve the limitations of the prior art, the present invention proposes a cross-resolution pedestrian re-identification method based on dual-stream input feature reconstruction, aiming to complete feature reconstruction in the shallow layer of the framework and obtain depth-invariant features in the deep layer of the framework. A dual-stream input structure is designed for feature reconstruction of low-resolution images, and an image feature enhancement module is used to obtain depth-invariant features for cross-resolution pedestrian matching, so that image information is effectively recovered, and the distributions of pedestrian features at different resolutions are closer, improving the accuracy of cross-resolution pedestrian re-identification.

[0005] The present invention adopts the following technical solutions: A cross-resolution pedestrian re-identification method based on dual-stream input feature reconstruction, the method comprising the following steps:

[0006] (1) Construct cross-resolution pedestrian image pairs. First, obtain batches of HR images from the standard dataset according to the pedestrian identity labels, and then randomly downsample the HR images with a scale factor of {2, 3, 4}. The obtained LR images and HR images are used as the image pairs for training;

[0007] (2) Input the pedestrian image (pair) into the dual-stream reconstruction network. During this process, different input branches will be selected according to the resolution of the image. The low-resolution branch will adaptively reconstruct the features of the low-resolution image. The two branches will respectively obtain and During the training process, the feature degradation loss and the reconstruction loss are combined with the lightweight residual decoder (RLDecoder) to learn the feature reconstruction ability;

[0008] (3) Input and into the weight-sharing siamese network structure. First, pass through the fourth and fifth residual blocks of the backbone to obtain the preliminarily embedded resolution-invariant features Then input it into the feature enhancement module (MGLIF) to obtain the pedestrian depth-invariant features with different scales and different fine-grainedness and

[0009] (4) Use the maximum average fusion pooling operation to perform pooling on and to obtain the vectorized pedestrian representations corresponding to different features and And splice them as the final feature vector for pedestrian matching. During the training process, the cross-resolution triplet loss, cross-resolution center loss, and identity loss are used to constrain the distribution of pedestrian features and reduce the distance between the features of different resolution images of the same pedestrian.

[0010] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0011] 1. The present invention introduces a low-resolution feature reconstruction strategy. At the shallow layer of the backbone, the feature information and resolution information of the high-resolution image are used to simultaneously constrain the reconstruction process, so that the information loss in the finally obtained depth-invariant features is smaller;

[0012] 2. The present invention adds a pedestrian feature enhancement module to perform correlation modeling and fusion on the multi-level and different-scale features of pedestrians to obtain a more discriminative pedestrian representation;

[0013] 3. The present invention introduces a new pooling operation, which can make more full use of the information in the feature map. Brief Description of the Drawings

[0014] Figure 1 It is the network structure diagram of the cross-resolution pedestrian re-identification based on dual-stream input feature reconstruction of the present invention;

[0015] Figure 2 It is the lightweight residual decoder in the feature reconstruction of the present invention;

[0016] Figure 3 This is the feature enhancement module in the invariant feature learning of the present invention;

[0017] Figure 4 This is the maximum average fusion pooling module in the feature vectorization of the present invention; Detailed implementation manners

[0018] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to explain the present invention, but not to limit the present invention.

[0019] As Figure 1 shown, a cross-resolution pedestrian re-identification method based on dual-stream input feature reconstruction includes the following steps:

[0020] (1) Construct cross-resolution pedestrian image pairs. First, obtain a batch of HR images from the standard dataset according to the pedestrian identity label, and then randomly downsample the HR images with a scale factor of {2, 3, 4}. The obtained LR images and HR images are used as the image pairs for training;

[0021] (2) Input the pedestrian images (pairs) into the dual-stream reconstruction network. In this process, different input branches will be selected according to the resolution of the images. The low-resolution branch will adaptively reconstruct the features of the low-resolution images. The two branches respectively obtain and During the training process, the feature degradation loss and the reconstruction loss are combined with the lightweight residual decoder (RLDecoder) to learn the feature reconstruction ability;

[0022] (3) Input and into the weight-sharing siamese network structure. First, pass through the fourth and fifth residual blocks of the backbone to obtain the initially embedded resolution-invariant features Then input it into the feature enhancement module (MGLIF) to obtain the pedestrian depth-invariant features with different scales and different fine-grainedness and

[0023] (4) Use the maximum average fusion pooling operation to pool and to obtain the vectorized pedestrian representations corresponding to different features and And splice them to form a feature vector finally used for pedestrian matching. During the training process, cross-resolution triplet loss, cross-resolution center loss, and identity loss are used to constrain the distribution of pedestrian features and reduce the distance between the features of different-resolution images of the same pedestrian.

[0024] The detailed steps are as follows:

[0025] Step (1): Obtain cross-resolution image pairs. From a standard pedestrian re-identification dataset, select P pedestrian identities, with K images for each pedestrian, to obtain PK high-resolution (HR) images. For each image, randomly select a scaling factor from {2, 3, 4} to downsample it, obtaining PK low-resolution (LR) images. After this operation, each HR image corresponds to one HR image, facilitating subsequent feature reconstruction and identity representation learning.

[0026] Step (2): Input the image pairs into the dual-stream input and feature reconstruction network.

[0027] For the network structure, as shown in the "dual-stream input and feature reconstruction" part of Figure 1 , divide the backbone network into 5 residual blocks {R 1 , R 2 , R 3 , R 4 , R 5}. Define the output of the last layer of each residual block as {f 1 , f 2 , f 3 , f 4 , f 5}, where and d is the number of channels of the corresponding feature map. Under this setting, the feature extraction and reconstruction function of the dual-stream input branch of the present invention can be expressed as:

[0028]

[0029]

[0030] where is the CBAM module to enhance the feature extraction ability of the LR branch, making feature reconstruction more stable and reliable. Moreover, the weights of the residual blocks in the HR branch and the LR branch are independent and not shared. Given an HR image x h and a corresponding LR image x l , the outputs of the HR and LR input branches can be expressed as and To ensure that the LR branch has the ability of feature reconstruction and recovery, two losses are used to constrain the output features, namely the feature degradation loss and the reconstruction loss where Directly constrain f hr and f lr while indirectly constrain through the RL Decoder which can be expressed as:

[0031]

[0032] where N represents the number of HR and LR sample pairs in the current batch, i represents the i-th pair of samples, and stopgrad represents the "gradient truncation" operation. The output f of the LR input branch lr is indirectly calculated through the RL Decoder which is expressed as follows:

[0033]

[0034] where is the RL Decoder module. Through the above operations, the output f of the two-stream input branch can be made lr as close as possible to f hr so as to achieve the restoration of LR image information at the feature level.

[0035] Regarding the RL Decoder, a simple and lightweight decoder is more suitable for information restoration in cross-resolution person re-identification tasks. The present invention constructs a lightweight residual decoder (RL Decoder) to constrain the reconstruction of LR image features, and its overall structure is as Figure 2 shown. The input feature map f lr passes through a lightweight convolutional block and is added to the dimension-reduced f lr . Then, it is upsampled through bilinear interpolation, followed by non-linear activation. Finally, the feature map is upsampled again through pixel shuffling to restore the image channels, obtaining the reconstructed image.

[0036] Step (3): Obtain depth-invariant features using a weight-sharing Siamese network.

[0037] For the network structure, as shown in the "Invariant Feature Learning and Embedding" part of Figure 1 , this part is the depth-invariant feature extraction strategy proposed by the present invention. A weight-sharing Siamese network is used to extract pedestrian features at different resolutions, so that the distribution of the finally obtained pedestrian feature representations at different resolutions should be as similar as possible, having better matching performance. First, the 4th and 5th residual blocks of the backbone are used to obtain invariant features with similar distributions Then, a feature enhancement module is used to obtain depth-invariant features.

[0038] Regarding the feature enhancement module, as shown in Figure 3As shown, where Bottleneck (denoted as ) is the basic convolutional block of the backbone, and Reduction (denoted as ) compresses the feature map in the channel dimension. Bottleneck processes the features finally output by the backbone as high-level features, and the corresponding unprocessed ones as low-level features. The high-level features are cut into blocks horizontally as local-scale features, and the entire feature map is used as the global-scale feature. The low-level features are compressed in the channel dimension to obtain low-level global features:

[0039]

[0040] Where is the input feature, that is, the invariant feature is the obtained low-level global feature, C in and C glo are the number of channels of the input feature and the compressed feature respectively. The high-level features are processed differently to obtain high-level global features and local features. First, the high-level global feature f hl is similar to f ll , and is also obtained by compressing in the channel dimension:

[0041]

[0042] As for the local relationship feature f loc , first, the feature map is cut into blocks horizontally, and then concatenated along the channel dimension. A 1×1 convolutional layer is used to establish the correlation between the features of different parts of the pedestrian, as shown below:

[0043]

[0044] Where represents the "cutting and concatenating" operation. To establish its connection with the global feature of the pedestrian, the present invention uses the "stacking" operation to make its size consistent with the global feature, and then uses a 1×1 convolutional layer to fuse f ll , f hl and f loc to establish its relationship with the global feature:

[0045]

[0046] Where Cat[·] represents concatenation along the channel dimension, represents the stacking operation along the height direction of the feature map, is the result of the interaction and fusion of features at different levels and scales.

[0047] Step (4): Vectorize the obtained depth-invariant features for pedestrian matching.

[0048] Regarding the maximum average fusion pooling, as Figure 4 shown, the results of average pooling and maximum pooling are adaptively fused to compress and refine the feature map information. The present invention simultaneously uses adaptive average pooling (AAP) and adaptive maximum pooling (AMP) operations to respectively ensure the retention of information and the extraction of significant features, and finally uses a 1×1 convolutional layer to perform adaptive fusion on them to obtain the final feature vector Seq:

[0049]

[0050] where Seq t , t∈{ll, fus, loc, hl} respectively represent f ll , f fus , f loc , f hl corresponding feature vectors, n controls the output size of the adaptive pooling layer, and C out controls the dimension of the final feature vector. The feature vector finally used for pedestrian matching is obtained by fusing and splicing the feature vectors obtained from the corresponding resolution images, that is:

[0051]

[0052] Regarding the cross-resolution triplet loss Existing triplet losses only mine hard triplets at a single resolution level. However, for cross-resolution pedestrian matching, in addition to considering the differences between pedestrians within the resolution, special attention needs to be paid to the influence of the differences between the same pedestrian at different resolutions. In each training batch, P pedestrian categories are randomly selected, and K HR and LR images are selected for each pedestrian category, a total of 2PK images are selected, and its formula is as follows:

[0053]

[0054] where D(·) is the Euclidean distance, [z] + is equivalent to max(z, 0), t represents different feature vectors, the anchor is selected from the set of HR and LR features, is the positive sample with the same category as the anchor, is the negative sample with a different category from the anchor. The selection of all hard positive and negative samples takes into account the HR and LR feature spaces, rather than simply selecting samples in the feature space of a single resolution.

[0055] Regarding the cross-resolution center loss The feature centers established by the existing center loss are based on a single resolution level. In cross-resolution pedestrian matching, the feature centers should be established between the HR and LR feature spaces to better constrain the distribution of pedestrian features. This can simultaneously penalize pedestrian features at different resolution levels, making their distributions more compact, thereby better performing cross-resolution pedestrian matching, which is expressed as:

[0056]

[0057] where and are the pedestrian features finally obtained by the model, is the processing function of the entire network model for the input, y i is the label of the i-th image pair in the current batch, represents the center of the y i -th class of pedestrian image features established between the HR and LR feature spaces.

[0058] Regarding the identity loss the cross-entropy label smoothing loss is adopted, which is expressed as follows:

[0059]

[0060] where M represents the number of pedestrian identities in the training set, p(y pre ) represents the probability that the predicted label is y pre , and in addition, the definition of q(y pre ) is:

[0061]

[0062] where y is the true label of the input image, ε is a small constant (set to 0.1), which can prevent the re-identification model from overfitting on the training set, and the total training loss is:

[0063]

[0064] where λ is the balancing weight of the center loss, which is set to 0.001 in the present invention.

[0065] To verify the effectiveness of the method of the present invention, the present invention is verified on three datasets, namely CAVIAR, MLR-VIPeR, and MLR-Market-1501, which are commonly used in the field of cross-resolution pedestrian re-identification. Four deep learning-based cross-resolution pedestrian re-identification methods are selected as comparison methods, specifically:

[0066] Method 1: INTACT proposed by Cheng et al., reference "Cheng Z, Dong Q, Gong S, et al. Inter-Task Association Critic for Cross-Resolution Person Re-Identification[C]IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020: 2602-2612."

[0067] Method 2: PS-HRNet proposed by Zhang et al., reference "Zhang G, Ge Y, Dong Z, et al. Deep High-Resolution Representation Learning for Cross-Resolution Person Re-Identification[J]. IEEE Transactions on Image Processing, 2021, 30: 8913-8925."

[0068] Method 3: JBIM proposed by Zheng et al., reference "Zheng W S, Hong J, Jiao J, et al. Joint Bilateral-Resolution Identity Modeling for Cross-Resolution Person Re-Identification[J]. International Journal of Computer Vision, 2022, 130(1): 136-156."

[0069] Method 4: LRAR proposed by Wu et al., reference "Wu L Y, Liu L, Wang Y, et al. Learning Resolution-Adaptive Representations for Cross-Resolution Person Re-Identification[J]. IEEE Transactions on Image Processing, 2023, 32: 4800-4811."

[0070] As shown in Table 1, the performance of the method proposed in the present invention using Rank1, Rank5, and Rank10 as evaluation indicators on three datasets has significant advantages compared with the other four methods.

[0071] Table 1 Comparison of Rank1, Rank5, and Rank10 indicators with other methods

[0072]

[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A cross-resolution pedestrian re-identification method based on dual-stream input feature reconstruction, Characterized in that, It includes the following steps: (1) Construct cross-resolution pedestrian image pairs. First, obtain a batch of HR images from the standard dataset according to the pedestrian identity label, and then randomly downsample the HR images with the scale factors of {2, 3, 4}. The obtained LR images and HR images are used as the image pairs for training; (2) Input the pedestrian image (pair) into the "Dual-Stream Input and Feature Reconstruction" network. During this process, different input branches will be selected according to the resolution of the image. The low-resolution branch will adaptively reconstruct the features of the low-resolution image, and the two branches will respectively obtain and During the training process, the feature degradation loss and the reconstruction loss are combined with the lightweight residual decoder (RL Decoder) to learn the feature reconstruction ability; (3) Input and into the weight - shared siamese network structure. First, pass through the fourth and fifth residual blocks of the backbone to obtain the initially embedded resolution - invariant features . Then, input them into the feature enhancement module to obtain the pedestrian depth - invariant features with different scales and different fine - grainedness and .​​​​​​​​​ (4) Use the maximum average fusion pooling operation for and to perform pooling, obtaining the vectorized pedestrian representations corresponding to different features and and concatenating them as the final feature vector for pedestrian matching. And during the training process, cross-resolution triplet loss, cross-resolution center loss, and identity loss are used to constrain the distribution of pedestrian features and reduce the distance between the features of different-resolution images of the same pedestrian.

2. The cross-resolution pedestrian re-identification method based on dual-stream input feature reconstruction according to claim 1, Characterized in that, The "dual-stream input and feature reconstruction" network structure in step (2); The backbone network is divided into 5 residual blocks {R 1 , R 2 , R 3 , R 4 , R 5}. The output of the last layer of each residual block is defined as {f 1 , f 2 , f 3 , f 4 , f 5}, where and d is the number of channels of the corresponding feature map. In this setting, the feature extraction and reconstruction function of the dual-stream input branch of the present invention can be expressed as: Among them is the CBAM module, which enhances the feature extraction ability of the LR branch, making the feature reconstruction more stable and reliable, and the weights of the residual blocks in the HR branch and the LR branch are independent and not shared; given an HR image x h and a corresponding LR image x l , the outputs are and 3. The cross-resolution pedestrian re-identification method based on dual-stream input feature reconstruction according to claim 1, Characterized in that, The feature degradation loss in step (2) and the reconstruction loss To ensure that the LR branch has the ability to reconstruct and recover features, two losses are used to constrain the output features, namely the feature degradation loss and the reconstruction loss Among them directly constrains f hr and f lr while is indirectly constrained by RLDecoder which can be expressed as: where N represents the number of HR and LR sample pairs in the current batch, i represents the i-th pair of samples, stopgrad represents the "gradient truncation" operation, and the output f of the LR input branch lr is indirectly calculated through the RL Decoder is expressed as follows: Among them is the RL Decoder module.

4. Reconstruction loss according to claim 3 Characterized in that, The network structure of the RL Decoder in the above calculation process, with the input feature map f lr is added to the f that has gone through the lightweight convolution block and dimensionality reduction lr Then, it is upsampled through bilinear interpolation, followed by non-linear activation. Finally, pixel shuffling is used to upsample the feature map again and restore the image channels to obtain the reconstructed image.

5. The cross-resolution pedestrian re-identification method based on dual-stream input feature reconstruction according to claim 1, Characterized in that, In the feature enhancement module in step (3), its network structure consists of Bottleneck (denoted as ), Reduction (denoted as ), and a 1×1 convolutional layer. Bottleneck is the basic convolutional block of the backbone, and Reduction compresses the feature map along the channel dimension; The Bottleneck processes the input features as high-level features, and the corresponding unprocessed ones are used as low-level features. The high-level features are cut into blocks in the horizontal direction as features at the local scale, and the entire feature map is used as the feature at the global scale. The low-level features are compressed in the channel dimension to obtain low-level global features: Among them is the input feature is the obtained low-level global feature, C in and C glo are the number of channels of the input feature and the compressed feature channels respectively. The high-level features are processed differently to obtain high-level global features and local features. First, the high-level global feature f hl is similar to f ll and is also obtained by compressing the channel dimension: Regarding the local relationship feature f loc , first, the feature map is sliced along the horizontal direction and then concatenated along the channel dimension. A 1×1 convolutional layer is used to establish the correlation between the features of different parts of the pedestrian, which is expressed as follows: Among them represents the "cutting and splicing" operation, uses the "stacking" operation to make its size consistent with the global feature, and then uses a 1×1 convolutional layer to fuse f ll , f hl and f loc to establish the relationship between it and the global feature: where Cat[·] represents concatenation along the channel dimension, represents the stacking operation along the height direction of the feature map, is the result of the interactive fusion of features at different levels and scales.

6. The cross-resolution pedestrian re-identification method based on dual-stream input feature reconstruction according to claim 1, Characterized in that, In step (4), the maximum average fusion pooling uses both adaptive average pooling (AAP) and adaptive max pooling (AMP) operations to ensure the retention of information and the extraction of significant features respectively. Finally, a 1×1 convolutional layer is used to adaptively fuse them to obtain the final feature vector Seq: Among them, Seq t , t ∈ {ll, fus, loc, hl} respectively represent f ll , f fus , f loc , f hl The corresponding feature vectors, n controls the output size of the adaptive pooling layer, C out Controls the dimension of the final feature vector.

7. The cross-resolution pedestrian re-identification method based on dual-stream input feature reconstruction according to claim 1, Characterized in that, The cross-resolution triplet loss in step (4) Cross-resolution center loss Randomly select P pedestrian categories in each training batch, and for each pedestrian category, select K HR and LR images respectively, so a total of 2PK images are selected. The formula is as follows: where D(·) is the Euclidean distance, [z] + is equivalent to max(z, 0), t represents different feature vectors, and the anchor point is selected from the sets of HR and LR features is a positive sample with the same category as the anchor point is a negative sample with a different category from the anchor point. The selection of all difficult positive and negative samples takes into account the feature spaces of HR and LR, rather than simply selecting samples in the feature space of a single resolution; Regarding Cross-Resolution Center Loss The feature center of each class of pedestrians is established between the HR and LR feature spaces to better constrain the distribution of pedestrian features. At the same time, pedestrian features at different resolution levels are penalized to make their distributions more compact, so as to better perform cross-resolution pedestrian matching, which is expressed as: Among them and are the pedestrian features finally obtained by the model is the processing function of the entire network model for the input, y i is the label of the i-th image pair in the current batch represents the center of the y i -th class of pedestrian image features established between the HR and LR feature spaces

Citation Information

Patent Citations

  • Picture pedestrian re-identification system and method based on resolution irrelevant characteristics

    CN110765864A

  • Deep learning network and segmentation method for text picture character segmentation

    CN110895695A

  • Face deep counterfeiting detection method based on traditional features and neural network

    CN114202782A

  • Image super-resolution reconstruction method based on mixed attention and double-layer supervision

    CN114897694A

  • Blind super-resolution reconstruction method based on kernel uncertainty learning and degradation embedding

    CN116843553A