Unsupervised pedestrian retrieval method based on component-guided fusion transformation network

By designing components to guide the fusion transformation network, we achieve two-way interaction and collaborative optimization of global and local features, solving the problem of insufficient local feature consistency in the existing Transformer model in pedestrian retrieval, and improving the accuracy and robustness of pedestrian retrieval.

CN120723938AActive Publication Date: 2025-09-30SHANDONG JIANZHU UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511172541.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-09-30
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

The existing Transformer model is limited to global feature interaction in pedestrian retrieval and lacks explicit modeling of local region structured relationships. Existing technical means: through the design of an unsupervised method of component-guided fusion transformation network, the bidirectional interaction of global and local features is achieved by designing a component-guided fusion transformation network, and the development of component-guided contrast loss to strengthen local discriminability through the dual constraints of identity attributes and component semantics. And through component-guided regularization, the consistency of multi-granular features is ensured from the two dimensions of feature space and cluster distribution. This solves the problem of lack of local feature consistency and fine-grained discrimination ability in the global feature optimization process in existing methods.

Method used

A component-guided fusion transformation network is designed, including a global encoder and a component encoder. Through component-guided contrastive learning and regularization modules, a dual-stream dynamic memory library is constructed. The component-guided contrastive loss and regularization module are used to optimize feature expression, achieving coordinated optimization of global and local features.

Benefits of technology

It significantly improves the representation ability of pedestrian features and the generalization performance of the model, and improves the accuracy and robustness of pedestrian retrieval in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723938A_ABST
    Figure CN120723938A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and particularly discloses an unsupervised pedestrian retrieval method based on a component-guided fusion transformation network, which comprises the following steps of: extracting global and component characteristics of pedestrians through the component-guided fusion transformation network; generating a component-global fusion feature for component-guided comparative learning by adopting a component-guided fusion strategy; averagely aggregating the component features to obtain component aggregated features; a corresponding clustering center is obtained through a clustering algorithm based on the global features and the component aggregation features, and a double-flow dynamic memory bank is constructed; constructing a component guide regularization module, and constraining the consistency between the global feature and the component aggregation feature by a dynamic weighting coefficient; and finally, iteratively optimizing the parameters of the component guide fusion transformation network and the component guide regularization module through a total loss calculation module. According to the method, feature distribution is optimized through unsupervised learning, dependence on labeled data is reduced, and pedestrian feature characterization capacity and model generalization performance are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to an unsupervised pedestrian retrieval method based on a component-guided fusion transformation network. Background Art

[0002] Person retrieval, a key technology in computer vision, aims to accurately match pedestrian identities across multiple camera scenarios. It holds significant application value in scenarios such as intelligent security and video surveillance. With the advancement of deep learning, supervised learning-based person retrieval methods have made significant progress by training deep models with large-scale labeled data. However, these methods face bottlenecks in practical applications, such as high labeling costs and poor cross-domain generalization. In recent years, researchers have attempted to incorporate the Transformer model into the person retrieval task, leveraging its global attention mechanism to model long-range dependencies. For example, He et al. proposed TransReID, which learns discriminative representations of pedestrian features using a pure Transformer architecture. Li et al.'s Diversified Compact Transformer (DC-Former) improves feature representation by constructing multiple embedding subspaces. Additionally, some methods embed Transformer modules within convolutional neural networks, fusing local and global features to enhance model performance. While these supervised learning methods perform well on specific datasets, their heavy reliance on labeled data leads to high deployment costs. Furthermore, the models are susceptible to data distribution discrepancies in cross-domain scenarios, significantly reducing generalization performance.

[0003] To reduce reliance on labeled data, unsupervised person retrieval methods have become a research hotspot, primarily categorized into unsupervised domain adaptation (UDA) and pure unsupervised learning (USL). Unsupervised domain adaptation methods optimize a model by pre-training it in the source domain and transferring it to the target domain. Typical approaches include style transfer techniques based on generative adversarial networks (GANs) and dynamic optimization strategies based on pseudo-labels. For example, some studies utilize generative adversarial networks (GANs) to generate person images in the target domain style to narrow inter-domain discrepancies, while others generate pseudo-labels through clustering to guide target domain model training. However, these methods require that the data distributions of the source and target domains be highly correlated. In practical cross-scenario applications, inter-domain discrepancies often lead to a sharp decline in model performance. Pure unsupervised learning methods rely entirely on unlabeled data in the target domain, improving feature discriminability through iterative optimization of pseudo-labels and dynamic updates of a memory bank. For example, the SpCL framework uses self-paced contrastive learning to distinguish source domain categories from target domain clusters, while the HCM model combines identity-level and image-level contrastive learning to mine information from difficult samples. Recently, researchers have attempted to incorporate the Transformer's global interaction mechanism into unsupervised frameworks, for example, by aggregating global features to enhance pedestrian representation. However, existing methods often focus on learning global features and lack fine-grained modeling of local pedestrian regions (such as body parts and clothing details). This results in insufficient feature discrimination in complex scenarios such as occlusion and posture changes.

[0004] Meanwhile, contrastive learning, a core technology in unsupervised learning, optimizes the spatial distribution of features by constructing pairs of positive and negative examples. Existing methods, such as cluster contrastive learning (C-Contrast), improve discriminability by mining difficult examples, while camera contrastive learning (CCL) utilizes camera information to construct proxy tasks and optimize feature representation. Furthermore, the AICES scheme enhances fine-grained knowledge learning by aggregating instance relationships perceived by the camera. However, existing contrastive learning methods have significant limitations: First, they primarily rely on global feature comparisons and neglect modeling the correlations between local regions of pedestrians (such as the head, torso, and stripes), resulting in loss of local semantic information. Second, the feature optimization process lacks explicit constraints on the consistency of local features, making cross-domain matching susceptible to background interference. Third, memory bank update strategies are often based on the global feature mean, failing to effectively capture the spatiotemporal dynamic correlations of local features, limiting fine-grained discriminative capabilities. For example, in pedestrian images, the local features of shirt stripes and trouser patterns may be highly discriminative, but existing methods struggle to specifically model and optimize such local relationships.

[0005] The shortcomings of existing technologies can be attributed to the following core issues: the application of Transformer models in pedestrian retrieval is mostly limited to global feature interactions, and lacks explicit modeling of structured relationships in local regions, resulting in insufficient utilization of detailed information; the unsupervised contrastive learning framework relies too much on global feature similarity measurement, fails to construct contrast relationships between local regions, and is unable to cope with posture and perspective changes in complex scenarios; the lack of semantic consistency constraints between global and local features during feature learning makes it difficult to coordinate the optimization of multi-granularity features, affecting the robustness of the model. Summary of the Invention

[0006] The purpose of the present invention is to provide an unsupervised pedestrian retrieval method based on a component-guided fusion transformation network. By designing a component-guided fusion transformation network, the two-way interaction between global and local features is realized. The component-guided contrast loss is developed to enhance local discriminability through the dual constraints of identity attributes and component semantics. The consistency of multi-granularity features is ensured from the two dimensions of feature space and cluster distribution through component-guided regularization, thereby significantly improving the representation ability of pedestrian features and the generalization performance of the model in the unsupervised pedestrian retrieval task.

[0007] To achieve the above objectives, the present invention provides an unsupervised pedestrian retrieval method based on a component-guided fusion transformation network, comprising the following steps: S1. Build a component-guided fusion transformation network based on the pre-trained visual Transformer model to extract the global features and component features of pedestrians; S2. Use component-guided fusion strategy to generate component-global fusion features for component-guided contrastive learning; perform average aggregation processing on the pedestrian component features to obtain component aggregation features; S3. Based on global features and component aggregation features, a clustering algorithm is used to obtain the corresponding cluster centers and build a dual-stream dynamic memory library; S4. Construct a component-guided regularization module to constrain the consistency between global features and component aggregate features by introducing dynamic weighting coefficients; S5. Construct a total loss calculation module to calculate the total loss value, and iteratively optimize the parameters of the component-guided fusion transformation network and the component-guided regularization module based on the total loss value.

[0008] Preferably, in S1, the component-guided fusion transformation network constructed includes Transformation layers, a global encoder and a component encoder, based on the pre-trained ImageNet dataset containing The ViT model is a visual Transformer model with a transformation layer. The parameters of the component-guided fusion transformation network are initialized. The initialization process is as follows: Before ViT model The parameters of the corresponding components of the layer guide the front of the fusion transformation network The parameters of the layer, ViT model The parameters of the layers correspond to the parameters of the global encoder and component encoder in the component-guided fusion transformer network.

[0009] Preferably, in S1, the pedestrian image is scaled and normalized to obtain a pre-processed pedestrian image. ; The preprocessed pedestrian image Input the initialized components to guide the fusion transformation network to obtain the global features of the pedestrian and assembly features .

[0010] Preferably, S2 comprises the following steps: S21. Use component-guided fusion strategy to fuse the global features and component features of pedestrians to generate component-global fusion features , the fusion of global features and component features is as follows: ; in, is the coupling factor, Indicates the The first pedestrian image Component-global fusion features, Indicates the The first pedestrian image component features, Indicates the number of components; Component-guided contrastive loss in component-guided contrastive learning for: ; in, 、 and Represent the cosine similarity between positive sample components, negative sample components, and weak positive sample components respectively; is the smoothing factor, defined as: ; in, , ,and , Indicates taking the absolute value; S22, perform average aggregation processing on the pedestrian's component features to obtain component aggregation features , whose expression is: ; in, Indicates the The first pedestrian image component features, Indicates the number of parts.

[0011] Preferably, in S21, for the component-global fusion feature corresponding to each component in the pedestrian image, the positive sample component is defined as the same component of the pedestrian image with the same identity, the negative sample component is defined as the component of the pedestrian image with different identities, and the weak positive sample component is defined as other components of the same pedestrian image and non-corresponding components of the pedestrian image with the same identity.

[0012] Preferably, S3 includes: The global features and component aggregation features are clustered using a clustering algorithm, and the corresponding cluster centers are obtained to construct a dual-stream dynamic memory library. The features stored in the dual-stream dynamic memory library are randomly initialized. The momentum update of the features stored in the two-stream dynamic memory bank is performed as follows: ; in, is the momentum coefficient, and Represent global features and component aggregation features respectively, and The first Cluster centers, and Represents the sample sets under the same identity corresponding to the global features and component aggregation features, Indicates the number of samples in the set.

[0013] Preferably, S3 further includes: Unsupervised feature learning using batch data and cluster centers in a two-stream dynamic memory library, and global loss in unsupervised feature learning and component loss The formal expression is as follows: ; ; in, and Represents global features Positive and negative pairwise similarities between cluster centers in the two-stream dynamic memory library, and Represents component aggregation features Positive and negative pairwise similarities between cluster centers in the two-stream dynamic memory library, is the temperature coefficient, is the size of the batch data.

[0014] Preferably, in S4, the global features and component aggregation features are first mapped to the same feature space using the fully connected layer; then the component guided regularization loss is designed. Constrain the consistency between global features and component aggregate features in the following form: ; ; in, represents the fully connected layer, express norm, is the dynamic weighting coefficient, represents the number of samples in the set, is the adjustment factor, and Represents component aggregation features With global features The corresponding clusters.

[0015] Preferably, in S5, the contrast loss is guided by the calculation component , global loss , component loss , and component-guided regularization loss The weighted sum of : ; in, is the weight factor; Based on the total loss value, the parameters of the component-guided fusion transformation network and the component-guided regularization module are iteratively updated through gradient backpropagation in the following form: ; in, represents the learning rate, are the parameters of the current round in the optimization, are the parameters of the previous round.

[0016] Therefore, the present invention adopts the above-mentioned unsupervised pedestrian retrieval method based on component-guided fusion transformation network, and the beneficial effects are as follows: The present invention realizes the bidirectional interaction between global and local features by designing a component-guided fusion transformation network, develops a component-guided contrast loss to enhance local discriminability through the dual constraints of identity attributes and component semantics, and ensures the consistency of multi-granularity features from the two dimensions of feature space and cluster distribution through component-guided regularization, thereby significantly improving the representation ability of pedestrian features and the generalization performance of the model in unsupervised pedestrian retrieval tasks.

[0017] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is an overall flow chart of an embodiment of an unsupervised pedestrian retrieval method based on a component-guided fusion transformation network of the present invention. DETAILED DESCRIPTION

[0019] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0020] Unless otherwise defined, technical or scientific terms used in the present invention shall have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.

[0021] like Figure 1 As shown, an unsupervised pedestrian retrieval method based on component-guided fusion transformation network includes the following steps: S1. Based on the pre-trained visual Transformer model, a component-guided fusion transformation network is constructed to extract the global features and component features of pedestrians.

[0022] In this embodiment, the constructed component guided fusion transformation network includes Transformation layers, a global encoder and a component encoder, based on the pre-trained ImageNet dataset containing The ViT model, a visual Transformer model with a transformation layer, initializes the parameters of the component-guided fusion transformation network. The initialization process is as follows: Before ViT model The parameters of the corresponding components of the layer guide the front of the fusion transformation network The parameters of the layer, ViT model The parameters of the layers correspond to the parameters of the global encoder and component encoder in the component-guided fusion transformation network. In this embodiment, .

[0023] Perform scaling and normalization operations on the pedestrian image to obtain the preprocessed pedestrian image In this embodiment, Indicates the Pedestrian images are trained using the public MSMT17 dataset. Pedestrian images are scaled to a predefined Size, and normalize all pixel values ​​of the pedestrian image, that is, reduce them to between 0 and 1.

[0024] The preprocessed pedestrian image Input the initialized components to guide the fusion transformation network to obtain the global features of the pedestrian and assembly features In this embodiment, the outputs of the global encoder and the component encoder correspond to the global features and assembly features ,Pick Indicates the number of parts.

[0025] S2. A component-guided fusion strategy is used to generate component-global fusion features for component-guided contrastive learning. At the same time, the pedestrian component features are averaged and aggregated to obtain component aggregate features, including the following steps: S21. Generate component-global fusion features by fusing the global features and component features of pedestrians , and the component-global fusion features are used as the input of component-guided contrastive learning. The fusion method of global features and component features is as follows: ; in, is the coupling factor, Indicates the The first pedestrian image Component-global fusion features, in this embodiment, .

[0026] The component-guided contrastive loss in component-guided contrastive learning Designed to: ; in, 、 and They represent the cosine similarity between positive sample components, negative sample components, and weak positive sample components respectively; for the component-global fusion feature corresponding to each component in the pedestrian image, the positive sample component is defined as the same component in the pedestrian image with the same identity, the negative sample component is defined as the component in the pedestrian image with different identities, and the weak positive sample component is defined as other components in the same pedestrian image and non-corresponding components in the pedestrian image with the same identity. is the smoothing factor, defined as: ; in, , ,and , Indicates taking the absolute value; S22, perform average aggregation processing on the pedestrian's component features to obtain component aggregation features , whose expression is: ; in, Indicates the The first pedestrian image component features, Indicates the number of parts.

[0027] S3. Based on global features and component aggregation features, a clustering algorithm is used to obtain the corresponding cluster centers and build a dual-stream dynamic memory library, including: A clustering algorithm is used to cluster the global features and component aggregation features respectively, and the corresponding cluster centers are obtained to build a dual-stream dynamic memory library. The features stored in the dual-stream dynamic memory library are randomly initialized. In this embodiment, the clustering algorithm used is the DBSCAN algorithm, and other clustering algorithms such as K-means can also be used.

[0028] During the iterative training process, the features stored in the dual-stream dynamic memory library are updated with momentum in the following way: ; in, is the momentum coefficient. In this embodiment, ; and Represent global features and component aggregation features respectively, and The first Cluster centers, and Represents the sample sets under the same identity corresponding to the global features and component aggregation features, Indicates the number of samples in the set.

[0029] Next, we use batch data and the cluster centers in the two-stream dynamic memory library to perform unsupervised feature learning. The global loss in unsupervised feature learning is and component loss The formal expression is as follows: ; ; in, and Represents global features Positive and negative pairwise similarities between cluster centers in the two-stream dynamic memory library, and Represents component aggregation features Positive and negative pairwise similarities between cluster centers in the two-stream dynamic memory library, is the temperature coefficient, is the size of the batch data; in this embodiment, , .

[0030] S4. Build a component-guided regularization module to constrain the consistency between global features and component aggregate features by introducing dynamic weighting coefficients. The specific process is as follows: First, the global features and component aggregation features are mapped to the same feature space using a fully connected layer; then the component guided regularization loss is designed. Constrain the consistency between global features and component aggregate features in the following form: ; ; in, represents the fully connected layer, express norm, is the dynamic weighting coefficient, represents the number of samples in the set, is the adjustment factor. In this embodiment, . and Represents component aggregation features With global features The corresponding clusters.

[0031] S5. Construct a total loss calculation module to calculate the total loss value, and iteratively optimize the parameters of the component-guided fusion transformation network and the component-guided regularization module based on the total loss value. Specifically: By calculating the component-guided contrastive loss , global loss , component loss , and component-guided regularization loss The weighted sum of : ; in, is the weight factor; Based on the total loss value, the parameters of the component-guided fusion transformation network and the component-guided regularization module are iteratively updated through gradient backpropagation in the following form: ; in, represents the learning rate, are the parameters of the current round in the optimization, are the parameters of the previous round.

[0032] In a preferred embodiment of the present invention, , SGD optimizer is used as the training optimizer of the model, and the initial learning rate It is set to 0.01, and the cosine decay strategy is used to adjust the learning rate during iteration. The total number of training iterations is 80.

[0033] It is worth noting that in the preferred embodiment of the present invention, the network training method and the parameter configuration of the learning rate optimization strategy are non-restrictive and preferred choices. Those skilled in the art can select the model training method and parameter configuration based on various indicators such as classification accuracy and efficiency.

[0034] The pedestrian retrieval model obtained by training in the preferred embodiment of the present invention is tested on the test set of the MSMT17 dataset, and the results are shown in Table 1.

[0035] Table 1 Comparison of model results on the MSMT17 dataset ;

[0036] In Table 1, the baseline method uses the ViT model for global feature extraction, and the component-guided method uses the component-guided fusion transformation network to extract global features and component features.

[0037] As shown in Table 1, the average precision and ranking accuracy of the pedestrian retrieval model in this embodiment reached 46.3% and 71.5%, respectively, compared to the baseline method's 39.5% and 66.6%, an improvement of 6.8% and 4.9%, respectively. Furthermore, the introduction of component-guided contrastive loss and component-guided regularization enhances feature discriminability, further improving retrieval performance.

[0038] Therefore, the present invention adopts the above-mentioned unsupervised pedestrian retrieval method based on component-guided fusion transformation network, realizes the two-way interaction of global and local features by designing component-guided fusion transformation network, develops component-guided contrast loss to enhance local discriminability through the dual constraints of identity attributes and component semantics, and ensures the consistency of multi-granularity features from the two dimensions of feature space and cluster distribution through component-guided regularization, thereby significantly improving the representation ability of pedestrian features and the generalization performance of the model in the unsupervised pedestrian retrieval task.

[0039] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. An unsupervised pedestrian retrieval method based on component-guided fusion transformation network, characterized in that: The following steps are involved: S1. Build a component-guided fusion transformation network based on the pre-trained visual Transformer model to extract the global features and component features of pedestrians; S2. Use component-guided fusion strategy to generate component-global fusion features for component-guided contrastive learning; perform average aggregation processing on the pedestrian component features to obtain component aggregation features; S3. Based on global features and component aggregation features, a clustering algorithm is used to obtain the corresponding cluster centers and build a dual-stream dynamic memory library; S4. Construct a component-guided regularization module to constrain the consistency between global features and component aggregate features by introducing dynamic weighting coefficients; S5. Construct a total loss calculation module to calculate the total loss value, and iteratively optimize the parameters of the component-guided fusion transformation network and the component-guided regularization module based on the total loss value.

2. The unsupervised pedestrian retrieval method based on component-guided fusion transformation network according to claim 1 is characterized in that: In S1, the constructed component-guided fusion transformation network includes Transformation layers, a global encoder and a component encoder, based on the pre-trained ImageNet dataset containing The ViT model is a visual Transformer model with a transformation layer. The parameters of the component-guided fusion transformation network are initialized. The initialization process is as follows: Before ViT model The parameters of the corresponding components of the layer guide the front of the fusion transformation network The parameters of the layer, ViT model The parameters of the layers correspond to the parameters of the global encoder and component encoder in the component-guided fusion transformer network.

3. The unsupervised pedestrian retrieval method based on component-guided fusion transformation network according to claim 2 is characterized in that: In S1, the pedestrian image is scaled and normalized to obtain the preprocessed pedestrian image. ; The preprocessed pedestrian image Input the initialized components to guide the fusion transformation network to obtain the global features of the pedestrian and assembly features .

4. The unsupervised pedestrian retrieval method based on component-guided fusion transformation network according to claim 3 is characterized in that: S2 includes the following steps: S21. Use component-guided fusion strategy to fuse the global features and component features of pedestrians to generate component-global fusion features , the fusion of global features and component features is as follows: ; in, is the coupling factor, Indicates the The first pedestrian image Component-global fusion features, Indicates the The first pedestrian image component features, Indicates the number of components; Component-guided contrastive loss in component-guided contrastive learning for: ; in, 、 and Represent the cosine similarity between positive sample components, negative sample components, and weak positive sample components respectively; is the smoothing factor, defined as: ; in, , ,and , Indicates taking the absolute value; S22, perform average aggregation processing on the pedestrian's component features to obtain component aggregation features , whose expression is: ; in, Indicates the The first pedestrian image component features, Indicates the number of parts.

5. The unsupervised pedestrian retrieval method based on component-guided fusion transformation network according to claim 4 is characterized in that: In S21, for the component-global fusion feature corresponding to each component in the pedestrian image, the positive sample component is defined as the same component of the pedestrian image with the same identity, the negative sample component is defined as the component of the pedestrian image with different identities, and the weak positive sample component is defined as other components of the same pedestrian image and non-corresponding components of the pedestrian image with the same identity.

6. The unsupervised pedestrian retrieval method based on component-guided fusion transformation network according to claim 5 is characterized in that S3 include: The global features and component aggregation features are clustered using a clustering algorithm, and the corresponding cluster centers are obtained to construct a dual-stream dynamic memory library. The features stored in the dual-stream dynamic memory library are randomly initialized. The momentum update of the features stored in the two-stream dynamic memory bank is performed as follows: ; in, is the momentum coefficient, and Represent global features and component aggregation features respectively, and The first cluster centers, and Represents the sample sets under the same identity corresponding to the global features and component aggregation features, Indicates the number of samples in the set.

7. The unsupervised pedestrian retrieval method based on component-guided fusion transformation network according to claim 6 is characterized in that: S3 also includes: Unsupervised feature learning using batch data and cluster centers in a two-stream dynamic memory library, and global loss in unsupervised feature learning and component loss The formal expression is as follows: ; ; in, and Represents global features Positive and negative pairwise similarities between cluster centers in the two-stream dynamic memory library, and Represents component aggregation features Positive and negative pairwise similarities between cluster centers in the two-stream dynamic memory library, is the temperature coefficient, is the size of the batch data.

8. The unsupervised pedestrian retrieval method based on component-guided fusion transformation network according to claim 7 is characterized in that: In S4, the global features and component aggregation features are first mapped to the same feature space using the fully connected layer; then the component guided regularization loss is designed. Constrain the consistency between global features and component aggregate features in the following form: ; ; in, represents the fully connected layer, express norm, is the dynamic weighting coefficient, represents the number of samples in the set, is the adjustment factor, and Represents component aggregation features With global features The corresponding clusters.

9. The unsupervised pedestrian retrieval method based on component-guided fusion transformation network according to claim 8 is characterized in that: In S5, the contrast loss is guided by the calculation component , global loss , component loss , and component-guided regularization loss The weighted sum of : ; in, is the weight factor; Based on the total loss value, the parameters of the component-guided fusion transformation network and the component-guided regularization module are iteratively updated through gradient backpropagation in the following form: ; in, represents the learning rate, are the parameters of the current round in the optimization, are the parameters of the previous round.

Citation Information

Patent Citations

  • Pedestrian re-identification method based on global-local feature dynamic alignment

    CN113408492A

  • Unsupervised target re-identification method and system based on perceptual assisted learning Transform model

    CN116403015A

  • Unsupervised domain adaptive pedestrian re-identification method and system

    CN116884052A

  • Transform and fusion clustering-based contrast learning unsupervised pedestrian re-identification method

    CN118692114A

  • Person re-identification method and apparatus, electronic device, and storage medium

    WO2021203801A1