Methods, devices, servers, and media for identifying pedestrians changing clothes

By using a complementary expert network model to process pedestrian re-identification data under clothing change scenarios, the accuracy and deployment cost issues of pedestrian identification methods under clothing change scenarios are solved, and efficient identification in complex scenarios is achieved.

CN122493486APending Publication Date: 2026-07-31TIANJIN POLYTECHNIC UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN POLYTECHNIC UNIV
Filing Date
2026-04-07
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing methods for recognizing pedestrians changing clothes are insufficient in terms of accuracy and deployment cost in scenarios involving clothing changes, making it difficult to effectively solve the problem of clothing changes in pedestrian re-identification.

Method used

Three complementary expert network models are used to process the basic features. The first expert network model learns identity features unrelated to clothing, the second expert network model learns clothing detail features, and the third expert network model integrates identity and detail features to generate comprehensive features. The re-identification result is then generated through a weighted network.

Benefits of technology

It significantly improves robustness and accuracy in complex real-world scenarios, avoids the need for additional noise interference and additional labeled data or sensors, and ensures the comprehensiveness and complementarity of feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493486A_ABST
    Figure CN122493486A_ABST
Patent Text Reader

Abstract

This invention discloses a method, device, server, and medium for recognizing pedestrians changing clothes, belonging to the field of visual recognition. The method includes: inputting the re-identification data of the pedestrian changing clothes to be identified into a feature extraction module to obtain basic features; directly or indirectly processing the basic features using three complementary expert network models; inputting the basic features into a weighted network to obtain the weights of the output features of the first, second, and third expert network models; generating comprehensive features based on the weights of each network; and generating a re-identification result using the comprehensive features. This significantly improves robustness and accuracy in complex real-world scenarios. Furthermore, it eliminates the need for additional noise interference and additional labeled data or sensors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of person recognition methods, and more particularly to a method, device, server, and medium for recognizing pedestrians changing clothes. Background Technology

[0002] Pedestrian re-identification is one of the core tasks in the field of computer vision. Its main goal is to quickly and accurately retrieve images of pedestrians with the same identity from a large-scale image database, given a pedestrian query image, under multiple non-overlapping camera viewpoints. This technology plays a crucial role in fields such as intelligent security, smart cities, and public safety, and is widely used in scenarios such as suspect tracking, missing persons retrieval, and intelligent passenger flow analysis.

[0003] Thanks to the development of deep learning technology, supervised person re-identification methods have achieved significant performance improvements on public benchmark datasets. However, most traditional methods are based on the assumption of pedestrian appearance consistency, which means that pedestrians with the same identity are assumed to wear the same or similar clothing under different cameras. In real-world applications, pedestrians may change their clothing due to seasonal changes, daily changes, or special job requirements, resulting in significant changes in appearance features and causing a sharp decline in model performance.

[0004] Currently, existing methods for identifying clothing changes in pedestrians who have changed clothes mainly employ the following approaches: Methods based on clothing-invariant feature learning attempt to separate identity semantic features from irrelevant features such as clothing and accessories in pedestrian images, using only identity-related features for matching. However, such methods struggle to completely eliminate clothing interference in complex scenes and may lose valuable clothing identification information in scenarios where clothing matches. Data augmentation and synthesis-based methods: To simulate real-world clothing changes, generative adversarial networks (GANs) are used to generate images of pedestrians changing clothes, or enhancement strategies such as color jittering and texture transformation are used to expand the training data to improve the model's generalization ability to appearance changes. However, the realism and diversity of the generated images are limited, making it difficult to cover the complex textures and lighting changes in real-world scenes, and they are also prone to introducing noise interference.

[0005] Multimodal fusion-based methods incorporate auxiliary information such as human contours, skeletal key points, and gait sequences to compensate for feature loss caused by clothing variations. However, these methods rely on additional labeled data or sensors (such as depth cameras and infrared devices), resulting in high costs in practical deployments and difficulty in widespread adoption in multi-camera heterogeneous systems.

[0006] In summary, existing methods for identifying pedestrians changing clothes are insufficient in terms of accuracy or deployment cost, and they fail to meet the technical requirements for re-identification accuracy and implementation difficulty. Summary of the Invention

[0007] This invention provides a method, apparatus, server, and storage medium for recognizing pedestrians changing clothes, in order to solve the technical problems of accuracy and performance in existing methods for recognizing pedestrians changing clothes.

[0008] In a first aspect, embodiments of the present invention provide a method for identifying pedestrians changing clothes, including: Input the re-identification data of the pedestrian changing clothes to be identified into the feature extraction module to obtain the basic features; The basic features are processed directly or indirectly using three complementary expert network models, which include: The first expert network model outputs clothing-independent human identity features. The first expert network model is used to learn clothing-independent identity features based on the basic features. The second expert network model outputs clothing detail features, and the second expert network model is used to learn and enhance clothing detail features based on the basic features. The third expert network model is used to receive human identity features and clothing detail features unrelated to clothing, and integrate them along the feature dimension as input to learn the collaborative representation of multi-view features. The basic features are input into the weight network to obtain the weights of the output features of the first expert network model, the second expert network model, and the third expert network. A comprehensive feature is generated based on the weights of each network, and the re-identification result is generated using the comprehensive feature.

[0009] Secondly, embodiments of the present invention also provide a pedestrian re-identification device for changing clothes, comprising: The basic feature acquisition module is used to input the re-identification data of the pedestrian changing clothes to be identified into the feature extraction module to obtain basic features; The processing module is used to directly or indirectly process the basic features using three complementary expert network models, wherein the three expert network models include: The first expert network model outputs clothing-independent human identity features. The first expert network model is used to learn clothing-independent identity features based on the basic features. The second expert network model outputs clothing detail features, and the second expert network model is used to learn and enhance clothing detail features based on the basic features. The third expert network model is used to receive human identity features and clothing detail features unrelated to clothing, and integrate them along the feature dimension as input to learn the collaborative representation of multi-view features. The generation module is used to input basic features into the weight network to obtain the weights of the output features of the first expert network model, the second expert network model, and the third expert network, and to generate comprehensive features based on the weights of each network, and to generate re-identification results using the comprehensive features.

[0010] Thirdly, embodiments of the present invention also provide a server, comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the pedestrian recognition method for changing clothes as provided in the above embodiments.

[0011] Fourthly, embodiments of the present invention also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the pedestrian identification method for changing clothes as provided in the above embodiments.

[0012] The present invention provides a method, apparatus, server, and storage medium for identifying pedestrians changing clothes. The method involves inputting the re-identification data of the pedestrian changing clothes to be identified into a feature extraction module to obtain basic features. Three complementary expert network models are used to directly or indirectly process these basic features. These three expert network models include: a first expert network model, which outputs clothing-independent human identity features and learns clothing-independent identity features based on the basic features; a second expert network model, which outputs clothing detail features and learns enhanced clothing detail features based on the basic features; and a third expert network model, which receives clothing-independent human identity features and clothing detail features, integrates them along the feature dimension as input, and learns a collaborative expression of multi-view features. The basic features are then input into a weighted network to obtain the weights of the output features from the first, second, and third expert network models. A comprehensive feature is generated based on the weights of each network, and the comprehensive feature is used to generate a re-identification result. By utilizing multiple complementary network models, persistent features unrelated to clothing and fine-grained features related to clothing are extracted separately, and then integrated along the feature dimension. This ensures the comprehensiveness and complementarity of feature extraction, overcoming the difficulty of a single feature extractor to handle multi-perspective information. Furthermore, the weighted network can learn which features are more important in different scenarios, effectively resolving the core contradiction in pedestrian re-identification after clothing changes: needing to forget the interference of clothing while utilizing clothing details for identification. This significantly improves robustness and accuracy in complex real-world scenarios. Moreover, it requires no additional noise interference or additional labeled data or sensors. Attached Figure Description

[0013] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart illustrating the pedestrian identification method for changing clothes provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart illustrating the pedestrian identification method for changing clothes provided in Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of the pedestrian re-identification device for changing clothes provided in Embodiment 3 of the present invention; Figure 4 This is a schematic diagram of the server structure provided in Embodiment 4 of the present invention. Detailed Implementation

[0014] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0015] Example 1 Figure 1 This is a flowchart illustrating the pedestrian recognition method for changing clothes provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where pedestrians changing clothes need to be re-identified in different scenarios. The method can be executed by a pedestrian re-identification device for changing clothes, and specifically includes the following steps: Step 110: Input the re-identification data of the pedestrian changing clothes to be identified into the feature extraction module to obtain basic features.

[0016] In this embodiment, the re-identification data of the person changing clothes to be identified can be a collected image. The collected image is input into the feature extraction module. For example, the feature extraction module can use a network model based on a residual network to extract basic features. For example, the basic visual features of the input image are extracted using a network model based on a residual network to generate preliminary basic features. In this way, it is possible to avoid each expert network extracting features independently in the later stage, thereby reducing redundant calculations.

[0017] Step 120: Use three complementary expert network models to process the basic features directly or indirectly.

[0018] Based on the preliminary basic features obtained from the above steps, they are processed to obtain features that facilitate the re-identification of pedestrians changing clothes. Optionally, three expert network models with different and complementary functions can be used to process the basic features directly or indirectly to obtain refined features suitable for the corresponding functions.

[0019] For example, the three expert network models include: a first expert network model that outputs clothing-independent human identity features, and the first expert network model is used to learn clothing-independent identity features based on the basic features; a second expert network model that outputs clothing detail features, and the second expert network model is used to learn enhanced clothing detail features based on the basic features; and a third expert network model that receives clothing-independent human identity features and clothing detail features, integrates them along the feature dimension as input, and learns the collaborative expression of multi-view features.

[0020] For example, the first expert network model can be composed of lightweight convolutional modules containing bottleneck structures to extract high-level semantics while preserving spatial structure information. Basic features are input into the first expert network model for transformation, outputting a robust feature representation independent of clothing. Its calculation can be expressed as: ,in, For robust feature expert networks, The basic features extracted for the backbone network are then used. Adversarial training is then performed, using gradient inversion to force the first-expert network model to optimize in a direction that increases the error of the clothing classifier during parameter updates, thereby suppressing clothing information in the features. This enables the learning of identity features unrelated to clothing. The final output is a highly robust feature that, due to the adversarial training mechanism, actively filters out clothing-related interference information, focusing on encoding human identification information that does not change with clothing, thus providing the model with a stable basis for identity discrimination in clothing-changing scenarios.

[0021] Correspondingly, the second expert network model focuses on learning identification features related to clothing, and its core is to enhance the modeling of clothing details through a multi-label clothing attribute recognition task. For example, the process of directly or indirectly processing the basic features using three complementary expert network models can include: inputting the basic features into the second expert network model, which processes the basic features through a spatial attention mechanism network to obtain a high-dimensional basic vector of clothing; inputting the high-dimensional basic vector of clothing into a multi-label clothing attribute classifier to obtain the predicted probability of each attribute; and using the predicted probability of each attribute to obtain the output clothing detail features.

[0022] The specific process is as follows: First, the feature extraction operation takes the basic features output by the backbone network and inputs them into the feature extraction network for feature transformation, outputting a discriminative feature representation containing rich details. Its calculation can be expressed as: ,in, As a second-expert network model, a spatial attention mechanism is employed to enhance the focusing and representation capabilities of key local areas of clothing (such as texture, pattern, and color). Then, attribute recognition supervision is performed. This is done in the feature identification phase. A multi-label clothing attribute classifier is connected in parallel above. A second expert network model is then guided to explicitly encode distinguishable clothing details in its features. The final output is a feature set with high discriminative power. This feature, supervised by the attribute recognition task, deeply encodes highly discriminative information such as clothing color, style, and local modifications, enabling high-precision identity recognition even when the clothing remains unchanged, utilizing these subtle visual cues. While the first and second expert network models output their respective features, these features may be dimensionally unrelated. A third expert network model aligns their dimensions, allowing the features to complement each other. The third expert network model learns how to fuse identity and clothing information into a unified feature vector. Thus, even if the first and second expert network models output seemingly unrelated information, the third expert network model can combine them into a complete feature vector that includes both the person and the clothing.

[0023] The third expert network model is responsible for integrating the outputs of the first two. By fusing robust features and discriminative features through feature splicing and fully connected layers, it achieves complementarity and enhancement of multi-perspective information. For example, it may include: splicing the human identity features and clothing detail features unrelated to clothing along the feature dimension to form a joint feature vector; inputting the joint feature vector into a fully connected network containing multiple nonlinear transformations for dimensionality reduction and deep fusion, and outputting a comprehensive feature that simultaneously contains identity information and clothing discriminative information.

[0024] Optionally, in the specific process, firstly, through feature concatenation, the robust features output by the first expert network model are combined. Discrimination features output by the second expert network model Concatenate along the feature dimension to form a joint feature vector. Its calculation is expressed as: , where || represents the feature concatenation operation.

[0025] Subsequently, a fully connected layer mapping operation is performed, inputting the concatenated joint feature vector into a fully connected network (FC) containing multiple nonlinear transformations for dimensionality reduction and deep fusion. The computation is represented as follows: FC is the connection layer. and The sum is a trainable parameter.

[0026] Finally, the output representation operation is performed, and the final output is the comprehensive feature. It contains both robust identity information and clothing identification information, enabling it to improve its discrimination ability by utilizing clothing details in scenarios where clothing is consistent, and to maintain recognition stability by relying on robust features in scenarios where clothing is changed.

[0027] Step 130: Input the basic features into the weight network to obtain the weights of the output features of the first expert network model, the second expert network model, and the third expert network. Generate comprehensive features based on the weights of each network and use the comprehensive features to generate the re-identification result.

[0028] While the three expert network models mentioned above can output corresponding features according to requirements, it's impossible to determine which model contributes more in a clothing-changing re-identification scenario. Therefore, this embodiment also includes a weighted network. This weighted network is used to determine the current re-identification scenario based on basic features, and then determine the contribution of each model to the current re-identification scenario. This allows for the extraction of features that fit the current scenario.

[0029] For example, a weighted network can be used to dynamically calculate the weights of each expert network model based on the features of the input image. The weighted network can consist of fully connected layers and a softmax function. ,in, and For trainable parameters, The normalized weights for each expert are used. The final feature is a weighted sum of the output features from the three experts. The final fused feature is: After obtaining the fused features, they can be input into a classification model, such as a softmax model, to output the re-identification result.

[0030] This embodiment inputs the re-identification data of pedestrians changing clothes to a feature extraction module to obtain basic features. Three complementary expert network models are then used to directly or indirectly process these basic features. These three expert network models include: a first expert network model, which outputs clothing-independent human identity features and learns clothing-independent identity features based on the basic features; a second expert network model, which outputs clothing detail features and learns to enhance these clothing detail features based on the basic features; and a third expert network model, which receives both clothing-independent human identity features and clothing detail features, integrates them along the feature dimension as input, and learns a collaborative representation of multi-perspective features. The basic features are then input into a weighted network to obtain the weights of the output features from the first, second, and third expert network models. A comprehensive feature is generated based on the weights of each network, and the re-identification result is generated using this comprehensive feature. By utilizing multiple complementary network models to extract both clothing-independent permanent features and clothing-related fine-grained features, and integrating them along the feature dimension, the comprehensiveness and complementarity of feature extraction are ensured, overcoming the difficulty of a single feature extractor to handle multi-perspective information. Furthermore, by leveraging weighted networks, it can learn which features are more important in different scenarios, effectively resolving the core contradiction in pedestrian re-identification after clothing changes: needing to both forget the interference of clothing and utilize clothing details for identification, thus significantly improving robustness and accuracy in complex real-world scenarios. Moreover, it requires no additional noise interference or additional labeled data or sensors.

[0031] Example 2 Figure 2This is a flowchart illustrating the pedestrian identification method for changing clothes provided in Embodiment 2 of the present invention. This embodiment is an optimization based on the above embodiment, and the method may further include the following steps: inputting the re-identification data of pedestrians changing clothes into the feature extraction module for training; inputting the human identity features output by the first expert network model into the first fully connected classifier, and establishing a first expert network loss function based on the re-identification result output by the first fully connected classifier and the real identity label; inputting the clothing detail features output by the second expert network model into the second fully connected classifier, and establishing a second expert network loss function based on the clothing detail re-identification result output by the second fully connected classifier and the real clothing label; inputting the comprehensive features output by the third expert network model into the third fully connected classifier, and establishing a third expert network loss function based on the re-identification result output by the third fully connected classifier and the real identity label; and according to the comprehensive features in the sample... The system identifies positive and negative samples based on their labels. Positive samples are those with the same identity but potentially different clothing, while negative samples are those with different identities but potentially similar clothing. The system calculates the comprehensive features of both positive and negative samples. It then calculates the distances between the comprehensive features of each sample and the comprehensive features of both positive and negative samples, and constructs a learning feature loss function based on distance thresholds between positive and negative samples. A first specificity loss function is determined based on adversarial training of the first expert network model, and a second specificity loss function is determined based on attribute classification supervision of the second expert network model. A comprehensive loss function is constructed using the first, second, and third expert network loss functions, the learning feature loss function, the first specificity loss function, and the second specificity loss function. Finally, the comprehensive loss function is used to optimize and update the parameters in the human feature expert network, the discrimination feature expert network, the comprehensive expert network, and the weight network.

[0032] See Figure 2 The method for identifying pedestrians changing clothes includes: Step 210: Input the pedestrian re-identification data for changing clothes into the feature extraction module for training.

[0033] Step 220: Input the human identity features output by the first expert network model into the first fully connected classifier, and establish the first expert network loss function based on the re-identification results output by the first fully connected classifier and the real identity label.

[0034] Before using the model to re-identify pedestrians changing clothes, the model needs to be trained. In this embodiment, due to the complex model structure, which includes multiple expert network models and weight models, traditional cross-entropy loss is used for training, making it impossible to optimize the model parameters.

[0035] For example, a fully connected classifier can be set up for the first expert network. Based on the pedestrian recognition results output by the fully connected classifier, the loss function of the first expert network can be established using the cross-entropy operation method, based on the re-identification results output by the first fully connected classifier and the real identity label.

[0036] Step 230: Input the clothing detail features output by the second expert network model into the second fully connected classifier, and establish the second expert network loss function based on the clothing detail re-identification results output by the second fully connected classifier and the real clothing labels.

[0037] Correspondingly, a fully connected classifier can be set up for the second expert network model. The clothing detail features output by the second expert network model are input into the second fully connected classifier. The second expert network loss function is established based on the re-identification results output by the second fully connected classifier and the real identity label.

[0038] Step 240: Input the comprehensive features output by the third expert network model into the third fully connected classifier, and establish the third expert network loss function based on the re-identification results output by the third fully connected classifier and the real identity label.

[0039] A third fully connected classifier is also set up for the third expert network model. The pedestrian recognition results output by the third expert network model are used to establish the third expert network loss function based on the re-recognition results output by the third fully connected classifier and the real identity labels using the cross-entropy operation method.

[0040] Step 250: Based on the labels corresponding to the comprehensive features in the samples, determine the corresponding positive and negative samples. The positive samples are samples with the same identity but different clothes, and the negative samples are samples with different identities but similar clothes. Calculate the comprehensive features of the positive and negative samples respectively. Calculate the distance between the comprehensive features in the samples and the comprehensive features of the positive and negative samples respectively. Based on the distance threshold between the positive and negative sample distances, construct the learning feature loss function.

[0041] Since the first, second, and third expert network models not only have independent classification and recognition capabilities but also complement and cooperate with each other in classification and recognition, other loss functions are also provided in this embodiment.

[0042] For example, a learning feature loss function can also be set to ensure that the features extracted by the third-party expert network model should map images of the same person to very close positions in the feature space, and images of different people should map to very far positions in the feature space. This ensures good separability of the different features. Therefore, based on the labels corresponding to the comprehensive features in the samples, it is necessary to determine the corresponding positive and negative samples. The positive samples are samples with the same identity but different clothing, and the negative samples are samples with different identities but similar clothing. The comprehensive features of the positive samples and the comprehensive features of the negative samples are calculated separately. The distances between the comprehensive features in the samples and the comprehensive features of the positive and negative samples are calculated separately, and the learning feature loss function is constructed based on the distance threshold between the positive and negative samples. In each batch of data, an intra-batch hard sample mining strategy is adopted to select the farthest positive sample and the closest negative sample of different identities for each anchor sample, and to construct the learning feature loss function, which is set with a margin parameter of 0.3.

[0043] Step 260: Determine the first specific loss function based on the adversarial training of the first expert network model, and determine the second specific loss function based on the attribute classification supervision of the second expert network model.

[0044] For example, the process may include: inputting the features output by the first expert network model into a gradient inversion layer, which inverts the received gradients during backpropagation; inputting the features output by the gradient inversion layer into a clothing classifier, which predicts the true clothing category; constructing a preliminary loss based on the features output by the gradient inversion layer and the cross-entropy loss of the clothing classifier, and inverting the preliminary loss to obtain an adversarial loss; updating the parameters of the clothing classifier along the direction of minimizing the error according to the adversarial loss, and updating the parameters of the first expert network model along the direction of maximizing the error to generate a first specificity loss function. Optionally, the clothing classifier may be a classifier primarily composed of softmax layers.

[0045] The first specificity loss function can be expressed as follows: ,in, For clothing classifiers, For the final fused features, E represents the average calculation, which is used to average the features.

[0046] During training, the gradient inversion layer inverts the classifier loss gradient during backpropagation and feeds it back to the first expert network model. The learned feature loss function is expressed as: Where GRL(·) represents the gradient reversal layer operation, And for the actual clothing category label, To counteract the loss, this operation forces the first-expert network model to optimize in a direction that increases the error of the clothing classifier during parameter updates, thereby suppressing clothing information in the features. This forces the output to ultimately possess highly robust features. Due to the adversarial training mechanism, these features actively filter out clothing-related interference information, focusing on encoding human identification information that does not change with clothing, thus providing the model with a stable basis for identity determination in clothing-changing scenarios.

[0047] Accordingly, determining the second specificity loss function based on the attribute classification supervision of the second expert network model may include: generating the second specificity loss function based on the number of independent attributes into which the clothing is decomposed, the binary label corresponding to each independent attribute, and the discrimination result of each attribute. Optionally, the discriminative features output by the second expert network model... A multi-label clothing attribute classifier is connected in parallel with the above. This classifier decomposes clothing into multiple specific attributes (such as color, texture, style, accessories, etc.), with each attribute treated as an independent classification task. By optimizing the attribute classification loss, we force the feature experts to explicitly encode rich clothing details in their features. The second specificity loss function is expressed as: Where M is the total number of attributes, It is the real binary tag of the m-th attribute. It is the probability predicted by the model. The loss is for multi-label attribute classification. This supervisory signal guides the second expert network model to explicitly encode distinguishable clothing details in its features. Finally, the output is a feature with high discriminative power. This feature, supervised by the attribute recognition task, deeply encodes highly discriminative information such as clothing color, style, and local modifications, enabling high-precision identity recognition even when the clothing remains unchanged, by utilizing these subtle visual cues.

[0048] Step 270: Construct a comprehensive loss function using the first expert network loss function, the second expert network loss function, the third expert network loss function, the learning feature loss function, the first specificity loss function, and the second specificity loss function. Use the comprehensive loss function to optimize and update the parameters in the human feature expert network, the discrimination feature expert network, the comprehensive expert network, and the weight network.

[0049] Multi-task joint training was conducted, and a multi-task loss function was used to supervise model training. This ensured the performance balance and overall optimization of the expert network models in the division of labor and collaboration. The total loss function was defined as: ,in, For the total loss function, The loss function of the first expert network. For the second expert network loss function, For the loss function of the third expert network, The second specificity loss function, To learn the feature loss function, This is the first specificity loss function.

[0050] Step 280: Input the re-identification data of the pedestrian changing clothes to be identified into the feature extraction module to obtain basic features. Then, use three complementary expert network models to process the basic features directly or indirectly.

[0051] Step 290: Input the basic features into the weight network to obtain the weights of the output features of the first expert network model, the second expert network model, and the third expert network. Generate comprehensive features based on the weights of each network and use the comprehensive features to generate the re-identification result.

[0052] This embodiment adds the following steps: Inputting pedestrian re-identification data for changing clothes into a feature extraction module for training; inputting the human identity features output by the first expert network model into a first fully connected classifier, and establishing a first expert network loss function based on the re-identification results output by the first fully connected classifier and the real identity label; inputting the clothing detail features output by the second expert network model into a second fully connected classifier, and establishing a second expert network loss function based on the clothing detail re-identification results output by the second fully connected classifier and the real clothing label; inputting the comprehensive features output by the third expert network model into a third fully connected classifier, and establishing a third expert network loss function based on the re-identification results output by the third fully connected classifier and the real identity label; determining the corresponding positive and negative samples based on the labels corresponding to the comprehensive features in the samples, wherein the positive samples... For samples with the same identity but potentially different clothing, and for negative samples with different identities but potentially similar clothing, comprehensive features are calculated for both positive and negative samples. The distances between the comprehensive features of each sample and the comprehensive features of both positive and negative samples are calculated, and a learning feature loss function is constructed based on distance thresholds between positive and negative samples. A first specificity loss function is determined based on adversarial training of the first expert network model, and a second specificity loss function is determined based on attribute classification supervision of the second expert network model. A comprehensive loss function is constructed using the first, second, and third expert network loss functions, the learning feature loss function, the first specificity loss function, and the second specificity loss function. The parameters in the human feature expert network, the discrimination feature expert network, the comprehensive expert network, and the weight network are optimized and updated using this comprehensive loss function. By uniformly constraining all models through the comprehensive loss function, it ensures that the model can adaptively adjust feature representation under different conditions, avoiding feature conflicts and improving the overall robustness and convergence speed of the model. This further improves the accuracy of re-identifying pedestrians changing clothes.

[0053] In a preferred embodiment of this example, the method may further include the following step: inputting the comprehensive features into a clothing classifier to obtain clothing classification results, and using the clothing classification results to maximize the error of the clothing classifier, thereby ensuring that the comprehensive features remain stable under different clothing conditions. Although a first expert network model extracts features without clothing information, the final features are obtained by fusing three expert network models. If the final features are not constrained, the fused features may still be affected by the clothing information extracted by the second expert network model. Therefore, by using the above-mentioned added steps, an adversarial strategy can be employed to ensure that the comprehensive features output by the model can maintain the purity of identity features under the interference of clothing changes, thereby significantly improving the accuracy of pedestrian re-identification under clothing changes.

[0054] Example 3 Figure 3 This is a schematic diagram of the pedestrian re-identification device for changing clothes provided in Embodiment 3 of the present invention. See also... Figure 3 The pedestrian re-identification device for changing clothes includes: The basic feature acquisition module 310 is used to input the re-identification data of the pedestrian changing clothes to be identified into the feature extraction module to obtain basic features; Processing module 320 is used to directly or indirectly process basic features using three complementary expert network models, wherein the three expert network models include: The first expert network model outputs clothing-independent human identity features. The first expert network model is used to learn clothing-independent identity features based on the basic features. The second expert network model outputs clothing detail features, and the second expert network model is used to learn and enhance clothing detail features based on the basic features. The third expert network model is used to receive human identity features and clothing detail features unrelated to clothing, and integrate them along the feature dimension as input to learn the collaborative representation of multi-view features. The generation module 330 is used to input basic features into the weight network to obtain the weights of the output features of the first expert network model, the second expert network model, and the third expert network, and to generate comprehensive features based on the weights of each network, and to generate re-identification results using the comprehensive features.

[0055] The pedestrian re-identification device for changing clothes provided in this embodiment obtains basic features by inputting the re-identification data of the pedestrian changing clothes to be identified into a feature extraction module. Three complementary expert network models are used to directly or indirectly process the basic features. These three expert network models include: a first expert network model, which outputs clothing-independent human identity features and learns clothing-independent identity features based on the basic features; a second expert network model, which outputs clothing detail features and learns enhanced clothing detail features based on the basic features; and a third expert network model, which receives clothing-independent human identity features and clothing detail features, integrates them along the feature dimensions as input, and learns the collaborative expression of multi-view features. The basic features are input into a weighted network to obtain the weights of the output features of the first, second, and third expert network models. A comprehensive feature is generated based on the weights of each network, and the re-identification result is generated using the comprehensive feature. By utilizing multiple complementary network models, persistent features unrelated to clothing and fine-grained features related to clothing are extracted separately, and then integrated along the feature dimension. This ensures the comprehensiveness and complementarity of feature extraction, overcoming the difficulty of a single feature extractor to handle multi-perspective information. Furthermore, the weighted network can learn which features are more important in different scenarios, effectively resolving the core contradiction in pedestrian re-identification after clothing changes: needing to forget the interference of clothing while utilizing clothing details for identification. This significantly improves robustness and accuracy in complex real-world scenarios. Moreover, it requires no additional noise interference or additional labeled data or sensors.

[0056] Based on the above embodiments, the processing module includes: The input unit is used to input basic features into the second expert network model, which processes the basic features through a spatial attention mechanism network to obtain a high-dimensional basic vector of clothing. The prediction unit is used to input the high-dimensional basic vector of clothing into the multi-label clothing attribute classifier to obtain the predicted probability of each attribute. The garment detail feature derivation unit is used to obtain the output garment detail features by utilizing the predicted probability of each attribute.

[0057] Based on the above embodiments, the third expert network model is used for: The human identity features unrelated to clothing and the clothing detail features are concatenated along the feature dimension to form a joint feature vector; The joint feature vector is input into a fully connected network containing multiple nonlinear transformations for dimensionality reduction and deep fusion, and the output is a comprehensive feature that simultaneously contains identity information and clothing identification information.

[0058] Based on the above embodiments, the device further includes: The training module is used to input the re-identification data of pedestrians changing clothes into the feature extraction module for training; The module for establishing the first expert network loss function is used to input the human identity features output by the first expert network model into the first fully connected classifier, and to establish the first expert network loss function based on the re-identification results output by the first fully connected classifier and the real identity label. The second expert network loss function establishment module is used to input the clothing detail features output by the second expert network model into the second fully connected classifier, and establish the second expert network loss function based on the clothing detail re-identification results output by the second fully connected classifier and the real clothing labels. The module for establishing the third expert network loss function is used to input the comprehensive features output by the third expert network model into the third fully connected classifier, and to establish the third expert network loss function based on the re-identification results output by the third fully connected classifier and the real identity label. The positive and negative sample determination module is used to determine the corresponding positive and negative samples based on the labels corresponding to the comprehensive features in the samples. The positive samples are samples with the same identity but different clothes, and the negative samples are samples with different identities but similar clothes. The module calculates the comprehensive features of the positive samples and the comprehensive features of the negative samples respectively. The module for constructing the learning feature loss function is used to calculate the distance between the comprehensive feature in the sample and the comprehensive feature of the positive sample and the comprehensive feature of the negative sample, respectively, and to construct the learning feature loss function based on the distance threshold between the positive sample distance and the negative sample distance. The loss function determination module is used to determine the first specific loss function based on the adversarial training of the first expert network model and to determine the second specific loss function based on the attribute classification supervision of the second expert network model. The comprehensive loss function construction module is used to construct a comprehensive loss function using the first expert network loss function, the second expert network loss function, the third expert network loss function, the learned feature loss function, the first specificity loss function, and the second specificity loss function; The update module is used to optimize and update the parameters in the human feature expert network, the discriminative feature expert network, the comprehensive expert network, and the weight network using the comprehensive loss function.

[0059] Based on the above embodiments, the comprehensive loss function construction module includes: The inversion processing unit is used to input the features output by the first expert network model into the gradient inversion layer, which inverts the received gradient during backpropagation. The clothing classifier input unit is used to input the features output by the gradient inversion layer into the clothing classifier, which is used to predict the true clothing category. The adversarial loss unit is used to construct a preliminary loss based on the features output by the gradient inversion layer and the cross-entropy loss of the clothing classifier. The preliminary loss is then inverted to obtain the adversarial loss. The first specific loss function generation unit is used to update the parameters of the clothing classifier in the direction of minimizing error and update the parameters of the first expert network model in the direction of maximizing error, based on the adversarial loss, to generate the first specific loss function.

[0060] Based on the above embodiments, the comprehensive loss function construction module includes: The second specificity loss function generation unit is used to generate a second specificity loss function based on the number of independent attributes into which the clothing is decomposed, the binary label corresponding to each independent attribute, and the discrimination result of each attribute.

[0061] Based on the above embodiments, the device further includes: The error maximization module is used to input the comprehensive features into the clothing classifier to obtain the clothing classification result, and to maximize the error of the clothing classifier using the clothing classification result, so that the comprehensive features remain stable under different clothes.

[0062] The pedestrian re-identification device for changing clothes provided in the embodiments of the present invention can execute the pedestrian re-identification method for changing clothes provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.

[0063] Example 4 Figure 4 This is a schematic diagram of the structure of a server provided in Embodiment 4 of the present invention. Figure 4 A block diagram of an exemplary server 12 suitable for implementing embodiments of the present invention is shown. Figure 4 The server 12 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0064] like Figure 4 As shown, server 12 is presented in the form of a general-purpose computing server. The components of server 12 may include, but are not limited to: one or more processors or processing units 16, memory 28, and bus 18 connecting different system components (including memory 28 and processing unit 16).

[0065] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0066] Server 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by server 12, including volatile and non-volatile media, removable and non-removable media.

[0067] Memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache 32. Server 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (… Figure 4 Not shown; usually referred to as a "hard drive"). Although Figure 4 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.

[0068] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of the present invention.

[0069] Server 12 can also communicate with one or more external servers 14 (e.g., keyboard, pointing server, display 24, etc.), one or more servers that enable users to interact with server 12, and / or any server (e.g., network card, modem, etc.) that enables server 12 to communicate with one or more other computing servers. This communication can be performed via input / output (I / O) interface 22. Furthermore, server 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of server 12 via bus 18. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with server 12, including but not limited to: microcode, server drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0070] Processing unit 16 executes various functional applications and data processing by running programs stored in memory 28, such as implementing the pedestrian recognition method for changing clothes provided in this embodiment of the invention. Example 5 Embodiment 5 of the present invention also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform any of the clothing-changing pedestrian recognition methods provided in the above embodiments.

[0071] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0072] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0073] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0074] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0075] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.

Claims

1. A method for identifying pedestrians changing clothes, characterized in that, include: Input the re-identification data of the pedestrian changing clothes to be identified into the feature extraction module to obtain the basic features; The basic features are processed directly or indirectly using three complementary expert network models, which include: The first expert network model outputs clothing-independent human identity features. The first expert network model is used to learn clothing-independent identity features based on the basic features. The second expert network model outputs clothing detail features, and the second expert network model is used to learn and enhance clothing detail features based on the basic features. The third expert network model is used to receive human identity features and clothing detail features unrelated to clothing, and integrate them along the feature dimension as input to learn the collaborative representation of multi-view features. The basic features are input into the weight network to obtain the weights of the output features of the first expert network model, the second expert network model, and the third expert network. A comprehensive feature is generated based on the weights of each network, and the re-identification result is generated using the comprehensive feature.

2. The method according to claim 1, characterized in that, The process of directly or indirectly processing basic features using three complementary expert network models includes: The basic features are input into the second expert network model, which processes the basic features through a spatial attention mechanism network to obtain the high-dimensional basic vector of the clothing. Input the high-dimensional basic vector of clothing into a multi-label clothing attribute classifier to obtain the predicted probability of each attribute; The predicted probability of each attribute is used to obtain the output clothing detail features.

3. The method according to claim 1, characterized in that, The third expert network model is used to receive human identity features and clothing detail features unrelated to clothing, and integrate them along the feature dimension as input to learn the collaborative representation of multi-view features, including: The human identity features unrelated to clothing and the clothing detail features are concatenated along the feature dimension to form a joint feature vector; The joint feature vector is input into a fully connected network containing multiple nonlinear transformations for dimensionality reduction and deep fusion, and the output is a comprehensive feature that simultaneously contains identity information and clothing identification information.

4. The method according to claim 1, characterized in that, The method further includes: Input the re-identification data of pedestrians changing clothes into the feature extraction module for training; The human identity features output by the first expert network model are input into the first fully connected classifier. The first expert network loss function is established based on the re-identification results output by the first fully connected classifier and the real identity label. The clothing detail features output by the second expert network model are input into the second fully connected classifier. The second expert network loss function is established based on the clothing detail re-identification results output by the second fully connected classifier and the real clothing labels. The comprehensive features output by the third expert network model are input into the third fully connected classifier, and the loss function of the third expert network is established based on the re-identification results output by the third fully connected classifier and the real identity label. Based on the labels corresponding to the comprehensive features in the samples, the corresponding positive and negative samples are determined. The positive samples are samples with the same identity but different clothes, and the negative samples are samples with different identities but similar clothes. The comprehensive features of the positive samples and the comprehensive features of the negative samples are calculated respectively. Calculate the distances between the comprehensive features of the sample and the comprehensive features of the positive sample and the comprehensive features of the negative sample, respectively, and construct the learning feature loss function based on the distance thresholds between the positive sample distance and the negative sample distance; The first specificity loss function is determined based on the adversarial training of the first expert network model, and the second specificity loss function is determined based on the attribute classification supervision of the second expert network model. A comprehensive loss function is constructed using the first expert network loss function, the second expert network loss function, the third expert network loss function, the learned feature loss function, the first specificity loss function, and the second specificity loss function. The parameters in the human feature expert network, the discriminative feature expert network, the comprehensive expert network, and the weight network are optimized and updated using the comprehensive loss function.

5. The method according to claim 4, characterized in that, The step of determining the first specificity loss function based on adversarial training of the first expert network model includes: The features output by the first expert network model are input into the gradient inversion layer, which inverts the received gradients during backpropagation. The features output by the gradient inversion layer are input into a clothing classifier, which is used to predict the true clothing category. A preliminary loss is constructed based on the features output by the gradient inversion layer and the cross-entropy loss of the clothing classifier. The preliminary loss is then inverted to obtain the adversarial loss. Based on the adversarial loss, the parameters of the clothing classifier are updated along the direction of minimizing the error, and the parameters of the first expert network model are updated along the direction of maximizing the error to generate the first specificity loss function.

6. The method according to claim 4, characterized in that, The step of determining the second specificity loss function based on attribute classification supervision of the second expert network model includes: A second specificity loss function is generated based on the number of independent attributes into which the clothing is decomposed, the binary label corresponding to each independent attribute, and the discrimination result of each attribute.

7. The method according to claim 1, characterized in that, The method further includes: The comprehensive features are input into the clothing classifier to obtain the clothing classification result. The error of the clothing classifier is maximized using the clothing classification result, thereby making the comprehensive features stable under different clothes.

8. A device for re-identifying pedestrians changing clothes, characterized in that, include: The basic feature acquisition module is used to input the re-identification data of the pedestrian changing clothes to be identified into the feature extraction module to obtain basic features; The processing module is used to directly or indirectly process the basic features using three complementary expert network models, wherein the three expert network models include: The first expert network model outputs clothing-independent human identity features. The first expert network model is used to learn clothing-independent identity features based on the basic features. The second expert network model outputs clothing detail features, and the second expert network model is used to learn and enhance clothing detail features based on the basic features. The third expert network model is used to receive human identity features and clothing detail features unrelated to clothing, and integrate them along the feature dimension as input to learn the collaborative representation of multi-view features. The generation module is used to input basic features into the weight network to obtain the weights of the output features of the first expert network model, the second expert network model, and the third expert network, and to generate comprehensive features based on the weights of each network, and to generate re-identification results using the comprehensive features.

9. A server, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the pedestrian recognition method for changing clothes as described in any one of claims 1-7.

10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the pedestrian identification method for changing clothes as described in any one of claims 1-7.