Cross-dataset reloading pedestrian re-identification method and system based on self-adaptive matching during testing

By building a backbone network and the adaptive matching module during testing, the problem of clothing changes and data distribution differences in pedestrian re-identification across data sets is solved, and pedestrian re-identification with high accuracy and interpretability is achieved, adapting to different environmental conditions, and simplifying the deployment process.

CN120236300APending Publication Date: 2025-07-01HENAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510308924.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing pedestrian re-identification method has reduced recognition accuracy due to data distribution differences and clothing changes across data set scenarios, lacks adaptive adjustment capabilities, and lacks limitations and explanatory features.

Method used

The backbone network and the adaptive matching module during testing are built, and the impact of clothing changes is reduced through the clothing weakening module, the adaptive matching module is introduced to adjust feature extraction, and the total loss function optimization model is adopted to achieve adaptability and interpretability across data sets.

Benefits of technology

It improves the identification accuracy and generalization ability of the model on the unknown data set, reduces the sensitivity to clothing changes, enhances the interpretability of the model, simplifies the deployment process, and improves the efficiency and reliability of the monitoring system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236300A_ABST
    Figure CN120236300A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-dataset reloading pedestrian re-identification method and system based on self-adaptive matching during testing, and the method comprises the steps: 1, carrying out the preprocessing of a plurality of given pedestrian images, and obtaining a preprocessed image set corresponding to the plurality of pedestrian images; step 2, inputting images in the plurality of preprocessed image sets into a backbone network to obtain feature maps corresponding to the input images; 3, inputting the feature map corresponding to the input image into a self-adaptive matching module during testing to obtain a predicted matching score; 4, training the backbone network and a self-adaptive matching module during testing based on the predicted matching score and a preset total loss function; and step 5, identifying the plurality of target images according to the optimal backbone network and the self-adaptive matching module during testing, and completing cross-set reloading pedestrian re-identification. According to the method, the performance of the model on an unseen data set is improved, so that the model has good interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cross-dataset dressed pedestrian re-identification, and particularly to a cross-dataset dressed pedestrian re-identification method and system based on adaptive matching during testing. Background Art

[0002] With the rapid development of video surveillance technology, pedestrian re-identification has become one of the key technologies in intelligent security and smart city construction. Pedestrian re-identification aims to determine whether the people in the images captured by different cameras are of the same identity, which is of great significance for tracking, positioning, and identifying individuals in a multi-camera system. However, in practical applications, due to factors such as individual clothing changes, camera view differences, and lighting condition changes, cross-camera pedestrian re-identification faces huge challenges. Especially in the scenario of cross-datasets, due to the differences in the distributions of training data and test data, the performance of existing pedestrian re-identification methods has decreased significantly. Therefore, researching a pedestrian re-identification method that can effectively handle clothing changes and adapt to different dataset distributions has important practical significance for improving the accuracy and practicality of the monitoring system.

[0003] In the field of pedestrian re-identification, existing technologies face a series of challenges, among which the most prominent are the impacts of data distribution differences and clothing changes. Most existing methods are designed under the assumption that the training and test data come from the same distribution, but this assumption often does not hold in real-world applications. Images captured by different cameras may have different feature distributions due to factors such as environment, lighting, and season, which poses a huge challenge to models that rely on the features of the source dataset. In addition, an individual's appearance can change significantly due to wearing different clothes, which is a difficult problem for re-identification systems based on visual features. Existing methods often cannot effectively distinguish the feature changes caused by clothing changes from the true differences between individual identities, resulting in a decrease in recognition accuracy.

[0004] On the other hand, existing methods lack the ability to adaptively adjust during testing, have limitations in feature representation, and have the risk of overfitting. When a model is applied to a new, unseen dataset, it needs to be able to adapt to the feature distribution of the new dataset, but many existing models are not designed with this in mind. At the same time, the limitations of feature representation are also a problem. The features extracted by existing methods are often highly correlated with clothing information, which limits the recognition performance of the model under clothing change conditions. Due to relying on a large amount of labeled data for training, existing models may overfit the specific features of the training dataset, resulting in a decrease in performance on unseen datasets. In addition, many deep learning-based pedestrian re-identification methods lack interpretability and are regarded as "black boxes", making it difficult to clearly show the decision-making process and basis of the model to users, which is an obvious defect in application scenarios that require high transparency. Summary of the Invention

[0005] In order to at least partially solve the problems that existing pedestrian re-identification methods cannot adapt to different data distributions, are sensitive to clothing changes, have low recognition accuracy, and lack interpretability, the present invention provides a cross-dataset dressed pedestrian re-identification method and system based on test-time adaptive matching. The present invention realizes pedestrian re-identification by constructing a backbone network and a test-time adaptive matching module as an overall model. The last layer in the last two ResNet Blocks in the backbone network is replaced with a clothing weakening module, which is specifically used to reduce the impact of clothing changes on feature extraction, enhance the generalization ability of features, enable it to adapt to different data distributions, improve the sensitivity of the model to clothing changes, and improve recognition accuracy. The present invention adaptively adjusts the matching model through the test-time adaptive matching module to cope with the changes in the pedestrian images to be matched, improves the performance of the model on unseen datasets, and makes it have good interpretability. The present invention can be trained and learned on the source dataset, and can effectively extract general features irrelevant to clothing, enabling it to perform well on unseen target datasets, which plays a crucial role in improving the efficiency and reliability of the monitoring system.

[0006] To achieve the above object, the technical solution of the present invention is as follows:

[0007] The first aspect of the present invention proposes a cross-dataset dressed pedestrian re-identification method based on test-time adaptive matching, including:

[0008] Step 1: Preprocess each of the given multiple pedestrian images to obtain a set of preprocessed images corresponding to the multiple pedestrian images, which is convenient for training the model;

[0009] Step 2: Input the images in the multiple sets of preprocessed images into the backbone network to obtain the feature maps corresponding to the input images, which are used to extract image features;

[0010] Step 3: Input the feature maps corresponding to the input images into the test-time adaptive matching module to obtain the predicted matching scores;

[0011] Step 4: Train the backbone network and the test-time adaptive matching module based on the predicted matching scores and a preset total loss function to obtain the optimal backbone network and test-time adaptive matching module, which is convenient for improving recognition accuracy;

[0012] Step 5: Identify multiple target images according to the optimal backbone network and test-time adaptive matching module to complete cross-set dressed pedestrian re-identification.

[0013] Further, the preprocessing includes cropping the image to a certain size, random cropping, random horizontal flipping, random erasing, and random color jittering, which is convenient for obtaining the image set.

[0014] Further, the backbone network includes five sequentially connected ResNet Blocks; among them, the activation function layers of the last layer in the fourth ResNet Block and the fifth ResNet Block are replaced with clothing weakening modules;

[0015] The clothing weakening module includes a channel feature amplifier, a spatial feature amplifier, and a connection layer for extracting clothing features;

[0016] The channel feature amplifier is used to convert the features of different images into a single distribution to facilitate adaptation to different data distributions;

[0017] The spatial feature amplifier is used to generate a spatial feature map;

[0018] The connection layer is used to connect the output of the channel feature amplifier, the output of the spatial feature amplifier, and the input features of the clothing weakening module.

[0019] Further, the channel feature amplifier is represented by the following formula:

[0020]

[0021] where σ is the sigmoid function, MLP is a multi-layer perceptron, AvgPool is average pooling, MaxPool is max pooling, Q is the input of the clothing weakening module, and both W1 and W0 are the weights of the multi-layer perceptron, is the feature after average pooling in the channel feature amplifier, is the feature after max pooling in the channel feature amplifier.

[0022] Further, the spatial feature amplifier is represented by the following formula:

[0023] F s (Q) = σ(f 7×7 (AvgPool(Q); MaxPool(Q))) = σ(f 7×7 ([Q avg ; Q max ))

[0024] where F s (Q) is the output of the spatial feature amplifier, f 7×7 is a convolution operation with a filter size of 7×7, Q avg is the feature after cross-channel average pooling in the spatial feature amplifier, and Q max is the feature after cross-channel max pooling in the spatial feature amplifier.

[0025] Further, the connection layer is used to connect the output of the channel feature amplifier, the output of the spatial feature amplifier, and the input of the clothing weakening module, which specifically includes:

[0026] Perform a matrix multiplication operation on the input of the clothing weakening module and the output of the channel feature amplifier to obtain a first processed feature;

[0027] Perform a matrix multiplication operation on the input of the clothing weakening module and the output of the spatial feature amplifier to obtain a second processed feature;

[0028] Perform an element-wise addition on the first processed feature and the second processed feature to obtain a third processed feature;

[0029] Perform an element-wise addition on the input of the clothing weakening module and the third processed feature to obtain the feature map corresponding to the input image.

[0030] Further, the adaptive matching module during testing includes a matching part, a partial perceptual similarity adaptation sub-module, a most similar part selection sub-module, and a scoring sub-module, which are convenient for adaptively adjusting the matching model to cope with the changes in the images to be matched;

[0031] The matching part is used to search for the matching corresponding blocks of the feature maps corresponding to two input images within a horizontal strip of a preset height, and obtain the matching regions of the feature maps corresponding to the two input images according to the matching corresponding blocks, which is convenient for subsequent image matching;

[0032] The partial perceptual similarity adaptation sub-module is used to reshape a corresponding block in the matching regions of the feature maps corresponding to the two input images into column vectors respectively, perform a convolution calculation on the two column vectors, and finally repeat the operation for all blocks within the matching regions to obtain two similarity matrices;

[0033] The most similar part selection sub-module is used to determine the most similar regions of each block according to the maximum similarity value, and generate two similarity vectors;

[0034] The scoring sub-module is used to calculate the predicted matching score according to the two similarity vectors; among them, the scoring sub-module includes a splicing layer, a first batch normalization layer, a fully connected layer, a summation layer, a second batch normalization layer, and an activation function layer connected in sequence, which is convenient for completing person re-identification according to the predicted matching score.

[0035] Further, the total loss function is expressed by the following formula:

[0036] L total =L ID +L C +λL ver +γL CA

[0037]

[0038]

[0039] Among them, L total is the total loss function, L ID is the recognition loss, L C is the clothing classification loss, L ver is the predicted matching score loss, L CA is the clothing-agnostic adversarial loss, λ is the weight of the validation loss for learning domain-invariant representations, γ is the weight of learning clothing-agnostic features in the unseen target domain, y i is the identity label, N is the batch size, P(y i |z i ) is the probability predicted by the identity classifier for the i-th image from the feature map Z i , B is the number of images, q ij is the judgment parameter, s ij is the predicted matching score between the i-th image and the j-th image, N C is the number of clothing categories in the training set, is the predicted probability that the feature map z i belongs to the i-th piece of clothing, is the clothing label, f i is the normalized feature vector of the sample, is the normalized weight vector of the c-th type of clothing, τ is the temperature parameter of the smooth distribution, and q(c) is the weight function.

[0040] In the second aspect of the present invention, a cross-dataset clothing-changing pedestrian re-identification system based on test-time adaptive matching is proposed, including:

[0041] A preprocessing module for preprocessing a given plurality of pedestrian images respectively to obtain a set of preprocessed images corresponding to the plurality of pedestrian images, facilitating the training of the model;

[0042] A feature extraction module for inputting the images in the set of preprocessed images into the backbone network to obtain the feature maps corresponding to the input images, for extracting image features;

[0043] A scoring module for inputting the feature maps corresponding to the input images into the test-time adaptive matching module to obtain the predicted matching scores;

[0044] A training module for training the backbone network and the test-time adaptive matching module based on the predicted matching scores and a preset total loss function to obtain the optimal backbone network and test-time adaptive matching module, facilitating the improvement of the recognition accuracy;

[0045] The recognition module is used to recognize multiple target images based on the optimal backbone network and the test-time adaptive matching module, and complete cross-set clothing-changing person re-identification.

[0046] Advantages of the present invention:

[0047] (1) By introducing the adaptive matching (TEAM) module and the clothing weakening module (CVA) module, the present invention significantly improves the generalization ability across datasets, enabling the person re-identification model to not only perform excellently on the training set but also maintain a high accuracy rate on unseen target datasets. This cross-dataset adaptability is crucial for practical monitoring and security applications because it allows the model to stably identify individuals in different environments and conditions, even when there are significant clothing changes, and still maintain the accuracy of identification. The present invention can adapt to different data distributions, reduce sensitivity to clothing changes, improve identification accuracy, and has good interpretability, which is crucial for improving the efficiency and reliability of monitoring systems.

[0048] (2) The present invention simplifies the deployment process. Since it does not rely on additional auxiliary information such as gait, body shape, or contour sketches, and can achieve efficient person re-identification only using the RGB modality, this reduces the technical threshold and cost, making the system easier to deploy and apply in various scenarios. This simplification not only improves the usability of the system but also provides new possibilities for the development of intelligent monitoring systems.

[0049] (3) The present invention has been extensively experimentally verified on multiple publicly available clothing-changing datasets, and the results show that its performance on cross-dataset tasks is significantly better than existing clothing-changing person re-identification methods. These experimental results not only prove the effectiveness of the method of the present invention but also demonstrate its potential and value in practical applications. Especially in complex scenarios that require dealing with long time spans and variable environmental conditions, the present invention can provide a more reliable and robust solution. Description of the drawings

[0050] Figure 1 It is a flowchart of the cross-dataset clothing-changing person re-identification method based on test-time adaptive matching provided by an embodiment of the present invention.

[0051] Figure 2 It is a schematic diagram of the overall structure of the cross-dataset clothing-changing person re-identification method based on test-time adaptive matching provided by an embodiment of the present invention.

[0052] Figure 3 It is a flowchart of the training process of the cross-dataset clothing-changing person re-identification method based on test-time adaptive matching provided by an embodiment of the present invention.

[0053] Figure 4 It is a schematic diagram of the visualization result provided by an embodiment of the present invention.

[0054] Figure 5 Schematic diagram of the matching part provided by the embodiment of the present invention.

[0055] Figure 6 Architecture diagram of the cross-dataset clothing-changing pedestrian re-identification system based on adaptive matching during testing provided by the embodiment of the present invention. Detailed implementation manners

[0056] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0057] Embodiment 1

[0058] As Figure 1 and Figure 2 shown, the cross-dataset clothing-changing pedestrian re-identification method based on adaptive matching during testing includes:

[0059] S101: Preprocess each of the given multiple pedestrian images to obtain a set of preprocessed images corresponding to the multiple pedestrian images.

[0060] Specifically, given two pedestrian images with different clothing style attributes, denoted as P1 and P2, and preprocess these two pedestrian images P1 and P2 respectively. The preprocessing process includes cropping the images to a certain size, random cropping, random horizontal flipping, random erasing, and random color jittering.

[0061] S102: Input the images in the multiple sets of preprocessed images into the backbone network to obtain the feature maps corresponding to the input images.

[0062] Specifically, the backbone network includes 5 cascaded ResNet Blocks (residual blocks, which are conventional technical settings in this field) from shallow to deep, and replace the last layer of the latter 2 ResNet Blocks with a Clothing Variation Alleviation Module (CVA). The CVA module aims to learn a set of distribution parameters and convert the features of different images into the same distribution so that the model can reduce the differences caused by clothing variations.

[0063] S103: Input the feature maps corresponding to the input images into the adaptive matching module during testing to obtain the matching scores.

[0064] S104: Train the backbone network and the adaptive matching module at test time based on the predicted matching scores and a preset total loss function to obtain the optimal backbone network and the adaptive matching module at test time.

[0065] Specifically, the specific training process is as Figure 3 shown.

[0066] S105: Identify multiple target images according to the optimal backbone network and the adaptive matching module at test time to complete cross-set clothing-changing person re-identification.

[0067] Specifically, the result is as Figure 4 shown.

[0068] The present invention constructs an identification model through a backbone network and an adaptive matching module at test time. By replacing the last layer of the last two Resnet Blocks in the backbone network with a clothing variation alleviation module, which is specifically used to reduce the impact of clothing changes on feature extraction and enhance the generalization ability of features. Then, the predicted matching scores are obtained through the adaptive matching module at test time to complete cross-set clothing-changing person re-identification. The present invention adaptively adjusts the matching model through the adaptive matching module at test time to cope with the changes in the pedestrian images to be matched, improves the performance of the model on unseen data sets, and makes it have good interpretability.

[0069] Embodiment 2

[0070] Based on the above embodiment, the present invention proposes the specific structure of a clothing variation alleviation module (Clothing Variation Alleviation Module, CVA), which specifically includes:

[0071] The clothing variation alleviation module (CVA) mainly consists of a channel feature amplifier (CFA), a spatial feature amplifier (SFA) and a connection layer. The channel feature amplifier aims to learn a suitable common distribution and then transform the features of different images into this distribution.

[0072] First, aggregate the spatial information of the feature map by using average pooling and max pooling operations to generate two different spatial context descriptors: and They represent the average pooling feature and the max pooling feature in the channel feature amplifier respectively. Then, these two descriptors are forwarded to a shared network to generate a channel feature map F c ∈R C×1×1 . The shared network consists of a multi-layer perceptron (MLP) with a hidden layer. To reduce the parameter overhead, the hidden activation size is set to R C / r×1×1, where r is the reduction ratio. After the shared network is applied to each descriptor, the output feature vectors are merged by element-wise summation. The above process is represented by the following formula:

[0073]

[0074] where σ is the sigmoid function, MLP is the multi-layer perceptron, AvgPool is the average pooling, MaxPool is the max pooling, Q is the input of the clothing weakening module, and both W1 and W0 are the weights of the multi-layer perceptron, W0 ∈ R C / r×C , W1 ∈ R C×C / r , C is the number of channels, r is the reduction ratio, is the feature after average pooling in the channel feature amplifier, is the feature after max pooling in the channel feature amplifier. Among them, the weights W0 and W1 of the MLP (multi-layer perceptron) are shared for both inputs, and the ReLU activation function follows W0.

[0075] For the spatial feature amplifier, a spatial feature map is generated by leveraging the mutual spatial relationships of the features. To calculate the spatial features, first, average pooling and max pooling operations are applied along the channel axis, and they are concatenated to generate a valid feature descriptor. Applying pooling operations along the channel axis has been proven effective in highlighting information-rich regions. On the concatenated feature descriptor, a convolutional layer is applied to generate a spatial attention map F s (Q) ∈ R H×W , which encodes the locations that should be emphasized or suppressed. Specifically, by using two pooling operations to aggregate the channel information of the feature map, two 2D maps are generated: Q avg ∈ R 1×H×W and Q max ∈ R 1×H×W . Each map represents the average pooling feature and the max pooling feature across channels, respectively. Then they are concatenated and convolved through a standard convolutional layer to generate a 2D spatial attention map. The above process is represented by the following formula:

[0076] F s (Q) = σ(f 7×7 (AvgPool(Q); MaxPool(Q))) = σ(f 7×7 ([Q avg ; Q max ))

[0077] where F s (Q) is the output of the spatial feature amplifier, f 7×7 is the convolutional operation with a filter size of 7×7, and Q avgQ is the feature after cross-channel average pooling in the spatial feature amplifier. max max is the feature after cross-channel max pooling in the spatial feature amplifier.

[0078] The connection layer is used to connect the output of the channel feature amplifier, the output of the spatial feature amplifier, and the input of the clothing weakening module, specifically including:

[0079] Perform a matrix multiplication operation on the input of the clothing weakening module and the output of the channel feature amplifier to obtain the first processed feature.

[0080] Perform a matrix multiplication operation on the input of the clothing weakening module and the output of the spatial feature amplifier to obtain the second processed feature.

[0081] Perform an element-wise addition on the first processed feature and the second processed feature to obtain the third processed feature.

[0082] Perform an element-wise addition on the input of the clothing weakening module and the third processed feature to obtain the feature map corresponding to the input image.

[0083] Embodiment 3

[0084] Based on the above embodiments, the present invention proposes the specific structure of the Test-time Adaptive Matching module (TEAM), specifically including:

[0085] The Test-time Adaptive Matching module takes the feature maps extracted by the backbone network (denoted as T1 and T2, corresponding to the pedestrian images P1 and P2 input to the backbone network respectively). The Test-time Adaptive Matching module includes a matching part, a partial perception similarity adaptation sub-module, a most similar part selection sub-module, and a scoring sub-module.

[0086] The matching part is used to search for the matching corresponding blocks of the feature maps corresponding to the two input images within a horizontal strip of a preset height, and obtain the matching regions of the feature maps corresponding to the two input images according to the matching corresponding blocks. The partial perception similarity adaptation sub-module is used to reshape each corresponding block in the matching regions of the feature maps corresponding to the two input images into column vectors respectively, perform convolution calculations on the two column vectors, and finally repeat the operation for all blocks in the matching regions to obtain two similarity matrices. The most similar part selection sub-module is used to determine the most similar regions of each block according to the maximum similarity value, and generate two similarity vectors. The scoring sub-module is used to calculate the matching score according to the two similarity vectors; wherein, the scoring sub-module includes a splicing layer, a first batch normalization layer, a fully connected layer, a summation layer, a second batch normalization layer, and an activation function layer connected in sequence.

[0087] Specifically, the TEAM module adaptively adjusts the matching model according to the pedestrian images to be matched. First, the features of each image are separately fed into the cross-image parameter predictor to generate dynamic convolution kernel parameters for other images. Subsequently, using the obtained kernel parameters, partial perception features are derived for the comparison images through the dynamic convolution network. Finally, based on the matching perception features of these two images, an adaptive matching score can be calculated.

[0088] The working principle of this algorithm is as follows: Let [c, h, w] represent the dimensions (number of channels, height, and width) of the feature map obtained from the CVA module. For each block of size s×s, the matching part is dynamically calculated according to its position. Specifically, a matching part ratio m is defined, and the search for the matching part is restricted within a horizontal strip of height h / m (i.e., 1 / m of the total feature map height). For the block located in the i-th row, where i ∈ [0, h - s], its matching corresponding block is searched between the [low, high] rows. Specifically as Figure 5 shown.

[0089] To calculate the similarity between pedestrian A (T1) and pedestrian B (T2), sliding kernel convolution is applied within the matching region. For the i-th block of size s×s of pedestrian A, the feature map is reshaped into a column vector x i , with a size of [c×s×s, 1]. The same process is applied to the corresponding j-th block of pedestrian B, reshaping it into y j . The convolution calculation between these vectors is This represents the cosine similarity between the two blocks because both vectors are l2-normalized. By repeating this operation for all blocks, a similarity matrix S of size M×N is obtained x , which represents the relationship between the corresponding regions of pedestrian A and pedestrian B, and a similar matrix S y , representing the relationship from pedestrian B to the corresponding regions of pedestrian A. Using these matrices, the most similar regions of each block are determined by selecting the maximum similarity value, generating two similarity vectors, representing the relationships from A to B and from B to A respectively. These similarity vectors are concatenated and passed through the first batch normalization layer, fully connected layer, summation layer, second batch normalization layer, and activation function layer to produce the final matching score. The sigmoid function is applied to ensure that the matching score can be directly used for pedestrian re-identification. For n1 images of pedestrian A and n2 images of pedestrian B, a similarity score matrix of size [n1, n2] is generated, representing the similarity between each pair of images.

[0090] Embodiment 4

[0091] Based on the above embodiments, the present invention proposes a total loss function, specifically including:

[0092] The total loss function is constructed by building multiple types of loss functions as follows:

[0093] The loss function during model training is:

[0094]

[0095] Among them, L ver is the matching score loss, B is the number of images, q ij is the judgment parameter, s ij is the predicted matching score between the i-th image and the j-th image. Among them, q ij indicates whether image i and j belong to the same person. If image i and j are from the same pedestrian, then q ij = 1, otherwise q ij = 0.

[0096] By minimizing the recognition loss L ID , the ID feature can be obtained through the identity classifier as follows:

[0097]

[0098] Among them, L ID is the recognition loss, y i is the identity label, N is the batch size, P(y i |z i ) is the probability predicted by the identity classifier for the i-th image from the feature map Z i .

[0099] Then, a clothing classifier trained by the clothing classification loss L C is adopted to utilize the real clothing label while maintaining clothing information in the feature space. L C can be expressed as:

[0100]

[0101] Among them, L C is the clothing classification loss, N C is the number of clothing categories in the training set, is the predicted probability that the feature map z i belongs to the i-th clothing, is the clothing label.

[0102] The clothing-agnostic adversarial loss L CA aims to reduce the impact of clothing-related features on the feature extraction process while maintaining the model's ability to recognize pedestrians under clothing variations. Different from L C which focuses on clothing classification, L CAThe goal is to distinguish samples with the same identity but wearing different outfits. For a given input feature z i , the loss treats all clothing categories belonging to the same identity as positive classes and all other clothing categories as negative classes. This multi-positive-class classification loss is defined as:

[0103]

[0104]

[0105] where L CA is the clothing-agnostic adversarial loss, f i is the normalized feature vector of the sample, is the normalized weight vector for the c-th clothing category, τ is the temperature parameter for the smooth distribution, q(c) is the weight function, K is the number of positive clothing categories, and ∈ is a hyperparameter in the range 0 < ∈ ≤ 1.

[0106] The total loss function is represented as follows:

[0107] L total = L ID + L C + λL ver + γL CA

[0108] where L total is the total loss function, λ is the weight of the validation loss for learning domain-invariant representations, and γ is the weight of the clothing-agnostic feature learning in the unseen target domain.

[0109] Example 5

[0110] Based on the above embodiments, the present invention provides the verification results of the present invention, specifically including:

[0111] To verify the effectiveness of this experimental method, the following experimental data is provided.

[0112] Evaluation metrics: The cumulative matching characteristic (CMC) with rank 1 and the mean average precision (mAP) are used to evaluate the performance.

[0113] Experimental settings: The present invention sets two test settings, defined as follows: (1) General setting (training and testing on the same dataset); (2) Cross-dataset setting (training and testing on different datasets). The performance of existing person re-identification methods with outfit changes (such as CAL, CAMC, DCR-ReID) is compared with the method of the present invention under the cross-dataset setting, and the comparison results are shown in Table 1.

[0114] Table 1 Experimental results of model recognition performance under the cross-dataset setting

[0115]

[0116]

[0117] As shown in Table 1, the present invention demonstrated excellent performance in the experiment. Compared with the existing methods, significant improvements were achieved in multiple key indicators. Especially in the cross-dataset person re-identification task, whether in the general mode or the clothing change mode, the method of the present invention can achieve higher recognition accuracy and average precision. These results fully prove the adaptability and effectiveness of the present invention in dealing with different datasets and clothing changes, significantly improving the performance of person re-identification and reaching the leading level in the industry.

[0118] Example 6

[0119] Based on the above embodiments, as Figure 6 shown, the present invention provides a cross-dataset clothing-changing person re-identification system based on test-time adaptive matching, including:

[0120] A preprocessing module for preprocessing a given plurality of pedestrian images respectively to obtain a set of preprocessed images corresponding to the plurality of pedestrian images.

[0121] A feature extraction module for inputting the images in the plurality of preprocessed image sets into a backbone network to obtain feature maps corresponding to the input images.

[0122] A scoring module for inputting the feature maps corresponding to the input images into a test-time adaptive matching module to obtain matching scores.

[0123] A training module for training the backbone network and the test-time adaptive matching module based on the predicted matching scores and a preset total loss function to obtain an optimal backbone network and test-time adaptive matching module.

[0124] An identification module for identifying a plurality of target images according to the optimal backbone network and the test-time adaptive matching module to complete cross-set clothing-changing person re-identification.

[0125] It should be noted that the cross-dataset clothing-changing person re-identification system based on test-time adaptive matching provided in the embodiments of the present invention is to implement the above-mentioned cross-dataset clothing-changing person re-identification method based on test-time adaptive matching. Its functions can be specifically referred to the above method embodiments and will not be elaborated here.

[0126] In summary, by introducing the Adaptive Matching (TEAM) module and the Clothing Variation Adaptation (CVA) module, the present invention significantly improves the generalization ability across datasets, enabling the person re-identification model to not only perform excellently on the training set but also maintain a high accuracy rate on unseen target datasets. This cross-dataset adaptability is crucial for practical surveillance and security applications as it allows the model to stably identify individuals in different environments and conditions, even when there are significant clothing variations, while maintaining the accuracy of identification. The present invention can adapt to different data distributions, reduce sensitivity to clothing variations, improve identification accuracy, and has good interpretability, which is crucial for improving the efficiency and reliability of surveillance systems. The present invention simplifies the deployment process. Since it does not rely on additional auxiliary information such as gait, body shape, or silhouette sketches and can achieve efficient person re-identification using only the RGB modality, this reduces the technical threshold and cost, making the system easier to deploy and apply in various scenarios. This simplification not only improves the usability of the system but also provides new possibilities for the development of intelligent surveillance systems. The present invention has been extensively experimentally verified on multiple publicly available clothing-changing datasets, and the results show that its performance on cross-dataset tasks is significantly better than existing clothing-changing person re-identification methods. These experimental results not only prove the effectiveness of the method of the present invention but also demonstrate its potential and value in practical applications, especially in complex scenarios that require handling long time spans and variable environmental conditions, where the present invention can provide a more reliable and robust solution.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A cross-dataset costume-changing pedestrian re-identification method based on test-time adaptive matching, characterized in that: include: Step 1: Preprocessing the given multiple pedestrian images respectively to obtain a preprocessed image set corresponding to the multiple pedestrian images; Step 2: Input multiple preprocessed images in the image set into the backbone network to obtain the feature map corresponding to the input image; Step 3: Input the feature map corresponding to the input image into the adaptive matching module during testing to obtain the predicted matching score; Step 4: Train the backbone network and the test-time adaptive matching module based on the predicted matching score and the preset total loss function to obtain the optimal backbone network and test-time adaptive matching module; Step 5: Identify multiple target images based on the optimal backbone network and the adaptive matching module during testing to complete cross-set pedestrian re-identification.

2. The method for cross-dataset costume-changing pedestrian re-identification based on test-time adaptive matching according to claim 1, characterized in that: The preprocessing includes cutting the image to a certain size, random cropping, random horizontal flipping, random erasing and random color jittering.

3. The method for cross-dataset costume-changing pedestrian re-identification based on test-time adaptive matching according to claim 1, characterized in that: The backbone network includes five ResNet Blocks connected in sequence; wherein the activation function layer of the last layer in the fourth ResNet Block and the fifth ResNet Block is replaced with a clothing weakening module; The clothing weakening module includes a channel feature amplifier, a spatial feature amplifier and a connection layer; The channel feature amplifier is used to convert the features of different images into a distribution; The spatial feature amplifier is used to generate a spatial feature map; The connection layer is used to connect the output of the channel feature amplifier, the output of the spatial feature amplifier and the input features of the clothing weakening module.

4. The method for cross-dataset costume-changing pedestrian re-identification based on test-time adaptive matching according to claim 3, characterized in that: The channel characteristic amplifier is expressed by the following formula: Among them, σ is the sigmoid function, MLP is the multi-layer perceptron, AvgPool is the average pooling, MaxPool is the maximum pooling, Q is the input of the clothing weakening module, W1 and W0 are the weights of the multi-layer perceptron, is the feature after average pooling in the channel feature amplifier, It is the feature after maximum pooling in the channel feature amplifier.

5. The method for cross-dataset costume-changing pedestrian re-identification based on test-time adaptive matching according to claim 4, characterized in that: The spatial characteristic amplifier is expressed by the following formula: F s (Q)=σ(f 7×7 (AvgPool(Q);MaxPool(Q)))=σ(f 7×7 ([Q avg ;Q max ])) Among them, F s (Q) is the output of the spatial characteristic amplifier, f 7×7 is a convolution operation with a filter size of 7×7 and Q avg is the average pooled feature across channels in the spatial feature amplifier, Q max It is the feature after the maximum pooling across channels in the spatial feature amplifier.

6. The method for cross-dataset costume-changing pedestrian re-identification based on test-time adaptive matching according to claim 5, characterized in that: The connection layer is used to connect the output of the channel feature amplifier, the output of the spatial feature amplifier and the input of the clothing weakening module, specifically including: Performing a matrix multiplication operation on the input of the clothing weakening module and the output of the channel feature amplifier to obtain a first processing feature; Performing a matrix multiplication operation on the input of the clothing weakening module and the output of the spatial feature amplifier to obtain a second processing feature; Add the first processing feature and the second processing feature element by element to obtain a third processing feature; The input of the clothing weakening module is added element by element to the third processing feature to obtain a feature map corresponding to the input image.

7. The method for cross-dataset costume-changing pedestrian re-identification based on test-time adaptive matching according to claim 1, characterized in that: The test-time adaptive matching module includes a matching part, a part perception similarity adaptation submodule, a most similar part selection submodule and a scoring submodule; The matching part is used to search for matching corresponding blocks of feature maps corresponding to the two input images in a horizontal strip of a preset height, and obtain matching areas of the feature maps corresponding to the two input images according to the matching corresponding blocks; The partial perception similarity adaptation submodule is used to reshape a corresponding block of the matching area of ​​the feature map corresponding to the two input images into a column vector, and perform convolution calculation on the two column vectors, and finally repeat the operation on all blocks in the matching area to obtain two similarity matrices; The most similar part selection submodule is used to determine the most similar area of ​​each block according to the maximum similarity value, and generate two similarity vectors; The scoring submodule is used to calculate the predicted matching score according to the two similarity vectors; wherein the scoring submodule includes a concatenation layer, a first batch normalization layer, a fully connected layer, a summation layer, a second batch normalization layer and an activation function layer connected in sequence.

8. The method for cross-dataset costume-changing pedestrian re-identification based on test-time adaptive matching according to claim 1, characterized in that: The total loss function is expressed as follows: THE total =L ID +L C +λL ver +γL CA Among them, L total is the total loss function, L ID To identify the loss, L C is the clothing classification loss, L ver is the predicted matching score loss, L CA is the clothing-independent adversarial loss, λ is the weight of the verification loss for learning domain-invariant representations, γ is the weight of learning clothing-independent features in unseen target domains, and y i is the identity label, N is the batch size, P(y i |z i ) is the identity classifier for the i-th image from the feature map Z i The predicted probability, B is the number of images, q ij is the judgment parameter, s ij is the predicted matching score between the i-th image and the j-th image, N C is the number of clothing categories in the training set, For the map z i The predicted probability of belonging to the i-th clothing item, For clothing labels, f i is the normalized feature vector of the sample, is the normalized weight vector of the c-th type of clothing, τ is the temperature parameter of the smooth distribution, and q(c) is the weight function.

9. A cross-dataset costume-changing pedestrian re-identification system based on test-time adaptive matching, characterized in that: include: A preprocessing module, used for preprocessing a plurality of given pedestrian images respectively to obtain a preprocessed image set corresponding to the plurality of pedestrian images; A feature extraction module is used to input images in a plurality of preprocessed image sets into the backbone network to obtain a feature map corresponding to the input image; The scoring module is used to input the feature map corresponding to the input image into the adaptive matching module during testing to obtain a predicted matching score; A training module, used to train the backbone network and the test-time adaptive matching module based on the predicted matching score and a preset total loss function to obtain the optimal backbone network and the test-time adaptive matching module; The recognition module is used to recognize multiple target images based on the optimal backbone network and the adaptive matching module during testing, and complete cross-set pedestrian re-identification.