Cross-Domain Vehicle Re-Identification Method Based on Domain Adaptation

By performing parameter weighted averaging and pseudo-label generation on three Net network models, combined with the cyclic learning framework, the problems of vehicle re-identification and pseudo-label noise are solved, and the recognition accuracy is improved.

CN114091510BActive Publication Date: 2025-07-01NANJING UNIV OF INFORMATION SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111090479.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-17
Publication Date
2025-07-01
Estimated Expiration
2041-09-17

AI Technical Summary

Technical Problem

In the prior art, the vehicle image feature distribution in the target domain and the source domain is inconsistent, resulting in low vehicle re-identification accuracy, and the unsupervised field adaptive methods are susceptible to pseudo-label noise interference, resulting in training crashes.

Method used

Three Net network models are used for parameter weighted averaging, soft pseudo-labels and hard pseudo-labels are generated, combined with the circular average learning framework, and trained through soft classification loss functions and hard classification loss functions, and style conversion and feature extraction are used for SNR modules to reduce the influence of domain differences and pseudo-label noise.

Benefits of technology

It improves the accuracy of vehicle re-identification, reduces the interference of pseudo-label noise on the model, enhances the model's adaptability in the target domain, and improves the recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114091510B_ABST
    Figure CN114091510B_ABST
Patent Text Reader

Abstract

Cross-domain vehicle re-identification method based on domain adaptation. The present invention discloses that pictures are input into three pre-trained Net network models, and the parameters of the three Net networks are weighted and averaged to obtain three Mean-Net network models; the predicted values output by inputting source domain pictures into the Mean-Net network model are marked as soft pseudo-labels; the target domain pictures are obtained with hard pseudo-labels through a clustering algorithm; the soft classification loss functions and hard classification loss functions of the three Net network models are calculated through the soft pseudo-labels and hard pseudo-labels respectively; a cyclic average learning framework is constructed through the soft classification loss functions and hard classification loss functions corresponding to the three Net network models, namely Net1, Net2, and Net3; vehicle data pictures are input into the constructed cyclic average learning framework to output pictures of the same vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a cross-domain vehicle re-identification method based on domain adaptation, belonging to the technical field of image processing. Background Art

[0002] Vehicle re-identification refers to the technology of using computer vision technology to determine whether a specific vehicle exists in an image or video sequence. This technology plays an important role in video surveillance, intelligent transportation, maintaining social security, etc. However, there are still great challenges in vehicle re-identification in actual traffic scenarios. For example, the vehicle picture dataset used in the source domain during training is different from the vehicle picture dataset used in the target domain during testing, and the vehicle picture features in the target domain and the source domain often have inconsistent distributions and differences. In addition, the appearance of the same vehicle varies significantly under different monitoring perspectives, and the appearances of different vehicles are sometimes similar, which brings great challenges to vehicle re-identification.

[0003] Unsupervised domain adaptation methods can apply a model trained in one traffic scenario to a new scenario, which can solve the problem of cross-domain vehicle re-identification to a certain extent. This method can be roughly divided into two categories: style transfer-based methods and pseudo-label-based methods. The former uses methods such as generative adversarial networks to achieve the style transfer between different vehicle datasets to improve the application ability of vehicle re-identification technology in different scenarios. However, some useful feature information is often lost during the style transfer process. The latter proposes an adaptive module to generate "pseudo-target images". This module learns the style of the unlabeled domain and retains the identity information of the labeled domain. Using a dynamic sampling strategy, reliable pseudo-label data is selected from the clustering results, and the "pseudo-target images" and reliable pseudo-label data are combined for training the re-identification model to adapt to the changes in the new domain dataset. However, the training of the model is often interfered by pseudo-label noise, and in severe cases, it will lead to training collapse. The noise of the pseudo-labels mainly comes from the unknown number of target categories, the weak performance of the network pre-trained in the source domain in the target domain, and the low probability of the clustering algorithm itself generating correct labels, etc. Summary of the Invention

[0004] The purpose of the present invention is to provide a cross-domain vehicle re-identification method based on domain adaptation to solve the defect that the vehicle picture features in the existing target domain and source domain often have inconsistent distributions and differences, resulting in low re-identification accuracy.

[0005] A cross-domain vehicle re-identification method based on domain adaptation, the method includes:

[0006] Input the pictures into three pre-trained Net network models, and perform weighted averaging on the parameters of the three Net networks to obtain three Mean-Net network models;

[0007] The predicted values output from the Mean-Net network model for the source domain images are marked as soft pseudo-labels;

[0008] The hard pseudo-labels are obtained for the target domain images through a clustering algorithm;

[0009] The soft classification loss function and the hard classification loss function of the three Net network models are calculated respectively through the soft pseudo-labels and the hard pseudo-labels;

[0010] A cyclic average learning framework is constructed through the soft classification loss function and the hard classification loss function corresponding to the three Net network models, namely Net1, Net2, and Net3;

[0011] The vehicle data images are input into the constructed cyclic average learning framework to output the same vehicle images.

[0012] Furthermore, the training method of the Net network model includes:

[0013] The transformer is used as the feature extraction network, and the pre-trained SNR module and multi-classifier are inserted into the feature extraction network to construct the Net network model.

[0014] Furthermore, the training steps of the SNR module:

[0015] The input source domain images are instance-normalized, then the difference is calculated between the normalized feature map and the original feature map, and the information related to identity is extracted through a mask in the difference and the extracted information is added to the instance-normalized feature map;

[0016] The feature map obtained by adding the instance-normalized feature map and the extracted information feature map is further processed through group normalization to obtain the final feature map.

[0017] Furthermore, the expression of instance normalization is as follows:

[0018]

[0019] Among them, F is the input feature, μ is the calculation of the average value for each sample channel, σ is the standard deviation calculation in the same way, γ and β are the parameters learned from the samples, and S i represents the pixel set used for calculating the mean and variance.

[0020] Furthermore, the expression for calculating the difference between the normalized feature map and the original feature map is as follows:

[0021]

[0022] R is the difference between the original feature and the normalized feature, and then R is divided into two parts (R + and R -Among them, R + is identity-related feature information, and R - is information unrelated to features, and each part is restored differently according to the differences; then the extracted identity-related feature information R + is added to obtain the final output of the module;

[0023]

[0024] Furthermore, on the premise of style normalization, identity-related feature information is further fully extracted. Finally, group normalization is used to reduce the impact of batch size on training;

[0025]

[0026] Among them, both γ and β are parameters learned from samples, and G is a hyperparameter representing the size of the group.

[0027] Furthermore, the calculation of hard pseudo-labels includes:

[0028] Select a clustering method based on probability density to cluster the target domain samples and determine the initial pseudo-labels of the target domain samples;

[0029] Use the initial pseudo-labels to supervise the network to learn on the target domain, and extract high-confidence samples from the common feature distribution, thereby extracting labeled samples,

[0030] The labeled samples use a clustering algorithm to generate initial hard pseudo-labels.

[0031] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: The present invention performs style conversion through a network model, and outputs pseudo-labels through the Mean-Net network model, which not only reduces the style difference between the source domain and the target domain but also considers samples with high confidence in the target domain and supplements the data distribution; the probability soft labels in the cyclic average learning framework model effectively reduce the impact of pseudo-label noise on vehicle re-identification and improve the vehicle identification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is the flowchart of the present invention for domain adaptive cross-domain vehicle re-identification;

[0033] Figure 2 is the SNR module of the present invention;

[0034] Figure 3 is the feature extraction network structure of NET1, NET2, and NET3 of the present invention;

[0035] Figure 4 is the cyclic average learning framework of the present invention. Detailed implementation manners

[0036] To make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific implementation manners.

[0037] As shown in the figure, a cross-domain vehicle re-identification method based on domain adaptation is disclosed, which is characterized in that the method includes:

[0038] Input the picture into three pre-trained Net network models, and perform weighted averaging on the parameters of the three Net networks to obtain three Mean-Net network models;

[0039] Mark the predicted value output by inputting the source domain picture into the Mean-Net network model as a soft pseudo-label;

[0040] Obtain a hard pseudo-label for the target domain picture through a clustering algorithm;

[0041] Calculate the soft classification loss function and the hard classification loss function of the three Net network models respectively through the soft pseudo-label and the hard pseudo-label;

[0042] Construct a cyclic average learning framework through the soft classification loss function and the hard classification loss function corresponding to the three Net network models, namely Net1, Net2 and Net3;

[0043] Input the vehicle data picture into the constructed cyclic average learning framework to output the same vehicle picture.

[0044] The specific steps are as follows:

[0045] Construct an SNR module so that it normalizes the style of the input source domain picture, such as illumination, hue, color contrast and saturation, to the style of the target domain picture. Inserting the SNR module after the block layer of the feature extraction backbone network (swin-transformer) can transform the style, which is very convenient to plug and play.

[0046] Construction of the SNR module

[0047] Step 1: First, build the style normalization part. During the process of inputting the image into the feature extraction network with swin-transformer as the backbone, perform instance normalization (IN) on the output after feature extraction, and take out each HW (H and W are the length and width of the feature map respectively) for separate standardization, compression and translation, so that the style of the source domain picture is basically the same as that of the target domain picture and the independence of each image is maintained, thereby reducing the domain difference of features; the specific mathematical expression is as follows:

[0048]

[0049] Among them, F is the input feature, μ is the calculation of the average value for each sample channel, and σ is the standard deviation calculation in the same way. Both γ and β are parameters learned from the samples, and S i represents the pixel set used for calculating the mean and variance.

[0050] Step 2: Secondly, build the style restoration part. The difference between the normalized feature map and the original feature map is obtained, and the difference between the two is obtained. Then, information related to identity is further extracted from the residual features. Then, the feature information related to identity is restored to the normalized feature to make the identification ability of the feature information stronger. Then, Group Normalization (GN) is used to accelerate convergence. The channels are divided into many groups, and normalization is performed on each group, avoiding the batch dimension and thus reducing the dependence on batch size; the obtained in the first part is subtracted from the original F:

[0051]

[0052] R is the difference between the original feature and the normalized feature. Then, R is divided into two parts (R + and R - ). Among them, R + is the feature information related to identity, and R - is the information unrelated to the feature, and each part is restored differently according to the difference. Then, the extracted feature information R + related to identity is added to to obtain the final output of the module.

[0053]

[0054] On the premise of style normalization, the feature information related to identity is further fully extracted. Finally, Group Normalization (GN) is used to reduce the impact of batch size on training.

[0055]

[0056] Among them, both γ and β are parameters learned from the samples. G is a hyperparameter representing the size of the group.

[0057] Generation of hard pseudo-labels

[0058] The initial hard labels such as [0, 1, 0, …, 0] are generated for the target domain images through the DBSCAN clustering algorithm. These hard pseudo-labels will be used in the cyclic average learning framework to supervise the learning of the network. These pseudo-labels are updated separately before each training epoch;

[0059] The clustering-based pseudo-labeling method is adopted to solve the problem that the target domain has no labels. After generating pseudo-labels through clustering, the samples with high confidence in the clustering results contribute to the network's learning on target domain images. This stage can be specifically divided into two steps:

[0060] Step 1: Initially generate hard pseudo-labels for the unlabeled target domain images using a clustering algorithm. Considering that the target domain samples have no class labels beforehand, the density-based spatial clustering of applications with noise (DBSCAN) is selected to cluster the target domain samples and determine the initial labels of the target domain samples.

[0061] Step 2: Use the initial pseudo-labels to supervise the network's learning on the target domain. Extract some high-confidence samples from the common feature distribution, thereby extracting labeled samples. After sample merging, it can play a certain role in supplementing the data distribution situation and improving the model's distribution learning ability.

[0062] Establish a cyclic average learning framework

[0063] Step 1: Pre-train three networks with swin-transformer as the backbone on ImageNet. To enhance the complementarity of the outputs of the cyclic network, during the pre-training process, we will take the following measures:

[0064] · Adopt different random augmentation methods for the images input to the three networks, such as random cropping, random erasing, random inversion, etc.;

[0065] · Use different initialization parameters for Net1, Net2, and Net3;

[0066] · Use different soft pseudo-labels for supervision when training NET1-3. The soft supervision of NET1 comes from the average model of NET3, the soft supervision of NET2 comes from the average model of NET1, and the soft supervision of NET3 comes from the average model of NET1

[0067] Step 2: The E[θ] of the network average model is the average value of the corresponding network parameter θ, that is, the parameter update of the average model does not rely on backpropagation, but is updated using exponential weighted average after each backpropagation of the loss function. The update formula of the network parameters is as follows:

[0068] E (T) [θ k =βE (T-1) [θ k +(1-β)θ k (k=1,2,3) (5)

[0069] where (T) refers to the Tth iteration, θ k Represents the current parameters of the kth network. Due to the accumulation of time, the three average models have stronger decoupling and more independent and complementary outputs, thereby alleviating the noise problem in the process of generating pseudo labels.

[0070] Step 3: The vehicle re-identification network achieves training effect through the convergence of the loss function. In the invention, F(·|θ) is used to represent the encoder, C is used to represent the classifier, and each NET consists of an encoder and a classifier. We use subscripts 1, 2, and 3 to distinguish NET1, NET2, and NET3, and use subscripts s and t to distinguish the source domain from the target domain. The triplet loss function is used to shorten the distance between similar features and increase the distance between different features; the hard triplet loss function is:

[0071]

[0072] However, the triplet loss cannot support the training of probabilistic soft labels. The loss relationship between the features in the triplet corresponding to the soft pseudo labels of the present invention is expressed as:

[0073]

[0074] The classification loss uses the general multi-classification cross entropy loss function l ce To express:

[0075]

[0076] in, For the target domain image The hard pseudo labels are generated by clustering.

[0077] Unlike hard false labels, the soft false labels in the soft classification loss are the classification prediction values ​​of the average model. In order to reduce the distance between the two distributions, the soft false label classification loss is as follows:

[0078]

[0079] Testing Phase

[0080] During testing, remove the data augmentation used in the training phase, such as image inversion, cropping, etc. Input the test vehicle data image into the trained vehicle re-identification model, and then the network calculates the similarity between the vehicle in the input image and the vehicle in the gallery image, and the final softmax layer of the network outputs the probability of the same vehicle.

[0081]

[0082] where zi is the output of the last fully-connected layer of the neural network, a i represents the i-th output of softmax, that is, the similarity between the i-th picture and the input picture. Thus, the neural network model identifies all vehicle pictures with high similarity to it.

[0083] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A cross-domain vehicle re-identification method based on domain adaptation, characterized in that The method includes: Input the pictures into three pre-trained Net network models, and perform weighted averaging on the parameters of the three Net networks to obtain three Mean-Net network models; Mark the predicted values output by inputting the source domain pictures into the Mean-Net network model as soft pseudo-labels; Obtain hard pseudo-labels for the target domain pictures through a clustering algorithm; Calculate the soft classification loss function and the hard classification loss function of the three Net network models through the soft pseudo-labels and the hard pseudo-labels respectively; Construct a cyclic average learning framework through the soft classification loss function and the hard classification loss function corresponding to the three Net network models, namely Net1, Net2, and Net3; Input the vehicle data pictures into the cyclic average learning framework to output the same vehicle pictures; The training method of the Net network model includes: Use swin-transformer as the feature extraction network, and insert the pre-trained SNR module and multi-classifier into the feature extraction network to construct the Net network model; Training steps of the SNR module: Input the source domain pictures, perform instance normalization, then calculate the difference between the instance-normalized feature map and the original feature map, extract identity-related information from the difference through a mask, and add the extracted information to the instance-normalized feature map; The feature map obtained by adding the instance-normalized feature map and the extracted information feature map is further processed by group normalization to obtain the final feature map.

2. The cross-domain vehicle re-identification method based on domain adaptation according to claim 1, wherein, The expression of instance normalization is as follows: (1) Among them, F is the input feature, is the calculation of the average value for each sample channel, Similarly, it is the calculation of the standard deviation, and are the parameters learned from the samples, represents the pixel set used for calculating the mean and variance.

3. The cross-domain vehicle re-identification method based on domain adaptation according to claim 1, characterized in that The expression for calculating the difference between the instance-normalized feature map and the original feature map is as follows: (2) R is the difference between the original feature and the instance-normalized feature, and then R is divided into two parts and , where is the identity-related feature information is the information unrelated to the feature, and each part is restored differently according to the difference; then the extracted identity-related feature information is added to to obtain the final output of the module (3)。 4. The cross-domain vehicle re-identification method based on domain adaptation according to claim 1, characterized in that, On the premise of style normalization, further extract identity-related feature information, and finally reduce the impact of batch size on training through group normalization; (4) Among them, and are both parameters learned from the samples. G is a hyperparameter representing the size of the group.

5. The cross-domain vehicle re-identification method based on domain adaptation according to claim 1, characterized in that The calculation of hard pseudo-labels includes: Select a clustering method based on probability density to cluster the target domain samples to determine the initial pseudo-labels of the target domain samples; Use the initial pseudo-labels to supervise the network to learn on the target domain, extract high-confidence samples from the common feature distribution, thereby extracting label samples, and the label samples use a clustering algorithm to generate initial hard pseudo-labels.

Citation Information

Patent Citations

  • Unsupervised domain adaptive pedestrian re-recognition algorithm based on pseudo label optimization

    CN113378632A