Continuous self-evolving target re-identification modeling method based on parameter correction
By adopting parameter correction and dynamic convolution network knowledge distillation technology in the target recognition model, combined with the classifier linear fusion strategy, the problem of "catastrophic forgetting" of the model when learning in the new data domain is solved, and the model's continuous learning ability and generalization ability in the distribution of multiple data domains is improved.
Patent Information
- Application Number
- CN202410788114.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-19
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-06-19
AI Technical Summary
Existing target re-identification models are prone to ‘catastrophic forgetting’ when facing new data domains, that is, when learning new data sets, the model will significantly reduce its performance on the previous data sets.
The continuous self-evolution target re-identification modeling method based on parameter correction is adopted. The model parameters of the new task are corrected by using the model trained on the old task, and combined with the parameter generator and classifier linear fusion strategy of the dynamic convolution network to constrain and correct model parameter updates.
It effectively curbed the catastrophic forgetting phenomenon that the target re-identification model has continuously learned the distribution of multiple data domains, allowing the model to retain memory of the feature distribution of the old data domain while adapting to the distribution of the new data domain while expanding the model's continuous learning ability.
Smart Images

Figure CN118658181B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a continuous self-evolving target re-identification modeling method. Background Art
[0002] The object re-identification model aims to retrieve images with the same identity as the object in a large-scale image library after a given object is queried. The performance of the object re-identification model is affected by many factors, such as the lighting conditions of the image, the resolution, and the object's clothing. A considerable amount of work has been conducted to study the above problems and enable the object re-identification model to perform well on a single data domain.
[0003] However, the distribution of data domains that the model encounters in real scenarios is inconsistent with what it has learned before. Therefore, it is necessary to improve the generalization ability of the target re-identification model so that it can perform better in unseen data domains.
[0004] In real-world application scenarios, the data that the object re-identification model needs to learn may not be available all at once, so the model is often trained online, that is, the data set arrives in batches at different time points. In this context, when the model learns the newly arrived data set, due to the change of its own parameters, the performance of the model on the previous data set will drop significantly. This phenomenon is called "catastrophic forgetting."
[0005] Therefore, it is necessary to constrain and correct the parameter updates of the model when it learns on the new data domain, so that it can adapt to the distribution characteristics of the new data domain while retaining the memory of previous tasks as much as possible to avoid "catastrophic forgetting". Summary of the invention
[0006] In order to overcome the shortcomings of the prior art, the present invention provides a method for continuous self-evolution target re-identification modeling based on parameter correction. The present invention discloses a method for continuous self-evolution target re-identification modeling based on parameter correction, which uses a model trained on an old task to correct the model parameters of the current new task, uses the output of the parameter generator of the dynamic convolutional network in the target re-identification model as the constraint indicator of the two models, and uses the output difference of the parameter generators of the two models as a loss function item, thereby realizing parameter constraint. Afterwards, at the end of the feature extraction backbone network, a classifier fusion strategy is adopted to further retain the knowledge in the previous model, thereby realizing the continuous autonomous evolution of the target re-identification model.
[0007] The technical solution adopted by the present invention to solve its technical problem is:
[0008] In continuous learning, the process of learning a new data set is divided into the following steps:
[0009] Step 1: First, freeze the model obtained on the previously learned data set as the original model to obtain the model parameters Frozen_Model. The model structure of Frozen_Model is consistent with the training model structure, but the model weight parameters of Frozen_Model are obtained based on the old training data and are used for parameter correction in the current model training.
[0010] Step 2: Obtain the image samples of the current data domain, copy the model structure and parameters of Frozen_Model as a model participating in the training of the new data set, recorded as Dynamic_Model, and pass the image samples of the current data domain into Frozen_Model and the current training model Dynamic_Model respectively. After the sample enters Dynamic_Model, it is passed into a 3×3 convolution layer to extract features, and then the preprocessing operations of normalization, activation and pooling are completed in Dynamic_Model before being passed into the dynamic normalization processing stage; in the dynamic normalization processing stage, the input sample features are processed by instance normalization, that is, normalization is performed on each channel of each sample, and the feature vector processed by instance normalization is passed into the dynamic convolution network together with the input sample feature vector for processing, while retaining part of the original feature pattern, the feature normalization operation is completed;
[0011] Step 3: Calculate KL divergence;
[0012] Step 4: Calculate the loss function of the Dynamic_Model model;
[0013] Step 5: Classifier fusion strategy;
[0014] In order to keep the memory of the old task, the classifier part in the Frozen_Model model is linearly combined with the classifier parameters of the Dynamic_Model model to obtain the classifier used on the current task. The specific implementation method is:
[0015] H=θ 1 H Frozen +θ 2 H Dynamic
[0016] Where H is the final classifier parameter, H Frozen and H Dynamic are the classifier parameters of the frozen model and the dynamically trained model, θ 1 and θ 2 They are the weights corresponding to the Frozen_Model model classifier and the Dynamic_Model model classifier respectively;
[0017] Step 6: Obtain the complete target re-identification model;
[0018] Use the classifier obtained in step 5 to replace the classifier part in Dynamic_Model to obtain a complete target re-identification model for subsequent target re-identification tasks.
[0019] In step 2, there is a parameter prediction network in the dynamic homogenization processing stage, which is used to generate corresponding weight parameters and bias for the dynamic convolution network; it is expressed as:
[0020] F OUT =DyConv(IN(F IN ), W, B)
[0021] {W, B} = FC(ReLU(Pooling(F IN )),θ)
[0022] Among them, F OUT is the final output feature after dynamic uniformization processing, DyConv is the dynamic convolution operation function, IN is the instance uniformization operation function, and F IN is the input image feature after the preprocessing stage, {W, B} is the weight parameter and bias of the dynamic convolutional network, the FC layer is the fully connected layer, ReLU is the activation function, Pooling is the pooling operation function, and θ is the fully connected layer coefficient.
[0023] In step 3, the steps of calculating KL divergence are:
[0024] Extract the weight parameters and bias of the dynamic convolutional network in Frozen_Model and Dynamic_Model respectively, and then calculate the KL divergence of the network weight parameters and the KL divergence of the bias as the standard for measuring the parameter and performance differences between the models, and use it for the subsequent parameter correction; the KL divergence of the network weight parameters KL Weight KL divergence with bias Bias The calculation is as follows:
[0025] KL Weight (W Frozen ||W Dynamic )=∑F(W Frozen )log F(W Frozen )F(W Dynamic )
[0026] KL Bias (B Frozen ||B Dynamic )=∑F(B Frozen )log F(B Frozen)F(B Dynamic )
[0027] Among them, W Frozen With B Frozen are the weight parameters and bias of the dynamic convolutional network in Frozen_Model, W Dynamic With B Dynamic They are the weight parameters and bias of the dynamic convolutional network in Dynamic_Model, and F is the size normalization function for the network parameters;
[0028] The steps of calculating the loss function of the Dynamic_Mode1 model in step 4 are:
[0029] The overall loss function of the Dynamic_Model model includes two items. The first item is the target re-identification model loss function, including cross entropy loss and triple loss, which is used to measure the processing ability of Dynamic_Model for the current data task; the second item is KL divergence, which is used to measure the memory ability of Dynamic_Model for old tasks; therefore, the overall loss function Ltota l It is expressed as follows:
[0030]
[0031] where x i is the input sample data, y i is the corresponding sample true label, x po i is a randomly selected positive sample, x ne i is a randomly selected negative sample, is the classification function, ψ is the feature extraction function, L c.e. is the cross entropy loss function, L tri. is the triple loss function, λ 1 With λ 2 are the weight coefficients of the two groups of KL divergence respectively.
[0032] Before the continuous learning and training of the model, the target re-identification model is pre-trained on the ImageNet dataset, and the weight parameters of the pre-trained model are used as Frozen_Model in subsequent training.
[0033] Four datasets, VIPeR, Market-1501, CUHK-SYSU and MSMT17, are used to construct continuous learning training sets respectively. The model learns the four datasets in turn. After training on each dataset, new pre-training model weight parameters are obtained, and then training is performed on the new dataset. After completing the continuous learning training, the obtained model performs well on multiple continuously learned target re-identification datasets.
[0034] An electronic device includes: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be implemented by the method.
[0035] A computer-readable storage medium stores program codes, and the program codes can be called by a processor to execute the method as described above.
[0036] The beneficial effect of the present invention is that it utilizes dynamic convolution knowledge distillation technology and classifier linear fusion technology to curb the catastrophic forgetting phenomenon that occurs in the target re-identification model during the continuous learning of multiple data domain distributions. The parameter correction method enables the model to reduce the forgetting of the feature distribution of the old data domain while learning the feature distribution of the new data domain, thereby expanding the continuous learning ability of the target re-identification model with domain generalization capability. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a model framework structure diagram of the present invention. DETAILED DESCRIPTION
[0038] The present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0039] In order to expand the continuous learning ability of the target re-identification model in the online learning process and avoid catastrophic forgetting, this paper proposes a continuous self-evolution target re-identification model based on parameter correction. The specific structure is as follows: Figure 1 shown.
[0040] 1. Dynamic Convolutional Network Knowledge Distillation
[0041] The dynamic convolution structure serves as a regulatory network for correcting the uniformity of instance forces. It is responsible for fusing the normalized feature vector with the feature vector directly extracted by the feature extraction backbone network, thereby enhancing the generalization ability of the model while retaining the sensitivity of the feature extraction backbone network to certain instances or features.
[0042] The implementation of the dynamic convolutional network depends on the existence of a parameter prediction network, which takes the extracted original feature vector as input and outputs the weights and biases of the dynamic convolutional network. The dynamic convolutional network replaces several traditional convolutional structures in the backbone network. Therefore, in order to pass the feature extraction capabilities stored in the frozen model to the dynamic model currently being trained, the parameter prediction network in the frozen model can be used as a teacher model to distill its own ability to generate dynamic convolutional structure weights and bias parameters into the dynamic training model.
[0043] Based on this, the relative entropy between the weight parameters and bias parameters of the two sets of dynamic convolution structures is used as the loss function term of the knowledge distillation process. At the same time, this loss function term is an item in the overall model loss function that characterizes the model memory stability. The calculation method of the loss function term is as follows:
[0044] KL Weight (W Frozen ||W Dynamic )=∑F(W Frozen )log F(W Frozen )F(W Dynamic ) (5)
[0045] KL Bias (B Frozen ||B Dynamic )=∑F(B Frozen )log F(B Frozen )F(B Dynamic )
[0046] Among them, W Frozen With B Frozen are the weight parameters and bias of the dynamic convolutional network in Frozen_Model, W Dynamic With B Dynamic They are the weight parameters and bias of the dynamic convolutional network in Dynamic_M0del, and F is the data processing function for the network parameters.
[0047] By taking the difference between the parameters of the current dynamic training model and the parameters of the frozen model as the loss function term, the update strategy of the dynamic convolution structure parameters can be effectively changed so that it can adapt to the new data distribution while retaining the ability to process feature vectors in the old data domain. This can effectively perform instance homogenization operations of appropriate scale on the current data, avoiding the risk of catastrophic forgetting.
[0048] In addition to the stability term, the model's loss function also includes an adaptability term for learning the current data domain distribution. In order to effectively improve the model classification effect and improve the metric learning effect, the cross entropy loss function and the triple loss function are used to measure the model's adaptability to the new data domain distribution. The overall loss function of the model is calculated as follows:
[0049]
[0050] where x i is the input sample data, y i is the corresponding sample true label, x po i is a randomly selected positive sample, x ne i is a randomly selected negative sample, is the classification function, ψ is the feature extraction function, L c.e is the cross entropy loss function, L tri. is the triple loss function, λ 1 With λ 2 is the weight coefficient of the two groups of KL divergence.
[0051] 2. Linear fusion of classifiers
[0052] The knowledge distillation of the dynamic convolutional network is responsible for forwarding the knowledge contained in the model feature extraction backbone network, which is directly related to the image features and belongs to the front network with a lower level of abstraction in the neural network structure. The classifier network, as a part with a higher level of abstraction, has the characteristic of being sample-independent compared to the feature extraction backbone network. Therefore, the forward transfer of knowledge of the classifier part does not require the process of knowledge distillation. The present invention adopts the classifier linear fusion method to further ensure that the model's rear network retains its memory ability in the previous data domain.
[0053] Since the classifier faces a task of classifying image features at a higher level of abstraction, the difference in parameter distribution is less affected by the difference between different training domains. In order to keep the memory of the old task, linear parameter fine-tuning can be performed on the basis of the frozen classifier to obtain the classifier used in the current task. The specific implementation method is:
[0054] H=θ 1 H Frozen +θ 2 H Dynamic (7)
[0055] Where H is the final classifier parameter, H Frozen and H Dynamic are the classifier parameters of the frozen model and the dynamically trained model, θ 1 and θ 2are the corresponding weights respectively. Specific embodiment:
[0057] 1. Dataset selection
[0058] In order to evaluate the performance of the proposed object re-identification network model with continuous learning ability in continuous learning tasks, the model is evaluated on four continuous datasets on challenging benchmarks: VIPeR, Market1501, CUHK-SYSU and MSMT17. VIPeR contains images from 632 identities and two cameras, with two images for each identity; Market1501 contains pedestrian images from 1501 identities, including images from multiple cameras; CUHK-SYSU contains pedestrian images from 554 identities with complex backgrounds and perspective changes; MSMT17 contains pedestrian images from 1261 identities with multiple cameras, complex backgrounds and large scale changes.
[0059] 2. Implementation details setting
[0060] As for pre-training, before starting online training, the object re-identification network model is pre-trained on the ImageNet dataset and used for training. The ImageNet dataset contains a large amount of general image data, which helps to give the initial model better generalization ability.
[0061] In terms of data augmentation, the image sizes used in training are adjusted to 265×128, and strategies such as image rotation and random erasing are used to further improve the diversity of image samples in the dataset.
[0062] As for the hyperparameter settings, since it has been pre-trained on the ImageNet dataset, the initial learning rate is set to 3.5×10 -5 , and then decreases by 0.1 after the 40th training cycle of each task. Subsequent tasks are initialized to 3.5×10 -6 , for the last task it is initialized to 1.75×10 -6 The batch size of each training task is set to 128.
[0063] 3. Implementation environment
[0064] The present invention is deployed in Sugon-W580-G20 server and Linux Ubuntu 18.04.4LTS operating system, using two NVIDIA GeForce GTX 1080Ti graphics processors for training, and using NVIDIA CUDA10.2 platform to accelerate training. The software environment uses Python3.7.6, PyTorch1.6.0, Torchvision0.7.0 and other dependent libraries.
[0065] 4. Model application
[0066] The present invention deploys the initial model to the actual application scenario, continuously accepts new scenes and image data distributed in the data domain, uses its own dynamic convolutional feature extraction backbone network to extract features, and combines the frozen model to generate parameters according to the new input to correct the parameter update of the dynamic model.
[0067] The classifier accepts the extracted features as input, and during training, the parameters of the dynamic model and the frozen model are fused as the final classifier to classify the results. The classification process calculates the Euclidean distance between the extracted image features and the query image features, sorts the results, and selects the closest image as the output result. If the label of the output result is consistent with the query image, it means that the query is successful.
Claims
1. A continuous self-evolution target re-identification modeling method based on parameter correction, characterized in that The steps include: Step 1: First, freeze the model obtained on the previously learned data set as the original model to obtain the model Frozen_Model. The model structure of Frozen_Model is consistent with the training model structure, but the model weight parameters of Frozen_Model are obtained based on the old training data and are used for parameter correction in the current model training; Step 2: Obtain the image samples of the current data domain, copy the model structure and parameters of Frozen_Model as a model participating in the training of the new data set, recorded as Dynamic_Model, and pass the image samples of the current data domain into Frozen_Model and the current training model Dynamic_Model respectively. After the sample enters Dynamic_Model, it is passed into a 3×3 convolution layer to extract features, and then the preprocessing operations of normalization, activation and pooling are completed in Dynamic_Model before being passed into the dynamic normalization processing stage; in the dynamic normalization processing stage, the input sample features are processed by instance normalization, that is, normalization is performed on each channel of each sample, and the feature vector processed by instance normalization is passed into the dynamic convolution network together with the input sample feature vector for processing, while retaining part of the original feature pattern, the feature normalization operation is completed; Step 3: Calculate KL divergence; in step 3, the steps of calculating KL divergence are: Extract the weight parameters and bias of the dynamic convolutional network in Frozen_Model and Dynamic_Model respectively, and then calculate the KL divergence of the network weight parameters and the KL divergence of the bias as the standard for measuring the parameter and performance differences between the models, and use it for the subsequent parameter correction; the KL divergence of the network weight parameters KL Weight KL divergence with bias Bias The calculation is as follows: KL Weight (W Frozen |||W Dynamic )=∑F(W Frozen )logF(W Frozen )F(W Dynamic ) (2) KL Bias (B Frozen ||B Dynamic )=∑F(B Frozen )logF(B Frozen )F(B Dynamic ) Among them, W Frozen With B Frozen are the weight parameters and bias of the dynamic convolutional network in Frozen_Model, W Dynamic With B Dynamic They are the weight parameters and bias of the dynamic convolutional network in Dynamic_Model, and F is the size normalization function for the network parameters; Step 4: Calculate the loss function of the Dynamic_Model model; The steps of calculating the loss function of the Dynamic_Model model in step 4 are: The overall loss function of the Dynamic_Model model includes two items. The first item is the target re-identification model loss function, including cross entropy loss and triple loss, which is used to measure the processing ability of Dynamic_Model for the current data task; the second item is KL divergence, which is used to measure the memory ability of Dynamic_Model for old tasks; therefore, the overall loss function L total It is expressed as follows: where x i is the input sample data, y i is the corresponding sample true label, x po i is a randomly selected positive sample, x ne i is a randomly selected negative sample, is the classification function, ψ is the feature extraction function, L c.e. is the cross entropy loss function, L tri. is the triple loss function, λ1 and λ2 are the weight coefficients of the two groups of KL divergence respectively; Step 5: Classifier fusion strategy; In order to keep the memory of the old task, the classifier part in the Frozen_Model model is linearly combined with the classifier parameters of the Dynamic_Model model to obtain the classifier used on the current task. The specific implementation method is: H=θ1H Frozen +θ2H Dynamic (4) Where H is the final classifier parameter, H Frozen and H Dynamic are the classifier parameters of the frozen model and the dynamic training model, respectively, θ1 and θ2 are the weights of the corresponding Frozen_Model model classifier and Dynamic_Model model classifier; Step 6: Obtain the complete target re-identification model; Use the classifier obtained in step 5 to replace the classifier part in Dynamic_Model to obtain a complete target re-identification model for subsequent target re-identification tasks.
2. The method for continuous self-evolution target re-identification modeling based on parameter correction according to claim 1, characterized in that: In step 2, there is a parameter prediction network in the dynamic homogenization processing stage, which is used to generate corresponding weight parameters and bias for the dynamic convolution network; it is expressed as: F OUT =DyConv(IN(F IN ),W,B) (1) {W,B}=FC(ReLU(Pooling(F IN )),θ) Among them, F OUT is the final output feature after dynamic uniformization processing, DyConv is the dynamic convolution operation function, IN is the instance uniformization operation function, and F IN is the input image feature after the preprocessing stage, {W, B} is the weight parameter and bias of the dynamic convolutional network, the FC layer is the fully connected layer, ReLU is the activation function, Pooling is the pooling operation function, and θ is the fully connected layer coefficient.
3. The method for continuous self-evolution target re-identification modeling based on parameter correction according to claim 1, characterized in that: Before the continuous learning and training of the model, the target re-identification model is pre-trained on the ImageNet dataset, and the weight parameters of the pre-trained model are used as Frozen_Model in subsequent training.
4. The method for continuous self-evolution target re-identification modeling based on parameter correction according to claim 1, characterized in that: Four datasets, VIPeR, Market-1501, CUHK-SYSU and MSMT17, are used to construct continuous learning training sets respectively. The model learns the four datasets in turn. After training on each dataset, new pre-training model weight parameters are obtained, and then training is performed on the new dataset. After completing the continuous learning training, the obtained model performs well on multiple continuously learned target re-identification datasets.
5. An electronic device, characterized in that: include: one or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1 to 4.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, and the program codes can be called by a processor to execute the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Event detection model construction method and device, electronic equipment and storage medium
CN111813931A
Target re-identification model anti-forgetting training method, target re-identification method and target re-identification device
CN115439878A