Lightweight method for pedestrian re-identification in video scene based on knowledge distillation
Patent Information
- Application Number
- CN202410402201.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-03
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-04-03
AI Technical Summary
然而,由于轻量化模型通常采用简化的结构,可能难以有效地捕捉这些细粒度特征,从而导致在具有挑战性的场景下性能下降
[0043]本发明方法,克服了现有模型在行人重识别下参数量较大,训练数据要求高的不足,通过对大模型进行蒸馏学习的方法,在保持较高性能的同时,显著减少了模型的参数量,使得在资源受限的环境下仍具有较好的计算效率;
Smart Images

Figure CN118351415B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of information technology, and more specifically, relates to a lightweight method for pedestrian re-identification in video scenes based on knowledge distillation. Background Technology
[0002] Pedestrian re-identification refers to retrieving a specific person from a large number of pedestrian images captured by different cameras. It is an important sub-topic in computer vision after facial recognition. In recent years, with the improvement of surveillance equipment and the public's increased awareness of safety, more public places, especially those with high pedestrian traffic, have begun to focus on pedestrian re-identification applications. Compared with image-based pedestrian re-identification, video-based pedestrian re-identification can provide richer gait and perspective information and mitigate the negative impact of occlusion. Therefore, video-based pedestrian re-identification is attracting increasing attention from researchers.
[0003] However, video-based person re-identification methods still face many challenges. Most popular person re-identification models have a large number of parameters, limiting their applicability on resource-constrained devices. For example, to improve the performance of person feature extraction in video scenes, a Transformer backbone architecture is generally adopted, with ViT being the most widely used. One of the latest trends in ViT is to achieve higher performance on standard benchmarks while the model size continues to grow. For example, the V-MoE model, trained on 305 million images, has achieved state-of-the-art performance in image classification with 14.7 billion parameters. The Swin Transformer, using 7 billion parameters, has achieved significant results in downstream detection and segmentation tasks. Even the smaller PiT model in the field of person re-identification uses 1 billion parameters. Such a large parameter scale makes it difficult to train and deploy on terminal and edge devices with limited hardware resources, which greatly affects the application of video-based person re-identification.
[0004] To deploy models to edge devices or endpoints and run them in resource-constrained environments, model lightweighting is necessary. This typically involves reducing the number of parameters and computational complexity, which can lead to a decrease in model performance. Existing lightweight models often employ modified CNN models for person re-identification. For example, the ResNet-based HA-CNN model achieves good performance with 1.09 billion computations and 2.7M parameters. Modified CNN models achieve better performance than ResNet with fewer parameters and computational costs. However, due to the complex scenes, diverse pedestrian poses, and dynamic changes in videos, these lightweight models may not be able to fully capture this complexity, resulting in insufficient extraction of pedestrian features and affecting the accuracy of re-identification. Secondly, the limitations of lightweight models are also reflected in their insufficient capture of fine-grained features. In videos, detailed features of pedestrians, such as clothing texture and shoe style, are crucial for person re-identification. However, because lightweight models typically employ simplified structures, they may struggle to effectively capture these fine-grained features, leading to performance degradation in challenging scenarios. Therefore, the feature fusion method described above for pedestrian re-identification cannot meet the accuracy requirements of the application and is not suitable for pedestrian re-identification in video scenarios using a lightweight model.
[0005] Therefore, there is an urgent need for a new lightweight method for pedestrian re-identification in video scenarios based on knowledge distillation. Summary of the Invention
[0006] To overcome the above limitations while retaining the performance of Transformer-based person re-identification algorithms and the lightweight characteristics of CNN-based person re-identification algorithms, this invention employs knowledge distillation to perform lightweight knowledge distillation on the model. Unlike pruning and quantization in model compression, knowledge distillation involves constructing a lightweight smaller model and training it using the supervision information from a higher-performance larger model, aiming to achieve better performance and accuracy. This larger model is called the teacher model, and the smaller model is called the student model. Essentially, the lightweight student model mimics a pre-trained complex teacher network model, maintaining its lightweight characteristics while possessing the performance of a complex model, thus achieving the goal of model compression.
[0007] Knowledge distillation achieves an ideal balance between performance and lightweight design, providing an innovative solution for our pedestrian re-identification system and yielding some positive results. However, lightweight student models employing only distillation cannot fully meet the application requirements for feature extraction in video scenes. Video scenes present numerous challenges for pedestrian feature extraction, including variations in lighting, occlusion, and the influence of different viewpoints. In these complex situations, student models exhibit limitations, resulting in less than satisfactory pedestrian feature recognizability. Therefore, we adopted TinyViT as a representative lightweight student model, which excels in lightweight design. TinyViT is a lightweight model based on the ViT architecture, possessing superior parameter and computational efficiency. Its structure considers efficient pedestrian feature extraction in resource-constrained video scenes, providing feasibility for lightweight deployment. It uses PiT as a baseline model and introduces fast pre-trained distillation, achieving a good balance between computation and accuracy, and demonstrating good transferability to downstream tasks.
[0008] This invention, based on the TinyViT student model, introduces dataset augmentation during the distillation learning process. This augmentation is applied simultaneously to both the teacher and student models, fully utilizing the information in the datasets, accelerating the distillation learning process, and enhancing model reliability. Simultaneously, this invention proposes a novel MCT model to replace the attention module in ViT, including a Convolutional Attention (DCSA) module employing depthwise separable convolutions and a Reactive Alternative Convolutional Array (RASA) module based on dilated convolutions. DCSA combines convolution with attention mechanisms, significantly reducing the number of parameters while enhancing attention to local contextual information. RASA further reduces the number of parameters using dilated convolutions and enhances attention to overall information. The combination of these two convolutions enables feature fusion of pedestrian images, resulting in more efficient pedestrian feature extraction in resource-constrained video scenarios. Furthermore, to further reduce the number of parameters, this invention replaces the standard 3x3 convolution with a depthwise separable convolution in the DCSA module. After the above improvements, the lightweight method for pedestrian re-identification in video scenes based on knowledge distillation proposed in this invention can effectively improve the identifiability of pedestrian feature extraction in video scenes while maintaining the model's lightweight nature, so as to meet the application requirements of pedestrian re-identification in video scenes, and can be effectively deployed on mobile edge devices and terminals with limited hardware resources.
[0009] To address at least one of the aforementioned technical problems, according to one aspect of the present invention, a lightweight method for pedestrian re-identification in video scenes based on knowledge distillation is provided, comprising the following steps:
[0010] S1: Download and process the MARS and iLIDS-VID datasets for the video scene;
[0011] S2: Training the teacher model PiT mainly consists of four parts: dataset preprocessing, teacher model feature extraction, loss calculation, and distillation feature extraction. The specific steps are as follows:
[0012] S2.1: Before training, perform data augmentation on the input video images. Data augmentation methods include horizontal flipping, scaling, and cropping. Making the dataset as diverse as possible can improve the generalization ability of the trained re-identification model. Simultaneously, save the data augmentation results for later use.
[0013] S2.2: The multi-directional, multi-scale pyramid module in the teacher model is used to extract features from video images, enhancing the robustness of the model and obtaining key feature information of pedestrians from multiple directions.
[0014] S2.3: Calculate the loss function, train the network, and update the network parameters through backpropagation. The loss function consists of triplet loss and additive angular interval loss (ArcFace loss), resulting in a well-trained pedestrian re-identification network model.
[0015] The function is as follows:
[0016]
[0017] in, and These represent the proportions of triplet loss and additive angular interval loss, respectively. and These are the triplet loss and the additive angular spacing loss (ArcFace loss), respectively. In this invention, Take 0.4, Take 0.6.
[0018] S2.4: Extract distilled features from the teacher model. This includes extracting the model's self-attention weight matrix and feature embeddings. These features contain key information about the pedestrians to be passed to the student model.
[0019] S3: Construct a student model TinyViT with a small number of parameters and insert a convolutional attention module MCT into the model. This module includes a convolutional attention module DCSA based on depthwise separable convolution and a dilated convolution-based RASA module, etc., to extract fusion features from the video and enhance the network's accuracy in extracting image features.
[0020] S4: Distillation learning is performed on the student model, mainly consisting of four parts: model initialization and iterative optimization, calculation of distillation loss, joint optimization, and pedestrian re-identification matching. The specific steps are as follows:
[0021] S4.1: The prediction results of the teacher model (i.e., the distilled features extracted in S2.4) are used as a supervision signal to train the student model. Through multiple iterations, the student model gradually optimizes its performance, continuously approaching the knowledge representation of the teacher model. The input image used in the forward propagation process of this iterative optimization is the same as that used in S2.1.
[0022] S4.2: Adjust the temperature parameter output by Softmax as needed to ensure a smooth probability distribution.
[0023] S4.3: Calculate the distillation loss between the softmax probability distributions of the teacher and student models using the KD loss function. This involves creating an overall loss function that combines the distillation loss function with the softmax loss function used for person re-identification. An Adam optimizer is configured to jointly optimize the overall loss function, and the parameters of the student model are updated using backpropagation to minimize the overall loss. The KD loss function is as follows:
[0024]
[0025] in, and These are the weighted proportions corresponding to the distillation loss and the student loss, respectively. and These are the distillation loss function and the Softmax loss function used for pedestrian re-identification, respectively. For teacher models in The Softmax output at temperature is at the 1st Values on a class For student models in The Softmax output at temperature is at the 1st Values on a class In the first The tag value on the class, Positive labels are assigned a value of 1, and negative labels are assigned a value of 0. This represents the total number of tags.
[0026] S4.4: Perform pedestrian re-identification and matching. Extract pedestrian features from the query dataset and the matching image dataset. Calculate the similarity between each image in the query dataset and each image in the matching image dataset using the Euclidean distance between their feature vectors. Sort the images in the matching image dataset by similarity to obtain the Rank-1 hit rate and the average precision mAP. Finally, achieve the re-identification of pedestrian samples.
[0027] Furthermore, in S1 above, the MARS and iLIDS-VID datasets are classified, with the training and test sets split in a 1:1 ratio. MARS is a video dataset for pedestrian re-identification research, including pedestrian videos from multiple perspectives, covering different behaviors and scenes, including over 1,700 video sequences from 6 cameras across indoor and outdoor perspectives. iLIDS-VID is a dataset specifically designed for video-based pedestrian re-identification research, focusing on pedestrians from different perspectives. The dataset includes 600 short video sequences captured by 2 cameras, each sequence capturing images of pedestrians from different camera perspectives. Due to its unique camera setup and scenes, iLIDS-VID demands high robustness and generalization ability from the model, making it suitable for generalization research.
[0028] Furthermore, in S2 above, the teacher model PiT is trained. First, the input video image is data augmented, including horizontal flipping, scaling, and cropping. This augmented image is then used in the distillation training of the student model, enhancing the generalization ability of the teacher model and making the student model's knowledge source more reliable. The PiT teacher model is selected for training, and multi-scale feature representation fusion is performed through a pyramid structure that includes global-level information and local-level information from different scales. Specifically, in this Transformer-based architecture, each pedestrian image is segmented into many blocks. These blocks are then fed into a Transformer layer to obtain the image's feature representation. Vertical and horizontal partitioning is applied to these blocks to generate human body parts in different orientations. These parts provide more fine-grained information. The network is then trained using an objective function consisting of triplet loss and additive angular margin loss. Finally, information from the teacher model is extracted, including the model self-attention weight matrix and feature embeddings. These features contain key information about the pedestrian. The self-attention weight matrix reflects the importance of the model at each position in the input sequence, while the feature embedding captures the rich features of pedestrians in the input image. The extracted features include key identity information, appearance features, and spatial relationships in the pedestrian re-identification task.
[0029] Furthermore, in S3, a convolutional attention module MCT (Mobile Conv Trans) layer is inserted. The MCT module layer is used as the Transformer attention layer for the student model. This module includes one downsampling layer, one MobileNetV2Blocks layer, and four Transformer Blocks. Finally, the blocks are merged (i.e., a PatchMerging layer without convolutional downsampling technology) to perform downsampling.
[0030] The following is an introduction to each module.
[0031] Downsample Block:
[0032] Composed of convolution (Conv) and max pooling operations, it refers to a 2×2 standard convolutional layer plus a pooling layer, used to reduce the spatial resolution of the input. This operation helps to extract more abstract features while reducing computational costs.
[0033] MobileNetV2 Blocks (where λ=4):
[0034] Compared to traditional CNN networks, Transformer layers have weaker spatial awareness and heavily rely on large-scale datasets. Therefore, a MobileNetV2 inverse residual layer module is inserted before the Transformer, where λ=4 represents the width factor. The number of channels is adjusted to control the model size and computational requirements. This helps extract mid-level features, allowing the model to adapt to the different characteristics of different pedestrians.
[0035] Transformer module:
[0036] It includes Depthwise Separable Convolution Self-Attention (DCSA) and Recursive Atrous Self-Attention (RASA) layers. DCSA introduces depthwise separable convolution operations, which can better capture local relationships while reducing the number of parameters. RASA introduces atrous convolutions to handle information at different scales.
[0037] Patch Merging layer:
[0038] The process of integrating information from different locations without losing too much information during downsampling.
[0039] Further, in S4, the student model undergoes distillation. First, the model is initialized and iteratively optimized. Distilled features extracted by the teacher model are used as supervision signals to train the student model. Through multiple iterations, the student model gradually optimizes its performance, approximating the knowledge representation of the teacher model. During the forward propagation of iterative optimization, the input image is the same as previously used. Then, the temperature parameter of the Softmax output is adjusted as needed to ensure the smoothness of the probability distribution. The distillation loss between the Softmax probability distributions of the teacher and student models is calculated using the Knowledge Distillation (KD) Loss function. Next, the distillation loss is combined with the Triplet Loss function for the person re-identification task to create an overall loss function. The Adam optimizer is set up, and the parameters of the student model are updated using the backpropagation algorithm to minimize the overall loss. This step aims to simultaneously optimize distillation learning and the person re-identification task. Finally, person re-identification matching is performed. Pedestrian features are extracted from the query dataset and the gallery dataset to be matched, respectively. The similarity is obtained by calculating the Euclidean distance between the feature vectors. Images in the dataset to be matched are sorted by similarity to obtain the Rank-1 hit rate and the mean accuracy (mAP). This ultimately enables efficient re-identification of pedestrian samples.
[0040] According to another aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the lightweight method for pedestrian re-identification in video scenes based on knowledge distillation of the present invention.
[0041] According to another aspect of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the lightweight method for pedestrian re-identification in video scenes based on knowledge distillation of the present invention.
[0042] Compared with existing technologies, the beneficial effects of the above-described method of the present invention are as follows:
[0043] The method of this invention overcomes the shortcomings of existing models in pedestrian re-identification, which have a large number of parameters and high requirements for training data. By using distillation learning on a large model, the number of model parameters is significantly reduced while maintaining high performance, so that it still has good computational efficiency in resource-constrained environments.
[0044] The method of this invention introduces dataset augmentation processing during the distillation learning process, applying the dataset augmentation processing in distillation learning to both the teacher model and the student model simultaneously. This fully utilizes the information in the dataset, accelerates the distillation learning process, and enhances the reliability of the model.
[0045] This invention, based on the TinyViT student model, proposes a novel MCT module and applies depthwise separable convolution to the convolutional attention module, including a Convolutional Attention (DCSA) module using depthwise separable convolution and a Dilated Convolution-based Reactive Attention (RASA) module. This module not only effectively captures key information in pedestrian images but also significantly reduces computational costs through the application of depthwise separable convolution, making the model more lightweight and suitable for various environments. Attached Figure Description
[0046] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of the present invention and are not intended to limit the present invention.
[0047] Figure 1 This is the overall training flowchart of the present invention;
[0048] Figure 2 This is a model framework diagram of the present invention;
[0049] Figure 3 This is a flowchart of the data enhancement and reuse process of the present invention;
[0050] Figure 4 This is a pyramid segmentation diagram of the teacher model of the present invention;
[0051] Figure 5 This is a structural diagram of the student model of the present invention;
[0052] Figure 6 This is a block diagram of the DCSA-RASA module of the present invention. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention.
[0054] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.
[0055] Example 1:
[0056] like Figure 1-6 As shown, this invention proposes a lightweight method for pedestrian re-identification in video scenes based on knowledge distillation, comprising the following steps:
[0057] The entire process can be divided into a pedestrian re-identification model training process and a testing process. The specific training process is as follows: Figure 1 As shown.
[0058] The first step is to download the video datasets MARS and iLIDS-VID. Classify the MARS and iLIDS-VID datasets, splitting the training and test sets in a 1:1 ratio. MARS is a video dataset for pedestrian re-identification research, including pedestrian videos from multiple perspectives, covering different behaviors and scenes, including over 1,700 multi-camera video sequences from 6 cameras, both indoor and outdoor. iLIDS-VID is a dataset specifically designed for video-based pedestrian re-identification research, focusing on pedestrians from different perspectives. The dataset includes 600 short video sequences captured by 2 cameras, each sequence capturing images of pedestrians from different camera perspectives.
[0059] The second step is to train the teacher model PiT, which mainly consists of four parts: dataset preprocessing, teacher model feature extraction, loss calculation, and distillation feature extraction. The specific steps are as follows:
[0060] S2.1: Before training, perform data augmentation on the input video images. Data augmentation methods include horizontal flipping, scaling, and cropping. Making the dataset as diverse as possible can improve the generalization ability of the trained re-identification model. Simultaneously, save the data augmentation results for later use, such as... Figure 3 As shown.
[0061] S2.2: Given some pedestrian videos and pedestrian IDs, after patched embedding, the images are first transformed into embedding vectors, which contain the local features and details of the images. Specifically, after performing a convolution operation on the pedestrian images, a feature map is obtained. Then the feature map is flattened to generate N features, where The size of each feature is In this way, each feature can be viewed as a feature embedding for each image patch, and the size of each image patch is the same as the convolution kernel size k. The convolution stride s determines the spacing between adjacent image patches. Then, a multi-directional, multi-scale pyramid module from the teacher model is used to extract features from the video images, enhancing the model's robustness and obtaining key feature information of pedestrians from multiple directions. The pyramid module here refers to... Figure 2The `pyramid` module in the code uses a four-layer feature pyramid for segmentation: overall segmentation, vertical segmentation, horizontal segmentation, and small block segmentation. For example... Figure 4 As shown, the processing formula is as follows:
[0062]
[0063] in, It is a class tag that represents the overall information of the pedestrian. This indicates the block markers after video image segmentation.
[0064] The following is the specific processing procedure:
[0065] First, the image patches are marked and rearranged, and then copied four times, each as a separate section for segmentation in each direction.
[0066] For the first part, no segmentation strategy is used; the class is labeled. The image patch markers are flattened into a new marker sequence. It is related to the overall image Equivalent. Then discard all image patch tags and keep only the class tags. and will As a representation of the entire image.
[0067] For the second part, vertical segmentation is used, dividing it in the vertical direction to generate... Each part includes [number] sections. Each image patch is labeled "patched," and the class is labeled "copy." Then, it is flattened into a one-dimensional tensor in the vertical direction to form a new vector sequence. .
[0068] For the third part, horizontal partitioning is used, dividing it in the horizontal direction to generate... Each part includes [number] sections. Each image patch is labeled "patched," and the class is labeled "copy." Then, it is flattened horizontally into a one-dimensional tensor, forming a new vector sequence. .
[0069] For the fourth part, both horizontal and vertical segmentation are used simultaneously to form a block-based partition, generating... Each part includes [number] sections. Each image patch is labeled patched, where Simultaneously generate A sequence of tags .
[0070] S2.3: The loss function consists of triplet loss and additive angular spacing loss (ArcFace loss). The loss function is calculated, the network is trained, and backpropagation updates the network parameters to obtain the trained person re-identification network model. The function is as follows:
[0071]
[0072] in and These represent the proportions of triplet loss and additive angular interval loss, respectively. and These are the triplet loss and the additive angular spacing loss (ArcFace loss), respectively. In this invention, Take 0.4, Take 0.6.
[0073] S2.4: Extracting distilled features from the teacher model. This step encompasses the extraction of the model's self-attention weight matrix and feature embeddings. These distilled features carry key information from pedestrian images and are intended to be transferred to the student model. Through distillation learning, the teacher model's knowledge is effectively transferred, helping to improve the performance of the student model and enabling it to better capture important features of pedestrian images.
[0074] The third step is to construct and improve the student model network TinyViT, whose model structure is as follows: Figure 5 As shown.
[0075] S3.1: First, the image is passed through the Patch Embedding layer. The Patch Embedding layer divides the input pedestrian image into multiple non-overlapping patches. Each patch is treated as an image segment, and each patch is mapped to a feature vector with a fixed dimension through convolution operation. At the same time, it accepts image patches transmitted from the teacher model to enhance the learning effect of the student model input, thereby enhancing the robustness and generalization of the model.
[0076] S3.2: Input the segmented image into the designed MCT module. Use the MCT module layer as the Transformer attention layer of the student model. This module includes one downsampling layer, one MobileNetV2 block layer, and four Transformer blocks, and finally perform Patch Merging. The following is a description of each module:
[0077] (1) Downsample Block:
[0078] Composed of convolution (Conv) and max pooling operations, it refers to a 2×2 standard convolutional layer plus a pooling layer, used to reduce the spatial resolution of the input. This operation helps to extract more abstract features while reducing computational costs.
[0079] (2) MobileNetV2 Blocks:
[0080] Compared to traditional CNN networks, Transformer layers have weaker spatial awareness and heavily rely on large-scale datasets. Therefore, a MobileNetV2 inverse residual layer module is inserted before the Transformer, where λ=4 represents the width factor, adjusting the number of channels to control the model size and computational requirements. The module primarily uses direct bottleneck connections. Since the bottleneck actually contains all the necessary information, while the expansion layers are merely implementation details of the nonlinear transformations accompanying the tensors, MobileNetV2 layers use shortcut connections directly between bottlenecks instead of expansion layers. This helps extract mid-level features, allowing the model to adapt to the different characteristics of different pedestrians.
[0081] (3) Transformer module:
[0082] It includes Depthwise Separable Convolution Self-Attention (DCSA) and Recursive Atrous Self-Attention (RASA) layers. DCSA introduces depthwise separable convolution operations, which can better capture local relationships while reducing the number of parameters. RASA introduces atrous convolutions to handle information at different scales.
[0083] DCSA and RASA architecture diagrams are as follows: Figure 6 As shown. The module uses a 1-layer DCSA layer and a 3-layer RASA module.
[0084] First is the DCSA module. This invention applies depthwise separable convolution to the attention mechanism and designs a convolutional attention block based on depthwise separable convolution.
[0085] Depthwise separable convolution breaks down standard convolution into two steps: depthwise convolution and pointwise convolution. Depthwise convolution performs convolution only on each input channel, reducing the number of parameters. This is important for lightweight model design, as it reduces computational and storage overhead while maintaining model performance. Depthwise convolution can capture the spatial correlation of input features at the channel level, while pointwise convolution, by combining the outputs of depthwise convolution, can capture a wider range of feature information. This helps to increase the model's receptive field, enabling it to better understand the global structure of the input data. Furthermore, the combination of depthwise and pointwise convolution introduces non-linear transformations, increasing the model's expressive power. This is beneficial for the model to learn complex feature representations and improve its generalization performance. Finally, depthwise separable convolution has fewer parameters than standard convolution, making it more computationally efficient. This makes this invention more suitable for maintaining good performance under conditions of limited parameter performance.
[0086] The module uses depthwise separable convolutions to extract low-level features, and global self-attention for higher-level feature extraction. This method aims to find a balance between these two approaches by designing a self-attention layer based on depthwise separable convolutions, combining self-attention with the convolution process to improve feature extraction performance.
[0087] The module steps are as follows:
[0088] 3.2.1. Vector input: The input is a four-dimensional tensor representing the image, with shape [B, H, W, C], where B is the batch size, H and W are the height and width, respectively, and C is the number of channels.
[0089] 3.2.2 Unfold Operation: The Unfold operation is used to process the tensor after the channel dimension transformation, dividing the image into blocks and flattening them. This reduces the number of parameters and facilitates subsequent depthwise separable convolution multiplication operations.
[0090] 3.2.3 Depthwise Separable Convolution: Depthwise separable convolution breaks down standard convolution into two steps: depthwise convolution and pointwise convolution. Specifically, the flattened blocks are linearly transformed with weights to facilitate convolution, followed by depthwise convolution and pointwise convolution to obtain the final v matrix. The v matrix functions similarly to the weight matrix v in attention.
[0091] 3.2.4 Calculating Similarity: Here, a similarity matrix s is used instead of the q and k matrices in ordinary attention. First, pooling is performed on the tensor to reduce the spatial dimensionality of the input. This step aims to reduce computational complexity while retaining key information.
[0092] Next, the correlation between different parts of the similarity matrix (attn) input data is calculated through linear transformation and activation functions, providing weight information for subsequent weighting operations. Then, a linear transformation is performed on the pooled data. This step uses a linear layer to transform the input and learn the corresponding weight parameters. Finally, the result of the linear transformation is reshaped to obtain a five-dimensional tensor. Among them, the second dimension The first dimension represents the spatial location of the data, the second dimension represents the number of attention heads, and the last two dimensions are used to construct a similarity matrix.
[0093] Then, a nonlinear mapping is applied to the result of the linear transformation using an activation function, introducing a nonlinear relationship. This invention employs the ReLU6 function. Finally, the matrix is scaled and normalized to prevent overfitting and gradient explosion.
[0094] 3.2.5 Weighted Multiplication: The weight matrix v and similarity matrix attn obtained in 3.2.3 and 3.2.4 respectively are multiplied by weight. By calculating the weight of each local region, the model’s focus on important features is strengthened.
[0095] 3.2.6 Fold operation: Perform a Fold operation on the weighted output to restore it to the same dimensions as the input and integrate local information.
[0096] 3.2.7 Output: The output is mapped to the final output dimension through a linear transformation. Meanwhile, to improve the model's generalization ability, Dropout is used to prevent overfitting.
[0097] (4) Patch Merging layer:
[0098] The process of integrating information from different locations without losing too much information during downsampling.
[0099] S3.3: Pre-train the student model.
[0100] S4: Distillation learning is performed on the student model, mainly consisting of four parts: model initialization and iterative optimization, calculation of distillation loss, joint optimization, and pedestrian re-identification matching. The specific steps are as follows:
[0101] S4.1: The prediction results of the teacher model (i.e., the distilled features extracted in S2.4) are used as a supervision signal to train the student model. Through multiple iterations, the student model gradually optimizes its performance, continuously approaching the knowledge representation of the teacher model. The input image used in the forward propagation process of this iterative optimization is the same as that used in S2.1.
[0102] S4.2: Adjust the temperature parameter of the Softmax output as needed to ensure a smooth probability distribution. Calculate the distillation loss between the Softmax probability distributions of the teacher and student models using the KD Loss function.
[0103] S4.3: Calculate the distillation loss between the softmax probability distributions of the teacher and student models using the KD Loss function. This involves creating an overall loss function and combining the distillation loss function with the softmax loss function used for person re-identification. An Adam optimizer is configured to jointly optimize the overall loss function, and the parameters of the student model are updated using backpropagation to minimize the overall loss. The KD function is as follows:
[0104]
[0105] in, and These are the weighted proportions corresponding to the distillation loss and the student loss, respectively. and These are the distillation loss function and the Softmax loss function used for pedestrian re-identification, respectively. For teacher models in The Softmax output at temperature is at the 1st Values on a class For student models in The Softmax output at temperature is at the 1st Values on a class In the first Ground truth value on class, Positive labels are assigned a value of 1, and negative labels are assigned a value of 0. This represents the total number of tags.
[0106] S4.4: Perform pedestrian re-identification and matching. Extract pedestrian features from the query dataset and the matching image dataset. Calculate the similarity between each image in the query dataset and each image in the matching image dataset using the Euclidean distance between their feature vectors. Sort the images in the matching image dataset by similarity to obtain the Rank-1 hit rate and the average precision mAP. Finally, achieve the re-identification of pedestrian samples.
[0107] Example 2:
[0108] The computer-readable storage medium of this embodiment stores a computer program that, when executed by a processor, implements the steps in the lightweight method for pedestrian re-identification in a video scene based on knowledge distillation in Embodiment 1.
[0109] The computer-readable storage medium in this embodiment can be an internal storage unit of the terminal, such as the terminal's hard disk or memory; the computer-readable storage medium in this embodiment can also be an external storage device of the terminal, such as a plug-in hard disk, smart memory card, secure digital card, flash memory card, etc. equipped on the terminal; furthermore, the computer-readable storage medium can include both the terminal's internal storage unit and external storage devices.
[0110] The computer-readable storage medium of this embodiment is used to store computer programs and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0111] Example 3:
[0112] The computer device of this embodiment includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the lightweight method for pedestrian re-identification in video scenes based on knowledge distillation in Embodiment 1.
[0113] In this embodiment, the processor can be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The memory can include read-only memory and random access memory, and provides instructions and data to the processor. A portion of the memory can also include non-volatile random access memory. For example, the memory can also store device type information.
[0114] Those skilled in the art will understand that the content disclosed in the embodiments can be provided as a method, system, or computer program product. Therefore, this solution can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this solution can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) that include computer-usable program code.
[0115] This solution is described with reference to flowchart illustrations and / or block diagrams of methods and computer program products according to embodiments of this solution. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0116] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0117] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0118] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0119] The examples described herein are merely preferred embodiments of the invention and are not intended to limit the concept and scope of the invention. Any modifications and improvements made by those skilled in the art to the technical solutions of the invention without departing from the design concept of the invention should fall within the protection scope of the invention.
Claims
1. A lightweight method for pedestrian re-identification in video scenes based on knowledge distillation, characterized in that, S1. Download and process the MARS and iLIDS-VID datasets for the video scene; S2. Train the teacher model PiT, including dataset preprocessing, teacher model feature extraction, loss calculation, and distillation feature extraction; S3. Construct a student model TinyViT with a small number of parameters and insert a convolutional attention module MCT into the model; the convolutional attention module MCT includes a convolutional attention DCSA module with depthwise separable convolution and a RASA module based on dilated convolution, which extracts fusion features from the video and enhances the network's accuracy in extracting image features. S4. Perform distillation learning on the student model; including model initialization and iterative optimization, calculation of distillation loss, joint optimization and pedestrian re-identification matching; The specific steps of S2 are as follows: S2.1: Before training, perform data augmentation on the input video images. Data augmentation methods include horizontal flipping, scaling, and cropping. S2.2: Given some pedestrian videos and pedestrian IDs, after patched embedding, the images are first transformed into embedding vectors, which contain the local features and details of the images; specifically, after performing a convolution operation on the pedestrian images, a feature map is obtained. Then the feature map is flattened to generate N features, where The size of each feature is In this way, each feature can be regarded as a feature embedding of each image patch, and the size of each image patch is the same as the convolution kernel size k; the convolution stride s determines the interval between adjacent image patches; then, the multi-directional, multi-scale pyramid module in the teacher model is used to extract features from the video image, enhancing the robustness of the model and obtaining key feature information of pedestrians in multiple directions; specifically, a four-layer feature pyramid is used for segmentation, namely, overall segmentation, vertical segmentation, horizontal segmentation, and small patch segmentation; the processing formula is as follows: ; in, It is a class tag that represents the overall information of the pedestrian. Indicates block markers after video image segmentation; S2.3: Calculate the loss function, train the network, and update the network parameters through backpropagation to obtain the trained person re-identification network model; the loss function consists of triplet loss and additive angular interval loss (ArcFace loss), as follows: ; in and These represent the proportions of triplet loss and additive angular interval loss, respectively. Tri and L Arc These are the triplet loss and the additive angular interval loss (ArcFace loss), respectively. S2.4: Extract distilled features from the teacher model, including the extraction of the self-attention weight matrix and feature embeddings.
2. The method according to claim 1, characterized in that, The specific steps for S1 are as follows: The MARS and iLIDS-VID datasets were classified, and the training and test sets were divided in a 1:1 ratio.
3. The method according to claim 2, characterized in that, In S3, a convolutional attention module (MCT) is inserted, and the MCT module is used as the Transformer attention layer of the student model; The MCT module includes: one downsampling layer, one MobileNetV2 Blocks, and four Transformer Blocks. Finally, the blocks are merged to perform downsampling.
4. The method according to claim 3, characterized in that, The specific steps for S4 are as follows: S4.
1. Use the prediction results of the teacher model as a supervision signal to train the student model; through multiple iterations, the student model gradually optimizes its performance and continuously approaches the knowledge representation of the teacher model. S4.
2. Adjust the temperature parameters output by Softmax as needed to ensure a smooth probability distribution; S4.
3. Calculate the distillation loss between the Softmax probability distributions of the teacher model and the student model using the KD loss function. That is, create an overall loss function, combine the distillation loss function with the Softmax loss function used for the person re-identification task, set up the Adam optimizer to jointly optimize the overall loss function, and use the backpropagation algorithm to update the parameters of the student model to minimize the overall loss. The KD loss function is as follows: ; in, and These are the weighted proportions corresponding to the distillation loss and the student loss, respectively. and These are the distillation loss function and the Softmax loss function used for person re-identification, respectively. For teacher models in The Softmax output at temperature is at the 1st Values on a class For student models in The Softmax output at temperature is at the 1st Values on a class In the first The tag value on the class, Positive labels are assigned a value of 1, and negative labels are assigned a value of 0. This represents the total number of tags; S4.
4. Perform pedestrian re-identification and matching. Extract pedestrian features from the query dataset (Query) and the matching image dataset (Gallery). Calculate the similarity between each image in the query dataset and each image in the matching image dataset using the Euclidean distance between their feature vectors. Sort the images in the matching image dataset according to their similarity to obtain the Rank-1 hit rate and the average precision mAP. Finally, achieve the re-identification of pedestrian samples.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by the processor, it implements the steps in the lightweight method for pedestrian re-identification in video scenes based on knowledge distillation as described in any one of claims 1 to 4.
6. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the lightweight method for pedestrian re-identification in video scenes based on knowledge distillation as described in any one of claims 1 to 4.