Unsupervised pedestrian re-identification method based on local feature complementary denoising

By combining global and local features for complementary noise removal and introducing teacher model guidance, the problem of limited feature expression ability and noise in unsupervised pedestrian re-identification is solved, and the recognition accuracy and generalization ability of the model are improved.

CN120260141APending Publication Date: 2025-07-04NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510746146.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing unsupervised pedestrian recognition method is overly dependent on global features, resulting in limited feature expression capabilities and lack of effective mechanisms when dealing with clustering noise and pseudo-label noise, which affects the generalization ability and recognition accuracy of the model.

Method used

By combining global features and local features for complementary denoising, a pseudo-label is generated using Gaussian hybrid model and clustering algorithm, and a teacher model is introduced for knowledge distillation, optimizing feature learning and pseudo-label quality.

Benefits of technology

It improves the quality of pseudo-labels and clustering stability, improves the early learning efficiency and robustness of the model, and enhances the generalization ability in real scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260141A_ABST
    Figure CN120260141A_ABST
Patent Text Reader

Abstract

The invention discloses an unsupervised pedestrian re-identification method based on local feature complementary denoising, and belongs to the technical field of computer vision and deep learning. The method comprises the following steps: extracting global features of a pedestrian image through a reference model, and fusing local features to obtain combined features; de-noising of global features and local features is cooperatively completed; generating a pseudo label by using a clustering algorithm; updating parameters of the reference model by using a back propagation method; through knowledge distillation loss, a teacher model is used to guide a reference model to extract features, and a student model is used to realize pedestrian re-identification. According to the method, the local view is switched into and is combined with the global view for learning, so that the most significant appearance clue of the whole can be concerned, and the local fine-grained clue can be explored; gaussian prior hypothesis is carried out by using mixed features, the difference of classification features is improved, and the purpose of pseudo tag denoising is achieved; and the trained teacher model is adopted to guide the student model, so that the denoising effect is optimal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a pedestrian re-identification method, specifically an unsupervised pedestrian re-identification method based on local feature complementary denoising, belonging to the technical field of computer vision and deep learning. Background Art

[0002] Pedestrian re-identification (Person Re-identification, ReID) aims to retrieve specific pedestrian targets across camera scenes. With the popularization of applications such as security monitoring and intelligent retail, pedestrian re-identification technology has attracted much attention, and its main applications include public security monitoring, crowd analysis, intelligent transportation, and unmanned retail scenarios.

[0003] Existing pedestrian re-identification is mainly divided into two types: supervised learning and unsupervised learning. Supervised learning methods rely on a large amount of manually labeled data to learn the identity information of pedestrians through deep neural networks. However, the cost of manual labeling is expensive and time-consuming, which severely restricts the development of supervised learning. Therefore, in recent years, fully unsupervised pedestrian re-identification (FullyUnsupervised Learning, USL-ReID) methods have received more and more attention.

[0004] Unsupervised pedestrian re-identification methods mainly extract discriminative features from unlabeled data through techniques such as pseudo-label generation and contrast learning. Currently, most mainstream unsupervised methods use clustering algorithms to group images and use pseudo-labels to guide model training. These methods assume that images of the same person have higher similarity and are thus more likely to be clustered into the same class. However, existing fully unsupervised methods mainly rely on global features for clustering and are easily affected by feature biases, resulting in incorrect matching of different images of the same identity. In addition, the noise in the pseudo-labels is also an important factor affecting the model performance. Since unsupervised pedestrian re-identification relies on pseudo-labels for model training, the quality of the pseudo-labels directly determines the final performance of the model. However, during the clustering process, due to the similarity calculation between samples being affected by some factors, some samples are misclassified, thus introducing label noise. These noisy samples will affect the stability of feature learning, making it difficult for the model to converge during training and even leading to a decline in feature representation ability. Therefore, how to effectively identify and reduce pseudo-label noise and how to obtain more discriminative features have become key issues in improving the performance of unsupervised pedestrian re-identification.

[0005] In existing methods, SPCL (Yixiao Ge, et al. "Self-paced contrastive learning with hybrid memory for domain adaptive object re-id." Advances in neural information processing systems 33 (2020): 11309-11321.) adopts a self-paced clustering strategy to gradually refine the clustering results and combines contrastive learning to optimize the feature representation. However, it relies on global features when dealing with noisy pseudo-labels and ignores the importance of local features. CCL (Zuozhuo Dai, et al. "Cluster contrast for unsupervised person re-identification." Proceedings of the Asian conference on computer vision. 2022.) enhances the robustness of clustering through cluster-level contrastive learning. However, its feature update strategy is vulnerable to noise accumulation, resulting in a decline in model performance. In addition, PPLR (Yoonki Cho, et al. "Part-based pseudo label refinement for unsupervised person re-identification." Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022.) enhances the robustness of the model through cross-consistency learning of local and global features, but fails to fully utilize local features for clustering optimization.

[0006] In summary, the existing USL-ReID methods mainly have the following key problems: over-reliance on global features or simple use of local features, resulting in limited feature expression ability; lack of effective mechanisms for dealing with clustering noise and pseudo-label noise, affecting the generalization ability and recognition accuracy of the model. Summary of the Invention

[0007] Objective of the Invention: Aiming at the above problems, the objective of the present invention is to provide an unsupervised person re-identification method based on local feature complementary denoising, which starts from the local view and jointly learns with the global view, so as to not only focus on the most significant overall appearance clues but also explore local fine-grained clues to improve the accuracy of person re-identification.

[0008] Technical solution: An unsupervised pedestrian re-identification method based on local feature complementarity denoising of the present invention includes the following steps: Step 1, obtain a pedestrian image, extract the global feature of the pedestrian image through a reference model, divide the global feature in the horizontal direction to obtain local features, and fuse the local features to obtain a combined feature; Step 2, use the combined feature and the Gaussian mixture model to jointly complete the denoising of the global feature and the local feature; Step 3, use a clustering algorithm to generate pseudo-labels, and construct a loss function using the soft triple loss; Step 4, combine the pseudo-labels and the loss function, use the backpropagation method to update the parameters of the reference model, and use the trained reference model as the teacher model; Step 5, through the knowledge distillation loss, use the teacher model to guide the reference model to extract features; repeat Steps 1 to 3 to train the reference model, use the trained reference model as the student model, and use the student model to achieve pedestrian re-identification.

[0009] Further, Step 1 includes: Denote the global feature as , and divide the obtained global feature into two equal parts in the horizontal direction to obtain the first local feature and the second local feature , and use a set of adaptive weights to fuse and recombine the first local feature and the second local feature to obtain a combined feature , and the formula is: , In the formula, are trainable weight parameters, which are updated during each backpropagation process, and .

[0010] Further, Step 2 includes: Calculate the recognition loss of the combined feature , and the formula is: , In the formula, is the i-th sample, is the pseudo-label of the sample , represents the softmax function, is the combined feature learned by the sample ; Calculate the confidence of each sample by using the Gaussian mixture model, and the formula is: , Wherein, is the confidence value of the th sample, is the th component coefficient, GMM is the Gaussian mixture model, and n represents the total number of samples; Through the confidence threshold , the global features and local features higher than the confidence threshold are divided into a high-confidence feature set , and the global features and local features lower than the confidence threshold are divided into low-confidence features ; Let the sample features with the same pseudo-label follow a Gaussian distribution, and use the samples higher than the confidence threshold under the same pseudo-label to calculate the Gaussian prior distribution , and represent the mean vector and covariance matrix respectively, and the calculation formula is: , , Wherein, , where represents the global feature, t represents the first local feature, b represents the second local feature, is the pseudo-label class in which the number of samples higher than the confidence threshold ; represents the feature vector of the th sample higher than the confidence threshold in the part; Random sampling is performed using the Gaussian prior distribution to obtain the auxiliary denoising feature , and the low-confidence samples are denoised using the auxiliary denoising feature to obtain the final feature, and the formula is as follows: , to form the denoised feature set .

[0011] Furthermore, step 3 includes: Using the denoised feature set , calculate the Jaccard distance matrices of the global features and local features respectively ; Fuse the distance matrices of the three different parts to obtain the fused Jaccard distance matrix , and the formula is: , Wherein, is the balance factor, The Jaccard distance matrix representing global features The Jaccard distance matrix representing the first local feature The Jaccard distance matrix representing the second local feature; Construct three cluster-level memory banks, corresponding to global features, the first local feature, and the second local feature respectively; In each feature space, cluster the pedestrian image samples without any identity labels to generate pseudo-labels; Then calculate the cluster centroids corresponding to each pseudo-label category and store them in the corresponding memory bank. Finally, obtain three sets of cluster centroid memory banks, which record the pseudo-label category representations from three feature perspectives; Initialize the cluster centers in the corresponding memory bank using the corresponding cluster centroids. The formula is: , In the formula, represents the cluster centroid of the th cluster, represents the th cluster, represents calculating the number of instances in the corresponding cluster class, represents the th denoised feature vector in the cluster; Construct a contrastive learning loss function. The formula is: , In the formula, represents the feature vector of the query instance in the corresponding part, represents the cluster centroid feature vector of the cluster to which the query instance belongs; represents the cluster centroid feature vector of the th cluster stored in the memory, is the temperature hyperparameter, is the number of clusters in the memory; Construct a local feature fusion contrastive learning loss function. The formula is: , In the formula, is the balance factor; Use the query features extracted currently to iteratively perform momentum updates on the cluster memory bank. The update formula is: , In the formula, is the momentum update factor, represents the feature vector of the query instance ; Identify the nearest neighbor as a positive sample using cosine similarity and identify the farthest neighbor as a negative sample . Calculate the soft triplet loss, the formula is as follows: , , In the formula, represents the soft triplet loss corresponding to the part, is the balance factor, represents the th positive sample, represents the th negative sample; Then the total loss function of the baseline model is expressed as: , In the formula, is the temperature parameter.

[0012] Furthermore, step 5 includes: Calculate the knowledge distillation loss, the formula is: , In the formula, and are the normalized feature vectors of the query instance in the student model and the teacher model respectively; Calculate the total loss of the student model, the formula is: , In the formula, is the temperature parameter.

[0013] Beneficial effects: Compared with the prior art, the remarkable advantages of the present invention are: 1. The present invention realizes global-local joint learning by combining global features and local feature fusion; by calculating the local feature and global feature fusion distance matrix and through weight adjustment, a new clustering mechanism is constructed, effectively alleviating the problem that traditional global feature clustering is easily affected by biases such as background noise and redundant information, and improving the quality of pseudo-labels and clustering stability; 2. Use Gaussian features for denoising, reducing the interference of low-confidence features on clustering and improving the quality of pseudo-labels from the source; 3. Aiming at the unstable situation of pseudo-labels in the initial stage of training, introduce the structure of teacher model-student model, use the teacher model to guide the student model, improve the early learning efficiency and robustness of the model, accelerate network convergence and weaken the influence of noise; 4. Collaborate with Gaussian feature denoising and knowledge distillation to continuously correct pseudo-labels throughout the training cycle, and the denoising effect is better than traditional single pseudo-label filtering or confidence threshold methods; 5. Significantly improved performance is achieved on multiple challenging datasets such as Market-1501 and PersonX, demonstrating that the method proposed in the present invention has good generalization ability and adaptability in real scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a flowchart of the present invention; Figure 2 is a process framework diagram of the present invention; Figure 3 is a schematic diagram of the process for generating combined features. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments.

[0016] Combined Figures 1 to 2 As shown, an unsupervised pedestrian re-identification method based on local feature complementary denoising described in this embodiment includes the following steps: Step 1, obtain a pedestrian image, extract the global feature of the pedestrian image through a reference model, divide the global feature horizontally to obtain local features, and fuse the local features to obtain a combined feature.

[0017] Further, Step 1 includes: Denote the global feature as , divide the obtained global feature into two equal parts horizontally to obtain the first local feature and the second local feature , use a set of adaptive weights to fuse and recombine the first local feature and the second local feature to obtain a combined feature , and the formula is: , wherein, are trainable weight parameters, which are updated during each backpropagation process, and .

[0018] Although significant progress has been made in USL - ReID methods based on local features, these existing methods usually simply utilize local features to enhance feature learning, assuming that the importance of each local feature is the same. In fact, the number of feature cues contained in each local feature is different. To automatically capture more important local cues, this example proposes a brand - new way to extract combined features. In this example, the ResNet50 network is used as the baseline model to extract the global features of pedestrian images , the global features are split into two equal parts along the horizontal direction to obtain the first local feature and the second local feature , as Figure 3 shown, the first local feature corresponds to the upper half of the figure, and the second local feature corresponds to the lower half. Combining a set of adaptive weight parameters to fuse and reorganize them into combined features.

[0019] Step 2: Utilize the combined features and the Gaussian mixture model to jointly complete the denoising of global features and local features.

[0020] Furthermore, Step 2 includes: Calculate the recognition loss of the combined features , and the formula is: , wherein, is the i - th sample, is the pseudo - label of the sample , denotes the softmax function, is the combined feature learned by the sample ; Calculate the confidence of each sample by using the Gaussian mixture model, and the formula is: , wherein, is the confidence value of the - th sample, is the - th component coefficient, GMM is the Gaussian mixture model, and n represents the total number of samples; Through the confidence threshold , divide the global features and local features higher than the confidence threshold into the high - confidence feature set , and divide the global features and local features lower than the confidence threshold into the low - confidence features ; To better address the impact of noise on pseudo-labels, in this example, a Gaussian distribution assumption is utilized to obtain more robust features. Samples with the same pseudo-label are made to follow a Gaussian distribution, and samples above the confidence threshold under the same pseudo-label are used to calculate the Gaussian prior distribution , and represent the mean vector and covariance matrix respectively. The feature distribution is made smoother through the Gaussian prior distribution, and the calculation formula is: , , In the formula, , where represents the global feature, t represents the first local feature, b represents the second local feature, is the number of samples in the pseudo-label class above the confidence threshold ; represents the feature vector of the th sample above the confidence threshold in the part; Auxiliary denoising features are randomly sampled using the Gaussian prior distribution, and the low-confidence samples are denoised using the auxiliary denoising features to obtain the final features. The formula is as follows: , to form the denoised feature set .

[0021] In step 2, first, a Gaussian mixture model is used to calculate the confidence value , and then the threshold is used to flexibly control the range of feature denoising, and useful information in the original features is retained as much as possible during the entire denoising process. It should be noted that this step will not take effect in the first round of training.

[0022] Step 3, a clustering algorithm is used to generate pseudo-labels, and a loss function is constructed using the soft triple loss.

[0023] Furthermore, step 3 includes: Using the denoised feature set , the Jaccard distance matrices of the global features and local features are calculated respectively ; The distance matrices of the three different parts are fused to obtain the fused Jaccard distance matrix , and the formula is: , In the formula, is the balance factor, represents the Jaccard distance matrix of global features, represents the Jaccard distance matrix of the first local feature, represents the Jaccard distance matrix of the second local feature; Construct three cluster-level memory banks, corresponding to global features, the first local feature, and the second local feature respectively, so that the model can capture more detailed information during the learning process; In each feature space, cluster the pedestrian image samples without any identity labels to generate pseudo-labels, thereby reducing the bias caused by using only global features. It should be noted that in the first Epoch, the un-denoised feature groups are used when calculating the Jaccard distance matrix ; Then calculate the cluster centroids corresponding to each pseudo-label category and store them in the corresponding memory bank. Finally, three groups of cluster centroid memory banks are obtained, which record the pseudo-label category representations from three feature perspectives; Initialize the cluster centers in the corresponding memory bank with the corresponding cluster centroids. The formula is: , In the formula, represents the cluster centroid of the th cluster, represents the th cluster, represents calculating the number of instances in the corresponding cluster class, represents the th denoised feature vector in the cluster; Construct a contrastive learning loss function. The formula is: , In the formula, represents the feature vector of the query instance in the corresponding part, represents the cluster centroid feature vector of the cluster to which the query instance belongs; represents the cluster centroid feature vector of the th cluster stored in the memory, is the temperature hyperparameter, is the number of clusters in the memory; Construct a local feature fusion contrastive learning loss function. The formula is: , In the formula, is the balance factor; Iteratively perform momentum updates on the cluster memory bank using the query features extracted currently. The update formula is as follows: , wherein, is the momentum update factor, represents the query instance 's feature vector; For each picture, there will be many positive samples and negative samples. In this example, the cosine similarity is used to identify the nearest neighbor as a positive sample and the farthest neighbor as a negative sample , and calculate the soft triplet loss. The formula is as follows: , , wherein, represents the soft triplet loss corresponding to the part, is the balance factor, represents the th positive sample, represents the th negative sample; Then the total loss function of the baseline model is expressed as: , wherein, is the temperature parameter.

[0024] Step 4: Combine the pseudo labels and the loss function, and use the backpropagation method to update the parameters of the baseline model. Use the trained baseline model as the teacher model.

[0025] By using the total loss to iteratively train the baseline model, an offline and well-trained teacher model can be obtained.

[0026] Step 5: Through the knowledge distillation loss, use the teacher model to guide the baseline model to extract features; repeat Steps 1 to 3 to train the baseline model. Use the trained baseline model as the student model, and use the student model to achieve person re-identification.

[0027] After obtaining the offline and well-trained teacher model, in this example, the teacher model is used to retrain the student model. In this way, the student model can quickly obtain knowledge from the teacher model in the initial stage of training, so as to generate more accurate pseudo labels earlier.

[0028] Furthermore, Step 5 includes: Calculate the knowledge distillation loss. The formula is as follows: , In the formula, and are the normalized feature vectors of the query instance in the student model and the teacher model, respectively; Calculate the total loss of the student model, and the formula is: , In the formula, is the temperature parameter.

[0029] To further demonstrate the effectiveness and superiority of the unsupervised person re-identification method of the present invention, the following examples are used for illustration. On the datasets Market1501 and PersonX respectively, the method of the present invention is compared with the existing state-of-the-art USL-ReID method, and the comparison models include SPCL, CCL, DHCCN (from Yongxi Li, et al. "Distribution-Guided Hierarchical Calibration Contrastive Network for Unsupervised Person Re-Identification." IEEE Transactions on Circuits and Systems for Video Technology (2024).), AAMT (from Xiaofeng Qu, et al. "AAMT: Adversarial attack-driven mutual teaching for source-free domain-adaptive person reidentification." IEEE Transactions on Multimedia (2024).), DKL-MPL (from Wenjie Zhu, Bo Peng, and Wei Qi Yan. "Dual knowledge distillation on multiview pseudo labels for unsupervised person re-identification." IEEE Transactions on Multimedia (2024)), and the results are shown in Table 1 and Table 2 respectively. It can be seen from the results that on different datasets, using the mean average precision (mAP) and cumulative match characteristics as evaluation metrics for a fair comparison with the state-of-the-art methods, the person re-identification of the present invention has achieved remarkable results, further verifying the effectiveness of the method of the present invention. Among them, Rank-1, Rank-5, and Rank-10 in the cumulative match characteristics respectively represent the probabilities that the correct match appears in the top 1, top 5, and top 10 candidate results. The higher the value, the stronger the ability of the model to find the correct identity among the top several candidates.

[0030] Table 1 Comparison with the state-of-the-art methods on the Market1501 dataset

[0031] Table 2 Comparison with the state-of-the-art methods on the PersonX dataset

Claims

1. An unsupervised pedestrian re-identification method based on local feature complementarity for denoising, characterized in that It includes the following steps: Step 1: Obtain a pedestrian image, extract the global features of the pedestrian image through a baseline model, divide the global features horizontally to obtain local features, and fuse the local features to obtain combined features; Step 2: Use the combined features and the Gaussian mixture model to jointly denoise the global features and local features; Step 3: Use a clustering algorithm to generate pseudo-labels and construct a loss function using the soft triple loss; Step 4: Combine the pseudo-labels and the loss function, and use the backpropagation method to update the parameters of the baseline model. Use the trained baseline model as the teacher model; Step 5: Through the knowledge distillation loss, use the teacher model to guide the baseline model to extract features; Repeat Steps 1 to 3 to train the baseline model. Use the trained baseline model as the student model and use the student model to achieve pedestrian re-identification.

2. The unsupervised pedestrian re-identification method based on local feature complementarity denoising according to claim 1, wherein Step 1 includes: Denote the global feature as , and divide the obtained global feature into two equal parts along the horizontal direction to obtain the first local feature and the second local feature respectively. Use a set of adaptive weights to fuse and recombine the first local feature and the second local feature to obtain the combined feature , and the formula is as follows: , In the formula, are trainable weight parameters, which are updated during each backpropagation process, and .

3. An unsupervised pedestrian re-identification method based on local feature complementarity for denoising according to claim 2, characterized in that, Step 2 includes: Calculated combined feature recognition loss , the formula is: , Wherein, is the i-th sample, is the sample pseudo-label, is expressed as the softmax function, is the sample learned combined feature; Calculate the confidence of each sample by using the Gaussian mixture model. The formula is: , Wherein, is the confidence value of the th sample, is the th component coefficient, GMM is the Gaussian mixture model, and n represents the total number of samples; By a confidence threshold , the global features and local features higher than the confidence threshold are classified into a high-confidence feature set , and the global features and local features lower than the confidence threshold are classified into low-confidence features ; Let the sample features with the same pseudo-labels follow a Gaussian distribution, and use the samples above the confidence threshold under the same pseudo-label to calculate the Gaussian prior distribution , and represent the mean vector and the covariance matrix respectively, and the calculation formula is: , , In the formula, , where represents the global feature, t represents the first local feature, and b represents the second local feature, is the pseudo-label class with a number of samples higher than the confidence threshold ; represents the feature vector of the th sample with a part higher than the confidence threshold ; Randomly sample the auxiliary denoising features using the Gaussian prior distribution , and use the auxiliary denoising features to perform feature denoising on the low-confidence samples to obtain the final features. The formula is as follows: , Form a denoised feature set .

4. An unsupervised pedestrian re-identification method based on local feature complementarity denoising according to claim 3, characterized in that, Step 3 includes: Use the denoised feature set , and calculate the Jaccard distance matrices of the global features and the local features respectively ; Fuse the distance matrices of three different parts to obtain the fused Jaccard distance matrix , and the formula is as follows: , In the formula, is the balance factor, represents the Jaccard distance matrix of the global features, represents the Jaccard distance matrix of the first local features, represents the Jaccard distance matrix of the second local features; Construct three cluster-level memory banks, corresponding to the global features, the first local features, and the second local features respectively; In each feature space, cluster the pedestrian image samples without any identity labels to generate pseudo-labels; Then calculate the cluster centroids corresponding to each pseudo-label category and store them in the corresponding memory bank. Finally, obtain three groups of cluster centroid memory banks, which respectively record the pseudo-label category representations from three feature perspectives; Initialize the clustering centers in the corresponding memory banks with the corresponding cluster centroids. The formula is: , In the formula, represents the cluster centroid of the th cluster, represents the th cluster, represents calculating the number of instances in the corresponding cluster class, represents the th denoised feature vector in the cluster; Construct a contrastive learning loss function. The formula is: , In the formula, represents a query instance in the corresponding feature vector of the part, represents a query instance centroid feature vector of the cluster to which it belongs; represents the centroid feature vector of the th cluster stored in memory, is the temperature hyperparameter, is the number of clusters in memory; Construct a local feature fusion contrastive learning loss function. The formula is: , In the formula, is the balance factor; Use the extracted query features to iteratively perform momentum updates on the cluster memory bank. The update formula is: , In the formula, is the momentum update factor, represents the query instance 's eigenvector; Identify the nearest neighbors as positive samples using cosine similarity , and identify the farthest neighbors as negative samples , and calculate the soft triple loss, with the formula as follows: , , In the formula, represents the soft triple loss corresponding to part, is the balance factor, represents the th positive sample, represents the th negative sample; Then the total loss function of the baseline model is expressed as: , In the formula, is the temperature parameter.

5. An unsupervised pedestrian re-identification method based on local feature complementarity for denoising, characterized in that, Step 5 includes: Calculate the knowledge distillation loss. The formula is: , wherein, and are the normalized eigenvectors of query instance in the student model and the teacher model, respectively; Calculate the total loss of the student model. The formula is: , In the formula, is the temperature parameter.