Object Detection Method Based on Aligned Instance Knowledge Distillation

Through elastic selection of alignment instances and multiple distillation methods in the object detection task, the problem of student models being misguided is solved, and detection performance and prediction accuracy are improved.

CN118230037BActive Publication Date: 2025-05-13HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410306042.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-18
Publication Date
2025-05-13
Estimated Expiration
2044-03-18

AI Technical Summary

Technical Problem

The existing knowledge distillation method is difficult to effectively select aligned instances with high classification confidence in the object detection task, resulting in student models being incorrectly guided and affecting detection performance.

Method used

An elastic selection module based on alignment examples is proposed. By screening aligned examples with high classification confidence and low degree of dislocation, combining feature distillation, relationship distillation and response distillation, the characteristics, relationship and response knowledge in the alignment example are fully extracted and transferred to the student model.

Benefits of technology

Improve the prediction accuracy of student models, enhance the feature representation ability of aligned instances, and improve detection performance, especially when dealing with extremely unbalanced positive and negative instance ratios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118230037B_ABST
    Figure CN118230037B_ABST
Patent Text Reader

Abstract

The present invention discloses an object detection method based on alignment instance knowledge distillation. First, the ground truth box B GT , the classification scores C and the bounding boxes B predicted by the classification head and the regression head in the detection head of the teacher model are input into the alignment instance elastic selection module to obtain the alignment instances AIs. Secondly, the alignment instances AIs, the teacher feature map T and the student feature map S are input into the feature distillation module generated based on alignment instance masking; the alignment instances AIs, the teacher feature map T and the student feature map S are input into the relationship distillation module based on alignment instances; the alignment instances AIs are input into the response distillation module based on alignment instances. Finally, an overall distillation loss function is constructed to train the student model and output the object detection result. The present invention fully extracts the three types of knowledge of features, relationships and responses contained in the alignment instances and transfers them to the student model, improving the accuracy of the student model prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of target detection, and in particular relates to a target detection method based on aligned instance knowledge distillation. Background Art

[0002] Deep convolutional neural networks have been widely used in various computer vision tasks. Generally, larger models have better performance, but lower inference speed, making them difficult to deploy in industrial fields with limited hardware resources. To overcome this difficulty, researchers have conducted research in several directions, including pruning, quantization, deep convolution compression, and knowledge distillation. Among them, knowledge distillation is an effective solution for model compression. Knowledge distillation refers to training a lightweight student network under the supervision of a large, high-accuracy teacher network and expecting the student network to have the capabilities of the teacher network. The earliest knowledge distillation was mainly aimed at classification problems. The common point of classification problems is that the model needs to be processed by the Softmax layer at the end, and its output is the probability value of the corresponding category. According to the characteristics of the output results of classification problems, an efficient way to transfer knowledge is to use soft labels. Soft labels refer to the output of the Softmax layer obtained by the teacher model. Since soft labels have higher entropy and smaller gradient changes than hard labels, the student model converges faster than the teacher model.

[0003] At present, typical knowledge is divided into three categories, namely feature-based knowledge, relationship-based knowledge, and response-based knowledge. However, most distillation methods are designed for classification problems. It is very inefficient to directly transfer the distillation method used for classification to the detection model, because the ratio of positive and negative instances in the detection task is extremely unbalanced. Recently, some distillation methods designed for detection tasks have solved this problem and achieved good results. For example, the method of imitating efficient networks for distillation detection solves this problem by sampling a certain proportion of positive instances, and the fine-grained distillation detection method only focuses on the features near the real target area. However, the selection ratio of positive instances needs to be carefully designed, and there may be misplaced instances with high classification confidence and low regression quality in the selected positive instances. These misplaced instances will greatly limit the detection performance. If these instances are selected for distillation, they will cause incorrect guidance to the student model, thereby affecting the distillation effect. In addition, how to fully extract various types of knowledge from the instances and transfer them to the student model has become an urgent problem to be solved. Summary of the invention

[0004] Aiming at the problem that the misaligned instances predicted by the teacher model cause incorrect guidance to the student model, an object detection method based on aligned instance knowledge distillation is proposed. The present invention is mainly divided into four parts:

[0005] In the first part, in order to flexibly select alignment instances with high classification confidence from the detection instances of the teacher model, and then use them for the subsequent distillation of three kinds of knowledge, the present invention proposes an alignment instance elastic selection module. This module first uses the confidence of the classification score in the instance to obtain the mean μ and standard deviation σ, and selects the high classification confidence instances that fall in [μ+2σ, 1); secondly, based on the selected high classification confidence instances, the alignment instances are further selected, and the method is to use the absolute value of the difference between the intersection over Union (IoU) score and the classification confidence in the instance as the misalignment degree of the instance, and use the misalignment degree of the instance to obtain the mean μ and standard deviation σ, and select the alignment instances with low misalignment degrees that fall in (0, μ-2σ). The above two steps can complete the elastic selection of alignment instances with high classification confidence.

[0006] In the second part, in order to help the features corresponding to the aligned instances in the student model obtain higher representation capabilities, the present invention proposes feature distillation based on aligned instance masking generation. First, the position information of the aligned instances in the feature map is obtained through a binary mask, and then a binary mask is used to randomly select masking positions at these positions to cover the student feature map. Then a simple convolutional network is used to force the masked student feature map to generate a complete teacher feature map. Finally, the present invention proposes a new distillation loss based on the l2 loss to optimize the above-mentioned convolutional network, so that the generated student feature map is closer to the teacher feature map.

[0007] In the third part, in order to extract the correlation feature knowledge between different aligned instances to improve the performance of the student model, this paper proposes relational distillation based on aligned instances. First, for the features of the corresponding positions of different aligned instances on the feature map, the Euclidean distance is used to measure the correlation of the features, and then the Huber loss is used to transfer the correlation knowledge.

[0008] In the fourth part, in order to extract the reasoning knowledge in the aligned instances, the present invention proposes response distillation based on the aligned instances. By extracting the reasoning knowledge of classification and regression from the classification responses and regression responses corresponding to the aligned instances and passing it to the student model, the student model is gradually made more powerful and efficient.

[0009] The specific implementation of the object detection method based on aligned instance knowledge distillation includes the following steps:

[0010] Step (1). Set the real frame B GT , the classification score C and bounding box B predicted by the classification head and regression head in the detection head of the teacher model are input into the alignment instance elastic selection module to obtain the alignment instance AIs.

[0011] First, the classification confidence G is calculated based on the classification score C c , according to G cThe threshold τ is obtained by the mean and standard deviation of c , filter out G c Greater than τ c Instances of S c Then B and B GT The intersection over union (IoU) of G is used as the regression confidence, and the IoU and G c The absolute value of the difference is taken as the degree of misalignment G m , according to G m The threshold τ is obtained by the mean and standard deviation of m , filter out G m Greater than 0 and less than τ m Alignment instance AIs. Instance S c The calculation formula for the aligned instance AIs is:

[0012]

[0013] Step (2). Input the aligned instance AIs, teacher feature map T and student feature map S into the feature distillation module based on aligned instance masking generation to enhance the feature representation capability of the AIs position in the student feature map S.

[0014] First, the student feature map S is passed through a 1×1 convolutional layer f align Processing, then the features of the corresponding positions of AIs in the student feature map are randomly blocked by AIs random blocking mask M, and then the feature map G generated by the blocked student feature map is made closer to the teacher feature map through the convolution network Z. Finally, the generated feature map G and the teacher feature map T are distilled using l2 loss. align The processing can align the student feature map with the teacher feature map. Generate feature map G and distillation loss function L AMGD The calculation formula is:

[0015]

[0016] Step (3). Input the aligned instance AIs, teacher feature map T and student feature map S into the aligned instance-based relationship distillation module, pass the feature correlation information between AIs to the student model, and then assist the student model to converge.

[0017] First, the Euclidean distance ψ(·,·) is used to measure the feature correlation of different AIs in the teacher and student feature maps, respectively. i , T j represents the feature vectors corresponding to different AIs instances in the teacher feature map, S i , S jRepresents the feature vectors corresponding to different AIs instances in the student feature map. Secondly, in order to comprehensively consider the relative distances between other aligned instances, μ(·) is used to normalize the distances. Finally, Huber loss l(·,·) is used to transfer the correlation knowledge. The distillation loss function L ARELD The calculation formula is:

[0018]

[0019] Step (4). Input the aligned instance AIs into the aligned instance-based response distillation module and transfer the corresponding classification and regression reasoning information to the student model.

[0020] i is used to represent the i-th aligned instance in AIs, t is the teacher model, and s is the student model. Based on the reasoning information in the classification head response, the Softmax function is used to convert the original classification scores C output by the teacher and student models into t With C s Convert to probability distribution P t With P s , through the KL divergence classification response distillation loss, the inference knowledge of the classification response in the teacher model is transferred to the student model. For the inference information in the regression head response, each edge of the prediction box output by the teacher and student models is represented as a probability distribution B through the Softmax function. s , B t , and then the KL divergence regression response distillation loss is used to transfer the inference knowledge of the regression response in the teacher model to the student model. ARESD_cls With regression response distillation loss L ARESD_reg The calculation formula is:

[0021]

[0022] The distillation loss function based on the response distillation of the aligned instances is calculated as:

[0023] L ARESD =αL ARESD_cls +βL ARESD_reg

[0024] Among them, α and β are the hyperparameters of the balance loss.

[0025] Step (5). Construct an overall distillation loss function. The overall distillation loss function is composed of the three distillation losses proposed in steps (2), (3), and (4). The student model is trained through the overall distillation loss function, and the target detection result is output.

[0026] The overall distillation loss function calculation formula is:

[0027] L all =Lorigin +λ0L AMGD +λ1L ARELD +λ2L ARESD

[0028] Where λ0, λ1, λ2 are the hyperparameters of the balance loss. origin Represents the classification and regression losses used in the distilled student model.

[0029] Beneficial effects of the present invention:

[0030] The model of the present invention is composed of an alignment instance elastic selection module, feature distillation based on alignment instance masking generation, alignment instance-based relationship distillation, and alignment instance-based response distillation. The alignment instance elastic selection module selects alignment instances with high classification confidence from the detection instances of the teacher model through statistical characteristics. Feature distillation based on alignment instance masking generation randomly blocks student features and forces them to generate teacher features according to the positions provided by the alignment instances in the feature map. In order to mine potentially valuable relationship information of the alignment instances, the relationship distillation based on the alignment instances measures the feature correlation of the corresponding positions of different AIs in the feature map through the distance formula and transfers it to the student model. The response distillation based on the alignment instances transfers the classification and regression reasoning information of the teacher model from the classification response and regression response corresponding to the alignment instances to the student model. The present invention proposes a new feature, relationship, and response distillation loss to fully extract the three types of knowledge of features, relationships, and responses contained in the alignment instances and transfer them to the student model, thereby improving the accuracy of the student model prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 Architecture diagram of the object detection method based on aligned instance knowledge distillation;

[0032] Figure 2 The structure diagram of the screening process of the alignment instance elastic selection module;

[0033] Figure 3 Feature distillation structure diagram based on aligned instance masking generation;

[0034] Figure 4 The present invention predicts the distribution of SDFE before and after distillation;

[0035] Figure 5 The results of SDFE before and after distillation of the present invention are compared. DETAILED DESCRIPTION

[0036] The present invention proposes a target detection method based on aligned instance knowledge distillation. Figure 1As shown in the figure, the overall structure consists of an alignment instance elastic selection module, feature distillation based on alignment instance masking generation, relationship distillation based on alignment instance, and response distillation based on alignment instance. First, through the alignment instance elastic selection module, the confidence of the classification score in the instance is used to obtain the mean μ and standard deviation σ, and the high classification confidence instances falling in [μ+2σ, 1) are screened out. On the basis of the screened high classification confidence instances, the alignment instances are further screened out. The method is to use the absolute value of the difference between the IoU score and the classification confidence in the instance as the misalignment degree of the instance, and use the misalignment degree of the instance to obtain the mean μ and standard deviation σ, and the alignment instances with low misalignment degrees falling in (0, μ-2σ) are screened out. Secondly, the position information of AIs in the feature map is obtained through binary masking. At these positions, the binary mask is used to randomly select masking positions to cover the student feature map. A simple convolutional network is used to force the occluded student feature map to generate a complete teacher feature map, and the feature knowledge is transferred through the l2 distillation loss. Then, for the channel feature vectors of the corresponding positions of different AIs on the feature map, the Euclidean distance is used to measure the correlation of the features, and the Huber distillation loss is used to transfer the relational knowledge. Finally, the KL divergence distillation loss is used to extract the classification and regression reasoning knowledge from the classification responses and regression responses corresponding to the AIs and pass it to the student model, gradually making the student model more powerful and efficient.

[0037] Next, the implementation steps will be described in detail with reference to the accompanying drawings.

[0038] Step (1). Figure 2 As shown, elastically screening alignment instances with high classification confidence is divided into two processes, namely, the process of elastically screening instances with high classification confidence and the process of elastically screening alignment instances from the screened instances.

[0039] Step (1.1). Input the classification score C predicted by the classification head in the detection head of the teacher model into the alignment instance elastic selection module to obtain the instance S c First, the classification confidence G is calculated based on the classification score C c , then according to G c The threshold τ is obtained by the mean and standard deviation of c Finally, G c Greater than τ c Instances of S c , the calculation steps are:

[0040]

[0041] Where X represents the probability distribution of classification scores after Sigmoid processing, max(·) represents the maximum function, mean(·) represents the mean function, and std(·) represents the standard deviation function.

[0042] Step (1.2). Set the real frame B GT , the bounding box B predicted by the regression head in the detection head of the teacher model and the classification confidence G obtained in step (1.1) c , Instance S c Input into the alignment instance elastic selection module to obtain the alignment instance AIs. First, call bbox_overlaps(·,·) to calculate the predicted box B and the real box B GT The IoU of G is used as the regression confidence. c The absolute value of the difference is taken as the degree of misalignment G m Then according to G m The threshold τ is obtained by the mean and standard deviation of m , and finally filter out the misalignment degree greater than 0 and less than τ m The calculation steps of the aligned instance AIs are:

[0043]

[0044] Where abs(·) represents the absolute value function, mean(·) represents the mean function, and std(·) represents the standard deviation function.

[0045] Step (2). Figure 3 As shown, the aligned instance AIs, the teacher feature map T and the student feature map S are input into the feature distillation module based on aligned instance masking generation to enhance the feature representation capability of the AIs position in the student feature map.

[0046] Step (2.1). First, pass the 1×1 convolutional layer f align Process the student feature map S to obtain the feature map F A , so that the student feature map is aligned with the teacher feature map, the calculation formula is:

[0047] F A =f align (S)

[0048] Step (2.2). Set a binary mask to distinguish the position corresponding to AIs from other positions. The calculation formula is:

[0049]

[0050] Where α1 represents the position set corresponding to AIs, i, j represent the horizontal coordinate and vertical coordinate of the feature map respectively. If (i, j) falls in the position set corresponding to AIs, then Q i,j is 1 if the value is true, otherwise it is 0.

[0051] Step (2.3). Then set the random binary mask to cover the feature map, and the calculation formula is:

[0052]

[0053] Among them, R i,j represents a random number in (0, 1), i, j represent the horizontal and vertical coordinates of the feature map respectively, and θ represents the hyperparameter of the degree of masking.

[0054] Step (2.4). Considering the above two binary masks, the AIs random occlusion mask of the student feature map is obtained, and the calculation formula is:

[0055]

[0056] Among them, M i,j Indicates whether (i, j) is the occluded AIs position, i, j represent the horizontal and vertical coordinates of the feature map respectively.

[0057] Step (2.5). The feature map F calculated by step (2.1) A With AIs randomly occluding mask M to cover the student’s feature map, and then forcing the occluded student feature map to generate the predicted teacher feature map G through the convolutional network, the calculation expression is:

[0058]

[0059] where Z(·) represents the projection layer, which consists of two convolutional layers W v1 With W v2 , a ReLU activation layer, W v1 With W v2 represents a 3×3 convolutional layer, and F represents the input feature map of Z(·).

[0060] Step (2.6). Based on the above method, the present invention designs a new feature distillation loss L based on AIs masking generation AMGD , the calculation formula is:

[0061]

[0062] Where C, H, W represent the shape of the feature map, T c,i,j represents the feature of the cth channel (i, j) position in the teacher feature map, G c,i,j Represents the features of the c-th channel (i, j) position in the generated student feature map.

[0063] Step (3). Figure 1 As shown, the aligned instance AIs, the teacher feature map T and the student feature map S are input into the aligned instance-based relationship distillation module to transfer the feature correlation information between AIs to the student model, thereby assisting the student model to converge.

[0064] Step (3.1). First, the Euclidean distance ψ(·,·) is used to measure the feature correlation of the corresponding positions of different AIs in the teacher and student feature maps. Secondly, in order to comprehensively consider the relative distances between other aligned instances, μ(·) is used to normalize the distance, and the calculation formula is:

[0065]

[0066] where χ 2 ={(i, j)|i≠j, 1≤i, j≤K}, K represents the number of AIs, x i 、x j They represent the channel feature vectors corresponding to the i-th and j-th instances of AIs in the feature map, respectively.

[0067] Step (3.2). Use Huber loss l(·,·) to transfer correlation knowledge. The calculation formula of l(·,·) is:

[0068]

[0069] Where x and y represent the relationship between different instances.

[0070] Step (3.3). The calculation formula of the relation distillation loss function based on the aligned instance is:

[0071]

[0072] Where T i , T j denote the channel feature vectors of the i-th and j-th instances of AIs in the teacher feature map, respectively, and S i , S j They represent the channel feature vectors corresponding to the i-th and j-th instances of AIs in the student feature map, respectively.

[0073] Step (4). Figure 1 As shown, the aligned instance AIs are input into the aligned instance-based response distillation module to transfer the classification and regression reasoning information corresponding to the AIs to the student model.

[0074] Step (4.1). Use Softmax to convert the original score C output by the model into the inference information in the classification header response. t With C s Convert to probability distribution P t With P s , the calculation formula is:

[0075]

[0076] where τ is the softening P t With Ps The temperature factor of the probability distribution of . t is the teacher model and s is the student model.

[0077] Step (4.2). Calculate the classification response distillation loss L through the KL divergence function ARESD_cls , transfers the inference knowledge of the classification response in the teacher model to the student model, and the calculation formula is:

[0078]

[0079] Where K represents the number of AIs, L KL represents the KL divergence loss, represents the probability distribution of student classification responses corresponding to the i-th instance in AIs, represents the probability distribution of the teacher classification response corresponding to the i-th instance in AIs. Through the knowledge distillation of the classification response of AIs, the student detector gradually learns the classification knowledge aligned with the teacher detector.

[0080] Step (4.3). For the inference information in the regression head response, each edge of the prediction box output by the teacher and student models is represented as a probability distribution B through the Softmax function s , B t , the calculation formula is:

[0081]

[0082] in Respectively represent the probability distribution of the left, top, right, and bottom sides of the student model prediction box, They represent the probability distribution of the left, top, right, and bottom sides of the teacher model prediction box, and n represents the number of logits predicted for each side.

[0083] Step (4.4). Calculate the regression response distillation loss L through the KL divergence function ARESD_reg , transfers the inference knowledge of the regression response in the teacher model to the student model, and the calculation formula is:

[0084]

[0085] Where K represents the number of AIs, L KL represents the KL divergence loss, represents the probability distribution of the student regression response corresponding to the i-th instance in AIs, represents the probability distribution of the teacher regression response corresponding to the i-th instance in AIs.

[0086] Step (4.5). The overall distillation loss function L based on the response distillation of the aligned instances ARESD , the calculation formula is:

[0087] L ARESD =αL ARESD_cls +βL ARESD_reg

[0088] Among them, α and β are the hyperparameters of the balance loss.

[0089] Step (5). Construct the overall distillation loss. The overall distillation loss function is composed of the three distillation losses proposed in step (2), step (3), and step (4). The student model is trained through the overall distillation loss function and the target detection result is output.

[0090] Overall distillation loss L all The calculation formula is:

[0091] L all =L origin +λ0L AMGD +λ1L ARELD +λ2L ARESD

[0092] Among them, λ0, λ1, and λ2 are the hyperparameters of the balanced loss, and the default values ​​are set to 1.0, 0.5, and 0.25 respectively. origin Represents the classification and regression losses used in the distilled student model.

[0093] By training the model with this loss function, we can fully extract the three types of knowledge contained in the aligned instances, namely features, relationships and responses, and transfer them to the student model, thereby improving the accuracy of the student model's predictions.

[0094] The experimental results are shown in Table 1. AP (Average Precision) is called the average accuracy. Its calculation method is to use integration to calculate the area enclosed by the PR curve of each category and the coordinate axis. The larger the area, the higher the average accuracy of the model, and the overall precision and recall are relatively high. AP 50 and AP 75 They represent the AP values ​​corresponding to IoU thresholds of 0.5 and 0.75 respectively. S 、AP M ,AP L Respectively indicate that the target area is less than 32 2 Small target with an area greater than 32 2 Less than 96 2 Medium targets and areas larger than 96 2 The AP value corresponding to the large target.

[0095] Table 1

[0096]

[0097] The present invention uses the dense object detection method of generalized focal loss (GFL) as a benchmark model to compare the performance of the present invention with other existing knowledge distillation object detection methods. The training plan adopts a single scale (12 epochs), the student model uses ResNet-50 as the backbone network, and the teacher model uses ResNet-101 as the backbone network. The comparison results on the MS COCO val2017 validation set are shown in the figure. It can be observed that the present invention enables the GFL student model to achieve an absolute gain of +2.6 in AP value.

[0098] The performance improvement achieved by the present invention is better than the current knowledge distillation object detection methods, such as Defeat: object detection based on disentangled feature knowledge distillation; Fine-Grained: object detection based on fine-grained feature imitation knowledge distillation; GID: object detection based on general instance knowledge distillation; ERD: overcoming catastrophic forgetting in incremental object detection based on elastic response distillation; LD: object detection based on localization knowledge distillation. Compared with the recent LD, the present invention further improves the GFL student model by +0.6AP.

[0099] The present invention Figure 4 The predicted distribution of SDFE before and after distillation is visualized in the present invention, where SDFE is the current advanced dense object detection model. Figure 4 The two figures on the left show the classification response distribution of P4 (FPN second layer) of SDFE before and after distillation of the present invention, respectively. The z-axis represents the classification confidence score of each position, and the output of the P4 classification head is 50×76. Figure 4 The two figures on the right in the middle show the regression response distribution of P4 (the first layer of FPN) of SDFE before and after AIKD distillation, respectively. The z-axis represents the regression IoU score at each position, and the output of the P4 regression head is 50×76. The green arrow points to the peak of the distribution. It can be observed that before and after the distillation of the present invention, the predicted distributions of classification and regression in SDFE are still spatially aligned (i.e., at the same position), and because the present invention uses the aligned instances of the teacher model to guide the student model, the SDFE detection results are more accurate, i.e., higher classification confidence and IoU scores are obtained. Figure 5 FIG. 4 shows a comparison of the detection results of SDFE before and after the distillation of the present invention, from which it is observed that the present invention enables SDFE to obtain more accurate detection results and further suppresses the results with low IoU but high classification confidence.

Claims

1. The object detection method based on aligned instance knowledge distillation is characterized by: The steps include: Step 1. Set the real frame B GT , the classification score C and the bounding box B predicted by the classification head and the regression head in the detection head of the teacher model are input into the alignment instance elastic selection module to obtain the alignment instance AIs; the specific process is as follows: First, the classification confidence G is calculated based on the classification score C c , according to G c The threshold τ is obtained by the mean and standard deviation of c , filter out G c Greater than τ c Instances of S c ; Next, B and B GT The intersection over union (IoU) of G is used as the regression confidence, and the IoU and G c The absolute value of the difference is taken as the degree of misalignment G m ; Finally, according to G m The threshold τ is obtained by the mean and standard deviation of m , filter out G m Greater than 0 and less than τ m Alignment instance AIs; The threshold τ c The calculation process is as follows: X = Sigmoid(C) G c =max(X) τ c =mean(G c )+2×std(G c ) Where X represents the probability distribution of classification scores after Sigmoid processing, max(·) represents the maximum function, mean(·) represents the mean function, and std(·) represents the standard deviation function; The threshold τ m The calculation is: IoU=bbox_overlaps(B,B GT ) G m =abs(G c -IoU) τ m =mean(G m )-2×std(G m ) Where abs(·) represents the absolute value function, bbox_overlaps(·,·) represents the calculation of the predicted box B and the real box B GT IoU function; The example S c The calculation formula for the aligned instance AIs is: Step 2. Input the alignment instance AIs, the teacher feature map T and the student feature map S into the feature distillation module based on the alignment instance masking generation to enhance the feature representation capability of the AIs position in the student feature map S; Step 3. Input the alignment instance AIs, the teacher feature map T and the student feature map S into the alignment instance-based relation distillation module to pass the feature correlation information between AIs to the student model; Step 4. Input the aligned instance AIs into the aligned instance-based response distillation module to transfer the corresponding classification and regression reasoning information to the student model; Step 5. Construct an overall distillation loss function, train the student model through the overall distillation loss function, and output the target detection results.

2. The object detection method based on aligned instance knowledge distillation according to claim 1, characterized in that: The feature distillation module described in step 2 is specifically implemented as follows: First, the student feature map S is passed through a 1×1 convolutional layer f align , align the student feature map with the teacher feature map; Secondly, the features of the corresponding positions of AIs in the student feature map are randomly occluded by using AIs random occlusion mask M; Then, the occluded student feature map is transformed through the convolutional network Z to generate the feature map G; Finally, the l2 loss is used to distill the generated feature map G and the teacher feature map T, and the generated feature map G and the distillation loss function L AMGD The calculation formula is:

3. The object detection method based on aligned instance knowledge distillation according to claim 2 is characterized in that: The relationship distillation module described in step 3 is implemented as follows: First, the Euclidean distance ψ(·,·) is used to measure the feature correlation of different AIs in the teacher and student feature maps, respectively. i , T j represents the feature vectors corresponding to different AIs instances in the teacher feature map, S i , S i Represents the feature vectors corresponding to different AIs instances in the student feature graph; Secondly, μ(·) is used to normalize the distance; Finally, the Huber loss l(·,·) is used to transfer the correlation knowledge, and the distillation loss function L ARELD The calculation is:

4. The object detection method based on aligned instance knowledge distillation according to claim 3 is characterized in that: The implementation process of the response distillation module described in step 4 is as follows: i is used to represent the i-th aligned instance in AIs, t is the teacher model, and s is the student model. Based on the reasoning information in the classification head response, the Softmax function is used to convert the original classification scores C output by the teacher and student models into t With C s Convert to probability distribution P t With P s , the reasoning knowledge of the classification response in the teacher model is transferred to the student model through the KL divergence classification response distillation loss; For the inference information in the regression head response, each edge of the prediction box output by the teacher and student models is represented as a probability distribution B through the Softmax function s , B t , and then the KL divergence regression response distillation loss is used to transfer the reasoning knowledge of the regression response in the teacher model to the student model; the classification response distillation loss L ARESD_cls With regression response distillation loss L ARESD_reg The calculation formula is: Distillation loss function L for response distillation based on aligned instances ARESD , the calculation formula is: L ARESD =αL ARESD_cls +βL ARESD_reg Among them, α and β are the hyperparameters of the balance loss.

5. The object detection method based on aligned instance knowledge distillation according to any one of claims 1 to 4, characterized in that: The overall distillation loss function L all The calculation formula is: L all =L origin +λ0L AMGD +λ1L ARESD +λ2L ARESD Where λ0, λ1, λ2 are the hyperparameters of the balanced loss, L origin Represents the classification and regression losses used in the distilled student model.

Citation Information

Patent Citations

  • Flower classification method, system and equipment based on knowledge distillation and medium

    CN117058437A

  • Target detection training method, electronic equipment and computer readable storage medium

    CN117576381A