Multi-teacher model consistency knowledge distillation method
By adopting the multi-teacher model consistency knowledge distillation method in the field of object detection, and integrating the head and neck network consistency knowledge of the multi-teacher model, the problem of inconsistent detection results of a single teacher model is solved, and the target detection accuracy and prediction consistency are improved.
Patent Information
- Application Number
- CN202510585548.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The existing method of distilling knowledge through a single teacher model has different detection results between teacher models due to uncertainty during the training process, which affects the distillation effect and reduces the target detection accuracy.
The multi-teacher model consistency knowledge distillation method is adopted to build an architecture including core teacher model, auxiliary teacher model and student model. Through the feature fusion module and consistency knowledge of the multi-teacher model, the final loss function of the student model is constructed to improve model consistency and detection accuracy.
The classification and positioning prediction consistency of the object detection model on complex scenarios has been improved, the adverse effects of knowledge bias of a single teacher model has been alleviated, and the detection accuracy and prediction consistency of the student model has been improved.
Smart Images

Figure CN120107567A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a knowledge distillation method, and in particular to the field of computer vision, and in particular to a knowledge distillation method for consistency of multi-teacher models. Background Art
[0002] When predicting, the target detection model may have inconsistent classification and positioning predictions, resulting in reduced detection accuracy. There are two types of inconsistent predictions: 1. The positioning prediction is accurate, that is, the intersection of the predicted box and the true box is high, but the category prediction is wrong; 2. The classification of the predicted box is correct and the score is high, but the intersection of the predicted box and the true box is low. Existing methods to solve this problem through knowledge distillation technology usually use a single teacher model to teach students. However, due to the uncertainty in the training process, there are also differences in the detection results of the teacher model with the same architecture. This difference will be passed to the student model in the knowledge distillation, affecting the distillation effect and ultimately affecting the target detection effect of the image. Summary of the invention
[0003] In order to solve the problems existing in the background technology, the present invention provides a multi-teacher model consistency knowledge distillation method.
[0004] The technical solution adopted by the present invention is: The multi-teacher model consistency knowledge distillation method of the present invention comprises: Step S1: Build a multi-teacher model consistency knowledge distillation architecture including a core teacher model, several auxiliary teacher models and a student model, construct a training image set for target detection and input it into the multi-teacher model consistency knowledge distillation architecture for training, obtain the multi-teacher head and neck network consistency knowledge distillation module, and use it together with the training loss function of the student model as the final loss function of the student model until the final loss function converges to obtain the trained student model.
[0005] Step S2: Input the image to be detected into the trained student model, and after processing, output the target detection result of the image and display it on the display.
[0006] In the step S1, the multi-teacher model consistency knowledge distillation architecture also includes a feature fusion module, which connects the outputs of the core teacher model and the teacher neck networks of each auxiliary teacher model in the channel dimension; the core teacher model, each auxiliary teacher model and the student model have the same structure, including a backbone network, a neck network and a head network connected in sequence, and the training image set includes several training images. Each time the same training image is input into the core teacher model, each auxiliary teacher model and the student model, the classification output and positioning output of the core teacher model, the classification output and positioning output of each auxiliary teacher model and the classification output and positioning output of the student model are output after processing, and at the same time, the outputs of the neck networks of the core teacher model and each auxiliary teacher model are processed by the feature fusion module together to output a fusion feature map, and the fusion feature map is then processed by the head network of the core teacher model to output a fusion classification output and a fusion positioning output; the teacher head network consistency knowledge is constructed according to the classification output and positioning output of each auxiliary teacher model and the core teacher model, and the student head network is constructed according to the classification output and positioning output of the student model. The core teacher model and each auxiliary teacher model have network consistency knowledge. According to the classification output and positioning output of the core teacher model and each auxiliary teacher model, the multi-teacher consistency knowledge fusion weight and fusion consistency knowledge are constructed. The teacher head network consistency knowledge of each auxiliary teacher model, the teacher head network consistency knowledge of the core teacher model, the student head network consistency knowledge of the student model, the multi-teacher consistency knowledge fusion weight and fusion consistency knowledge are jointly constructed into a multi-teacher head network consistency knowledge distillation module; According to the output, classification output and positioning output of the neck network of the core teacher model and the output of the neck network of the student model, the core teacher neck network consistency knowledge distillation module is constructed; According to the fusion feature map, fusion classification output and fusion positioning output and the output of the neck network of the student model, the auxiliary teacher neck network consistency knowledge distillation module is constructed; Based on the core and auxiliary teacher neck network consistency knowledge distillation module, a multi-teacher neck network consistency knowledge distillation module is constructed; The multi-teacher head network consistency knowledge distillation module, the multi-teacher neck network consistency knowledge distillation module and the training loss function of the student model are jointly used as the final loss function of the student model. Both the teacher model and the student model are target detection models, and the backbone network depth of the teacher model is higher than that of the student model.
[0007] The feature fusion module is as follows: f ama = M ama ( f c ) f c =Concat ( fj ), j =1, 2, …, J M ama ( f c )= W A ( f c )+ W C ( f c ) W G ( f c )) W A ( f c )= Conv 2 d ( ReLU ( Conv 2 d ( f c ))) W C ( f c )= MLP ( AvgPool ( f c ))+ MLP ( MaxPool ( f c )) W G ( f c )= MLP ( Softmax ( Conv 2 d ( f c ))× f c ) in, f ama Represents the fused feature map; f c represents the connection feature graph; Concat ( ) represents the channel dimension connection operation; M ama ( ) represents a feature module; f j Indicatesj The output of the neck network of the teacher model, J Represents the total number of core teacher models and auxiliary teacher models; W A ( ), W C ( )and W G ( ) represent the first, second and third characteristics respectively; Conv 2 d ( ) represents a two-dimensional convolution operation; ReLU ( ) represents the activation operation of the rectified linear unit (Rectified Linear Unit); MLP ( ) represents a multilayer perceptron; AvgPool ( ) represents the average pooling operation; MaxPool ( ) represents the maximum pooling operation; Softmax ( ) represents a vector activation operation.
[0008] The multi-teacher head network consistency knowledge distillation module L MTCD as follows:
[0009] in, λ 1 and λ 2 denote the first and second hyperparameters for balancing the loss, respectively; L D and L A They represent the association knowledge distillation loss function of the multi-teacher model and the angle knowledge distillation loss function of the multi-teacher model respectively; J represents the total number of teacher models; M and N Represents the total number of detection categories and the total number of true boxes respectively; w n,j Indicates j The first n The multi-teacher consistent knowledge fusion weight of the ground truth frame; CM m,n S Represents the student model m Detection categories, n Head network consistency knowledge under real frames, CM j,m,n T Indicates jThe first m Detection categories, n Head network consistency knowledge under real frames; CM n Indicates n The head network consistency knowledge matrix under the real frame, CM 1 , CM 2 and CM 3 Respectively represent n The head network consistency knowledge matrix under the ground truth frame CM n The three head networks in the same row have consistent knowledge; l δ ( ) represents the smooth loss Huber function; ψ A ( ) represents the angle relationship calculation function; CM 1,n ~T , CM 2,n ~T and CM 3,n ~T They represent the first n The three fused consistent knowledge in the same row of the fused consistent knowledge matrix of the real box, CM 1,n S , CM 2,n S and CM 3,n S They represent the first n The three head network consistency knowledge in the same row in the head network consistency knowledge matrix of the real frame; CM 1 CM 2 CM 3 Indicates n The head network consistency knowledge matrix under the ground truth frame CM n The angular relationship between the consistent knowledge of the three head networks in the same row; e 12 and e 32 denote the first and second unit vectors respectively.
[0010] The head network consistency knowledge of the teacher model and the student model is as follows:
[0011] in, CM m,n Represents the first m Detection categories, n Head network consistency knowledge under real frames; P Represents the total set of head network output indexes of the teacher model or student model, A n and B n Respectively represent n The head network of the teacher model or the student model under the real frame outputs the first and second index sets; k m,b ´ cls Indicates m The first b The classification output corresponding to the positioning output is k m " cls Indicates m The average of the classification outputs under the detection categories, k n,b ´ loc Indicates n The ground truth box and the head network b The intersection-and-parallel ratio of the positioning outputs, k n " loc Indicates n The average of the intersection-over-union ratios of the ground-truth boxes and the localization output of the head network; T 1 Indicates softening temperature; k n,p ´ loc Indicates n The total set of ground-truth boxes and head network output indices P The p The intersection-and-parallel ratio of the positioning outputs, k p ´ loc Represents the total set of true boxes and head network output indexes P The p The intersection and union ratio set of positioning outputs; TopK ( ) indicates before selection K The index of the element; k n,a ´ loc Indicates n The ground-truth box and the first index set of the head network output A n The aThe intersection-and-union ratio of the positioning outputs.
[0012] The multi-teacher consistent knowledge fusion weights are as follows:
[0013] in, w n,j Indicates j The first n The multi-teacher consistent knowledge fusion weight of the ground truth frame; Softmax ( ) represents a vector activation operation; w n,j c and w n,j h Respectively represent j The first n The first and second weight parameters in the multi-teacher consistent knowledge fusion weight of the ground-truth frame; k b,j ´ cls Indicates j The ground truth box of the teacher model and the second index set of the head network output B n The b The classification output set corresponding to the positioning output, k n,b,j ´ cls Indicates j The first n The real box and the head network output the second index set B n The b The classification output set corresponding to the positioning output; k b,j ´ loc Indicates j The ground truth box of the teacher model and the second index set of the head network output B n The b The intersection-and-union ratio set of positioning outputs, k n,b,j ´ loc Indicates j The first n The real box and the head network output the second index set B n The b The intersection and union ratios of the positioning outputs are as follows; ( )´ indicates the averaging operation.
[0014] The fusion consistency knowledge is as follows:
[0015] in, CM m,n ~T Indicates m The first n The fusion consistency knowledge of the ground-truth boxes; P 1 , P 2 , …, P j , …, P J Respectively represent the first, second, ..., j , …, J The total set of head network output indexes output by the teacher model, J Represents the total number of teacher models.
[0016] The multi-teacher neck network consistency knowledge distillation module L MTFPD as follows:
[0017] in, λ 3 represents the third hyperparameter for balancing the loss; L CT and L AT denote the core and auxiliary teacher neck network consistency knowledge distillation modules respectively; w A represents adaptive weight; L Det AT represents the loss of training the head network that feeds the fused feature map into the core teacher model; L Det S Represents the training loss function of the student model. The final loss function of the student model L = L MTCD + L MTFPD + L Det S .
[0018] The core and auxiliary teacher neck network consistency knowledge distillation modules are as follows:
[0019] in, I Represents the core teacher model CT Or assistant teacher model AT, L I represents the core or auxiliary teacher neck network consistency knowledge distillation module; C , H and W Represent the core teacher model CT The feature maps output by the neck network of the student model and the number of channels, height, and width of the fused feature maps; Mask h,w cls,I Represents the core teacher model CT The feature map or fusion feature map output by the neck network is in the h Height and w Classification feature mask at width position, Mask h,w loc,I Represents the core teacher model CT The feature map or fusion feature map output by the neck network is in the h Height and w Position feature mask at width position; f c,h,w ,I Represents the core teacher model CT The feature map or fusion feature map output by the neck network is in the c Channel, h Height and w The response value of the width position, f c,h,w ,S Represents the student model S The feature map of the neck network output is in c Channel, h Height and w The response value of the width position; Softmax ( ) represents a vector activation operation; k h,w,m ´ cls,I Represents the teacher model I In the h Height and w The width position positioning output or the fusion positioning output corresponding to the m The classification output or fusion classification output of each detection category, k h,w,n ´ loc,I Represents the teacher model I In the h Height and w The positioning output or fusion positioning output of the width position and the n The intersection-over-union ratio of the ground-truth boxes.
[0020] The multi-teacher model consistency knowledge distillation system of the present invention comprises: The data acquisition module acquires images for target detection.
[0021] The model building module builds a multi-teacher model consistency knowledge distillation architecture including a core teacher model, several auxiliary teacher models and a student model, and trains the student model.
[0022] The target detection display module uses the trained student model to process the target detection image, outputs the target detection result of the image and displays it on the display.
[0023] The electronic device of the present invention comprises: a memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method as described above.
[0024] The readable storage medium of the present invention stores program data thereon, and the method described above is implemented when the program data is executed by a processor.
[0025] The present invention first calculates the teacher and student head network consistency knowledge and fusion consistency knowledge based on the head network classification output and the head network positioning output, realizes the quantification of the model prediction consistency, and constructs the multi-teacher consistency knowledge fusion weight based on the teacher head network classification output and the teacher head network positioning output. Then, based on the fusion consistency knowledge, the teacher head network consistency knowledge, the student head network consistency knowledge and the multi-teacher consistency knowledge fusion weight, constructs the multi-teacher head network consistency knowledge distillation module to realize the full fusion and utilization of the head network knowledge of the multi-teacher model; secondly, the core teacher model and the auxiliary teacher model are selected, and the fusion feature map is obtained through the feature fusion module. The combined feature map is input into the core teacher model for training, and the fused head network classification output and the fused head network positioning output are obtained to realize the fusion of the teacher model neck network knowledge. Then, based on the teacher neck network output, the teacher head network classification output and the teacher head network positioning output of the core teacher model, a core teacher neck network consistency knowledge distillation module is constructed. Based on the teacher model fusion feature map, the fused head network classification output and the fused head network positioning output, an auxiliary teacher neck network consistency knowledge distillation module is constructed. Finally, based on the core teacher neck network consistency knowledge distillation module and the auxiliary teacher neck network consistency knowledge distillation module, a multi-teacher neck network consistency knowledge distillation module is constructed.
[0026] The beneficial effects of the present invention are: The method of the present invention improves the classification and positioning prediction consistency of the target detection model for complex scene detection by fusing the head network consistency knowledge and neck network consistency knowledge of multiple teacher models, and alleviates the adverse effects of the knowledge bias of a single teacher model. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a flow chart of the method of the present invention; Figure 2 A schematic diagram of a knowledge distillation architecture for consistency of a multi-teacher model in an embodiment of the present invention; Figure 3 This is a schematic diagram of inconsistent predictions, where Figure 3 (a) is a schematic diagram of the first inconsistent prediction. Figure 3 (b) is a schematic diagram of the second inconsistent prediction; Figure 4 This is a comparison chart of the impact of teacher model detection differences on the effect of knowledge distillation, where: Figure 4 (a) is a schematic diagram of the teacher model’s error detection. Figure 4 (b) is a schematic diagram of the correct detection of the teacher model. Figure 4 (c) is a schematic diagram of the incorrect prediction results of the student model that inherits the incorrectly detected teacher model. Figure 4 (d) is a schematic diagram of the correct prediction results of the student model that inherits the correctly detected teacher model; Figure 5 This is a comparison chart of the effects of the student model on the artifact detection dataset before and after distillation, where: Figure 5 (a) is a schematic diagram of the detection results of the initial student model before distillation. Figure 5 (b) is a schematic diagram of the detection results of the initial student model after distillation. DETAILED DESCRIPTION
[0028] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.
[0029] The embodiment of the present invention is carried out on a self-built workpiece detection dataset, which contains 6 types of workpieces, named with numbers 1-6, workpiece 1 is the first fixed block, workpiece 2 is the first connecting block, workpiece 3 is the first flange, workpiece 4 is the second fixed block, workpiece 5 is the second connecting block, and workpiece 6 is the second flange. Under the above workpiece dataset, the multi-teacher model consistency knowledge distillation method of the present invention is implemented to solve the problem of inconsistent predictions. Figure 3 It shows the prediction inconsistency problem when using RetinaNet-R18 detection, such as Figure 3As shown in (a), this is the first type of prediction inconsistency problem. Its positioning prediction is accurate, that is, the intersection of the predicted box and the true box is relatively high, but the category prediction is wrong. It can be seen that the position of the predicted box of the workpiece and the target true box in the figure basically coincides, but the model mistakenly predicts the 6th type of workpiece as the 3rd type of workpiece; Figure 3 (b) shows the second type of prediction inconsistency problem. The classification of the prediction box is correct and the score is high, but the intersection-union ratio between the prediction box and the true box is low. It can be seen that the category prediction of the workpiece in the figure is correct and the classification score reaches 0.69, but there is a large difference between the position of the prediction box and the target true box, and the intersection-union ratio between the two is only 0.46, which is low.
[0030] like Figure 1 As shown, the multi-teacher model consistency knowledge distillation method of the present invention specifically includes the following steps: First, we build a multi-teacher model consistency knowledge distillation architecture including a core teacher model, several auxiliary teacher models, and a student model. Both the teacher model and the student model are target detection models, and the backbone network depth of the teacher model is higher than that of the student model. Figure 2 As shown in FIG. 1 , the multi-teacher model consistency knowledge distillation architecture also includes a feature fusion module, which connects the outputs of the teacher neck network of the core teacher model and each auxiliary teacher model in the channel dimension. The feature fusion module is as follows: f ama = M ama ( f c ) f c =Concat ( f j ), j =1, 2, …, J M ama ( f c )= W A ( f c )+ W C ( f c ) W G ( f c )) W A ( f c )= Conv 2 d ( ReLU ( Conv 2 d ( f c ))) W C ( f c )= MLP ( AvgPool ( f c ))+ MLP ( MaxPool ( f c )) W G ( f c )= MLP ( Softmax ( Conv 2 d ( f c ))× f c ) MLP ( f c )= Conv 2 d ( ReLU ( LN ( Conv 2 d ( f c )))) in, f ama Represents the fused feature map; f c represents the connection feature graph; Concat ( ) represents the channel dimension connection operation; M ama ( ) represents a feature module; f j Indicates j The output of the neck network of the teacher model, J Represents the total number of core teacher models and auxiliary teacher models; W A ( ), W C ( )and W G ( ) represent the first, second and third characteristics respectively; Conv 2 d ( ) represents a two-dimensional convolution operation using a 1×1 convolution kernel; ReLU ( ) represents the activation operation of the rectified linear unit (RectifiedLinear Unit); MLP ( ) represents a multi-layer perceptron, LN ( ) represents layer normalization operation; AvgPool ( ) represents the average pooling operation; MaxPool ( ) represents the maximum pooling operation; Softmax ( ) represents a vector activation operation.
[0031] The structures of the core teacher model, each auxiliary teacher model and the student model are the same, including a backbone network, a neck network and a head network connected in sequence. The training image set includes several training images. Each time, the same training image is input into the core teacher model, each auxiliary teacher model and the student model. After processing, the classification output and positioning output of the core teacher model, the classification output and positioning output of each auxiliary teacher model and the classification output and positioning output of the student model are output. At the same time, the outputs of the neck networks of the core teacher model and each auxiliary teacher model are processed by the feature fusion module to output a fusion feature map. The fusion feature map is then processed by the head network of the core teacher model to output a fusion classification output and a fusion positioning output. The teacher head network consistency knowledge is constructed according to the classification output and positioning output of each auxiliary teacher model and the core teacher model, and the student head network consistency knowledge is constructed according to the classification output and positioning output of the student model, as follows:
[0032] in, CM m,n Represents the first m Detection categories, n Head network consistency knowledge under real frames; P Represents the total set of head network output indexes of the teacher model or student model, A n and B n Respectively represent n The head network of the teacher model or the student model under the real frame outputs the first and second index sets, P Include A n , A n Include B n ; km,b ´ cls Indicates m The first b The classification output corresponding to the positioning output is k m " cls Indicates m The average of the classification outputs under the detection categories, k n,b ´ loc Indicates n The ground truth box and the head network b The intersection-and-parallel ratio of the positioning outputs, k n " loc Indicates n The average of the intersection-over-union ratios of the ground-truth boxes and the localization output of the head network; T 1 Indicates softening temperature; k n,p ´ loc Indicates n The total set of ground-truth boxes and head network output indices P The p The intersection-and-parallel ratio of the positioning outputs, k p ´ loc Represents the total set of true boxes and head network output indexes P The p The intersection and union ratio set of positioning outputs; TopK ( ) indicates before selection K The index of the element; k n,a ´ loc Indicates n The ground-truth box and the first index set of the head network output A n The a The intersection and union ratio of the positioning outputs. The multi-teacher consistent knowledge fusion weights and fusion consistent knowledge are constructed based on the classification outputs and positioning outputs of the core teacher model and each auxiliary teacher model. The multi-teacher consistent knowledge fusion weights are as follows:
[0033] in, w n,j Indicates j The first n The multi-teacher consistent knowledge fusion weight of the ground truth frame; Softmax ( ) represents a vector activation operation; w n,j c andw n,j h Respectively represent j The first n The first and second weight parameters in the multi-teacher consistent knowledge fusion weight of the ground-truth frame; k b,j ´ cls Indicates j The ground truth box of the teacher model and the second index set of the head network output B n The b The classification output set corresponding to the positioning output, k n,b,j ´ cls Indicates j The first n The real box and the head network output the second index set B n The b The classification output set corresponding to the positioning output; k b,j ´ loc Indicates j The ground truth box of the teacher model and the second index set of the head network output B n The b The intersection-and-union ratio set of positioning outputs, k n,b,j ´ loc Indicates j The first n The real box and the head network output the second index set B n The b The intersection and union ratios of the positioning outputs are as follows; ( )´ indicates the averaging operation.
[0034] The fusion consistency knowledge is as follows:
[0035] in, CM m,n ~T Indicates m The first n The fusion consistency knowledge of the ground-truth boxes; P 1 , P 2 , …, P j , …, P J Respectively represent the first, second, ...,j , …, J The total set of head network output indexes output by the teacher model, J Represents the total number of teacher models.
[0036] The teacher head network consistency knowledge of each auxiliary teacher model, the teacher head network consistency knowledge of the core teacher model, the student head network consistency knowledge of the student model, the multi-teacher consistency knowledge fusion weight and the fusion consistency knowledge are jointly constructed into a multi-teacher head network consistency knowledge distillation module L MTCD , as follows:
[0037] in, λ 1 and λ 2 denote the first and second hyperparameters for balancing the loss, respectively; L D and L A They represent the association knowledge distillation loss function of the multi-teacher model and the angle knowledge distillation loss function of the multi-teacher model respectively; J represents the total number of teacher models; M and N Represents the total number of detection categories and the total number of true boxes respectively; w n,j Indicates j The first n The multi-teacher consistent knowledge fusion weight of the ground truth frame; CM m,n S Represents the student model m Detection categories, n Head network consistency knowledge under real frames, CM j,m,n T Indicates j The first m Detection categories, n Head network consistency knowledge under real frames; CM n Indicates n The head network consistency knowledge matrix under the real frame, CM 1 , CM 2 and CM 3 Respectively represent n The head network consistency knowledge matrix under the ground truth frame CM nThe three head networks in the same row have consistent knowledge; l δ ( ) represents the smooth loss Huber function; ψ A ( ) represents the angle relationship calculation function; CM 1,n ~T , CM 2,n ~T and CM 3,n ~T They represent the first n The three fused consistent knowledge in the same row of the fused consistent knowledge matrix of the real box, CM 1,n S , CM 2,n S and CM 3,n S They represent the first n The three head network consistency knowledge in the same row in the head network consistency knowledge matrix of the real frame; CM 1 CM 2 CM 3 Indicates n The head network consistency knowledge matrix under the ground truth frame CM n The angular relationship between the consistent knowledge of the three head networks in the same row; e 12 and e 32 They represent the first and second unit vectors, respectively, and their directions are from CM 1 arrive CM 2 and from CM 3 arrive CM 2 .
[0038] According to the output, classification output and positioning output of the neck network of the core teacher model and the output of the neck network of the student model, the core teacher neck network consistency knowledge distillation module is constructed. According to the fused feature map, fused classification output and fused positioning output and the output of the neck network of the student model, the auxiliary teacher neck network consistency knowledge distillation module is constructed as follows:
[0039] in,I Represents the core teacher model CT Or assistant teacher model AT , L I represents the core or auxiliary teacher neck network consistency knowledge distillation module; C , H and W Represent the core teacher model CT The feature maps output by the neck network of the student model and the number of channels, height, and width of the fused feature maps; Mask h,w cls,I Represents the core teacher model CT The feature map or fusion feature map output by the neck network is in the h Height and w Classification feature mask at width position, Mask h,w loc,I Represents the core teacher model CT The feature map or fusion feature map output by the neck network is in the h Height and w Position feature mask at width position; f c,h,w ,I Represents the core teacher model CT The feature map or fusion feature map output by the neck network is in the c Channel, h Height and w The response value of the width position, f c,h,w ,S Represents the student model S The feature map of the neck network output is in c Channel, h Height and w The response value of the width position; Softmax ( ) represents a vector activation operation; k h,w,m ´ cls,I Represents the teacher model I In the h Height and w The width position positioning output or the fusion positioning output corresponding to the m The classification output or fusion classification output of each detection category, k h,w,n ´ loc,I Represents the teacher model I In the h Height and wThe positioning output or fusion positioning output of the width position and the n The intersection-over-union ratio of the ground-truth boxes.
[0040] Constructing a multi-teacher neck network consistency knowledge distillation module based on the core and auxiliary teacher neck network consistency knowledge distillation module L MTFPD ,as follows:
[0041] in, λ 3 represents the third hyperparameter for balancing the loss; L CT and L AT denote the core and auxiliary teacher neck network consistency knowledge distillation modules respectively; w A represents adaptive weight; L Det AT represents the loss of training the head network that feeds the fused feature map into the core teacher model; L Det S Represents the training loss function of the student model.
[0042] The constructed training image set for target detection is input into the multi-teacher model consistency knowledge distillation architecture for training, and the multi-teacher head and neck network consistency knowledge distillation module is obtained. The training loss function of the student model is used together with the student model as the final loss function of the student model. The final loss function of the student model L = L MTCD + L MTFPD + L Det S , until the final loss function converges and the trained student model is obtained.
[0043] The present invention performs inheritance verification, such as Figure 4 (a) and Figure 4As shown in (b), the detection results of the same artifact by RetinaNet-R101 teacher model 1 and RetinaNet-R101 teacher model 2 obtained using two different training seeds. It can be seen that RetinaNet-R101 teacher model 1 mistakenly detects artifact 6 as artifact 3, while RetinaNet-R101 teacher model 2 can correctly identify the artifact. The above two teacher models correspond to RetinaNet-R18 student model 1 and RetinaNet-R18 student model 2, respectively. The knowledge distillation of the student models of the above two teacher models is performed with the same random number seed, as shown in Figure 4 (c) and Figure 4 As shown in (d), it can be seen that the RetinaNet-R18 student model 1 inherits the incorrect prediction results of its teacher model, while the RetinaNet-R18 student model 2 is able to correctly identify artifact 6.
[0044] After the training is completed, the image to be detected is input into the trained student model, and the target detection result of the image is output after processing and displayed on the display. In the specific implementation of the present invention, the target detection RetinaNet-R101 model is used as the teacher model, and the target detection RetinaNet-R18 model is used as the student model. The student model after distillation is obtained by using the multi-teacher model consistency knowledge distillation method of the present invention, and the existing target detection task balanced distillation method TBD (Task-balanced distillation for object detection) is used for comparison. The evaluation is performed on a self-built workpiece detection data set. In the evaluation, the evaluation result is judged according to the detection accuracy evaluation index mAP (mean Average Precision) and the prediction consistency index SPCC (Spearman Correlation Coefficient). The mean average precision mAP is the average value of the average precision over all categories. For multi-category target detection tasks, the average precision of each category is usually calculated, and then the average value is taken. The higher the value of the mean average precision, the better the detection performance of the model; the prediction consistency index SPCC is used to test the correlation between the prediction box classification and the positioning sorting. The higher the value, the higher the model prediction consistency. The evaluation results are shown in Table 1. It can be seen that the method of the present invention can effectively improve the detection accuracy and prediction consistency of the student model, and is superior to the existing target detection task balanced distillation method TBD.
[0045] Table 1
[0046] like Figure 5As shown in Figure 2, the comparison of the detection effect of the student model before and after partial distillation is shown. Figure 5 As shown in (a), it is the detection result of the initial student model. Figure 5 (b) shows the detection result of the student model after distillation. The numbers in the upper left corner of the prediction box represent the workpiece category and classification confidence respectively. After distillation by the method of the present invention, the false detection and missed detection results of the student model are reduced, and the prediction consistency is improved.
[0047] The present invention also designs a multi-teacher model consistency knowledge distillation system, which includes a data acquisition module, a model construction module and a target detection display module. The data acquisition module acquires images for target detection; the model construction module builds a multi-teacher model consistency knowledge distillation architecture including a core teacher model, several auxiliary teacher models and a student model and trains the student model; the target detection display module uses the trained student model to process the target detection image, outputs the target detection result of the image and displays it on the display.
[0048] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some or all of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.
[0049] It should be understood by those skilled in the art that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program codes. The schemes in the embodiments of the present application may be implemented in various computer languages. The present application is described according to the flowcharts of the methods, systems and computer program products of the embodiments of the present application.
[0050] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the present invention is intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0051] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of equivalent technologies of the present invention, the present application is also intended to include these modifications and variations.
Claims
1. A knowledge distillation method for multi-teacher model consistency, characterized in that: include: Step S1: Build a multi-teacher model consistency knowledge distillation architecture including a core teacher model, several auxiliary teacher models and a student model, construct a training image set for target detection and input it into the multi-teacher model consistency knowledge distillation architecture for training, obtain a multi-teacher head and neck network consistency knowledge distillation module, and use it together with the training loss function of the student model as the final loss function of the student model until the final loss function converges to obtain a trained student model; Step S2: Input the image to be detected into the trained student model, and after processing, output the target detection result of the image and display it on the display.
2. The method for knowledge distillation based on multi-teacher model consistency according to claim 1, characterized in that: In the step S1, the multi-teacher model consistency knowledge distillation architecture also includes a feature fusion module, which connects the outputs of the core teacher model and the teacher neck networks of each auxiliary teacher model in the channel dimension; the core teacher model, each auxiliary teacher model and the student model have the same structure, including a backbone network, a neck network and a head network connected in sequence, and the training image set includes several training images. Each time the same training image is input into the core teacher model, each auxiliary teacher model and the student model, the classification output and positioning output of the core teacher model, the classification output and positioning output of each auxiliary teacher model and the classification output and positioning output of the student model are output after processing, and at the same time, the outputs of the neck networks of the core teacher model and each auxiliary teacher model are processed by the feature fusion module together to output a fusion feature map, and the fusion feature map is then processed by the head network of the core teacher model to output a fusion classification output and a fusion positioning output; the teacher head network consistency knowledge is constructed according to the classification output and positioning output of each auxiliary teacher model and the core teacher model, and the student head network is constructed according to the classification output and positioning output of the student model. Network consistency knowledge, according to the classification output and positioning output of the core teacher model and each auxiliary teacher model, the multi-teacher consistency knowledge fusion weight and fusion consistency knowledge are constructed, and the teacher head network consistency knowledge of each auxiliary teacher model, the teacher head network consistency knowledge of the core teacher model, the student head network consistency knowledge of the student model, the multi-teacher consistency knowledge fusion weight and fusion consistency knowledge are jointly constructed into a multi-teacher head network consistency knowledge distillation module; according to the output, classification output and positioning output of the neck network of the core teacher model and the output of the neck network of the student model, the core teacher neck network consistency knowledge distillation module is constructed, according to the fusion feature map, the fusion classification output and the fusion positioning output and the output of the neck network of the student model, the auxiliary teacher neck network consistency knowledge distillation module is constructed, and the multi-teacher neck network consistency knowledge distillation module is constructed based on the core and auxiliary teacher neck network consistency knowledge distillation modules; the multi-teacher head network consistency knowledge distillation module, the multi-teacher neck network consistency knowledge distillation module and the training loss function of the student model are jointly used as the final loss function of the student model.
3. The method for knowledge distillation based on multi-teacher model consistency according to claim 2, characterized in that: The feature fusion module is as follows: f ama = M ama ( f c ) f c =Concat ( f j ), j =1、2、…、 J M ama ( f c )= W A ( f c )+ W C ( f c )· W G ( f c )) W A ( f c )= Conv 2 d ( ReLU ( Conv 2 d ( f c ))) W C ( f c )= MLP ( AvgPool ( f c ))+ MLP ( MaxPool ( f c )) W G ( f c )= MLP ( Softmax ( Conv 2 d ( f c ))× f c ) in, f ama Represents the fused feature map; f c represents the connection feature graph; Concat ( ) represents the channel dimension connection operation; M ama () indicates feature module; f j Indicates j The output of the neck network of the teacher model, J Represents the total number of core teacher models and auxiliary teacher models; W A (), W C ()and W G () represent the first, second and third features respectively; Conv 2 d ( ) represents a two-dimensional convolution operation; ReLU ( ) represents the activation operation of the rectified linear unit; MLP ( ) represents a multilayer perceptron; AvgPool () represents the average pooling operation; MaxPool () represents the maximum pooling operation; Softmax ( ) represents a vector activation operation.
4. The method for knowledge distillation based on multi-teacher model consistency according to claim 2, characterized in that: The multi-teacher head network consistency knowledge distillation module L MTCD as follows: in, λ 1 and λ 2 denotes the first and second hyperparameters respectively; L D and L A They represent the association knowledge distillation loss function of the multi-teacher model and the angle knowledge distillation loss function of the multi-teacher model respectively; J represents the total number of teacher models; M and N Represents the total number of detection categories and the total number of true boxes respectively; w n,j Indicates j The first n The multi-teacher consistent knowledge fusion weight of the ground truth frame; CM m,n S Represents the student model m Detection categories, n Head network consistency knowledge under real frames, CM j,m,n T Indicates j The first m Detection categories, n Head network consistency knowledge under real frames; CM n Indicates n The head network consistency knowledge matrix under the real frame, CM 1. CM 2 and CM 3 respectively represent the n The head network consistency knowledge matrix under the ground truth frame CM n The three head networks in the same row have consistent knowledge; l δ ( ) represents the smooth loss Huber function; ψ A ( ) represents the angle relationship calculation function; CM 1,n ~T , CM 2,n ~T and CM 3,n ~T They represent the first n The three fused consistent knowledge in the same row of the fused consistent knowledge matrix of the real box, CM 1,n S , CM 2,n S and CM 3,n S They represent the first n The three head network consistency knowledge in the same row in the head network consistency knowledge matrix of the real frame; CM 1 CM 2 CM 3 means the n The head network consistency knowledge matrix under the ground truth frame CM n The angular relationship between the consistent knowledge of the three head networks in the same row; e 12 and e 32 denote the first and second unit vectors respectively.
5. The method for knowledge distillation based on multi-teacher model consistency according to claim 2, characterized in that: The head network consistency knowledge of the teacher model and the student model is as follows: in, CM m,n Represents the first m Detection categories, n Head network consistency knowledge under real frames; P Represents the total set of head network output indexes of the teacher model or student model, A n and B n Respectively represent n The head network of the teacher model or the student model under the real frame outputs the first and second index sets; k m,b ´ cls Indicates m The first b The classification output corresponding to the positioning output is k m " cls Indicates m The average of the classification outputs under the detection categories, k n,b ´ loc Indicates n The ground truth box and the head network b The intersection-and-parallel ratio of the positioning outputs, k n " loc Indicates n The average of the intersection-over-union ratios of the ground-truth boxes and the localization output of the head network; T 1 indicates softening temperature; k n,p ´ loc Indicates n The total set of ground-truth boxes and head network output indices P The p The intersection-and-parallel ratio of the positioning outputs, k p ´ loc Represents the total set of true boxes and head network output indexes P The p The intersection and union ratio set of positioning outputs; TopK ( ) indicates before selection K The index of the element; k n,a ´ loc Indicates n The ground-truth box and the first index set of the head network output A n The a The intersection-and-union ratio of the positioning outputs.
6. The method for knowledge distillation based on multi-teacher model consistency according to claim 2, characterized in that: The multi-teacher consistent knowledge fusion weights are as follows: in, w n,j Indicates j The first n The multi-teacher consistent knowledge fusion weight of the ground truth frame; Softmax ( ) represents a vector activation operation; w n,j c and w n,j h Respectively represent j The first n The first and second weight parameters in the multi-teacher consistent knowledge fusion weight of the ground-truth frame; k b,j ´ cls Indicates j The ground truth box of the teacher model and the second index set of the head network output B n The b The classification output set corresponding to the positioning output, k n,b,j ´ cls Indicates j The first n The real box and the head network output the second index set B n The b The classification output set corresponding to the positioning output; k b,j ´ loc Indicates j The ground truth box of the teacher model and the second index set of the head network output B n The b The intersection-and-union ratio set of positioning outputs, k n,b,j ´ loc Indicates j The first n The real box and the head network output the second index set B n The b The intersection and union ratios of the positioning outputs are as follows; ( )´ indicates the averaging operation.
7. The method for knowledge distillation based on multi-teacher model consistency according to claim 5, characterized in that: The fusion consistency knowledge is as follows: in, CM m,n ~T Indicates m The teacher model under the detection category n The fusion consistency knowledge of the ground-truth boxes; P 1. P 2. … P j , …, P J Respectively represent the first, second, ..., j , …, J The total set of head network output indexes output by the teacher model, J Represents the total number of teacher models.
8. The method for knowledge distillation based on multi-teacher model consistency according to claim 2, characterized in that: The multi-teacher neck network consistency knowledge distillation module L MTFPD as follows: in, λ 3 represents the third hyperparameter; L CT and L AT denote the core and auxiliary teacher neck network consistency knowledge distillation modules respectively; w A represents adaptive weight; L Det AT represents the loss of training the head network that feeds the fused feature map into the core teacher model; L Det S Represents the training loss function of the student model.
9. The method for knowledge distillation based on multi-teacher model consistency according to claim 2, characterized in that: The core and auxiliary teacher neck network consistency knowledge distillation modules are as follows: in, I Represents the core teacher model CT Or assistant teacher model AT , L I represents the core or auxiliary teacher neck network consistency knowledge distillation module; C , H and W Represent the core teacher model CT The feature maps output by the neck network of the student model and the number of channels, height, and width of the fused feature maps; Mask h,w cls,I Represents the core teacher model CT The feature map or fusion feature map output by the neck network is in the h Height and w Classification feature mask at width position, Mask h,w loc,I Represents the core teacher model CT The feature map or fusion feature map output by the neck network is in the h Height and w Position feature mask at width position; f c,h,w ,I Represents the core teacher model CT The feature map or fusion feature map output by the neck network is in the c Channel, h Height and w The response value of the width position, f c,h,w ,S Represents the student model S The feature map of the neck network output is in c Channel, h Height and w The response value of the width position; Softmax ( ) represents a vector activation operation; k h,w,m ´ cls,I Represents the teacher model I In the h Height and w The width position positioning output or the fusion positioning output corresponding to the m The classification output or fusion classification output of each detection category, k h,w,n ´ loc,I Represents the teacher model I In the h Height and w The positioning output or fusion positioning output of the width position and the n The intersection-over-union ratio of the ground-truth boxes.
10. A multi-teacher model consistency knowledge distillation system applicable to the method according to any one of claims 1 to 9, characterized in that: include: Data acquisition module, acquiring images for target detection; Model building module, builds a multi-teacher model consistency knowledge distillation architecture including a core teacher model, several auxiliary teacher models and a student model, and trains the student model; The target detection display module uses the trained student model to process the target detection image, outputs the target detection result of the image and displays it on the display.
Citation Information
Patent Citations
Multi-teacher self-adaptive joint knowledge distillation
CN112418343A
Knowledge distillation-based lightweight SAR (Synthetic Aperture Radar) image target detection method
CN116935213A
Yolov7 grapefruit counting method fusing multi-teacher knowledge distillation
CN117496509A
Industrial anomaly detection method and system based on multi-teacher network knowledge distillation
CN117609925A
Continuous learning target detection method and system based on knowledge distillation
CN118114724A