A target detection method based on knowledge distillation

By constructing the teacher-taught assistant-student model, and using the teaching assistant model to process and remove redundant features extracted by the teacher model, the problem of insufficient in-depth and overfitting of the teacher model features is solved, and the target detection effect is improved.

CN117292192BActive Publication Date: 2025-06-13HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311282250.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-28
Publication Date
2025-06-13
Estimated Expiration
2043-09-28

AI Technical Summary

Technical Problem

In the existing object detection algorithm based on knowledge distillation, the features extracted by the teacher model are not deep enough and are easily disturbed by noise. The teacher model is similar to the network structure of the student model, and it is prone to overfitting, affecting the detection effect.

Method used

The teacher-taught assistant-student model is constructed, where the teacher model includes P-conv and four-layer stacked convolutional layers, the teaching assistant model includes feature processing layer, feature splicing layer, spatial attention module and Bottleneck convolutional block, and the student model includes four-layer stacked deconvolutional layers. The features extracted by the teacher model are processed and de-redundant by the teacher model, so as to enhance the learning ability of the student model and avoid overfitting.

Benefits of technology

By introducing the teaching assistant model, the teacher model can extract deeper features, and the teaching assistant model further processes and streamlinesses the features, so that the student model can better learn the knowledge of the teacher model, avoid overfitting, and improve the target detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117292192B_ABST
    Figure CN117292192B_ABST
Patent Text Reader

Abstract

The present invention discloses an object detection method based on knowledge distillation, which includes constructing a teacher-assistant-student model, and using a pre-trained teacher model to obtain multi-level feature maps of an image; using an assistant module to perform upsampling and splicing on the feature maps, and after spatial attention processing, fusing the multi-scale feature maps extracted by the teacher model into one feature map; using a student model to perform multiple deconvolution calculations on the transmitted feature map to obtain a feature map with the same feature dimension as that extracted by the teacher model; calculating the cosine similarity between the feature maps obtained by the teacher model and the student model, and restoring the calculation result map to the size of the input image through bilinear interpolation sampling, and then taking the product to obtain a feature map Ma1, calculating the feature scores of each value in the feature map Ma1, and drawing an image heat map. The present invention improves the object detection effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a target detection method, and more particularly to a target detection method based on knowledge distillation. Background Art

[0002] With the continuous development of deep learning technology, image target detection algorithms have been continuously optimized and iterated, and have begun to be applied to industrial production. However, most target detection algorithms require a large amount of space and resources, and the emergence of lightweight models has well solved this problem. The target detection algorithm based on knowledge distillation uses the originally bulky and complex network as the teacher model and a smaller model as the student model. By having the student model learn the knowledge of the teacher model, a detection effect similar to that of the teacher model can be achieved. When deploying the model, only the lightweight student model can be deployed, or the teacher model and the student model can be deployed simultaneously for comparison and verification. However, the lightweight teacher-student model still has the following problems: the features extracted by the teacher model are not deep enough and are easily interfered by noise. The network structures of the teacher model and the student model are similar, and overfitting is likely to occur, thus affecting the detection effect. Summary of the Invention

[0003] Object of the Invention: The object of the present invention is to provide a target detection method based on knowledge distillation with good detection effect.

[0004] Technical Solution: The target detection method based on knowledge distillation according to the present invention includes:

[0005] (1) Construct a teacher-assistant-student model, where the teacher model includes P-conv and four stacked convolutional layers Layer1 to Layer4; the assistant model includes four feature processing layers Af1 to Af4, a feature splicing layer, a spatial attention module, and a Bottleneck convolutional block. The feature processing layer Af1 includes three convolutional blocks, the feature processing layer Af2 includes two convolutional blocks, and the feature processing layers Af3 and Af4 each include one convolutional block. The convolutional block is sequentially connected by a convolutional layer, a batch normalization layer BatchNorm, and a Relu activation function layer; the student model includes four stacked deconvolutional layers Slayer1 to Slayer4;

[0006] The pre-trained teacher model extracts two sets of deep feature maps F1, F2 and two sets of shallow feature maps F3, F4 of the image and inputs them into the teaching assistant model; the sizes of the feature maps F1 to F4 increase in sequence; the deep feature map F1 is input into the feature processing layer Af1 to obtain the feature map Al1; the deep feature map F2 is input into the feature processing layer Af2 to obtain the feature map Al2; the shallow feature map F3 is input into the feature processing layer Af3 to obtain the feature map Al3; the shallow feature map F4 is downsampled to obtain the feature map Al4; the feature maps Al1 to Al4 have the same dimension; after pairwise splicing of the feature maps Al1, Al2, Al2, Al3, Al3, Al4, the feature maps A1, A2, A3 are obtained; after the feature maps A1, A2, A3 are respectively processed by the feature processing layer Af4 and the spatial attention module, they are horizontally spliced to obtain a feature map At; after the feature map At is accelerated by the Bottleneck convolution block, it is input into the student model; the feature map At is input into the deconvolution layer Slayer1 to obtain the feature map F5; the feature map F5 is input into the deconvolution layer Slayer2 to obtain the feature map F6; the feature map F6 is input into the deconvolution layer Slayer3 to obtain the feature map F7; the feature map F7 is input into the deconvolution layer Slayer4 to obtain the feature map F8;

[0007] (2) Train the teacher-teaching assistant-student model, where the parameters of the teacher model are fixed and do not participate in the training; the student model and the teaching assistant module participate in the training; calculate the cosine similarities of F1 and F8, F2 and F7, F3 and F6, F4 and F5 respectively, and take the mean and add them as the loss function;

[0008] (3) Input the image to be detected into the trained teacher-teaching assistant-student model, and the teacher model and the student model respectively obtain a feature map, and calculate the cosine similarities of the four groups of feature maps respectively; restore the feature map to the original image size through bilinear interpolation sampling, and perform a product operation on the four groups of cosine similarity calculation results after bilinear sampling upsampling to obtain the feature map Ma1. After smoothing Ma1 using Gaussian filtering, calculate the feature score of each value in Ma1 to obtain the feature score map; draw an image heat map according to the feature score map. The higher the feature score, the brighter it is in the image heat map, so as to achieve object detection.

[0009] Further, P-conv includes a convolutional layer, a batch normalization layer, a Relu activation function layer and a pooling layer. The convolutional layer has a stride of 2, a padding of 3, and a convolutional kernel size of 7×7; the pooling layer has a stride of 2, a padding of 1, and a pooling window of 3×3.

[0010] Further, the convolutional layers Layer1 to Layer4 include a set of t_conv convolutional block groups and a parallel downsampling layer DownSample, followed by a feature concatenation layer; where Layer1 includes 3 t_conv convolutional blocks, Layer2 includes 4 t_conv convolutional blocks, Layer3 includes 6 t_conv convolutional blocks, and Layer4 includes 3 t_conv convolutional blocks; the t_conv convolutional block is composed of a 3×3 convolution conv2 with a stride of 1 and two 1×1 convolutions conv1 and conv3 connected in series, and a BatchNorm layer needs to be connected after each convolution; the downsampling layer includes a 1×1 conv convolution and a BatchNorm layer.

[0011] Further, the transposed convolutional layers Slayer1 to Slayer4 include a tr_conv transposed convolutional block, multiple connected t_conv convolutional block groups, and a parallel upsampling layer UpSample, followed by a feature concatenation layer; where Slayer1 includes 1 tr_conv transposed convolutional block and 2 t_conv convolutional blocks, Slayer2 includes 1 tr_conv transposed convolutional block and 5 t_conv convolutional blocks, Slayer3 includes 1 tr_conv transposed convolutional block and 3 t_conv convolutional blocks, and Slayer4 includes 1 tr_conv transposed convolutional block and 2 t_conv convolutional blocks; the tr_conv transposed convolutional block includes a 2×2 convtrans transposed convolution with a stride of 2 and a BatchNorm layer; the upsampling layer UpSample includes a 2×2 convtrans transposed convolution with a stride of 2 and a BatchNorm layer.

[0012] Further, the feature maps output by the teacher model and the student model have exactly the same dimensional size.

[0013] Further, the size of feature map F1 is 256×64×64, the size of feature map F2 is 512×32×32, the size of feature map F3 is 1024×16×16, and the size of feature map F4 is 2048×8×8; the size of feature map F5 is 2048×8×8, the size of feature map F6 is 1024×16×16, the size of feature map F7 is 512×32×32, and the size of feature map F8 is 256×64×64.

[0014] Further, the size of the image input to the teacher model is 256×256×3, the sizes of feature maps Al1 to Al4 are all 1024×8×8, the sizes of feature maps A1 to A3 are 2048×8×8, and the size of feature map At is 6144×8×8.

[0015] Further, the formula for calculating the feature score is: Ma1 = (ma1 - ma1.min) / (ma1.max - ma1.min)

[0016] Where ma1.min is the minimum value in the feature matrix Ma1, and ma1.max is the maximum value in the feature matrix Ma1.

[0017] Further, the teacher - teaching assistant - student model is trained using the MvTec AD dataset.

[0018] Further, the teacher model is pre - trained on the ImageNet dataset.

[0019] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages:

[0020] The present invention introduces a teaching assistant model and constructs a new teacher - teaching assistant - student model. The teacher model can extract deeper - level features. The teaching assistant model further processes the features extracted by the teacher model, removes redundant features, makes the feature map concise but contains multi - level features, and amplifies the features through spatial attention, enabling the student model to better learn the knowledge of the teacher model. In addition, the teacher model and the student model have non - similar structures, avoiding the problem of overfitting. Thus, the present invention improves the target detection effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a schematic diagram of the teacher - teaching assistant - student model structure in an embodiment of the present application;

[0022] Figure 2 is a schematic diagram of the Layer layer structure in an embodiment of the present application;

[0023] Figure 3 is a schematic diagram of the Slayer layer structure in an embodiment of the present application;

[0024] Figure 4 is a flowchart of the loss calculation during the training process in an embodiment of the present application;

[0025] Figure 5 is a flowchart of calculating the feature score of the image to be detected in an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION

[0026] The present invention will be further described below with reference to the accompanying drawings.

[0027] A target detection method based on knowledge distillation includes the following steps:

[0028] (1) Construct a teacher - teaching assistant - student model, as Figure 1As shown, the teacher model includes a P-conv (pre-convolution block) and four stacked convolutional layers Layer1 to Layer4; the teaching assistant model includes four feature processing layers Af1 to Af4, a feature concatenation layer, a spatial attention module, and a Bottleneck convolution block. The feature processing layer Af1 includes three convolution blocks, the feature processing layer Af2 includes two convolution blocks, and the feature processing layers Af3 and Af4 each include one convolution block. The convolution block is sequentially connected by a convolutional layer, a batch normalization layer BatchNorm, and a Relu activation function layer; the student model includes four stacked transposed convolutional layers Slayer1 to Slayer4.

[0029] Specifically, the P-conv (pre-convolution block) includes a convolutional layer, a batch normalization layer, a Relu activation function layer, and a pooling layer. The convolutional layer has a stride of 2, a padding of 3, and a kernel size of 7×7; the pooling layer has a stride of 2, a padding of 1, and a pooling window of 3×3.

[0030] As Figure 2 shown, Layer1 to Layer4 include a group of t_conv convolution block groups and a parallel downsampling layer DownSample, and then connect to a feature concatenation layer; where Layer1 includes 3 t_conv convolution blocks, Layer2 includes 4 t_conv convolution blocks, Layer3 includes 6 t_conv convolution blocks, and Layer4 includes 3 t_conv convolution blocks; the t_conv convolution block includes a 3×3 convolution conv2 with a stride of 1 and two 1×1 convolutions conv1 connected together. After each layer of convolution, a BatchNorm layer needs to be connected. The downsampling layer DownSample includes a 1×1 conv convolution and a BatchNorm layer.

[0031] As Figure 3As shown, corresponding to Layer1 to Layer4, Slayer1 to Slayer4 include a tr_conv transposed convolution block, multiple connected groups of t_conv convolution blocks, and a parallel upsampling layer UpSample, and then connect to a feature concatenation layer; among them, Slayer1 includes 1 tr_conv transposed convolution block and 2 t_conv convolution blocks, Slayer2 includes 1 tr_conv transposed convolution block and 5 t_conv convolution blocks, Slayer3 includes 1 tr_conv transposed convolution block and 3 t_conv convolution blocks, and Slayer4 includes 1 tr_conv transposed convolution block and 2 t_conv convolution blocks; the tr_conv transposed convolution block includes a 2×2 convtrans transposed convolution with a stride of 2 and a BatchNorm layer; the upsampling layer UpSample includes a 2×2 convtrans transposed convolution with a stride of 2 and a BatchNorm layer.

[0032] The preprocessed image with a size of 256×256×3 is input into the teacher model pre-trained on ImageNet. First, it sequentially passes through a convolutional layer with a stride of 2, a padding of 3, and a kernel size of 7×7, a batch normalization layer, a Relu activation function layer, and a pooling layer with a stride of 2, a padding of 1, and a pooling window of 3×3 to obtain a feature map F0 of 64×64×64. The feature map F0 is input into Layer1 to obtain a deep feature map F1 of 256×64×64, the deep feature map F1 is input into Layer2 to obtain a deep feature map F2 of 512×32×32, the deep feature map F2 is input into Layer3 to obtain a shallow feature map F3 of 1024×16×16, and the shallow feature map F3 is input into Layer4 to obtain a shallow feature map F4 of 2048×8×8. The feature maps F1, F2, F3, and F4 obtained by the teacher model are input into the teaching assistant model.

[0033] The deep feature map F1 is input into the feature processing layer Af1 to obtain a feature map Al1 of 1024×8×8; the deep feature map F2 is input into the feature processing layer Af2 to obtain a feature map Al2 of 1024×8×8; the shallow feature map F3 is input into the feature processing layer Af3 to obtain a feature map Al3 of 1024×8×8; the shallow feature map F4 is downsampled to obtain a feature map Al4 of 1024×8×8. The purpose of downsampling is to make the dimension of the feature map Al4 consistent with those of the feature maps Al1 to Al3. After pairwise concatenation of the feature maps Al1, Al2, Al2, Al3, Al3, and Al4, feature maps A1, A2, and A3 with dimensions of 2048×8×8 are obtained. Each of the feature maps A1, A2, and A3 is multiplied by a set of parameters obtained after passing through the feature processing layer Af4 and then through the spatial attention module to amplify important features, and the dimension size of the feature map remains unchanged; subsequently, the three feature maps are horizontally concatenated into a feature map At of 6144×8×8. After being accelerated by the Bottleneck convolution block, the feature map At is input into the student model.

[0034] The feature map At is input into the transposed convolution layer Slayer1 to obtain a feature map F5 of 2048×8×8; the feature map F5 is input into the transposed convolution layer Slayer2 to obtain a feature map F6 of 1024×16×16; the feature map F6 is input into the transposed convolution layer Slayer3 to obtain a feature map F7 of 512×32×32; the feature map F7 is input into the transposed convolution layer Slayer4 to obtain a feature map F8 of 256×64×64.

[0035] (2) The teacher - teaching assistant - student model is trained using the MvTecAD dataset, where the parameters of the teacher model are fixed and do not participate in the training; the student model and the teaching assistant module participate in the training; as Figure 4 shown, the cosine similarities between F1 and F8, F2 and F7, F3 and F6, and F4 and F5 are calculated respectively, and the mean values are added as the loss function to continuously optimize the teaching assistant model and the student model.

[0036] The MvTecAD dataset used for training is a comprehensive / multi - category object detection and localization dataset.

[0037] After training is completed, a teacher - teaching assistant - student model for object detection is obtained.

[0038] (3) As Figure 5As shown, the preprocessed test image with a size of 256×256×3 is input into the trained teacher-assistant-student model to obtain the feature maps T1, T2, T3, T4 extracted by the teacher model and the feature maps S1, S2, S3, S4 learned and extracted by the student model; the cosine similarities cos_similarity of the four groups of feature maps are calculated respectively; first, the dimensions of the four groups of results are expanded to the size of the input image, i.e., 256×256×3, by bilinear interpolation sampling, and then the product operation is performed on the four groups of cosine similarity calculation results after bilinear sampling dimensionality increase to obtain the feature matrix Ma1. After smoothing Ma1 using Gaussian filtering, the feature scores of each value ma1 in the feature matrix Ma1 are calculated to obtain the feature score matrix;

[0039] The formula for calculating the feature score is: Ma1 = (ma1 - ma1.min) / (ma1.max - ma1.min)

[0040] where ma1.min is the minimum value in the feature matrix Ma1, and ma1.max is the maximum value in the feature matrix Ma1.

[0041] The values in the feature score matrix are converted to rbg values, and the conversion formula is: Ma1 = Ma1 * 255

[0042] An image heat map is drawn based on the converted feature score matrix Ma1, and the target feature area is displayed in bright red in the image. Table 1 shows the comparison of the pixel Auroc detection metrics when the present invention and other algorithms are verified on the MvTecAD dataset. It can be seen from the data comparison in Table 1 that the method of the present invention has a high detection accuracy under various classifications.

[0043] Table 1

[0044]

[0045]

Claims

1. A target detection method based on knowledge distillation, characterized in that, it includes: (1) Construct a teacher-assistant-student model, where the teacher model includes P-conv and four stacked convolutional layers Layer1 to Layer4; the assistant model includes four feature processing layers Af1 to Af4, a feature splicing layer, a spatial attention module, and a Bottleneck convolutional block. The feature processing layer Af1 includes three convolutional blocks, the feature processing layer Af2 includes two convolutional blocks, and the feature processing layers Af3 and Af4 each include one convolutional block. The convolutional block is sequentially connected by a convolutional layer, a batch normalization layer BatchNorm, and a Relu activation function layer; the student model includes four stacked deconvolutional layers Slayer1 to Slayer4; The pre-trained teacher model extracts two groups of deep feature maps F1, F2 and two groups of shallow feature maps F3, F4 of the image and inputs them into the assistant model; The sizes of the feature maps F1 to F4 increase in sequence; the deep feature map F1 is input into the feature processing layer Af1 to obtain the feature map Al1; the deep feature map F2 is input into the feature processing layer Af2 to obtain the feature map Al2; the shallow feature map F3 is input into the feature processing layer Af3 to obtain the feature map Al3; the shallow feature map F4 is downsampled to obtain the feature map Al4; the feature maps Al1 to Al4 have the same dimension; after the feature maps Al1, Al2, Al2, Al3, Al3, Al4 are pairwise spliced, the feature maps A1, A2, A3 are obtained; after the feature maps A1, A2, A3 are respectively processed by the feature processing layer Af4 and the spatial attention module, they are horizontally spliced to obtain a feature map At; after the feature map At is accelerated by the Bottleneck convolutional block, it is input into the student model; the feature map At is input into the deconvolutional layer Slayer1 to obtain the feature map F5; the feature map F5 is input into the deconvolutional layer Slayer2 to obtain the feature map F6; the feature map F6 is input into the deconvolutional layer Slayer3 to obtain the feature map F7; the feature map F7 is input into the deconvolutional layer Slayer4 to obtain the feature map F8; (2) Train the teacher-assistant-student model, where the parameters of the teacher model are fixed and do not participate in the training; the student model and the assistant module participate in the training; calculate the cosine similarity of F1 and F8, F2 and F7, F3 and F6, F4 and F5 respectively, and take the mean and add them as the loss function; (3) Input the image to be detected into the trained teacher-assistant-student model. The teacher model and the student model respectively obtain a feature map, and calculate the cosine similarity of the four groups of feature maps respectively; the feature map is restored to the original image size by bilinear interpolation sampling, and the product operation of the four groups of cosine similarity calculation results after bilinear sampling and upsampling is used to obtain the feature map Ma1. After smoothing Ma1 with Gaussian filtering, calculate the feature score of each value in Ma1 to obtain the feature score map; Draw an image heat map according to the feature score map. The higher the feature score, the brighter it is in the image heat map, thereby realizing target detection.

2. The target detection method based on knowledge distillation according to claim 1, characterized in that, The P-conv includes a convolutional layer, a batch normalization layer, a Relu activation function layer, and a pooling layer. The convolutional layer has a stride of 2, a padding of 3, and a convolutional kernel size of 7×7; the pooling layer has a stride of 2, a padding of 1, and a pooling window of 3×3.

3. The object detection method based on knowledge distillation according to claim 2, characterized in that The convolutional layers Layer1 to Layer4 include a group of t_conv convolutional block groups and a parallel downsampling layer DownSample, and then connect to a feature splicing layer; among them, Layer1 includes 3 t_conv convolutional blocks, Layer2 includes 4 t_conv convolutional blocks, Layer3 includes 6 t_conv convolutional blocks, and Layer4 includes 3 t_conv convolutional blocks; the t_conv convolutional block is composed of a 3×3 convolution conv2 with a stride of 1 and two 1×1 convolutions conv1 and conv3 connected. After each layer of convolution, a BatchNorm layer needs to be connected; the downsampling layer includes a 1×1conv convolution and a BatchNorm layer.

4. The object detection method based on knowledge distillation according to claim 1, characterized in that The transposed convolutional layers Slayer1 to Slayer4 include a tr_conv transposed convolutional block, multiple connected t_conv convolutional block groups, and a parallel upsampling layer UpSample, and then connect to a feature splicing layer; wherein Slayer1 includes 1 tr_conv transposed convolutional block and 2 t_conv convolutional blocks, Slayer2 includes 1 tr_conv transposed convolutional block and 5 t_conv convolutional blocks, Slayer3 includes 1 tr_conv transposed convolutional block and 3 t_conv convolutional blocks, and Slayer4 includes 1 tr_conv transposed convolutional block and 2 t_conv convolutional blocks; the tr_conv transposed convolutional block includes a 2×2convtrans transposed convolution with a stride of 2 and a BatchNorm layer; the upsampling layer UpSample includes a 2×2convtrans transposed convolution with a stride of 2 and a BatchNorm layer.

5. The object detection method based on knowledge distillation according to any one of claims 1 to 4, characterized in that The feature maps output by the teacher model and the student model have exactly the same dimensional size.

6. The object detection method based on knowledge distillation according to claim 5, characterized in that The size of the feature map F1 is 256×64×64, the size of the feature map F2 is 512×32×32, the size of the feature map F3 is 1024×16×16, and the size of the feature map F4 is 2048×8×8; the size of the feature map F5 is 2048×8×8, the size of the feature map F6 is 1024×16×16, the size of the feature map F7 is 512×32×32, and the size of the feature map F8 is 256×64×64.

7. The object detection method based on knowledge distillation according to claim 6, characterized in that The input image size for the teacher model is 256×256×3, the sizes of the feature maps Al1 to Al4 are all 1024×8×8, the sizes of the feature maps A1 to A3 are 2048×8×8, and the size of the feature map At is 6144×8×8.

8. The object detection method based on knowledge distillation according to claim 1, wherein, the feature score calculation formula is: Ma1 = (ma1 - ma1.min) / (ma1.max - ma1.min) where ma1.min is the minimum value in the feature matrix Ma1, and ma1.max is the maximum value in the feature matrix Ma1.

9. The object detection method based on knowledge distillation according to claim 1, wherein, the teacher - teaching assistant - student model is trained using the MvTecAD dataset.

10. The object detection method based on knowledge distillation according to claim 1, wherein, the teacher model is pre - trained on the ImageNet dataset.