Decompression Image Object Detection Method and System Based on Attention Mechanism Distillation

By building a network of teachers and students and introducing attention distillation losses, we solve the problem of understanding the accuracy of object detection on compressed images, and improve detection accuracy and generalization capabilities.

CN115731447BActive Publication Date: 2025-07-11STATE GRID FUJIAN ELECTRIC POWER RES INST +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211420783.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-13
Publication Date
2025-07-11
Estimated Expiration
2042-11-13

AI Technical Summary

Technical Problem

The existing deep learning-based object detection methods perform poorly on decompressed images, resulting in missed and missed detection, making it difficult to work effectively in practical application scenarios.

Method used

Using knowledge distillation technology based on attention mechanism, we use attention distillation loss to train students' networks by building teachers and students' networks, thereby improving the target detection performance of low-quality images.

Benefits of technology

The object detection accuracy and generalization ability of low-quality images are improved, and the detection effect on decompressed images is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115731447B_ABST
    Figure CN115731447B_ABST
Patent Text Reader

Abstract

The present invention relates to a decompressed image object detection method based on attention mechanism distillation, which comprises the following steps: Step S1: Obtain a high-quality image dataset, and obtain a corresponding decompressed low-quality image dataset by compressing and decompressing the high-quality image dataset; Step S2: Construct an object detection teacher network and an object detection student network; Step S3: Train the object detection teacher network based on the high-quality image dataset; Step S4: Based on the trained object detection teacher network, add an attention-based distillation loss to train the object detection student network; Step S5: Perform object detection on the decompressed image based on the trained student network. The present invention realizes the extraction of higher-quality image features from decompressed images, and effectively improves the object detection performance of low-quality images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image target detection, and in particular to a decompressed image target detection method and system based on attention mechanism distillation. Background Art

[0002] As a key technology in the fields of autonomous driving, intelligent monitoring, etc., the target detection algorithm is one of the most popular research directions in the field of computer vision today. In recent years, with the rapid development of deep learning, the target detection method based on deep learning has achieved remarkable performance. These target detection networks are trained on high-quality clean images. However, in some practical application scenarios, it is difficult to obtain high-quality clean images (bandwidth limitation, high-quality images are difficult to transmit), and most images are decompressed images. The compression of images will inevitably lead to a certain decline in image quality. Even in some specific occasions, such as field monitoring detection, a large amount of data needs to be detected. Due to the limitations of equipment and bandwidth, a large compression ratio needs to be used to compress the data. When the compression ratio is large, the quality of the decompressed image drops sharply, which makes the target detection network often have serious missed detections and false detections when detecting such decompressed low-quality images, and almost fails in practical application scenarios. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a decompressed image target detection method and system based on attention mechanism distillation, which can extract higher-quality image features from decompressed images and effectively improve the target detection performance of low-quality images.

[0004] To achieve the above purpose, the present invention adopts the following technical solutions:

[0005] A decompressed image target detection method based on attention mechanism distillation includes the following steps:

[0006] Step S1: Obtain a high-quality image dataset, and obtain the corresponding decompressed low-quality image dataset by compressing and decompressing the high-quality image dataset;

[0007] Step S2: Construct a target detection teacher network and a target detection student network;

[0008] Step S3: Train the target detection teacher network based on the high-quality image dataset;

[0009] Step S4: Based on the trained target detection teacher network, add an attention-based distillation loss to train the target detection student network;

[0010] Step S5: Perform target detection on the low-quality image based on the trained student network.

[0011] Further, the target detection teacher network is constructed based on YOLOv3 or YOLOv5s, and the backbone network is fixed and the detection head is removed during the training process.

[0012] Further, the target detection student network is based on YOLOv3 or YOLOv5s, and an attention learning module is added before each branch detection head at different scales of YOLOv3 or YOLOv5s.

[0013] Further, the attention learning module includes a transposed convolutional layer, N residual blocks, and an average pooling layer.

[0014] Further, the specific steps of S4 are as follows:

[0015] Use the high-quality image dataset as the input of the trained target detection teacher network, and use the corresponding decompressed low-quality image dataset as the input of the target detection student network. After fixing the parameters of the target detection teacher network, use the high-quality feature z t extracted by the target detection teacher network and the low-quality feature z s extracted by the target detection student network to calculate the distillation loss, and add the detection loss of the target detection network itself to train the student network.

[0016] Further, promoting the degenerate features of the decompressed image to be close to the high-quality image features based on the knowledge distillation technology is expressed as the following formula:

[0017]

[0018] where t and s represent the teacher network and the student network respectively, f represents the backbone network with parameters θ, and z t = f t (x; θ t ) represents the high-quality feature extracted from the high-quality image x, represents the degenerate feature extracted from the decompressed image ; d represents a certain distance or divergence measure in the feature space.

[0019] Further, the detection loss is expressed as:

[0020]

[0021] where ω represents the attention map with a size of 1×C×H×w, and the latter term is a regularization term due to sparsity; and set R(ω)=||ω||1;

[0022] The detection loss L det is composed of three parts:

[0023]

[0024]

[0025]

[0026] Among them, λ in each formula represents three different weight sizes of parts; S 2 represents the size of the feature map output by the detection network; B represents the number of detection boxes allocated for each grid; represents 1 when there is an object in the detection box with subscripts i and j, and 0 otherwise; p i (c) is the probability that the object is of class c;

[0027] The detection loss can be briefly expressed as:

[0028] L det = L box + L cls + L obj

[0029] Then the loss for finally training the student network is:

[0030] L = L det + λ * L dis .

[0031] A decompressed image object detection device based on attention mechanism distillation, comprising:

[0032] A data acquisition module, configured to acquire a high-quality image data set, and obtain a corresponding decompressed low-quality image data set by compressing and decompressing the high-quality image data set;

[0033] A model construction module, configured to construct a target detection teacher network and a target detection student network;

[0034] A model training module, which will train the target detection teacher network based on the high-quality image data set, and based on the trained target detection teacher network, add an attention-based distillation loss to train the target detection student network;

[0035] A detection module, which performs object detection on the low-quality image based on the trained student network.

[0036] A decompressed image object detection system based on attention mechanism distillation, comprising a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the decompressed image object detection method described above.

[0037] A computer-readable storage medium, comprising a program that can be executed by a processor to implement the method described above.

[0038] The present invention has the following beneficial effects compared with the prior art:

[0039] Through the knowledge distillation technology based on the self-attention mechanism proposed by the present invention, it is possible to focus on important regions in image features, prompting the network to extract higher-quality image features from decompressed images, improving the detection accuracy of existing deep learning-based object detection methods on low-quality images (decompressed images), enhancing their generalization and promotion capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is the overall architecture of the present invention;

[0041] Figure 2 is the detection results, extracted features and difference map of high-quality pictures and decompressed pictures in an embodiment of the present invention, where (a) is the detection result of the high-quality image, (b) is the detection result of the low-quality decompressed image, (c) is the features extracted from the high-quality image, (d) is the features extracted from the low-quality decompressed image, and (e) is the difference between the high-quality features and the low-quality features;

[0042] Figure 3 is the structure diagram of yolov3-tiny in an embodiment of the present invention;

[0043] Figure 4 is the structure diagram of yolov5s in an embodiment of the present invention;

[0044] Figure 5 is the feature maps extracted by different algorithms in an embodiment of the present invention, where (a) is the features extracted from the high-quality image, (b) is the features extracted from the low-quality decompressed image, (c) is the features extracted from the detector after enhanced training by the Aug algorithm, and (d) is the features restored from the model after l2-norm distillation. (e) is the restored function of the present invention, the first row is the visualization result of the features, and the second row is the difference map between the high-quality features and the corresponding restored features;

[0045] Figure 6 is the detection visualization results of different algorithms in an embodiment of the present invention, where (a) is the detection result of the high-quality image, (b) is the detection result of the low-quality decompressed image, (c) is the detection result of the l2-norm distillation method on the low-quality decompressed image, and (d) is the detection result of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0047] Refer to Figures 1-5 , this embodiment provides an object detection method for decompressed images based on attention mechanism distillation, including the following steps:

[0048] Step S1: Obtain a high-quality image dataset, and obtain the corresponding decompressed low-quality image dataset by decompressing the high-quality image dataset;

[0049] Step S2: Construct a target detection teacher network and a target detection student network;

[0050] Step S3: Train the target detection teacher network based on the high-quality image dataset;

[0051] Step S4: Based on the trained target detection teacher network, add an attention-based distillation loss to train the target detection student network;

[0052] Step S5: Perform target detection on the low-quality image based on the trained student network.

[0053] In this embodiment, the COCO2017 dataset is selected as the high-quality image dataset, and this image dataset consists of 118,287 training images and 5,000 test images.

[0054] In this embodiment, please refer to Figure 1 , a YOLO series single-stage detector is adopted. The input of the teacher network is a high-quality image, and the structure of the student network is the same as that of the teacher, but the decompressed image is used as the input. As Figure 1 shown, YOLOv3 consists of two parts. The backbone network is used to extract features, and the detection head is used for classification and bounding box regression. In this embodiment, the backbone of the teacher is used to extract high-quality features. Therefore, the backbone network is fixed and the head is removed during training. For the student model, the backbone and the head are retained, and it is initialized with pre-trained parameters to better converge. Pairwise images, that is, high-quality images and corresponding decompressed images, are respectively input into the teacher and student networks.

[0055] In this embodiment, an attention-aware feature extraction method is proposed, and the learned attention map is used as the weight of the l2 norm. Since the distillation weights of different regions are different, the method of the present invention enables the degraded features extracted from the decompressed images to better align with the corresponding high-quality features.

[0056] The knowledge distillation technology that promotes the degraded features of the decompressed images to approach the high-quality image features can be expressed as the following formula:

[0057]

[0058] where t and s respectively represent the teacher network and the student network, f represents the backbone network with parameters θ, and z t = f t (x; θ t ) represents the high-quality features extracted from the high-quality image x, Represents the degraded features extracted from the decompressed image. d represents a certain distance (or divergence) metric in the feature space. For the object detection task, the importance of different regions of the feature map is not the same. Similarly, as shown in

[0059] (c), (d), and (e), by visualizing the features, the results show the difference map between high-quality features and degraded features. Therefore, it is not suitable to consider the importance of each region as a constant with the same value. Figure 2 (c), (d), and (e) show that by visualizing the features, the results indicate the difference map between high-quality features and degraded features. Therefore, it is not appropriate to regard the importance of each region as a constant with the same value.

[0060] The present invention represents the importance of different regions in the feature map by learning the attention map and weights (denoted as ω), and applies it to the distillation loss. Assume that the importance of the image features is defined as the magnitude of the difference between z s and z t , that is, the greater the difference in a certain region of the feature (such as the edge and texture regions shown in Figure 2 (e)), the more important it is in the feature extraction process, and the greater the value of ω should be.

[0061] The present invention learns the attention map ω by adding an attention learning module branch to the student model (as shown in Figure 1 ), and the proposed attention-aware feature extraction loss can be expressed as:

[0062]

[0063] where ω represents the attention map with a size of 1×C×H×w, and the latter term is a regularization term due to sparsity

[0064] Set R(ω)=||ω||1.

[0065] The loss for finally training the student network is:

[0066] L = L det + λ * L dis

[0067] The value of the learned attention map ω measures the difficulty / importance of feature reconstruction. If there is a large gap between z t and z s , the student network tends to learn a larger weight ω to reduce the loss. Conversely, once ω increases, the second term in the loss function will also increase, which will prompt the model to optimize and reduce the difference between z t and z s , which makes the student model pay more attention to the difficult / important regions in the feature map. Therefore, the student network can better enhance the features under the guidance of the teacher network and the attention map, and improve the detection accuracy.

[0068] Among them, the attention learning module includes a transposed convolutional layer for magnifying the feature map, N residual blocks, and an average pooling layer for restoring the original resolution.

[0069] Preferably, N is 3, and the network structure of the attention learning module is as follows:

[0070] Input layer → First deconvolution layer (upsampling) → First activation function layer → Second convolutional layer → Second activation function layer → Third convolutional layer → Third activation function layer → Fourth convolutional layer → Fourth activation function layer → Fifth convolutional layer → Fifth activation function layer → Sixth convolutional layer → Sixth activation function layer → Seventh convolutional layer → Seventh activation function layer → Eighth convolutional layer → Eighth activation function layer → Ninth convolutional layer → Ninth activation function layer → Tenth convolutional layer → Tenth activation function layer → First average pooling layer (downsampling) → Eleventh convolutional layer → Eleventh activation function layer → Twelfth convolutional layer → Twelfth activation function layer → Thirteenth convolutional layer → Output layer

[0071] Example 1:

[0072] In this embodiment, the experimental quantization results with yolov3-tiny as the detector are shown in Table 1, and the experimental quantization results with yolov5s as the detector are shown in Table 2. Various comparative experimental methods are in the order from top to bottom as described above.

[0073] Table 1 Comparative experimental results of yolov3-tiny

[0074]

[0075] Table 2 Comparative experimental results of yolov5s

[0076]

[0077] From the above experimental results, it can be found that the present invention proposes a new distillation loss function based on the attention mechanism to train the object detection network, which is mainly used for object detection of low-quality decompressed images. By introducing the attention mechanism into the distillation loss of the object detection network, the importance of different regions of the high-resolution feature image feature map can be learned simultaneously, prompting the network to pay more attention to learning more important regions such as the edges of objects. In addition, the method of the present invention not only has better effects than the existing optimal method Aug, but is also more easily generalized to other object detection tasks. In addition, referring to Figure 5 , the present invention proves for the first time that this distillation loss based on the attention mechanism can better restore features than the MSE distillation loss, and at the same time has better detection effects. The experimental results on two commonly used object detection networks show that the distillation loss based on the attention mechanism proposed by the present invention can achieve better detection results in the object detection task of decompressed images. Some object detection visualization results are asFigure 6 as shown

[0078] Embodiment 2

[0079] Based on the same inventive concept, the present application also provides a decompressed image object detection device based on attention mechanism distillation, including:

[0080] A data acquisition module, configured to acquire a high-quality image data set, and obtain a corresponding decompressed low-quality image data set by compressing and decompressing the high-quality image data set;

[0081] A model construction module, configured to construct an object detection teacher network and an object detection student network;

[0082] A model training module, which trains the object detection teacher network based on the high-quality image data set, and adds an attention-based distillation loss to train the object detection student network based on the trained object detection teacher network;

[0083] A detection module, which performs object detection on the low-quality image based on the trained student network.

[0084] Embodiment 3

[0085] Based on the same inventive concept, the present application also provides a decompressed image object detection system based on attention mechanism distillation, including a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the decompressed image object detection method as described above.

[0086] Embodiment 4

[0087] Based on the same inventive concept, the present application also provides a computer-readable storage medium, including a program that can be executed by a processor to implement the method as described above.

[0088] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0089] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0090] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0091] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0092] As mentioned above, it is only the preferred embodiments of the present invention, and it is not a limitation of the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A decompressed image object detection method based on attention mechanism distillation, characterized in that, It includes the following steps: Step S1: Obtain a high-quality image dataset, and obtain the corresponding decompressed low-quality image dataset by decompressing the high-quality image dataset; Step S2: Construct a target detection teacher network and a target detection student network; Step S3: Train the target detection teacher network based on the high-quality image dataset; Step S4: Based on the trained target detection teacher network, add an attention-based distillation loss to train the target detection student network; Step S5: Perform target detection on the low-quality image based on the trained student network; The specific content of step S4 is as follows: Use the high-quality image dataset as the input to the trained object detection teacher network, and the corresponding decompressed low-quality image dataset as the input to the object detection student network. After fixing the parameters of the object detection teacher network, use the high-quality features z extracted by the object detection teacher network t and the low-quality features z extracted by the object detection student network s to calculate the distillation loss L dis , and the distillation loss is expressed as: where z t represents the high-quality features extracted from the high-quality image x, and z s represents the degraded features extracted from the decompressed image ; ω represents the attention map of size 1×C×H×w, and the latter term is a regularization term due to sparsity; and set R(ω)=||ω||1; d represents a certain distance or divergence metric in the feature space; the student network is trained by adding the detection loss of the object detection network itself on the basis of the distillation loss; Based on the knowledge distillation technology, promoting the degraded features of the decompressed image to be close to the feature expression of the high-quality image is expressed as the following formula: Among them, θ s represents the parameters of the student network, t and s represent the teacher network and the student network respectively, f represents the backbone network with parameters θ, z t = f t (x; θ t ) represents the high-quality features extracted from the high-quality image x, represents the degraded features extracted from the decompressed image ; Detection loss L det It is composed of three parts: where λ in each expression represents three different weights; S 2 represents the size of the feature map output by the detection network; B represents the number of detection boxes assigned to each grid; represents 1 when there is an object in the detection box with subscripts i and j, and 0 otherwise; p i (c) is the probability that the object is of class c; The detection loss is expressed as: L det = L box + L cls + L obj Then the loss for finally training the student network is: L = L det + λ·L dis .

2. The decompressed image object detection method based on attention mechanism distillation according to claim 1, wherein, The target detection teacher network is constructed based on YOLOv3 or YOLOv5s, and the backbone network is fixed and the detection head is removed during the training process.

3. The decompressed image object detection method based on attention mechanism distillation according to claim 1, characterized in that The target detection student network is based on YOLOv3 or YOLOv5s, and an attention learning module is added in front of each branch detection head at different scales of YOLOv3 or YOLOv5s.

4. The decompressed image object detection method based on attention mechanism distillation according to claim 3, wherein, The attention learning module includes a transposed convolutional layer, N residual blocks, and an average pooling layer.

5. A decompressed image object detection device based on attention mechanism distillation, characterized in that, The decompressed image target detection device is implemented by using the decompressed image target detection method based on attention mechanism distillation as described in any one of claims 1-4, and includes: A data acquisition module, configured to obtain a high-quality image dataset, and obtain the corresponding decompressed low-quality image dataset by decompressing the high-quality image dataset; A model construction module, configured to construct a target detection teacher network and a target detection student network; A model training module, configured to train the target detection teacher network based on the high-quality image dataset, and based on the trained target detection teacher network, add an attention-based distillation loss to train the target detection student network; A detection module, configured to perform target detection on the low-quality image based on the trained student network.

6. A decompressed image object detection system based on attention mechanism distillation, characterized in that, It includes a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the decompressed image target detection method as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, It includes a program that can be executed by a processor to implement the method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Underwater target detection method and system based on self-attention distillation and image enhancement

    CN115049815A

  • Target detection compression method based on knowledge distillation

    CN115063663A