A remote sensing image small target detection method and device
Patent Information
- Application Number
- CN202311253673.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-26
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-09-26
AI Technical Summary
但是高分辨图像也意味着更高的计算成本,该方法需要更高的硬件支持
[0033]本发明实施例中提供的一种遥感图像小目标检测方法及装置,包括:S1、准备数据集,并根据实验需求对数据集进行预处理;S2、利用高分辨率图像训练教师网络;S3、构建学生网络并训练,学生网络使用与教师网络相同的检测器,使用低分辨率图像进行训练,利用跨尺度蒸馏机制学习教师横向特征,亚像素卷积金字塔网络解析融合高层语义信息;S4、使用训练好的目标检测网络对遥感图像进行目标检测。本发明中提出了跨尺度蒸馏机制,使用高分辨率图像训练教师网络。利用训练好的教师横向特征作为监督信号指导学生网络学习目标区域特征,本发明中提出的亚像素超分金字塔网络充分利用了卷积神经网络的学习和解析能力,在解析高层语义信息的同时学习跨尺度蒸馏机制引入的教师横向特征信息,在自上而下的路径上充分融合了高层语义和低层语义,提高了遥感图像小目标的检测精度。
Smart Images

Figure CN117409314B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method and apparatus for detecting small targets in remote sensing images. Background Technology
[0002] Target detection technology is widely used in remote sensing, playing a crucial role in numerous tasks such as traffic planning, disaster early warning, and military reconnaissance. Currently, deep learning-based target detection methods have made significant progress in natural image detection. Compared to traditional target detection techniques that utilize manually designed features, deep learning, with its powerful learning capabilities, significantly improves the accuracy of target detection. However, compared to natural images, remote sensing images are characterized by complex backgrounds, large target scale variations, and variable target orientations. This presents considerable challenges to the research of remote sensing image target detection algorithms, especially those for small targets. Designing deep learning-based small target detection algorithms tailored to the characteristics of remote sensing images is of great significance for enhancing the practical value of this field.
[0003] The biggest difference between remote sensing images and natural images is their unique bird's-eye view. Secondly, remote sensing images are taken by high-altitude aircraft, resulting in small and blurry targets. This is one reason why deep learning object detection methods are not performing well when directly applied to remote sensing images. Another reason for the increased difficulty in detecting small targets is that the extracted small target information is further lost after downsampling through convolutional neural networks. In recent years, some scholars have proposed methods using high-resolution remote sensing images to address the problem of small targets. However, high-resolution images also mean higher computational costs, requiring more advanced hardware support. Other scholars have focused on the rational application of extracted target features, such as multi-scale detection techniques for small targets and contextual feature pyramids that fuse semantic information from different resolutions. However, these methods do not fully analyze and utilize high-level semantic features.
[0004] In summary, there is an urgent need for a method and system for detecting small targets in remote sensing images that can fully exploit high-level semantic features while maintaining low computational cost. Summary of the Invention
[0005] To address the aforementioned problems, this invention provides a method and apparatus for detecting small targets in remote sensing images.
[0006] In a first aspect, embodiments of the present invention provide a method for detecting small targets in remote sensing images, comprising:
[0007] A dataset is constructed, and the image data in the dataset is preprocessed. The dataset includes high-resolution images and low-resolution images.
[0008] A teacher network is constructed using the high-resolution image, and the output feature map of the trained teacher network is used as a supervision signal for the student network.
[0009] The student network is trained using the low-resolution image, the teacher lateral features of the teacher network are learned through a cross-scale distillation mechanism, and high-level semantic information is fused through a sub-pixel convolutional pyramid network. The trained teacher lateral features are used as a supervision signal to guide the student network to learn target region features.
[0010] The trained student network is used as a target detection network to perform target detection on the remote sensing image to be detected.
[0011] As an optional approach, the dataset is constructed, and the image data in the dataset is preprocessed. The dataset includes high-resolution images and low-resolution images, including:
[0012] A dataset is constructed, and interpolation techniques are used to set up one-to-one high- and low-resolution image pairs in the dataset, wherein the size of the high-resolution image is twice the size of the low-resolution image.
[0013] As an optional approach, the construction of a teacher network using the high-resolution image, and the use of the output feature map of the trained teacher network as a supervision signal for the student network, includes:
[0014] The teacher network is trained using high-resolution images. Training is stopped once the teacher network test loss stabilizes. The weights with the best test results are selected as the weights of the teacher network and then frozen.
[0015] As an optional approach, the student network is trained using the low-resolution image, the teacher lateral features of the teacher network are learned through a cross-scale distillation mechanism, and high-level semantic information is fused using a sub-pixel convolutional pyramid network. The trained teacher lateral features are then used as a supervision signal to guide the student network in learning target region features, including:
[0016] The high-resolution image and the low-resolution image are respectively input into the trained teacher network and the untrained student network for preliminary feature extraction. The teacher network obtains the teacher's lateral feature C. T3 C T2 C T1 The student network obtains the student's lateral features C. S3 C S2 C S1 ;
[0017] The teacher's horizontal characteristic C T3 C T2 C T1 The teacher network features P are obtained by fusing the feature pyramids. T3 P T2P T1 Using the teacher network feature P T3 P T2 P T1 As a supervisory signal, the student's lateral feature C S3 C S2 C S1 The input is fused into a sub-pixel feature super-resolution pyramid network to obtain the student network feature P. S3 P S2 P S1 The student network feature P S3 P S2 P S1 Used for target detection.
[0018] As an alternative, the teacher network and the student network use the same detector.
[0019] Secondly, embodiments of the present invention provide a remote sensing image small target detection device, comprising:
[0020] A dataset construction unit is used to construct a dataset and preprocess the image data in the dataset, which includes high-resolution images and low-resolution images;
[0021] The network construction unit is used to construct a teacher network using the high-resolution image and use the output feature map of the trained teacher network as a supervision signal for the student network.
[0022] The training unit is used to train the student network using the low-resolution image, learn the teacher lateral features of the teacher network through a cross-scale distillation mechanism, and parse and fuse high-level semantic information through a sub-pixel convolutional pyramid network. The trained teacher lateral features are used as supervision signals to guide the student network to learn target region features.
[0023] The detection unit is used to perform target detection on the remote sensing image to be detected by using the trained student network as a target detection network.
[0024] As an optional approach, the dataset construction unit is specifically used for:
[0025] A dataset is constructed, and interpolation techniques are used to set up one-to-one high- and low-resolution image pairs in the dataset, wherein the size of the high-resolution image is twice the size of the low-resolution image.
[0026] As an optional solution, the network building unit is specifically used for:
[0027] The teacher network is trained using high-resolution images. Training is stopped once the teacher network test loss stabilizes. The weights with the best test results are selected as the weights of the teacher network and then frozen.
[0028] As an optional approach, the training unit is specifically used for:
[0029] The high-resolution image and the low-resolution image are respectively input into the trained teacher network and the untrained student network for preliminary feature extraction. The teacher network obtains the teacher's lateral feature C. T3 C T2 C T1 The student network obtains the student's lateral features C. S3 C S2 C S1 ;
[0030] The teacher's horizontal characteristic C T3 C T2 C T1 The teacher network features P are obtained by fusing the feature pyramids. T3 P T2 P T1 Using the teacher network feature P T3 P T2 P T1 As a supervisory signal, the student's lateral feature C S3 C S2 C S1 The input is fused into a sub-pixel feature super-resolution pyramid network to obtain the student network feature P. S3 P S2 P S1 The student network feature P S3 P S2 P S1 Used for target detection.
[0031] As an alternative, the teacher network and the student network use the same detector.
[0032] Compared with the prior art, the present invention can achieve the following beneficial effects:
[0033] This invention provides a method and apparatus for detecting small targets in remote sensing images, comprising: S1, preparing a dataset and preprocessing the dataset according to experimental requirements; S2, training a teacher network using high-resolution images; S3, constructing and training a student network, which uses the same detector as the teacher network, is trained using low-resolution images, learns teacher lateral features using a cross-scale distillation mechanism, and uses a sub-pixel super-resolution pyramid network to parse and fuse high-level semantic information; S4, using the trained target detection network to perform target detection on the remote sensing image. This invention proposes a cross-scale distillation mechanism and uses high-resolution images to train the teacher network. The trained teacher lateral features are used as a supervisory signal to guide the student network in learning target region features. The sub-pixel super-resolution pyramid network proposed in this invention fully utilizes the learning and parsing capabilities of convolutional neural networks, learning teacher lateral feature information introduced by the cross-scale distillation mechanism while parsing high-level semantic information, and fully fusing high-level and low-level semantics in a top-down path, thereby improving the detection accuracy of small targets in remote sensing images. Attached Figure Description
[0034] Figure 1 This is a flowchart illustrating a method for detecting small targets in remote sensing images according to an embodiment of the present invention.
[0035] Figure 2 This is a schematic diagram of the overall model structure in a remote sensing image small target detection method provided by an embodiment of the present invention;
[0036] Figure 3 This is a schematic diagram of a sub-pixel feature super-resolution module in a remote sensing image small target detection method provided by an embodiment of the present invention;
[0037] Figure 4 This is a structural block diagram of a remote sensing image small target detection device provided according to an embodiment of the present invention. Detailed Implementation
[0038] In the following description, embodiments of the invention will be described with reference to the accompanying drawings. In the description below, the same modules are denoted by the same reference numerals. Where the same reference numerals are used, their names and functions are also the same. Therefore, their detailed description will not be repeated.
[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not constitute a limitation thereof.
[0040] Combination Figure 1 As shown, this embodiment of the invention provides a method for detecting small targets in remote sensing images, including:
[0041] S101. Construct a dataset and preprocess the image data in the dataset, wherein the dataset includes high-resolution images and low-resolution images.
[0042] The dataset is constructed, and the preprocessing operation includes using interpolation technology to set one-to-one high- and low-resolution image pairs in the dataset. The size of the high-resolution image is twice the size of the low-resolution image. In this embodiment, the dataset is the DIOR dataset and the NWPU VHR-10. Those skilled in the art can choose flexibly as needed, and there is no limitation.
[0043] S102. Construct a teacher network using the high-resolution image, and use the trained network weights as the weights for constructing the student network.
[0044] The teacher network is trained using high-resolution images. Training is stopped once the teacher network test loss stabilizes. The weights with the best test results are selected as the weights of the teacher network and then frozen.
[0045] It should be noted that the teacher network and the student network can use the same detector. In this embodiment, the detector is the RetinaNet detector, but there is no limitation on this.
[0046] S103. The student network is trained using the low-resolution image. The teacher lateral features of the teacher network are learned through a cross-scale distillation mechanism. High-level semantic information is fused through a sub-pixel convolutional pyramid network. The trained teacher lateral features are used as a supervision signal to guide the student network in learning target region features.
[0047] The high-resolution image and the low-resolution image are respectively input into the trained teacher network and the untrained student network for preliminary feature extraction. The teacher network obtains the teacher's lateral feature C. T3 C T2 C T1 The student network obtains the student's lateral features C. S3 C S2 C S1 ;
[0048] The teacher's horizontal characteristic C T3 C T2 C T1 The teacher network features P are obtained by fusing the feature pyramids. T3 P T2 P T1 Using the teacher network feature P T3 P T2 P T1 As a supervisory signal, the student lateral feature CS3 C S2 C S1 The input is fused into a sub-pixel feature super-resolution pyramid network to obtain the student network feature P. S3 P S2 P S1 The student network feature P S3 P S2 P S1 Used for target detection.
[0049] like Figure 3 As shown, specifically, the sub-pixel feature super-resolution pyramid network replaces the traditional interpolation method in the feature pyramid network with a sub-pixel super-resolution module. Taking the student's lateral feature P... S3 To P S2 Taking the fusion process as an example, the following fusion steps are described. Before being input into the sub-pixel super-resolution module, the features P are... S3 ∈R B×C×H×W The transformation feature tmp∈R is obtained through dimensional transformation. (B×C)×1×H×W The transformed feature tmp is then input into the subpixel convolution module, first passing through three convolutional layers, including one 5×5 convolutional kernel and two 3×3 convolutional kernels, to obtain the convolutional feature tmp′∈R. (B×C)×4×H×W The convolutional feature tmp′ is used to obtain feature P after sub-pixel convolution. S3 ′∈P B×C×2H×2W This process uses a teacher network P T2 For monitoring signals, P S3 ′Amplified by interpolation and P T2 With the same size, the target region between the two features is selected using the Mask technique, and then the network loss is calculated. P S3 ′ and student lateral characteristics C S2 Fusion to obtain student network characteristics P S2 Student network characteristics P S2 Features are used for subsequent object detection. Student network features P S2 To P S1 The fusion process is the same, so it will not be described in detail here.
[0050] Optionally, the three convolutional layers in the subpixel convolution module consist of a 5×5 convolutional kernel and two 3×3 convolutional kernels, and then the subpixel convolutions are concatenated.
[0051] S104. Using the trained student network as a target detection network, target detection is performed on the remote sensing image to be detected.
[0052] The trained object detection network is used to detect objects in the remote sensing image to be detected.
[0053] This invention provides a method for detecting small targets in remote sensing images, comprising: S1, preparing a dataset and preprocessing it according to experimental requirements; S2, training a teacher network using high-resolution images; S3, constructing and training a student network, which uses the same detector as the teacher network, is trained using low-resolution images, learns teacher lateral features using a cross-scale distillation mechanism, and uses a sub-pixel super-resolution pyramid network to parse and fuse high-level semantic information; S4, using the trained target detection network to detect targets in the remote sensing image. This invention proposes a cross-scale distillation mechanism and uses high-resolution images to train the teacher network. The trained teacher lateral features are used as a supervisory signal to guide the student network in learning target region features. The sub-pixel super-resolution pyramid network proposed in this invention fully utilizes the learning and parsing capabilities of convolutional neural networks, learning teacher lateral feature information introduced by the cross-scale distillation mechanism while parsing high-level semantic information, and fully fusing high-level and low-level semantics in a top-down path, thereby improving the detection accuracy of small targets in remote sensing images.
[0054] This invention also provides a method for detecting small targets in remote sensing images, comprising:
[0055] S1. Prepare training and validation set images and preprocess them according to experimental requirements;
[0056] The dataset used in this embodiment of the invention is the DIOR dataset and the NWPU VHR-10. Interpolation techniques are used to create one-to-one pairs of high- and low-resolution images in the dataset. The size of the high-resolution image is twice the size of the low-resolution image.
[0057] S2. Train the teacher network using high-resolution images;
[0058] This invention can be applied to a variety of detectors; the following description uses its application on RetinaNet as an example.
[0059] Step S2 includes the following sub-steps:
[0060] S21. Train the teacher network using high-resolution images, and stop training once the teacher network test loss is stable.
[0061] S22. Select the best test result weight as the initial weight of the teacher network and freeze it.
[0062] S3. Construct and train the student network. The student network uses the same detector as the teacher network and is trained using low-resolution images. It learns the teacher's lateral features using a cross-scale distillation mechanism, and the sub-pixel convolutional pyramid network parses and fuses high-level semantic information.
[0063] Step S3 includes the following sub-steps:
[0064] S31. First, the high-resolution image and the low-resolution image are input into the trained teacher network and the untrained student network, respectively, for preliminary feature extraction. The teacher network obtains the teacher's lateral features C. T3 C T2 C T1 The student network obtains the student's lateral characteristics C S3 C S2 C S1 .
[0065] S32, Teacher Horizontal Characteristics Obtained from Teacher Networks (C) T3 C T2 C T1 The teacher network features P are obtained by fusing the feature pyramids. T3 P T2 P T1 Student lateral characteristics C of the student network S3 C S2 C S1 The input is fed into a Sub-pixel Super-Resolution Feature Pyramid Network (SSRFPN). The student network features P... S3 To P S2 The following steps are described using the fusion process as an example.
[0066] S33. Before inputting the feature P into the sub-pixel super-resolution module, S3 ∈R B×C×H×W The feature tmp∈R is obtained through dimensionality transformation. (B×C)×1×H×W .
[0067] S34. Then input tmp into the subpixel convolution module, and obtain tmp′∈R after three convolutional layers. (B×C)×4×H×W tmp′ obtains feature P after subpixel convolution. S3 ′∈R B×C×2H×2W This process uses teacher network characteristics P T2 For monitoring signals, P S3 Interpolation amplification and teacher network characteristics P T2 With the same size, the mask technique is used to filter out the target region between the two features, and then the loss is calculated. Student network feature P S2 To P S1 The process is the same.
[0068] S35, P S3 ′ and S31 student lateral characteristics C S2 Fusion to obtain student network characteristics P S2 Student network characteristics PS2 The features are used for subsequent target detection.
[0069] In some embodiments, the sub-pixel feature super-resolution pyramid network in S32 replaces the traditional interpolation method in the feature pyramid network with a sub-pixel super-resolution module.
[0070] In some embodiments, the three convolutional layers in the subpixel convolution module of S34 consist of a 5×5 convolutional kernel and two 3×3 convolutional kernels, and then the subpixel convolutions are linked together.
[0071] S4: Use the trained object detection network to perform object detection on the remote sensing image to be detected.
[0072] To demonstrate the detection performance of the proposed solution, experiments were conducted on the NWPU VHR-10 and DIOR datasets using three detectors: YOLOv3, RetinaNet, and Faster R-CNN. Targets in the datasets were classified according to the following criteria: targets with fewer than 32×32 pixels were classified as small targets, those larger than 96×96 pixels were classified as large targets, and the rest as medium-sized targets. The experimental hardware platform consisted of a server equipped with two RTX 3090Ti 24G graphics cards and an Intel i9-10920x processor. The server software platform was based on the Linux 20.04 operating system, using CUDA 11.1.0, CUDNN 11.1, and PyTorch 1.8.0.
[0073] The robustness of the algorithm is evaluated by the mean of the average precision of 10 IOUs between the IOU threshold of 0.5 and 0.95 (interval of 0.05).
[0074]
[0075] Table 1 Target detection results on NWPU VHR-10
[0076]
[0077] Table 2 shows the target detection results on DIOR.
[0078] As can be seen from the results in Tables 1 and 2, the average accuracy of the present invention has been significantly improved across multiple detectors and data, especially in the performance of small targets.
[0079] Combination Figure 4 As shown, this embodiment of the invention provides a small target detection device for remote sensing images, comprising:
[0080] Dataset construction unit 401 is used to construct a dataset and preprocess the image data in the dataset, the dataset including high-resolution images and low-resolution images;
[0081] The network construction unit 402 is used to construct a teacher network using the high-resolution image and use the output feature map of the trained teacher network as a supervision signal for the student network.
[0082] Training unit 403 is used to train the student network using the low-resolution image, learn the teacher lateral features of the teacher network through a cross-scale distillation mechanism, and parse and fuse high-level semantic information through a sub-pixel convolutional pyramid network. The trained teacher lateral features are used as supervision signals to guide the student network to learn target region features.
[0083] The detection unit 404 is used to perform target detection on the remote sensing image to be detected by using the trained student network as a target detection network.
[0084] In some embodiments, the dataset construction unit 401 is specifically used for:
[0085] A dataset is constructed, and interpolation techniques are used to set up one-to-one high- and low-resolution image pairs in the dataset, wherein the size of the high-resolution image is twice the size of the low-resolution image.
[0086] In some embodiments, the network construction unit 402 is specifically used for:
[0087] The teacher network is trained using high-resolution images. Training is stopped once the teacher network test loss stabilizes. The weights with the best test results are selected as the weights of the teacher network and then frozen.
[0088] In some embodiments, the training unit 403 is specifically used for:
[0089] The high-resolution image and the low-resolution image are respectively input into the trained teacher network and the untrained student network for preliminary feature extraction. The teacher network obtains the teacher's lateral feature C. T3 C T2 C T1 The student network obtains the student's lateral features C. S3 C S2 C S1 ;
[0090] The teacher's horizontal characteristic C T3 C T2 C T1 The teacher network features P are obtained by fusing the feature pyramids. T3 P T2 PT1 Using the teacher network feature P T3 P T2 P T1 As a supervisory signal, the student's lateral feature C S3 C S2 C S1 The input is fused into a sub-pixel feature super-resolution pyramid network to obtain the student network feature P. S3 P S2 P S1 The student network feature P S3 P S2 P S1 Used for target detection.
[0091] In some embodiments, the teacher network and the student network use the same detector.
[0092] This invention provides a method for small target detection in remote sensing images, comprising: a dataset construction unit for constructing a dataset and preprocessing the image data in the dataset, the dataset including high-resolution images and low-resolution images; a network construction unit for constructing a teacher network using the high-resolution images, and using the output feature map of the trained teacher network as a supervision signal for a student network; a training unit for training the student network using the low-resolution images, learning the teacher's lateral features through a cross-scale distillation mechanism, and fusing high-level semantic information through a sub-pixel convolutional pyramid network, using the trained teacher's lateral features as a supervision signal to guide the student network in learning target region features; and a detection unit for using the trained student network as a target detection network to perform target detection on the remote sensing image to be detected. This invention proposes a cross-scale distillation mechanism and uses high-resolution images to train the teacher network. By utilizing the trained teacher lateral features as a supervisory signal to guide the student network in learning target region features, the sub-pixel super-resolution pyramid network proposed in this invention fully leverages the learning and parsing capabilities of convolutional neural networks. While parsing high-level semantic information, it learns teacher lateral feature information introduced by the cross-scale distillation mechanism, fully integrating high-level and low-level semantics on a top-down path, thereby improving the detection accuracy of small targets in remote sensing images.
[0093] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.
[0094] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for detecting small targets in remote sensing images, characterized in that, include: A dataset is constructed, and the image data in the dataset is preprocessed. The dataset includes high-resolution images and low-resolution images. A teacher network is constructed using the high-resolution image, and the output feature map of the trained teacher network is used as a supervision signal for the student network. The student network is trained using the low-resolution image, the teacher lateral features of the teacher network are learned through a cross-scale distillation mechanism, and high-level semantic information is fused through a sub-pixel convolutional pyramid network. The trained teacher lateral features are used as a supervision signal to guide the student network to learn target region features. This process includes: inputting the high-resolution image and the low-resolution image into the trained teacher network and the untrained student network, respectively, to perform preliminary feature extraction, wherein the teacher network obtains the teacher's lateral features. , , The student network obtains the lateral characteristics of students. , , The teacher's horizontal characteristics , , Teacher network features are obtained by fusing features from the feature pyramid. 、 、 Utilizing the aforementioned teacher network characteristics 、 、 As a supervisory signal, the student lateral features , , The input is fused into a sub-pixel feature super-resolution pyramid network to obtain student network features. 、 、 The student network characteristics 、 、 Used for target detection; the sub-pixel feature super-resolution pyramid network replaces the interpolation method in the feature pyramid network with a sub-pixel super-resolution module; The trained student network is used as a target detection network to perform target detection on the remote sensing image to be detected.
2. The method for detecting small targets in remote sensing images according to claim 1, characterized in that, The dataset is constructed, and the image data in the dataset is preprocessed. The dataset includes high-resolution images and low-resolution images, including: A dataset is constructed, and interpolation techniques are used to set up one-to-one high- and low-resolution image pairs in the dataset, wherein the size of the high-resolution image is twice the size of the low-resolution image.
3. The method for detecting small targets in remote sensing images according to claim 2, characterized in that, The step of constructing a teacher network using the high-resolution image and using the output feature map of the trained teacher network as a supervision signal for the student network includes: The teacher network is trained using high-resolution images. Training is stopped once the teacher network test loss stabilizes. The weights with the best test results are selected as the weights of the teacher network and then frozen.
4. The method for detecting small targets in remote sensing images according to claim 1, characterized in that, The teacher network and the student network use the same detector.
5. A device for detecting small targets in remote sensing images, characterized in that, include: A dataset construction unit is used to construct a dataset and preprocess the image data in the dataset, which includes high-resolution images and low-resolution images; The network construction unit is used to construct a teacher network using the high-resolution image and use the output feature map of the trained teacher network as a supervision signal for the student network. The training unit is used to train the student network using the low-resolution image, learn the teacher lateral features of the teacher network through a cross-scale distillation mechanism, and parse and fuse high-level semantic information through a sub-pixel convolutional pyramid network. The trained teacher lateral features are used as supervision signals to guide the student network to learn target region features. The sub-pixel feature super-resolution pyramid network replaces the interpolation method in the feature pyramid network with a sub-pixel super-resolution module. The training unit is specifically used for: The high-resolution image and the low-resolution image are respectively input into the trained teacher network and the untrained student network for preliminary feature extraction. The teacher network obtains the teacher's lateral features. , , The student network obtains the lateral characteristics of students. , , ; The horizontal characteristics of teachers , , Teacher network features are obtained by fusing features from the feature pyramid. , , Utilizing the aforementioned teacher network characteristics , 、 As a supervisory signal, the student lateral features , , The input is fused into a sub-pixel feature super-resolution pyramid network to obtain student network features. 、 、 The student network characteristics 、 、 Used for target detection; The detection unit is used to perform target detection on the remote sensing image to be detected by using the trained student network as a target detection network.
6. The remote sensing image small target detection device according to claim 5, characterized in that, The dataset construction unit is specifically used for: A dataset is constructed, and interpolation techniques are used to set up one-to-one high- and low-resolution image pairs in the dataset, wherein the size of the high-resolution image is twice the size of the low-resolution image.
7. The remote sensing image small target detection device according to claim 5, characterized in that, The network construction unit is specifically used for: The teacher network is trained using high-resolution images. Training is stopped once the teacher network test loss stabilizes. The weights with the best test results are selected as the weights of the teacher network and then frozen.
8. The remote sensing image small target detection device according to claim 5, characterized in that, The teacher network and the student network use the same detector.
Citation Information
Patent Citations
Convolutional neural network optimization method based on knowledge distillation
CN108764462A
Neural network acceleration method based on cross-resolution knowledge distillation
CN111160533A