Method and system for detecting tiny ship target in thermal infrared image based on hierarchical feature attention network

By applying a hierarchical feature attention network method in thermal infrared images, the problems of information loss and insufficient feature extraction are solved, and high-precision micro-ship target detection is achieved, which significantly improves detection accuracy and robustness.

CN120107819AActive Publication Date: 2025-06-06SHANGHAI INSTITUTE OF TECHNICAL PHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510311385.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-06
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The prior art has problems of information loss and insufficient feature extraction in the detection of micro-ship targets in thermal infrared images, resulting in low detection accuracy, especially in complex background environments that are prone to missed detection and false detection.

Method used

Using a method based on hierarchical feature attention network, a three-channel infrared ship object detection data set with pixel-level annotation is constructed, combining multi-level detail enhancement and multi-level large-nuclear attention mechanism, multi-level long-distance information is extracted and feature fusion is performed to achieve high-precision extraction and positioning of ship contours.

Benefits of technology

It significantly improves the accuracy and robustness of ship target detection, improves detection accuracy, reduces missed and missed detection, and enhances the model's adaptability to complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107819A_ABST
    Figure CN120107819A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for detecting a tiny ship target in a thermal infrared image based on a hierarchical feature attention network, and belongs to the technical field of satellite remote sensing image processing and application. According to the method, a first three-channel infrared ship target detection image data set based on SDGSAT-1 pixel-level annotation is constructed; the data set covers more scenes and contains abundant heat information, and is more suitable for a micro ship detection task in a thermal infrared image. Meanwhile, a network model based on a hierarchical feature attention mechanism is constructed, and the model generates feature maps of different scales through multi-stage detail enhancement processing; extracting multi-level long-distance information through a multi-level big kernel attention mechanism; a prediction target probability graph is generated through multi-level fusion of feature graphs of different scales, fusion and interaction of features of different scales are realized, and a ship can be accurately positioned. In a ship target detection task, the detection accuracy, the false alarm rate and the intersection-to-parallel ratio index are obviously superior to those of an existing model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of satellite remote sensing image processing and application technology, and in particular to a method and system for detecting tiny ship targets in thermal infrared images based on a hierarchical feature attention network. Background Art

[0002] Accurate and reliable ship target detection is of great significance for ensuring the steady development of the economy and maintaining the safety of maritime transportation. Remote sensing images can provide a wide range of information, which provides a potential solution for space-based ship detection. However, different earth observation satellites often carry sensors with different band settings and large image sizes. In an image, the detection of tiny ship targets becomes a challenging task due to the complex background environment, the diversity of ship appearance and surrounding targets, and the low image resolution. Therefore, the establishment of intelligent ship detection methods combined with remote sensing images is of great significance in sustainable marine management.

[0003] At present, images used for ship detection are mainly concentrated in the visible light band and the infrared band. Among them, visible light images are easily affected by weather changes. In SAR images, due to the constant changes in natural factors such as weather and wind speed, uneven sea surface height may occur, which is not conducive to the application of SAR images. The infrared band can penetrate the atmosphere more effectively, especially in the marine environment, where ships usually produce obvious heat differences, so thermal infrared images may be easier to identify ships and are not easily affected by the surrounding environment. Using the satellite's infrared band to identify ships is a typical infrared small target detection task (IRSTD). Traditional IRSTD methods include filter-based background methods (see, “Infrared small target detection via low-rank tensor completion with top-hat regularization,” IEEE Transactions on Geoscience and Remote Sensing, vol.58, no.2, pp.1004-1016, 2019); local contrast-based methods (see, “A local contrast method for infrared small-target detection utilizing a tri-layer window,” IEEE Geoscience and Remote Sensing Letters, vol.17, no.10, pp.1822-1826, 2020) and local rank-based methods (see, “Infrared small target detection via non-convex rank approximation minimization joint l 2,1norm,” Remote Sensing, vol.10, no.11, pp.1821, 2018), which have made some progress. Due to the need to manually design parameters and extract thresholds, the performance is limited in remote sensing images with complex environments and exhibits weak generalization ability.With the rapid development of deep learning, Hou et al. used fully connected layers in skip connections to enhance the contrast between targets and backgrounds in infrared images (see, "ISTDU-Net: Infrared Small-Target Detection U-Net," IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1-5, 2022); in order to achieve the interaction between low-level and high-level features, Li et al. proposed a dense nested attention network (DNANet) and introduced cascaded channels and spatial attention modules (CSAM) to enhance multi-level features (see, "Dense nested attention network for infrared small target detection," IEEE Transactions on Image Processing, vol. 32, no., pp. 1745-1758, 2023). Although convolutional neural networks (CNNs) are indeed effective in extracting local information, the extraction of remote information is crucial in the context of complex structures in remote sensing images. Wu et al. proposed the Space-based Infrared Ship Small Target Dataset (NUDT-SIRST-Sea), introduced the hybrid encoder structure of VisionTransformer (ViT) and CNN, and constructed a multi-level TransUNet (MTU-Net) to extract features at all levels (see, "MTU-Net: Multilevel TransUNet for space-based infrared tiny ship detection," IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1-15, 2023). However, when extracting multi-level global information, the relationship between the features at each level is not utilized.

[0004] So far, domestic and foreign scholars have proposed many ship detection methods based on deep learning, but there are still some obvious defects: (1) The model does not fully call on the feature information of different scales in the global remote sensing image, which is not conducive to the segmentation of small target objects from the image. (2) Limited by the quantity and quality of the data set and the selection of different satellite bands, the detection accuracy of the model will also be affected. (3) Due to the complex background environment of the image and the diversity of surrounding targets, missed detection and false detection often occur during detection. Summary of the invention

[0005] In view of the demand for rich data set information in the task of detecting tiny ship targets in thermal infrared images, and the problems of information loss and feature information extraction in target detection models, the present invention provides a method and system for detecting tiny ship targets in thermal infrared images based on a hierarchical feature attention network. The present invention provides rich thermal information by constructing a three-channel infrared ship target detection data set with pixel-level annotations, thereby providing richer and more accurate data support for feature extraction of subsequent deep learning models. The constructed network model based on the hierarchical feature attention mechanism plays a key role in the input, extraction, aggregation and target positioning of multi-scale features, thereby achieving full extraction and efficient utilization of multi-level features, thereby achieving high-precision extraction and positioning of ship contours.

[0006] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:

[0007] A method for detecting tiny ship targets in thermal infrared images based on a hierarchical feature attention network comprises the following steps:

[0008] S1: Collect and preprocess the detection area image data of the three thermal infrared bands of SDGSAT-1 to establish a three-channel infrared ship target detection dataset with pixel-level annotations;

[0009] S2: Construct a network model based on the hierarchical feature attention mechanism and use it to identify ships in thermal infrared images. The recognition process of this network model is divided into four stages: input stage, feature extraction stage, prediction stage and positioning stage, which include:

[0010] S21: In the input stage, the obtained three-channel infrared ship target data set is subjected to multi-level detail enhancement processing to generate feature maps of different scales;

[0011] S22: In the feature extraction stage, multi-level long-distance information is extracted from the multi-scale feature maps, and the long-distance information is fused into the feature maps of different scales to generate feature maps of different scales containing the long-distance information;

[0012] S23: In the prediction stage, feature maps of different scales containing long-range information are fused to generate a predicted target probability map;

[0013] S24: In the positioning stage, the generated predicted target probability map is used to locate the ship target using the eight-connected aggregation algorithm.

[0014] Preferably, for step S1, the detection area image data of three thermal infrared bands of SDGSAT-1 are collected and preprocessed to establish a three-channel infrared ship target detection data set with pixel-level annotations, including the following steps:

[0015] S11: Collect thermal infrared image data of three thermal infrared bands from SDGSAT-1 simultaneously, and combine the three thermal infrared bands to form a three-channel thermal infrared image dataset;

[0016] S12: performing correction processing on the collected three-channel thermal infrared image data set through radiometric calibration to generate a three-channel thermal infrared image data set with amplitude brightness values;

[0017] S13: Use the SAM model to annotate the ship targets in the thermal infrared image dataset and obtain a three-channel infrared ship target detection dataset with pixel-level annotations. Each image in the dataset is a three-channel image composed of three thermal infrared images.

[0018] Preferably, for step S21, in the input stage, the obtained three-channel infrared ship target data set is subjected to multi-level detail enhancement processing to generate feature maps of different scales, specifically:

[0019] S211: Construct a multi-level detail enhancement mode, with each level adopting a single-input dual-output processing structure;

[0020] S212: In each different level, the first output path will not process the input image and directly send it to the next level; the second output path will first obtain images of different sizes by downsampling, and then extract features of the input images of different sizes by using different numbers of residual blocks to obtain local features of different scales;

[0021] S213: local features of different scales are fused into images of different sizes, and feature maps of different scales are generated at different levels.

[0022] Preferably, the second output path adopts a dual-branch structure, each branch structure includes a different number of residual blocks for respectively extracting semantic information and detail information of the image.

[0023] Preferably, for step S22, in the feature extraction stage, a multi-level encoder-large core attention mechanism-decoder structure mode is adopted to extract multi-level long-distance information from multi-scale feature maps, and then the long-distance information is fused into feature maps of different scales to generate feature maps of different scales containing long-distance information, specifically:

[0024] S221: The multi-level encoder encodes feature maps of different scales to generate multi-level features;

[0025] S222: The multi-level features output by the multi-level encoder are fused with the feature maps of different scales obtained in the input stage according to the corresponding levels to obtain a new multi-scale feature map;

[0026] S223: Extract multi-level long-distance information from multi-scale feature maps through a multi-level large-core attention mechanism;

[0027] S224: Multi-level long-distance information is aggregated into feature maps of different scales through feature concatenation operations, and then sent to multi-level decoders for decoding and restoration into feature maps of different scales, which contain long-distance information.

[0028] Preferably, for step S223, each level of the large core attention mechanism includes the following processing operations:

[0029] S2231: Reconstruct the multi-scale feature maps through a convolution layer with a step size equal to the convolution kernel size to unify the dimensions of the feature maps. The calculation expression is:

[0030] V i =Conv(E i ; s=ks=RF i )

[0031] Among them, V i Represents the feature map of the i-th scale of the reconstructed output; E i Represents the i-th scale feature map of the input, RF i Represents the reduction factor of the i-th level feature. The reduction factors used in feature maps of different scales are different. Conv represents standard convolution. s and ks represent the step size and convolution kernel size respectively.

[0032] S2232: The reconstructed multi-scale feature map is first batch normalized, and then enters the large kernel attention mechanism, and performs point-by-point convolution, large kernel convolution, and point-by-point convolution again in sequence to extract multi-level long-distance information; among them, the large kernel convolution is decomposed into three parts, and the large kernel convolution LKA calculation expression is:

[0033]

[0034] Among them, I represents the input of the large kernel convolution; DW represents the depth convolution with a convolution kernel size of (2d-1)×(2d-1); DDW represents the convolution kernel size of PW represents the point-by-point convolution with a kernel size of 1×1; d is the convolution expansion rate; K is the preset conventional parameter;

[0035] S2233: Transmit multi-level long-distance information to a multi-layer perceptron for information enhancement. The multi-layer perceptron includes 1x1 point convolution, 3x3 depth convolution, Gaussian error linear unit and 1x1 point convolution.

[0036] Preferably, for step S23, in the prediction stage, feature maps of different scales containing long-distance information are fused to generate a predicted target probability map, specifically:

[0037] S231: Upsampling the feature maps of different scales to generate feature maps of different scales with unified dimensions;

[0038] S232: Use the attention mechanism to assign different weights to feature maps of different scales, and perform element-by-element addition processing on the weighted feature maps to generate the final predicted target probability map. The weight α calculation formula is:

[0039] β=Conv(Conv(GAP(D);ks=1);ks=1)

[0040] γ=Conv(Conv(GMP(D);ks=1);ks=1)

[0041] α=γ+β

[0042] Among them, D represents feature maps of different scales; Conv represents standard convolution; ks=1 represents the convolution kernel size of 1x1; GAP represents global average pooling; GMP represents global maximum pooling; β and γ represent adaptive parameters for feature fusion.

[0043] Preferably, for step S24, in the positioning stage, the generated predicted target probability map is used to locate the ship target using an eight-connected aggregation algorithm, specifically:

[0044] S241: Binarizing the predicted target probability map using a preset threshold to generate a binary map;

[0045] S242: Using the eight-connected aggregation algorithm to count the pixel points contained in each ship target in the binary image, the ship target is positioned.

[0046] Preferably, on the basis of the determination of the ship target, the center of mass coordinates of the ship target are calculated and determined by a weighted center of mass positioning algorithm. The weighted center of mass positioning algorithm is specifically as follows: first, based on the located ship target, the horizontal and vertical coordinate information of all the pixel points contained therein is obtained; then the pixel values ​​of all the pixel points of the ship target are normalized to obtain weights, and finally the center of mass of the ship target is calculated based on the weights and the horizontal and vertical coordinate information.

[0047] Preferably, the constructed network model based on the hierarchical feature attention mechanism adopts a deep supervision strategy to evaluate the model. By adding a loss function in each layer of the network model, the network model can learn discriminative features at different scales, aggregate the loss values ​​of the final outputs of different layers, and then adjust the network model parameters to reduce the loss value to improve the model performance.

[0048] The present invention also provides a system for detecting tiny ship targets in thermal infrared images based on a hierarchical feature attention network, using the aforementioned method for detecting tiny ship targets in thermal infrared images based on a hierarchical feature attention network, the system includes a thermal infrared image data processing module and a ship target recognition module;

[0049] The thermal infrared image data processing module pre-processes the image data of the three thermal infrared bands of SDGSAT-1 and outputs a three-channel infrared ship target detection dataset with pixel-level annotations;

[0050] The ship target recognition module is used to build a network model based on the hierarchical feature attention mechanism to perform thermal infrared image ship target recognition. The ship target recognition module includes a multi-level detail enhancement module, a multi-level large core attention module, a multi-level feature fusion module and a ship target positioning module;

[0051] The multi-level detail enhancement module is used in the input stage to generate feature maps of different scales by performing multi-level detail enhancement processing on the input three-channel infrared ship target data set;

[0052] The multi-level large core attention module is used in the feature extraction stage. It uses the multi-level large core attention mechanism to extract multi-level long-distance information from the input multi-scale feature maps, and then fuses the long-distance information into feature maps of different scales to generate feature maps of different scales containing long-distance information.

[0053] The multi-level feature fusion module is used in the prediction stage to generate a predicted target probability map by fusing the input feature maps of different scales containing long-range information;

[0054] The ship target positioning module is used in the positioning stage to use the generated predicted target probability map, adopt the eight-connected aggregation algorithm to locate the ship target, and locate the center of mass coordinates of the ship through the weighted center of mass positioning algorithm.

[0055] Compared with the prior art, the beneficial effects of the present invention are:

[0056] 1) The present invention combines the thermal infrared image data of SDGSAT-1 to construct the first three-channel infrared ship target detection image dataset based on SDGSAT-1 with pixel-level annotations. Compared with other mainstream datasets, this dataset has the highest imaging width, the smallest target size, the largest image size, the smallest average target-background ratio, the most complex observation scene, and the largest number of bands, making it more suitable for the task of detecting small ship targets in thermal infrared images. The three thermal infrared bands of this dataset can be displayed synchronously in the image, providing rich thermal information, which provides strong support for the subsequent deep learning model to fully extract thermal information.

[0057] 2) The present invention constructs a network model based on a hierarchical feature attention mechanism. In the input stage of the network model, local features of different scales are fused with images of different sizes through multi-level detail enhancement processing to generate feature maps of different scales, thereby enhancing the representation of detail information of different scales, effectively compensating for the useful information that may be lost during the downsampling process, and providing richer and more accurate data support for the next step of feature extraction. In the feature extraction stage, multi-level long-distance information is extracted through a multi-level large-core attention mechanism. The multi-layer large-core attention mechanism can comprehensively capture and utilize global information at multiple levels while reducing the computational complexity, thereby compensating for the disadvantage that the convolutional neural network model only focuses on local features, enabling the model to learn the relationship between targets and backgrounds at different scales, and effectively suppressing the interference of complex background environments and the diversity of surrounding targets on ship detection. In the prediction stage, feature maps of different scales are fused to generate a final predicted target probability map, realizing feature fusion and interaction of different scales, which can significantly improve the segmentation accuracy and enhance the comprehensive representation capability of details and global information, effectively improving the adaptability of the model to complex scenes, and finally in the positioning stage, the eight-connected aggregation algorithm and the weighted centroid positioning algorithm are used to realize the accurate extraction and positioning of the ship shape. Compared with the existing convolutional neural network model, the network model proposed in the present invention shows better performance in the ship target detection task, and its key indicators such as detection accuracy (Pd), false alarm rate (Fa) and intersection over union (IoU) are significantly better than those of the existing network models. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 A flow chart of a method for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network according to an embodiment of the present invention;

[0059] Figure 2 is a three-channel image of the SDG-IRSTD data set of Example 2 of the present invention;

[0060] Figure 3 is a single-channel image of the NUDT-SIRST-Sea dataset of Example 2 of the present invention;

[0061] Figure 4 The target size distribution histogram of the six data sets in Example 2 of the present invention;

[0062] Figure 5 The target quantity distribution histogram of the six data sets in Example 2 of the present invention;

[0063] Figure 6 The target-background ratio distribution histogram of the six data sets in Example 2 of the present invention;

[0064] Figure 7 This is an overall structural diagram of a system for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network according to an embodiment of the present invention. DETAILED DESCRIPTION

[0065] The following will be combined with the attached embodiment of the present invention Figure 1 -Attached Figure 7 , the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0066] Example 1

[0067] Combination Figure 1 As shown, an embodiment of the present invention provides a method for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network, comprising the following steps:

[0068] S1: Collect and preprocess the detection area image data of the three thermal infrared bands of SDGSAT-1 to establish a three-channel infrared ship target detection dataset with pixel-level annotations;

[0069] S2: Construct a network model based on the hierarchical feature attention mechanism and use it to identify small targets of ships in thermal infrared images. The recognition process of this network model is divided into four stages, namely, input stage, feature extraction stage, prediction stage and positioning stage, which include:

[0070] S21: In the input stage, the obtained three-channel infrared ship target data set is subjected to multi-level detail enhancement processing to generate feature maps of different scales;

[0071] S22: In the feature extraction stage, multi-level long-distance information is extracted from the multi-scale feature maps, and the long-distance information is fused into the feature maps of different scales to generate feature maps of different scales containing the long-distance information;

[0072] S23: In the prediction stage, feature maps of different scales containing long-range information are fused to generate a predicted target probability map;

[0073] S24: In the positioning stage, the generated predicted target probability map is used to locate the ship target using the eight-connected aggregation algorithm.

[0074] Example 2

[0075] Based on Example 1, this example describes in detail the implementation method of generating a three-channel infrared ship target detection dataset with pixel-level annotations in step S1.

[0076] In this embodiment, for step S1, the detection area image data of the three thermal infrared bands of SDGSAT-1 are collected and preprocessed to establish a three-channel infrared ship target detection data set with pixel-level annotations, which specifically includes the following steps:

[0077] S11: Collect thermal infrared image data of three thermal infrared bands from SDGSAT-1 at the same time, combine the three thermal infrared bands to form a three-channel thermal infrared image data set, which is convenient for the subsequent data set production. Among them, SDGSAT-1 is the Sustainable Development Satellite No. 1 launched by my country, which can carry three payloads of thermal infrared, low-light and multi-spectral imagers at the same time. SDGSAT-1 has three thermal infrared detection bands, and the three thermal infrared work synchronously. The spatial resolution of the three thermal infrared bands data collected by SDGSAT-1 is 30m, which can distinguish the temperature difference of 0.2 degrees Celsius in a large dynamic range, which is helpful for the infrared small target detection task (Infrared Small Target Detection, IRSTD).

[0078] S12: The collected three-channel thermal infrared image data set is corrected by radiometric calibration to generate a three-channel thermal infrared image data set with amplitude and brightness values. Because the thermal infrared image data acquired by the SDGSAT-1 sensor is a kind of original, unprocessed pixel value, which cannot directly reflect any physical meaning, that is, it has no dimension, therefore, it is necessary to convert the obtained pixel value into an amplitude and brightness value with actual physical meaning through radiometric calibration. The calculation expression is:

[0079] L=GAIN·DN+BIAS

[0080] Where L represents the radiation brightness value; GAIN and BIAS represent the gain coefficient and bias coefficient of the corresponding band respectively.

[0081] S13: Use the SAM model to annotate the ship targets in the thermal infrared image dataset using a semi-automatic annotation method. The SAM model is a large model tool that supports multiple segmentation types. The SAM model can automatically identify ship targets by point selection or box selection, and generate corresponding masks and bounding boxes, but the SAM model may have annotation errors when processing some images. Therefore, the problematic data is manually refined to ensure the quality of the dataset, and finally a three-channel infrared ship target detection dataset with pixel-level annotations is obtained. Each image in the dataset is a three-channel image composed of three thermal infrared images.

[0082] The following will further illustrate the implementation method of generating a three-channel infrared ship target detection dataset with pixel-level annotations by listing specific examples, and compare it with a ship target detection dataset obtained by other conventional methods to illustrate the technical effects brought about by the present invention:

[0083] The collected thermal infrared image data comes from the Thermal Infrared Spectrometer (TIS) band images of SDGSAT-1. SDGSAT-1 is the Sustainable Development Satellite No. 1, and is the first scientific satellite launched by my country to meet the data coverage requirements for sustainable development applications. SDGSAT-1 carries three payloads: thermal infrared, low-light and multi-spectral imagers, which can achieve all-day, multi-payload collaborative observation. The revisit period of SDGSAT-1 is about 11 days, the spatial resolution of the TIS band is 30m, the imaging width can reach 300km, and it has 3 detection bands. Its specific parameters are shown in Table 1.

[0084] Table 1

[0085] Band wavelength Gain Bias TIS-1 8.0-10.5μm 0.003947 0.167126 TIS-2 10.3-11.3μm 0.003946 0.124622 TIS-3 11.5-12.5μm 0.005329 0.222530

[0086] Among the three bands in the table, TIS-1 (8.0–10.5 μm) is mainly used to monitor surface temperature and heat distribution. This spectral range is highly sensitive to surface radiation and can accurately reflect changes in surface temperature. The TIS-2 (10.3–11.3 μm) spectral range has high penetration of water vapor and clouds in the atmosphere, making it suitable for monitoring atmospheric temperature and humidity. The TIS-3 (11.5–12.5 μm) spectral range is highly sensitive to greenhouse gases such as carbon dioxide on the surface and in the atmosphere, and can be used to monitor and study greenhouse gas concentrations. The comprehensive application of long-wave infrared bands enables thermal infrared imaging to accurately capture thermal radiation information.

[0087] In the specific implementation process, the three synchronous TIS bands in the collected thermal infrared image data are combined, and then combined with the values ​​in Table 1, radiation calibration processing and ship target annotation are performed, and finally a three-channel infrared ship target detection image dataset with pixel-level annotations is established, referred to as the SDG-IRSTD dataset. The SDG-IRSTD dataset of this embodiment collects a total of 329 images from SDGSAT-1, including 3492 ship targets, and the size of each image is 1024x1024, wherein each image is composed of TIS1-TIS3.

[0088] The basic parameters of the SDG-IRSTD dataset are compared with those of four existing public datasets (NUST-SIRST, IRSTD-1k, NUDT-SIRST, and NUDT-SIRST-Sea). NUST-SIRST (see Wang et al., 2019, Miss detection vs. false alarm: adversarial learning for small objects segmentation in infrared images.) is a dataset of 428 real and synthetic infrared images with resolutions between 256x256 and 512x512 pixels; IRSTD-1k (see Zhang et al., 2022, ISNet: shapematters for infrared small target detection.) is a dataset of 1000 images with a resolution of 512x512 for infrared small target detection; NUDT-SIRST (see Li et al., 2022, Dense nested attention network for infrared small target detection) is a dataset of 1000 images with a resolution of 512x512 for infrared small target detection. detection.) is an infrared small target dataset containing 1327 images with a resolution of 256x256; NUDT-SIRST-Sea (see, Wu et al., 2023, MTU-Net: multilevel TransUNet for space-based infrared tiny ship detection.) is an infrared dataset containing 17598 images with a resolution of 1024x1024, which is used to detect small ship targets. The comparison results are shown in Table 2.

[0089] Table 2

[0090]

[0091] As can be seen from Table 2, compared with the other five mainstream datasets, SDG-IRSTD has the highest imaging width, the smallest target size, the largest image size, and the largest number of bands, making it more suitable for IRSTD tasks. Most of the other mainstream infrared small target detection datasets are single-channel images, that is, they only contain information from one band.

[0092] Combination Figure 2 and 3 As shown in the figure, by comparing the effects of the two figures, it can be seen that the three thermal infrared bands of the SDG-IRSTD three-channel image data set proposed by the present invention can be displayed synchronously, and the thermal information in the image is richer, so the target brightness is stronger. Multiple thermal infrared bands are helpful to discover the thermal characteristics of small ship targets, and more bands are conducive to ship detection.

[0093] Combination Figure 4 As shown in Table 2, the size of the targets in the SDG-IRSTD dataset is concentrated between 10 and 30, accounting for 69% of the total targets. The other four mainstream datasets account for 67% (NUDT-SIRST-Sea), 63% (NUDT-SIRST), 65% (SIRST), 50% (NUST-SIRST), and 48% (IRSTD-1k). Combined with Table 2, it can be seen that the average target size of the SDG-IRSTD dataset is 23, which has the smallest average size. The average target size of the other four mainstream datasets is about 1.3 to 2.2 times the average size of the SDG-IRSTD dataset. Based on the above analysis, the targets of the SDG-IRSTD dataset are generally smaller than those of other datasets, which is more conducive to the model detecting small space-based ship targets.

[0094] Combination Figure 5 As shown in the figure, compared with other datasets, the ships in the images of the SDG-IRSTD dataset are more evenly distributed. The number of targets in the other datasets is between 0 and 105, and there are almost no cases of dense targets or multiple targets. However, the SDG-IRSTD dataset has more dense ship scenes and more even target distribution, which is crucial for model learning.

[0095] Combination Figure 6As shown in the figure, about 92% of the target-background ratios of the SDG-IRSTD dataset are between 0% and 0.005%, and the average target-background ratio of the SDG-IRSTD dataset is 0.02654%. Compared with other mainstream datasets, the target size in the SDG-IRSTD dataset is relatively small, but the overall image size is larger, indicating that the SDG-IRSTD dataset has the smallest average target-background ratio. In addition, the background of the SDG-IRSTD dataset accounts for a larger proportion, the environmental information is more complex, and it can cover more application scenarios. The SDG-IRSTD dataset covers various scenes, such as offshore scenes, nearshore scenes, thick cloud interference scenes, and weak target scenes. This is because the specific heat capacity of water is large. As heat is absorbed, the sea surface will appear warm in thermal infrared images, making the radiation difference between the ship and the sea surface smaller, resulting in a relatively weak target brightness. This minimum average target-background ratio and complex environmental information undoubtedly increase the difficulty of ship detection and pose a higher challenge to accurate ship detection. However, under such complex data set conditions, the ship target detection method of the present application accurately identifies ship targets, which indirectly proves that this method has better performance than other methods.

[0096] It can be seen that the present invention, by combining the thermal infrared image data of SDGSAT-1, establishes the first three-channel infrared ship target detection image dataset based on SDGSAT-1 with pixel-level annotations. Compared with other mainstream datasets, SDG-IRSTD has the highest imaging width, the smallest target size, the largest image size and the largest number of bands, which can cover more application scenarios and is more suitable for infrared small target ship detection tasks. The three thermal infrared bands can not only be displayed synchronously in the image, but also provide more sufficient thermal information, which is helpful for the subsequent deep learning model to fully extract thermal information and perform deep learning.

[0097] Example 3

[0098] Based on Example 1, this embodiment takes the construction of a 5-layer feature attention mechanism network model as an example, and for step S2, constructs a network model based on the hierarchical feature attention mechanism, and describes in detail the implementation method of using this network model to identify small targets of thermal infrared images of ships.

[0099] In this embodiment, for step S21, in the input stage, the obtained three-channel infrared ship target data set is subjected to multi-level detail enhancement processing to generate feature maps of different scales, specifically:

[0100] S211: Construct a multi-level detail enhancement mode, with each level adopting a single-input dual-output processing structure;

[0101] S212: In each different level, the first output path will not process the input image and directly send it to the next level; the second output path will first obtain images of different sizes by downsampling, and then extract features of the input images of different sizes by using different numbers of residual blocks to obtain local features of different scales;

[0102] S213: local features of different scales are fused into images of different sizes, and feature maps of different scales are generated at different levels.

[0103] Among them, the second output path specifically adopts a dual-branch structure, each branch structure includes a different number of residual blocks, and the two branches are used to extract the semantic information and detail information of the image respectively to enrich the features of the corresponding level.

[0104] Specifically: Given an input image X∈R C×H×W , where C, H, and W represent the number of channels, height, and width of the input image, respectively. The input image of the i-th level can be expressed as At each level, the input images of different scales are used to extract the semantic information and detail information of different levels of images through different numbers of residual blocks, and then fused into images of different scales by element addition to generate feature maps of different scales, where the feature map of the i-th scale is represented as D i , the generated multi-scale feature map is used for the next step of feature extraction.

[0105] The residual block is a neural network structure that directly passes the input to the deep layer through a jump connection, allowing the network to learn the residual more easily. Its jump connection allows the gradient to be directly passed back to the shallow layer, effectively alleviating the gradient vanishing problem of the deep network. In addition, the residual block is used here to extract multi-scale features and enhance the ability to retain detail information.

[0106] This design can enhance the detail information of the image, which is crucial for the subsequent segmentation of the target. In previous IRSTD tasks, the traditional approach is that the input image will be directly sent to the backbone network to extract multi-level features. This method makes insufficient use of the input image information. In order to reduce the amount of calculation and extract high-level features, the backbone network often uses the pooling layer to sample the feature map multiple times. Although this can reduce the size of the feature map, it inevitably causes the loss of image information that is important for the segmentation task. In view of this problem, the present invention performs multi-level detail enhancement processing on the input image data set, thereby generating feature maps of different scales, and supplementing the low-frequency information and high-frequency information of images of different scales through a dual-branch structure. This method can fully mine and utilize image information, and can also effectively make up for the useful information that may be lost during the downsampling process, providing richer and more accurate data support for the next step of feature extraction.

[0107] In this embodiment, for step S22, in the feature extraction stage, a multi-level encoder-large core attention mechanism-decoder structure mode is adopted to extract multi-level long-distance information from multi-scale feature maps, and then the long-distance information is fused into feature maps of different scales to generate feature maps of different scales containing long-distance information. The main steps include:

[0108] S221: The multi-level encoder encodes feature maps of different scales to generate multi-level features to enrich feature representation. The multi-level features include high-resolution low-level features and low-resolution high-level features, which can achieve high-precision image segmentation.

[0109] S222: The multi-level features output by the multi-level encoder are fused with the feature maps of different scales obtained in the input stage according to the corresponding levels to obtain a new multi-scale feature map. Specifically, given an input image represented as X∈R C×H×W , the fused i-th scale feature map is expressed as Where i=1,2,3,4,5.

[0110] S223: Extract multi-level long-distance information from multi-scale feature maps through multi-level large-core attention mechanisms. Each level of large-core attention mechanisms mainly includes batch normalization (BN) mechanism, large-core attention mechanism (Attention module), multi-layer perceptron (MLP), etc. The large-core attention mechanism refers to the use of convolution with a large convolution kernel to extract long-distance information. Long-distance information refers to the dependency between points that are far apart in the feature map. Among them, each level of the large-core attention mechanism includes the following processing operations:

[0111] S2231: Reconstruct the multi-scale feature map through a convolution layer with a step size equal to the convolution kernel size to unify the dimension of the feature map to facilitate subsequent fusion. In this embodiment, the specific expression formula is:

[0112] V i =Conv(E i ; s=ks=RF i ),i=1,2,3,4,5

[0113] Among them, V i Represents the reconstructed feature map of the i-th scale; RF i It represents the reduction factor of the i-th level feature. The reduction factors used in feature maps of different scales are different. Conv represents standard convolution. s and ks represent the step size and convolution kernel size respectively.

[0114] S2232: The reconstructed multi-scale feature map is first batch normalized, and then enters the large core attention mechanism, where point-by-point convolution, large core convolution, and point-by-point convolution are performed in sequence to extract multi-level long-distance information. In the large core attention mechanism, the large core convolution attention will model the long-distance relationship, and the multi-level long-distance information output by the attention mechanism is represented as T i , and its calculation formula is expressed as:

[0115] T i =PW(LKA(GELU(PW(BN(V i ))))))+V i ,i=1,2,3,4,5

[0116] Wherein, PW represents point-wise convolution; LKA represents large kernel convolution; BN represents batch normalization; GELU represents Gaussian Error Linear Unit. In this embodiment, the large kernel convolution is decomposed into three parts, and the specific calculation formula of the large kernel convolution LKA is expressed as:

[0117]

[0118] Among them, I represents the input of the large kernel convolution; DW represents the depth convolution with a convolution kernel size of (2d-1)×(2d-1); DDW represents the convolution kernel size of PW represents the depthwise convolution with a kernel size of 1×1; d is the convolution expansion rate; K is the preset conventional parameter.

[0119] S2233: Transmit the multi-level long-distance information to a multi-layer perceptron for information enhancement. In this embodiment, the multi-layer perceptron (MLP) includes 1x1 point convolution, 3x3 depth convolution, Gaussian error linear unit and 1x1 point convolution. The multi-level long-distance information Oi enhanced by the multi-layer perceptron is specifically expressed as:

[0120] O i =BN(MLP(T i ))+T i

[0121] Among them, O i represents the multi-level long-distance information after multi-layer perceptron information enhancement; BN represents batch normalization; MLP is the multi-layer perceptron information enhancement operation; Ti is the multi-level long-distance information of the input.

[0122] S224: Multi-level long-distance information is aggregated into feature maps of different scales through feature concatenation operations, and then sent to multi-level decoders for decoding and restoration into feature maps of different scales, which contain long-distance information.

[0123] The present invention combines the large-core attention mechanism with multi-scale features. Feature maps of different scales are used to extract multi-level long-distance information through a multi-level large-core attention mechanism. The multi-layer large-core attention mechanism can comprehensively capture and utilize global information at multiple levels while reducing computational complexity, and can better utilize local information and long-distance information to improve the segmentation accuracy of the model.

[0124] In this embodiment, for step S23, in the prediction stage, feature maps of different scales containing long-distance information are fused to generate a predicted target probability map. The specific steps are:

[0125] S231: before fusion, in order to maintain the consistency of dimension, the feature maps of different scales are upsampled to generate feature maps of different scales with unified dimension;

[0126] S232: Use the attention mechanism to assign different weights to feature maps of different scales, and perform element-by-element addition processing on the weighted feature maps to generate the final predicted target probability map. In the specific implementation process, the specific expression formula is:

[0127]

[0128] in, represents the feature map of the i-th scale after upsampling; D i represents the feature map of the i-th scale; Upsample represents the upsampling operation; sf represents the upsampling rate; RF i represents the reduction factor of the i-th level feature; P represents the output predicted target probability map; α i Represents the weight corresponding to the feature map of the i-th scale; Sigmoid is the activation function, which maps the input data between 0 and 1, that is, obtains the probability that each pixel is the predicted target.

[0129] Weight α i It is generated by the self-attention mechanism. In the specific implementation process, α i It can be generated by the convolution layer calculation of the 1x1 convolution kernel, and the weight α calculation formula is:

[0130] β=Conv(Conv(GAP(D);ks=1);ks=1)

[0131] γ=Conv(Conv(GMP(D);ks=1);ks=1)

[0132] α=γ+β

[0133] Among them, D represents feature maps of different scales; Conv represents standard convolution; ks=1 represents the convolution kernel size of 1×1; GAP represents global average pooling; GMP represents global maximum pooling; β and γ represent adaptive parameters for feature fusion.

[0134] Compared with the traditional method that only relies on the final output of the decoder as the final probability likelihood map, the present invention fuses the feature maps of different scales output by the decoder to generate the final predicted target probability map, realizes the fusion and interaction of features of different scales, can significantly improve the segmentation accuracy and enhance the comprehensive representation ability of details and global information, and effectively improves the adaptability of the model to complex scenes.

[0135] In this embodiment, for step S24, in the positioning stage, the generated predicted target probability map is used to locate the ship target using the eight-connected aggregation algorithm. Specifically:

[0136] S241: Binarize the predicted target probability map by a preset threshold to generate a binary map. In the specific implementation process, a suitable threshold is first pre-set, and each pixel in the predicted target probability map is traversed. If the probability value of the pixel is greater than or equal to the threshold, it is assigned a value of 1 (representing the foreground, that is, it may be a ship); if it is less than the threshold, it is assigned a value of 0 (representing the background), thereby converting the probability map into a binary map containing only 0 and 1 values.

[0137] S242: Use the eight-connected aggregation algorithm to count the pixel points contained in each ship target in the binary image to complete the ship target positioning. The specific expression formula is:

[0138] N 8 (pixel 1 )∩N 8 (pixel 2 )≠φ

[0139] Among them, N 8 (pixel 1 ), N 8 (pixel 2 ) represent pixel points 1 ,pixel 2 Specifically, if two pixels have an intersection point in their eight neighborhoods, the pixel point pixel 1 ,pixel 2 can be considered as adjacent pixels; if the pixel values ​​of adjacent pixels are equal, the pixel 1 ,pixel 2 They can be regarded as pixel points of the same target. After obtaining the pixel points corresponding to the target, the ship target can be positioned.

[0140] In addition, unlike the traditional method of obtaining the centroid by averaging the coordinates corresponding to each pixel of the target, the present invention further calculates the centroid coordinates of the ship target through a weighted centroid positioning algorithm. The weighted centroid positioning algorithm is specifically as follows: first, based on the located ship target, the horizontal and vertical coordinate information of all the pixels contained in it is obtained; then the pixel values ​​of all the pixels of the ship target are normalized to obtain the weights, and finally the centroid of the ship target is calculated based on the weights and the horizontal and vertical coordinate information; the calculation formula is:

[0141] c=(w T x,w T y)

[0142]

[0143] Among them, x and y represent the horizontal and vertical coordinates of the pixel points contained in the ship target, v represents the pixel value of the pixel point in the ship target, w represents the weight corresponding to the pixel point; c represents the center of mass of the ship target.

[0144] After HFA-Net generates the predicted target probability map, a threshold is used to binarize the probability map and the eight-connected aggregation algorithm is used to count the pixels contained in each target. If two pixels have an intersection in their eight domains, then the two pixels can be considered as adjacent points. If the pixel values ​​of the pixels are equal, then the two pixels can be regarded as pixels of the same target. After obtaining the pixel points corresponding to the target, the target can be located and the center of mass coordinates of the target can be calculated by the weighted center of mass positioning algorithm.

[0145] In this embodiment, the hierarchical feature attention network model of the present invention adopts a deep supervision strategy to evaluate the model, and a loss function is added to each level of the network model. By using deep supervision, the network model can learn discriminative features at different scales. By measuring the similarity between the model's predicted target probability map and the true value, the segmentation head is allowed to consider the label information at each level. In the specific implementation process, the overall loss function expression formula is:

[0146]

[0147] Among them, L represents the final total loss value; L i Represents the output loss value of the i-th level decoder.

[0148] This paper uses focal loss L F and dice loss L D is the loss function of each level of decoder, and the specific calculation formula is as follows:

[0149] L F= -α(1-p) γ log(p)

[0150]

[0151] L i =L F +L D

[0152] Among them, GT represents the true value map; P represents the predicted target probability map; α and γ represent the hyperparameters used to control the weights; ε represents the smoothing factor to avoid oscillation during training; and p represents the detection probability corresponding to any pixel.

[0153] The present invention can aggregate the final output loss values ​​of different network layers through a deep supervision strategy, and pass the loss values ​​to each layer of the model through back propagation, guiding the model to adjust parameters to reduce the loss, making the detection results closer to the true value, improving model performance, enhancing robustness, and making ship detection more accurate.

[0154] Example 4

[0155] Combination Figure 7 As shown, the embodiment of the present invention also provides a system for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network. The system uses a method for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network described in Example 3. The system includes a thermal infrared image data processing module and a ship target recognition module.

[0156] The thermal infrared image data processing module pre-processes the image data of the three thermal infrared bands of SDGSAT-1 and outputs a three-channel infrared ship target detection dataset with pixel-level annotations;

[0157] The ship target recognition module is used to build a network model based on the hierarchical feature attention mechanism to identify small ship targets in thermal infrared images. The ship target recognition module includes a multi-level detail enhancement module (Multi Level Detail Enhance Module, MLDEM), a multi-level large kernel module (Multi Level Large Kernel Module, MLLKAM), a multi-level feature fusion module (Multi Level Feature Fusion Module, MLFFM) and a ship target positioning module;

[0158] The multi-level detail enhancement module (MLDEM) is used in the input stage to generate feature maps of different scales by performing multi-level detail enhancement processing on the input three-channel infrared ship target data set;

[0159] The Multi-Level Large Kernel Attention Module (MLLKAM) is used in the feature extraction stage. It uses a multi-level large kernel attention mechanism to extract multi-level long-distance information from the input multi-scale feature maps, and then fuses the long-distance information into feature maps of different scales to generate feature maps of different scales containing long-distance information.

[0160] The multi-level feature fusion module (MLFFM) is used in the prediction stage to generate a predicted target probability map by fusing the input feature maps of different scales containing long-range information;

[0161] The ship target positioning module is used in the positioning stage to use the generated predicted target probability map, adopt the eight-connected aggregation algorithm to locate the ship target, and locate the center of mass coordinates of the ship through the weighted center of mass positioning algorithm.

[0162] In this embodiment, the multi-level detail enhancement module is composed of multiple detail enhancement modules (Detail Enhance Module, DEM) in parallel and introduced into the input stage of the network model. Each detail enhancement module is a single-input dual-output structure. Among them, one output path does not process the image, but sends the image to the detail enhancement module of the next level, and the other output path is used to supplement local features of different scales through a dual-branch structure to enhance the detail information of the model. Each branch structure includes a different number of residual blocks, which are used to extract the semantic information and detail information of the image respectively to enrich the features of the corresponding stage.

[0163] In this embodiment, the multi-level large core attention module is a symmetrical multi-level encoder-large core attention mechanism-decoder. In the multi-level large core attention module, the multi-level features output by the multi-level encoder are first fused with the feature maps of different scales obtained in the input stage according to the hierarchical correspondence to obtain a new multi-scale feature map. Then the multi-scale feature map enters the multi-level large core attention module to extract multi-level long-distance information. The multi-level large core attention module is composed of multiple large core units (LKABlock) in parallel, and feature maps of different scales enter multiple large core units (LKABlock) respectively to extract multi-level long-distance information, which is used to fully capture global information at multiple levels. Each large core unit (LKABlock) includes a batch normalization (Batch Normalization, BN) mechanism, a large core attention mechanism (Attention module), a multi-layer perceptron (MLP), etc. In the large core unit, the multi-scale feature map is first reconstructed, and then the long-distance information is extracted through the large core attention mechanism, and finally the information is enhanced through the multi-layer perceptron.

[0164] Finally, multi-level long-distance information will be aggregated into feature maps of different scales through feature concatenation operations, and then enter the multi-level decoder to be decoded and restored to feature maps of different scales. At this time, the feature maps contain long-distance information.

[0165] In this embodiment, the multi-level feature fusion module generates feature maps of different scales with unified dimensions by upsampling the information of the input feature maps of different scales; then uses the attention mechanism to assign different weights to the feature maps of different scales, and adds the weighted feature maps element by element to generate the final predicted target probability map.

[0166] In this embodiment, the ship target positioning module uses the generated predicted target probability map and adopts the eight-connected aggregation algorithm to locate the ship target. On this basis, the center of mass coordinates of the ship are located through the weighted center of mass positioning algorithm.

[0167] The following compares the detection method of the network model based on the hierarchical feature attention mechanism with the detection method of the existing publicly available convolutional neural network model, and compares and analyzes the detection results of each method. The specific results are shown in Table 3.

[0168] Table 3

[0169] method <![CDATA[IoU(×10 -2 )]]> <![CDATA[P d (×10 -2 )]]> <![CDATA[F a (×10 -6 )]]> ACM 46.50 88.97 26.91 ALC-Net 46.09 76.21 46.34 ResU-Net 50.42 84.31 13.10 DNANet 59.07 76.72 8.71 ISTDU-Net 56.07 89.66 36.19 MTU-Net 63.26 90.34 10.54 HFA-Net 65.61 94.31 7.25

[0170] Among them, ACM (see Dai et al., 2021, Asymmetric contextual modulation for infrared small target detection.) mainly uses asymmetric contextual modulation mechanism to realize small target recognition; ALC-Net (see Dai et al., 2021, Attentional local contrast networks for infrared small target detection.) mainly uses attention mechanism and local contrast enhancement to realize small target recognition; ResU-Net (see Diakogiannis et al., 2020, ResUNet-a: A deep learning framework for semantic segmentation of remotely sensed data.) mainly realizes small target recognition through a deep learning model combining U-Net and residual network (ResNet); DNANet (see Li et al., 2022, Dense nested attention network for infrared small target detection.) mainly realizes small target recognition through densely nested attention mechanism; ISTDU-Net (see Hou et al., 2022, ISTDU-net: infrared small-target detection U-Net.) mainly uses U-Net and infrared image characteristics to detect and locate infrared small targets; MTU-Net (see, Wu et al., 2023, MTU-Net: multilevel TransUNet for space-based infrared tiny ship detection.) mainly realizes ship target recognition through a multi-level self-attention U-shaped network; HFA-Net (Hierarchy feature attention network) is a network model based on a hierarchical feature attention mechanism proposed in this application. Through a hierarchical feature attention mechanism, it effectively captures and utilizes multi-scale features and contextual information in the image to complete the detection task. IoU is the intersection over union ratio, which indicates the degree of overlap between the predicted box and the true box; Pd is the detection probability, which refers to the probability of correctly detecting the target; Fa is the false alarm rate, which indicates the probability of incorrectly detecting the target.

[0171] From the quantitative experimental results in Table 3, it can be observed that the detection performance of the HFA-Net model of the present invention is greatly improved compared with other methods. d 、F a They are 65.61%, 94.31%, and 7.25×10 -6 From the suboptimal comparison, the HFA-Net model is better than MTU-Net in terms of IoU and P d The increase was 2.35% and 3.97% respectively. a Reduced by 3.29×10 -6 , the detection performance has been greatly improved. This is because the HFA-Net model fully extracts and utilizes multi-level features. In the input stage of the network model, the local features of different scales are fused with images of different sizes output from the corresponding stage of the backbone network through the multi-level detail enhancement module (MLDEM) to generate feature maps of different scales, thereby enhancing the representation of detail information of different scales. In the feature extraction stage, multi-level long-distance information is extracted through the multi-level large kernel attention module (MLLKAM), so that the model learns the relationship between the target and the background at different scales. In the final prediction stage, the present invention can learn the influence of different scales on the probability likelihood map through the multi-level feature fusion module (MLFFM).

[0172] In addition, in order to further prove and analyze the influence of various modules introduced in the HFA-Net model on the ship segmentation performance, the present invention uses different parameter configurations to conduct experiments on the SDG-IRSTD dataset, and analyzes the influence of each module (MLDEM, MLLKAM, MLFFM) on the accuracy according to the results of different network configurations. The specific results are shown in Table 4.

[0173] Table 4

[0174] method <![CDATA[IoU(×10 -2 )]]> <![CDATA[P d (×10 -2 )]]> <![CDATA[F a (×10 -6 )]]> No MLDEM 64.16 92.06 7.63 No MLLKAM 63.98 93.97 11.77 No MLFFM 64.51 92.93 10.05 HFA-Net 65.61 94.31 7.25

[0175] As shown in Table 4, removing any module will reduce the overall performance of the HFA-Net model to varying degrees. Among them, removing MLLKAM has the greatest impact on the IoU indicator, causing the IoU of the HFA-Net model to drop by 1.63%. Other indicators also show varying degrees of decline, among which P d Reduced by 0.34%, F a Rise 4.52×10 -6This is because MLLKAM can extract multi-level long-distance information from remote sensing images and effectively utilize it, making up for the shortcomings of the convolutional neural network model that only focuses on local features, effectively suppressing the interference of the background on ship detection, and fully proving the importance of multi-scale long-distance information for infrared ship detection. Removing other modules will also affect the performance of HFA-Net to varying degrees, which further proves the reliability of MLDEM, MLLKAM and MLFFM in IRSTD tasks.

[0176] Through the above experiments, it can be known that the present invention combines all the bands of SDGSAT-1TIS to establish a three-channel infrared ship target detection dataset (SDG-IRSTD) with pixel-level annotations. Compared with other mainstream datasets, SDG-IRSTD has the highest imaging width, the smallest target size, the largest image size, the smallest average target-to-background ratio, the most complex observation scene, and the largest number of bands, and is more suitable for small ship target detection tasks in thermal infrared images. This dataset can fully extract thermal information from remote sensing images and provide assistance for deep learning networks. At the same time, the proposed HFA-Net generates feature maps of different scales through MLDEM, and then uses MLLKAM to extract long-distance information of multi-level features, and finally realizes the fusion and interaction of features of different scales through MLFFM. Compared with the MTU-Net method, HFA-Net has the highest IoU, P d The increase was 2.35% and 3.97% respectively. a Reduced by 3.29×10 -6 ,The experimental results verify the superiority of the method of the ,invention, which can simultaneously obtain the overall shape profile of the ,ship while achieving precise positioning of the target, providing a reliable ,method for sustainable development goals and subsequent research in the ,field of ocean management.

[0177] The above description is only an embodiment of the present invention and is not intended to limit the present invention. Any modification, equivalent replacement and improvement made within the application scope of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network, characterized in that: The following steps are involved: S1: Collect and preprocess the detection area image data of the three thermal infrared bands of SDGSAT-1 to establish a three-channel infrared ship target detection dataset with pixel-level annotations; S2: Construct a network model based on the hierarchical feature attention mechanism and use it to identify ships in thermal infrared images. The recognition process of this network model is divided into four stages: input stage, feature extraction stage, prediction stage and positioning stage, which include: S21: In the input stage, the obtained three-channel infrared ship target data set is subjected to multi-level detail enhancement processing to generate feature maps of different scales; S22: In the feature extraction stage, multi-level long-distance information is extracted from the multi-scale feature maps, and the long-distance information is fused into the feature maps of different scales to generate feature maps of different scales containing the long-distance information; S23: In the prediction stage, feature maps of different scales containing long-range information are fused to generate a predicted target probability map; S24: In the positioning stage, the generated predicted target probability map is used to locate the ship target using the eight-connected aggregation algorithm.

2. According to claim 1, a method for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network is characterized in that: For step S1, the detection area image data of the three thermal infrared bands of SDGSAT-1 are collected and preprocessed to establish a three-channel infrared ship target detection dataset with pixel-level annotations, including the following steps: S11: Collect thermal infrared image data of three thermal infrared bands from SDGSAT-1 simultaneously, and combine the three thermal infrared bands to form a three-channel thermal infrared image dataset; S12: performing correction processing on the collected three-channel thermal infrared image data set through radiometric calibration to generate a three-channel thermal infrared image data set with amplitude brightness values; S13: Use the SAM model to annotate the ship targets in the thermal infrared image dataset and obtain a three-channel infrared ship target detection dataset with pixel-level annotations. Each image in the dataset is a three-channel image composed of three thermal infrared images.

3. According to claim 1, a method for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network is characterized in that: For step S21, in the input stage, the obtained three-channel infrared ship target data set is subjected to multi-level detail enhancement processing to generate feature maps of different scales, specifically: S211: Construct a multi-level detail enhancement mode, with each level adopting a single-input dual-output processing structure; S212: In each different level, the first output path will not process the input image and directly send it to the next level; the second output path will first obtain images of different sizes by downsampling, and then extract features of the input images of different sizes by using different numbers of residual blocks to obtain local features of different scales; S213: local features of different scales are fused into images of different sizes, and feature maps of different scales are generated at different levels.

4. According to claim 3, a method for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network is characterized in that: The second output path adopts a dual-branch structure, each branch structure includes a different number of residual blocks, which are used to extract semantic information and detail information of the image respectively.

5. According to claim 1, a method for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network is characterized in that: For step S22, in the feature extraction stage, a multi-level encoder-large core attention mechanism-decoder structure mode is adopted to extract multi-level long-distance information from multi-scale feature maps, and then the long-distance information is fused into feature maps of different scales to generate feature maps of different scales containing long-distance information, specifically: S221: The multi-level encoder encodes feature maps of different scales to generate multi-level features; S222: The multi-level features output by the multi-level encoder are fused with the feature maps of different scales obtained in the input stage according to the corresponding levels to obtain a new multi-scale feature map; S223: Extract multi-level long-distance information from multi-scale feature maps through a multi-level large-core attention mechanism; S224: Multi-level long-distance information is aggregated into feature maps of different scales through feature concatenation operations, and then sent to multi-level decoders for decoding and restoration into feature maps of different scales, which contain long-distance information.

6. According to claim 5, a method for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network is characterized in that: For step S223, each level of the large core attention mechanism includes the following processing operations: S2231: Reconstruct the multi-scale feature maps through a convolution layer with a step size equal to the convolution kernel size to unify the dimensions of the feature maps. The calculation expression is: V i =Conv(E i ;s=ks=RF i ) Among them, V i Represents the feature map of the i-th scale of the reconstructed output; E i Represents the i-th scale feature map of the input, RF i Represents the reduction factor of the i-th level feature. The reduction factors used in feature maps of different scales are different. Conv represents standard convolution. s and ks represent the step size and convolution kernel size respectively. S2232: The reconstructed multi-scale feature map is first batch normalized, and then enters the large kernel attention mechanism, and performs point-by-point convolution, large kernel convolution, and point-by-point convolution again in sequence to extract multi-level long-distance information; among them, the large kernel convolution is decomposed into three parts, and the large kernel convolution LKA calculation expression is: Among them, I represents the input of the large kernel convolution; DW represents the depth convolution with a convolution kernel size of (2d-1)×(2d-1); DDW represents the convolution kernel size of PW represents the point-by-point convolution with a kernel size of 1×1; d is the convolution expansion rate; K is the preset conventional parameter; S2233: Transmit multi-level long-distance information to a multi-layer perceptron for information enhancement. The multi-layer perceptron includes 1x1 point convolution, 3x3 depth convolution, Gaussian error linear unit and 1x1 point convolution.

7. According to claim 1, a method for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network is characterized in that: For step S23, in the prediction stage, feature maps of different scales containing long-distance information are fused to generate a predicted target probability map, specifically: S231: Upsampling the feature maps of different scales to generate feature maps of different scales with unified dimensions; S232: Use the attention mechanism to assign different weights to feature maps of different scales, and perform element-by-element addition processing on the weighted feature maps to generate the final predicted target probability map. The weight α calculation formula is: β=Conv(Conv(GAP(D);ks=1);ks=1) γ=Conv(Conv(GMP(D);ks=1);ks=1) α=γ+β Among them, D represents feature maps of different scales; Conv represents standard convolution; ks=1 represents the convolution kernel size of 1x1; GAP represents global average pooling; GMP represents global maximum pooling; β and γ represent adaptive parameters for feature fusion.

8. According to claim 1, a method for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network is characterized in that: For step S24, in the positioning stage, the generated predicted target probability map is used to locate the ship target using the eight-connected aggregation algorithm, specifically: S241: Binarizing the predicted target probability map using a preset threshold to generate a binary map; S242: Using the eight-connected aggregation algorithm to count the pixel points contained in each ship target in the binary image, the ship target is positioned.

9. According to claim 1, a method for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network is characterized in that: On the basis of the determination of the ship target, the center of mass coordinates of the ship target are calculated by the weighted center of mass positioning algorithm. The weighted center of mass positioning algorithm is specifically as follows: first, based on the located ship target, the horizontal and vertical coordinate information of all the pixel points contained in it is obtained; then the pixel values ​​of all the pixel points of the ship target are normalized to obtain the weights; finally, the center of mass of the ship target is calculated according to the weights and the horizontal and vertical coordinate information.

10. A small ship target detection system in thermal infrared images based on a hierarchical feature attention network, characterized in that: A method for detecting small ship targets in thermal infrared images based on a hierarchical feature attention network as described in any one of claims 1 to 9, comprising a thermal infrared image data processing module and a ship target recognition module; The thermal infrared image data processing module pre-processes the image data of the three thermal infrared bands of SDGSAT-1 and outputs a three-channel infrared ship target detection dataset with pixel-level annotations; The ship target recognition module is used to build a network model based on the hierarchical feature attention mechanism to perform thermal infrared image ship target recognition. The ship target recognition module includes a multi-level detail enhancement module, a multi-level large core attention module, a multi-level feature fusion module and a ship target positioning module; The multi-level detail enhancement module is used in the input stage to generate feature maps of different scales by performing multi-level detail enhancement processing on the input three-channel infrared ship target data set; The multi-level large core attention module is used in the feature extraction stage. It uses the multi-level large core attention mechanism to extract multi-level long-distance information from the input multi-scale feature maps, and then fuses the long-distance information into feature maps of different scales to generate feature maps of different scales containing long-distance information. The multi-level feature fusion module is used in the prediction stage to generate a predicted target probability map by fusing the input feature maps of different scales containing long-range information; The ship target positioning module is used in the positioning stage to use the generated predicted target probability map, adopt the eight-connected aggregation algorithm to locate the ship target, and locate the center of mass coordinates of the ship through the weighted center of mass positioning algorithm.

Citation Information

Patent Citations

  • Remote sensing image high-quality automatic instance segmentation method based on SAM large model fine tuning

    CN118691815A

  • Ship target detection method, device and equipment and storage medium

    CN119559374A

  • Method for detecting infrared ship target based on improved yolov7

    US20250078541A1

  • Scene text detection method and apparatus, and storage medium

    WO2025044534A1