Infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement

By adopting synthetic fog and multi-level feature fusion and enhancement methods in infrared ship target detection, the problem of degradation of infrared ship image quality in fog environments is solved, and the detection accuracy and real-time performance are significantly improved.

CN120198653APending Publication Date: 2025-06-24CHINA THREE GORGES UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510524544.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the foggy environment, the image quality of infrared ships is degraded, resulting in blurred image, less obvious texture of ship target feature, and the infrared image itself has defects such as low resolution and low contrast, which affects the detection performance.

Method used

The infrared ship object detection method based on synthetic fog and multi-level feature fusion and enhancement is adopted. Through technical means such as data set preparation, adding small object detection layers, designing C2f-DynamicConv structure, multi-level feature fusion and enhancement module and lightweight shared convolution detection head, the YOLOv8n model is improved to improve detection accuracy.

Benefits of technology

It effectively improves the detection accuracy of small and medium-sized infrared ships in fog environments, reduces the number of parameters and calculations of the model, and improves the real-time and accuracy of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198653A_ABST
    Figure CN120198653A_ABST
Patent Text Reader

Abstract

An infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement comprises the following steps: S1, data set preparation: simulating an infrared ship target in a fog environment by using an atmospheric scattering model, and constructing a fog environment infrared ship data set; s2, adding a small target detection layer, and reconstructing a neck network to improve the shallow feature weight; s3, a C2f-DynamicConv structure is designed, and the C2f structure of the YOLOv8n backbone network is replaced by the C2f structure of the YOLOv8n backbone network; s4, designing a multi-level feature fusion and enhancement (MLFFAE) module, and adding the MLFFAE module between the feature extraction and the feature fusion of the YOLOv8n model, and carrying out the feature extraction and the feature fusion of the YOLOv8n model; s5, designing a lightweight shared convolution detection head to improve an original YOLOv8n detection head; and S6, training and testing the data set by using the improved IB-YOLO model. The method solves the problems that in the fog environment, the infrared ship imaging quality is reduced, the image is blurred, the ship target feature texture is not obvious, and the infrared ship image has the defects of low resolution, low contrast ratio and the like, so that the detection performance is poor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, in particular to an infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement. Background Art

[0002] Marine resources have brought marine economy to our country. With the continuous development of the marine economy, maritime traffic has become increasingly complex, and there are more and more ports and ships at sea. This not only increases the pressure of port ship supervision, but also affects the safety guarantee during ship navigation. Secondly, the Chinese sea area is a barrier to national security, and the intrusion of illegal ships from other countries poses a threat to our country's security and marine fishery. Implementing the recognition of image targets through computer intelligence is of great significance for improving the supervision efficiency of port ships, ensuring the safety of maritime ship navigation, and preventing the intrusion of illegal ships into the Chinese sea area.

[0003] Compared with visible light, infrared imaging technology has been widely used in military, industrial, medical, environmental protection, security and other fields due to its advantages of strong anti-interference ability, long detection distance, all-weather operation, good concealment, etc. These fields also mainly achieve their goals through the recognition of infrared imaging targets. Moreover, with the continuous maturity of infrared imaging technology, many researchers have applied infrared imaging technology to the detection of maritime ship targets. However, frequent foggy weather at sea will degrade the quality of infrared ship images, resulting in problems such as low image saturation, few edge detail texture features, and blurred imaging. In addition, infrared images themselves have defects such as low resolution and low contrast, which all bring difficulties to the detection of maritime infrared ship targets.

[0004] In the field of infrared ship image target detection algorithms, many researchers have proposed their own innovative single-stage target detection algorithms. Liu Fen et al. based on YOLOv5 architecture, used K-means++ to make the anchor boxes better match the ship targets in the dataset, and proposed the MioU regression loss function to improve the ship detection accuracy in complex marine backgrounds and avoid missing detections. Zhang Shen et al. improved the Backbone network and Neck network in the YOLOv7 algorithm network structure through the MobileNetv3 convolutional neural network and the SE attention mechanism respectively, and introduced the WiseIoU loss function, reducing the model's parameter quantity and computational complexity and improving the ship detection accuracy. Yang Shi et al. redesigned the backbone network of the YOLOv5 algorithm, which consists of four multi-scale residual blocks and is connected by the CBAM attention mechanism, and then added shallow feature fusion in the FPN structure of the model, enabling the model to achieve good performance in the accuracy of ship detection. The above research on infrared ship image target detection algorithms has improved both the detection performance of ships and the lightweight of the detection model. However, affected by the frequent thick fog environment at sea, the quality of infrared images deteriorates, resulting in a decline in the accuracy and robustness of infrared ship detection models. To further improve the ship target detection ability in foggy environments, some researchers combined dehazing algorithms with detection models to achieve target detection. However, after using the dehazing algorithm, it is necessary to evaluate the quality of the dehazed image to avoid affecting the detection effect, and the method of combining dehazing with detection will affect the real-time performance of detection. Therefore, a single-stage target detection algorithm can be directly used to achieve infrared ship image target detection in the thick fog environment at sea. Such algorithms can maintain high detection speed while keeping the accuracy, and are more suitable for real-time monitoring of ship targets at sea. In response to this, there are also many related studies. Ma Haowei et al. improved the backbone network and multi-scale feature fusion network of YOLOv5 through the Swin Transformer and CA attention mechanism respectively, and introduced the CIoU loss function, effectively improving the infrared ship detection accuracy in foggy environments. Deng et al. first used the synthetic fog algorithm to construct an infrared ship dataset in a foggy environment, and then improved YOLOv7 by designing a spatial and channel hybrid attention mechanism and a weighted feature fusion pyramid network based on dilated convolution, effectively reducing the impact of the fog environment on the detection of small and medium-sized targets in infrared ships and improving the overall performance of infrared ship detection.

[0005] In summary, the research on infrared ship detection in fog environment effectively improves the detection accuracy of ships. However, the existing methods still have the problem of high model complexity. Moreover, in the frequent heavy fog environment at sea, the infrared ship images will be blurred, and the spatial and semantic information of the infrared ship target features extracted by the target detection algorithm will become weak, affecting the detection performance. In addition, small targets at a long distance are submerged in the sensor device noise and the complex background at sea, and dense targets overlap and occlude, which will make it difficult to extract target features and affect the accuracy of target positioning and classification. The existing research on these problems is insufficient. Therefore, when dealing with infrared ship target detection in fog environment, more appropriate algorithms need to be designed to meet the high precision and high efficiency of detection. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide an infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement, which solves the problem that the imaging quality of infrared ships in fog environment decreases, makes the image blurred, the texture of ship target features is not obvious, and the infrared ship images themselves have defects such as low resolution and low contrast, which will lead to poor detection performance.

[0007] To solve the above technical problems, the technical solution adopted by the present invention is: an infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement, including the following steps: S1. Dataset preparation: Use the atmospheric scattering model to simulate the infrared ship target in the fog environment, and construct an infrared ship dataset in the fog environment; S2. Add a small target detection layer and reconstruct the neck network to improve the weight of shallow features; S3. Design a C2f-DynamicConv structure to replace the C2f structure of the YOLOv8n backbone network; S4. Design a multi-level feature fusion and enhancement (Multi level feature fusion and enhancement module, MLFFAE) module and add it between the feature extraction and feature fusion of the YOLOv8n model; S5. Design a lightweight shared convolution detection head to improve the original YOLOv8n detection head; S6. Use the improved IB-YOLO model to train and test the dataset.

[0008] Preferably, in step S1, the publicly available maritime infrared ship dataset is used as a sample, and the infrared ship imaging in the fog environment is simulated through the atmospheric scattering model to construct an infrared ship dataset in the fog environment for subsequent training and testing of the model. The specific steps are as follows: S101. Collect infrared ship images, including 7 types of ships: sailing ships, fishing boats, cruise ships, boat-type ships, bulk carriers, warships, and container ships. S102. Use the atmospheric scattering model to simulate the imaging of infrared ships in a haze environment and construct an infrared ship dataset in a fog environment. The formula is as follows: ; ; Where, is the original image; is the image with fog; is the atmospheric light value; is the transmittance. In this formula, is the atmospheric scattering coefficient, and the magnitude of its value is positively correlated with the fog concentration, which is used to measure the fog concentration; is the distance from the pixel point of the image to the fog center, and its calculation formula is: ; Where, represents the Euclidean distance from the current pixel point in the image to the center pixel point; and represent the number of pixel rows in the vertical direction and the number of pixel columns in the horizontal direction respectively; S103. Divide the constructed infrared ship dataset in a fog environment into a training set, a validation set, and a test set according to the ratio of 7:2:1 for model training and testing.

[0009] Preferably, step S2 is to improve the multi-scale features of the original YOLOv8n model. The original YOLOv8n model has undergone five downsampling operations in the backbone network part, namely 2×, 4×, 8×, 16×, and 32×. Through these five downsampling operations, five different-scale feature maps are obtained respectively, and then the feature maps with sizes of 80×80, 40×40, and 20×20 are used to detect small, medium, and large targets in the detection target; the improved algorithm adds a small target detection layer of 160×160 according to the actual scene of infrared ships at sea during navigation and changes the multi-scale feature extraction and fusion relationship.

[0010] Preferably, step S3 includes the following steps: S301: Introduce DynamicConv to replace the second convolution in the Bottleneck part of C2f to form the Bottleneck_DynamicConv module; DynamicConv draws on the channel attention mechanism SE to generate attention weights, then aggregates features through the attention weights, and finally uses batch normalization (BN) and activation function (ReLu) for the aggregated features to construct a dynamic convolution layer; S302: Replace the Bottleneck part in the backbone network C2f with Bottleneck_DynamicConv, effectively improving the feature extraction ability of the backbone network and the detection accuracy.

[0011] Preferably, the step S4 includes the following steps: S401. Design a multi-level feature fusion module (Multi level feature fusion module, MLFF) to fully fuse the information of the shallow and deep layers of the backbone network; S402. Then, through a multi-scale channel attention mechanism module (Multi scale channel attention module, MSCA), better learn the deep semantic information and shallow spatial information after fusion, and enhance the expression ability of the features after fusion; S403. Design a multi-level feature fusion and enhancement module.

[0012] Preferably, the design method of the multi-level feature fusion module is as follows: First, for the shallow features Use the ADown module for downsampling to obtain the feature map , making the size of the feature map consistent with . The ADown module adopts a dual-branch structure combined with average pooling and max pooling, which can obtain global features and local significant features, and reduce the information loss caused by downsampling; Secondly, for the deep features Use the Upsample module for downsampling to obtain the feature map , making its spatial dimension consistent with ; Then, splice the three features , and in the channel dimension; Finally, realize the fusion of the spliced features through a convolution with a convolution kernel size of 1×1 to obtain the final output feature , and this process can be expressed by the following formula: .

[0013] Preferably, the specific design steps of the multi-scale channel attention mechanism module are as follows: Parallelly learn the feature information output by the multi-level feature fusion module through depth convolutions with four different convolution kernel sizes , use to represent the convolution with a convolution kernel size of , where , and the features extracted by different depth convolutions are , where the expression is: ; Depth convolutions with different convolutional kernel sizes obtain multi-scale feature information through receptive fields of different sizes; then, global average pooling operations are performed on the features obtained from different receptive fields to compress spatial information and obtain channel description vectors; Denote the channel description vectors corresponding to different receptive fields, which correspond to the features one by one, , and can be expressed as: ; The obtained channel description vectors are learned for channel weights through one-dimensional convolution. Compared with the fully connected layer, one-dimensional convolution can reduce the number of parameters and avoid the negative impact caused by channel dimensionality reduction. Use Conv1D to represent one-dimensional convolution. Then use the Sigmoid activation function to normalize the weights of each channel to obtain , which corresponds to the features , vector one by one, , and its expression is: ; Finally, multiply the features by the corresponding channel weights respectively, and fuse the obtained features by element-wise addition to obtain different receptive field features containing rich information, and its expression is: .

[0014] Preferably, the design steps of the multi-level feature fusion and enhancement module are as follows: the features obtained after feature fusion and the features obtained through multi-scale channel attention are integrated through residual connection and convolution with a convolutional kernel size of 1×1 to obtain the final output , and the specific expression is as follows: .

[0015] Preferably, in step S5, the two tasks share a convolution, and then use the GroupNormalization normalization technique to reduce the accuracy loss caused by parameter sharing.

[0016] Preferably, step S6 includes the following steps: S601: The data set is put into the IB-YOLO model for training according to the training strategy. The training strategy is to set the image size of the input network training to 640×640, the training round is 200 epochs, the batch size is set to 16, the loss function optimizer is the stochastic gradient optimizer (SGD), and the weight decay, momentum and initial learning rate used by the model are set to 0.937, 0.0005 and 0.01 respectively; finally, in order to make the data set richer, the weights of the Mosaic and MixUp data enhancements that come with the model are set to 1.0 and 0.15 respectively; S602: After training, the IB-YOLO model is evaluated using the test set. In order to test the performance of the improved model in this paper, the experiment uses precision (Precision, P), recall (Recall, R), mean average detection accuracy (meanAveragePrecision, mAP), average precision (Average Precision, AP), model parameters (Params) and gigaflops per second (GFLOPs) as evaluation indicators of the model.

[0017] The present invention provides an infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement, improves the backbone network of the initial YOLOv8n model, adds a small target detection layer, and mainly increases the utilization of 4× downsampling, and improves the neck network at the same time, and this optimization improves the detection accuracy of small targets in infrared ships in foggy environment. In addition, the present invention designs a C2f-DynamicConv structure, replaces the C2f structure of the YOLOv8n backbone network, and this optimization significantly improves the feature extraction capability of the model backbone network, thereby improving the detection accuracy. Then, a multi-level feature fusion and enhancement (Multi level feature fusion and enhancement module, MLFFAE) module is designed and added between the feature extraction and feature fusion of the YOLOv8n model, effectively improving the utilization rate of shallow and deep feature information in the backbone network, and enhancing the expression ability of the fused features, and this optimization effectively improves the detection accuracy of the model. Finally, a lightweight shared convolution detection head is designed to effectively reduce the calculation amount and parameter amount of the IB-YOLO model, and improve the accuracy by a small amount. The IB-YOLO algorithm proposed in the present invention effectively improves the detection accuracy of infrared ship targets in foggy environments, and has lower parameter and calculation complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The present invention will be further described below in conjunction with the accompanying drawings and embodiments: Figure 1 is a flow chart of the method of the present invention; Figure 2 Schematic diagram of the algorithm structure of the present invention; Figure 3 Structural diagram of the C2f-DynamicConv module designed by the present invention; Figure 4 Structural diagram of the MLFF module designed by the present invention; Figure 5 Structural diagram of the MSCA module designed by the present invention; Figure 6 Structural diagram of the MLFFAE module designed by the present invention; Figure 7 Synthetic fog images of different concentrations of the dataset of the present invention; Figure 8 Structural diagram of the LSCD module designed by the present invention; Figure 9 Effect diagram of target detection before the algorithm improvement of the present invention; Figure 10 Effect diagram of target detection after the algorithm improvement of the present invention. Detailed implementation manners

[0019] The present invention designs an infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement, which is implemented by Python. Taking YOLOv8n as the benchmark model, the overall network architecture of the improved IB-YOLO is as Figure 2 shown. To realize the detection of infrared ship targets in a fog environment by means of deep learning, sufficient relevant data is required. The dataset of the present invention is screened and processed from the dataset provided by the infiRay company, and this dataset is used as the dataset for training the YOLOv8n model. After improving the YOLOv8n model and then training, the weights of the improved IB-YOLO are obtained. The improved weights are used to test the detection effect.

[0020] As Figure 1 shown, an infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement includes the following steps: S1. Dataset preparation: Use the atmospheric scattering model to simulate infrared ship targets in a fog environment and construct an infrared ship dataset in a fog environment; Further, the process of S1 is that the present invention takes the publicly available maritime infrared ship dataset as a sample, uses the atmospheric scattering model to simulate the imaging of red ships in a fog environment, and constructs an infrared ship dataset in a fog environment for subsequent model training and testing. Specifically, there are the following sub-steps: S101: The dataset samples of the present invention are from the open-source maritime infrared ship dataset of infiRay Company. This dataset has 8,402 infrared ship images, including 7 types of ships: sailboats, fishing boats, cruise ships, boat-type ships, bulk carriers, warships, and container ships.

[0021] S102: The present invention uses the open-source maritime infrared ship dataset of infiRay Company as a sample, and uses the atmospheric scattering model to simulate the imaging of infrared ships in a haze environment to construct an infrared ship dataset in a fog environment. The formula is as follows: ; ; Where, is the original image; is the image with fog; is the atmospheric light value; is the transmittance. In this formula, is the atmospheric scattering coefficient, and the value is positively correlated with the fog concentration, which is used to measure the fog concentration; is the distance from the pixel point of the image to the fog center, and its calculation formula is: ; Where, represents the Euclidean distance from the current pixel point in the image to the center pixel point; and represent the number of pixel rows in the vertical direction and the number of pixel rows in the horizontal direction respectively. By adjusting different atmospheric scattering coefficients and atmospheric light values , different fog concentration infrared ship datasets in a fog environment are constructed. In this paper, is taken as 0.02 - 0.08, and fog images with different concentrations are shown in Figure 7 .

[0022] S103: The constructed infrared ship dataset in a fog environment is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1 for model training and testing.

[0023] S2. Add a small target detection layer and reconstruct the neck network to improve the weight of shallow features; Furthermore, the process of S2 is to improve the multi-scale features of the original YOLOv8n model. The original YOLOv8n model has gone through five downsampling operations in the backbone network part, which are 2×, 4×, 8×, 16×, and 32× respectively. Through these five downsampling operations, five feature maps with different scale sizes are obtained, and then the feature maps with sizes of 80×80, 40×40, and 20×20 are used to detect small, medium, and large targets in the detected objects. The improved algorithm adds a small target detection layer of 160×160 according to the actual scene of maritime infrared ships during navigation, and changes the multi-scale feature extraction and fusion relationship.

[0024] S3. Design the C2f-DynamicConv structure to replace the C2f structure of the YOLOv8n backbone network. The C2f-DynamicConv structure diagram is as Figure 3 shown; S301: Introduce DynamicConv to replace the second convolution in the Bottleneck part of C2f to form the Bottleneck_DynamicConv module. DynamicConv draws on the channel attention mechanism SE to generate attention weights, then aggregates features through the attention weights, and finally uses batch normalization (BN) and activation function (ReLu) for the aggregated features to construct a dynamic convolution layer.

[0025] S302: Replace the Bottleneck part in the backbone network C2f with Bottleneck_DynamicConv to effectively improve the feature extraction ability of the backbone network and improve the detection accuracy.

[0026] S4. Design the Multi level feature fusion and enhancement (MLFFAE) module and add it between the feature extraction and feature fusion of the YOLOv8n model. The MLFFAE module structure diagram is as Figure 6 shown; Furthermore, the process of S4 is to design the Multi level feature fusion and enhancement (MLFFAE) module and add it between the feature extraction and feature fusion of the YOLOv8n model. The specific sub-steps are as follows: S401: First, design the Multi level feature fusion module (MLFF) to fully fuse the information of the shallow and deep layers of the backbone network part. The MLFF structure diagram is as Figure 4As shown below, the specific process of MLFF is as follows: First, for the shallow features downsampling is performed using the ADown module to obtain the feature map , making the size of the feature map consistent with . The ADown module adopts a double-branch structure combining average pooling and max pooling, which can obtain global features and local significant features, reducing information loss caused by downsampling. Second, for the deep features downsampling is performed using the Upsample module to obtain the feature map , making its spatial dimension consistent with . Then, the three features , and are concatenated in the channel dimension. Finally, the concatenated features are fused through a convolution with a kernel size of 1×1 to obtain the final output feature , and this process can be expressed by the following formula: ; S402: Then, through a multi-scale channel attention mechanism (Multi scale channel attention module, MSCA), the deep semantic information and shallow spatial information after fusion are better learned, enhancing the expression ability of the fused features. The structure of MSCA is as Figure 5 shown. The specific details of the MSCA module are as follows: The feature information output by the multi-level feature fusion module is learned in parallel through depthwise convolutions with four different kernel sizes , and is used to represent the convolution with a kernel size of , where , and the features extracted by different depthwise convolutions are , where the expression is: ; Depthwise convolutions with different kernel sizes obtain multi-scale feature information through different receptive fields. Then, global average pooling operations are performed on the features obtained from different receptive fields to compress the spatial information, obtaining channel description vectors. represents the channel description vectors corresponding to different receptive fields, which correspond one-to-one with the features , , and can be expressed as: ; The obtained channel description vectors are learned for channel weights through the one-dimensional convolution adopted in this paper. Compared with the fully connected layer, the one-dimensional convolution can reduce the number of parameters while avoiding the negative impact caused by channel dimensionality reduction. Conv1D is used to represent the one-dimensional convolution. Then, the Sigmoid activation function is used to normalize the weights of each channel to obtain , which corresponds one-to-one with the feature , vector , one by one, , and its expression is: ; Finally, multiply the feature by the corresponding channel weights respectively, and fuse the obtained features by element-wise addition to obtain different receptive field features containing rich information , and its expression is: ; S403: Combining the advantages of the proposed multi-level feature fusion module and multi-scale channel attention module, a multi-level feature fusion and enhancement module MLFFAE is designed. The structure diagram is as shown in Figure 6 . The specific process is as follows: The feature obtained after feature fusion and the feature obtained through multi-scale channel attention are integrated through residual connection and convolution with a convolution kernel size of 1×1 to obtain the final output , and the specific expression is as follows: ; S5. Design a lightweight shared convolutional detection head to improve the detection head of the original YOLOv8n. The structure diagram is as shown in Figure 8 .

[0027] Furthermore, the process of S5 is to design a lightweight shared convolutional detection head (lightweight shared convolutional detection head, LSCD) to improve the detection head of the original YOLOv8n. In the original detection head, the classification and regression tasks exist independently, and each task needs to go through two convolutions with a convolution kernel size of , which will result in a large number of parameters at the detection head. By designing a shared convolutional detection head, the two tasks share a convolution, and then using the Group Normalization normalization technique to reduce the accuracy loss caused by parameter sharing.

[0028] S6. Use the improved IB-YOLO model to train and test the dataset. The comparison results of different algorithms are shown in the following figure.

[0029] Furthermore, the process of S6 is to train and test the infrared ship dataset in fog environment using the improved IB-YOLO model. The specific sub-steps are as follows: S601: The data set is put into the IB-YOLO model for training according to a certain training strategy. The training strategy of the present invention is mainly to set the image size of the input network training to 640×640, the training round is 200 epochs, the BatchSize size is set to 16, the loss function optimizer is the stochastic gradient optimizer (SGD), and the weight decay, momentum and initial learning rate used by the model are set to 0.937, 0.0005 and 0.01 respectively. Finally, in order to make the data set richer, the weights of the Mosaic and MixUp data enhancements that come with the model are set to 1.0 and 0.15 respectively.

[0030] S602: After training, the IB-YOLO model is evaluated using the test set. In order to test the performance of the improved model in this paper, the experiment uses precision (Precision, P), recall (Recall, R), mean average detection accuracy (meanAveragePrecision, mAP), model parameters (Params), and computational complexity (GFLOPs).

[0031]

[0032] The above table compares the comprehensive performance of different algorithms in the infrared ship dataset in foggy environment, comparing the indicators such as precision (Precision, P), recall (Recall, R), mean average precision (mean Average Precision, mAP), model parameters (Params), and computational complexity (GFLOPs). From the results in the table, it can be seen that the IB-YOLO algorithm achieves the best in P, R, and mAP, and the computational complexity and parameter quantity are in a relatively low range, achieving a balance between model performance and complexity.

[0033] like Figure 9 , Figure 10 As shown in Figure 2, in order to intuitively reflect the detection effect of the algorithm proposed in this paper, we selected multi-scale targets with different fog concentrations, complex background targets, occluded targets, and long-distance dense small targets, which are difficult to detect, for detection comparison. The detection comparison results are shown in Figure 2. Figure 9 , 10As shown in the figure, (a) shows the comparison results of multi-scale targets with different fog concentrations. Among them, for larger-scale targets, the accuracy slightly decreases as the fog concentration increases. The IB-YOLO algorithm of the present invention can effectively improve the detection accuracy of larger targets in a fog environment, and can also detect tiny targets, reducing the missed detection of targets. (b) shows the comparison of targets under complex backgrounds. Due to the high similarity between infrared ship targets and background noise under complex backgrounds, the contrast between the two is low, and the noise caused by the fog environment weakens the feature expression of infrared ship targets, resulting in low detection accuracy of targets by YOLOv8n and missed detection of infrared ship targets. The algorithm of the present invention can correctly detect targets and improve the target detection accuracy. (c) shows the comparison of the detection results of occluded targets. It can be seen from the comparison that the model of the present invention effectively improves the detection accuracy of occluded targets. Moreover, in the figure, YOLOv8n wrongly identifies the target in the shore background as a container ship, and the algorithm of the present invention effectively reduces the occurrence of this misdetection. (d) and (c) are both comparisons of the detection results of distant and dense small targets. It can be clearly seen in (c) that the IB-YOLO algorithm of the present invention has better performance in detecting small targets, can detect more distant infrared ship small targets, and effectively reduces the missed detection rate of small targets. In (d), the detection accuracy of a few targets decreases, but the overall detection effect is better. Therefore, compared with YOLOv8n, the IB-YOLO algorithm of the present invention can significantly improve the detection performance of infrared ships in a fog environment, has better performance in detecting multi-scale targets, complex backgrounds, target occlusion, distant and dense small targets and other scenarios, and can further reduce the missed detection and misdetection of targets.

[0034] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as limitations on the present invention. The protection scope of the present invention should be the technical solutions recorded in the claims, including equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, equivalent replacement improvements within this scope are also within the protection scope of the present invention.

Claims

1. An infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement, characterized in that: The following steps are involved: S1. Dataset preparation: Using the atmospheric scattering model to simulate infrared ship targets in foggy environment, a foggy environment infrared ship dataset is constructed. S2, add a small target detection layer, reconstruct the neck network and improve the weight of shallow features; S3. Design the C2f-DynamicConv structure to replace the C2f structure of the YOLOv8n backbone network. S4. Design a multi-level feature fusion and enhancement module and add it to the feature extraction and feature fusion of the YOLOv8n model; S5. Design a lightweight shared convolutional detection head to improve the original YOLOv8n detection head; S6. Use the improved IB-YOLO model to train and test the dataset.

2. The infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement according to claim 1 is characterized in that: The step S1 uses a public marine infrared ship dataset as a sample, simulates the infrared ship imaging in a foggy environment through an atmospheric scattering model, and constructs a foggy environment infrared ship dataset for subsequent model training and testing. The specific steps are as follows: S101, collect infrared ship images, including sailboats, fishing boats, cruise ships, boat-type ships, bulk carriers, warships and container ships, a total of 7 types of ships; S102. Use the atmospheric scattering model to simulate the imaging of infrared ships in a haze environment and construct a foggy environment infrared ship dataset. The formula is as follows: ; ; in, is the original image; For images containing fog; is the atmospheric light value; is the transmittance, in which is the atmospheric scattering coefficient, the value of which is positively correlated with the fog concentration and is used to measure the fog concentration; is the distance from the pixel of the image to the center of the fog, and its calculation formula is: ; in, Represents the Euclidean distance from the current pixel to the center pixel in the image; and Respectively represent the number of pixel rows in the vertical direction and the number of pixel rows in the horizontal direction; S103, dividing the constructed foggy environment infrared ship dataset into a training set, a validation set, and a test set in a ratio of 7:2:1 for model training and testing.

3. The infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement according to claim 1 is characterized in that: The step S2 is to improve the multi-scale features of the original YOLOv8n model. The original YOLOv8n model has undergone five downsampling operations in the backbone network part, which are 2×, 4×, 8×, 16×, and 32× respectively. Five feature maps of different scales are obtained through these five downsampling operations, and then the feature maps of sizes 80×80, 40×40 and 20×20 are used to detect small, medium and large targets in the target. The improved algorithm adds a 160×160 small target detection layer according to the actual scene of infrared ships at sea during navigation, and changes the multi-scale feature extraction and fusion relationship.

4. The infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement according to claim 1 is characterized in that: The step S3 comprises the following steps: S301: Introduce DynamicConv to replace the second convolution of the Bottleneck part in C2f to form the Bottleneck_DynamicConv module; DynamicConv uses the channel attention mechanism SE to generate attention weights, then aggregates features through the attention weights, and finally uses batch normalization and activation functions to construct a dynamic convolution layer for the aggregated features; S302: Replace the Bottleneck part in the backbone network C2f with Bottleneck_DynamicConv, effectively improving the feature extraction capability of the backbone network and improving the detection accuracy.

5. The infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement according to claim 1 is characterized in that: The step S4 comprises the following steps: S401. Design a multi-level feature fusion module to fully integrate the shallow and deep information of the backbone network; S402, a multi-scale channel attention mechanism module is then used to better learn the fused deep semantic information and shallow spatial information, thereby enhancing the expression capability of the fused features; S403. Design a multi-level feature fusion and enhancement module.

6. The infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement according to claim 5 is characterized in that: The design method of the multi-level feature fusion module is as follows: First, shallow features Use the ADown module to downsample and get the feature map , so that the feature map size is The ADown module adopts a dual-branch structure combined with average pooling and maximum pooling, which can obtain global features and local significant features and reduce the information loss caused by downsampling; secondly, deep features Use the Upsample module to downsample and get the feature map , so that it The spatial dimensions of , and Perform splicing of channel dimensions; finally, the spliced ​​features are fused through convolution with a convolution kernel size of 1×1 to obtain the final output features ,This process can be expressed by the following formula: 。 7. The infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement according to claim 5 is characterized in that: The specific design steps of the multi-scale channel attention mechanism module are as follows: The feature information output by the multi-level feature fusion module is learned in parallel through four deep convolutions with different convolution kernel sizes. ,use Indicates that the convolution kernel size is The convolution of , the features extracted by convolution with different depths are , where the expression is: ; Deep convolution with different convolution kernel sizes obtains multi-scale feature information through receptive fields of different sizes. Then, a global average pooling operation is performed on the features obtained from different receptive fields to compress the spatial information and obtain a channel description vector. Represents the channel description vector corresponding to different receptive fields, which is related to the feature One to one correspondence, , which can be expressed as: ; The obtained channel description vector is subjected to one-dimensional convolution to learn the channel weight. Compared with the fully connected layer, the one-dimensional convolution can reduce the number of parameters while avoiding the negative impact caused by channel dimensionality reduction. Conv1D represents the one-dimensional convolution. Then the Sigmoid activation function is used to normalize the weights of each channel to obtain , and its characteristics ,vector One to one correspondence, , whose expression is: ; Finally, the features Multiply by the corresponding channel weights , and fuse the obtained features by element-wise addition to obtain different receptive field features containing rich information , whose expression is: 。 8. The infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement according to claim 5 is characterized in that: The design steps of the multi-level feature fusion and enhancement module are as follows: And the features obtained by multi-scale channel attention Through residual connection and convolution with kernel size of 1×1 Integrate to get the final output , the specific expression is as follows: 。 9. The infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement according to claim 1 is characterized in that: In step S5, two tasks share one The convolution is performed, and then the GroupNormalization normalization technique is used to reduce the accuracy loss caused by parameter sharing.

10. The infrared ship target detection method based on synthetic fog and multi-level feature fusion and enhancement according to claim 1, characterized in that: The step S6 comprises the following steps: S601: The data set is put into the IB-YOLO model for training according to the training strategy. The training strategy is to set the image size of the input network training to 640×640, the training round is 200 epochs, the batch size is set to 16, the loss function optimizer is the stochastic gradient optimizer, and the weight decay, momentum and initial learning rate used by the model are set to 0.937, 0.0005 and 0.01 respectively; finally, in order to make the data set richer, the weights of the Mosaic and MixUp data enhancements that come with the model are set to 1.0 and 0.15 respectively; S602: After training, the IB-YOLO model is evaluated using the test set. In order to test the performance of the improved model in this paper, the experiment uses accuracy, recall rate, average detection accuracy mean, average precision, model parameter number and Giga floating point operations per second as evaluation indicators of the model.

Citation Information

Cited By

  • Infrared ship detection method for complex marine environment based on bionic hierarchical feature optimization

    CN121354077A

  • Real-time anti-interference ship detection method and system for low-visibility environment

    CN121686251A

  • Double-branch ship classification and identification method and system fused with contrast learning

    CN122135073A