Infrared ship image detection method based on improved YOLOV7
By improving the multi-scale feature processing, anchor frame adaptation and multi-scale residual cavity convolution module design of YOLOV7 model, the problem of poor detection performance of small and dense targets in infrared ship images is solved, and the lightweight and efficient detection performance of the model is achieved.
Patent Information
- Application Number
- CN202510524541.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-06-27
AI Technical Summary
The existing infrared ship image detection algorithms have poor detection performance and high model complexity when detecting scenarios containing many small targets and dense targets, so they are not suitable for deployment in hardware devices with limited resources.
Improve the YOLOV7 model, optimize the model structure and training process through multi-scale feature improvement, anchor box reclustering, designing multi-scale residual void convolution module (RDC) and using MPDIoU loss function.
It significantly improves the detection accuracy and performance of the model, reduces the model size and parameter volume, and reduces the calculation volume, making it suitable for deployment in hardware devices with limited resources.
Smart Images

Figure CN120219723A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and particularly to an infrared ship image detection method for improving YOLOV7. Background Art
[0002] Marine resources have brought marine economy to our country. With the continuous development of the marine economy, maritime traffic has become increasingly complex, and there are more and more ports and ships at sea. This not only increases the pressure of port ship supervision, but also affects the safety guarantee during ship navigation. Secondly, the Chinese sea area is a barrier to national security, and the intrusion of illegal ships from other countries poses a threat to our country's security and marine fishery. Implementing the recognition of image targets through computer intelligence is of great significance for improving the supervision efficiency of port ships, ensuring the safety of maritime ship navigation, and preventing the intrusion of illegal ships into the Chinese sea area.
[0003] In recent years, the rise of artificial intelligence has strengthened the field of computer vision. In the field of infrared image target detection, infrared target detection algorithms based on deep learning have emerged continuously. The target detection algorithms based on deep learning are mainly divided into two-stage algorithms (RCNN, fast RCNN, faster RCNN) and one-stage algorithms (SSD, YOLO series). For the target detection of infrared ship images, the two-stage algorithms first generate candidate regions and then perform regression and classification of the target's bounding box through a convolutional neural network. However, they have problems such as large computational amount, complex training, and difficulty in optimization. The one-stage target detection algorithm directly generates class predictions and bounding box regression from the input image, which greatly improves the efficiency of target detection. Therefore, it is now also applied to the related research of infrared ship image target detection.
[0004] In the field of infrared ship image target detection algorithms, many researchers have proposed their own innovative one-stage target detection algorithms. Miao Chuankai et al. first used ResNet50 as the backbone network to extract infrared ship features, then introduced dilated convolution to further process the extracted feature maps, and finally used the anchor-free CenterNet algorithm for classification and regression, improving the detection accuracy of ships. Liu Fen et al. based on YOLOv5, used K-means++ to make the anchor boxes more matched with the ship targets in the dataset, and proposed the MioU regression loss function to improve the ship detection accuracy in complex marine backgrounds and avoid missing detections of targets. Zhang Shen et al. improved the Backbone network and Neck network in the YOLOv7 algorithm network structure through the MobileNetv3 convolutional neural network and the SE attention mechanism respectively, and introduced the WiseIoU loss function, reducing the number of model parameters and computational complexity and improving the ship detection accuracy. Yang Shi et al. redesigned the backbone network of the YOLOv5 algorithm, which consists of four multi-scale residual blocks and is connected by the CBAM attention mechanism, and then added shallow feature fusion to the FPN structure of the model, enabling the model to achieve good performance in the accuracy of ship detection. The above research on infrared ship image target detection algorithms has improved the detection accuracy of ships and the lightweight of the detection model. However, the research on the problem that the detection performance is poor when there are many small targets and dense targets in the detection targets is insufficient, and there is still the problem of high model complexity. These problems are not conducive to deploying the model on hardware devices with limited resources and achieving high detection performance for infrared ship images. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide an improved infrared ship image detection method for YOLOV7 to solve the problem that the detection performance is poor when there are many small targets and dense targets in ships at a relatively long distance at sea.
[0006] To solve the above technical problem, the technical solution adopted by the present invention is: an improved infrared ship image detection method for YOLOV7, including the following steps: S1. Dataset preparation: Divide the dataset into a training set and a validation set according to a ratio of 8:2, and take another part of the data as a test set to verify the generalization of the improved model; specifically, 400 images.
[0007] S2. Perform multi-scale feature improvement on the overall YOLOv7 network; S3. Re-cluster the preset anchor box parameters; S4. Design a multi-scale residual dilated convolution module (RDC) and add it after the 1×1 convolution; S5. Use the MPDIoU loss function as the bounding box regression loss function; S6. Use the improved YOLOv7 model to train and test the dataset.
[0008] Preferably, the dataset selected in step S1 includes images of fishing boats, boat-type ships, and sailboats.
[0009] Preferably, step S2 is specifically as follows: The YOLOv7 model undergoes five downsampling operations in the backbone network part, namely 2×, 4×, 8×, 16×, and 32×. Through these five downsampling operations, five different scales of feature maps are obtained respectively. Use the feature maps with sizes of 80×80, 40×40, and 20×20 to correspond to small, medium, and large targets in the detection targets; After improvement, according to the actual scene of marine infrared ships during navigation, the 32× downsampling is removed, that is, the output of the 20×20 scale is removed, the output feature size is adjusted to 80×80 and 40×40, and the feature fusion network is redesigned.
[0010] Preferably, step S3 includes the following steps: S301. Evaluate the initial preset anchor boxes. When the BPR (Best Possible Recall) between the model's initial preset anchor boxes and the targets in the ship dataset is less than that of the original model, it means that the preset anchor box size does not match the dataset used in the experiment; The BPR value of the original model is selected as 0.98, which is obtained by analyzing the model's original anchor boxes through k-means clustering and genetic algorithms.
[0011] S302. The anchor box adaptive mechanism of YOLOv7 uses k-means clustering and genetic algorithms to analyze and obtain the optimal anchor box size; S303. According to the improvement in step S2, use k-means clustering and genetic algorithms to analyze and obtain two different sizes of anchor boxes corresponding to medium and small receptive fields respectively. At this time, the BPR value is 0.993, which is better than the default anchor box size. 0.993 is obtained by analyzing the improved anchor boxes through k-means clustering and genetic algorithms.
[0012] Preferably, step S4 includes the following steps: S401. The residual dilated convolution module represents the given input feature map as , means dividing the input feature map into four parts according to the number of channels to form four feature maps; S402. Process the four input features in parallel using different convolutions. The different convolutions are a 1×1 convolution and three dilated convolutions with a kernel size of 3×3 and different dilation rates. Represents the dilation rate of the dilated convolution. The three values among them are 1, 2, and 3 respectively. Represents the input feature After passing through a dilated convolution operation with a kernel size of and a dilation rate of ; S403. Concatenate (concat) the processed features in step S402 to obtain a new output, and then achieve channel feature fusion through a 1×1 convolution; Represent the further processed feature map as Z, and feature Z is represented by the following formula: ; S404: According to the residual connection method, perform an element-wise addition operation (Element-Wise Sum) on the original input feature X and the fused feature Z to obtain the final output feature; In summary, the residual dilated convolution module is represented by the following formula:
[0013] ;
[0014] Preferably, in step S5, the initial loss function of YOLOv7 is replaced with MPDIoU, which is used as the target bounding box regression loss function; By minimizing the distances between the upper-left and lower-right points of the predicted bounding box and the ground truth bounding box, it can effectively improve the convergence speed and accuracy while simplifying the calculation. The calculation formula of MPDIoU is as follows: ; ; ; Among them, w and h are the width and height of the input image, that is, the ground truth bounding box, and A and B represent the predicted bounding box and the ground truth bounding box respectively. Is the distance between the upper-left corner points of bounding box A and bounding box B. Is the distance between the lower-right corner points of bounding box A and bounding box B.
[0015] Preferably, step S6 includes the following steps: S601. Put the dataset into the improved model for training according to the training strategy. S602. After the training is completed, use the test set to evaluate the improved YOLOv7 model.
[0016] Preferably, the training strategy is: the Size is set to 8, the loss function optimizer is the stochastic gradient optimizer (SGD), and the weight decay, momentum and initial learning rate used by the model are set to 0.937, 0.0005 and 0.01 respectively.
[0017] Preferably, in order to enrich the data set, the weights of the Mosaic and MixUp data enhancements provided by the model are set to 1.0 and 0.15, respectively.
[0018] Preferably, in order to test the performance of the improved model, the experiment uses precision (Precision, P), recall (Recall, R), mean average detection accuracy (mean Average Precision, mAP), average precision (Average Precision, AP) model parameters (Params), model size (Model Size) and giga floating point operations per second (GFLOPs) as evaluation indicators of the model.
[0019] The present invention provides an improved YOLOV7 infrared ship image detection method. First, the present invention improves the backbone network of the initial YOLOv7 model according to the research object (long-distance small and medium-sized infrared ship targets), removes 32× downsampling, and reduces the redundant structure in the model. This optimization greatly reduces the number of model parameters, reduces the amount of calculation, reduces the complexity of the model, and also improves the detection accuracy of the model. In addition, the present invention designs a multi-scale residual hole convolution module (RDC) to effectively utilize the correlation and complementarity of features of different scales, thereby enhancing the semantic information of the features and reducing the loss caused by information loss. This optimization significantly improves the detection accuracy of the model. The algorithm proposed by the present invention greatly reduces the model size and parameter amount, reduces the amount of calculation, and improves the detection performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The present invention will be further described below in conjunction with the accompanying drawings and embodiments: Figure 1 is a flow chart of the method of the present invention; Figure 2 This is the improved algorithm structure diagram of the present invention; Figure 3 This is a structural diagram of the RDC module of the present invention; Figure 4 (a) shows the detection effect of small targets in complex scenes of YOLOv7 before the improvement of the present invention; FIG4( b ) is a diagram showing the effect of detecting small targets in complex scenes after improvement of the present invention; FIG4 (c) is a diagram showing the effect of YOLOv7 dense target detection before the improvement of the present invention; Figure 4 (d) shows the effect diagram of dense target detection after the improvement of the present invention. Specific Embodiments
[0021] The present invention designs a method for detecting medium and small infrared ship targets at a long distance based on the improvement of the YOLOv7 model. The algorithm is implemented in Python and uses deep learning to detect medium and small ship targets at a long distance on the sea. What is needed is sufficient relevant data. The dataset of the present invention is screened and processed from the dataset provided by a certain optoelectronic company, and this dataset is used as the training dataset for the YOLOv7 model. After improving the YOLOv7 model and then training, the improved weights are obtained. The improved weights are used to test the detection effect.
[0022] As Figure 1 shown, a method for detecting medium and small infrared ship targets at a long distance based on the improvement of the YOLOv7 model includes the following steps: S1. Dataset preparation: The dataset is divided into a training set and a validation set according to a ratio of 8:2, and another part of the data is taken as a test set to verify the generalization of the improved model. Preferably, the process of S1 is that the present invention mainly focuses on the safety during ship navigation, which often requires detecting the road conditions ahead in advance, so there are mostly medium and small targets in the detection targets. The experimental dataset selects three categories of fishing boats, boat-type ships, and sailing boats from the dataset provided by Arctis Optoelectronics Co., Ltd. There are mostly medium and small targets in these three types of data, and the targets are in many different complex scenarios, which better meets the requirements of the actual application scenario. The specific distribution of the dataset is shown in Table 1 below.
[0023] Table 1 Dataset category and label distribution
[0024] S2. As Figure 2 shown, multi-scale feature improvement is performed on the overall YOLOv7 network, the downsampling rate is reduced, and the number of detection heads is correspondingly changed from three to two. Preferably, the process of S2 is to perform multi-scale feature improvement on the overall YOLOv7. The YOLOv7 model undergoes five downsampling operations in the backbone network part, which are 2×, 4×, 8×, 16×, and 32× respectively. Through these five downsampling operations, five different scales of feature maps are obtained respectively, and then the feature maps with sizes of 80×80, 40×40, and 20×20 are used to correspond to small, medium, and large targets in the detection targets. The improved algorithm eliminates the 32× downsampling according to the actual scenario of infrared ships at sea during navigation, that is, removes the output of the 20×20 scale, adjusts the output feature size to 80×80 and 40×40, and correspondingly redesigns the feature fusion network.
[0025] S3. Re - cluster the preset anchor box parameters. The sizes of the anchor boxes after re - clustering are shown in Table 2; Preferably, the process of S3 is to re - cluster the preset anchor boxes, and specifically has the following sub - steps: S301: First, evaluate the initial preset anchor boxes. When the BPR between the model's initial preset anchor boxes and the targets in the ship dataset is less than 0.98, it indicates that the sizes of the preset anchor boxes do not match the dataset used in the experiment.
[0026] S302: The anchor box adaptive mechanism of YOLOv7 uses k - means clustering and genetic algorithm analysis to obtain the optimal anchor box sizes.
[0027] S303: According to the improvement in step S2, use k - means clustering and genetic algorithm analysis to obtain two different sizes of anchor boxes corresponding to medium and small receptive fields respectively. At this time, the value of BPR is 0.993, which is better than the default anchor box sizes.
[0028] Table 2 Optimized Anchor Box Sizes
[0029] S4. Design a multi - scale residual dilated convolution module (RDC, Residual Dilated convolution) and add it after the 1×1 convolution. The RDC module is as Figure 3 shown.
[0030] Preferably, the process of S4 is to design a residual dilated convolution module (RDC, Residual Dilated convolution) and add this module after the 1×1 convolution used to change the number of channels from feature extraction to feature fusion stage. Specifically, it has the following sub - steps: S401: The RDC module represents the given input feature map as , where
[0031] means dividing the input feature map into four parts by the number of channels to form four feature maps. where the three values in are 1, 2, and 3 respectively, representing the dilation rates of the dilated convolutions. means the input feature passes through a dilated convolution operation with a kernel size of and a dilation rate of
[0032] S403: Concatenate (concat) the processed features in step S402 to obtain a new output, and then implement channel feature fusion through a 1×1 convolution. Represent the further processed feature map as Z, then the feature Z can be expressed by the following formula: Z =
[0033] S404: According to the residual connection method, perform an element-wise addition operation (Element-Wise Sum) on the original input feature X and the fused feature Z to obtain the final output feature.
[0034] In summary of the above steps, the RDC module designed by the present invention can be expressed by the following formula:
[0035]
[0036] S5: Use the MPDIoU loss function as the bounding box regression loss function.
[0037] Preferably, the process of S5 is to replace the initial loss function of YOLOv7 with MPDIoU to be used as the target bounding box regression loss function. The calculation formula of MPDIoU is as follows:
[0038]
[0039]
[0040] Among them, w and h are the width and height of the input image, that is, the real bounding box, and A and B respectively represent the predicted bounding box and the real bounding box. is the distance between the upper left corner points of bounding box A and bounding box B. is the distance between the lower right corner points of bounding box A and bounding box B.
[0041] S6: Use the improved YOLOv7 model to train and test the dataset. The ablation experiment results are shown in Table 3.
[0042] Preferably, the process of S6 is to use the improved YOLOv7 model to train and test the dataset, and specifically has the following sub-steps: S601: The data set is put into the improved model for training according to a certain training strategy. The training strategy of the present invention is mainly to set the image size of the input network training to 640×640, the training round is 300 epochs, the BatchSize size is set to 8, the loss function optimizer is the stochastic gradient optimizer (SGD), and the weight decay, momentum and initial learning rate used by the model are set to 0.937, 0.0005 and 0.01 respectively. Finally, in order to make the data set richer, the weights of the Mosaic and MixUp data enhancements that come with the model are set to 1.0 and 0.15 respectively.
[0043] S602: After training, the improved YOLOv7 model is evaluated using the test set. In order to test the performance of the improved model in this paper, the experiment uses precision (Precision, P), recall (Recall, R), mean average detection accuracy (mean Average Precision, mAP), average precision (Average Precision, AP), model parameters (Params), model size (Model Size) and giga floating point operations per second (GFLOPs) as evaluation indicators of the model.
[0044] In order to verify the effectiveness of the various improvements to YOLOv7, this paper conducts an ablation experiment on the infrared ship dataset used under the same experimental environment and training hyperparameters. The experiment uses indicators such as accuracy, recall rate, average detection accuracy, model size, parameter amount and computational amount to reflect the effectiveness of the improved model. The ablation experiment results are shown in Table 3 below.
[0045] From the experimental data in Table 3, it can be seen that the multi-scale improvement of YOLOv7 alone and the elimination of the feature map of the scale of 20×20 can effectively reduce the model size, parameters and calculation amount, and effectively improve the model detection accuracy. On the basis of multi-scale improvement, the residual dilated convolution module (RDC) is added, which increases a small amount of parameters and calculation amount, but the accuracy, recall rate and average detection accuracy are improved. On the basis of multi-scale improvement, the loss function is changed to MPDIoU, and the performance is slightly improved. Overall, compared with the original YOLOv7 model, the improved model has improved all indicators, the model size is reduced by 53.02%, the number of parameters is reduced by 52.18%, and the calculation amount is reduced by 2.8GFLOPs. On the infrared ship dataset, the accuracy is improved by 0.4% to 88.6%, the recall rate is improved by 2.9% to 84.8%, and the mAP is improved by 1.2% to 89.5%.
[0046] Table 3 Ablation experiment results
[0047] To truly reflect the effectiveness of the improved algorithm, it is verified on the test set, and the actual detection results of the original algorithm and the improved algorithm are compared, as shown in Figures 4(a) - 4(d). The detection results of the original algorithm are on the left, and the detection effect of the improved algorithm is on the right. From the detection results shown in the pictures, for the detection of small targets in complex scenes, the improved algorithm can effectively reduce missed detections and improve the detection accuracy of small targets; for the detection of dense targets, the improved algorithm has both increases and decreases in the target detection accuracy, but can reduce target missed detections; for the detection of multi-class and multi-scale targets, the improved algorithm can effectively detect small targets among them and also improve the accuracy of other scale targets.
[0048] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as limitations on the present invention. The protection scope of the present invention should be the technical solutions recorded in the claims, including equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, equivalent replacement improvements within this scope are also within the protection scope of the present invention.
Claims
1. An improved YOLOV7 infrared ship image detection method, characterized in that: The following steps are involved: S1. Dataset preparation: divide the dataset into training set and validation set in a ratio of 8:2, and take another part of the data as the test set to verify the generalization of the improved model; S2, improve the multi-scale features of the YOLOv7 overall network; S3, re-clustering the preset anchor box parameters; S4. Design a multi-scale residual dilated convolution module and add it to the back of the 1×1 convolution. S5. Use the MPDIoU loss function as the bounding box regression loss function; S6. Use the improved YOLOv7 model to train and test the dataset.
2. According to claim 1, an improved YOLOV7 infrared ship image detection method is characterized in that: The data set selected in step S1 includes images of fishing boats, small boats and sailing boats.
3. According to claim 1, an improved YOLOV7 infrared ship image detection method is characterized in that: The step S2 is specifically as follows: The YOLOv7 model has undergone five downsampling operations in the backbone network, namely 2×, 4×, 8×, 16×, and 32×. Through these five downsampling operations, five feature maps of different scales are obtained respectively, among which the feature map sizes of 80×80, 40×40 and 20×20 are used to detect small, medium and large targets. After improvement, according to the actual scene of infrared ships at sea during navigation, the 32× downsampling is eliminated, that is, the output of the 20×20 scale is removed, the output feature size is adjusted to 80×80 and 40×40, and the feature fusion network is redesigned.
4. According to claim 1, an improved YOLOV7 infrared ship image detection method is characterized in that: The step S3 comprises the following steps: S301, evaluating the initial preset anchor frame. When the BPR of the model's initial preset anchor frame and the target of the ship dataset is smaller than the BPR of the original model, the preset anchor frame size does not match the dataset used in the experiment. S302,YOLOv7's anchor box adaptation mechanism uses k-means clustering and genetic algorithm analysis to obtain the optimal anchor box size; S303. According to the improvement in step S2, k-means clustering and genetic algorithm analysis are used to obtain two anchor boxes of different sizes corresponding to medium and small receptive fields respectively.
5. According to claim 1, an improved YOLOV7 infrared ship image detection method, characterized in that: The step S4 comprises the following steps: S401, residual hole convolution module represents the given input feature map as , It means that the input feature map is divided into four parts according to the number of channels to form four feature maps; S402, using different convolution pairs to process the four input features in parallel, where the different convolution pairs are a 1×1 convolution and three dilated convolutions with different dilation rates and a kernel size of 3×3; Indicates the size of the dilated convolution rate, The three values are 1, 2, and 3; Represents input features After a convolution kernel size of The void rate is The dilated convolution operation; S403, concatenating the features processed in step S402 to obtain a new output, and then implementing channel feature fusion through a 1×1 convolution; the further processed feature map is represented as Z, and the feature Z is represented by the following formula: ; S404: According to the residual connection method, the original input feature X and the fused feature Z are added element by element to obtain the final output feature; To sum up the above steps, the residual hole convolution module is expressed by the following formula: 。 6. According to claim 1, an improved YOLOV7 infrared ship image detection method is characterized in that: In step S5, the initial loss function of YOLOv7 is replaced by MPDIoU, which is used as the target border regression loss function; by minimizing the distance between the upper left and lower right points of the predicted bounding box and the true bounding box, the calculation is simplified and the convergence speed and accuracy can be effectively improved. The MPDIoU calculation formula is as follows: ; ; ; Among them, w and h are the width and height of the input image, i.e. the real bounding box, and A and B represent the predicted bounding box and the real bounding box respectively. is the distance between the upper left corners of bounding box A and bounding box B, is the distance between the lower right corners of bounding box A and bounding box B.
7. According to claim 1, an improved YOLOV7 infrared ship image detection method is characterized in that: The step S6 comprises the following steps: S601, putting the data set into the improved model for training according to the training strategy; S602: After the training is completed, the improved YOLOv7 model is evaluated using the test set.
8. According to claim 7, an improved YOLOV7 infrared ship image detection method is characterized in that: The training strategy is: the size is set to 8, the loss function optimizer is the stochastic gradient optimizer, and the weight decay, momentum and initial learning rate used by the model are set to 0.937, 0.0005 and 0.01 respectively.
9. According to claim 8, an improved YOLOV7 infrared ship image detection method is characterized in that: In order to enrich the data set, the weights of the Mosaic and MixUp data enhancements provided by the model are set to 1.0 and 0.15 respectively.
10. According to claim 7, an improved YOLOV7 infrared ship image detection method is characterized in that: In order to test the performance of the improved model, the experiment uses accuracy, recall rate, average detection accuracy, average accuracy model parameters, model size and gigaflop operations per second as evaluation indicators of the model.