Ship detection method, device, medium and equipment

By introducing receptive field enhancement feature extraction RFEFM module and high- and low-dimensional feature fusion HLF-FPN network in the YOLOv7 model, and using the F-MPDIoU loss function, the problem of low detection accuracy in ship detection is solved, and higher detection accuracy and fewer missed detection and false detection are achieved.

CN120032250AActive Publication Date: 2025-05-23WUXI UNIV

Patent Information

Application Number
CN202510121960.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-23
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

The prior art has the problem of low detection accuracy in ship detection, especially when complex terrain interference and multi-scale ships exist at the same time.

Method used

The RFEFM module is extracted by replacing the ELAN module in the YOLOv7 model for the receptive field enhancement feature, and adding the high and low-dimensional feature fusion module HLF to the feature pyramid network to form an HLF-FPN network. At the same time, an F-MPDIoU loss function is introduced to improve detection accuracy.

Benefits of technology

The model's feature extraction ability of multi-scale targets is improved, the ability to filter complex background information is enhanced, the accuracy of ship detection is significantly improved, and missed detection and missed detection problems are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032250A_ABST
    Figure CN120032250A_ABST
Patent Text Reader

Abstract

The invention discloses a ship detection method and device, a medium and equipment, and relates to the technical field of target detection, and the method comprises the steps: obtaining a remote sensing image of a ship, and constructing a ship image data set; an improved YOLOv7 model is constructed; on the basis of an original YOLOv7 model, an ELAN module in an original backbone network is replaced by a receptive field enhanced feature extraction RFEFM module, ship features of different scales in a ship image data set are extracted, and on the basis of an original PAFPN feature pyramid structure, a high-low dimension feature fusion module HLF module is added, so that the ship image data set is obtained. Obtaining a high and low dimension fusion feature pyramid HLF-FPN network, and carrying out feature fusion; training the improved YOLOv7 model to obtain a ship detection model for ship detection; the to-be-recognized ship image is obtained, the ship image is input into the ship detection model for ship recognition, and the ship detection result can be accurately obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection, and in particular to a ship detection method, device, medium and equipment. Background Art

[0002] As a large country with a vast sea area, ship detection technology plays a vital role in ensuring national security and maritime transportation. Synthetic Aperture Radar (SAR) is a popular remote sensing technology. SAR images overcome the influence of the complex and changeable climate environment at sea and have the characteristics of all-weather detection, making it widely used in military reconnaissance, maritime traffic control and maritime rescue.

[0003] In recent years, with the vigorous development of deep learning and SAR imaging technology, SAR ship image detection technology has made great progress. Many scholars have proposed various SAR ship image detection algorithms for different scenarios, which have achieved good results in detection accuracy and speed. In view of the irregularity of the ship shape, in order to better extract the geometric features of the ship, Ke et al. improved the Faster R-CNN network using deformable convolution. Liu et al. introduced coordinate attention in the YOLOv7-tiny model, improved the spatial pyramid pooling (SPP) and SIoU loss functions, and enhanced the detection performance. In view of the large scale difference caused by the size of the ship itself and the imaging distance, Zhang et al. proposed a detection network containing a new spatial cross-scale attention (SCSA) module to eliminate the interference of noise and complex background information. Zhao et al. designed a multi-scale denoising network based on the Laplacian operator, combined it with the Yolox model, and obtained a lightweight detection model.

[0004] Although the above method has improved target detection in bad weather, the ship images collected by SAR are affected by complex terrain such as ports and islands, which will cause interference in the ship detection process; when ships of different scales appear in the same image, the above method still has the problem of low detection accuracy for densely occluded targets. Summary of the invention

[0005] The present invention provides a ship detection method, device, medium and equipment to solve the above-mentioned problem existing in the prior art, that is, the problem of how to improve the accuracy of ship detection in the prior art. The present invention provides a ship detection method, which includes:

[0006] Acquire remote sensing images of ships and build a ship image dataset;

[0007] An improved YOLOv7 model is constructed; based on the original YOLOv7 model, the ELAN module in the original backbone network is replaced with a receptive field enhancement feature extraction RFEFM module to extract ship features of different scales in the ship image data set; on the basis of the original PAFPN feature pyramid structure, a high- and low-dimensional feature fusion module HLF module is added to obtain a high- and low-dimensional fusion feature pyramid HLF-FPN network, and high-level features and low-level features are fused through a top-down path; wherein the receptive field enhancement feature extraction RFEFM module includes three parallel branches, including a residual network branch, a partial convolution PConv network branch, and a convolutional layer and a lightweight RFCA attention network branch, wherein the RFCA attention network adds a receptive field attention RFA module on the basis of the lightweight CA attention;

[0008] Based on the ship image dataset, the improved YOLOv7 model is trained to obtain a ship detection model for ship detection;

[0009] Obtain the ship image to be identified, input the ship image into the ship detection model for ship identification, and obtain the ship detection result.

[0010] Optionally, the acquiring of the remote sensing image of the ship specifically includes:

[0011] The remote sensing images of ships are obtained by using synthetic aperture radar SAR.

[0012] Optionally, the partially convolutional PConv network specifically includes:

[0013] The 1x1, 3x3, and 5x5 partial convolution PConv networks that only convolve some channels of the input features have a computational cost of h×w×k 2 ×c p 2 ; Among them, c p is the part of the channel for convolution, h is the height of the input feature, w is the width of the input feature, and k is the size of the convolution kernel.

[0014] Also includes:

[0015] Based on the original YOLOv7 model, the F-MPDIoU loss function is used to replace the original CIoU loss function; wherein the F-MPDIoU loss function is constructed by combining the MPDIoU loss function with the idea of ​​Focaler-IoU.

[0016] Optionally, the MPDIoU loss function specifically includes:

[0017]

[0018] Among them, A gt , B gt , A prd , B prd Represents the points in the upper left and lower right corners of the real box and the predicted box, w and h represent the width and height of the input feature respectively, ρ 2 (A prd ,A gt ) and ρ 2 (B prd ,B gt ) represents the Euclidean distance between two points.

[0019] Optionally, the F-MPDIoU loss function is constructed by combining the MPDIoU loss function with the idea of ​​Focaler-IoU, specifically including:

[0020] L F-MPDIoU =L MPDIoU +IoU-IoU Focaler

[0021] L MPDIoU =1-MPDIoU

[0022]

[0023] Among them, L F-MPDIoU represents the F-MPDIoU loss function, L MPDIoU Represents the MPDIoU loss function; IoU Focaler is the reconstructed intersection-over-union ratio, IoU is the original intersection-over-union ratio, and MPDIoU is the minimum perimeter distance intersection-over-union ratio; d and u are numbers between [0,1]. By adjusting the values ​​of d and u, IoU Focaler Focus on different regression samples.

[0024] The present invention provides a ship detection device, comprising:

[0025] The acquisition module is used to acquire remote sensing images of ships and construct a ship image dataset;

[0026] A construction module is used to replace the ELAN module in the original backbone network with the receptive field enhancement feature extraction RFEFM module based on the original YOLOv7 model, and extract ship features of different scales in the ship image data set; on the basis of the original PAFPN feature pyramid structure, by adding a high- and low-dimensional feature fusion module HLF module, a high- and low-dimensional fusion feature pyramid HLF-FPN network is obtained, and high-level features and low-level features are fused through a top-down path; wherein the receptive field enhancement feature extraction RFEFM module includes three parallel branches, including a residual network branch, a partial convolution PConv network branch, and a convolutional layer and a lightweight RFCA attention network branch, wherein the RFCA attention network adds a receptive field attention RFA module on the basis of the lightweight CA attention;

[0027] A training module is used to train the improved YOLOv7 model based on the ship image dataset to obtain a ship detection model for ship detection;

[0028] The detection module is used to obtain the ship image to be identified, input the ship image into the ship detection model for ship identification, and obtain the ship detection result.

[0029] The present invention provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned ship detection method is implemented.

[0030] The present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned ship detection method when executing the program.

[0031] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention provides a ship detection method, which reconstructs the backbone network through the RFEFM module, replaces the feature extraction network in the original YOLOv7 model, effectively enhances the scale adaptability of the model and the receptive field of the extracted target, and finally improves the feature extraction ability of the backbone network for multi-scale targets; by replacing the original feature pyramid network with the HLF-FPN feature pyramid network, it is possible to filter invalid background information based on the high and low dimensional feature fusion module HLF, and interact with different dimensional feature information across scales, so that the neck network can obtain more accurate information such as ship position and texture; in addition, the present invention improves the original CIoU loss function, incorporates the idea of ​​Focaler-IoU into the MPDIoU loss function, and proposes a new F-MPDIoU loss function, which can take into account the sample imbalance problem in the training process, make the gradient distribution more reasonable, improve the problems of missed detection and false detection during ship detection, and significantly improve the accuracy of ship detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0033] Figure 1 A flow chart of a ship detection method provided by an embodiment of the present invention;

[0034] Figure 2 A structural diagram of the RFEFM module provided in an embodiment of the present invention;

[0035] Figure 3 A PConv structure diagram provided for an embodiment of the present invention;

[0036] Figure 4 RFCA attention mechanism diagram provided for an embodiment of the present invention;

[0037] Figure 5 HLF-FPN network diagram provided for an embodiment of the present invention;

[0038] Figure 6 A structural diagram of an HLF module provided in an embodiment of the present invention;

[0039] Figure 7 An improved YOLOv7 model structure diagram provided by an embodiment of the present invention;

[0040] Figure 8 A comparison chart of experimental test results of the original YOLOv7 model and the improved YOLOv7 model provided in an embodiment of the present invention in various scenarios;

[0041] Fig. 9 A comparison chart of the loss functions of the original YOLOv7 model and the improved YOLOv7 model provided in an embodiment of the present invention;

[0042] Fig.10 A schematic diagram of a computer device for a ship detection method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0044] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.

[0045] Figure 1 is a flow chart of a ship detection method provided by an embodiment of the present invention, such as Figure 1 As shown, a ship detection method shown in this embodiment includes:

[0046] S1: Obtain remote sensing images of ships and build a ship image dataset.

[0047] For example, by using the high-resolution synthetic aperture radar (SAR) image HRSID dataset released by a certain school, it contains SAR images of different sizes and dimensions with different resolutions, polarization modes, sea conditions, sea areas and ports; the original images are 136 large-scale SAR satellite images, which are cropped into 5604 800×800 SAR images, including 16951 ships, of which the proportions of small, medium and large ships are 54.5%, 43.5% and 2% respectively. The image resolution ranges from 0.5m to 3m. The dataset is randomly divided into training set, test set and validation set according to 7:2:1.

[0048] S2: Based on the original YOLOv7 model, the ELAN module in the original backbone network is replaced by the receptive field enhancement feature extraction RFEFM module to extract ship features of different scales in the ship image dataset; on the basis of the original PAFPN feature pyramid structure, a high- and low-dimensional feature fusion module HLF module is added to obtain a high- and low-dimensional fusion feature pyramid HLF-FPN network, and high-level features and low-level features are fused through a top-down path.

[0049] Among them, the receptive field enhancement feature extraction RFEFM module includes three parallel branches, including a residual network branch, a partial convolution PConv network branch, and a convolutional layer and a lightweight RFCA attention network branch. The RFCA attention network adds a receptive field attention RFA module on the basis of the lightweight CA attention.

[0050] Exemplary, the receptive field enhancement feature extraction network architecture RFEFM module is as follows Figure 2As shown in the figure, it consists of three branches. The first branch is a residual connection, which can increase the transmission capacity of network information and improve the generalization ability of the model. The second branch first uses partial convolution (PConv) with convolution kernel sizes of 1x1, 3x3 and 5x5 to extract features of targets of different scales, while reducing redundant model calculations and memory access, and then concatenates the output by channel for feature fusion. The third branch consists of a traditional convolution layer and a lightweight RFCA attention layer to capture the effective receptive field in the contextual information. Finally, the output feature maps of the three branches are added in the spatial dimension as the output of RFEFM.

[0051] After multiple convolutions, the extracted feature maps have a large number of similar redundant features, so we use Figure 3 The PConv shown here only convolves some channels of the input features, leaving the rest unchanged. In this way, the redundant computation of the model during feature extraction is greatly reduced, thereby extracting spatial features more efficiently. For an input feature layer with a channel number, height, and width of c×h×w, the conventional convolution computation is h×w×k. 2 ×c 2 , where k is the size of the convolution kernel, and the computational complexity of PConv is h×w×k 2 ×c p 2 , where c p is the part of the channel that is convolved. If c p =1 / 4, the computational complexity of PConv is only 1 / 16 of that of conventional convolution.

[0052] like Figure 4 As shown in the figure, RFCA (Receptive-Field Coordinate Attention) adds a receptive field attention module (RFA) on the basis of lightweight CA attention. The RFA module not only emphasizes the importance of different features within the receptive field, but also prioritizes the spatial features of the receptive field, capturing long-distance information in a similar way to the self-attention mechanism. Through this method, the limitation of the CA attention mechanism that only focuses on spatial features is solved at a very low computational cost, the problem of convolution kernel parameter sharing is perfectly solved, and attention to important contextual information is improved.

[0053] The process of constructing the HLF-FPN feature pyramid network in the present invention is described below.

[0054] Exemplarily, in the multi-scale feature map generated by the Backbone network, low-scale features have higher resolution and provide accurate target location information; high-scale features have lower resolution and rich semantic information. By fusing shallow features and deep features, the feature expression ability of the model can be enhanced, thereby improving the performance of the model. The original YOLOv7 model uses the PAFPN feature pyramid structure, which simply adds multiple feature layers, and in the process of feature information being transmitted at different intermediate scales, multiple upsampling will cause the loss and degradation of semantic information at different levels. In order to solve this limitation, in the present invention, the PAFPN feature pyramid network is improved to obtain a high-low dimensional fusion feature pyramid network (High-Low Dimension Fusion Feature Pyramid Network, HLF-FPN), and the structure diagram is shown as follows. Figure 5 shown.

[0055] The HLF-FPN network is constructed by Figure 6 The high-low dimensional feature fusion module (HLF) shown in the figure establishes an efficient connection between feature layers of different scales, so that low-dimensional position information and high-dimensional semantic information can interact better and better fuse feature information of different dimensions. The HLF module first uses the RFCA attention mechanism to filter out interference information such as noise and complex background in the low-scale feature map and expand the receptive field; then, in order to unify the dimensions of high-level features and low-scale features, bilinear interpolation is used to up- or down-sample high-level features; finally, the filtered low-scale features are fused with high-scale features to enhance the feature representation ability of the model. In order to improve the feature pyramid network to fuse more accurate target position information, the P2 feature layer extracted by the Backbone network is introduced into the feature pyramid network to provide more comprehensive feature information for the model.

[0056] The improved network model is as follows Figure 7 As shown in the figure, the backbone network Backbone is mainly composed of RFEFM modules, which inherit the gradient diversion strategy and use PConv of different sizes to control the model parameters and realize feature extraction of targets of different scales. In order to make up for the decrease in accuracy caused by the reduction in parameters, a receptive field enhancement branch consisting of RFCA attention and ordinary convolution is designed. The neck network Neck innovatively proposes the HLF module to improve the feature fusion network, reduce the semantic degradation problem in the process of feature information transmission, and introduces the P2 feature layer with richer ship information into the feature fusion network.

[0057] Optionally, based on the original YOLOv7 model, an F-MPDIoU loss function is used to replace the original CIoU loss function; wherein the F-MPDIoU loss function is constructed by combining the MPDIoU loss function with the idea of ​​Focaler-IoU.

[0058] For example, the YOLOv7 algorithm uses the CIoU loss function to calculate the prediction box regression loss, and the formula is as follows:

[0059]

[0060]

[0061]

[0062]

[0063] Among them, IoU represents the real box B gt and prediction box B prd The intersection-over-union ratio, ρ 2 (B gt ,B prd ) is the Euclidean distance between the center point of the predicted box and the true box, C 2 represents the diagonal length of the minimum outer rectangle, is the balance coefficient, V is used to measure the aspect ratio of the real box and the predicted box, w gt 、w prd 、h gt and h prd They represent the true box width, predicted box width, true box height and predicted box height respectively.

[0064] The CIoU loss function uses a relative value to describe the aspect ratio of the true box and the predicted box. When the aspect ratio of the true box is equal to that of the predicted box, the CIoU penalty term will degenerate and become invalid, which will lead to slow model convergence and poor results.

[0065] In order to make full use of the geometric characteristics of the bounding box in the loss function, the MPDIoU loss function is proposed in the prior art. MPDIOU simplifies the similarity comparison between the predicted box and the real box, directly predicts the minimum distance between the upper left corner and the lower right corner, and can adapt to overlapping or non-overlapping bounding box regression, improving the convergence speed and detection accuracy. The MPDIoU loss function formula is as follows:

[0066]

[0067] In the formula, A gt , B gt , A prd , B prdRepresents the points in the upper left and lower right corners of the real box and the predicted box, w and h represent the width and height of the image, ρ 2 (A prd ,A gt ) and ρ 2 (B prd ,B gt ) represents the Euclidean distance between two points.

[0068] In order to further reasonably distribute gradients and improve detection accuracy, this paper incorporates the idea of ​​Focaler-IoU into the MPDIoU loss function and proposes the F-MPDIoU loss function, the formula is as follows:

[0069] L F-MPDIoU =L MPDIoU +IoU-IoU Focaler

[0070] L MPDIoU =1-MPDIoU

[0071]

[0072] Among them, L F-MPDIoU represents the F-MPDIoU loss function, L MPDIoU Represents the MPDIoU loss function; IoU Focaler is the reconstructed intersection-over-union ratio, IoU is the original intersection-over-union ratio, and MPDIoU is the minimum perimeter distance intersection-over-union ratio; d and u are numbers between [0,1]. By adjusting the values ​​of d and u, IoU Focaler Focus on different regression samples.

[0073] The F-MPDIoU loss function not only includes factors such as overlapping or non-overlapping areas, center point distance, width and height deviation considered in the MPDIoU loss function, but also takes into account the problem of sample imbalance during training. By introducing the idea of ​​Focaler-IoU, the F-MPDIoU loss function can divide simple samples and difficult samples according to the difficulty of detection, and pay more attention to difficult samples during training, so that the F-MPDIoU loss function can allocate gradients more reasonably during training, thereby improving the model's detection ability for difficult samples. Experimental results show that the model using the F-MPDIoU loss function achieves higher detection accuracy in SAR ship detection tasks.

[0074] In this step, the RFEFM module is used to replace the ELAN module in the YOLOv7 backbone network, which effectively enhances the scale adaptability of the model and the receptive field of the extracted target, and improves the backbone network's feature extraction capability for multi-scale targets. In addition, by performing feature fusion through the high- and low-dimensional feature fusion module HLF in the multi-scale feature map generated by the backbone network Backbone, the P2 feature layer is introduced to enrich the ship information, which can filter out invalid background information, interact with feature information of different dimensions across scales, and enable the neck network Neck to obtain more accurate information such as ship position and texture, thereby improving the accuracy of ship recognition and detection.

[0075] S3: Based on the ship image dataset, the improved YOLOv7 model is trained to obtain a ship detection model for ship detection.

[0076] Exemplarily, while training the improved YOLOv7 model, the original YOLOv7 model is also trained to facilitate the comparison of subsequent test results. The experimental environment configuration selected for training the YOLOv7 model before and after the improvement is shown in Table 1.

[0077] Table 1 Experimental environment platform configuration

[0078]

[0079] Using the transfer learning training method, the pre-trained weights of YOLOv7 are used to train the improved YOLOv7 model for 200 epochs on the HRSID dataset, with a batch-size of 4 and the SGD optimizer.

[0080] As shown in Table 2, by comparing the improved YOLOv7 algorithm with other algorithms, precision (P), recall (R), mean average precision (mAP@0.5) and frames per second (FPS) are used as evaluation indicators to analyze the model performance; by observing the performance during the training process and continuously adjusting the network structure parameters until the weight with the best effect after improvement is obtained.

[0081] Table 2 Comparison of improved YOLOv7 algorithm with other algorithms

[0082]

[0083] S4: Obtain a ship image to be identified, input the ship image into a ship detection model for ship identification, and obtain a ship detection result.

[0084] For example, the trained weights of YOLOv7 before and after the improvement are loaded, and the verification images are inferred and predicted. Then the ship detection situation is output, and the detection effect of the model in four scenes, near the coastal port, near the islands and reefs, dense ship groups and the ocean scene, is compared. The detection effect diagram is as follows: Figure 8 To verify the effectiveness of the F-MPDIoU loss function proposed in this invention, CIoU, MPDIoU and F-MPDIoU were trained for 200 rounds in the HRSID dataset and the same experimental environment. The comparison results of the visualized loss functions are shown in Figure 2. Fig. 9 shown.

[0085] The results of ship identification using the ship detection model proposed in the present invention show that the detection accuracy is significantly improved, and the problems of false detection and missed detection in SAR ship detection under complex interference scenarios are improved.

[0086] The above is a ship detection method provided by one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding ship detection device, including:

[0087] The acquisition module is used to acquire remote sensing images of ships and construct a ship image dataset;

[0088] A construction module is used to replace the ELAN module in the original backbone network with the receptive field enhancement feature extraction RFEFM module based on the original YOLOv7 model, and extract ship features of different scales in the ship image data set; on the basis of the original PAFPN feature pyramid structure, by adding a high- and low-dimensional feature fusion module HLF module, a high- and low-dimensional fusion feature pyramid HLF-FPN network is obtained, and high-level features and low-level features are fused through a top-down path; wherein the receptive field enhancement feature extraction RFEFM module includes three parallel branches, including a residual network branch, a partial convolution PConv network branch, and a convolutional layer and a lightweight RFCA attention network branch, wherein the RFCA attention network adds a receptive field attention RFA module on the basis of the lightweight CA attention;

[0089] A training module is used to train the improved YOLOv7 model based on the ship image dataset to obtain a ship detection model for ship detection;

[0090] The detection module is used to obtain the ship image to be identified, input the ship image into the ship detection model for ship identification, and obtain the ship detection result.

[0091] For the specific definition of the ship detection device, please refer to the definition of the ship detection method above, which will not be repeated here. Each module in the above-mentioned ship detection device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0092] The present invention also provides a computer-readable storage medium, which stores a computer program. The computer program can be used to execute the above-mentioned ship detection method.

[0093] The present invention also provides Fig.10 The structural diagram of the computer device shown in FIG. Fig.10 As shown, at the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the ship detection method provided in the above embodiment.

[0094] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0095] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present invention.

Claims

1. A ship detection method, characterized in that: include: Acquire remote sensing images of ships and build a ship image dataset; An improved YOLOv7 model is constructed; based on the original YOLOv7 model, the ELAN module in the original backbone network is replaced with a receptive field enhancement feature extraction RFEFM module to extract ship features of different scales in the ship image data set; on the basis of the original PAFPN feature pyramid structure, a high- and low-dimensional feature fusion module HLF module is added to obtain a high- and low-dimensional fusion feature pyramid HLF-FPN network, and high-level features and low-level features are fused through a top-down path; wherein the receptive field enhancement feature extraction RFEFM module includes three parallel branches, including a residual network branch, a partial convolution PConv network branch, and a convolutional layer and a lightweight RFCA attention network branch, wherein the RFCA attention network adds a receptive field attention RFA module on the basis of the lightweight CA attention; Based on the ship image dataset, the improved YOLOv7 model is trained to obtain a ship detection model for ship detection; Obtain the ship image to be identified, input the ship image into the ship detection model for ship identification, and obtain the ship detection result.

2. The ship detection method according to claim 1, characterized in that: The acquiring of the remote sensing image of the ship specifically includes: The remote sensing images of ships are obtained by using synthetic aperture radar SAR.

3. The ship detection method according to claim 1, characterized in that: The partial convolution PConv network specifically includes: The 1x1, 3x3, and 5x5 partial convolution PConv networks that only convolve some channels of the input features have a computational cost of h×w×k 2 ×c p 2 ; Among them, c p is the part of the channel for convolution, h is the height of the input feature, w is the width of the input feature, and k is the size of the convolution kernel.

4. The ship detection method according to claim 1, characterized in that: The method further comprises: Based on the original YOLOv7 model, the F-MPDIoU loss function is used to replace the original CIoU loss function; wherein the F-MPDIoU loss function is constructed by combining the MPDIoU loss function with the idea of ​​Focaler-IoU.

5. The ship detection method according to claim 4, characterized in that: The MPDIoU loss function specifically includes: Among them, A gt , B gt , A prd , B prd Represents the points in the upper left and lower right corners of the real box and the predicted box, w and h represent the width and height of the input feature respectively, ρ 2 (A prd ,A gt ) and ρ 2 (B prd ,B gt ) represents the Euclidean distance between two points.

6. The ship detection method according to claim 4, characterized in that: The F-MPDIoU loss function is constructed by combining the MPDIoU loss function with the idea of ​​Focaler-IoU, specifically including: L F-MPDIoU =L MPDIoU +IoU-IoU Focaler L MPDIoU =1-MPDIoU Among them, L F-MPDIoU represents the F-MPDIoU loss function, L MPDIoU Represents the MPDIoU loss function; IoU Focaler is the reconstructed intersection-over-union ratio, IoU is the original intersection-over-union ratio, and MPDIoU is the minimum perimeter distance intersection-over-union ratio; d and u are numbers between [0,1]. By adjusting the values ​​of d and u, IoU Focaler Focus on different regression samples.

7. A ship detection device, characterized in that: include: The acquisition module is used to acquire remote sensing images of ships and construct a ship image dataset; A construction module is used to construct an improved YOLOv7 model; based on the original YOLOv7 model, the ELAN module in the original backbone network is replaced with a receptive field enhancement feature extraction RFEFM module to extract ship features of different scales in a ship image data set; on the basis of the original PAFPN feature pyramid structure, a high- and low-dimensional feature fusion module HLF module is added to obtain a high- and low-dimensional fusion feature pyramid HLF-FPN network, and high-level features and low-level features are fused through a top-down path; wherein the receptive field enhancement feature extraction RFEFM module includes three parallel branches, including a residual network branch, a partial convolution PConv network branch, and a convolutional layer and a lightweight RFCA attention network branch, wherein the RFCA attention network adds a receptive field attention RFA module on the basis of the lightweight CA attention; A training module is used to train the improved YOLOv7 model based on the ship image dataset to obtain a ship detection model for ship detection; The detection module is used to obtain the ship image to be identified, input the ship image into the ship detection model for ship identification, and obtain the ship detection result.

8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the ship detection method according to any one of claims 1 to 6 is implemented.

9. A computer device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the ship detection method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • YOLOv8-based ship target rotation detection method

    CN117830622A

  • Solar cell surface defect detection method and system based on improved YOLOv8

    CN117994210A

  • Ship detection method, device and equipment based on synthetic aperture radar image

    CN118691964A

  • KR20240101284A

Cited By

  • Lightweight visible light ship target detection method based on edge feature guidance

    CN120656032A

  • Lightweight visible ship target detection method based on edge feature guidance

    CN120656032B