Ship target detection method and system based on attention mechanism and mask mechanism

By introducing a method based on attention mechanism and mask mechanism in ship object detection, using the mask attention module to fuse space and channel features, and filtering information through the mask mechanism, the problem of insufficient accuracy of ship object detection in complex environments is solved, and more efficient detection performance and robustness are achieved.

CN120088636APending Publication Date: 2025-06-03NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411386861.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art still has room for improvement in detection accuracy in ship target detection in complex environments, especially when dealing with multi-size ships and complex backgrounds.

Method used

The ship object detection method based on attention mechanism and mask mechanism is adopted, and the spatial characteristics and channel characteristics are fusion through the mask attention module, and the mask mechanism is used to filter useless information to enhance the feature extraction ability of ship detection.

Benefits of technology

It improves the accuracy and detection performance of ship target detection, enhances the robustness of the model, simplifies the network structure, and reduces the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088636A_ABST
    Figure CN120088636A_ABST
Patent Text Reader

Abstract

The invention discloses a ship target detection method and system based on an attention mechanism and a mask mechanism, and the method comprises the steps: carrying out the ship target detection of a to-be-detected target image through a target detection neural network based on a multi-scale fusion strategy, and achieving the ship target detection. A mask attention module in the target detection neural network realizes fusion of spatial features and channel features; firstly, Query, Key and Value are generated through two times of convolution, and the Query and the Key are subjected to matrix multiplication through dimension transformation to obtain channel features and spatial features respectively; then, mask mechanism filtering is carried out on the channel features and the spatial features, and channel sparse attention and spatial sparse attention are obtained; matrix multiplication is carried out on the two branches and Value to obtain sparse space features and sparse channel features, output of the two branches is spliced and then subjected to convolution once, and the output is the features subjected to masking and attention extraction. The accuracy of ship detection is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In the field of ship detection, the working environment of ships is very complex. For example, during the ship's voyage, containers, high-rise buildings, goods, sand piles, etc. in the background are very similar to the outline of the ship itself and are often misidentified as ships. Secondly, due to the vastness of the sea and the variety of ship types including yachts, small fishing boats, etc., the size of ships varies greatly. Therefore, the neural network needs to have good multi-size processing capabilities and effectively filter background information.

[0003] Using a trained neural network for object detection is a commonly used object detection method; however, due to the large variation in the size of ships themselves and the complex background information, it is often difficult to consider background information during convolution and unable to use background information for object recognition.

[0004] To address the challenges of ship object detection, various strategies have been proposed in the prior art to solve the problem of ship object detection, including multi-scale fusion strategies, context information extraction strategies, etc. These methods have all significantly improved the performance of the model.

[0005] Among them, for the multi-scale fusion strategy, FPN (Feature Pyramid Networks) is one of the pioneering works, which proposed constructing a feature pyramid for feature fusion; for example, in the YOLO series of networks, multi-scale fusion can be found from YOLOv3 to the current new version YOLOv9.

[0006] Figure 1 Shows the structural schematic diagram of the existing FPN; as can be seen from Figure 1 The existing FPN realizes multi-scale feature fusion by first upsampling the low-level features (Upsample) and then adding them to the high-level features to form a residual structure, thereby achieving the fusion of high-level and low-level features.

[0007] For the context information extraction strategy, self-attention expands the receptive field to the global, enabling it to consider the global background information for object recognition using background information.

[0008] However, although the above multi-scale fusion strategy and context information extraction strategy have shown good performance in the face of ship objects, their performance in detection accuracy still needs to be further improved. Summary of the Invention

[0009] The technical problem to be solved by the present invention is to provide a ship object detection method and system based on an attention mechanism and a mask mechanism to address the low ship object detection in complex environments in view of the above deficiencies in the prior art.

[0010] The present invention adopts the following technical solutions: A ship target detection method based on an attention mechanism and a masking mechanism, comprising the following steps: Using an object detection neural network based on a multi-scale fusion strategy to perform ship target detection on the target image to be detected, realizing ship target detection; Among them, in the object detection neural network, a masked attention module is used to realize the fusion of spatial features and channel features; the masked attention module includes a spatial attention extraction branch and a channel attention extraction branch, which are respectively used to extract spatial features and channel features; The masked attention module first generates Query, Key, and Value. After the matrix multiplication of the dimension-transformed Query and Key, channel features and spatial features are respectively obtained; The extracted channel features and spatial features are respectively filtered by the masking mechanism to retain the useful information for ship detection, obtaining channel sparse attention and spatial sparse attention; The channel sparse attention and the spatial sparse attention are respectively multiplied by Value in matrix form to obtain sparse spatial features and sparse channel features. After splicing the outputs of the two branches, a convolution is performed again, and the output is the feature refined by masking and attention.

[0011] Preferably, the object detection neural network includes a masked attention module, which is used to extract the channel features and spatial features in the object detection neural network and then select and strengthen the features through masking.

[0012] Preferably, the masked attention module first generates Query, Key, and Value specifically as follows:

[0013] Among them, represents the input feature, , , represents the convolution operator, , , represent the Query matrix, the Key matrix, and the Value matrix.

[0014] Preferably, after the matrix multiplication of the dimension-transformed Query and Key, channel features and spatial features are respectively obtained and as follows:

[0015] Among them, and respectively represent the extracted channel features and spatial features, respectively transform and to dimensions of and .

[0016] Preferably, the feature outputs after filtering the extracted channel features and spatial features through a mask mechanism respectively are:

[0017] wherein, is the element in the th row and jth column of the input matrix, is the k-th largest value in the j-th row of the attention matrix , and k is a learnable parameter.

[0018] Preferably, the process of fusing the weighted spatial features and channel features includes:

[0019] wherein, and are the spatial features and channel features after mask selection, represents the concatenation operation, represents a 1×1 cross-channel convolution operation and a 3×3 depthwise separable convolution.

[0020] Preferably, the final output feature after mask selection and weighted by Value is as follows:

[0021] wherein, represents the final output feature after mask selection and weighted by Value, and represent the new Value obtained by dimension transformation of the Value obtained after preprocessing, represents matrix multiplication.

[0022] Preferably, the object detection neural network adopts a Feature Pyramid Network FPN and / or a Path Aggregation Network PAN.

[0023] Preferably, the mask attention module implements YOLOX, YOLOv7 or ResNet50 for fusing channel features and spatial features.

[0024] In a second aspect, an embodiment of the present invention provides a ship target detection system based on an attention mechanism and a mask mechanism, which is characterized by including: A data module to obtain a target image to be detected; A detection module to perform ship target detection on the target image to be detected by using a target detection neural network based on a multi-scale fusion strategy, thereby achieving ship target detection; In the target detection neural network, a mask attention module is used to implement the fusion of spatial features and channel features; the mask attention module includes a spatial attention extraction branch and a channel attention extraction branch, which are respectively used to extract spatial features and channel features; The mask attention module first generates Query, Key, and Value. After matrix multiplication of the dimension-transformed Query and Key, channel features and spatial features are respectively obtained; The extracted channel features and spatial features are respectively filtered by a mask mechanism to retain the useful information for ship detection, and channel sparse attention and spatial sparse attention are obtained; The channel sparse attention and the spatial sparse attention are respectively multiplied by Value through matrix multiplication to obtain sparse spatial features and sparse channel features. After the outputs of the two branches are concatenated and then convolved once, the features refined by mask and attention are output.

[0025] In a third aspect, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned ship target detection method based on the attention mechanism and the mask mechanism are implemented.

[0026] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, including a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned ship target detection method based on the attention mechanism and the mask mechanism are implemented.

[0027] Compared with the prior art, the present invention has at least the following beneficial effects: A ship target detection method based on an attention mechanism and a mask mechanism. In the target detection neural network based on a multi-scale fusion strategy used, the fusion of channel features and spatial features is implemented based on a mask attention module; the mask attention module obtains spatial features and channel features through convolution and matrix multiplication; then, mask selection is performed on the extracted channel features and spatial features, effectively retaining the useful information of the ship and filtering out the useless information; then, the sparse spatial features and sparse channel features after mask selection will be concatenated and convolved once again to further fuse channel and spatial information; based on the above operations, ship target detection based on this target detection neural network can further improve the detection performance for ships.

[0028] Furthermore, the masked attention module first generates Query, Key, and Value, which helps improve detection performance, enhance model robustness, reduce computational complexity, simplify the network structure, and facilitate cross-scale feature fusion.

[0029] Furthermore, after the dimension-transformed Query and Key are subjected to matrix multiplication, the channel features and spatial features obtained respectively help improve detection performance, enhance model robustness, reduce computational complexity, simplify the network structure, and facilitate cross-scale feature fusion.

[0030] Furthermore, the feature outputs obtained by filtering the extracted channel features and spatial features through a masking mechanism respectively help improve detection performance, enhance model robustness, reduce computational complexity, simplify the network structure, and facilitate cross-scale feature fusion.

[0031] Furthermore, fusing the weighted spatial features and channel features helps improve detection performance, enhance model robustness, reduce computational complexity, simplify the network structure, and facilitate cross-scale feature fusion.

[0032] Furthermore, the object detection neural network adopting the Feature Pyramid Network (FPN) and / or the Path Aggregation Network (PAN) helps achieve the balance of multi-scale feature extraction, computational efficiency, and model performance, simplify and optimize the network structure, facilitate cross-scale feature fusion, and enhance model robustness.

[0033] Furthermore, the YOLOX, YOLOv7, or ResNet50 that the masked attention module uses to achieve the fusion of channel features and spatial features helps improve detection performance, enhance model robustness, reduce computational complexity, simplify the network structure, and facilitate cross-scale feature fusion.

[0034] It can be understood that the beneficial effects of the second aspect above can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here.

[0035] In summary, the present invention realizes the fusion of channel features and spatial features based on a masked attention module, and shows obvious advantages in aspects such as multi-scale feature extraction, effective fusion of channel features and spatial features, balance of computational efficiency and model performance, simplification and optimization of the network structure, and facilitation of cross-scale feature fusion, providing new ideas and directions for the development of the ship target detection field.

[0036] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings to be used in the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0038] Figure 1 It is a schematic structural diagram of the existing ResNet50 network; Figure 2 It is a flowchart of the ship target detection method based on the attention mechanism and the mask mechanism of the present invention; Figure 3 For Figure 2 It is a schematic structural diagram of the mask attention module used in the target detection neural network in the method shown; Figure 4 It is in the embodiments of the present invention Figure 1 The schematic structural diagram of the mask attention module shown is used in the ResNet50 shown; Figure 3 It is a schematic structural diagram of the mask attention module shown; Figure 5 It is a diagram showing the label distribution of the dataset used in the simulation verification of the embodiments of the present invention; Figure 6 It is a schematic diagram of a computer device provided by an embodiment of the present invention; Figure 7 It is a block diagram of an electronic device provided according to an embodiment of the present invention. Detailed implementation manners

[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0040] In the description of the present invention, it should be understood that the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0041] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0042] It should be further understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the present invention, the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0043] It should be understood that although terms such as first, second, and third may be used in the embodiments of the present invention to describe preset ranges, etc., these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0044] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0045] Schematic diagrams of various structures according to the disclosed embodiments of the present invention are shown in the drawings. These figures are not drawn to scale, where for the purpose of clear expression, some details are enlarged and some details may be omitted. The shapes of various regions and layers shown in the figures and their relative sizes and positional relationships are only exemplary, and may actually deviate due to manufacturing tolerances or technical limitations, and those skilled in the art can additionally design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0046] The present invention provides a ship target detection method based on an attention mechanism and a masking mechanism. First, a target image to be detected is obtained; a target detection neural network based on a multi-scale fusion strategy is used to perform target detection on the target image. Among them, a masked attention module in the target detection neural network realizes the fusion of spatial features and channel features. The masked attention module is divided into a spatial attention extraction branch and a channel attention extraction branch, which are respectively used to extract spatial features and channel features. The masked attention module first generates Query, Key, and Value through two convolutions. After the dimension transformation and matrix multiplication of Query and Key, channel features and spatial features are respectively obtained. The extracted channel features and spatial features are respectively filtered through the masking mechanism to retain the useful information for ship detection, and channel sparse attention and spatial sparse attention are obtained. The two kinds of attention will be respectively multiplied by Value through matrix multiplication to obtain sparse spatial features and sparse channel features. After splicing the outputs of the two branches and then passing through another convolution, the output is the feature refined by masking and attention. The present invention further improves the accuracy of ship detection.

[0047] A ship target detection method based on an attention mechanism and a masking mechanism according to the present invention includes the following steps: S1. Obtain a target image to be subjected to ship detection; Here, there can be multiple specific types of target images, such as images taken by an unmanned aerial vehicle at high altitude, or remote sensing images taken by a radar, etc. In such target images, the ship categories can be diverse.

[0048] S2. Use a preset target detection neural network based on a multi-scale fusion strategy to perform ship detection on the target image.

[0049] Among them, a masked attention module in the target detection neural network realizes the fusion of channel features and spatial features; see Figure 3 As shown, the masked attention module is divided into two branches of spatial features and channel features. The two branches first generate , , ; Then, and need to be respectively subjected to dimension transformation and then matrix multiplication to obtain and ; Then, through a masking selection mechanism, important information is screened out and useless information is filtered out to obtain sparse spatial attention and sparse channel attention; Then, the sparse spatial attention and the sparse channel attention are respectively multiplied by through matrix multiplication to obtain an enhanced spatial feature map and an enhanced channel feature map; Next, after the two feature maps are concatenated, they are fused through a single convolution to obtain the final output.

[0050] Among them, the mask selection mechanism has strong flexibility, and the threshold in the implementation process is set to be trainable, that is, the network decides which features to retain and which features to discard.

[0051] Exemplarily, the object detection neural network that implements masked attention based on the above SMAM includes: YOLOX, YOLOv7 or ResNet50, and of course it is not limited to this.

[0052] In order to make the layout of the specification clear, the following will give an example of the method for fusing channel features and spatial features in the object detection neural network based on the masked attention module.

[0053] In the above masked attention module, the ways to preprocess the input features include: ; Among them, represents the input feature, represents the convolution operator, and represent the learning parameters of the 1st to th convolutional kernels of the object detection neural network, , among which, and have convolutional kernel sizes of 1 and 3 respectively, represents the learning parameters of the s th convolutional kernel of the object detection neural network, represents the feature value corresponding to the s th convolutional kernel channel in the low-level features; represents the number of convolutional kernel channels of the input feature; represents convolution, represents the SiLU function, , , represent the preprocessed Query, Key and Value.

[0054] It can be understood that since the output of a single neuron in the neural network is generated by weighted summation of the outputs of all previous neurons, the mutual relationship between channels is in .

[0055] Here, in order to make full use of the cross-channel information of the input features without losing rich local information in the embodiments of the present invention, we adopt two different sizes of convolutional kernels when generating Query, Key, and Value. The first 1×1 convolution is used to extract cross-channel information, and the second depthwise separable 3×3 convolution is used to extract local information.

[0056] Exemplarily, the way to preprocess the input features in the above mask attention module may include: preprocessing the input features by using the layer normalization method, but of course it is not limited thereto.

[0057] In the above mask attention module, the way to extract spatial features includes: ; wherein, represents the extracted spatial features, so represents the eigenvalue of the corresponding h th convolutional kernel channel in the preprocessed channel features, represents the new Query obtained by dimension transformation of the preprocessed Query, where , represents the number of pixels of the input features, represents the new Key obtained by dimension transformation of the preprocessed Key, represents matrix multiplication.

[0058] In the above mask attention module, the way to extract channel features includes: ; wherein, represents the extracted channel features, so represents the eigenvalue of the corresponding h th convolutional kernel channel in the preprocessed channel features, represents the new Query obtained by dimension transformation of the preprocessed Query, represents the number of pixels of the input features, represents the new Key obtained by dimension transformation of the preprocessed Key, represents matrix multiplication.

[0059] It can be understood that in order to retain the useful information of the ship target to the greatest extent and filter out useless information such as the background at the same time, so we need to filter the and extracted.

[0060] In the above mask attention module, the process of mask screening for features includes:

[0061] Among them, represents the feature output after mask selection, which will be multiplied by the V matrix to obtain the enhanced features, namely and , is the element in the j-th column of the i-th row of the input matrix, and is the k-th largest value in the j-th row of the attention matrix and , where k is a learnable parameter.

[0062] In the above mask attention module, the process of weighting the extracted spatial features and channel features includes:

[0063] Among them, represents the final output feature after being weighted with Value after mask selection, and represent the new Value obtained by dimension transformation of the Value after preprocessing. Among them, represents the number of pixels of the input feature, represents matrix multiplication.

[0064] In the above mask attention module, the process of fusing the weighted spatial features and channel features includes:

[0065] and are the spatial features and channel features after mask selection, represents the concatenation operation, represents a 1×1 cross-channel convolution operation and a 3×3 depthwise separable convolution.

[0066] Next, an example is given of the method for fusing channel features and spatial features based on the mask selection module in the object detection neural network.

[0067] Exemplarily, Figure 4 shows a schematic structural diagram of a ResNet50 that realizes the fusion of channel features and spatial features based on the mask attention module; compared with Figure 1 the structure of the existing ResNet50 shown in, it can be seen that the mask attention module in the embodiment of the present invention further strengthens the semantic correlation of C5 and uses spatial information and channel information to select features useful for the target.

[0068] In addition, the mask attention module can also be used in the PAN module of ResNet50 to achieve the fusion of channel features and spatial features, and the usage is basically the same as that in FPN.

[0069] Similarly, for the method of implementing the fusion of channel features and spatial features based on SMAM in YOLOX or YOLOv7, reference can be made to the above embodiments of implementing the fusion of channel features and spatial features based on SMAM in ResNet50, and the embodiments of the present invention will not be elaborated herein.

[0070] In summary, in YOLOX, YOLOv7 or ResNet50, the module for implementing the fusion of channel features and spatial features based on SMAM can include FPN module and / or PAN module.

[0071] It should be noted that the application scenarios of the embodiments of the present invention are relatively wide, not limited to the feature input of one scale, and even have good performance when facing the feature input of other scales.

[0072] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, a method or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation manner, a complete software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, which can be collectively referred to as "circuit", "module" or "platform" here.

[0073] In another embodiment of the present invention, a ship target detection system based on an attention mechanism and a mask mechanism is provided. This system can be used to implement the above ship target detection method based on an attention mechanism and a mask mechanism. Specifically, the ship target detection system based on an attention mechanism and a mask mechanism includes a data module and a detection module.

[0074] Among them, the data module obtains the target image to be detected; The detection module uses a target detection neural network based on a multi-scale fusion strategy to perform ship target detection on the target image to be detected, so as to achieve ship target detection; In the target detection neural network, a mask attention module is used to achieve the fusion of spatial features and channel features; the mask attention module includes a spatial attention extraction branch and a channel attention extraction branch, which are respectively used to extract spatial features and channel features; The mask attention module first generates Query, Key and Value, and after the dimension-transformed Query and Key are subjected to matrix multiplication, channel features and spatial features are respectively obtained; The extracted channel features and spatial features are respectively filtered by a masking mechanism to retain the useful information for ship detection, obtaining channel sparse attention and spatial sparse attention; The channel sparse attention and spatial sparse attention are respectively multiplied by Value in matrix multiplication to obtain sparse spatial features and sparse channel features. After concatenating the outputs of the two branches, a convolution is performed again to output the features refined by masking and attention.

[0075] In another embodiment of the present invention, a terminal device is provided. The terminal device includes a processor and a memory. The memory is used to store a computer program. The computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of the ship target detection method based on the attention mechanism and the masking mechanism, including: Using an object detection neural network based on a multi-scale fusion strategy to perform ship target detection on the target image to be detected, realizing ship target detection; Among them, in the object detection neural network, a masked attention module is used to realize the fusion of spatial features and channel features; the masked attention module includes a spatial attention extraction branch and a channel attention extraction branch, which are respectively used to extract spatial features and channel features; the masked attention module first generates Query, Key, and Value. After matrix multiplication of the dimension-transformed Query and Key, channel features and spatial features are respectively obtained; the extracted channel features and spatial features are respectively filtered by a masking mechanism to retain the useful information for ship detection, obtaining channel sparse attention and spatial sparse attention; the channel sparse attention and spatial sparse attention are respectively multiplied by Value in matrix multiplication to obtain sparse spatial features and sparse channel features. After concatenating the outputs of the two branches, a convolution is performed again to output the features refined by masking and attention.

[0076] In another embodiment of the present invention, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a terminal device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. It can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. The computer-readable storage medium provides a storage space that stores the operating system of the terminal. And, in this storage space, there is also stored one or more instructions suitable for being loaded and executed by a processor. These instructions can be one or more computer programs (including program codes). It should be noted that more specific examples (non-exhaustive list) of the computer-readable storage medium here include: electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0077] The computer-readable storage medium also includes data signals propagated in a baseband or as part of a carrier wave, which carry readable program codes. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, device, or component. The program codes contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0078] The program codes for performing the operations of the present invention can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages - such as Java, C++, etc., and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program codes can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0079] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the ship target detection method based on the attention mechanism and the masking mechanism in the above embodiments; the one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps: Use a target detection neural network based on a multi-scale fusion strategy to perform ship target detection on a target image to be detected, and achieve ship target detection; Among them, in the target detection neural network, a masked attention module is used to implement the fusion of spatial features and channel features; the masked attention module includes a spatial attention extraction branch and a channel attention extraction branch, which are respectively used to extract spatial features and channel features; the masked attention module first generates Query, Key, and Value, and after matrix multiplication of the dimension-transformed Query and Key, channel features and spatial features are obtained respectively; the extracted channel features and spatial features are respectively filtered by the masking mechanism to retain useful information for ship detection, and channel sparse attention and spatial sparse attention are obtained; the channel sparse attention and spatial sparse attention are respectively multiplied by Value through matrix multiplication to obtain sparse spatial features and sparse channel features, and the outputs of the two branches are concatenated and then convolved once to output features refined by masking and attention.

[0080] Please refer to Figure 6 , the terminal device is a computer device. The computer device 60 in this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When the computer program 63 is executed by the processor 61, it implements the ship target detection method based on the attention mechanism and the masking mechanism in the embodiment. To avoid repetition, it will not be elaborated here one by one. Alternatively, when the computer program 63 is executed by the processor 61, it implements the functions of each model / unit in the ship target detection system based on the attention mechanism and the masking mechanism in the embodiment. To avoid repetition, it will not be elaborated here one by one.

[0081] The computer device 60 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art can understand that Figure 6 merely an example of the computer device 60, which does not constitute a limitation on the computer device 60, and may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.

[0082] The so-called processor 61 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, central processors, graphics processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, data processing logic units based on quantum computing, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0083] The memory 62 may be an internal storage unit of the computer device 60, such as the hard disk or memory of the computer device 60. The memory 62 may also be an external storage device of the computer device 60, such as a plug-in hard disk equipped on the computer device 60, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0084] Furthermore, the memory 62 may also include both the internal storage unit of the computer device 60 and the external storage device. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 may also be used to temporarily store data that has been output or is to be output.

[0085] In each of the embodiments provided in the present application, any reference to a memory, a database, or other media may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0086] In each of the embodiments provided in the present application, the database involved may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., without limitation. In each of the embodiments provided in the present application, the processor involved may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without limitation.

[0087] Please refer to Figure 7 , the terminal device is an electronic device 600, and the electronic device is presented in the form of a general computing device. The components of the electronic device may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.

[0088] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present invention described in the above method part of this specification. For example, the processing unit 610 can execute steps as shown in Figure 2 .

[0089] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.

[0090] The storage unit 620 may also include a program / utilities 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0091] The bus 630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0092] The electronic device 600 may also communicate with one or more external devices 700 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or may communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be through an input / output (I / O) interface 650. Also, the electronic device 600 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 660. The network adapter 660 may communicate with other modules of the electronic device 600 through the bus 630. It should be understood that, although not shown in the figures, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.

[0093] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the drawings here may be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present invention provided in the drawings below is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0094] The beneficial effects of the embodiments of the present invention are further described below by means of simulation experiments.

[0095] Experimental environment The experimental platform is PyTorch version 1.13.0, CUDA 11.6, the GPU (graphics processing unit) model is RTX4090, the hardware memory is 24G, and the compilation language is Python 3.8.

[0096] Experimental dataset The dataset used is the Seaships7000 dataset; this dataset selects 7000 images from coastal surveillance videos with a total of 6 types of targets, and the image resolution is 1920×1080. In this dataset, there are six types of ships, including: ore carrier, container ship, bulk cargo carrier, general cargo ship, fishing boat, and passenger ship. Since the images in the Seaships7000 dataset are relatively large in size, they are not suitable for direct detection. Therefore, they are cropped to 640×640.

[0097] Then, the data subset is split according to the ratio of training set: validation set: test set of 1:1:2. Among them, 1750 are used for training and validation, and 3500 are for testing. The label distribution of the training set is as Figure 5 shown, where the horizontal axis corresponds to ore carrier, container ship, bulk cargo carrier, general cargo ship, fishing boat, and passenger ship respectively.

[0098] Evaluation metrics To evaluate the effectiveness of the object detection method, accuracy, recall rate, average precision (AP), mean average precision (mAP), and model inference time are used to evaluate the algorithm.

[0099] Accuracy is the proportion of correct samples in the total test, and the calculation formula is:

[0100] Among them, true positive (TP) is the number of positive samples accurately predicted, and false positive (FP) is the number of negative samples misdetected as positive samples. The recall rate represents the proportion of accurately predicted positive samples, and its definition is:

[0101] Among them, the false negative (FN) is the number of positive samples predicted as negative samples.

[0102] AP is the area enclosed by the precision curve and the recall on the x-axis, and it is expressed as:

[0103] Among them, y is the recall curve under different intersection ratio thresholds, is the recall rate y corresponding precision.

[0104] mAP refers to the average value of AP for all classes, and the calculation formula is:

[0105] Among them, N is the number of classes, and N = 6 in the experiment.

[0106] Experimental process Based on YOLOX and YOLOv7, the ship target detection performance with and without using SMAM was compared. The experimental results on the test dataset are shown in the following table:

[0107] Among them, YoloX and Yolov7 respectively represent the existing YoloX and Yolov7 without using the mask attention module; YoloX-SMAM and Yolov7-SMAM respectively represent YoloX and Yolov7 using the mask attention module.

[0108] As can be seen from the above table, after using the mask attention module, the mAP on YOLOX increased from 96.1% to 97.0%, and the inference time increased from 20.4 ms for processing 32 batches of images to 23.6 ms, and the increased time consumption is almost negligible. After using the mask attention module, the mAP on YOLOv7 increased from 93.1% to 94.0%, and the inference speed increased from 13 ms for processing 32 batches of images to 14.2 ms, and the increased time consumption is almost negligible.

[0109] In summary, for the ship target detection method and system based on the attention mechanism and the masking mechanism of the present invention, through one cross-channel convolution and one depthwise separable convolution on the input features, a global receptive field is obtained while rich local information is retained; then, dimensional transformation and matrix multiplication are performed on the generated matrix to obtain spatial features and channel features, enabling the two branches to focus on different information respectively; then, the masking mechanism is used to screen the features. Specifically, the k-th largest value in the j-th column is selected as the threshold, and the features greater than the threshold, that is, the useful features, are retained, where k is adaptive and is selected by the model itself; then, the screened features are weighted with the Value matrix to obtain a new feature output, and the new feature output strengthens the spatial features and channel features useful for detection; finally, the new features output by the two branches are concatenated and subjected to one cross-channel convolution and one depthwise separable convolution for further fusion. Therefore, the relationship between channels and the correlation between spaces are considered in the object detection neural network used in the embodiments of the present invention. Therefore, ship target detection based on this object detection neural network can further improve the detection performance for ship targets. In particular, the embodiments of the present invention can preferably solve the problem of ship target detection in complex environments.

[0110] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0111] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0112] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present invention can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0113] In the embodiments provided by the present invention, it should be understood that the disclosed device / terminal and method can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0114] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0115] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0116] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0117] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the specified function in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0118] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the specified function in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the steps specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 in one block or a plurality of blocks.

[0120] The above is only to illustrate the technical idea of the present invention and should not be used to limit the protection scope of the present invention. Any modifications made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.

Claims

1. A ship target detection method based on attention mechanism and mask mechanism, characterized in that: The following steps are involved: The target detection neural network based on multi-scale fusion strategy is used to detect the target image to be detected, so as to realize the ship target detection; Among them, in the target detection neural network, a mask attention module is used to realize the fusion of spatial features and channel features; the mask attention module includes a spatial attention extraction branch and a channel attention extraction branch, which are used to extract spatial features and channel features respectively; The masked attention module first generates Query, Key and Value. The dimensionally transformed Query and Key are matrix multiplied to obtain channel features and spatial features respectively. The extracted channel features and spatial features are filtered by mask mechanism respectively, and the useful information for ship detection is retained to obtain channel sparse attention and spatial sparse attention. The channel sparse attention and spatial sparse attention are matrix multiplied with Value respectively to obtain sparse spatial features and sparse channel features. The outputs of the two branches are concatenated and then convolved once to output features refined by mask and attention.

2. The ship target detection method based on attention mechanism and mask mechanism according to claim 1 is characterized in that: The target detection neural network includes a masked attention module, which is used to extract channel features and spatial features in the target detection neural network and then select enhanced features through masks.

3. The ship target detection method based on attention mechanism and mask mechanism according to claim 1 is characterized in that: The masked attention module first generates Query, Key and Value as follows: in, represents the input features, , , represents the convolution operator, , , Represents Query matrix, Key matrix, and Value matrix.

4. The ship target detection method based on attention mechanism and mask mechanism according to claim 1 is characterized in that: After the dimension transformation, the query and key are matrix multiplied to obtain channel features and spatial features respectively. and as follows: in, and Represent the extracted channel features and spatial features respectively, Respectively and Transform the dimension to and .

5. The ship target detection method based on attention mechanism and mask mechanism according to claim 1 is characterized in that: The feature output of the extracted channel features and spatial features after filtering by the mask mechanism is: in, The input matrix is The element in row j and column j, is the attention matrix The kth largest value in the jth row, where k is a learnable parameter.

6. The ship target detection method based on attention mechanism and mask mechanism according to claim 1, characterized in that: The weighted spatial features and channel features are fused as follows: in, and are the spatial features and channel features after mask selection, Represents a splicing operation, It represents a 1×1 cross-channel convolution operation and a 3×3 depth-wise separable convolution.

7. The ship target detection method based on attention mechanism and mask mechanism according to claim 6 is characterized in that: The final output feature after mask selection and Value weighting as follows: in, Represents the final output feature after mask selection and Value weighting. and It indicates the new value obtained by dimension transformation after the value obtained after preprocessing. Represents matrix multiplication.

8. The ship target detection method based on attention mechanism and mask mechanism according to claim 1 is characterized in that: The target detection neural network uses a feature pyramid network FPN and / or a path aggregation network PAN.

9. The ship target detection method based on attention mechanism and mask mechanism according to claim 1 is characterized in that: The masked attention module implements YOLOX, YOLOv7 or ResNet50 that fused channel features and spatial features.

10. A ship target detection system based on attention mechanism and mask mechanism, characterized in that: include: Data module, obtaining the target image to be detected; The detection module uses a target detection neural network based on a multi-scale fusion strategy to perform ship target detection on the target image to be detected, thereby realizing ship target detection; In the target detection neural network, the mask attention module is used to achieve the fusion of spatial features and channel features; The mask attention module includes a spatial attention extraction branch and a channel attention extraction branch, which are used to extract spatial features and channel features respectively; The masked attention module first generates Query, Key and Value. The dimensionally transformed Query and Key are matrix multiplied to obtain channel features and spatial features respectively. The extracted channel features and spatial features are filtered by mask mechanism respectively, and the useful information for ship detection is retained to obtain channel sparse attention and spatial sparse attention. The channel sparse attention and spatial sparse attention are matrix multiplied with Value respectively to obtain sparse spatial features and sparse channel features. The outputs of the two branches are concatenated and then convolved once to output features refined by mask and attention.