CNN and GNN fusion-based network architecture and image detection method
By adopting a network architecture that integrates CNN and GNN in SAR image processing, combining sparse visual image attention blocks and maximum relative image convolution processing, the problems of complex background, noise interference and feature extraction imbalance in SAR images are solved, and more efficient ship detection accuracy and recall rate are achieved.
Patent Information
- Application Number
- CN202510344367.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-13
AI Technical Summary
The complex background in SAR images, a large number of spot noise interference, and global-local feature extraction imbalance, resulting in low ship detection accuracy.
Using a network architecture based on the fusion of CNN and GNN, local features are extracted through convolutional neural networks, and global information is captured using graph neural network encoder, combining sparse visual graph attention blocks and maximum relative graph convolution processing to achieve the balance between global features and local features.
It effectively suppresses speckle noise and interference from complex geographical environments, improves the accuracy and recall of ship detection, and significantly improves the ability to position and overlap target separation.
Smart Images

Figure CN120147753A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and object detection, and in particular to a network architecture and an image detection method based on the fusion of CNN and GNN. Background Art
[0002] The background of SAR images is complex and variable, which not only contains a large amount of speckle noise, but also contains a complex geographical environment, all of which bring strong interference to the accurate detection of ships and pose a severe challenge to accurately identifying ship targets in images. Especially in the near-shore scenario, ships are densely arranged, making it a difficult task to accurately distinguish the position and boundary of each ship.
[0003] Most current methods rely on Convolutional Neural Networks (CNNs) and Transformers. CNNs can accurately capture the local features of images, but due to the lack of long-distance dependencies and correlations between pixels and image patches, the network is vulnerable to interference from complex backgrounds and speckle noise in SAR images, limiting the overall performance. Transformers use many modules with high computational complexity, resulting in slow inference speed. Most networks have unbalanced global and local feature extraction for SAR images, leading to weakened object detection capabilities.
[0004] In view of this, there is an urgent need to propose a detection system and method that can solve problems such as complex backgrounds, a large amount of speckle noise interference, and unbalanced global-local feature extraction faced in ship detection of SAR images. Summary of the Invention
[0005] The purpose of the present invention is to provide a network architecture and an image detection method based on the fusion of CNN and GNN, which can solve problems such as complex backgrounds, a large amount of speckle noise interference, and unbalanced global-local feature extraction faced in ship detection of SAR images.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a network architecture based on the fusion of CNN and GNN, including:
[0008] A convolutional neural network that receives an input image and outputs N local feature maps after at least two rounds of convolution and local information extraction;
[0009] A graph neural network encoder that receives N local feature maps and performs one convolution, n 1 times of global information capture, two convolutions, n 2After the second global information capture, M multi-scale feature maps are output. The M multi-scale feature maps pass through the PAFPN feature fusion network and the detection head to output the final result. The graph neural network encoder is configured with a convolutional branch and a graph attention branch. The convolutional branch is used to retain the original information in the local feature maps. The graph attention branch, namely the GNN block, includes a convolutional trunk and a sparse visual graph attention block. The sparse visual graph attention block captures the global information in each local feature map based on the graph attention mechanism, rolls and aligns pixels along the row and column dimensions to calculate the relationship between pixels, calculates the difference between the rolled feature map and the pre-rolled feature map, and takes the maximum value of the difference as the relative feature. The maximum relative graph convolution is used to process the relative feature, and the convolution result of the relative feature and the input image feature is used as the output of the Grapher in the sparse visual graph attention block. At the same time, the FNN in the sparse visual graph attention block is applied to enhance the feature representation ability of the convolution result.
[0010] As a possible implementation, the convolutional neural network sequentially includes a Focus layer, a first convolutional kernel, a first CSP layer, a second convolutional kernel, and a second CSP layer;
[0011] Among them, the Focus layer receives the input image, performs slicing, channel splicing, and convolution processing on it, and then inputs it to the first convolutional kernel. After being compressed by the first convolutional kernel and the feature channels are expanded, it is input to the first CSP layer. After the first CSP layer extracts local information, it is input to the second convolutional kernel. After being compressed again by the second convolutional kernel and the feature channels are expanded, it is input to the second CSP layer. After the second CSP layer extracts local information again, N local feature maps are output.
[0012] As a possible implementation, both the first CSP layer and the second CSP layer extract local information from the feature maps based on parallel convolutional branches and residual structures; that is, both the first CSP layer and the second CSP layer include: a third convolutional kernel and a fourth convolutional kernel, a residual block that are parallel to the third convolutional kernel and serially connected to each other; the outputs of the parallel convolutional branches are jointly input to the first splicing block for splicing and then input to the fifth convolutional kernel.
[0013] As a possible implementation, the residual block includes two serially connected sixth convolutional kernels. The output of the fourth convolutional kernel is processed by the two serially connected sixth convolutional kernels and then added to the output of the fourth convolutional kernel in the Add block and then input to the second splicing block.
[0014] As a possible implementation, the graph neural network encoder includes a seventh convolutional kernel, an L-GNN layer, an eighth convolutional kernel, and an S-GNN layer;
[0015] Among them, the seventh convolutional kernel receives N local feature maps, adjusts their dimensions and then inputs them to the L-GNN layer, and after passing through n 1After the first global information capture, it is input into the eighth convolutional kernel for dimension adjustment again, and then input into the S-GNN layer. After n 2 times of global information capture, M multi-scale feature maps are output.
[0016] As a possible implementation, the architectures of the L-GNN layer and the S-GNN layer are the same, and both include: the ninth convolutional kernel, the Split block, the GNN block, the third splicing block, and the tenth convolutional kernel;
[0017] Among them, the output of the seventh convolutional kernel or the eighth convolutional kernel is adjusted by the ninth convolutional kernel to double the number of channels, and is evenly divided into two branches by the Split block. One branch is input into the GNN block for global information capture, and the other branch is input into the third splicing block. After the splicing process with the output of the GNN block in the third splicing block, it is input into the tenth convolutional kernel.
[0018] As a possible implementation, the GNN block at least includes a convolutional backbone and a sparse visual graph attention block; among them, the convolutional backbone is used for channel expansion; the sparse visual graph attention block includes a Grapher and an FFN; the GNN block captures the global information in the feature map in the following way:
[0019] S10. Initialize the parameters, and input the feature map X ∈ R C×H×W ; the parameters include the total number of repetitions N of the sparse visual graph attention block, the number of repetitions is n, the preset value n = 0, the rolling distance K,
[0020] S11. Receive the feature map X of one branch evenly divided by the Split block; adjust the number of channels of the feature map through the convolutional backbone (two-layer convolution) and then output X s ;
[0021] S12. X s After another convolution process, output X i ;
[0022] S13. Initialize the parameter m for controlling the loop to 0;
[0023] S14. Each row element of the feature map X i scrolls down by mK distance and aligns, calculates the difference between the rolled feature map and the feature map X i , and takes the maximum value of each element in the difference feature map and the feature map X i as the relative feature map X j ;
[0024] S15. Increase the parameter m for controlling the loop by 1;
[0025] S16. Determine whether the product of m and the rolling distance K in step S15 is less than the height H of the input feature map X. If so, repeat steps S13 - S16 until the judgment result is greater than or equal to, then execute S17;
[0026] S17. Reset m to 0;
[0027] S18. For each column element of the feature map X i roll to the right by mK distance and align, calculate the difference between the rolled feature map and the feature map X i take the maximum value of each element in the difference feature map and the relative feature map X j as the relative feature map X q ;
[0028] S19. Increase the parameter m that controls the loop by 1;
[0029] S20. Determine whether the product of m and the rolling distance K in step S19 is less than the width W of the input feature map X. If so, repeat steps S17 - S20 until the judgment result is greater than or equal to, then execute S21;
[0030] S21. Concatenate the relative feature map X q and the feature map X i and fuse them through a 1×1 convolution, and finally enhance the features through FFN;
[0031] S22. Increase the number of repetitions n by 1;
[0032] S23. Determine whether n in S22 is less than N + 1. If so, loop and execute steps S12 - S23; until the judgment result is false, then execute S25;
[0033] S25. Adjust the number of channels of the input feature map X through a convolutional layer and then end.
[0034] As a possible implementation, Grapher includes the eleventh convolutional kernel, MRGrapherCnov, and the twelfth convolutional kernel.
[0035] As a possible implementation, n 1 = 6, n 2 = 2.
[0036] In a second aspect, the present invention provides an image detection method, including the following steps:
[0037] The convolutional neural network receives the input image and outputs N local feature maps after at least two rounds of convolution and local information extraction;
[0038] The graph neural network encoder receives the N local feature maps and after one convolution, n 1 times of global information capture, secondary convolution, n2 After the second global information capture, M multi-scale feature maps are output;
[0039] The M multi-size feature maps are output through the aggregation detection head to obtain the detection results.
[0040] Beneficial effects:
[0041] A network architecture and an image detection method based on the fusion of CNN and GNN proposed by the present invention have the following beneficial effects compared with the prior art:
[0042] 1. The network architecture based on the fusion of CNN and GNN proposed by the present invention realizes the balance between global features and local features by extracting global context information and capturing the relationships between pixels and between targets, and can fully exploit the global and local features in the image, suppressing the influence of interference factors such as speckle noise and complex geographical environments;
[0043] 2. In the network architecture based on the fusion of CNN and GNN proposed by the present invention, a single GNN block calculates the relationships between rolling distance pixels, and by iteratively calculating the relationships between different pixels, the global information of the feature map is enhanced, thereby improving the ability of this network architecture to locate targets and separate overlapping targets;
[0044] 3. In the network architecture based on the fusion of CNN and GNN proposed by the present invention, the sparse visual graph attention block captures the global information in the feature map based on the graph attention mechanism, calculates the relationships between pixels by rolling and aligning pixels along the row and column dimensions, and then calculates the difference between the original input image and the rolled image and takes the maximum value as the relative feature; then, these features are processed by the maximum relative graph convolution, and the relative feature and the convolution result of the original feature map are used as the output of the Grapher; the FFN further enhances the representation ability of the features. Since the sparse visual graph attention block obtains the relationships between pixels in a fixed relationship calculation mode, avoiding the K-nearest neighbor (KNN) algorithm and the shape change operation of the feature map, the calculation is efficient and lightweight. Description of the drawings
[0045] The drawings described herein are used to provide a further understanding of the present invention, form a part of the present invention, and the schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention. In the drawings:
[0046] Figure 1 is a schematic structural diagram of the network architecture based on the fusion of CNN and GNN provided by the embodiment of the present invention;
[0047] Figure 2 is a schematic structural diagram of the first CSP layer and the second CSP layer in the network architecture provided by the embodiment of the present invention;
[0048] Figure 3 Schematic diagram of the structures of the L-GNN layer and the S-GNN layer in the network architecture provided by an embodiment of the present invention;
[0049] Figure 4 Schematic diagram of the structure of the GNN block in the network architecture provided by an embodiment of the present invention;
[0050] Figure 5 Flowchart of the method for the GNN block in the network architecture provided by an embodiment of the present invention to capture global information in the feature map;
[0051] Figure 6 Schematic diagram of the structure of the Grapher in the network architecture provided by an embodiment of the present invention;
[0052] Figure 7 Schematic diagram of the processing flow of the input image based on the network architecture structure of the fusion of CNN and GNN provided by an embodiment of the present invention;
[0053] Figure 8 Flowchart of the image detection method provided by an embodiment of the present invention;
[0054] Figure 9 Visualization result image obtained by comparing image detection using YOLO V8 and this solution in an embodiment of the present invention.
[0055] Reference numerals
[0056] 1 - Convolutional neural network, 10 - Focus layer, 11 - First convolutional kernel, 12 - First CSP layer, 120 - Third convolutional kernel, 121 - Fourth convolutional kernel, 122 - Residual block, 1220 - Sixth convolutional kernel, 13 - Second convolutional kernel, 14 - Second CSP layer, 15 - First splicing block, 16 - Fifth convolutional kernel, 17 - Second splicing block, 2 - Graph neural network encoder, 20 - Seventh convolutional kernel, 21 - L-GNN layer, 210 - Ninth convolutional kernel, 211 - Split block, 212 - GNN block, 2120 - Convolutional trunk, 2121 - Sparse visual graph attention block, 21210 - Grapher, 212100 - Eleventh convolutional kernel, 212101 - MR Grapher Cnov, 212102 - Twelfth convolutional kernel, 21211 - FFN, 213 - Third splicing block, 214 - Tenth convolutional kernel, 22 - Eighth convolutional kernel, 23 - S-GNN layer. Detailed implementation manners
[0057] In order to clearly describe the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. For example, the first threshold and the second threshold are only used to distinguish different thresholds, and do not limit their order. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily mean different.
[0058] It should be noted that in the present invention, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0059] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. The following at least one (item) or its similar expression refers to any combination of these items, including any combination of single (item) or plural items (items). For example, at least one (item) of a, b or c can represent: a, b, c, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b and c, where a, b and c can be single or multiple.
[0060] The embodiments of the present invention aim to provide a network architecture and an image detection method based on the fusion of CNN and GNN to solve problems such as complex backgrounds, a large amount of speckle noise interference, and unbalanced global-local feature extraction in SAR image ship detection. The specific implementation is as follows:
[0061] In a first aspect, the embodiments of the present invention provide a network architecture based on the fusion of CNN and GNN. See Figure 1 , including: a convolutional neural network 1 and a graph neural network encoder 2.
[0062] The convolutional neural network 1 receives the input image and outputs N local feature maps after at least two rounds of convolution and local information extraction; the graph neural network encoder 2 receives the N local feature maps, and after one convolution, n 1 times of global information capture, secondary convolution, n 2After the first global information capture, M multi-scale feature maps are output. The M multi-scale feature maps pass through the PAFPN feature fusion network and the detection head to output the final result. Both M and N are positive integers. The graph neural network encoder is configured with a convolutional branch and a graph attention branch. The convolutional branch is used to retain the original information in the local feature map. The graph attention branch, i.e., the GNN block, includes a convolutional trunk and a sparse visual graph attention block. The sparse visual graph attention block captures the global information in each local feature map based on the graph attention mechanism, rolls and aligns pixels along the row and column dimensions to calculate the relationship between pixels, calculates the difference between the rolled feature map and the unrolled feature map, and takes the maximum value of the difference as the relative feature. The maximum relative graph convolution is used to process the relative feature, and the convolution result of the relative feature and the input image feature is used as the output of the Grapher in the sparse visual graph attention block. At the same time, the FNN in the sparse visual graph attention block is applied to enhance the feature representation ability of the convolution result.
[0063] See Figure 1 , As a possible implementation, the convolutional neural network 1 sequentially includes a Focus layer 10, a first convolutional kernel 11, a first CSP layer 12, a second convolutional kernel 13, and a second CSP layer 14;
[0064] Among them, the Focus layer 10 receives the input image, performs slicing, channel splicing, and convolution processing on it, and then inputs it to the first convolutional kernel 11. After being compressed by the first convolutional kernel 11 and the feature channels are expanded, it is input to the first CSP layer 12. After the first CSP layer 12 extracts local information, it is input to the second convolutional kernel 13. After being compressed again by the second convolutional kernel 13 and the feature channels are expanded, it is input to the second CSP layer 14. After the second CSP layer 14 extracts local information again, N local feature maps are output.
[0065] As an example, both the first convolutional kernel 11 and the second convolutional kernel 13 are 3×3 convolutional kernels.
[0066] See Figure 2 , As a possible implementation, both the first CSP layer 12 and the second CSP layer 14 extract local information from the feature map based on parallel convolutional branches and residual structures; that is, both the first CSP layer 12 and the second CSP layer 14 include: a third convolutional kernel 120 and a fourth convolutional kernel 121 and a residual block 122 that are parallel to the third convolutional kernel 120 and serially connected to each other; the outputs of the parallel convolutional branches are jointly input to the first splicing block 15 to complete splicing and then input to the fifth convolutional kernel 16.
[0067] As an example, the third convolutional kernel 120, the fourth convolutional kernel 121, and the fifth convolutional kernel 16 are all 1×1 convolutional kernels.
[0068] See Figure 2, as a possible implementation, the residual block 122 includes two serially connected sixth convolutional kernels 1220. The output of the fourth convolutional kernel 121 is processed by the two serially connected sixth convolutional kernels 1220 and then added to the output of the fourth convolutional kernel 121 in the Add block and then input to the second splicing block 17.
[0069] As an example, the two serially connected sixth convolutional kernels 1220 are both 1×1 convolutional kernels.
[0070] See Figure 1 , as a possible implementation, the graph neural network encoder 2 includes a seventh convolutional kernel 20, an L-GNN layer 21, an eighth convolutional kernel 22, and an S-GNN layer 23; wherein, the seventh convolutional kernel 20 receives N local feature maps, adjusts their dimensions and then inputs them to the L-GNN layer 21. After n 1 times of global information capture, it is input to the eighth convolutional kernel 22 to adjust the dimensions again and then input to the S-GNN layer 23. After n 2 times of global information capture, M multi-scale feature maps are output.
[0071] As a possible implementation, n 1 = 6, n 2 = 2.
[0072] As an example, both the seventh convolutional kernel 20 and the eighth convolutional kernel 22 are 3×3 convolutional kernels.
[0073] See Figure 3 , as a possible implementation, the architectures of the L-GNN layer 21 and the S-GNN layer 23 are the same, and both include: a ninth convolutional kernel 210, a Split block 211, a GNN block 212, a third splicing block 213, and a tenth convolutional kernel 214;
[0074] Among them, the output of the seventh convolutional kernel 20 or the eighth convolutional kernel 22 is adjusted by the ninth convolutional kernel 210 to double the number of channels, and is evenly divided into two branches by the Split block 211. One branch is input to the GNN block 212 for global information capture, and the other branch is input to the third splicing block 213. In the third splicing block 213, it is spliced with the output of the GNN block 212 and then input to the tenth convolutional kernel 214.
[0075] As an example, both the ninth convolutional kernel 210 and the tenth convolutional kernel 214 are 1×1 convolutional kernels.
[0076] See Figure 4, As a possible implementation, the GNN block 212 at least includes a convolutional trunk 2120 and a sparse visual graph attention block 2121; among them, the convolutional trunk 2120 is used for channel expansion; the sparse visual graph attention block 2121 includes a Grapher 21210 and an FFN 21211.
[0077] See Figure 5 , The GNN block captures the global information in the feature map in the following way:
[0078] S10. Initialize the parameters, and input the feature map X ∈ R C×H×W ; The parameters include the total number of repetitions N of the sparse visual graph attention block, the number of repetitions is n, the preset value n = 0, the rolling distance K,
[0079] S11. Receive the feature map X of a branch evenly divided by the Split block; adjust the number of channels of the feature map through the convolutional trunk (two layers of convolution) and then output X s ;
[0080] S12. X s Output X after being processed by convolution again i ;
[0081] S13. Initialize the parameter m for controlling the loop to 0;
[0082] S14. Each row element of the feature map X i Scrolls down by mK distance and aligns, calculates the difference between the rolled feature map and the feature map X i , and takes the maximum value of each element in the difference feature map and the feature map X i as the relative feature map X j ;
[0083] S15. Increase the parameter m for controlling the loop by 1;
[0084] S16. Judge whether the product of m and the rolling distance K in step S15 is less than the height H of the input feature map X. If so, repeat steps S13 - S16 until the judgment result is greater than or equal to, then execute S17;
[0085] S17. Reset m to 0;
[0086] S18. Each column element of the feature map X i Scrolls to the right by mK distance and aligns, calculates the difference between the rolled feature map and the feature map X i , and takes the maximum value of each element in the difference feature map and the relative feature map X j as the relative feature map X q ;
[0087] S19. Increase the parameter m for controlling the loop by 1;
[0088] S20. Determine whether the product of m and the rolling distance K in step S19 is less than the W of the input feature map X. If so, repeat S17 - S20 until the judgment result is greater than or equal to, then execute S21;
[0089] S21. For the relative feature map X q and the feature map X i concatenate and fuse through a 1×1 convolution, and finally enhance the features through FFN;
[0090] S22. Increase the repetition count n by 1;
[0091] S23. Determine whether n in S22 is less than N + 1. If so, loop and execute S12 - S23; until the judgment result is no, then execute S25;
[0092] S25. End after adjusting the number of channels of the input feature map X through a convolutional layer.
[0093] The sparse visual graph attention block captures the global information in the feature map based on the graph attention mechanism, calculates the relationship between pixels by rolling and aligning pixels along the row and column dimensions, then calculates the difference between the original input image and the rolled image and takes the maximum value as the relative feature. Subsequently, the maximum relative graph convolution is used to process these features, and the convolution result of the relative feature and the original feature map is used as the output of the Grapher. FFN further enhances the representation ability of the features. Since the sparse visual graph attention block obtains the relationship between pixels in a fixed relationship calculation mode, it avoids the K-nearest neighbor (KNN) algorithm and the shape change operation of the feature map, making the calculation efficient and lightweight.
[0094] A single GNN block calculates the relationship between pixels at the rolling distance. By iteratively calculating the relationships between different pixels, the global information of the feature map is enhanced, thereby improving the ability of this network architecture to locate targets and separate overlapping targets.
[0095] See Figure 6 , as a possible implementation, Grapher21210 includes the eleventh convolution kernel 212100, MRGrapherCnov212101, and the twelfth convolution kernel 212102.
[0096] As an example, both the eleventh convolution kernel 212100 and the twelfth convolution kernel 212102 are 1×1 convolution kernels.
[0097] See Figure 7, as an example, the convolutional neural network 1 receives an input image. The input image first enters the Focus layer. After the Focus layer slices, splices channels, and performs convolution processing on it, it enters a 3×3 convolutional kernel. After the convolutional kernel compresses it and expands the feature channels, it enters the CSP layer to extract local information, and then enters another 3×3 convolutional kernel. After the 3×3 convolutional kernel compresses it and expands the feature channels again, it enters another CSP layer. After the CSP layer extracts local information again, N local feature maps are output.
[0098] The graph neural network encoder 2 receives N local feature maps. The N local feature maps first enter a 3×3 convolutional kernel. After dimension adjustment, they are input into multiple L-GNN layers. The multiple L-GNN layers perform multiple global information captures on them, and then enter another 3×3 convolutional kernel for dimension adjustment. After the dimension adjustment is completed, they enter multiple S-GNN layers. The multiple S-GNN layers perform multiple global information captures on them and then output M multi-size feature maps.
[0099] Among them, the numbers of L-GNN layers and S-GNN layers are different. There are 6 L-GNN layers and 2 S-GNN layers. However, the structures of the L-GNN layers and S-GNN layers are the same, and both include two 1×1 convolutional kernels, a Split block, a GNN block, and a splicing block. Specifically, the N local feature maps after dimension adjustment by the 3×3 convolution double their number of channels through the first 1×1 convolutional kernel, and then are evenly divided into two branches by the Split block. One branch enters the GNN block for global information capture and then reaches the splicing block, and the other branch directly enters the splicing block. The two complete splicing processing in the splicing block and then enter another 1×1 convolutional kernel. Among them, the GNN block includes a convolutional trunk, a sparse visual graph attention block, and a 1×1 convolutional kernel, and its purpose is to capture global information of the received image. Let X represent the feature map entering the GNN block, then X first adjusts the number of channels through the convolutional trunk. The convolutional trunk includes two layers of 3×3 convolutional kernels, and after the channel adjustment, it becomes X s , then enters the sparse visual graph attention block to iteratively capture the global information in the feature map multiple times, and then enters the 1×1 convolutional kernel to adjust the number of channels and output.
[0100] The M multi-size feature maps output by the graph neural network encoder 2 enter the aggregation detection head SPP, and the detection result is output after aggregation.
[0101] In a second aspect, an embodiment of the present invention provides an image detection method. Refer to Figure 8 , including the following steps:
[0102] After receiving the input image, the convolutional neural network outputs N local feature maps after at least two rounds of convolution and local information extraction;
[0103] After receiving N local feature maps, the graph neural network encoder performs one convolution, n 1 times of global information capture, two convolutions, and n 2 times of global information capture, and then outputs M multi-scale feature maps;
[0104] The M multi-size feature maps are output through the aggregation detection head to obtain the detection results.
[0105] In this embodiment, under the background of rotated bounding box ship detection, the algorithm effect is verified based on the RSSDD dataset. The network architecture of the present invention is compared with YOLO V8, and the comparison results are shown in Table 1:
[0106] Table 1 Verification of algorithm effect based on RSSDD dataset
[0107] Accuracy Recall <![CDATA[mAP 50-95 > <![CDATA[mAP 50 > <![CDATA[mAP 75 > YOLO V8 0.965 0.950 0.765 0.986 0.922 CG-Encoder 0.969 0.971 0.776 0.990 0.928
[0108] It can be seen that compared with YOLO V8, the network architecture proposed by the present invention improves each index. The accuracy is improved by 0.4%, and the recall rate is improved by 2.1%, reaching 96.9% and 97.1%. The accuracy rate and recall rate are more balanced. mAP 50-95 、mAP 50 、mAP 75 are improved by 1.1%, 0.4% and 0.6% respectively, indicating that the network architecture proposed by the present invention can better perceive the global and local features in the image, effectively improving the accuracy of ship detection and making the target positioning more accurate.
[0109] For the visualization effect, see Figure 9 . The red bounding boxes are the detection results, the yellow bounding boxes are the missed detections, and the false alarms are in the blue ellipses. From the above results, it can be seen that under the influence of complex geographical environment and speckle noise, YOLO V8 has a small number of missed detections and false alarms, while the CG-Encoder Network of the present application can accurately detect all ships in the image, proving the effectiveness of this method.
[0110] Although the present invention has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed invention, those skilled in the art can understand and realize other variations of the disclosed embodiments by viewing the drawings, the disclosure content, and the accompanying drawings. In the specification, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude the case of multiple. A single processor or other unit can implement several functions listed in the specification. Certain measures are recited in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0111] Although the present invention has been described in connection with specific features and their embodiments, it will be apparent that various modifications and combinations can be made without departing from the spirit and scope of the invention. Accordingly, this specification and the accompanying drawings are merely illustrative of the invention and are considered to cover any and all modifications, variations, combinations or equivalents within the scope of the invention. Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A network architecture based on the fusion of CNN and GNN, characterized in that: Including convolutional neural network and graph neural network encoder; Convolutional neural network, receives input image, and outputs N local feature maps after at least two rounds of convolution and local information extraction; The graph neural network encoder receives N local feature maps, and after one convolution, n1 global information capture, two convolutions, and n2 global information capture, it outputs M multi-scale feature maps. The M multi-scale feature maps pass through the PAFPN feature fusion network and the detection head to output the final result. Among them, the graph neural network encoder is configured with a convolution branch and a graph attention branch, wherein the convolution branch is used to retain the original information in the local feature map, and the graph attention branch is the GNN block, including a convolution stem and a sparse visual graph attention block. The sparse visual graph attention block captures the global information in each local feature map based on the graph attention mechanism, scrolls and aligns pixels along the row and column dimensions to calculate the relationship between pixels, calculates the difference between the feature map after scrolling and the feature map before scrolling, and takes the maximum value of the difference as the relative feature; the maximum relative graph convolution is used to process the relative feature, and the convolution result of the relative feature and the input image feature is used as the output of the Grapher in the sparse visual graph attention block, and the FNN in the sparse visual graph attention block is applied to enhance the feature representation ability of the convolution result.
2. The network architecture based on CNN and GNN fusion according to claim 1, characterized in that: The convolutional neural network includes the Focus layer, the first convolution kernel, the first CSP layer, the second convolution kernel and the second CSP layer in sequence; Among them, the Focus layer receives the input image, slices it, performs channel splicing and convolution processing on it, and then inputs it into the first convolution kernel. After being compressed by the first convolution kernel and expanding the feature channel, it is input to the first CSP layer. After the local information is extracted by the first CSP layer, it is input to the second convolution kernel. After being compressed again by the second convolution kernel and expanding the feature channel, it is input to the second CSP layer. After the local information is extracted again by the second CSP layer, N local feature maps are output.
3. The network architecture based on CNN and GNN fusion according to claim 2, characterized in that: The first CSP layer and the second CSP layer both extract local information from the feature map based on parallel convolution branches and residual structures; that is, the first CSP layer and the second CSP layer both include: a third convolution kernel and a fourth convolution kernel parallel to the third convolution kernel and serially connected to each other, and a residual block; the outputs of the parallel convolution branches are jointly input to the first splicing block and then input to the fifth convolution kernel after splicing.
4. The network architecture based on CNN and GNN fusion according to claim 3, characterized in that: The residual block includes two serial sixth convolution kernels. The output of the fourth convolution kernel is processed by the two serial sixth convolution kernels, added to the output of the fourth convolution kernel in the Add block, and then input into the second splicing block.
5. The network architecture based on CNN and GNN fusion according to claim 1, characterized in that: The graph neural network encoder includes the seventh convolution kernel, the L-GNN layer, the eighth convolution kernel and the S-GNN layer; Among them, the seventh convolution kernel receives N local feature maps, adjusts their dimensions and inputs them into the L-GNN layer, and then inputs them into the eighth convolution kernel after n1 times of global information capture, and then adjusts their dimensions again and inputs them into the S-GNN layer, and then outputs M multi-size feature maps after n2 times of global information capture.
6. The network architecture based on CNN and GNN fusion according to claim 5, characterized in that: The architecture of the L-GNN layer is the same as that of the S-GNN layer, including the ninth convolution kernel, the Split block, the GNN block, the third concatenation block, and the tenth convolution kernel. Among them, the output of the seventh convolution kernel or the eighth convolution kernel is adjusted by the ninth convolution kernel to double the number of channels, and is evenly divided into two branches by the Split block. One branch is input to the GNN block for global information capture, and the other branch is input to the third splicing block. After the splicing process is completed with the output of the GNN block in the third splicing block, it is input to the tenth convolution kernel.
7. The network architecture based on CNN and GNN fusion according to claim 6, characterized in that: The GNN block at least includes a convolutional stem and a sparse visual graph attention block; wherein the convolutional stem is used for channel expansion; the sparse visual graph attention block includes a Grapher and FFN; the GNN block captures the global information in the feature map in the following way: S10. Initialize parameters and input feature map X∈R C×H×W ; The parameters include the total number of repetitions of the sparse visual graph attention block N, the number of repetitions is n, the default value n = 0, the scrolling distance K, S11. Receive a feature map X of a branch averaged by the Split block; adjust the number of channels of the feature map through convolution and output X s ; S12.X s After convolution again, the output is X i ; S13. Initialize the parameter m of the control loop to 0; S14. Feature map X i Each row of elements is scrolled down by mK distance and aligned, and the feature map after scrolling is calculated and the feature map X i The difference between the difference feature map and the feature map X i The maximum value of each element in is taken as the relative feature map X j ; S15. The parameter m of the control loop is increased by 1; S16. Determine whether the product of m and the scrolling distance K in step S15 is less than the height H of the input feature map X. If so, repeat S13 to S16 until the judgment result is greater than or equal to S17; S17. Reset m to 0; S18. Feature map X i Each column element of is scrolled to the right by mK distance and aligned, and the feature map after scrolling is calculated and the feature map X i The difference between the difference feature map and the relative feature map X j The maximum value of each element in is taken as the relative feature map X q ; S19. The parameter m of the control loop is increased by 1; S20. Determine whether the product of m and the scroll distance K in step S19 is less than W of the input feature map X. If so, repeat S17 to S20 until the judgment result is greater than or equal to S21; S21. Relative feature map X q With feature map X i The concatenation is then fused through 1×1 convolution, and finally the features are enhanced through FFN; S22. The number of repetitions n increases by 1; S23. Determine whether n in S22 is less than N+1. If so, execute S12 to S23 in a loop; until the result of the judgment is negative, execute S25; S25. The process ends after adjusting the number of channels of the input feature map X through the convolution layer.
8. The network architecture based on CNN and GNN fusion according to claim 7, characterized in that: The Grapher includes an eleventh convolution kernel, MRGrapherCnov and a twelfth convolution kernel.
9. The network architecture based on CNN and GNN fusion according to claim 5, characterized in that: n1=6, n2=2.
10. An image detection method, characterized in that: The steps include: After receiving the input image, the convolutional neural network performs at least two rounds of convolution and local information extraction and then outputs N local feature maps; The graph neural network encoder receives N local feature maps, performs one convolution, n1 global information capture, two convolutions, and n2 global information capture, and then outputs M multi-scale feature maps. M multi-scale feature maps are output by the aggregation detection head to obtain the detection results.
Citation Information
Cited By
Tobacco field contour extraction method based on graph neural network
CN120852798A
A tobacco field contour extraction method based on a graph neural network
CN120852798B