Marine Life Detection Method Based on Bidirectional Cooperative Guidance Network of CNN and Transformer

By adopting a two-way collaborative guidance network based on CNN and Transformer in marine biological detection, the accuracy of marine biological detection in complex underwater environments is solved, and higher detection accuracy and morphological integrity are achieved.

CN116152650BActive Publication Date: 2025-06-17NINGBO UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211553252.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-06
Publication Date
2025-06-17
Estimated Expiration
2042-12-06

AI Technical Summary

Technical Problem

The prior art is difficult to accurately detect marine organisms in complex underwater environments, and is affected by color bias, uneven light, low contrast, low visibility and camouflage characteristics of marine organisms.

Method used

A two-way collaborative guidance network based on CNN and Transformer is adopted to extract high-level semantic information of images through the backbone network, and spatial and texture information are extracted in combination with the Transformer network. The global feature enhancement module and the dual-branch cross-level progressive feature fusion module are enhanced to enhance the fusion of feature information and features at different scales.

Benefits of technology

It significantly improves the accuracy and morphological integrity of marine biological detection, and can accurately segment binary images of marine biological organisms in complex underwater environments, improving the integrity and accuracy of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116152650B_ABST
    Figure CN116152650B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting marine organisms based on a bidirectional collaborative guidance network of CNN and Transformer. A deep neural network is built, which consists of a backbone network, a global feature enhancement module for enhancing the features extracted from the backbone network, and a dual-branch cross-level progressive feature fusion module for fusing features of different scales in a progressive manner, as the bidirectional collaborative guidance network based on CNN and Transformer. The deep neural network is trained using an extended training set to obtain a trained model of the deep neural network. Each image in the test set is tested using the trained model of the deep neural network, and a detected image of camouflaged marine organisms corresponding to each image in the test set is obtained. The advantage is that it can accurately and efficiently detect marine organisms underwater.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a target detection technology, and particularly to a method for detecting marine organisms based on a two-way collaborative guidance network of CNN (Convolutional Neural Network) and Transformer. Background Art

[0002] With the continuous development of technologies such as deep learning, image processing, and digital imaging, computer vision technology can simulate human vision to solve practical problems in the directions of target detection, video surveillance, polyp segmentation, lung infection area segmentation, industrial defect detection, pest detection, culture and art, etc. Among them, target detection in practical applications has always been emphasized by researchers and has made great progress in recent years. Underwater marine organism detection in target detection occupies a very important position in the development of the marine economy.

[0003] The marine economy is playing an increasingly important role, and underwater marine organism detection has also received more and more attention. With the development of artificial intelligence technology, underwater marine organism detection has been applied to various research topics related to marine organisms, such as fish recognition, marine animal segmentation, etc. However, due to (1) underwater images will inevitably be affected by the absorption and scattering of wavelengths, such as forward scattering and backward scattering. Therefore, the original underwater images often have problems such as color bias, uneven illumination, low contrast, and low visibility, which seriously affect various applications of underwater images, such as underwater marine organism detection; (2) in the underwater environment, it often happens that part of the information of an object is blocked by other objects, such as seahorses being blocked by seagrass, clownfish being blocked by corals, etc. Therefore, when performing underwater marine organism detection, it is difficult to clearly find the complete detection target; (3) the turbid underwater environment makes it difficult to accurately locate the position of the detection target, resulting in difficult underwater marine organism detection; (4) marine organisms have camouflage characteristics. In order to survive, marine organisms have developed rich camouflage abilities. They try to change their appearance to perfectly blend into the surrounding environment and avoid the attention of other organisms. Therefore, performing underwater marine organism detection has become a very challenging task. How to effectively remove the color bias of underwater images, how to effectively improve the visibility of underwater images, and how to improve the accuracy of underwater marine organism detection are all research directions worthy of study.

[0004] Underwater target detection can be roughly divided into three categories, namely underwater general target detection, underwater salient target detection, and underwater camouflaged target detection. Underwater general target detection is classification plus detection, locating the position of the target and giving the category; underwater salient target detection is finding the most prominent object in a given image; underwater camouflaged target detection is finding the object hidden in a given image.

[0005] Underwater marine life detection aims to identify marine life in a given image, which has received extensive attention in recent years for underwater salient object detection. Moreover, due to the complex underwater environment, underwater marine life often exhibits camouflage characteristics, so underwater camouflaged object detection is more suitable for underwater marine life detection. Underwater marine life detection, as a new field, has attracted much attention and plays a crucial role in fields such as underwater navigation and marine military, because it can provide important information for identifying marine life in a complex marine environment. However, there are currently few detection methods for underwater marine life. The existing underwater marine life detection methods mainly focus on saliency detection for natural datasets and camouflage detection for natural datasets. However, the complex underwater environment limits the performance of these methods for natural dataset detection; when light travels underwater, there is uneven attenuation, and in a harsh underwater environment, when the imaging device acquires an underwater image, the quality of the underwater image will be affected, such as problems like color deviation, low contrast, low light, uneven illumination, etc., and the camouflage characteristics of marine life further limit the performance of these methods for natural dataset detection. Therefore, it is particularly important to design a detection method for underwater marine life. After research, there are currently relatively large datasets, mainly the marine animal dataset mainly based on MAS3K and the natural dataset of camouflaged objects mainly based on COD10K. Among them, the MAS3K dataset mainly contains 1769 training images, 1141 test images, and a total of 37 species of marine animals; the COD10K dataset includes 5066 camouflaged pictures, 3000 background pictures, and 1934 non-camouflaged pictures, with more than 78 species, including 20 species of marine animals. The introduction of these datasets has greatly promoted the development of camouflaged object detection and marine life detection.

[0006] In recent years, the camouflage target detection method combining digital image processing and deep learning mainly focuses on the research of relatively obvious features such as the texture structure and color of the detection target. Traditional camouflage target detection methods mainly rely on low-level features, such as texture, geometric features, simple linear iterative clustering superpixels, etc. However, the detection performance of these methods is relatively poor and the generalization ability is not good. With the development of deep learning technology, remarkable progress has been made in camouflage target detection. In particular, the emergence of the U-Net network has received extensive attention due to its ability to reconstruct high-resolution prediction results using multi-level features. Recently, the SINet architecture uses a localization and recognition method to first preliminarily locate the position of the camouflage target and then perform fine-grained segmentation. The C2F-Net architecture proposes a context-based perception model that effectively fuses information at different levels. In addition, in recent years, researchers have successively proposed some underwater target detection methods, but basically they are mainly for underwater general target detection, and they all consider from the perspective of image enhancement such as the color and contrast of underwater images. There are few methods specifically for underwater marine life detection. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a marine life detection method based on a two-way collaborative guidance network of CNN and Transformer, which can accurately and efficiently detect underwater marine life.

[0008] The technical solution adopted by the present invention to solve the above technical problem is: a marine life detection method based on a two-way collaborative guidance network of CNN and Transformer, which is characterized by including the following steps:

[0009] Step 1: Select or construct a data set, which contains two types of original images, namely marine life images and non-marine life images; then preprocess each original image in the data set so that the size of the preprocessed image is H'×W'×3, and the mean value of the pixel values of all pixel points in the R channel of the preprocessed image is 0.485 and the variance is 0.229, the mean value of the pixel values of all pixel points in the G channel is 0.456 and the variance is 0.224, and the mean value of the pixel values of all pixel points in the B channel is 0.406 and the variance is 0.225; then divide all the preprocessed images into a training set and a test set, and both the training set and the test set contain two types of images, namely marine life images and non-marine life images;

[0010] Step 2: Build a deep neural network using a deep learning framework. The deep neural network consists of a backbone network, a global feature enhancement module for enhancing the features extracted from the backbone network, and a dual-branch cross-level progressive feature fusion module for fusing features of different scales in a progressive manner, thereby forming a bidirectional collaborative guidance network based on CNN and Transformer. The backbone network includes a CNN-based Resnet50-backbone network for extracting high-level semantic information of the image and a Transformer-backbone network for extracting the spatial and texture information of the image and supplementing the extracted spatial and texture information to the feature map captured by the Resnet50-backbone network. The Resnet50-backbone network has a total of five sequentially connected layers, and the Transformer-backbone network has a total of two sequentially connected layers. The Transformer-backbone network is composed of the first and second layers in the PVT network based on Transformer. The input ends of the first layer of the Resnet50-backbone network and the first layer of the Transformer-backbone network simultaneously receive an image with a size of H×W×3. In the Resnet50-backbone network, the input end of the second layer receives the feature map R1 with a size of output from the output end of the first layer, the input end of the third layer receives the feature map R2 with a size of output from the output end of the second layer, the input end of the fourth layer receives the feature map R3 with a size of output from the output end of the third layer, the input end of the fifth layer receives the feature map R4 with a size of output from the output end of the fourth layer, and the output end of the fifth layer outputs a feature map R5 with a size of In the Transformer-backbone network, the input end of the second layer receives the feature map T1 with a size of output from the output end of the first layer, and the output end of the second layer outputs a feature map T2 with a size of The first input end of the global feature enhancement module receives R4, the second input end receives R5, and the third input end receives T2. The global feature enhancement module generates a feature map with a size of The feature map P1' is such that at the output end of the global feature enhancement module, after performing an eight-fold upsampling operation on P1', a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 1 is passed through to obtain the feature map P1 of size H×W×1 for auxiliary training; the first input end of the dual-branch cross-level progressive feature fusion module receives T1, the second input end receives T2, and the third input end receives P1'. The dual-branch cross-level progressive feature fusion module generates a feature map P2' of size At the output end of the dual-branch cross-level progressive feature fusion module, after performing a four-fold upsampling operation on P2', a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 1 is passed through to obtain the feature map P2 of size H×W×1 for the training loss function;

[0011] Step 3: Based on the size of each image in the training set, each image in the training set is respectively shrunk by 0.75 times and enlarged by 1.25 times to expand the training set, forming an extended training set. The size of each image in the extended training set is H×W×3; where, H = H', 0.75H', 1.25H', and W = W', 0.75W', 1.25W';

[0012] Step 4: Use the extended training set to perform network training on the deep neural network built in Step 2. After each round of network training, the deep neural network outputs the feature map P1 for auxiliary training and the feature map P2 for the training loss function corresponding to each image in the extended training set. Then, the loss function Loss is calculated. Loss = L main +L aux , L main =L wbce (P2, GT)+L iou (P2, GT), L aux =L wbce (P1, GT)+L iou (P1, GT); where, L main represents the main loss function, L aux represents the auxiliary loss function, L wbce () represents the binary cross-entropy loss function, L iou () represents the weighted intersection over union loss function, GT represents the true label. The network training is implemented using Pytorch. The optimizer adopts AdaXW, the batch size is set to 24, the initial learning rate is set to 1e - 4, and the learning rate decays by 10 times every 30 epochs;

[0013] Step 5: Perform network training for a total of 150 epochs according to the process in Step 4 to obtain the deep neural network training model;

[0014] Step 6: Use the trained model of the deep neural network to test each image in the test set, and detect the corresponding camouflaged marine creature detection image P2 for each image in the test set.

[0015] In the said Step 1, the process of preprocessing an original image is as follows: First, scale the size of the original image to H'×W'×3; Second, perform normalization processing on the scaled image, and normalize the pixel values of all pixel points in the R channel of the scaled image to a mean of 0.485 and a variance of 0.229, the pixel values of all pixel points in the G channel to a mean of 0.456 and a variance of 0.224, and the pixel values of all pixel points in the B channel to a mean of 0.406 and a variance of 0.225.

[0016] In the said Step 2, the global feature enhancement module is a dual-path parallel structure, and its processing process is as follows:

[0017] The first path: Perform a two-fold upsampling operation on R4 to obtain a feature map R with a size of 41 ; Pass R 41 through a convolutional layer with a convolutional kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 1024, and number of output channels of 64. This convolutional layer reduces the number of channels of R 41 to 64 dimensions, and the output end of this convolutional layer outputs a feature map R with a size of 42 ; Pass R 42 through a Relu activation function to obtain a feature map R with a size of 43 ; Perform feature fusion on R 43 and T2 by means of channel splicing to obtain a feature map f1' with a size of ; Pass f1' through a convolutional layer with a convolutional kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 192, and number of output channels of 64. This convolutional layer reduces the number of channels of f1' to 64 dimensions, and the output end of this convolutional layer outputs a feature map f1'1 with a size of ; Pass f1'1 through a Relu activation function to obtain a feature map f1'2 with a size of ; Pass f1'2 through a feature enhancement module for feature enhancement, and the output end of this feature enhancement module outputs a feature map f1'3 with a size of ; Pass f1'3 through a convolutional layer with a convolutional kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 3. The output end of this convolutional layer outputs a size of The feature map f1; Input f1 into the Transf module based on Transformer, and the output end of the Transf module based on Transformer outputs a feature map with a size of The feature map P1'1'; Pass P1'1' through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 6, and number of output channels of 64. The output end of this convolutional layer outputs a feature map with a size of The feature map P1'1;

[0018] The second path: Perform a two-fold upsampling operation on R5 to obtain a feature map with a size of R 51 ; Through the method of channel splicing, perform feature fusion on R 51 And R4 to obtain a feature map with a size of R 52 ; Pass R 52 Through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 3072, and number of output channels of 64. This convolutional layer reduces the number of channels of R 52 To 64 dimensions, and the output end of this convolutional layer outputs a feature map with a size of R 53 ; Pass R 53 Through a Relu activation function to obtain a feature map with a size of R 54 ; Perform a two-fold upsampling operation on R 54 To obtain a feature map with a size of R 55 ; Pass R 55 Through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 64. The output end of this convolutional layer outputs a feature map with a size of R 56 ; Pass R 56 Through a Relu activation function to obtain a feature map with a size of R 57 ; Through the method of channel splicing, perform feature fusion on R 57 And T2 to obtain a feature map f2'; Pass f2' through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 192, and number of output channels of 64. This convolutional layer reduces the number of channels of f2' to 64 dimensions, and the output end of this convolutional layer outputs a feature map with a size of f2'; Pass f2'1 through a Relu activation function to obtain a feature map with a size of f2'1; Pass f2'1 through a Relu activation function to obtain a feature map with a size of The feature map f2'2; enhancing the features of f2'2 through a feature enhancement module, and the output end of the feature enhancement module outputs a feature map with a size of The feature map f2'3; passing f2'3 through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 3, and the output end of the convolutional layer outputs a feature map with a size of The feature map f2; inputting f2 into the Transf module based on Transformer, and the output end of the Transf module based on Transformer outputs a feature map with a size of The feature map P1'2'; passing P1'2' through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 6, and number of output channels of 64, and the output end of the convolutional layer outputs a feature map with a size of The feature map P1'2;

[0019] Performing feature fusion on P1'1 and P1'2 by feature addition to obtain a feature map γ with a size of Passing γ through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 64, and the output end of the convolutional layer outputs a feature map with a size of The feature map P1'; performing eight-fold upsampling operation on P1' to obtain a feature map with a size of H×W×64 Passing through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 1, and the output end of the convolutional layer outputs a feature map P1 for auxiliary training with a size of H×W×1.

[0020] The structures of the two-way feature enhancement modules are the same, and it includes five branches. For the first branch: passing the input feature map s through a convolutional layer with a kernel size of 1×1, padding of 0, stride size of 1, number of input channels of 64, and number of output channels of 64, and the output end of the convolutional layer outputs a feature map with a size of The feature map s 11 ; passing s 11 through a convolutional layer with a kernel size of 1×3, padding of (0,1), number of input channels of 64, and number of output channels of 64, and the output end of the convolutional layer outputs a feature map with a size of The feature map s 12 ; passing s 12 through a convolutional layer with a kernel size of 3×1, padding of (1,0), number of input channels of 64, and number of output channels of 64, and the output end of the convolutional layer outputs a feature map with a size of The feature map s13 ; Pass s 13 through a dilated convolutional layer with a dilation rate of 3, padding of 3, stride size of 3, number of input channels of 64, and number of output channels of 64. The output size at the output end of this dilated convolutional layer is feature map s 14 ; For the second branch: Pass the input feature map s through a convolutional layer with a kernel size of 1×1, padding of 0, stride size of 1, number of input channels of 64, and number of output channels of 64. The output size at the output end of this convolutional layer is feature map s 21 ; Pass s 21 through a convolutional layer with a kernel size of 1×5, padding of (0,2), number of input channels of 64, and number of output channels of 64. The output size at the output end of this convolutional layer is feature map s 22 ; Pass s 22 through a convolutional layer with a kernel size of 5×1, padding of (2,0), number of input channels of 64, and number of output channels of 64. The output size at the output end of this convolutional layer is feature map s 23 ; Pass s 23 through a dilated convolutional layer with a dilation rate of 5, padding of 5, stride size of 3, number of input channels of 64, and number of output channels of 64. The output size at the output end of this dilated convolutional layer is feature map s 24 ; For the third branch: Pass the input feature map s through a convolutional layer with a kernel size of 1×1, padding of 0, stride size of 1, number of input channels of 64, and number of output channels of 64. The output size at the output end of this convolutional layer is feature map s 31 ; Pass s 31 through a convolutional layer with a kernel size of 1×7, padding of (0,3), number of input channels of 64, and number of output channels of 64. The output size at the output end of this convolutional layer is feature map s 32 ; Pass s 32 through a convolutional layer with a kernel size of 7×1, padding of (3,0), number of input channels of 64, and number of output channels of 64. The output size at the output end of this convolutional layer is feature map s 33 ; Pass s 33 through a dilated convolutional layer with a dilation rate of 7, padding of 7, stride size of 3, number of input channels of 64, and number of output channels of 64. The output size at the output end of this dilated convolutional layer is feature map s 34; For the 4th branch, pass the input feature map s through a convolutional layer with a kernel size of 1×1, padding of 0, stride size of 1, number of input channels of 64, and number of output channels of 64. The output of this convolutional layer has a size of of the feature map s 41 ; For the 5th branch, pass the input feature map s through a convolutional layer with a kernel size of 1×1, padding of 0, stride size of 1, number of input channels of 64, and number of output channels of 64. The output of this convolutional layer has a size of of the feature map s 51 ; Through the method of channel concatenation, perform feature fusion on s 14 , s 24 , s 34 , s 41 to obtain a feature map s with a size of 1234 ; Pass s 1234 through a convolutional layer with a kernel size of 1×1, stride size of 1, number of input channels of 256, and number of output channels of 64. The output of this convolutional layer has a size of of the feature map s1' 234 ; Through the method of feature addition, perform feature fusion on s1' 234 and s 51 to obtain a feature map with a size of where s is f1'2 or f2'2. If s is f1'2, then is f1'3. If s is f2'2, then is f2'3.

[0021] The Transf module based on Transformer is composed of the first layer in the PVT network based on Transformer. The size of the input feature map is The size of the output feature map is

[0022] In the step 2 described above, the dual-branch cross-level progressive feature fusion module has a dual-path parallel structure. Each path has a camouflaged target recognition module and a cross-level feature fusion module. The processing process is as follows:

[0023] The first path: Pass T1 through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 64. The output of this convolutional layer has a size of of the feature map T1”; Pass T1” through the camouflaged target recognition module to remove the noise information irrelevant to the camouflaged target on the feature map. The output of this camouflaged target recognition module has a size of ​​The feature map T1'; Perform cross-level feature fusion on T1' and P1' through a cross-layer feature fusion module, and the output end of the cross-layer feature fusion module outputs a feature map with a size of The feature map P2'1;

[0024] The second path: Pass T2 through a convolutional layer with a kernel size of 3×3, a padding of 1, a stride size of 1, an input channel number of 128, and an output channel number of 64. This convolutional layer reduces the channel number of T2 to 64 dimensions, and the output end of this convolutional layer outputs a feature map with a size of The feature map T2”; Pass T2” through a camouflaged target recognition module to remove noise information unrelated to the camouflaged target on the feature map. The output end of the camouflaged target recognition module outputs a feature map with a size of The feature map T2'; The size of the feature map obtained after performing a two-fold upsampling operation on T2' is The feature map and P1' are subjected to cross-level feature fusion through a cross-layer feature fusion module, and the output end of the cross-layer feature fusion module outputs a feature map with a size of The feature map P2'2;

[0025] Perform feature fusion on P' 21 and P' 22 through feature addition to obtain a feature map β with a size of Pass β through a convolutional layer with a kernel size of 3×3, a padding of 1, a stride size of 1, an input channel number of 64, and an output channel number of 64. The output end of this convolutional layer outputs a feature map with a size of The feature map P2'; Perform a four-fold upsampling operation on P2' to obtain a feature map with a size of H×W×64 Pass through a convolutional layer with a kernel size of 3×3, a padding of 1, a stride size of 1, an input channel number of 64, and an output channel number of 1. The output end of this convolutional layer outputs a feature map P2 with a size of H×W×1 for the training loss function.

[0026] The structures of the camouflaged target recognition modules in the two paths are the same. It is mainly composed of a channel attention module and a spatial attention module connected in series. The processing process of the channel attention module is as follows: Pass T i ” through a global average pooling layer, and the output end of the global average pooling layer outputs the feature map T” i,1 ; Pass T” i,1 through a convolutional layer with a kernel size of 1×1, an input channel number of 64, and an output channel number of 4. The output end of this convolutional layer outputs the feature map T” i,2 ; Pass T” i,2 through a Relu activation function to obtain the feature map T” i,3 ; Pass T”i,3 Through a convolutional layer with a convolutional kernel size of 1×1, 4 input channels, and 64 output channels, the output end of this convolutional layer outputs a feature map T". i,4 ; Pass T i " through a global max pooling layer, and the output end of this global max pooling layer outputs a feature map T". i,5 ; Pass T" i,5 through a convolutional layer with a convolutional kernel size of 1×1, 64 input channels, and 4 output channels, and the output end of this convolutional layer outputs a feature map T". i,6 ; Pass T" i,6 through a Relu activation function to obtain a feature map T". i,7 ; Pass T" i,7 through a convolutional layer with a convolutional kernel size of 1×1, 4 input channels, and 64 output channels, and the output end of this convolutional layer outputs a feature map T". i,8 ; Perform feature fusion on T" i,4 and T" i,8 by feature addition to obtain a feature map T". i,9 ; Pass T" i,9 through a sigmoid function to obtain a feature map T". i,10 ; Perform feature fusion on T" i,10 and T i " by feature multiplication to obtain a feature map T". i,11 ; The processing process of the spatial attention module is as follows: Calculate the average value and maximum value of " i,11 " by channel, and correspondingly obtain an average value feature map T". i,12 and a maximum value feature map T". i,13 ; Perform feature fusion on T" i,12 and T" i,13 by channel concatenation to obtain a feature map T". i,14 ; Pass T" i,14 through a convolutional layer with a convolutional kernel size of 7×7, padding of 3, stride size of 1, 128 input channels, and 64 output channels, and the output end of this convolutional layer outputs a feature map T". i,15 ; Perform feature fusion on T" i,15 and T" i,11 by feature multiplication to obtain a feature map T i '; where i = 1, 2.

[0027] The structures of the two cross-layer feature fusion modules are the same.

[0028] Compared with the prior art, the advantages of the present invention are as follows:

[0029] 1) The method of the present invention uses a Resnet50-backbone network based on CNN to extract semantic information from the input image, and a PVT network based on Transformer to extract spatial and texture information from the input image. It fully utilizes the advantages of CNN in image feature extraction and the global information modeling advantage of Transformer, promotes and learns from each other between the two, so as to fully aggregate texture features and semantic information, and aggregates feature information at different scales in a two-way collaborative guidance manner. CNN ensures the extraction of local information and can easily locate the information of the area to be detected. The global information modeling ability of Transformer supplements the boundaries and content of the semantic information extracted by CNN, so that the finally detected binary image has more complete content and boundary information, that is, the morphology is more complete.

[0030] 2) The method of the present invention uses dual-path encoding of CNN and Transformer and uses the global information modeling ability of Transformer to decode the encoding results, greatly improving the recognition ability of marine organisms and the morphological integrity of the finally detected binary image.

[0031] 3) The method of the present invention uses a global feature enhancement module GFEM to perform global information modeling on the information output by the backbone network, thereby enhancing the feature information extracted from the backbone network and significantly enhancing the detection results of marine organisms.

[0032] 4) The method of the present invention uses a dual-branch cross-level progressive feature fusion module DPFFM to perform progressive fusion of different levels on the output features of the backbone network, effectively solving the problems brought by feature fusion between different scales and greatly improving the detection results of camouflaged marine organisms.

[0033] 5) The method of the present invention adopts a design of two-way collaborative guidance based on CNN and Transformer, and adopts a two-stage design of "localization-identification", which can accurately segment the binary image of marine organisms in a relatively complex underwater environment. The test results show that even when the object to be detected has low light, turbidity, occlusion, and marine organisms exhibit camouflage characteristics, the method of the present invention can still accurately segment marine organisms. Brief Description of the Drawings

[0034] Figure 1 is the overall flow block diagram of the method of the present invention;

[0035] Figure 2 is the overall network framework diagram of the two-way collaborative guidance network based on CNN and Transformer constructed by the method of the present invention;

[0036] Figure 3 The network structure diagram of the two-way collaborative guidance network based on CNN and Transformer constructed by the method of the present invention;

[0037] Figure 4 The diagram showing the detection effects of some data of the method of the present invention and the existing method on the MAS3K and COD10K datasets;

[0038] Figure 5 The comparison results of the ablation experiment. Detailed implementation manners

[0039] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments.

[0040] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments of the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0041] A marine organism detection method based on a two-way collaborative guidance network of CNN and Transformer proposed by the present invention, the overall implementation block diagram of which is as Figure 1 shown, and it includes the following steps:

[0042] Step 1: Select or construct a dataset. To increase the diversity of samples and improve the generalization ability of the model, this dataset contains two types of original images, namely marine biological images and non-marine biological images. Then, preprocess each original image in the dataset so that the size of the preprocessed image is H'×W'×3, and the mean of the pixel values of all pixel points in the R channel of the preprocessed image is 0.485 and the variance is 0.229, the mean of the pixel values of all pixel points in the G channel is 0.456 and the variance is 0.224, and the mean of the pixel values of all pixel points in the B channel is 0.406 and the variance is 0.225. Then, divide all the preprocessed images into a training set and a test set, and both the training set and the test set contain two types of images, namely marine biological images and non-marine biological images. Among them, H' = W' = 352.

[0043] In a specific embodiment, in Step 1, the process of preprocessing an original image is as follows: First, use the existing technology to scale the size of the original image to H'×W'×3. Second, perform normalization processing on the scaled image, and normalize the pixel values of all pixel points in the R channel of the scaled image to a mean of 0.485 and a variance of 0.229, the pixel values of all pixel points in the G channel to a mean of 0.456 and a variance of 0.224, and the pixel values of all pixel points in the B channel to a mean of 0.406 and a variance of 0.225. Since the input of the neural network usually has a fixed form, it is necessary to preprocess the original image. Here, the normalization processing only normalizes the original image, and no normalization operation is performed on the label corresponding to the original image.

[0044] Step 2: Build a deep neural network using a deep learning framework, such as Figure 2 and Figure 3As shown, the deep neural network consists of a backbone network, a global feature enhancement module GFEM for enhancing the features extracted from the backbone network, and a dual-branch cross-level progressive feature fusion module DPFFM for fusing features of different scales in a progressive manner, thus forming a bidirectional collaborative guidance network based on CNN and Transformer; the backbone network includes a Resnet50-backbone network based on CNN (Convolutional Neural Network) for extracting high-level semantic information of the image and a Transformer-backbone network for extracting spatial and texture information of the image and supplementing the extracted spatial and texture information to the feature map captured by the Resnet50-backbone network. The Resnet50-backbone network has a total of five sequentially connected layers, and the Transformer-backbone network has a total of two sequentially connected layers. The Transformer-backbone network is composed of the first and second layers in the Pyramid Vision Transformer (PVT) network (with a total of five layers) based on Transformer; the input ends of the first layer of the Resnet50-backbone network and the first layer of the Transformer-backbone network simultaneously receive an image with a size of H×W×3. In the Resnet50-backbone network, the input end of the second layer receives the feature map R1 with a size of output from the output end of the first layer, the input end of the third layer receives the feature map R2 with a size of output from the output end of the second layer, the input end of the fourth layer receives the feature map R3 with a size of output from the output end of the third layer, the input end of the fifth layer receives the feature map R4 with a size of output from the output end of the fourth layer, and the output end of the fifth layer outputs a feature map R5 with a size of In the Transformer-backbone network, the input end of the second layer receives the feature map T1 with a size of output from the output end of the first layer, and the output end of the second layer outputs a feature map T2 with a size of The first input end of the global feature enhancement module receives R4, the second input end receives R5, and the third input end receives T2. The global feature enhancement module generates a feature map with a size of The feature map P1' is upsampled by a factor of eight at the output end of the global feature enhancement module and then passed through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, input channel number of 64, and output channel number of 1 to obtain the feature map P1 of size H×W×1 for auxiliary training; the first input end of the dual-branch cross-level progressive feature fusion module receives T1, the second input end receives T2, and the third input end receives P1'. The dual-branch cross-level progressive feature fusion module generates a feature map P2' of size The feature map P2' is upsampled by a factor of four at the output end of the dual-branch cross-level progressive feature fusion module and then passed through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, input channel number of 64, and output channel number of 1 to obtain the feature map P2 of size H×W×1 for the training loss function; among them, the Resnet50-backbone network based on CNN (Convolutional Neural Network) is disclosed in K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.2 (Deep Residual Learning for Image Recognition), and the PVT (Pyramid Vision Transformer) network based on Transformer is disclosed in “Pvt v2: Improved baselines with pyramid vision transformer”, Computational Visual Media, vol. 8, no. 3, pp. 415–424, 2022. III-B (Improved Baseline Model Based on Pyramid Vision Transformer).

[0045] Here, the backbone network, the global feature enhancement module, and the dual-branch cross-level progressive feature fusion module are all dual-path parallel structures. Since the feature maps obtained from the first three layers of the Resnet50-backbone network in the backbone network contain relatively little camouflage information to be extracted, they are discarded and not used. Instead, the feature maps obtained from the last two layers of the Resnet50-backbone network are directly used for fusion with the spatial texture information extracted by the Transformer-backbone network.

[0046] Step 3: Based on the size of each image in the training set, each image in the training set is respectively reduced by 0.75 times and enlarged by 1.25 times to expand the training set and form an extended training set. The size of each image in the extended training set is H×W×3; where, H = H', 0.75H', 1.25H', and W = W', 0.75W', 1.25W'; During training, multi-scale training is adopted. For the input image, based on H'×W' = 352×352, scaling with different scales of 0.75 times and 1.25 times is performed.

[0047] Step 4: Use the extended training set to perform network training on the deep neural network constructed in Step 2. After each round of network training, the deep neural network outputs the feature map P1 for auxiliary training and the feature map P2 for the training loss function corresponding to each image in the extended training set. Then, calculate the loss function Loss, Loss = L main +L aux ,L main =L wbce (P2, GT)+L iou (P2, GT), L aux =L wbce (P1, GT)+L iou (P1, GT); where, L main represents the main loss function, L aux represents the auxiliary loss function, L wbce () represents the binary cross-entropy loss function, L iou () represents the weighted intersection over union loss function, Loss is the mixed loss function, which can effectively evaluate the differences between images from multiple perspectives of pixels, regions, and the whole, and effectively alleviate the negative impact on the segmentation performance caused by different object sizes. The main loss function L main and the auxiliary loss function L aux both use the binary cross-entropy loss function L wbce () and the weighted intersection over union loss function L iou () for calculation. L iou () is used to focus on the global structure information and form global information constraints on the network. L wbce () is used to calculate the loss of each pixel and form pixel constraints on the network. Both P1 and P2 are binary black-and-white images with a resolution of H×W, GT represents the true label, the network training is implemented using Pytorch, the optimizer adopts AdaXW, the batch size is set to 24, the initial learning rate is set to 1e-4, and the learning rate decays by 10 times every 30 epochs.

[0048] Step 5: Perform network training for a total of 150 epochs according to the process in Step 4 to obtain the deep neural network training model.

[0049] Step 6: Use the trained model of the deep neural network to test each image in the test set, and detect the camouflaged marine creature detection image corresponding to each image in the test set, namely P2.

[0050] In a specific embodiment, in step 2, the global feature enhancement module has a dual-path parallel structure, and its processing process is as follows:

[0051] The first path: Perform a two-fold upsampling operation on R4 to obtain a feature map R with a size of 41 ; Pass R 41 through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 1024, and number of output channels of 64. This convolutional layer reduces the number of channels of R 41 to 64 dimensions, and the output end of this convolutional layer outputs a feature map R with a size of 42 ; Pass R 42 through a Relu activation function to obtain a feature map R with a size of 43 ; Perform feature fusion on R 43 and T2 by means of channel splicing to obtain a feature map f1' with a size of ; Pass f1' through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 192, and number of output channels of 64. This convolutional layer reduces the number of channels of f1' to 64 dimensions, and the output end of this convolutional layer outputs a feature map f1'1 with a size of ; Pass f1'1 through a Relu activation function to obtain a feature map f1'2 with a size of ; Pass f1'2 through a feature enhancement module FEM for feature enhancement, and the output end of this feature enhancement module outputs a feature map f1'3 with a size of ; Pass f1'3 through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 3. The output end of this convolutional layer outputs a feature map f1 with a size of ; Input f1 into the Transf module based on Transformer. The output end of this Transf module based on Transformer outputs a feature map P1'1' with a size of ; Pass P1'1' through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 6, and number of output channels of 64. The output end of this convolutional layer outputs a feature map P1'1 with a size of 。

[0052] The second path: perform a two-fold upsampling operation on R5 to obtain a feature map R with a size of ; perform feature fusion on R 51 and R4 through channel concatenation to obtain a feature map R with a size of 51 ; pass R through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 3072, and number of output channels of 64. This convolutional layer reduces the number of channels of R 52 to 64 dimensions, and the output end of this convolutional layer outputs a feature map R with a size of 52 ; pass R 52 through a Relu activation function to obtain a feature map R with a size of ; perform a two-fold upsampling operation on R 53 to obtain a feature map R with a size of 53 ; pass R through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 64. The output end of this convolutional layer outputs a feature map R with a size of 54 ; pass R 54 through a Relu activation function to obtain a feature map R with a size of ; perform feature fusion on R 55 and T2 through channel concatenation to obtain a feature map f2'; pass f2' through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 192, and number of output channels of 64. This convolutional layer reduces the number of channels of f2' to 64 dimensions, and the output end of this convolutional layer outputs a feature map f2'1 with a size of 55 ; pass f2'1 through a Relu activation function to obtain a feature map f2'2 with a size of ; perform feature enhancement on f2'2 through a feature enhancement module FEM. The output end of this feature enhancement module outputs a feature map f2'3 with a size of 56 ; pass f2'3 through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 3. The output end of this convolutional layer outputs a feature map f2 with a size of 56 ; input f2 into the Transf module based on Transformer. The output end of this Transf module based on Transformer outputs a size of ; perform feature fusion on R 57 and T2 through channel concatenation to obtain a feature map f2'; pass f2' through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 192, and number of output channels of 64. This convolutional layer reduces the number of channels of f2' to 64 dimensions, and the output end of this convolutional layer outputs a feature map f2'1 with a size of 57 ; pass f2'1 through a Relu activation function to obtain a feature map f2'2 with a size of ; pass f2'2 through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 192, and number of output channels of 64. This convolutional layer reduces the number of channels of f2' to 64 dimensions, and the output end of this convolutional layer outputs a feature map f2'1 with a size of ; pass f2'1 through a Relu activation function to obtain a feature map f2'2 with a size of ; perform feature enhancement on f2'2 through a feature enhancement module FEM. The output end of this feature enhancement module outputs a feature map f2'3 with a size of ; pass f2'3 through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 3. The output end of this convolutional layer outputs a feature map f2 with a size of ; input f2 into the Transf module based on Transformer. The output end of this Transf module based on Transformer outputs a size of Feature map P1'2'; Pass P1'2' through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 6, and number of output channels of 64. The output at the output end of this convolutional layer has a size of Feature map P1'2.

[0053] Perform feature fusion on P1'1 and P1'2 by feature addition to obtain a feature map γ with a size of Pass γ through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 64. The output at the output end of this convolutional layer has a size of Feature map P1'; Perform eight-fold upsampling on P1' to obtain a feature map with a size of H×W×64 Pass through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 1. The output at the output end of this convolutional layer is a feature map P1 of H×W×1 for auxiliary training.

[0054] Further limit that the structures of the two-path feature enhancement modules are the same. It includes five branches. For the first branch: Pass the input feature map s through a convolutional layer with a kernel size of 1×1, padding of 0, stride size of 1, number of input channels of 64, and number of output channels of 64. The output at the output end of this convolutional layer has a size of Feature map s 11 ; Pass s 11 through a convolutional layer with a kernel size of 1×3, padding of (0,1), number of input channels of 64, and number of output channels of 64. The output at the output end of this convolutional layer has a size of Feature map s 12 ; Pass s 12 through a convolutional layer with a kernel size of 3×1, padding of (1,0), number of input channels of 64, and number of output channels of 64. The output at the output end of this convolutional layer has a size of Feature map s 13 ; Pass s 13 through an atrous convolutional layer with an atrous rate of 3, padding of 3, stride size of 3, number of input channels of 64, and number of output channels of 64. The output at the output end of this atrous convolutional layer has a size of Feature map s 14 ; For the second branch: Pass the input feature map s through a convolutional layer with a kernel size of 1×1, padding of 0, stride size of 1, number of input channels of 64, and number of output channels of 64. The output at the output end of this convolutional layer has a size of Feature map s21 ; Pass s 21 through a convolutional layer with a kernel size of 1×5, padding of (0, 2), 64 input channels, and 64 output channels. The output size at the output end of this convolutional layer is feature map s 22 ; Pass s 22 through a convolutional layer with a kernel size of 5×1, padding of (2, 0), 64 input channels, and 64 output channels. The output size at the output end of this convolutional layer is feature map s 23 ; Pass s 23 through an atrous convolutional layer with an atrous rate of 5, padding of 5, stride size of 3, 64 input channels, and 64 output channels. The output size at the output end of this atrous convolutional layer is feature map s 24 ; For the third branch: Pass the input feature map s through a convolutional layer with a kernel size of 1×1, padding of 0, stride size of 1, 64 input channels, and 64 output channels. The output size at the output end of this convolutional layer is feature map s 31 ; Pass s 31 through a convolutional layer with a kernel size of 1×7, padding of (0, 3), 64 input channels, and 64 output channels. The output size at the output end of this convolutional layer is feature map s 32 ; Pass s 32 through a convolutional layer with a kernel size of 7×1, padding of (3, 0), 64 input channels, and 64 output channels. The output size at the output end of this convolutional layer is feature map s 33 ; Pass s 33 through an atrous convolutional layer with an atrous rate of 7, padding of 7, stride size of 3, 64 input channels, and 64 output channels. The output size at the output end of this atrous convolutional layer is feature map s 34 ; For the fourth branch, pass the input feature map s through a convolutional layer with a kernel size of 1×1, padding of 0, stride size of 1, 64 input channels, and 64 output channels. The output size at the output end of this convolutional layer is feature map s 41 ; For the fifth branch, pass the input feature map s through a convolutional layer with a kernel size of 1×1, padding of 0, stride size of 1, 64 input channels, and 64 output channels. The output size at the output end of this convolutional layer is feature map s 51; Through the way of channel splicing for s 14 s 24 s 34 s 41 Perform feature fusion to obtain a feature map s with a size of 1234 ; Pass s 1234 through a convolutional layer with a kernel size of 1×1, a stride size of 1, an input channel number of 256, and an output channel number of 64. The output end of this convolutional layer outputs a feature map s1' with a size of 234 ; Through the way of feature addition, perform feature fusion on s1' 234 and s 51 to obtain a feature map with a size of where s is f1'2 or f2'2. If s is f1'2, then is f1'3. If s is f2'2, then is f2'3.

[0055] Further limit that the Transf module based on Transformer is composed of the first layer in the PyramidVision Transformer (PVT) network (with a total of five layers) based on Transformer. The size of the input feature map is and the size of the output feature map is

[0056] In a specific embodiment, in step 2, the dual-branch cross-level progressive feature fusion module is a dual-path parallel structure. Each path has a camouflaged target recognition module COTRM and a cross-level feature fusion module CFFM. The processing process is as follows:

[0057] The first path: Pass T1 through a convolutional layer with a kernel size of 3×3, a padding of 1, a stride size of 1, an input channel number of 64, and an output channel number of 64. The output end of this convolutional layer outputs a feature map T1” with a size of ; Pass T1” through the camouflaged target recognition module to remove the noise information irrelevant to the camouflaged target on the feature map. The output end of this camouflaged target recognition module outputs a feature map T1' with a size of ; Perform cross-level feature fusion on T1' and P1' through the cross-level feature fusion module. The output end of this cross-level feature fusion module outputs a feature map P2'1 with a size of

[0058] ​​The second path: Pass T2 through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 128, and number of output channels of 64. This convolutional layer reduces the number of channels of T2 to 64 dimensions, and the output size at the output end of this convolutional layer is of the feature map T2"; Pass T2" through the camouflage target recognition module to remove the noise information unrelated to the camouflage target on the feature map. The output size at the output end of this camouflage target recognition module is of the feature map T2'; After performing a two-fold upsampling operation on T2', the resulting feature map with a size of and P1' are subjected to cross-level feature fusion through the cross-layer feature fusion module. The output size at the output end of this cross-layer feature fusion module is of the feature map P2'2.

[0059] Perform feature fusion on P2'1 and P2'2 by means of feature addition to obtain a feature map β with a size of Pass β through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 64. The output size at the output end of this convolutional layer is of the feature map P2'; Perform a four-fold upsampling operation on P2' to obtain a feature map with a size of H×W×64 Pass through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 1. The output size at the output end of this convolutional layer is the feature map P2 of H×W×1 for training the loss function.

[0060] Further limit that the structures of the camouflage target recognition modules of the two paths are the same, which is mainly composed of a channel attention module and a spatial attention module connected in series. The processing process of the channel attention module is as follows: Pass T i " through a global average pooling layer, and the output end of this global average pooling layer outputs the feature map T" i,1 ; Pass T" i,1 through a convolutional layer with a kernel size of 1×1, number of input channels of 64, and number of output channels of 4. The output end of this convolutional layer outputs the feature map T" i,2 ; Pass T" i,2 through a Relu activation function to obtain the feature map T" i,3 ; Pass T" i,3 through a convolutional layer with a kernel size of 1×1, number of input channels of 4, and number of output channels of 64. The output end of this convolutional layer outputs the feature map T" i,4 ; Pass T i " through a global max pooling layer, and the output end of this global max pooling layer outputs the feature map T"i,5 ; The "T" i,5 passes through a convolutional layer with a kernel size of 1×1, 64 input channels, and 4 output channels. The output end of this convolutional layer outputs the feature map "T" i,6 ; The "T" i,6 passes through a Relu activation function to obtain the feature map "T" i,7 ; The "T" i,7 passes through a convolutional layer with a kernel size of 1×1, 4 input channels, and 64 output channels. The output end of this convolutional layer outputs the feature map "T" i,8 ; The feature map "T" i,4 and the feature map "T" i,8 are fused through feature addition to obtain the feature map "T" i,9 ; The "T" i,9 passes through a sigmoid function to obtain the feature map "T" i,10 ; The feature map "T" i,10 and the feature map "T" i " are fused through feature multiplication to obtain the feature map "T" i,11 ; The processing process of the spatial attention module is as follows: The average value and maximum value of the "T" i,11 are calculated channel by channel, corresponding to obtaining the average value feature map "T" i,12 and the maximum value feature map "T" i,13 ; The feature map "T" i,12 and the feature map "T" i,13 are fused through channel concatenation to obtain the feature map "T" i,14 ; The "T" i,14 passes through a convolutional layer with a kernel size of 7×7, padding of 3, stride of 1, 128 input channels, and 64 output channels. The output end of this convolutional layer outputs the feature map "T" i,15 ; The feature map "T" i,15 and the feature map "T" i,11 are fused through feature multiplication to obtain the feature map "T" i '; where i = 1, 2. When i = 1, the sizes of "T" i,1 , "T" i,2 , "T" i, ', "T" i,4 , "T" i,5 , "T" i,6 , "T" i,7 , "T" i,8 , "T" i,9 , "T" i,10 , "T" i,11 , "T" i,12 , "T" i,13 , "T" i,14 , "T" i,15 correspond to When i = 2, T” i,1 , T” i,2 , T” i,3 , T” i,4 , T” i,5 , T” i,6 , T” i,7 , T” i,8 , T” i,9 , T” i,10 , T” i,11 , T” i,12 , T” i,13 , T” i,14 , T” i,15 The dimensions of

[0061] Furthermore, it is specified that the structures of the cross-layer feature fusion modules of the two paths are the same, and it directly adopts the AFF module disclosed in Y. Dai, F. Gieseke, S. Oehmcke, Y. Wu, and K. Barnard, “Attentional feature fusion,” in Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision, 2021, pp. 3560–3569. III-D (Attentional Feature Fusion), which is used to fuse features of different sizes and avoid the disadvantages brought by direct feature addition or channel concatenation.

[0062] To further verify the feasibility and effectiveness of the method of the present invention, experiments are carried out on the method of the present invention.

[0063] During the experiment, existing natural datasets were directly selected for training and testing. The datasets include the MAS3K, COD10K, and CHAMELEON datasets. To improve the generalization ability and anti-interference ability of the network training model, the method of the present invention, during network training, based on the MAS3K training dataset, adds 3040 camouflage images from the COD10K dataset (which contain both camouflaged marine creature images and non-marine creature images), and combines the two into a training set containing 4809 images, so as to improve the anti-interference ability and generalization performance of the network training model. Performance testing is carried out on the MAS3K test dataset, and the generalization performance and anti-interference ability of the network training model are assisted and verified on the CHAMELEON dataset and the COD10K dataset to verify the generalization performance and anti-interference ability of the method of the present invention on natural datasets. Among them, the MAS3K dataset consists of 1769 training images (constituting the training dataset) and 1141 test images (constituting the test dataset), and contains a total of 37 species of marine creatures; the COD10K dataset consists of more than 78 species, as well as 5066 camouflage images, 3000 background images, and 1934 non-camouflage images, and contains a total of 20 species of marine creatures; the CHAMELEON dataset consists of 76 manually annotated camouflage images and contains a total of 10 species of marine creatures. Selecting different datasets for different tasks can train more pertinently to improve the detection accuracy. These natural datasets contain some non-camouflaged marine creature images and some camouflaged terrestrial natural images, which can improve the generalization performance and anti-interference ability of the network training model.

[0064] Four widely used evaluation metrics are adopted to evaluate the performance, including the mean absolute error MAE, the human visual perception evaluation metric Structural similarity index S α , and the weighted metric Specifically, the mean absolute error MAE metric is used to evaluate the pixel-level accuracy between the prediction result and the true label, and the human visual perception evaluation metric simultaneously evaluates the pixel-level matching degree and the image-level statistic to evaluate the overall and local accuracy of the model. The structural similarity index S α is used to measure the structural similarity between the prediction result and the true label, and the weighted metric combines the recall rate and the precision rate, eliminating the influence of equally considering each pixel in the traditional metrics.

[0065] Table 1 shows the quantitative comparison of the method of the present invention and common object segmentation methods in terms of evaluation metrics on the MAS3K, COD10K, and CHAMELEON datasets.

[0066] Table 1 Quantitative comparison table of the method of the present invention and common object segmentation methods in terms of evaluation metrics on MAS3K, COD10K and CHAMELEON datasets

[0067]

[0068]

[0069] As can be seen from Table 1, the four evaluation metrics S of the method of the present invention α 、 MAE、 reach 0.906, 0.865, 0.019, 0.945 respectively on the MAS3K dataset. The method of the present invention has achieved a certain degree of leadership in multiple public datasets and various evaluation metrics, and is superior to existing common object segmentation methods.

[0070] Figure 4 The detection effect display of partial data of the method of the present invention and existing methods on MAS3K and COD10K datasets is given. Figure 4 In [reference], the first column Image is the original image, the second column GT is the true label, the third column is the detection effect of the method of the present invention, the fourth column is the detection effect of the existing SINet-V2 method, the fifth column is the detection effect of the existing C2FNet method, the sixth column is the detection effect of the existing BASNet method, the seventh column is the detection effect of the existing F3Net method, the eighth column is the detection effect of the existing CPD method, the ninth column is the detection effect of the existing ECDNet method, the tenth column is the detection effect of the existing PFNet method, the eleventh column is the detection effect of the existing PraNet method, the twelfth column is the detection effect of the existing Rank-Net method, the thirteenth column is the detection effect of the existing C2FNet-V2 method, the fourteenth column is the detection effect of the existing PFSNet method, the fifteenth column is the detection effect of the existing Poly-PVT method, the sixteenth column is the detection effect of the existing PSGLoss method, the seventeenth column is the detection effect of the existing SINet-V1 method, and the eighteenth column is the detection effect of the existing SCRNet method. Comparing the detection effect of the method of the present invention with that of existing methods, it is found that in complex marine environments, such as backgrounds of turbidity, low light, occlusion, etc., the method of the present invention can clearly segment the position area where the camouflaged marine organisms are located.

[0071] At the same time, ablation experiments were carried out on the method of the present invention. Figure 5 The comparison results of the ablation experiments are given Figure 5In the first column Image, the original image is shown. In the second column GT, the true label is presented. In the third column, it shows the detection effect when only the backbone network of the deep neural network is retained. In the fourth column, it shows the detection effect when only the backbone network and the global feature enhancement module of the deep neural network are retained (that is, removing the dual-branch cross-level progressive feature fusion module). In the fifth column, it shows the detection effect of the method of the present invention. From Figure 5 it can be found that when only the backbone network and the global feature enhancement module of the deep neural network are retained, the segmented image will have blurred edges and incomplete content, while the morphological integrity of the image segmented by the method of the present invention is more excellent.

[0072] There are many specific implementation manners and ways for the method of the present invention. The above description is only a specific implementation manner of the method of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the method of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the method of the present invention. Each component not clearly defined in this embodiment can be implemented by using the prior art.

[0073] Those of ordinary skill in this technical field can realize that the units and steps of each example described in combination with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraint conditions of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of the method of the present invention.

[0074] Although the specific implementation manner of the method of the present invention has been described above in combination with the drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the method of the present invention, various modifications or deformations that can be made by those skilled in the art without creative labor are still within the protection scope of the method of the present invention.

Claims

1. A method for detecting marine organisms based on a bidirectional collaborative guidance network of CNN and Transformer, characterized in that Including the following steps: Step 1: Select or construct a dataset that contains two types of original images, namely marine organism images and non-marine organism images; then preprocess each original image in the dataset so that the size of the preprocessed image is H'×W'×3, and the mean of the pixel values of all pixel points in the R channel of the preprocessed image is 0.485 and the variance is 0.229, the mean of the pixel values of all pixel points in the G channel is 0.456 and the variance is 0.224, and the mean of the pixel values of all pixel points in the B channel is 0.406 and the variance is 0.225; then divide all the preprocessed images into a training set and a test set, and both the training set and the test set contain two types of images, namely marine organism images and non-marine organism images; Step 2: Build a deep neural network using a deep learning framework. This deep neural network consists of a backbone network, a global feature enhancement module for enhancing the features extracted from the backbone network, and a dual-branch cross-level progressive feature fusion module for fusing features of different scales in a progressive manner, thus forming a bidirectional collaborative guidance network based on CNN and Transformer. The backbone network includes a CNN-based Resnet50-backbone network for extracting high-level semantic information of the image and a Transformer-backbone network for extracting spatial and texture information of the image and supplementing the extracted spatial and texture information to the feature map captured by the Resnet50-backbone network. The Resnet50-backbone network has a total of five layers connected in sequence, and the Transformer-backbone network has a total of two layers connected in sequence. The Transformer-backbone network is composed of the first and second layers in the PVT network based on Transformer. The input ends of the first layer of the Resnet50-backbone network and the first layer of the Transformer-backbone network simultaneously receive an image with a size of H×W×3. In the Resnet50-backbone network, the input end of the second layer receives the feature map R1 with a size of output from the output end of the first layer, the input end of the third layer receives the feature map R2 with a size of output from the output end of the second layer, the input end of the fourth layer receives the feature map R3 with a size of output from the output end of the third layer, the input end of the fifth layer receives the feature map R4 with a size of output from the output end of the fourth layer, the output end of the fifth layer outputs a feature map R5 with a size of In the Transformer-backbone network, the input end of the second layer receives the feature map T1 with a size of output from the output end of the first layer, and the output end of the second layer outputs a feature map T2 with a size of ; The first input end of the global feature enhancement module receives R4, the second input end receives R5, and the third input end receives T2. The global feature enhancement module generates a feature map P1' with a size of . The output end of the global feature enhancement module outputs a feature map P1 for auxiliary training with a size of H×W×1, which is obtained by performing an eight-fold upsampling operation on P1' and then passing it through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 1; The first input end of the dual-branch cross-level progressive feature fusion module receives T1, the second input end receives T2, and the third input end receives P1'. The dual-branch cross-level progressive feature fusion module generates a feature map P2' with a size of . The output end of the dual-branch cross-level progressive feature fusion module outputs a feature map P2 for the training loss function with a size of H×W×1, which is obtained by performing a four-fold upsampling operation on P2' and then passing it through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 1; Step 3: Based on the size of each image in the training set, reduce each image in the training set by 0.75 times and enlarge it by 1.25 times respectively to expand the training set and form an extended training set, and the size of each image in the extended training set is H×W×3; where, H = H', 0.75H', 1.25H', and W = W', 0.75W', 1.25W'; Step 4: Use the extended training set to train the deep neural network constructed in Step 2. At the end of each round of network training, the deep neural network outputs the feature map P1 for auxiliary training and the feature map P2 for training the loss function corresponding to each image in the extended training set. Then, calculate the loss function Loss, where Loss = L main + L aux , L main = L wbce (P2, GT) + L iou (P2, GT), L aux = L wbce (P1, GT) + L iou (P1, GT); where L main represents the main loss function, L aux represents the auxiliary loss function, L wbce () represents the binary cross-entropy loss function, L iou () represents the weighted intersection over union loss function, GT represents the true label, the network training is implemented using Pytorch, the optimizer is AdaXW, the batch size is set to 24, the initial learning rate is set to 1e-4, and the learning rate decays by 10 times every 30 epochs; Step 5: Perform network training for 150 epochs in accordance with the process of Step 4 to obtain a deep neural network training model; Step 6: Use the deep neural network training model to test each image in the test set, and detect the corresponding camouflaged marine organism detection image P2 for each image in the test set.

2. The method for detecting marine organisms based on a bidirectional collaborative guidance network of CNN and Transformer according to claim 1, characterized in that In the said Step 1, the process of preprocessing an original image is as follows: First, scale the size of the original image to H'×W'×3; second, perform normalization processing on the scaled image, and normalize the pixel values of all pixel points in the R channel of the scaled image to a mean of 0.485 and a variance of 0.229, the pixel values of all pixel points in the G channel to a mean of 0.456 and a variance of 0.224, and the pixel values of all pixel points in the B channel to a mean of 0.406 and a variance of 0.

225.

3. The method for detecting marine organisms based on a bidirectional collaborative guidance network of CNN and Transformer according to claim 1, characterized in that In the said Step 2, the global feature enhancement module is of a dual-path parallel structure, and its processing process is: The first path: Perform a two-fold upsampling operation on R4 to obtain a feature map R with a size of ; Pass R 41 through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 1024, and number of output channels of 64. This convolutional layer reduces the number of channels of R 41 to 64 dimensions, and the output end of this convolutional layer outputs a feature map R with a size of 41 ; Pass R through a Relu activation function to obtain a feature map R with a size of 42 ; Perform feature fusion on R 42 and T2 through channel concatenation to obtain a feature map f1' with a size of ; Pass f1' through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 192, and number of output channels of 64. This convolutional layer reduces the number of channels of f1' to 64 dimensions, and the output end of this convolutional layer outputs a feature map f1'1 with a size of 43 ; Pass f1'1 through a Relu activation function to obtain a feature map f1'2 with a size of 43 ; Perform feature enhancement on f1'2 through a feature enhancement module, and the output end of this feature enhancement module outputs a feature map f1'3 with a size of ; Pass f1'3 through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 3. The output end of this convolutional layer outputs a feature map f1 with a size of ; Input f1 into the Transf module based on Transformer. The output end of this Transf module based on Transformer outputs a feature map P1'1' with a size of ; Pass P1'1' through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 6, and number of output channels of 64. The output end of this convolutional layer outputs a feature map P1'1 with a size of ; Pass P1'1' through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 3. The output end of this convolutional layer outputs a feature map P1'1 with a size of ; Input f1 into the Transf module based on Transformer. The output end of this Transf module based on Transformer outputs a feature map P1'1' with a size of ; Pass P1'1' through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 6, and number of output channels of 64. The output end of this convolutional layer outputs a feature map P1'1 with a size of ; Second path: perform a double upsampling operation on R5 to obtain a size of The feature map R 51 ; R is spliced ​​by channel splicing 51 After feature fusion with R4, the size is The feature map R 52 ; R 52 Through a convolution layer with a kernel size of 3×3, padding of 1, stride size of 1, input channels of 3072, and output channels of 64, R 52 The number of channels is reduced to 64 dimensions, and the output size of the output of the convolutional layer is The feature map R 53 ; R 53 Through a Relu activation function, the size is The feature map R 54 ; for R 54 Perform a double upsampling operation to obtain a size of The feature map R 55 ; R 55 Through a convolution layer with a convolution kernel size of 3×3, padding of 1, stride size of 1, input channels of 64, and output channels of 64, the output size of the output of the convolution layer is The feature map R 56 ; R 56 Through a Relu activation function, the size is The feature map R 57 ; R is spliced ​​by channel splicing 57 After feature fusion with T2, the size is The feature map f2' is taken as f2'; f2' is passed through a convolution layer with a convolution kernel size of 3×3, a padding of 1, a stride size of 1, an input channel number of 192, and an output channel number of 64. This convolution layer reduces the number of channels of f2' to 64 dimensions. The output size of the output of this convolution layer is The feature map f2'1 is obtained by passing f2'1 through a Relu activation function. The feature map f2'2 is enhanced by a feature enhancement module, and the output size of the output end of the feature enhancement module is The feature map f2'3 of f2'3 is obtained; f2'3 is passed through a convolution layer with a convolution kernel size of 3×3, a padding of 1, a stride size of 1, an input channel number of 64, and an output channel number of 3. The output size of the output of the convolution layer is The feature map f2 of f2 is input into the Transformer-based Transf module, and the output size of the output end of the Transformer-based Transf module is feature map P1'2'; passing P1'2' through a convolutional layer with a convolutional kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 6, and number of output channels of 64, and the output end of this convolutional layer outputs a size of feature map P1'2; Feature fusion of P1'1 and P1'2 is performed by adding features to obtain a feature map γ of size ; γ is passed through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 64. The output end of this convolutional layer outputs a feature map P1' of size ; An eight-fold upsampling operation is performed on P1' to obtain a feature map of size H×W×64 is passed through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 1. The output end of this convolutional layer outputs a feature map P1 of size H×W×1 for auxiliary training.​ 4. The method for detecting marine organisms based on a bidirectional collaborative guidance network of CNN and Transformer according to claim 3, characterized in that The structure of the feature enhancement module of the two paths is the same, which includes five branches. For the first branch: the input feature map s passes through a convolution layer with a convolution kernel size of 1×1, padding of 0, stride size of 1, input channel number of 64, and output channel number of 64. The output size of the output of the convolution layer is The feature map of 11 ; will s 11 Through a convolution layer with a convolution kernel size of 1×3, padding of (0,1), 64 input channels, and 64 output channels, the output size of the output of the convolution layer is The feature map of 12 ; will s 12 Through a convolution layer with a convolution kernel size of 3×1, padding of (1,0), 64 input channels, and 64 output channels, the output size of the output of the convolution layer is The feature map of 13 ; will s 13 Through a dilated convolution layer with a dilated convolution rate of 3, padding of 3, stride size of 3, input channels of 64, and output channels of 64, the output size of the output of the dilated convolution layer is The feature map of 14 ; For the second branch: the input feature map s is passed through a convolution layer with a convolution kernel size of 1×1, padding of 0, stride size of 1, input channels of 64, and output channels of 64. The output size of the output of the convolution layer is The feature map of 21 ; will s 21 Through a convolution layer with a convolution kernel size of 1×5, padding of (0,2), 64 input channels, and 64 output channels, the output size of the output of the convolution layer is The feature map of 22 ; will s 22 Through a convolution layer with a convolution kernel size of 5×1, padding of (2,0), 64 input channels, and 64 output channels, the output size of the output of the convolution layer is The feature map of 23 ; will s 23 Through a dilated convolution layer with a dilated convolution rate of 5, a padding of 5, a stride size of 3, 64 input channels, and 64 output channels, the output size of the output of the dilated convolution layer is The feature map of 24 ; For the third branch: the input feature map s is passed through a convolution layer with a convolution kernel size of 1×1, padding of 0, stride size of 1, input channels of 64, and output channels of 64. The output size of the output of the convolution layer is The feature map of 31 ; Pass s 31 through a convolutional layer with a kernel size of 1×7, padding of (0, 3), 64 input channels, and 64 output channels. The output size at the output end of this convolutional layer is feature map s 32 ; Pass s 32 through a convolutional layer with a kernel size of 7×1, padding of (3, 0), 64 input channels, and 64 output channels. The output size at the output end of this convolutional layer is feature map s 33 ; Pass s 33 through an atrous convolutional layer with an atrous rate of 7, padding of 7, stride size of 3, 64 input channels, and 64 output channels. The output size at the output end of this atrous convolutional layer is feature map s 34 ; For the 4th branch, pass the input feature map s through a convolutional layer with a kernel size of 1×1, padding of 0, stride size of 1, 64 input channels, and 64 output channels. The output size at the output end of this convolutional layer is feature map s 41 ; For the 5th branch, pass the input feature map s through a convolutional layer with a kernel size of 1×1, padding of 0, stride size of 1, 64 input channels, and 64 output channels. The output size at the output end of this convolutional layer is feature map s 51 ; Through the way of channel concatenation, perform feature fusion on s 14 , s 24 , s 34 , s 41 to obtain a feature map with a size of feature map s 1234 ; Pass s 1234 through a convolutional layer with a kernel size of 1×1, stride size of 1, 256 input channels, and 64 output channels. The output size at the output end of this convolutional layer is feature map s1' 234 ; Through the way of feature addition, perform feature fusion on s1' 234 and s 51 to obtain a feature map with a size of feature map where s is f1'2 or f2'2. If s is f1'2, then is f1'3. If s is f2'2, then is f2'3.

5. The method for detecting marine organisms based on a bidirectional collaborative guidance network of CNN and Transformer according to claim 3 or 4, characterized in that The described Transf module based on Transformer is composed of the first layer in the PVT network based on Transformer, and the size of the input feature map is The size of the output feature map is 6. The method for detecting marine organisms based on a bidirectional collaborative guidance network of CNN and Transformer according to claim 3, characterized in that In the said Step 2, the dual-branch cross-level progressive feature fusion module is of a dual-path parallel structure, and each path has a camouflage target recognition module and a cross-level feature fusion module, and the processing process is: The first path: Pass T1 through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 64. The output size at the output end of this convolutional layer is of the feature map T1”; Pass T1” through the camouflaged target recognition module to remove the noise information unrelated to the camouflaged target on the feature map. The output size at the output end of this camouflaged target recognition module is of the feature map T1'; Perform cross-level feature fusion on T1' and P1' through the cross-layer feature fusion module. The output size at the output end of this cross-layer feature fusion module is of the feature map P2'1; The second path: Pass T2 through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 128, and number of output channels of 64. This convolutional layer reduces the number of channels of T2 to 64 dimensions, and the output of this convolutional layer has a size of feature map T2"; Pass T2" through the camouflaged target recognition module to remove noise information unrelated to the camouflaged target on the feature map. The output of this camouflaged target recognition module has a size of feature map T2'; The feature map with a size of obtained after performing a two-fold upsampling operation on T2' and P1' are subjected to cross-level feature fusion through the cross-layer feature fusion module. The output of this cross-layer feature fusion module has a size of feature map P2'2; Feature fusion of P2'1 and P2'2 is performed by adding features to obtain a feature map β of size ; β is passed through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 64. The output end of this convolutional layer outputs a feature map P2' of size ; A four-fold upsampling operation is performed on P2' to obtain a feature map of size H×W×64 is passed through a convolutional layer with a kernel size of 3×3, padding of 1, stride size of 1, number of input channels of 64, and number of output channels of 1. The output end of this convolutional layer outputs a feature map P2 of size H×W×1 for the training loss function.​ 7. The method for detecting marine organisms based on a bidirectional collaborative guidance network of CNN and Transformer according to claim 6, characterized in that The structures of the two-channel camouflaged target recognition modules are the same. It is mainly composed of a channel attention module and a spatial attention module connected in series. The processing process of the channel attention module is as follows: Take T i " through a global average pooling layer, and the output end of the global average pooling layer outputs a feature map T i ' ,1 '; Take T i ' ,1 ' through a convolutional layer with a kernel size of 1×1, 64 input channels, and 4 output channels. The output end of this convolutional layer outputs a feature map T i ' , '2; Take T i ' , '2 through a Relu activation function to obtain a feature map T i ' , '3; Take T i ' , '3 through a convolutional layer with a kernel size of 1×1, 4 input channels, and 64 output channels. The output end of this convolutional layer outputs a feature map T i ' , '4; Take T i ” through a global max pooling layer, and the output end of the global max pooling layer outputs a feature map T i ' , '5; Take T i ' , '5 through a convolutional layer with a kernel size of 1×1, 64 input channels, and 4 output channels. The output end of this convolutional layer outputs a feature map T i ' , '6; Take T i ' , '6 through a Relu activation function to obtain a feature map T i ' , '7; Take T i ' , '7 through a convolutional layer with a kernel size of 1×1, 4 input channels, and 64 output channels. The output end of this convolutional layer outputs a feature map T i ' , '8; Through the way of feature addition, perform feature fusion on T i ' , '4 and T i ' , '8 to obtain a feature map T i ' , '9; Take T i ' , '9 through a sigmoid function to obtain a feature map T i, '1'0; Feature fusion is performed on T by multiplying features i ' ,1 '0 and T i ” to obtain the feature map T i ' ,1 '1; The processing process of the spatial attention module is as follows: For T i ' ,1 '1, calculate the average value and maximum value by channel, and correspondingly obtain the average value feature map T i ' ,1 '2 and the maximum value feature map T i ' ,1 '3; Feature fusion is performed on T i, '1'2 and T i ' ,1 '3 by channel concatenation to obtain the feature map T i ' ,1 '4; Pass T i, '1'4 through a convolutional layer with a convolutional kernel size of 7×7, padding of 3, stride size of 1, number of input channels of 128, and number of output channels of 64. The output end of this convolutional layer outputs the feature map T i, '1'5; Feature fusion is performed on T i ' ,1 '5 and T i ' ,1 '1 to obtain the feature map T i '; where, i = 1, 2.

8. The method for detecting marine organisms based on a bidirectional collaborative guidance network of CNN and Transformer according to claim 6, characterized in that The structures of the cross-level feature fusion modules of the two paths are the same.

Citation Information

Patent Citations

  • Underwater target detection method based on attention fusion

    CN114782798A

  • CNN and Transform fusion-based colonoscope polyp image segmentation method

    CN115018824A