A progressive detection enhancement method for marine benthic organisms guided by channel-spatial fuzzy logic attention under detail loss

By constructing a progressive detection enhancement model guided by fuzzy logic, the problems of comprehensive enhancement of feature maps and regional responses in fuzzy marine benthic organisms are solved, and efficient detection of marine benthic organisms is achieved.

CN119723312BActive Publication Date: 2025-08-19DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411791327.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-08-19
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively comprehensively enhance information feature maps in fuzzy marine benthic biological detection, single-scale convolutional kernels are difficult to enhance responses to different scales at the same time, and the fuzzy perceptual information introduced by underwater optical imaging makes it difficult to appropriately derivatize the channel and spatial attention map weights.

Method used

The method of channel-space fuzzy logic attention guidance is adopted to build a progressive detection enhancement model containing fuzzy channels and spatial fuzzy logic attention modules, and combine fuzzy inference systems to fusion discriminant and textured channels and spatial descriptors to design fuzzy channels and spatial attention modules to achieve comprehensive enhancement of feature maps.

Benefits of technology

The performance of fuzzy marine benthic biological detection is improved, background noise is suppressed from the channel and spatial levels, feature response is enhanced, and detection accuracy and recall rate is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723312B_ABST
    Figure CN119723312B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for progressive detection and enhancement of marine benthic organisms guided by channel-spatial fuzzy logic attention under detail loss, comprising the following steps: constructing a dataset based on images containing one or more marine benthic organisms selected from the group consisting of sea urchins, scallops, starfish, and sea cucumbers; dividing the dataset into a training set and a test set; constructing a progressive detection and enhancement model for marine benthic organisms guided by channel-spatial fuzzy logic attention under detail loss; training the progressive detection and enhancement model for marine benthic organisms guided by channel-spatial fuzzy logic attention based on the training set data to obtain a trained progressive detection and enhancement model for marine benthic organisms guided by channel-spatial fuzzy logic attention; and inputting the test set data into the trained progressive detection and enhancement model for marine benthic organisms guided by channel-spatial fuzzy logic attention to achieve detection and recognition of marine benthic organisms such as sea urchins, scallops, starfish, and sea cucumbers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of underwater visual intelligent perception, and relates to a method for progressive detection enhancement of marine benthic organisms guided by channel-spatial fuzzy logic attention under detail loss. Background Art

[0002] Compared with the terrestrial environment, the real underwater environment contains a large amount of dissolved matter and particulate matter. [1]-[3] , such as underwater micro-phytoplankton, soluble organic matter and non-algae particles, which directly change the original propagation direction of light and produce scattering [4] , resulting in the loss of detailed information of marine benthic organisms, which undoubtedly increases the difficulty of detecting fuzzy marine benthic organisms. Therefore, it is urgent to propose a high-precision detection method for fuzzy marine benthic organisms.

[0003] To address the above challenges, existing methods can be summarized as follows: integrating contextual features [5] , explore the attention mechanism [6] , build a multi-layer spatial pyramid [7] , jointly optimize image enhancement and object detection tasks [8] , fuzzy fusion feature map [9]-

[11] Specifically, in order to improve the detection accuracy of fuzzy marine benthic organisms, the literature [5] A multi-scale context feature fusion module and a multi-scale fuzzy sampling module were proposed to effectively construct the multi-scale features of marine benthic organisms.

[12] First, channel attention is used to enhance the feature maps related to marine benthic organisms and suppress the feature maps related to the seabed background; then, spatial attention is used to enhance the features of the area containing marine benthic organisms and suppress the features of the seabed background area; finally, the features of marine benthic organisms are enhanced from both the channel and spatial levels, improving the detection performance of fuzzy marine benthic organisms.

[13] A feature enhancement gating module is proposed to selectively suppress or enhance multi-layer features, reducing the impact of the complex underwater environment on subsequent feature fusion; a feature dynamic fusion module is designed to aggregate features of adjacent layers, establishing the relationship between the target scale in the input image and feature fusion; a spatial pyramid hybrid pooling module is developed to enhance the ability of deep convolutional neural networks to extract texture features and contour features, effectively reducing the missed detection rate of ambiguous marine benthic organisms.

[14] A sample-weighted hyperparameter network model based on dilated convolution was constructed to form multiple semantic feature maps, which enhanced the detection performance of fuzzy marine benthic organisms; a progressive learning model from clean samples to fuzzy noise samples was explored to improve the detection performance of fuzzy marine benthic organisms.

[0004] Generally speaking, jointly optimizing multiple different tasks can improve the generalization performance of convolutional neural network models.[8] A method for jointly optimizing image enhancement and marine benthic organism detection tasks is proposed. A guidance model based on channel and position dual attention is constructed. The degradation process is inverted using an underwater imaging model to obtain potential clear images, thereby enhancing the ability to express details of blurred marine benthic organisms. Similarly, the literature

[15] A deep convolutional neural network for image enhancement-guided marine benthic organism detection is proposed. The core idea is to enhance the branch guidance to optimize the detection network, so that the detection branch features have both good visual effects and robust expression capabilities. The results show that this method can achieve excellent detection results for fuzzy marine benthic organisms. In the weakly supervised learning mode, the literature

[16] A novel fuzzy marine benthic organism detection method is proposed. The core idea is to train two deep learning detectors simultaneously and design a batch sample selection strategy to encourage mutual learning between detectors, thereby improving the fuzzy marine benthic organism detection performance.

[0005] Through the review and analysis of existing work, the following deficiencies still exist in the fuzzy marine benthic organism detection:

[0006] (1) Using only a single discriminative channel descriptor or texture channel descriptor cannot comprehensively enhance the informative feature map, and a single hidden layer feedforward neural network is difficult to specifically and efficiently learn the aforementioned discriminative channel descriptor and texture channel descriptor.

[0007] (2) Based on the channel cascade feature map, it is difficult to simultaneously enhance the regional responses of blurred marine benthic organisms at different scales using a single-scale convolution kernel.

[0008] (3) When a large amount of fuzzy perceptual information is introduced into the underwater optical imaging system, it is often difficult to derive appropriate channel attention weights or spatial attention weights by directly fusing the discriminative and texture channel attention maps or spatial attention maps with different receptive fields in an algebraic sum manner. Summary of the Invention

[0009] In order to solve the above problems, the technical solution adopted by the present invention is: a method for progressive detection and enhancement of marine benthic organisms guided by channel-spatial fuzzy logic attention under detail loss, characterized by comprising the following steps:

[0010] Construct a dataset based on images of one or more of the following marine benthic organisms: sea urchins, scallops, starfish, and sea cucumbers;

[0011] Divide the dataset into: training set and test set;

[0012] Construct a progressive detection enhancement model for marine benthic organisms, including detail loss channels and spatial fuzzy logic attention guidance;

[0013] Based on the training set data, the channel-spatial fuzzy logic attention-guided marine benthic organisms progressive detection enhancement model is trained to obtain the trained channel-spatial fuzzy logic attention-guided marine benthic organisms progressive detection enhancement model;

[0014] The test set data is input into the trained channel-spatial fuzzy logic attention guided progressive detection enhancement model of marine benthic organisms to achieve the detection and recognition of marine benthic organisms such as sea urchins, scallops, starfish and sea cucumbers.

[0015] Furthermore, the detailed loss channel and spatial fuzzy logic attention-guided marine benthic organism progressive detection enhancement model includes:

[0016] Backbone network: extracts three different fine-grained feature maps from images of marine benthic organisms;

[0017] Detection head network: Based on the three different fine-grained feature maps output by the backbone network, feature fusion operations are performed to achieve the detection of large-sized, medium-sized and small-sized marine benthic organisms.

[0018] Further: the backbone network includes a first CBR3 module, a second CBR3 module, a CSFLA-1 module, a third CBR3 module, a CSFLA-2 module, a fourth CBR3 module, a first CSFLA-8 module, a fifth CBR3 module, a second CSFLA-8 module, a sixth CBR3 module and a CSFLA-4 module;

[0019] The first CBR3 module, the second CBR3 module, the CSFLA-1 module, the third CBR3 module, the CSFLA-2 module, the fourth CBR3 module, the first CSFLA-8 module, the fifth CBR3 module, the second CSFLA-8 module, the sixth CBR3 module and the CSFLA-4 module are connected in sequence;

[0020] The first CSFLA-8 module output, the second CSFLA-8 module output and the CSFLA-4 module output three different fine-grained feature maps respectively.

[0021] Further: the CSFLA-1 module, CSFLA-2 module, first CSFLA-8 module, second CSFLA-8 module and CSFLA-4 module respectively represent CSFLA modules reused 1, 2, 8, 8 and 4 times in sequence;

[0022] The CSFLA module includes 3×3 convolution, 1×1 convolution, fuzzy channel attention FCA module and fuzzy spatial attention FSA module, and adopts a skip connection strategy to prevent semantic information degradation;

[0023] The fuzzy channel attention FCA module is used to combine spatial global average pooling and maximum pooling methods to construct discriminative and textural channel descriptors for fuzzy marine benthic features, and a channel fuzzy inference system is designed to fuse the discriminative and textural channel attention maps derived from two single hidden layer networks.

[0024] The fuzzy spatial attention (FSA) module uses channel global average pooling and maximum pooling methods to construct discriminative and textural spatial descriptors of fuzzy marine benthic features, and designs a spatial fuzzy inference system to fuse spatial attention maps with different receptive fields.

[0025] Furthermore, the discriminative and texture channel descriptors are constructed as follows:

[0026] Two independent single-hidden-layer feedforward neural networks with the same network structure are used to learn the discriminative and texture channel descriptors derived from spatial global average pooling and maximum pooling operations respectively. The output of spatial global average pooling is expressed as:

[0027]

[0028] Among them, y a,k represents the average value of the feature map of the kth layer, i = 1, 2, ..., h, j = 1, 2, ..., w, h and w are the height and width of the current feature map respectively, represents the input features, represents the k-th layer feature map, represents the discriminative channel descriptor, k represents the feature layer index ranging from 1 to n c ;

[0029] The output of a single hidden layer feed-forward network that learns a discriminative channel descriptor is represented as:

[0030] o a =s(W a,2 r(W a,1 y a )) (2)

[0031] in, represents the discriminative channel attention map, and Represent the hidden layer weights and output layer weights respectively, r(·) represents the Leaky ReLU activation function, and s(·) represents the conversion of o a Sigmoid activation function limited to the range (0,1);

[0032] Accordingly, the output of the spatial global maximum pooling operation (i.e., the texture channel descriptor) is expressed as:

[0033] y m,k=max{X k (1, 1),…,X k (h, w)} (3)

[0034] in, represents the texture channel descriptor, max{·} represents the maximum value operation, y m,k Represents the maximum value in the k-th layer feature map;

[0035] Similarly, the output of a single hidden layer feed-forward network for learning texture channel descriptors is expressed as:

[0036] o m =s(W m,2 r(W m,1 y m )) (4)

[0037] in, represents the texture channel attention map, and Represent the hidden layer and output layer weights respectively, and the activation function s(·) converts o m Clamped to the range (0,1).

[0038] Furthermore: the design process of the channel fuzzy inference system is as follows:

[0039] Fusion of discriminative channel attention maps o a and texture channel attention map o m , the input variables and output variables are fuzzified by three fuzzy sets, namely "S", "M" and "L" (i.e. small, medium and large), and are characterized by Gaussian membership functions, expressed as:

[0040]

[0041] Among them, Θ∈{a,m,c}, μ Θ,k,S (·), μ Θ,k,M (·) and μ Θ,k,L (·) represents the small, medium and large membership of the channel attention map, σ represents the standard deviation, exp(·) represents the natural exponential function, c S =0 represents the Gaussian membership function center of the fuzzy set as "small", c M =0.5 represents the center of the Gaussian membership function of the fuzzy set as “medium”, c L =1.0 represents the center of the Gaussian membership function of the fuzzy set as “large”;

[0042] The following fuzzy rules are established:

[0043] Rule 1: If o a,k It is S and o m,k is S, then oc,k For S;

[0044] Rule 2: If o a,k It is S and o m,k is M, then o c,k For S;

[0045] Rule 3: If o a,k It is S and o m,k is L, then o c,k is M;

[0046] Rule 4: If o a,k It is M and o m,k is S, then o c,k For S;

[0047] Rule 5: If o a,k It is M and o m,k is M, then o c,k is M;

[0048] Rule 6: If o a,k It is M and o m,k is L, then o c,k is M;

[0049] Rule 7: If o a,k It is L and o m,k is S, then o c,k is M;

[0050] Rule 8: If o a,k It is L and o m,k is M, then o c,k is M;

[0051] Rule 9: If o a,k It is L and o m,k is L, then o c,k is L;

[0052] Among them, a,k represents the k-th layer weight in the discriminative channel attention map, o m,k Represents the k-th layer weight in the texture channel attention map, o c,k Represents the k-th layer weight in the channel attention map after fuzzy reasoning;

[0053] Fuzzy reasoning is performed using the smaller operation, and the aggregated credibility (i.e., the union of the credibility reasoning results of each rule) is obtained by the larger operation, which is expressed as:

[0054] μ agg,k (o c,k )=max{min{μ a,k,Γ (o a,k ), μm,k,Γ (o m,k ), μ c,k,Γ (o c,k )}|n cfr =1, 2, ..., 9}(8)

[0055] Among them, Γ∈{S,M,L}, n cfr =1,…,9 represents the fuzzy rule index;

[0056] Deblurring is performed using the centroid method, which is expressed as:

[0057]

[0058] in, Represents the recalibration weight of the kth channel.

[0059] Furthermore, the channel features are recalibrated as follows:

[0060] The output of the FCA module is represented as:

[0061]

[0062] in, represents the recalibrated channel weights, B(·) represents broadcasting a one-dimensional vector into a three-dimensional matrix, and “e” represents element-by-element multiplication. Represents the features recalibrated using the FCA module.

[0063] Furthermore: the discriminative and texture space descriptors are constructed as follows:

[0064] Channel global average pooling is expressed as:

[0065]

[0066] Where i = 1, ..., h, j = 1, ..., w, represents the discriminative spatial descriptor, n c Represents the number of channel layers;

[0067] The channel global maximum pooling is expressed as:

[0068]

[0069] in, stands for Texture Space Descriptor;

[0070] Then, the channel cascade feature map is expressed as:

[0071]

[0072] in, Represents the channel cascade feature map, symbol Represents the concatenation of features in the channel dimension;

[0073] Subsequently, the convolution operation under different receptive fields is expressed as:

[0074]

[0075] Among them, Ψ∈{p,q,l}, W p,i 、W q,i and W l,i Represent the weight of the i-th layer in the convolution kernel of 3×3, 5×5 and 7×7 respectively, Spatial attention maps representing different receptive fields, Represents the convolution operation;

[0076] In order to enhance the nonlinear mapping capability, the spatial attention map is activated and expressed as:

[0077] O p =s(U p ) (15)

[0078] O q =s(U q ) (16)

[0079] O l =s(U l ) (17)

[0080] in, Represents the spatial attention map after activation.

[0081] Furthermore: the design process of the spatial fuzzy reasoning system is as follows:

[0082] The activated spatial attention map O is fused using a fuzzy inference system. p , O q and O l , the input variables and output variables are fuzzified by three fuzzy sets, namely "S", "M" and "L" (i.e. small, medium and large), and described by Gaussian membership functions, expressed as:

[0083]

[0084] Among them, Φ∈{p,q,l,s}, μ Φ,i,j,S (·), μ Φ,i,j,M (·) and μ Φ,i,j,L (·) represents the small, medium, and large membership of the spatial attention map, respectively;

[0085] Create the following fuzzy rules:

[0086] Rule 1: If op,i,j It is S, o q,i,j It is S and o l,i,j is S, then o s,i,j For S;

[0087] Rule 2: If o p,i,j It is S, o q,i,j It is S and o l,i,j is M, then o s,i,j For S;

[0088] Rule 3: If o p,i,j It is S, o q,i,j It is S and o l,i,j is L, then o s,i,j For S;

[0089] Rule 4: If o p,i,j It is S, o q,i,j It is M and o l,i,j is S, then o s,i,j For S;

[0090] Rule 5: If o p,i,j It is S, o q,i,j It is M and o l,i,j is M, then o s,i,j For S;

[0091] Rule 6: If o p,i,j It is S, o q,i,j It is M and o l,i,j is L, then o s,i,j is M;

[0092] Rule 7: If o p,i,j It is S, o q,i,j It is L and o l,i,j is S, then o s,i,j For S;

[0093] Rule 8: If o p,i,j It is S, o q,i,j It is L and o l,i,j is M, then o s,i,j is M;

[0094] Rule 9: If o p,i,j It is S, o q,i,j It is L and o l,i,j is L, then o s,i,j is M;

[0095] Rule 10: If o p,i,j It's M, o q,i,j It is S and o l,i,j is S, then o s,i,j For S;

[0096] Rule 11: If o p,i,j It's M, o q,i,j It is S and o l,i,j is M, then o s,i,j For S;

[0097] Rule 12: If o p,i,j It's M, o q,i,j It is S and o l,i,j is L, then o s,i,j is M;

[0098] Rule 13: If o p,i,j It's M, o q,i,j It is M and o l,i,j is S, then o s,i,j For S;

[0099] Rule 14: If o p,i,j It's M, o q,i,j It is M and o l,i,j is M, then o s,i,j is M;

[0100] Rule 15: If o p,i,j It's M, o q,i,j It is M and o l,i,j is L, then o s,i,j is M;

[0101] Rule 16: If o p,i,j It's M, o q,i,j It is L and o l,i,j is S, then o s,i,j is M;

[0102] Rule 17: If o p,i,j It's M, o q,i,j It is L and o l,i,j is M, then o s,i,j is M;

[0103] Rule 18: If o p,i,j It's M, o q,i,j It is L and o l,i,j is L, then o s,i,j is L;

[0104] Rule 19: If o p,i,j It's L, o q,i,j It is S and o l,i,j is S, then o s,i,j For S;

[0105] Rule 20: If o p,i,j It's L, o q,i,j It is S and o l,i,j is M, then os,i,j is M;

[0106] Rule 21: If o p,i,j It's L, o q,i,j It is S and o l,i,j is L, then o s,i,j is M;

[0107] Rule 22: If o p,i,j It's L, o q,i,j It is M and o l,i,j is S, then o s,i,j is M;

[0108] Rule 23: If o p,i,j It's L, o q,i,j It is M and o l,i,j is M, then o s,i,j is M;

[0109] Rule 24: If o p,i,j It's L, o q,i,j It is M and o l,i,j is L, then o s,i,j is L;

[0110] Rule 25: If o p,i,j It's L, o q,i,j It is L and o l,i,j is S, then o s,i,j is M;

[0111] Rule 26: If o p,i,j It's L, o q,i,j It is L and o l,i,j is M, then o s,i,j is L;

[0112] Rule 27: If o p,i,j It's L, o q,i,j It is L and o l,i,j is L, then o s,i,j is L;

[0113] Similarly, fuzzy reasoning is performed by taking the smaller operation, and then the aggregate credibility (i.e., the union of the credibility reasoning results of each rule) in the spatial fuzzy reasoning system is obtained by taking the larger operation, which is expressed as:

[0114]

[0115] Among them, Γ∈{S,M,L}, n sfr =1,…,27 represents the fuzzy rule index;

[0116] Deblurring is performed using the centroid method, expressed as:

[0117]

[0118] in, Represents the recalibration weight at the coordinate position (i, j).

[0119] The output of the FSA module is represented as:

[0120]

[0121] in, represents the recalibrated spatial weight, B(·) represents the operation of broadcasting a two-dimensional matrix into a three-dimensional matrix, Represents the features recalibrated using the FSA module.

[0122] A channel-spatial fuzzy logic attention-guided progressive detection enhancement method for marine benthic organisms under detail loss: including:

[0123] Building Module I: Constructing a dataset based on images of one or more marine benthic organisms including sea urchins, scallops, starfish, and sea cucumbers;

[0124] Partitioning module: divide the data set into training set and test set;

[0125] Building Module II: Constructing a channel-spatial fuzzy logic attention-guided progressive detection enhancement model for marine benthic organisms;

[0126] Training module: Based on the training set data, the channel-spatial fuzzy logic attention-guided marine benthic organisms progressive detection enhancement model is trained to obtain the trained channel-spatial fuzzy logic attention-guided marine benthic organisms progressive detection enhancement model;

[0127] Detection module: The test set data is input into the trained channel-spatial fuzzy logic attention-guided progressive detection enhancement model of marine benthic organisms to achieve detection and recognition of marine benthic organisms such as sea urchins, scallops, starfish and sea cucumbers.

[0128] The present invention provides a method for progressive detection enhancement of marine benthic organisms guided by channel-spatial fuzzy logic attention under detail loss, which has the following advantages:

[0129] 1. Combining spatial global average pooling and maximum pooling techniques, a discriminative and textural channel descriptor for fuzzy marine benthic features was constructed. With the help of a fuzzy inference system, the discriminative and textural channel attention maps were integrated, and a fuzzy channel attention module was proposed. Comprehensive channel weights were derived, which enhanced the feature map response related to fuzzy marine benthic organisms and suppressed the seabed background noise feature map, thereby improving the detection performance of fuzzy marine benthic organisms at the channel level.

[0130] 2. A discriminative and textural spatial descriptor was constructed by combining channel global average pooling and maximum pooling techniques. With the help of a fuzzy inference system, spatial attention maps with different receptive fields were fused, and a fuzzy spatial attention module was proposed. Comprehensive spatial weights were derived, which enhanced the regional response of fuzzy marine benthic organisms and improved the detection performance of fuzzy marine benthic organisms from a spatial perspective.

[0131] 3. Combining the fuzzy channel attention module and the fuzzy spatial attention module, a fuzzy-channel controlled attention method with a residual mechanism is proposed, which suppresses the seabed background noise from both the channel and spatial levels, while enhancing the characteristics of marine benthic organisms, ultimately improving the detection performance of marine benthic organisms. BRIEF DESCRIPTION OF THE DRAWINGS

[0132] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0133] Figure 1 A channel-spatial fuzzy logic attention-guided progressive detection enhancement model for marine benthic organisms;

[0134] Figure 2 is the fuzzy channel attention module;

[0135] Figure 3 It is the fuzzy spatial attention module;

[0136] Figure 4 The performance comparison chart of the accuracy-recall curves of different categories, where (a) is the accuracy-recall curve of sea urchins; (b) is the accuracy-recall curve of scallops; (c) is the accuracy-recall curve of starfish; and (d) is the accuracy-recall curve of sea cucumbers.

[0137] Figure 5 The following is a qualitative comparison of detection performance in typical scenarios, where (a) is a dark light scene, (b) is a bright light scene, (c) is a multi-scale target coexistence scene, and (d) is a blurred scene. DETAILED DESCRIPTION

[0138] It should be noted that, unless there is any conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0139] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is in no way intended to limit the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0140] A method for progressive detection enhancement of marine benthic organisms guided by channel-spatial fuzzy logic attention under detail loss includes the following steps:

[0141] S1: Construct a dataset based on images of one or more marine benthic organisms including sea urchins, scallops, starfish, and sea cucumbers;

[0142] S2: Divide the dataset into training set and test set;

[0143] S3: Construct a channel-spatial fuzzy logic attention-guided progressive detection enhancement model for marine benthic organisms;

[0144] S4: training the channel-spatial fuzzy logic attention-guided marine benthic organism progressive detection enhancement model based on the training set data to obtain a trained channel-spatial fuzzy logic attention-guided marine benthic organism progressive detection enhancement model;

[0145] S5: Input the test set data into the trained channel-spatial fuzzy logic attention-guided progressive detection enhancement model of marine benthic organisms to achieve detection and recognition of marine benthic organisms such as sea urchins, scallops, starfish and sea cucumbers.

[0146] Steps S1 / S2 / S3 / S4 / S5 are performed sequentially;

[0147] Figure 1 A channel-spatial fuzzy logic attention-guided progressive detection enhancement model for marine benthic organisms;

[0148] The channel-spatial fuzzy logic attention-guided marine benthic organism progressive detection enhancement model includes:

[0149] Backbone network: Extracts three different fine-grained feature maps of 13×13×27, 26×26×27, and 52×52×27 for images of marine benthic organisms (416×416×3 resolution);

[0150] Detection head network: Based on the three different fine-grained feature maps output by the backbone network, feature fusion operations are performed to achieve the detection of large-sized, medium-sized and small-sized marine benthic organisms.

[0151] The backbone network includes a first CBR3 module, a second CBR3 module, a CSFLA-1 module, a third CBR3 module, a CSFLA-2 module, a fourth CBR3 module, a first CSFLA-8 module, a fifth CBR3 module, a second CSFLA-8 module, a sixth CBR3 module, and a CSFLA-4 module;

[0152] The first CBR3 module, the second CBR3 module, the CSFLA-1 module, the third CBR3 module, the CSFLA-2 module, the fourth CBR3 module, the first CSFLA-8 module, the fifth CBR3 module, the second CSFLA-8 module, the sixth CBR3 module and the CSFLA-4 module are connected in sequence;

[0153] The first CSFLA-8 module output, the second CSFLA-8 module output and the CSFLA-4 module output three different fine-grained feature maps respectively.

[0154] The first CBR3 module, the second CBR3 module, the third CBR3 module, the fourth CBR3 module, the fifth CBR3 module, and the sixth CBR3 module have the same structure; the CBR3 module is constructed by 3×3 convolution, batch normalization, and ReLU activation function;

[0155] Furthermore: the CSFLA-1 module, CSFLA-2 module, first CSFLA-8 module, second CSFLA-8 module and CSFLA-4 module respectively represent CSFLA modules reused 1, 2, 8, 8 and 4 times in sequence;

[0156] The CSFLA module includes 3×3 convolution, 1×1 convolution, fuzzy channel attention FCA module and fuzzy spatial attention FSA module, and adopts a skip connection strategy to prevent semantic information degradation;

[0157] By adopting the residual learning paradigm and combining the previously constructed fuzzy channel attention module and fuzzy spatial attention module, we finally proposed a systematic method for progressive detection enhancement of marine benthic organisms guided by Channel-Spatial Fuzzy Logic Attention (CSFLA).

[0158] The fuzzy channel attention (FCA) module is used to combine spatial global average pooling and maximum pooling methods to construct discriminative and textural channel descriptors for fuzzy marine benthic features, and a channel fuzzy inference system is designed to fuse the discriminative and textural channel attention maps derived from two single hidden layer networks.

[0159] The fuzzy spatial attention (FSA) module uses channel global average pooling and maximum pooling methods to construct discriminative and textural spatial descriptors of fuzzy marine benthic features, and designs a spatial fuzzy inference system to fuse spatial attention maps with different receptive fields.

[0160] The construction process of the fuzzy channel attention FCA (Fuzzy Channel Attention) module is as follows:

[0161] The construction process of the discriminative and texture channel descriptors is as follows:

[0162] Two independent single-hidden-layer feedforward neural networks with the same network structure are used to learn the discriminative and texture channel descriptors derived from spatial global average pooling and maximum pooling operations respectively. Figure 2 As shown, the output of spatial global average pooling is expressed as:

[0163]

[0164] Among them, y a,k represents the average value of the feature map of the kth layer, i = 1, 2, ..., h, j = 1, 2, ..., w, h and w are the height and width of the current feature map respectively, represents the input features, represents the k-th layer feature map, stands for Discriminative Channel Descriptor.

[0165] The output of a single hidden layer feed-forward network that learns a discriminative channel descriptor is represented as:

[0166] o a =s(W a,2 r(W a,1 y a )) (2)

[0167] in, represents the discriminative channel attention map, and Represent the hidden layer weights and output layer weights respectively, r(·) represents the Leaky ReLU activation function, and s(·) represents the conversion of o aSigmoid activation function restricted to the range (0,1).

[0168] Accordingly, the output of the spatial global maximum pooling operation (i.e., the texture channel descriptor) is expressed as:

[0169] y m,k =max{X k (1, 1),…,X k (h, w)} (3)

[0170] in, represents the texture channel descriptor, max{·} represents the maximum value operation, y m,k Represents the maximum value in the k-th layer feature map.

[0171] Similarly, the output of a single hidden layer feed-forward network for learning texture channel descriptors is expressed as:

[0172] o m =s(W m,2 r(W m,1 y m )) (4)

[0173] in, represents the texture channel attention map, and Represent the hidden layer and output layer weights respectively, and the activation function s(·) converts o m Clamped to the range (0,1).

[0174] The design process of the channel fuzzy inference system is as follows:

[0175] Inspired by the application of fuzzy inference systems in clustering, classification and segmentation tasks, we use fuzzy inference systems to fuse discriminative channel attention maps. a and texture channel attention map o m Specifically, the input and output variables are fuzzified by three fuzzy sets, namely "S", "M" and "L" (i.e. small, medium and large), and are characterized by Gaussian membership functions, expressed as:

[0176]

[0177] Among them, Θ∈{a,m,c}, μ Θ,k,S (·), μ Θ,k,M (·) and μ Θ,k,L (·) represents the small, medium and large membership of the channel attention map, σ represents the standard deviation, exp(·) represents the natural exponential function, c S =0, c M =0.5 and c L =1.0 represents the corresponding fuzzy set center.

[0178] The following fuzzy rules are established:

[0179] Rule 1: If o a,k It is S and o m,k is S, then o c,k For S;

[0180] Rule 2: If o a,k It is S and o m,k is M, then o c,k For S;

[0181] Rule 3: If o a,k It is S and o m,k is L, then o c,k is M;

[0182] Rule 4: If o a,k It is M and o m,k is S, then o c,k For S;

[0183] Rule 5: If o a,k It is M and o m,k is M, then o c,k is M;

[0184] Rule 6: If o a,k It is M and o m,k is L, then o c,k is M;

[0185] Rule 7: If o a,k It is L and o m,k is S, then o c,k is M;

[0186] Rule 8: If o a,k It is L and o m,k is M, then o c,k is M;

[0187] Rule 9: If o a,k It is L and o m,k is L, then o c,k is L;

[0188] Fuzzy reasoning is performed using the smaller operation (i.e., the conclusion is obtained by taking the “and” operation between the premises), and the aggregate credibility is obtained by taking the larger operation (i.e., the union of the credibility reasoning results of each rule), which is expressed as:

[0189] μ agg,k (o c,k )=max{min{μ a,k,Γ (o a,k ), μ m,k,Γ(o m,k ), μ c,k,Γ (o c,k )}|n cfr =1, 2, ..., 9}(8)

[0190] Among them, Γ∈{S,M,L}, n cfr =1,…,9 represents the fuzzy rule index.

[0191] Deblurring is performed using the centroid method, which is expressed as:

[0192]

[0193] in, Represents the recalibration weight of the kth channel.

[0194] Furthermore, the channel features are recalibrated as follows:

[0195] The output of the FCA module is represented as:

[0196]

[0197] in, represents the recalibrated channel weights, B(·) represents broadcasting a one-dimensional vector into a three-dimensional matrix, and “e” represents element-by-element multiplication. Represents the features recalibrated using the FCA module.

[0198] The design process of the fuzzy spatial attention module is as follows:

[0199] From a spatial perspective, the accuracy of fuzzy marine benthic organism detection depends on the response of the marine benthic organisms and the seabed background area in the feature map. In this case, the fuzzy spatial attention module is designed to adaptively adjust the spatial weights in the feature map. Figure 3 As shown in the figure, the channel global average pooling and maximum pooling techniques are used to construct discriminative and texture spatial descriptors.

[0200] The construction process of discriminative and texture space descriptors is as follows:

[0201] Channel global average pooling is expressed as:

[0202]

[0203] Where i = 1, ..., h, j = 1, ..., w, stands for Discriminative Spatial Descriptor.

[0204] The channel global maximum pooling is expressed as:

[0205]

[0206] in, Stands for Texture Space Descriptor.

[0207] Then, the channel cascade feature map is expressed as:

[0208]

[0209] in, Represents the channel cascade feature map, symbol Represents concatenating features in the channel dimension.

[0210] Subsequently, the convolution operation under different receptive fields is expressed as:

[0211]

[0212] Among them, Ψ∈{p,q,l}, W p,i 、W q,i and W l,i Represent the weight of the i-th layer in the convolution kernel of 3×3, 5×5 and 7×7 respectively, Spatial attention maps representing different receptive fields, Represents a convolution operation.

[0213] In order to enhance the nonlinear mapping capability, the spatial attention map is activated and expressed as:

[0214] O p =s(U p ) (15)

[0215] O q =s(U q ) (16)

[0216] O l =s(U l ) (17)

[0217] in, Represents the spatial attention map after activation.

[0218] The design process of spatial fuzzy reasoner is as follows:

[0219] The activated spatial attention map O is fused using a fuzzy inference system p , O q and O l Specifically, the input and output variables are fuzzified by three fuzzy sets, namely “S”, “M” and “L” (i.e. small, medium and large), and described by Gaussian membership functions, expressed as:

[0220]

[0221] Among them, Φ∈{p,q,l,s}, μ Φ,i,j,S (·), μ Φ,i,j,M (·) and μ Φ,i,j,L (·) represents the small, medium and large membership of the spatial attention map, respectively.

[0222] Create the following fuzzy rules:

[0223] Rule 1: If o p,i,j It is S, o q,i,j It is S and o l,i,j is S, then o s,i,j For S;

[0224] Rule 2: If o p,i,j It is S, o q,i,j It is S and o l,i,j is M, then o s,i,j For S;

[0225] Rule 3: If o p,i,j It is S, o q,i,j It is S and o l,i,j is L, then o s,i,j For S;

[0226] Rule 4: If o p,i,j It is S, o q,i,j It is M and o l,i,j is S, then o s,i,j For S;

[0227] Rule 5: If o p,i,j It is S, o q,i,j It is M and o l,i,j is M, then o s,i,j For S;

[0228] Rule 6: If o p,i,j It is S, o q,i,j It is M and o l,i,j is L, then o s,i,j is M;

[0229] Rule 7: If o p,i,j It is S, o q,i,j It is L and o l,i,j is S, then o s,i,j For S;

[0230] Rule 8: If o p,i,j It is S, o q,i,j It is L and o l,i,j is M, then o s,i,j is M;

[0231] Rule 9: If o p,i,j It is S, oq,i,j It is L and o l,i,j is L, then o s,i,j is M;

[0232] Rule 10: If o p,i,j It's M, o q,i,j It is S and o l,i,j is S, then o s,i,j For S;

[0233] Rule 11: If o p,i,j It's M, o q,i,j It is S and o l,i,j is M, then o s,i,j For S;

[0234] Rule 12: If o p,i,j It's M, o q,i,j It is S and o l,i,j is L, then o s,i,j is M;

[0235] Rule 13: If o p,i,j It's M, o q,i,j It is M and o l,i,j is S, then o s,i,j For S;

[0236] Rule 14: If o p,i,j It's M, o q,i,j It is M and o l,i,j is M, then o s,i,j is M;

[0237] Rule 15: If o p,i,j It's M, o q,i,j It is M and o l,i,j is L, then o s,i,j is M;

[0238] Rule 16: If o p,i,j It's M, o q,i,j It is L and o l,i,j is S, then o s,i,j is M;

[0239] Rule 17: If o p,i,j It's M, o q,i,j It is L and o l,i,j is M, then o s,i,j is M;

[0240] Rule 18: If o p,i,j It's M, o q,i,j It is L and o l,i,j is L, then o s,i,j is L;

[0241] Rule 19: If o p,i,j It's L, o q,i,j It is S and o l,i,j is S, then o s,i,j For S;

[0242] Rule 20: If o p,i,j It's L, o q,i,j It is S and o l,i,j is M, then o s,i,j is M;

[0243] Rule 21: If o p,i,j It's L, o q,i,j It is S and o l,i,j is L, then o s,i,j is M;

[0244] Rule 22: If o p,i,j It's L, o q,i,j It is M and o l,i,j is S, then o s,i,j is M;

[0245] Rule 23: If o p,i,j It's L, o q,i,j It is M and o l,i,j is M, then o s,i,j is M;

[0246] Rule 24: If o p,i,j It's L, o q,i,j It is M and o l,i,j is L, then o s,i,j is L;

[0247] Rule 25: If o p,i,j It's L, o q,i,j It is L and o l,i,j is S, then o s,i,j is M;

[0248] Rule 26: If o p,i,j It's L, o q,i,j It is L and o l,i,j is M, then o s,i,j is L;

[0249] Rule 27: If o p,i,j It's L, o q,i,j It is L and o l,i,j is L, then o s,i,j is L;

[0250] Similarly, fuzzy reasoning is performed by taking the smaller operation (i.e., taking the “and” operation between the premises to obtain the rule conclusion), and then the aggregate credibility in the spatial fuzzy reasoning system is obtained by taking the larger operation (i.e., the union of the credibility reasoning results of each rule), which is expressed as:

[0251]

[0252] Among them, Γ∈{S,M,L}, n sfr =1,…,27 represents the fuzzy rule index.

[0253] Deblurring is performed using the centroid method, expressed as:

[0254]

[0255] in, Represents the recalibration weight at the coordinate position (i, j).

[0256] The spatial response recalibration is as follows:

[0257] The output of the FSA module is represented as:

[0258]

[0259] in, represents the recalibrated spatial weight, B(·) represents the operation of broadcasting a two-dimensional matrix into a three-dimensional matrix, Represents the features recalibrated using the FSA module.

[0260] The detection head network includes a first detection head sub-network, a second detection head sub-network, and a third detection head sub-network;

[0261] The third detection head sub-network includes a CBR5 module, a CBR3 module, and a CBR1 module;

[0262] The CBR1 module is constructed by convolution (1×1), batch normalization, and ReLU activation function;

[0263] The CBR3 module is constructed by convolution (3×3), batch normalization, and ReLU activation function;

[0264] The CBR5 module is composed of a CBR1 module, a CBR3 module, a CBR1 module, a CBR3 module and a CBR1 module in sequence;

[0265] The first detection head sub-network and the second detection head sub-network have the same structure;

[0266] The first detection head sub-network includes a first CU module and a first detection sub-network;

[0267] The output ends of the first CSFLA-8 module and the first CU module are connected to the first detection sub-network through splicing;

[0268] The first detection sub-network has the same structure as the third detection head sub-network;

[0269] The first detection head sub-network includes a second CU module and a second detection sub-network;

[0270] The output ends of the second CSFLA-8 module and the second CU module are connected to the second detection sub-network through splicing;

[0271] The input end of the CBR5 module of the third detection head sub-network is connected to the input end of the CSFLA-4 module, and the output end of the CBR5 module is connected to the input end of the second CU module and the input end of the CBR3 module of the third detection head sub-network;

[0272] The CU module is constructed by convolution, batch normalization, ReLU activation function, and upsampling, and its function is to increase the resolution of feature maps. To enhance the generalization of the detection model, two different fine-grained feature maps are spliced in the channel dimension.

[0273] For an input image with a resolution of 416×416×3, feature maps with dimensions of 13×13×27, 26×26×27, and 52×52×27 are finally obtained, which are used to detect large-scale, medium-scale, and small-scale marine benthic organisms, respectively.

[0274] Example 1:

[0275] 1. Dataset and experimental configuration;

[0276] The present invention verifies the effectiveness and superiority of the CSFLA method on the URPC2019 and DUO datasets, both of which include four categories (i.e., sea urchins, scallops, starfish, and sea cucumbers). The URPC2019 and DUO datasets contain 4757 and 7782 images, respectively. For the URPC2019 dataset, the training set contains 3806 images and the test set contains 951 images; for the DUO dataset, the training set contains 6671 images and the test set contains 1111 images. The corresponding experimental configuration parameters are shown in Table 1. The present invention uses AP, mAP@.5, mAP@.75, and mAP@[.5,.95] as detection performance evaluation indicators. Using a five-fold cross-validation strategy, the average of the five test results is taken as the detection accuracy of different models; a video with a resolution of 1280×720 is used to obtain the inference speed of the model.

[0277] Table 1

[0278]

[0279] 2. Selection of fuzzy membership width parameter value

[0280] In order to determine the optimal value of the fuzzy membership function width parameter σ, the present invention selects from the candidate set {0.05, 0.1, 0.2, 0.3, 0.4, 0.5}. As can be seen from Tables 2 and 3, σ = 0.05 or σ = 0.1 makes the designed membership function too narrow, resulting in poor detection performance, while σ = 0.4 or σ = 0.5 makes the designed membership function too wide, which also leads to poor detection performance. The results show that choosing σ = 0.2 and σ = 0.3 is more desirable. When σ = 0.2, the most superior mAP@.75 and mAP@[.5,.95] indicator values can be achieved on both the URPC2019 dataset and the DUO dataset. For convenience, the present invention finally sets the parameter σ to 0.2.

[0281] Table 2

[0282]

[0283] Table 3

[0284]

[0285] 3. Fuzzy attention performance

[0286] To demonstrate the effectiveness and superiority of the FCA and FSA modules in fusing different types of channel attention maps and spatial attention maps with different receptive fields, we consider the Algebraic Summarization-based Channel Attention (ASCA) and Algebraic Summarization-based Spatial Attention (ASSA) methods. For comparison purposes, we use the YOLOv3 model to obtain baseline detection accuracy.

[0287] Tables 4 and 5 summarize the detection results on the URPC2019 and DUO datasets, respectively. Regarding the detection performance on the URPC2019 dataset, Table 4 shows that the baseline model's detection accuracy is significantly lower than that of the ASCA- and ASSA-based detection methods. Specifically, in terms of mAP@.5, mAP@.75, and mAP@[.5,.95] metrics, the ASCA-based detection accuracy is 1.17%, 1.29%, and 0.67% higher than the baseline model, respectively. This demonstrates that the ASCA module is able to enhance the characteristics of ambiguous marine benthic organisms at the channel level while suppressing seafloor background noise. Similarly, the ASSA-based detection accuracy is 0.91%, 1.20%, and 1.65% higher than the baseline model, respectively, demonstrating that the ASSA module is able to enhance the regional response of ambiguous marine benthic organisms at the spatial level. More importantly, the proposed FCA and FSA modules achieve higher detection accuracy than the ASCA and ASSA modules in terms of mAP@.5, mAP@.75, and mAP@[.5,.95] metrics. This demonstrates that, with the help of the fuzzy inference system, the FCA and FSA modules can more robustly enhance fuzzy marine benthic biosignatures and suppress seabed background noise at the channel and spatial levels, respectively. Regarding the DUO dataset, Table 5 shows that, except for the mAP@.75 and sea urchin metrics, the proposed FCA and FSA modules achieve superior detection performance compared to the ASCA and ASSA modules.

[0288] Table 4

[0289]

[0290] Table 5

[0291]

[0292]

[0293] 4. Quantitative comparison of classic attention detection methods

[0294] To demonstrate the superiority of the CSFLA method, the present invention conducts comprehensive comparative experiments with SENet, CBAM, and SKNet methods on the URPC2019 and DUO datasets. Tables 6 and 7 summarize the corresponding comparison results. It can be seen that the detection accuracy of the SENet model is significantly lower than that of the proposed CSFLA method. The main reason is that the SENet model only uses spatial global average pooling technology to extract discriminative information, making it difficult to construct a comprehensive channel descriptor. In addition, although the CBAM model uses both spatial attention and channel attention technologies, its detection performance is still weaker than that of the proposed CSFLA method. The important reasons are summarized as follows: the number of hidden layer nodes in the channel attention is less than that in the input and output layers, which destroys the direct connection between the channel weights and the corresponding feature maps; single-scale convolution kernels are difficult to simultaneously generate large-scale, medium-scale, and small-scale regional responses. In addition, although SKNet can adaptively select the receptive field of the convolution kernel, its detection accuracy is still lower than the proposed CSFLA method. The main reason is that the SKNet model adopts algebraic sum strategy to fuse feature maps with different receptive fields. Especially when the visual imaging system introduces fuzzy perceptual information, the robustness and generalization of the SKNet model are poor.

[0295] Table 6

[0296]

[0297] Table 7

[0298]

[0299] Figure 4 The performance comparison chart of the accuracy-recall curves of different categories, where (a) is the accuracy-recall curve of sea urchins; (b) is the accuracy-recall curve of scallops; (c) is the accuracy-recall curve of starfish; and (d) is the accuracy-recall curve of sea cucumbers.

[0300] In addition, the present invention carried out a precision-recall performance comparison experiment. Figure 4As can be seen from the figure, by combining channel attention and spatial attention in sequence, the CBAM model achieves better detection performance than SENet. The main reason is that SENet only enhances the feature map of marine benthic organisms at the channel level, while the CBAM model enhances the feature map and regional response related to marine benthic organisms at the channel and spatial levels respectively. In addition, the detection performance of the SKNet model surpasses the CBAM model. The key reason is that the SKNet model can adaptively select the receptive field and realize multi-scale feature extraction. Overall, the proposed CSFLA method can achieve an optimal balance between precision and recall. The main reasons include: 1) The FCA module recalibrates the channel weights, improving the detection performance of blurred marine benthic organisms at the channel level; 2) The FSA module recalibrates the spatial weights, suppresses the seabed background noise, and improves the detection performance of blurred marine benthic organisms at the spatial level.

[0301] 5. Quantitative comparison of specialized marine benthic organism detection methods

[0302] To further validate the superiority of the CSFLA method, we comprehensively considered specialized marine benthic biota detection methods, including TVFRD, Pruning-SSD, Quantized-SSD, YOLOv3-UW, DG-YOLO, and BBROABR. Tables 8 and 9 summarize the experimental results of different algorithms on the URPC2019 and DUO datasets, respectively. As shown in Table 8, at a resolution of 608×608, the proposed CSFLA method not only achieves the best mAP@.5 and mAP@[.5,.95] metrics, but also achieves the best performance in detecting sea urchins. By leveraging generalized intersection-over-union (GIU) and prior anchor boxes for marine benthic biota, the BBROABR method comprehensively considers both localization and recall performance, achieving the best performance in detecting scallops and starfish. However, the BBROABR method lacks an attention mechanism to enhance the blurred features of marine benthic biota, making it difficult to achieve the highest comprehensive evaluation metric (i.e., mAP@[.5,.95]). Furthermore, by utilizing a domain-invariant module, DG-YOLO achieves optimal performance in detecting sea cucumbers. However, due to the lack of spatial attention to enhance the spatial response of benthic organisms, the YOLOv3-UW method, which uses only channel attention, struggles to achieve optimal detection performance. Using a depthwise separable convolution strategy and weight mapping techniques, the Quantized-SSD method achieves the fastest detection speed (31 FPS) with an input image resolution of 512×512. By using kernel pruning to remove low-contributing neurons, the Pruning-SSD method, which does not utilize channel and spatial attention, also struggles to achieve superior detection performance. The TVFRD method utilizes a region proposal network to pre-generate a large number of candidate detection boxes. While this improves detection accuracy, it only achieves a detection speed of 3 FPS, which is insufficient for real-time detection. Regarding detection performance on the DUO dataset, Table 9 shows that, with the exception of the sea cucumber category, the CSFLA method achieves superior detection performance in all other metrics.

[0303] Table 8

[0304]

[0305] Table 9

[0306]

[0307] 6. Quantitative comparison of common target detection methods

[0308] To illustrate the superiority of the CSFLA method, this paper comprehensively considers two-stage detection methods (Faster R-CNN), single-stage detection methods (SSD, EfficientDet, YOLOv2, YOLOv3, YOLOv4, CEFP2N, and M2YOLOF), and anchor-free detection methods (CenterNet and FCOS). Tables 10 and 11 summarize the detection performance of different methods on the URPC2019 and DUO datasets. It can be seen that the anchor-free CenterNet and FCOS methods have difficulty achieving superior detection performance, mainly because they do not utilize prior anchor box information during training. Although the use of the context enhancement module and feature purification module can enhance texture information, the detection performance of the CEFP2N method is significantly weaker than the proposed CSFLA method. The key reason is that the CEFP2N method uses algebraic sum strategies to construct channel attention weights and spatial attention weights. In particular, when the visual imaging system introduces fuzzy perceptual information, it is difficult to derive appropriate channel weights and spatial weights. Furthermore, by designing a multi-input single-output encoder and a dynamic sample selection strategy, the M2YOLOF method effectively improves global representation capabilities and enhances the quality of positive samples, thereby improving small object detection performance on a general object detection dataset. Except for the mAP@.75 metric, the M2YOLOF method's detection performance is significantly weaker than the proposed CSFLA method, demonstrating the important role of fuzzy channel attention and fuzzy spatial attention in enhancing the characteristics of marine benthic organisms and suppressing seabed background noise. On both the URPC2019 and DUO datasets, the YOLOv4 model achieves the best detection accuracy for the sea urchin category. This is primarily due to the YOLOv4 model's Mosaic technique, which randomly mixes four images, which is particularly beneficial for increasing the diversity of training samples for a large number of categories. Furthermore, since the YOLOv3 model does not employ channel and spatial attention strategies, feature maps related to fuzzy benthic organisms cannot be enhanced, thus hindering superior detection performance. Compared to using a VGG-16 backbone network, SSD using MobileNet v2 achieves the fastest detection speed (i.e., 38 FPS). Unlike single-stage detection methods, Faster R-CNN relies on a region proposal network to generate a large number of candidate boxes that are independent of category, making real-time performance difficult to guarantee. Overall, the proposed CSFLA method achieves optimal results in most metrics, primarily due to the more comprehensive channel and spatial weights derived from the FCA and FSA modules.

[0309] Table 10

[0310]

[0311]

[0312] Table 11

[0313]

[0314] 7. Qualitative comparison of marine benthic organism detection

[0315] In order to verify the detection superiority of the proposed CSFLA method in real environments, this paper considers TVFRD, Pruning-SSD, Quantized-SSD, YOLOv3-UW and DG-YOLO methods. The corresponding experiments are carried out in dark light scenes, bright light scenes, multi-scale target coexistence scenes and blurred scenes. Figure 5 The following is a qualitative comparison of detection performance in typical scenarios, where (a) is a dark light scene, (b) is a bright light scene, (c) is a multi-scale target coexistence scene, and (d) is a blurred scene.

[0316] like Figure 5 As shown in the figure, in low-light scenarios, although the CSFLA method cannot detect all benthic organisms, its detection performance far exceeds that of other methods. The TVFRD method achieves better detection performance in bright-light scenarios, primarily due to the large number of candidate regions generated by the region proposal network in the first stage. Regarding performance comparisons in scenarios with multi-scale coexistence, particularly for detecting benthic organisms obscured by seagrass, the proposed CSFLA method achieves the best detection performance, primarily due to its ability to suppress seabed background noise while enhancing the characteristics of benthic organisms. Thanks to the CSFLA method's generalization, nearly consistent conclusions can be drawn in ambiguous scenarios.

[0317] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

[0318] [1]Li C, Anwar S, Hou J, et al. Underwater image enhancement via mediumtransmission-guided multi-color space embedding[J]. IEEE Transactions on ImageProcessing, 2021, 30: 4985-5000.

[0319] [2]Zhuang P,Wu J,Porikli F,et al.Underwater image enhancement withhyper-laplacian reflectance priors[J].IEEE Transactions on Image Processing,2022,31:5442-5455.

[0320] [3]Wang N,Chen T,Kong X,et al.Underwater attentional generativeadversarial networks for image enhancement[J].IEEE Transactions on Human-Machine Systems,2023,53(3):490-500.

[0321] [4]Zhou J,Yang T,Chu W,et al.Underwater image restoration viabackscatter pixel prior and color compensation[J].Engineering Applications ofArtificial Intelligence,2022,111:104785.

[0322] [5]Liu C,Wang Z,Wang S,et al.A new dataset,Poisson GAN and AquaNetfor underwater object grabbing[J].IEEE Transactions on Circuits and Systemsfor Video Technology,2022,32(5):2831-2844.

[0323] [6]Chen W,Fan B.Underwater object detection with mixed attentionmechanism and multi-enhancement strategy[C].Chinese Automation Congress,Shanghai,China,2020:2821-2826.

[0324] [7]Peng F,Miao Z,Li F,et al.S-FPN:A shortcut feature pyramid networkfor sea cucumber detection in underwater images[J].Expert Systems withApplications,2021,182:115306.

[0325] [8]Cheng N,Xie H,Zhu X,et al.Joint image enhancement learning formarine object detection in natural scene[J].Engineering Applications ofArtificial Intelligence,2023,120:105905.

[0326] [9]Saha S,Bovolo F,Bruzzone L.Building change detection in VHR SARimages via unsupervised deep transcoding[J].IEEE Transactions on Geoscienceand Remote Sensing,2021,59(3):1917-1929.

[0327]

[10] Huang K,Zhang Y,Cheng H,et al.Semantic segmentation of breastultrasound image with fuzzy deep learning network and breast anatomyconstraints[J].Neurocomputing,2021,450:319-335.

[0328]

[11] Zhou Y,Mao A,Huo S,et al.Salient object detection via fuzzytheory and object-level enhancement[J].IEEE Transactions on Multimedia,2018,21(1):74-85.

[0329]

[12] Zhang J,Zhu L,Xu L,et al.MFFSSD:An enhanced SSD for underwaterobject detection[C].ChineseAutomation Congress,Shanghai,China,2020:5938-5943.

[0330]

[13] Hua X,Cui X,Xu X,et al.Underwater object detection algorithmbased on feature enhancement and progressive dynamic aggregation strategy[J].Pattern Recognition,2023,139:109511.

[0331]

[14] Chen L,Zhou F,Wang S,et al.SWIPENET:Object detection in noisyunderwater scenes[J].Pattern Recognition,2022,132:108926.

[0332]

[15] Liu Z,Zhuang Y,Jia P,et al.A novel underwater image enhancementalgorithm and an improved underwater biological detection pipeline[J].Journalof Marine Science and Engineering,2022,10(9):1204.

[0333]

[16] Cai S,Li G,Shan Y.Underwater object detection using collaborativeweakly supervision[J].Computers and Electrical Engineering,2022,102:108159.

Claims

1. A method for progressive detection enhancement of marine benthic organisms guided by channel-spatial fuzzy logic attention under detail loss, characterized by: The following steps are involved: Construct a dataset based on images of one or more of the following marine benthic organisms: sea urchins, scallops, starfish, and sea cucumbers; Divide the dataset into: training set and test set; Construct a progressive detection enhancement model for marine benthic organisms, including detail loss channels and spatial fuzzy logic attention guidance; Based on the training set data, the channel-spatial fuzzy logic attention-guided marine benthic organisms progressive detection enhancement model is trained to obtain the trained channel-spatial fuzzy logic attention-guided marine benthic organisms progressive detection enhancement model; The test set data is input into the trained channel-spatial fuzzy logic attention-guided progressive detection enhancement model for marine benthic organisms to achieve detection and recognition of sea urchins, scallops, starfish, and sea cucumbers. The detailed loss channel and spatial fuzzy logic attention guided marine benthic organism progressive detection enhancement model includes: Backbone network: extracts three different fine-grained feature maps from images of marine benthic organisms; Detection head network: performs feature fusion based on the three different fine-grained feature maps output by the backbone network to detect large, medium, and small-sized marine benthic organisms; The backbone network includes a first CBR3 module, a second CBR3 module, a CSFLA-1 module, a third CBR3 module, a CSFLA-2 module, a fourth CBR3 module, a first CSFLA-8 module, a fifth CBR3 module, a second CSFLA-8 module, a sixth CBR3 module, and a CSFLA-4 module; The CSFLA-1 module, CSFLA-2 module, first CSFLA-8 module, second CSFLA-8 module and CSFLA-4 module represent CSFLA modules reused 1, 2, 8, 8 and 4 times in sequence respectively; The CSFLA module includes 3×3 convolution, 1×1 convolution, fuzzy channel attention FCA module and fuzzy spatial attention FSA module, and adopts a skip connection strategy to prevent semantic information degradation; The fuzzy channel attention FCA module is used to combine spatial global average pooling and maximum pooling methods to construct discriminative and textural channel descriptors for fuzzy marine benthic features, and a channel fuzzy inference system is designed to fuse the discriminative and textural channel attention maps derived from two single hidden layer networks. The fuzzy spatial attention (FSA) module uses channel global average pooling and maximum pooling methods to construct discriminative and textural spatial descriptors of fuzzy marine benthic features, and designs a spatial fuzzy inference system to fuse spatial attention maps with different receptive fields.

2. The method for progressive detection enhancement of marine benthic organisms guided by channel-spatial fuzzy logic attention under detail loss according to claim 1, characterized in that: The first CBR3 module, the second CBR3 module, the CSFLA-1 module, the third CBR3 module, the CSFLA-2 module, the fourth CBR3 module, the first CSFLA-8 module, the fifth CBR3 module, the second CSFLA-8 module, the sixth CBR3 module and the CSFLA-4 module are connected in sequence; The first CSFLA-8 module output, the second CSFLA-8 module output and the CSFLA-4 module output three different fine-grained feature maps respectively.

3. The method for progressive detection enhancement of marine benthic organisms guided by channel-spatial fuzzy logic attention under detail loss according to claim 1, characterized in that: The discriminative and texture channel descriptors are constructed as follows: Two independent single-hidden-layer feedforward neural networks with the same network structure are used to learn the discriminative and texture channel descriptors derived from spatial global average pooling and maximum pooling operations respectively. The output of spatial global average pooling is expressed as: (1) in, Representative The average value of the layer feature map, , , and are the height and width of the current feature map, respectively. represents the input features, Representative Layer feature map, represents the discriminative channel descriptor, Represents the feature layer index range from arrive ; The output of a single hidden layer feed-forward network that learns a discriminative channel descriptor is represented as: (2) in, represents the discriminative channel attention map, and Represent the hidden layer weights and output layer weights respectively, represents the Leaky ReLU activation function, Representatives will Sigmoid activation function limited to the range (0, 1); Accordingly, the output of the spatial global maximum pooling operation, the texture channel descriptor, is expressed as: (3) in, represents the texture channel descriptor, Represents the maximum value operation, Representative The maximum value in the layer feature map; Similarly, the output of a single hidden layer feed-forward network for learning texture channel descriptors is expressed as: (4) in, represents the texture channel attention map, and Represents the hidden layer and output layer weights, activation function Will Clamped to the range (0, 1).

4. The method for progressive detection enhancement of marine benthic organisms guided by channel-spatial fuzzy logic attention under detail loss according to claim 1, characterized in that: The design process of the channel fuzzy inference system is as follows: Fusion of discriminative channel attention maps and texture channel attention map , the input variables and output variables are fuzzified by three fuzzy sets, namely S, M and L, namely small, medium and large, and are characterized by Gaussian membership functions, expressed as: (5) (6) (7) in, , 、 and Represent the small, medium and large membership of the channel attention map respectively, represents the standard deviation, represents the natural exponential function, The representative fuzzy set is a small Gaussian membership function center, represents the center of the Gaussian membership function in the fuzzy set, The representative fuzzy set is a large Gaussian membership function center; The following fuzzy rules are established: in, Represents the discriminative channel attention map layer weights, Represents the texture channel attention map layer weights, Represents the channel attention map after fuzzy reasoning layer weights; Fuzzy reasoning is performed using the smaller operation, and the aggregated credibility is obtained by the larger operation, which is the union of the credibility reasoning results of each rule, expressed as: (8) in, , represents the fuzzy rule index; Deblurring is performed using the centroid method, which is expressed as: (9) in, Representative The recalibration weight of each channel; Furthermore, the channel features are recalibrated as follows: The output of the FCA module is represented as: (10) in, represents the recalibrated channel weights, Represents broadcasting a one-dimensional vector into a three-dimensional matrix, represents element-wise multiplication, Represents the features recalibrated using the FCA module.

5. The method for progressive detection enhancement of marine benthic organisms guided by channel-spatial fuzzy logic attention under detail loss according to claim 1, characterized in that: The discriminative and texture space descriptors are constructed as follows: Channel global average pooling is expressed as: (11) in, , , represents the discriminative space descriptor, Represents the number of channel layers; The channel global maximum pooling is expressed as: (12) in, stands for Texture Space Descriptor; Then, the channel cascade feature map is expressed as: (13) in, Represents the channel cascade feature map, symbol Represents the concatenation of features in the channel dimension; Subsequently, the convolution operation under different receptive fields is expressed as: (14) in, , 、 and Represents the convolution kernels of 3×3, 5×5 and 7×7 respectively. layer weights, Spatial attention maps representing different receptive fields, Represents the convolution operation; In order to enhance the nonlinear mapping capability, the spatial attention map is activated and expressed as: (15) (16) (17) in, Represents the spatial attention map after activation.

6. The method for progressive detection and enhancement of marine benthic organisms guided by channel-spatial fuzzy logic attention under detail loss according to claim 1, characterized in that: The design process of the spatial fuzzy reasoning system is as follows: Fuzzy inference system is used to fuse the activated spatial attention map 、 and , the input variables and output variables are fuzzified by three fuzzy sets, namely S, M and L, namely small, medium and large, and are described by Gaussian membership functions, expressed as: (18) (19) (20) in, , 、 and Represent the small, medium and large membership of the spatial attention map respectively; Create the following fuzzy rules: Fuzzy reasoning is performed by taking the smaller operation, and then the larger operation is used to obtain the aggregate credibility in the spatial fuzzy reasoning system, that is, the union of the credibility reasoning results of each rule, which is expressed as: (21) in, , represents the fuzzy rule index; Deblurring is performed using the centroid method, expressed as: (22) in, The representative coordinate position is The recalibration weight at ; The output of the FSA module is represented as: (23) in, represents the recalibrated spatial weight, Represents the operation of broadcasting a two-dimensional matrix into a three-dimensional matrix, Represents the features recalibrated using the FSA module.

7. A method and device for progressive detection enhancement of marine benthic organisms guided by channel-spatial fuzzy logic attention under detail loss, characterized by: include: Building Module I: Constructing a dataset based on images of one or more marine benthic organisms including sea urchins, scallops, starfish, and sea cucumbers; Partitioning module: divide the data set into training set and test set; Building Module II: Constructing a channel-spatial fuzzy logic attention-guided progressive detection enhancement model for marine benthic organisms; Training module: Based on the training set data, the channel-spatial fuzzy logic attention-guided marine benthic organisms progressive detection enhancement model is trained to obtain the trained channel-spatial fuzzy logic attention-guided marine benthic organisms progressive detection enhancement model; Detection module: The test set data is input into the trained channel-spatial fuzzy logic attention-guided progressive detection enhancement model for marine benthic organisms to detect and identify sea urchins, scallops, starfish, and sea cucumbers. The detailed loss channel and spatial fuzzy logic attention guided marine benthic organism progressive detection enhancement model includes: Backbone network: extracts three different fine-grained feature maps from images of marine benthic organisms; Detection head network: performs feature fusion based on the three different fine-grained feature maps output by the backbone network to detect large, medium, and small-sized marine benthic organisms; The backbone network includes a first CBR3 module, a second CBR3 module, a CSFLA-1 module, a third CBR3 module, a CSFLA-2 module, a fourth CBR3 module, a first CSFLA-8 module, a fifth CBR3 module, a second CSFLA-8 module, a sixth CBR3 module, and a CSFLA-4 module; The CSFLA-1 module, CSFLA-2 module, first CSFLA-8 module, second CSFLA-8 module and CSFLA-4 module represent CSFLA modules reused 1, 2, 8, 8 and 4 times in sequence respectively; The CSFLA module includes 3×3 convolution, 1×1 convolution, fuzzy channel attention FCA module and fuzzy spatial attention FSA module, and adopts a skip connection strategy to prevent semantic information degradation; The fuzzy channel attention FCA module is used to combine spatial global average pooling and maximum pooling methods to construct discriminative and textural channel descriptors for fuzzy marine benthic features, and a channel fuzzy inference system is designed to fuse the discriminative and textural channel attention maps derived from two single hidden layer networks. The fuzzy spatial attention (FSA) module uses channel global average pooling and maximum pooling methods to construct discriminative and textural spatial descriptors of fuzzy marine benthic features, and designs a spatial fuzzy inference system to fuse spatial attention maps with different receptive fields.