Method for efficiently detecting fish target in complex underwater environment based on priori knowledge guidance network
Through the recovery subnet of the prior knowledge guide network, the relationship reasoning attention and adaptive feature fusion module, the light scattering and environmental complexity problems of underwater fish target detection in the Yangtze River are solved, and high-precision and robust fish target detection are achieved.
Patent Information
- Application Number
- CN202510446075.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional underwater fish target detection methods are difficult to detect in complex environments like the Yangtze River, with severe light scattering and poor image quality. The existing imaging models are insufficiently generalized, resulting in low detection accuracy.
The method based on prior knowledge guidance network is adopted to remove underwater turbidity characteristics through the recovery subnet module, the relational reasoning attention module supplements visual characteristics, the adaptive feature fusion module optimizes feature expression, and combines the underwater scattering model and graph convolution network for fish target detection.
In complex underwater environments, the accuracy and robustness of fish target detection are significantly improved, false detection and missed detection are reduced, and detection performance is adapted to different underwater conditions.
Smart Images

Figure CN120259867A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and image processing, and particularly relates to an efficient method for detecting fish targets in complex underwater environments based on a prior knowledge-guided network. Background Art
[0002] The technology for detecting fish targets in complex underwater environments is the core key to realizing the ecological protection of the Yangtze River and the rational utilization of fishery resources, and is of great significance in many fields such as Yangtze River ecological research and intelligent fishery management. Most traditional underwater fish target detection methods rely on manually designed features, such as methods based on edge detection operators and shape descriptors, and complete target recognition by analyzing the appearance features of fish. However, the underwater environment of the Yangtze River is extremely complex and ever-changing. The turbidity of the water causes extremely serious light scattering. In addition, the fish have a wide range of activities and rich and diverse behavior patterns, which greatly increase the detection difficulty.
[0003] At the same time, various environmental factors have extremely adverse effects on the quality of underwater images. From the color level, the overall image shows an obvious greenish-blue hue; in terms of contrast and details, the contrast is low and the details are blurred. The physical imaging mechanism of underwater images has many similarities with fogged images. When light propagates in water, it is scattered by particles in the water, resulting in a deviation between the light captured by the camera and the actual reflected light of the object, thereby greatly reducing the quality of the acquired image. Some researchers have tried to apply the physical imaging model of fogged images to underwater images in the hope of achieving color correction. However, due to the completely different propagation characteristics of light in water and in the atmosphere, the existing imaging models do not fully consider the differences in the scattering effects of light with different wavelengths in water, resulting in unsatisfactory correction effects. It was not until researchers such as Akkaynak [4] conducted a large number of field underwater experiments, corrected the attenuation coefficient of the existing model, and proposed an improved underwater image formation model that the situation was improved to a certain extent. The above methods based on physical imaging models are often only applicable to specific images and do not have generalization. Summary of the Invention
[0004] Object of the Invention: The technical problem to be solved by the present invention is to provide an efficient method for detecting fish targets in complex underwater environments based on a prior knowledge-guided network in view of the deficiencies of the prior art, including the following steps:
[0005] Step 1, obtain an underwater image and perform data preprocessing, where the preprocessing includes simulating underwater condition enhancement and simulating underwater effects;
[0006] Step 2, establish a restoration subnet module, a relationship reasoning attention module, and an adaptive feature fusion module to detect fish targets in complex underwater environments.
[0007] The restoration subnet module is used to guide the network to learn underwater turbidity features through the underwater scattering model, generate clear images through the water body restoration decoder WR, and reconstruct underwater turbidity images by combining the background ambient light decoder BAL, the water body transmittance decoder WTM, and the backscattered light decoder BE, so as to enhance the model's understanding of underwater turbidity characteristics during the training process and reduce the impact of underwater turbidity on detection performance;
[0008] The relationship reasoning attention module is used to construct a co-occurrence relationship graph (such as the spatial or semantic association between objects), perform relationship reasoning using a graph convolutional network (GCN), dynamically adjust the attention weights, and guide the model to focus on potentially co-occurring targets (such as perch and grass carp), so as to supplement the detection omissions caused by blurred visual features in underwater turbidity scenes;
[0009] The adaptive feature fusion module is used to dynamically adjust the multi-scale feature fusion ratio of the output of the backbone network and the neck network through learnable parameters, combine the Sigmoid function and the per-channel weighting mechanism, optimize the feature expression, and enhance the model's integration ability for different levels of features (such as local details and global semantics).
[0010] In step 1, underwater images are selected from the public dataset (directly screen the required clear underwater images containing fish from the public dataset, such as images of grass carp and snakehead) to form the dataset train.
[0011] In step 1, the simulation of underwater conditions enhancement includes: adding simulated turbidity effects to the images according to the turbidity data measured in different regions of the Yangtze River; adjusting the illumination intensity and color tone of the images according to the illumination change rules in different seasons and time periods to simulate underwater scenes under different illumination conditions.
[0012] In step 1, the simulation of underwater effects includes: based on the dataset train, simulating the influence of the underwater environment on the images, simulating common color deviations underwater by adjusting the color channels of the images to make the images overall greenish and bluish; using an improved underwater image scattering model (an improvement of the Jaffe-McGlamery model), processing the clear images X1 times respectively to simulate different degrees of scattering and backscattering effects, changing the clarity and contrast of the images, and reflecting underwater environments with different turbidities. The formula of the improved underwater scattering model is:
[0013] I(x) = J(x)t(x) + A·(1 - t(x)) + B(x),
[0014] Where \(I(x)\) represents the observed image, \(J(x)\) represents the clear image without scattering, \(t(x)\) represents the transmittance, \(A\) represents the background light, and \(B(x)\) represents the backscattered light; \(x\) represents the spatial position coordinates of a pixel in the image, which is a two-dimensional coordinate \((i, j)\), where \(i\) and \(j\) correspond to the rows and columns of the image respectively;
[0015] The improved underwater scattering model separates the background light \(A\) and represents the relationship between the background light \(A\) and the transmittance by \(A\cdot(1 - t(x))\);
[0016] The backscattering \(B(x)\) of the traditional model usually includes the comprehensive effect of all scattered light. The improved underwater scattering model splits it into two parts: \(A\cdot(1 - t(x))\): representing the uniform scattering component generated by the attenuation of the background light due to the transmittance; \(B(x)\): reserved as the spatially varying local backscattering (such as the random reflection of suspended particles).
[0017] Step 2 includes:
[0018] Step 2-1, divide the dataset train into a training set, a validation set, and a test set;
[0019] Step 2-2, create a file named mydata.yaml, configure the storage paths of the training set, validation set, and test set in the file, and at the same time clarify the class labels of various types of fish and related underwater objects;
[0020] Generate the model.yaml file, determine the total number of classes of fish and related objects to be detected, improve the Yolov7 model. First, define the network structure of the Yolov7 model, including the parameter settings and connection methods of the backbone network backbone, the head, and the Feature Pyramid Network Neck (Neck uses the FPN structure); add a Graph Convolutional Network (GCN) module and a decoder module to the detection head of the Yolov7 model. The graph convolutional network module is used to perform graph convolutional operations on the feature map; the decoder module is used to perform upsampling and feature fusion on the feature map;
[0021] Step 2-3, establish a recovery subnet module in the backbone network backbone. The recovery subnet module uses the improved underwater scattering model to calculate the transmittance \(t(x)\) through the following formula:
[0022] \(t(x)=e\) -β·q ,
[0023] where \(e\) is the natural constant, \(\beta\) is the attenuation coefficient of the water body, which can be processed channel by channel and can be extended to \(\beta\) c (such as \(\beta_{red},\beta_{green},\beta_{blue}\)), \(q\) is the propagation distance from the target to the camera, and the calculation formula is:
[0024]
[0025] Among them, Center[0] represents the abscissa of the center point, and Center[1] represents the ordinate of the center point;
[0026] Assign values to the background light A adaptively according to the water body type:
[0027]
[0028] The calculation formula for the backscattered light B(x) is:
[0029]
[0030] Among them, l represents the scattering intensity coefficient of suspended particles, which can be processed by channels and can be extended to m c (such as m_red, m_green, m_blue);
[0031] Step 2-4, the restoration subnet module further includes a water body restoration decoder WR (Water Restoration, whose main function is to restore the clarity of the underwater image and remove the influence of the turbid water body on the image), a background ambient light decoder BAL (Background Ambient Light, used to estimate the background ambient light in the underwater scene), a water body transmittance decoder WTM (Water Transmission Map, estimating the light transmission in the water body and used to describe the attenuation degree of light in the water body), and a backscattered light decoder BE (Backscatter Estimation, used to estimate and process the backscattered light effect in underwater imaging);
[0032] The background ambient light decoder BAL uses an adaptive max pooling layer and a convolutional block ConvBlock composed of a convolutional layer, a batch normalization layer, and a Tanh activation function to estimate the background light A;
[0033] The water body restoration decoder WR uses more than two transposed convolutional blocks DeConvBlocks to estimate the clear image J(x); the transposed convolutional blocks DeConvBlocks include transposed convolution, batch normalization, and ReLU activation functions;
[0034] The water body transmittance decoder WTM estimates the transmittance t(x) through two convolutional blocks ConvBlock;
[0035] The backscattered light decoder BE uses a convolutional block ConvBlock to estimate the backscattered light B(x);
[0036] Use the upsampling operation to align feature maps of different resolutions to the same size for fusion. (In a convolutional neural network, the feature maps output at different levels (shallow, middle, deep) have different spatial resolutions, and these feature maps come from the backbone network and the Feature Pyramid Network Neck (Neck uses the FPN structure).) Through the adaptive feature fusion mechanism, the features extracted from the underwater scattering model and the co-occurrence relationship map are effectively fused with the features of the backbone network. (The backbone network is the main feature extraction network, which is the CSPDarknet structure derived from YOLOv7 and contains multiple CSP (Cross Stage Partial) modules, generating feature maps of different resolutions through convolution and downsampling.)
[0037] Combine the restoration subnet module with the detection head to form an end-to-end training framework. The output of the restoration subnet module serves as the input of the detection head, and the performance feedback of the detection head guides the restoration subnet module to further optimize the restoration effect.
[0038] Step 2-5: Establish a relationship reasoning attention module.
[0039] Step 2-6: Establish an adaptive feature fusion module.
[0040] In step 2-4, the restoration subnet module performs the following steps:
[0041] Step 2-4-1: Data processing and depth map generation: Process the input underwater image \(W_{Raw}\in R^{H,W,3}\), where \(H\) is the height, \(W\) is the width, 3 represents the three RGB channels, and \(R\) is the real number space; H×W×3 Calculate the image center coordinates. Then, for each pixel coordinate \((i,j)\), construct the Euclidean distance matrix \(D(i,j)\). Let the maximum value of the Euclidean distance matrix \(D(i,j)\) be \(D_{max}=\max\{D(i,j)\}\). Then define the normalized distance \(d(i,j)=D_{max} / D(i,j)\), and use the exponential mapping to generate the depth weight matrix \(d_{norm}(i,j)\):
[0042] Calculate the image center coordinates. Then, for each pixel coordinate \((i,j)\), construct the Euclidean distance matrix \(D(i,j)\). Let the maximum value of the Euclidean distance matrix \(D(i,j)\) be \(D_{max}=\max\{D(i,j)\}\). Then define the normalized distance \(d(i,j)=D_{max} / D(i,j)\), and use the exponential mapping to generate the depth weight matrix \(d_{norm}(i,j)\):
[0043] \(d_{norm}(i,j)=\exp(-\lambda\cdot d(i,j))\),
[0044] where \(\lambda\) is the scale factor; \(\exp\) represents the natural exponential function; the generated depth weight matrix \(d_{norm}(i,j)\) will be used as the input for the subsequent decoders (BAL, WR, WTM, BE) to estimate the transmittance, background light, and scattered light;
[0045] Step 2-4-2: The restoration subnet module contains the following four decoders:
[0046] The background ambient light decoder BAL, used to estimate the background light \(A\);
[0047] The water body restoration decoder WR is used to estimate the clear image J(x);
[0048] The water body transmittance decoder WTM is used to estimate the transmittance t(x);
[0049] The backscattered light decoder BE is used to estimate the backscattered light B(x);
[0050] The working process of each decoder is as follows: estimating the transmittance, background light, and scattered light: through the water body transmittance decoder WTM, the input feature map F (the input feature map F is the result of fusing multi-level features output by the backbone network) passes through two cascaded convolutional blocks ConvBlock (including a convolutional layer and an activation function), and the preliminary feature map Ft1 is output. The Ft1 is upsampled to the H×W resolution using the upsampling operation and mapped to a three-channel transmittance map
[0051] For each RGB channel, the transmittance is calculated using the following formula respectively:
[0052] t c (i,j) = exp(-β c ·dnorm(i,j)),
[0053] where the color channel parameter c ∈ {red, green, blue}; t c (i,j) represents the transmittance value when the color channel is c, which is used to model the attenuation difference of light with different wavelengths underwater; β c represents the attenuation coefficient of the color channel c;
[0054] The transmittance matrix t c ∈R H×W is obtained, and the transmittance calculation formula is:
[0055] t(x) = ConvBlock2(ConvBlock1(Concat(F, dnorm(i,j)))),
[0056] where Concat represents the concatenation operation; ConvBlock1 represents a convolutional block composed of a 3×3 convolution, a ReLU activation function, and batch normalization BatchNorm, and ConvBlock2 represents a convolutional block composed of a 1×1 convolution and a Sigmoid function;
[0057] Set fixed values for the three RGB channels, and then dynamically estimate the background light A through the background ambient light decoder BAL. Adaptive max pooling is performed on the feature map F of the input image to obtain a global feature vector, which is convolved, batch-normalized, and activated by Tanh to output an estimated vector Broadcast the estimated vector as a matrix of the same size as the original image, and perform a convolution operation on the feature map F or local features through the backscattering light decoder BE to obtain a preliminary backscattering feature map Fb. Use the upsampling operation to restore it to the H×W resolution and map the output to the corresponding estimated vector Set the backscattering light calculation formula for each channel:
[0058]
[0059] where m c is the scattering intensity coefficient related to the channel; represents the value of the backscattering light at the position (i, j) when the color channel is c;
[0060] The backscattering light calculation formula is:
[0061] B(x) = ConvBlock2(ConvBlock1(Concat(Fb, dnorm(i, j))));
[0062] Step 2-4-3, reconstruct the clear image: The water body restoration decoder WR uses more than two deconvolution modules DeConvBlock to upsample the feature map F to restore it to the original resolution, obtain the detail restoration feature Fj, and output the restored image Combine and to inversely deduce the per-pixel calculation formula for restoring the clear image from the underwater scattering model formula:
[0063]
[0064] where I c (i, j) represents the pixel value of the input degraded image in the color channel c and at the position coordinates (i, j), that is, the original underwater image to be restored; represents the pixel value of the restored clear image in the color channel c and at the position coordinates (i, j), that is, the clear image calculated through the inverse process of the underwater scattering model;
[0065] Step 2-4-4, synthesize the degraded image:
[0066]
[0067] where ⊙ represents element-wise multiplication;
[0068] Step 2-4-5, establish the joint loss function L:
[0069]
[0070] where ||I - Igt||1 is the reconstruction loss, representing the difference between the original clear image I in the input training set and the real clear image Igt generated by the water body restoration decoder WR; Jsyn represents the input turbid image, and Jnorm represents the reconstructed turbid image; λ1 and γ are weights used to adjust the contribution ratio of different loss terms in the joint optimization. represents the gradient of the transmittance map: By minimizing forcing the transmittance map to be smooth, which conforms to the law that the transmittance changes continuously with depth in the real underwater scene (avoiding local noise or checkerboard effects).
[0071] Step 2 - 4 - 6, the closed - loop in the training process: The training goal is to minimize the joint loss function L.
[0072] In step 2 - 5, the relationship reasoning attention module performs the following steps:
[0073] Step 2 - 5 - 1, constructing the co - occurrence relationship graph: Input all the images and annotations in the training set, and create a zero matrix C of size N×N ij , where the zero matrix C ij is the element in the i - th row and j - th column, representing the co - occurrence frequency of the corresponding categories in the i - th row and the j - th column. N is the total number of categories; for each image, if categories i and j exist simultaneously, then update C ij by adding 1: For the set of annotated categories Y of each image, traverse all category pairs (i, j) ∈ Y×Y, update the co - occurrence matrix, and calculate the conditional probability matrix P:
[0074]
[0075] where the element P ij in the i - th row and j - th column of the conditional probability matrix P represents the probability that the category in the j - th column appears given that the category in the i - th row appears; P ij ≠P j ;
[0076] Step 2 - 5 - 2, inferring the co - occurrence relationship through a two - layer graph convolutional network GCN: Use the pre - trained word embedding model GloVe to obtain the word embedding vectors E ∈ R N×d of the category labels, and map them to the detection head feature space through a linear layer:
[0077] X = ReLU(E·Wembed),
[0078] where d is the word embedding dimension, representing the length of the category label word vector generated by the GloVe pre-trained model; Wembed is the embedding mapping weight matrix, which is used to linearly transform the GloVe word embedding vector into the feature space (dimension 256) of the detection head; align the word embedding with the feature space of the detection head to facilitate subsequent graph convolution operations; Wembed ∈ R d×256 ; X is the mapped category feature matrix, which is the 256-dimensional feature representation of each category (a total of N categories), integrating semantic information (GloVe) and task-related features;
[0079] Take the conditional probability matrix P as the adjacency matrix Q and add a self-connection:
[0080]
[0081] where represents the updated adjacency matrix;
[0082] The first layer of the graph convolutional network GCN performs the following operations:
[0083]
[0084] where W (1) ∈ R 256×128 ; Z (1) represents the output of the first layer of graph convolution;
[0085] W (1) represents the weight matrix of the first layer of GCN, which is used to map the input features from 256 dimensions to 128 dimensions;
[0086] Z (2) represents the output of the second layer of graph convolution; W (2) represents the weight matrix of the second layer of GCN;
[0087] The second layer of the graph convolutional network GCN performs the following operations:
[0088]
[0089] where W (2) ∈ R 128×64 ;
[0090] Output feature Z:
[0091] Z = Z (2) ,
[0092] where Z ∈ R N×64 represents the high-order co-occurrence relationship encoding;
[0093] Step 2-5-3, fuse visual and co-occurrence features through the attention mechanism: The visual feature v of the multi-scale feature map of the detection head ∈ R B×H×W×C, perform global max pooling and global average pooling on each channel to obtain the channel description vector v max ∈R B ×C , v avg ∈R B×C , v max is the channel description vector obtained by taking the maximum value of each channel of the visual feature v through global max pooling, and v avg is the channel description vector obtained by taking the average value of each channel of the visual feature v through global average pooling; v is the multi-scale visual feature map output by the detection head;
[0094] Concatenate the visual features v max and v avg into the channel description vector V of the visual feature, and then concatenate it with the feature Z, and generate the attention weight matrix M through a linear layer:
[0095] M = Sigmoid(Linear(Concat(V, Z))),
[0096] where M ∈ R B×C , dynamically adjust the linear layer parameters according to the input feature dimension C, and Sigmoid is the activation function; Linear represents the linear transformation layer (fully connected layer), which is used to map the input features to the target dimension;
[0097] Broadcast the attention weights to the original feature map size and multiply them with the visual features channel by channel:
[0098] Ffused = v ⊙ Broadcast(M),
[0099] where Ffused ∈ R B×C×H×W ; Ffused is the fused feature map, which is the result of weighted fusion of the co-occurrence features of the visual feature V through the attention mechanism;
[0100] Broadcast is dimension broadcasting, which is used to expand the attention weight matrix M to the same spatial dimension as the visual feature V;
[0101] Step 2-5-4, dynamic guidance in the inference stage:
[0102] Take the preliminary prediction category set S = {s1, s2,..., sk} output by the detection head as the input; where sk represents the kth category label predicted by the detection head;
[0103] According to the conditional probability matrix P and the top m categories with the highest co-occurrence probability of the categories in the preliminary prediction category set S:
[0104] T = Top m (∑ sk∈S Ps,:);
[0105] where T is the set of target categories for dynamic guidance, and Topm represents the operation of taking the top m maximum values, that is, selecting the m categories with the highest sum of co-occurrence probabilities with the categories in S;
[0106] In Ps,:, the comma, is used to separate the row and column indices. For example, Pi,j represents the element in the i-th row and j-th column of the matrix; the colon : represents taking the entire row or column, and Ps,: represents the row corresponding to the category sk;
[0107] Attention enhancement: In the detection head, the confidence of the prediction boxes corresponding to the categories in T is weighted: Confidence new = Confidence orig ·(1 + α·∑P S,t / h),
[0108] where α is an intermediate parameter, usually set as α = 0.2; Confidence orig is the confidence of the prediction box in the original output of the detection head;
[0109] Confidence new is the confidence dynamically adjusted according to the co-occurrence probability;
[0110] h is the number of categories for dynamic guidance;
[0111] ∑P S,t is the sum of the co-occurrence probabilities of the target category t and all categories in the preliminary prediction category set S.
[0112] In step 2 - 6, the adaptive feature fusion module performs the following steps:
[0113] Step 2 - 6 - 1, extracting multi-scale features through the YOLOv7 backbone network:
[0114] Fbackbone = {F1, F2, F3},
[0115] where Fr ∈ R Hr×Wr×Cr, H1 = H / 8, W1 = W / 8, C1 = 256; H2 = H / 16, W2 = W / 16, C2 = 512; H3 = H / 32, W3 = W / 32, C3 = 1024; Fbackbone is a set of multi-scale feature maps extracted by the YOLOv7 backbone network; F1 is a shallow feature map (high resolution, low semantic information), F2 is a middle feature map, and F3 is a deep feature map (low resolution, high semantic information); Fr is the r-th layer feature map, which is a multi-scale feature map extracted by the YOLOv7 backbone network; Hr and Wr respectively represent the height and width of the feature map Fr; Cr is the number of channels of the feature map Fr (i.e., the number of convolutional kernels).
[0116] H1 is the height of the input image after being reduced by the backbone network, and W1 is the width of the input image after being reduced by the backbone network; C1 is the number of shallow feature channels, C2 is the number of middle feature channels, and C3 is the number of deep feature channels.
[0117] Upsample the deep features and concatenate them with the shallow features:
[0118] F fpn2 = Concat(Upsample(F3), F2),
[0119] F fpn1 = Concat(Upsample(F fpn2 ), F1);
[0120] where F fpn2 is the second layer fusion feature of the Feature Pyramid Network FPN, and Upsample represents upsampling;
[0121] F fpn1 is the first layer fusion feature of the Feature Pyramid Network FPN;
[0122] Step 2 - 6 - 2, Adaptive weight generation and feature alignment: Define a learnable weight parameter x ∈ R for each pair of fusion features, initialized to a uniform distribution: x ∼ U(-0.1, 0.1); U(-0.1, 0.1) represents a uniform distribution.
[0123] Dynamically update x through the semantic difference between the backbone and neck features:
[0124] Δx = Conv 1x1 (Concat(F backbone , F fpn ))
[0125] x new = x + η·Δx + μ·(x - x new )
[0126] where η represents the learning rate, μ is the momentum coefficient, typically set as η = 0.01 and μ = 0.9; Δx is the update amount of the weight parameter; Conv 1x1 is a 1x1 convolutional layer for compressing or expanding the feature channels; F backbone is the multi-scale features output by the backbone network, namely {F1, F2, F3}; F fpn is the multi-level features after fusion by the Feature Pyramid Network FPN; x new is the updated weight parameter;
[0127] Bilinear upsampling is performed on the low-resolution features (referring to the feature map output by the deep layer of the backbone network, such as F3) to make the size of the low-resolution features consistent with that of the high-resolution features, and the number of channels is adjusted through a 1x1 convolution to obtain the aligned feature map F aligned :
[0128] F aligned = Conv 1x1 (F backbone );
[0129] Step 2-6-4, Dynamic Feature Fusion and Post-processing: Batch normalization BN and SiLU activation are performed on the parameter x:
[0130] x norm = BN(x),
[0131] x act = SiLU(x norm );
[0132] where x norm represents the weight parameter after batch normalization (BatchNorm, BN) for eliminating the offset of the parameter distribution; x act represents the weight parameter after being processed by the SiLU activation function (Sigmoid Linear Unit);
[0133] Two complementary weight maps w1 and w2 are generated by weighting through the Sigmoid function:
[0134] w1 = Sigmoid(x act ),
[0135] w2 = 1 - w1,
[0136] Channel-wise weighting is performed on the aligned feature map:
[0137] F fused = w1 · F aligned + w2 · F fpn ;
[0138] where F fused represents the feature map after dynamic weighted fusion, Ffpn Represents the fused features output by the Feature Pyramid Network (FPN), such as F fpn1 or F fpn2 ;
[0139] Generate the fused feature pyramid:
[0140] F final ={F fused1 ,F fused2 ,F fused3};
[0141] Among them, F final represents the set of feature pyramids after multi-level fusion, including dynamic weighted fusion feature maps at three levels, corresponding to object detections at different scales; F fused1 is the highest-resolution fused feature (for detecting small objects); F fused2 is the medium-resolution fused feature (for detecting medium-sized objects); F fused3 is the lowest-resolution fused feature (for detecting large objects);
[0142] Perform LeakyReLU activation and batch normalization on the fused features:
[0143] F out =BN(LeakyReLU(F fused ))
[0144] Among them, F out represents the final output feature map.
[0145] The present invention also provides an electronic device, including a processor and a memory, where the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the above method.
[0146] The present invention also provides a storage medium storing a computer program or instruction, and when the computer program or instruction runs on a computer, it executes the steps of the above method.
[0147] The present invention first designs a restoration subnet module. Using the underwater scattering model, during the training phase, it guides the detector to learn to remove the adverse features generated by water environment interference. This module is only enabled during training and does not increase the time cost of the detection process. Introduce a relational reasoning attention module. With the help of the co-occurrence relationship graph, it supplements the visual features missing for fish targets in complex underwater environments due to factors such as light and turbidity, and guides the detector to focus on fish and related potential co-occurring objects in the same water area scene. Adopt an adaptive feature fusion module. Effectively fuse the key features of the backbone network and the neck network, provide adapted features for the restoration subnet module and the relational reasoning attention module, and improve the model's ability to extract and process the features of fish targets in complex underwater environments.
[0148] The method of the present invention can be applied to the following fields:
[0149] Yangtze River ecological protection and research: By means of the prior knowledge-guided network technology, in-depth analysis is carried out on the images and videos taken underwater in the Yangtze River. It can accurately identify the fish species existing in the Yangtze River Basin, such as crucian carp, grass carp, etc., master the changes in their population numbers and habitats, provide detailed data for the ecological protection and research of the Yangtze River, help formulate targeted protection strategies, and promote the restoration and sustainable development of the Yangtze River ecosystem.
[0150] Smart fishery management: In the fishery production in the Yangtze River, by using the equipment equipped with this technology, the distribution of fish resources can be monitored in real time, helping fishery practitioners reasonably plan the fishing areas and times, avoid overfishing, and at the same time optimize the aquaculture strategy according to the growth conditions of fish, improve the scientific nature and efficiency of fishery aquaculture, and realize the sustainable utilization of fishery resources in the Yangtze River.
[0151] Underwater facility safety guarantee: For facilities such as cables laid underwater and bridge pile foundations in the Yangtze River, the prior knowledge-guided network technology is used to monitor the activities of surrounding fish in real time. Once it is found that abnormal fish approach and may pose a threat to the facilities, an alarm will be issued immediately to ensure the safe and stable operation of the underwater facilities and maintain the normal operation of the Yangtze River shipping and related infrastructure.
[0152] The method of the present invention has crucial application value in the fields of ecological research, fishery resource management, and underwater facility safety maintenance in the Yangtze River Basin, especially significant for the ecological restoration of the Yangtze River and the sustainable development of fisheries. Relying on high-resolution underwater imaging equipment, this technology uses a comprehensive model that integrates prior knowledge of underwater scattering models, prior knowledge of co-occurrence relationship graphs, and deep learning algorithms to comprehensively extract features and classify the image and video data collected underwater in the Yangtze River. It can accurately identify a variety of fish species existing in the Yangtze River Basin, such as crucian carp, grass carp, perch, etc. Compared with traditional underwater target detection means, the prior knowledge-guided network technology can make full use of the prior knowledge such as the complex and changeable environmental characteristics underwater in the Yangtze River and the inherent characteristics of fish targets, and greatly improve the accuracy and stability of detection in complex scenarios such as turbid river water, large changes in light due to seasons and weather, and fish migration. This plays an important role in the long-term monitoring and protection of the Yangtze River Basin ecosystem, the scientific and reasonable development and utilization of fishery resources, and the safety protection of underwater facilities.
[0153] Beneficial effects: Compared with traditional underwater target detection methods, this solution can significantly improve the detection accuracy. In complex scenarios such as turbid river water, variable light, and fish migration, it can accurately identify various fish species in the Yangtze River Basin, such as crucian carp, grass carp, perch, etc., meeting the requirements for accurate fish monitoring in fields such as Yangtze River ecological research and fishery resource management. By integrating prior knowledge and deep learning algorithms, the model can better adapt to the changes in complex underwater environments. The lightweight design enables the model to still process information quickly and accurately while reducing computational resource consumption, effectively coping with sudden changes in water turbidity or complex backgrounds, and maintaining high detection performance under different underwater conditions. Brief Description of the Drawings
[0154] Figure 1 is the flowchart of the method of the present invention.
[0155] Figure 2 is the processing flowchart of the restoration subnet module.
[0156] Figure 3 is the processing flowchart of the relational reasoning attention module.
[0157] Figure 4 is the processing flowchart of the adaptive feature fusion module.
[0158] Figure 5 is an image of a bluestreak cleaner wrasse with low turbidity.
[0159] Figure 6 is an image of a regal angelfish with medium turbidity.
[0160] Figure 7 is an image of a regal angelfish with high turbidity. Detailed Embodiments
[0161] The following further specifically describes the present invention in conjunction with the drawings and detailed embodiments, and the above and / or other advantages of the present invention will become clearer.
[0162] As Figure 1 shown, an embodiment of the present invention provides an efficient method for detecting fish targets in complex underwater environments based on a prior knowledge-guided network, including the following steps:
[0163] Step 1, perform data preprocessing;
[0164] Image Screening: From publicly available datasets and images taken on-site, select clear underwater images that are not significantly degraded. These images should contain a variety of fish targets and may appropriately include some objects related to the fish's ecological environment, such as corals, aquatic plants, small underwater reefs, etc. Ensure that the types of target fish in the images are rich, covering individuals of different sizes, shapes, and colors, and construct the train dataset.
[0165] Underwater Condition Enhancement by Simulation: Specifically targeting the characteristics of the underwater environment, further simulate different underwater conditions for data augmentation. For example, according to the turbidity data measured in different regions of the Yangtze River, add corresponding degrees of simulated turbidity effects to the images; according to the light change rules in different seasons and time periods, adjust the light intensity and color tone of the images to simulate underwater scenes under different lighting conditions.
[0166] Underwater Effect Simulation: Based on the train dataset, simulate the impact of the underwater environment on the images. By adjusting the color channels of the images, simulate the common color deviations underwater, making the images overall greenish-blue; using the underwater image scattering model, process the clear images 10 times respectively to simulate different degrees of scattering and backscattering effects, changing the clarity and contrast of the images to reflect underwater environments with different turbidities, and form the improved dataset train1 after enhancement by the underwater scattering model.
[0167] Construct the test Dataset: Similar to the way of constructing the train dataset, select a part of the clear images from the publicly available underwater image dataset that have not been used for training to construct the test dataset. Ensure that these images have a similar distribution to the train in terms of target categories, scenes, etc., but are not exactly the same, to test the detection ability of the model for unseen clear underwater images, and generate test1 in a similar processing manner to train1.
[0168] Generate the Evaluation Dataset: Select fish datasets with publicly available underwater turbidity situations from other networks, and the evaluation dataset is used to evaluate the model.
[0169] Step 2: Perform fish target detection in complex underwater environments;
[0170] Step 2-1: Divide the dataset;
[0171] Data division: Determine that the ratio of train in the training set and test in the test set is 8:2. The evaluation set is a fish dataset that selects the publicly available underwater turbidity situation of other networks. The training set is used for the model to learn features and patterns, the test set is used to adjust the model hyperparameters and prevent overfitting, and the evaluation set is used to evaluate the final performance of the model. Note that the training set consists of two parts, one is the original dataset train, and the other is the improved dataset train1 enhanced by the underwater scattering model. It is composed of the mixture of these two datasets. The test set consists of two parts, one is the original dataset test, and the other is the improved dataset test1 enhanced by the underwater scattering model. It is composed of the mixture of these two datasets.
[0172] Step 2-2, Configure parameters;
[0173] Dataset configuration: Create a file named mydata.yaml, configure the storage paths of the training set, validation set, and test set in it, and at the same time clarify the class labels of various fish and related underwater objects to facilitate the model to identify different targets.
[0174] Model structure configuration: Generate a model.yaml file, determine the total number of classes of fish and related objects to be detected, and define the network structure of the model in detail, including parameter settings and connection methods of each layer, to build a model architecture suitable for underwater environment detection. The detection head (Head) is improved by adding a Graph Convolutional Network (GCN) module to perform graph convolution operations on the feature map to process the co-occurrence relationship graph and guide the model to focus on potential co-occurring objects. A decoder module is added to the detection head to perform upsampling and feature fusion on the feature map. The decoder module is used to convert the low-resolution feature map into a high-resolution feature map to improve the detection accuracy of the model.
[0175] Step 2-3, Build modules;
[0176] Design a recovery subnet module. As Figure 2 shown, the core is to introduce an improved underwater scattering model (improvement of the Jaffe-McGlamery model, ignoring the forward scattering component), and the formula is:
[0177] I(x) = J(x)t(x) + A·(1 - t(x)) + B(x),
[0178] Among them, I(x) represents the observed image, J(x) represents the clear image without scattering, t(x) represents the transmittance, A represents the background light, 1 - t(x) is the attenuated part, and B(x) represents the backscattered light. The part J(x)t(x) represents the light intensity directly reflected from the target object to the camera after attenuation by the water body. The part A·(1 - t(x)) represents the background light, that is, the part of the ambient light that enters the camera after scattering by the water body. 1 - t(x) represents the proportion of light that cannot reach the target object due to water body attenuation. The part B(x) represents the backscattered light, that is, the part of the ambient light that enters the camera after being scattered by the suspended particles in the water body.
[0179] The transmittance t(x) is calculated by the exponential attenuation formula, mainly used to simulate the absorption and scattering of light in the water body:
[0180] t(x) = e -β·q ,
[0181] where β is the attenuation coefficient of the water body, which can be processed by channels and can be extended to β c (such as βred R , βgreen G , βblue B ) reflects the absorption and scattering intensity of the water body to light; q is the propagation distance from the target to the camera, and the calculation formula is:
[0182]
[0183] where i and j are the row and column indices of the image, and center is the center coordinate of the image.
[0184] The background light A is a constant, representing the influence of the background light. In the underwater environment, due to the large number of suspended particles, the influence of the background light is more significant. It is usually set to a fixed value, such as 0.5. In this embodiment, it is adaptively assigned according to the water body type:
[0185]
[0186] The backscattered light B(x) is a distance-related term, representing the scattered light intensity caused by suspended particles, and the calculation formula is:
[0187]
[0188] where l represents the scattering intensity coefficient of suspended particles, which can be processed by channels and can be extended to m c (such as mred, mgreen, mblue) is positively correlated with the turbidity of the water quality.
[0189] Attenuation coefficient β calibration: Use an underwater spectrometer to measure the attenuation coefficients of different water areas, establish an NTU (turbidity) and β mapping table, and adapt to different water area conditions through β parameter control.
[0190] The main improvements to the underwater scattering model are as follows:
[0191] Ignore the forward scattering component: In underwater detection tasks, the influence of forward scattering on image quality is relatively small. This invention pays more attention to the image turbidity effect caused by backscattering. Ignoring the forward scattering component can simplify the model and at the same time concentrate computing resources on more critical backscattering and background light estimation.
[0192] Introduce multi-channel processing: The traditional Jaffe-McGlamery model usually processes the three RGB channels uniformly, but studies have found that there are significant differences in the scattering characteristics of light with different wavelengths in water. Therefore, this invention adds a mechanism for independent processing of each channel in the model, and sets different attenuation coefficients and background light values for each channel. This multi-channel processing method can more accurately simulate and restore the degradation process of underwater images.
[0193] Dynamic parameter adjustment: The restoration subnet module dynamically estimates parameters such as background light, transmittance, and backscattered light during the training process, rather than using fixed values. This dynamic adjustment mechanism enables the model to adapt to different underwater environments and image conditions, improving the quality and generality of the restored images.
[0194] When simulating the underwater scattering effect, different attenuation coefficients and background light values are set for the three RGB channels respectively. For the three RGB channels, different attenuation coefficients (such as red: 0.08, green: 0.03, blue: 0.01) and background light values (such as A = (0.7, 0.8, 0.9)) are set respectively. Apply wavelength-dependent attenuation coefficients to the three RGB channels (red > green > blue). Each color channel uses a different attenuation coefficient (red light attenuates the fastest, blue light the slowest) to accurately simulate the propagation characteristics of light with different wavelengths in water, and dynamic background light compensation: Dynamically adjust the weight of the background light A according to the depth to simulate the scattering characteristics of light with different wavelengths underwater.
[0195] Simulate the turbidity at different positions in the underwater scene by generating a depth map. According to the distance information of objects in the image, reasonably allocate the intensity of the scattering effect. Utilize the distance information in the depth map and combine it with the scattering model to calculate the transmittance and the proportion of background light influence for each pixel. Through the exponential attenuation formula, perform weighted fusion on the original value of the pixel and the background light value to simulate the attenuation and scattering process of light in water. Mainly set a higher depth value in the central area of the image to simulate the situation that the farther an object is from the camera in the underwater environment, the higher the turbidity. For each pixel, calculate its Euclidean distance from the center point, and based on the maximum depth value and the image size, calculate the final depth value through an exponential function.
[0196] During the training process, dynamically convert clear images into simulated turbid images, and generate enhanced images that conform to underwater optical characteristics in real time. By combining the scattering model generated from the depth map and applying the attenuation coefficients and background light values for each channel, highly realistic underwater turbid images can be generated. Provide clear learning objectives for the detector, that is, how to restore the true fish characteristics from the disturbed images. During the training process, the model continuously adjusts the parameters to make the prediction results closer to the results corrected based on the scattering model, thus gradually learning to remove the bad characteristics.
[0197] At the same time, a background ambient light decoder BAL (estimating the background light A), a water body restoration decoder WR (estimating the clear image J(x)), a water body transmittance decoder WTM (estimating the transmittance t(x)), and a backscattered light decoder BE (estimating the backscattered light B(x)) are designed to effectively utilize the underwater scattering model. The background ambient light decoder BAL and the water body transmittance decoder WTM mainly involve estimating the attributes related to the atmospheric effect and do not require high-resolution spatial information. The water body transmittance decoder WTM mainly uses two convblocks to estimate the transmittance. The water body restoration decoder WR mainly uses multiple DeConvBlocks to effectively restore the fine details in the degraded images. The background ambient light decoder BAL mainly uses an adaptive max pooling layer and a ConvBlock composed of a convolutional layer, a batch normalization layer, and a Tanh activation function to estimate the variable A of the global background light. The backscattered light decoder BE uses a convolutional block (ConvBlock) to estimate the backscattered light B(x). The convolutional block can extract the local features of the image and perform non-linear transformation through the activation function to estimate and compensate for the influence of backscattering on the image. Use the upsampling operation to align the feature maps of different resolutions to the same size for fusion. Through the adaptive feature fusion mechanism, effectively fuse the features extracted by the underwater scattering model guidance module and the fish co-occurrence relationship graph construction module with the features of the backbone network to enhance the network's ability to express fish characteristics.
[0198] At the same time, based on the underwater scattering model, the scattering and absorption characteristics of light in turbid water are modeled. Each decoder estimates the background light, global scattered light, and transmittance in the underwater environment respectively to generate a restored clear image. The estimation of these parameters not only helps to restore the clear image but also provides additional feature information for the detection head. For example, the transmittance can indicate the turbidity degree of different regions in the image, and the backscattered light can reflect the influence of suspended particles. By combining these parameters with the original image features, the restoration subnet module can enhance the feature expression ability and use it together with the original turbid image to guide the learning of the network, helping the network better understand the imaging principle in the underwater turbid environment, and thus extract more valuable fish features.
[0199] The restoration subnet module incorporates this prior knowledge of the underwater scattering model into the deep learning framework, enabling the model to not only learn the statistical laws in the data but also utilize the guidance of the physical model to more accurately understand the image degradation mechanism. This method combining the physical model and data-driven approach aims to make up for the deficiencies of pure data-driven models in complex environments. The restoration subnet module is combined with the detection head to form an end-to-end training framework. The output of the restoration subnet module (the restored clear image and the estimated parameters) serves as the input of the detection head, and at the same time, the feedback of the detection head also affects the parameter update of the restoration subnet module.
[0200] The restoration subnet module and the detection head share parameters and are co-optimized. The output of the restoration subnet module (the restored clear image and the estimated parameters) serves as the input of the detection head, and at the same time, the feedback of the detection head also affects the parameter update of the restoration subnet module. This two-way knowledge transfer mechanism enables the two modules to promote each other and jointly improve the overall performance of the model. The restoration subnet module helps the detection head better learn the target features by providing more accurate restoration results; while the performance feedback of the detection head guides the restoration subnet module to further optimize the restoration effect. Although the restoration subnet module plays an important role in the training stage, it will not be enabled in the actual detection stage. This is because once the model is trained, the detection head has learned how to accurately detect targets in the turbid environment and no longer requires the assistance of the restoration subnet module.
[0201] This method is similar to the process of using the atmospheric scattering model to guide the model to learn defogging in the foggy image defogging task. In a complex underwater environment, different water conditions (such as different turbidity degrees and water quality components) will lead to different degrees of light scattering and absorption. By incorporating these factors into the model training, the detector can be made more adaptable and accurately extract the features of fish targets. Since this module is only enabled during training, it will not affect the time cost of the actual detection process. After training, the detector has learned the ability to remove interference features and can quickly and accurately detect underwater fish targets in actual detection. The technical solution of the restoration subnet module is as follows:
[0202] (1) Data processing and depth map generation: Process the input underwater image \(W_{Raw}\in\mathbb{R}\) H×W×3 (where \(H\) is the height, \(W\) is the width, and 3 represents the three RGB channels), calculate the center coordinates of the image, and then for each pixel \((i, j)\), construct the Euclidean distance matrix \(D(i, j)\). Let \(D_{max}=\max\{D(i, j)\}\), and then define the normalized distance \(d(i, j)=\frac{D_{max}}{D(i, j)}\). Use exponential mapping to generate the depth weight matrix \(d_{norm}(i, j)=\exp(-\lambda\cdot d(i, j))\) to reflect the influence of distance on attenuation, where \(\lambda\) is the scale factor.
[0203] (2) Transmittance estimation, background light estimation, and scattered light estimation: Through the water body transmittance decoder \(W_{TM}\), the input feature map \(F\) passes through two cascaded convolutional blocks \(ConvBlock\) (including a convolutional layer and an activation function) to output a preliminary feature map \(F_t\). Use the upsampling operation to upsample \(F_t\) to the \(H\times W\) resolution and map it to a three-channel transmittance map
[0204] For each RGB channel, use the formula to calculate the transmittance \(t(x)\) respectively. The formula is:
[0205] t c (i, j)=\(\exp(-\beta\) c \(\cdot d_{norm}(i, j))\),
[0206] where \(c\in\{red, green, blue\}\), and the attenuation coefficient values for each channel are examples: \(\beta\) R = 0.08, \(\beta\) G = 0.03, \(\beta\) B = 0.01, to obtain the transmittance matrix \(t\) c \(\in\mathbb{R}\) H×W (calculated independently for each channel). So the transmittance calculation formula is:
[0207] t(x)=\(ConvBlock2(ConvBlock1(Concat(F, d_{norm}(i, j))))\),
[0208] Initially set fixed values for the three RGB channels: \(A = [A\) R , A\) G , A\) B = [0.7, 0.8, 0.9], and then dynamically estimate the background light \(A\) through the background ambient light decoder \(BAL\). Perform adaptive max pooling on the feature map of the input image to obtain a global feature vector. After convolution, batch normalization, and Tanh activation, output the estimated vector Then the estimated vector The broadcast is a matrix of the same size as the original image. Through the backscattering light decoder BE, a convolution operation is performed on the encoded feature map F or the local features to obtain the preliminary backscattering feature map Fb, which is restored to the H×W resolution using the upsampling operation and mapped for output. Set the backscattering light calculation formula for each channel: where m c is the scattering intensity coefficient related to the channel. So the backscattering light calculation formula is:
[0209] B(x) = ConvBlock2(ConvBlock1(Concat(Fb, dnorm(i, j))));
[0210] (3) Reconstruction of the clear image: The water body restoration decoder WR uses multiple deconvolution modules DeConvBlock to upsample the encoded feature map F to the original resolution to obtain the detail restoration feature Fj. Output the restored image. Combine the above estimated and for fusion, and inversely deduce the per-pixel calculation formula (calculated independently for each channel) for restoring the clear image from the underwater scattering model formula:
[0211]
[0212] where c ∈ {red, green, blue}, and this formula combines the observed value of each pixel with the background light, transmittance, and backscattering light for calculation to obtain the restored image for each channel.
[0213] (4) Synthesize the degraded image: The synthesized degraded image is generated based on the underwater scattering physical model through the real clear image and the transmittance background light backscattering light for calculation:
[0214]
[0215] where ⊙ represents element-wise multiplication, used to calculate the observed image formed after the real image passes through the scattering effect. The generated synthetic degraded image I c (i, j) is used as the training data and compared with the actual input image to verify the effect of the restoration network.
[0216] (5) Joint loss function: The joint loss L consists of three parts:
[0217]
[0218] Among them, ||I - Igt||1 is the reconstruction loss, which is the difference between the original input image I and the real clear image Igt generated by the water body restoration decoder WR. The loss part ||Jsyn - Jnorm||2 ensures the physical consistency between the generated image Jsyn and the real observed image Jnorm. That is, it forces the model to generate a synthetic degraded image as close as possible to the actual input image. This loss ensures the spatial smoothness of the transmittance t, preventing unreasonable discontinuities or mutations.
[0219] (6) Closed-loop during training: The training of the closed-loop synthetic degraded image generation is carried out using the joint loss function L. The training objective is to simultaneously minimize the reconstruction loss (to ensure visual quality), the physical consistency loss (to ensure that the degradation process conforms to physical laws), and the transmittance smoothness loss (to ensure the spatial consistency of the transmittance). Physical consistency closed-loop verification. The synthetic degraded image Jsyn is compared with the input image Jnorm to ensure that the restored clear image not only has good visual effects but also conforms to the physical degradation process. During the training process, the model will be gradually optimized according to the loss function to ensure that there is a minimum difference between the generated image I and Igt, and at the same time, the generated Jsyn is closer to the real input image.
[0220] Establish a relational reasoning attention module. By constructing an underwater ecological co-occurrence map, through analyzing a large number of annotated underwater image data, the frequency of different fish species appearing in the same scene is statistically analyzed, and a fish co-occurrence relationship graph is constructed. This graph takes fish species as nodes and their co-occurrence probabilities as edges, reflecting the correlation between fish in the underwater ecosystem. This module uses the fish co-occurrence relationship graph to reason about the co-occurrence relationship through a graph convolutional network to obtain the dependence relationship between fish. Then, these dependence relationships are applied to the attention mechanism to re-weight the visual features, making the detector YOLO pay more attention to potentially co-occurring fish and improving the detection accuracy.
[0221] Construct a co-occurrence relationship graph using the word embedding vectors of class labels and the co-occurrence matrix of object classes to guide the detector to pay more attention to potential co-occurring objects in the same scene. Co-occurrence matrix initialization: Create a matrix C of size N×N, where N is the total number of classes. Each element C ij represents the co-occurrence frequency of the class in the i-th row and the class in the j-th column. Co-occurrence frequency statistics: Traverse each image in the training set and count the co-occurrence frequency of each pair of class objects. If the class in the i-th row and the class in the j-th column appear in the same image, then C ij and C ji are both incremented by 1. Normalization processing: Normalize the co-occurrence matrix to obtain the conditional probability matrix P, where P ij represents the probability that the class in the j-th column appears given that the class in the i-th row appears.
[0222]
[0223] Use a two - layer graph convolutional network (GCN) to learn the dependencies in the co - occurrence relationship graph and use it to reason about the co - occurrence relationship graph. Extract the feature vectors of each category of objects to form the input feature matrix X. Use the graph convolutional layer to perform a convolution operation on the input feature matrix X and the co - occurrence matrix P to obtain a new feature matrix Z:
[0224] Z = GCN(X, P),
[0225] At the same time, in order to more deeply explore the high - order relationships between fish, use a multi - layer graph convolutional network GCN. Each layer of GCN further aggregates neighbor information on the basis of the previous layer and updates the node features.
[0226] Finally, take the visual feature V and the high - order feature Z of the graph convolutional network as inputs. Calculate the attention weight A through a linear transformation and a Sigmoid activation function. Multiply the attention weight by the visual feature to obtain the final fused feature F.
[0227] It is mainly the attention mechanism combined with prior knowledge. By constructing a co - occurrence relationship graph of fish and using a graph convolutional network to reason about the co - occurrence relationship, the prior knowledge is combined with the attention mechanism. Utilize the ability of the graph convolutional network to model relationships, and also strengthen the role of the correlation information between fish in feature learning through the attention mechanism. There is also an adaptive channel mapping, which adaptively selects different linear mapping operations according to the channel dimension of the input features and adjusts the output of the graph convolutional network to match the number of channels of the input features. This design enables the module to flexibly adapt to feature maps in different stages and with different numbers of channels, ensuring the coherence and effectiveness of the entire network structure.
[0228] In the training stage, the relationship reasoning attention module learns the dependencies between fish through the construction of the co - occurrence relationship graph and the reasoning of the graph convolutional network. This relationship information is encoded into the attention weights and applied to the visual features, enhancing the model's attention to potential co - occurring fish. In this way, the model can better learn the correlations between fish and improve its detection ability in complex scenarios.
[0229] In the actual detection stage, the relationship reasoning attention module applies the learned co - occurrence relationships of fish to the detection process. When the detector identifies a certain type of fish in the image, the module will guide the detector to pay attention to other fish that may co - occur according to the co - occurrence relationship graph. This guiding role can help the detector more comprehensively identify targets in a complex underwater environment and reduce the situations of missed detection and false detection.
[0230] This module can utilize the relationship between fish and surrounding environmental elements to assist in target detection. In an underwater scene, there are also close connections between different elements. For example, certain fish may gather in specific reef areas during the breeding season. By learning this co-occurrence relationship, when the detector identifies a reef area in an image, it can, based on the information in the co-occurrence map, pay more attention to whether relevant fish exist in that area. Even if some visual features of the fish are not obvious due to environmental interference, the detection accuracy can be improved through relationship reasoning. As Figure 3 shown, the relationship reasoning attention module performs the following steps:
[0231] (1) Co-occurrence relationship map construction (data-driven): Input all the images and their annotations (including class labels) in the training dataset, and create a zero matrix C of size N×N, where C ij represents the co-occurrence frequency of the class in the i-th row and the class in the j-th column, and N is the total number of classes. For each image, if classes i and j coexist, then update: For the set of annotated classes Y of each image, traverse all class pairs (i, j) ∈ Y×Y and update the co-occurrence matrix. Calculate the conditional probability matrix P:
[0232]
[0233] where P ij represents the probability that the class in the j-th column appears given that the class in the i-th row appears. Asymmetry retention: Retain P ij ≠P j (e.g., the probability of bass → grass carp is high, and vice versa is not necessarily true. Bass → grass carp means that when a bass is detected, there is a grass carp nearby).
[0234] (2) Graph Convolutional Network (GCN) reasoning of co-occurrence relationships: Use a pre-trained word embedding model (GloVe) to obtain the word embedding vectors E ∈ R N×d of class labels (such as "shark", "flying fish"). Map them to the detection head feature space through a linear layer, X = ReLU(E·Wembed), Wembed ∈ R d×256 ; Use the conditional probability matrix P as the adjacency matrix Q and add self-connections: The first layer of GCN operation:
[0235]
[0236] where W (1) ∈ R 256×128 ;
[0237] The second layer of GCN operation:
[0238]
[0239] where W(2) ∈R 128×64 ;
[0240] Output feature Z:
[0241] Z = Z (2) ,
[0242] where Z ∈ R N×64 represents the high - order co - occurrence relationship encoding, which is used to capture the high - order dependency relationships between categories.
[0243] (3) Attention mechanism to fuse visual and co - occurrence features: The multi - scale feature map visual feature v ∈ R of the detection head B ×H×W×C . Global max - pooling and global average - pooling are performed on each channel to obtain the channel description vector v max ∈R B×C , v avg ∈R B×C . The visual features v max , v avg are concatenated with the GCN output Z, and attention weights are generated through a linear layer:
[0244] M = Sigmoid(Linear(Concat(V, Z))),
[0245] where M ∈ R B×C , and the parameters of the linear layer are dynamically adjusted according to the input feature dimension C. The attention weights are broadcast to the original feature map size and multiplied with the visual features channel - by - channel:
[0246] Ffused = v ⊙ Broadcast(M) (Ffused ∈ R B×C×H×W ).
[0247] (4) Dynamic guidance in the inference stage: Input: The initial predicted category set S = {s1, s2,..., sk} output by the detection head. Co - occurrence probability retrieval, retrieving the top m categories with the highest co - occurrence probability with the categories in S according to the P matrix: T = Top m (∑ sk∈ S Ps,:). Attention enhancement: In the subsequent stages of the detection head, the confidence of the prediction boxes corresponding to T is weighted: Confidence new = Confidence orig ·(1 + α·∑P s,t / h);
[0248] where α = 0.2;
[0249] Build an adaptive feature fusion module. The YOLOv7 backbone network and neck network play different roles in object detection. As the backbone network, the Backbone is mainly responsible for extracting basic image features and uses operations such as convolutional layers (Conv), max pooling layers (MP), and skip connections (Concat) to extract multi-scale features f Backbone . The improved Feature Pyramid Network (FPN) as the neck network focuses on fusing and enhancing features at different scales and levels. Operations such as upsampling (Upsample) and skip connections (Concat) are used to fuse features f Neck .
[0250] In a complex underwater environment, different environmental factors will have different degrees of influence on the features output by these two networks. By designing an adaptive feature fusion strategy, introduce learnable parameter x to control the contribution ratio of different feature maps during the fusion process. These weight parameters can be set through random initialization or initialization methods based on prior knowledge. Design an adjustment mechanism for the weight parameters so that they can change dynamically according to the semantic information and importance of the feature maps, enabling the weight parameters to adaptively reflect the quality and relevance of the feature maps. If the resolutions or numbers of channels of different feature maps are inconsistent, operations such as upsampling, downsampling, or channel mapping need to be performed to align them in the spatial and channel dimensions. According to the generated weight parameters, perform weighted summation on different feature maps. The weight parameters determine the contribution ratio of each feature map in the fusion result, enabling the module to flexibly adjust the fusion strategy to adapt to different input features.
[0251] Perform batch normalization on x and use the SiLU activation function to perform a non-linear transformation on the generated parameter x. Generate two weight maps through the Sigmoid function: Sigmoid(x) and 1 - Sigmoid(x). Apply these two weight maps to the feature maps of f Backbone and f Neck respectively. Add the weighted feature maps element-wise to obtain the final fused feature map f Fused . Perform post-processing operations such as activation and normalization on the fused feature map:
[0252] f Fused = Sigmoid(x) × f Backbone +(1 - Sigmoid(x)) × f Neck ),
[0253] Mainly by introducing learnable weight parameters and a dynamic adjustment mechanism, the module can adaptively adjust the fusion strategy according to the semantic information and quality of the feature maps, avoiding the limitations of traditional fixed fusion methods. And through multi-scale feature fusion, it can process feature maps from different levels simultaneously, making full use of the detailed information of the low-level and the semantic information of the high-level to provide a richer feature representation for the detection task.
[0254] During the training phase, the adaptive feature fusion module dynamically adjusts the fusion strategy by learning the correlation and importance between different feature maps. The module adaptively assigns weights according to the semantic information of the feature maps and their contribution to the detection task, so that the fused feature maps can better capture the target features. By combining with the loss function, the module continuously optimizes the weight parameters to improve the learning efficiency and detection accuracy of the model.
[0255] During the detection phase, the adaptive feature fusion module performs real-time fusion on the input feature maps according to the learned weight parameters and fusion strategy. The fused feature maps contain multi-scale and multi-level feature information, which can provide a more comprehensive and accurate target representation for the detector.
[0256] This module can reasonably combine the features of the backbone network and the neck network according to the actual underwater environmental conditions. The introduction of the attention mechanism can further improve the effect of feature fusion. In the underwater environment, there is a large amount of background information and interference factors in the image. Through the attention mechanism, the model can automatically select the most valuable features for fish target detection, making the feature fusion more accurate and effective. As Figure 4 shown, the adaptive feature fusion module performs the following steps:
[0257] (1) Feature extraction of the backbone network: Extract multi-scale features through the YOLOv7 backbone network (including Conv, MP, and CSP modules):
[0258] Fbackbone = {F1, F2, F3},
[0259] where Fr ∈ R Hr×Wr×Cr , H1 = H / 8, W1 = W / 8, C1 = 256; H2 = H / 16, W2 = W / 16, C2 = 512;
[0260] H3 = H / 32, W3 = W / 32, C3 = 1024. Upsample the deep features and concatenate them with the shallow features:
[0261] F fpn2 = Concat(Upsample(F3), F2),
[0262] F fpn1= Concat(Upsample(F fpn2 ), F1);
[0263] (2) Adaptive weight generation and feature alignment: Define a learnable weight parameter x ∈ R for each pair of fused features (Backbone and FPN), initialized to a uniform distribution: x ∼ U(-0.1, 0.1), and dynamically update x through the semantic differences between the backbone and neck features:
[0264] Δx = Conv 1x1 (Concat(F backbone , F fpn ));
[0265] x new = x + η·Δx + μ·(x - x new ),
[0266] where the learning rate η = 0.01 and the momentum coefficient μ = 0.9; bilinearly upsample the low-resolution features to make their sizes consistent with the high-resolution features. Adjust the number of channels through 1x1 convolution:
[0267] F aligned = Conv 1x1 (F backbone );
[0268] (3) Dynamic feature fusion and post-processing: Apply batch normalization and SiLU activation to the parameter x:
[0269] x norm = BN(x), x act = SiLU(x norm );
[0270] Sigmoid weighting: Generate two complementary weight maps:
[0271] w1 = Sigmoid(x act ),
[0272] w2 = 1 - w1,
[0273] Perform channel-wise weighting on the aligned feature maps:
[0274] F fused = w1·F aligned + w2·F fpn ;
[0275] Repeat the above operations on multi-level features to generate a fused feature pyramid:
[0276] F final = {F fused1 , F fused2 , Ffused3};
[0277] Perform LeakyReLU activation and batch normalization on the fused features:
[0278] F out = BN(LeakyReLU(F fused ));
[0279] In this embodiment, several representative images are selected to show the detection effects under different turbidity levels.
[0280] The first example image 1 is as Figure 5 shown, which is an image of a low-turbidity blue parrotfish (Acanthurus_coeruleus) with a size of 512×512. During the training process, the model generates simulated turbid images through the restoration subnet module and enhances the feature representation through the relational reasoning attention module and the adaptive feature fusion module. The detection results show that even in relatively turbid underwater conditions, the model can accurately identify the species of blue parrotfish.
[0281] The second example image 2 is as Figure 6 shown, which is an image of a moderately turbid emperor angelfish (Pomacanthus_imperator) with a size of 512×512. The model estimates the transmittance, background light, and backscattered light through the restoration subnet module to generate a restored clear image. The detection results show that the model can still accurately detect emperor angelfish under moderate turbidity and reduce the cases of false detection and missed detection.
[0282] The third example image 3 is as Figure 7 shown, which is an image of a highly turbid emperor angelfish (Pomacanthus_imperator) with a size of 512×512. The model guides the detector to focus on potential fish through the relational reasoning attention module and the restoration subnet module. The detection results show that even under high turbidity, the model can accurately identify emperor angelfish and improve the detection accuracy through co-occurrence relationships.
[0283] Final result: Through detections under different turbidity levels, the model achieved a high mean average precision (mAP) on the test set, reaching over 80%. Compared with the prior art, the present invention significantly improves the accuracy and robustness of fish target detection in complex underwater environments and reduces the cases of false detection and missed detection.
[0284] The technical effects achieved by the embodiments of the present invention are as follows:
[0285] Improve detection accuracy: Through the restoration subnet module and the relational reasoning attention module, the model can more accurately identify fish targets and maintain high accuracy even in turbid environments.
[0286] Enhanced robustness: The adaptive feature fusion module enables the model to flexibly adapt to different underwater conditions and improves the robustness against environmental changes.
[0287] Reduced false positives and missed detections: Through the co-occurrence relationship graph and attention mechanism, the model can focus on potentially co-occurring fish, reducing false positives and missed detections.
[0288] Real-time performance: Although the model introduces multiple modules, through optimized design, the model can still maintain high real-time performance in actual detection and is suitable for practical applications.
[0289] The present invention provides an efficient method for detecting fish targets in complex underwater environments based on a prior knowledge-guided network. There are many methods and ways to specifically implement this technical solution. The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.
Claims
1. An efficient detection method for fish targets in complex underwater environments based on a prior knowledge-guided network, characterized in that It includes the following steps: Step 1: Obtain underwater images and perform data preprocessing. The preprocessing includes simulating underwater condition enhancement and simulating underwater effects; Step 2: Establish a restoration subnet module, a relationship reasoning attention module, and an adaptive feature fusion module to perform fish target detection in complex underwater environments; The restoration subnet module is used to guide the network to learn the underwater turbidity characteristics through the underwater scattering model, generate clear images through the water body restoration decoder WR, and reconstruct the underwater turbidity image in combination with the background ambient light decoder BAL, the water body transmittance decoder WTM, and the backscattered light decoder BE; The relationship reasoning attention module is used to construct a co-occurrence relationship graph, perform relationship reasoning using a graph convolutional network, dynamically adjust the attention weights, and guide the model to focus on potentially co-occurring targets; The adaptive feature fusion module is used to dynamically adjust the multi-scale feature fusion ratio of the outputs of the backbone network Backbone and the neck network Neck through learnable parameters, and optimize the feature expression by combining the Sigmoid function and the per-channel weighting mechanism; 2. The method according to claim 1, characterized in that, In Step 1, select underwater images from the public dataset to form the dataset train.
3. The method according to claim 2, wherein In Step 1, the simulating underwater condition enhancement includes: adding simulated turbidity effects to the images according to the turbidity data measured in different regions of the Yangtze River; adjusting the light intensity and color tone of the images according to the light change rules in different seasons and time periods to simulate underwater scenes under different lighting conditions.
4. The method according to claim 3, wherein In Step 1, the simulating underwater effects includes: based on the dataset train, simulating the influence of the underwater environment on the images, simulating the color deviation underwater by adjusting the color channels of the images to make the images overall greenish and bluish; using the improved underwater image scattering model to process the clear images X1 times respectively to simulate different degrees of scattering and backscattering effects, changing the clarity and contrast of the images to reflect underwater environments with different turbidities. The formula of the improved underwater scattering model is: I(x) = J(x)t(x) + A·(1 - t(x)) + B(x), where I(x) represents the observed image, J(x) represents the clear image without scattering, t(x) represents the transmittance, A represents the background light, B(x) represents the backscattered light; x represents the spatial position coordinates of a pixel in the image, which is a two-dimensional coordinate (i, j), and i and j correspond to the rows and columns of the image respectively.
5. The method according to claim 4, characterized in that, Step 2 includes: Step 2-1: Divide the dataset train into a training set, a validation set, and a test set; Step 2-2: Create a file named mydata.yaml, configure the storage paths of the training set, the validation set, and the test set in the file, and at the same time clarify the class labels of various fish and related underwater objects. Generate the model.yaml file, determine the total number of fish and related object categories to be detected, and improve the Yolov7 model. First, define the network structure of the Yolov7 model, including the parameter settings and connection methods of the backbone network, the head, and the Feature Pyramid Network (Neck); add a graph convolutional network module and a decoder module to the detection head of the Yolov7 model. The graph convolutional network module is used to perform graph convolutional operations on the feature map; the decoder module is used to perform upsampling and feature fusion on the feature map. Step 2-3, establish a restoration subnet module in the backbone network. The restoration subnet module adopts an improved underwater scattering model and calculates the transmittance t(x) through the following formula: t(x) = e -β·q , where e is the natural constant, β is the attenuation coefficient of the water body, q is the propagation distance from the target to the camera, and the calculation formula is: where Center[0] represents the abscissa of the center point, and Center[1] represents the ordinate of the center point. Assign the background light A adaptively according to the water body type: The calculation formula for the backscattered light B(x) is: where l represents the scattering intensity coefficient of suspended particles. Step 2-4, the restoration subnet module also includes a water body restoration decoder WR, a background ambient light decoder BAL, a water body transmittance decoder WTM, and a backscattered light decoder BE. The background ambient light decoder BAL uses an adaptive max pooling layer and a convolutional block ConvBlock composed of a convolutional layer, a batch normalization layer, and a Tanh activation function to estimate the background light A. The water body restoration decoder WR uses more than two transposed convolutional blocks DeConvBlocks to estimate the clear image J(x); the transposed convolutional blocks DeConvBlocks include transposed convolution, batch normalization, and ReLU activation functions. The water body transmittance decoder WTM estimates the transmittance t(x) through two convolutional blocks ConvBlock. The backscattered light decoder BE uses a convolutional block ConvBlock to estimate the backscattered light B(x). Use the upsampling operation to align feature maps of different resolutions to the same size for fusion. Through the adaptive feature fusion mechanism, fuse the features extracted by the underwater scattering model and the co-occurrence relationship graph with the features of the backbone network. Combine the restoration subnet module with the detection head to form an end-to-end training framework. The output of the restoration subnet module is used as the input of the detection head, and the performance feedback of the detection head guides the restoration subnet module to further optimize the restoration effect. Step 2-5, establish a relationship reasoning attention module. Step 2-6, establish an adaptive feature fusion module.
6. The method according to claim 5, characterized in that In step 2-4, the restoration subnet module performs the following steps: Step 2-4-1, data processing and depth map generation: Process the input underwater image \(W_{Raw}\in\mathbb{R}\) H×W×3 , where \(H\) is the height, \(W\) is the width, 3 represents the three RGB channels, and \(\mathbb{R}\) is the real number space; Calculate the image center coordinates. Then, for each pixel coordinate (i, j), construct an Euclidean distance matrix D(i, j). Let the maximum value of the Euclidean distance matrix D(i, j) be Dmax = max{D(i, j)}. Then define the normalized distance d(i, j) = Dmax / D(i, j), and generate the depth weight matrix dnorm(i, j) using the exponential mapping: dnorm(i,j) = exp(-λ·d(i,j)), where λ is the scale factor; exp represents the natural exponential function; Step 2-4-2, the restoration subnet module includes the following four decoders: The background ambient light decoder BAL is used to estimate the background light A; The water body restoration decoder WR is used to estimate the clear image J(x); The water body transmittance decoder WTM is used to estimate the transmittance t(x); The backscattered light decoder BE is used to estimate the backscattered light B(x); The collaborative working process of each decoder is as follows: estimating the transmittance, background light, and scattered light: through the water body transmittance decoder WTM, the input feature map F passes through two consecutive convolutional blocks ConvBlock, and the preliminary feature map Ft1 is output. The Ft1 is upsampled to the H×W resolution using the upsampling operation and mapped to a three-channel transmittance map For each RGB channel, the transmittance is calculated using the following formula respectively: t c (i, j) = exp(-β c ·dnorm(i, j)), where the color channel parameter c ∈ {red, green, blue}; t c (i, j) represents the transmittance value when the color channel is c; β c represents the attenuation coefficient of color channel c; Obtain the transmittance matrix t c ∈R H×W , and the transmittance calculation formula is: t(x) = ConvBlock2(ConvBlock1(Concat(F, dnorm(i,j)))), where Concat represents the concatenation operation; ConvBlock1 represents a convolutional block composed of a 3×3 convolution, a ReLU activation function, and batch normalization BatchNorm, and ConvBlock2 represents a convolutional block composed of a 1×1 convolution and a Sigmoid function; Set fixed values for the three RGB channels, and then dynamically estimate the background light A through the background ambient light decoder BAL. Perform adaptive max pooling on the feature map F of the input image to obtain a global feature vector, which is then convolved, batch-normalized, and activated by Tanh to output an estimated vector. Then broadcast the estimated vector into a matrix of the same size as the original image, and perform a convolution operation on the feature map F or local features through the backscattered light decoder BE to obtain a preliminary backscattered feature map Fb, which is restored to the H×W resolution by an upsampling operation and mapped to output the corresponding estimated vector. Set the backscattered light calculation formula for each channel: where m c is the scattering intensity coefficient related to the channel; represents the value of the backscattered light at the position (i, j) when the color channel is c; The formula for the scattered light is: B(x) = ConvBlock2(ConvBlock1(Concat(Fb, dnorm(i,j)))); Step 2-4-3, reconstruct the clear image: The water body restoration decoder WR uses more than two deconvolution modules DeConvBlock to upsample the feature map F to the original resolution, obtaining the detail restoration feature Fj, and outputting the restored image. Combine and to fuse, and inversely deduce the per-pixel calculation formula for restoring the clear image from the underwater scattering model formula: Among which I c (i,j) represents the pixel value of the input degraded image at the color channel c and the position coordinates (i,j); represents the pixel value of the restored clear image at the color channel c and the position coordinates (i,j); Step 2-4-4, synthesize the degraded image: where ⊙ represents element-wise multiplication; Step 2-4-5, establish the joint loss function L: where ||I - Igt||1 is the reconstruction loss, representing the difference between the original clear image I in the input training set and the real clear image Igt generated by the water body restoration decoder WR; Jsyn represents the input turbid image, and Jnorm represents the reconstructed turbid image; λ1 and γ are weights; represents the gradient of the transmission rate map; Step 2-4-6, the closed loop during training: The goal of training is to minimize the joint loss function L.
7. The method according to claim 6, characterized in that, In Step 2-5, the relation reasoning attention module performs the following steps: Step 2-5-1, construct a co-occurrence relationship graph: Input all the images and annotations in the training set, and create a zero matrix C of size N×N ij , where the zero matrix C ij is the element in the i-th row and j-th column, representing the co-occurrence frequency of the corresponding category in the i-th row and the category in the j-th column. N is the total number of categories; for each image, if categories i and j exist simultaneously, then increment C ij by 1 and update it: For the set of annotated categories Y of each image, traverse all category pairs (i, j) ∈ Y×Y, update the co-occurrence matrix, and calculate the conditional probability matrix P: where the element P in the i-th row and j-th column of the conditional probability matrix P ij represents the probability that the category in the j-th column appears given that the category in the i-th row appears; P ij ≠P j ; Step 2-5-2, infer the co-occurrence relationship through two-layer graph convolutional network GCN: Use the pre-trained word embedding model GloVe to obtain the word embedding vector E ∈ R N×d , and map it to the detection head feature space through a linear layer: X = ReLU(E·Wembed), where d is the dimension of word embeddings; Wembed is the weight matrix of the embedding mapping; Wembed ∈ R d×256 ; X is the matrix of mapped class features; Take the conditional probability matrix P as the adjacency matrix Q and add self-connections: Among them represents the updated adjacency matrix; The first layer of the graph convolutional network GCN performs the following operations: where W (1) ∈R 256×128 ; Z (1) represents the output of the first layer of graph convolution; W (1) Represents the weight matrix of the first-layer GCN, which is used to map the input features from 256 dimensions to 128 dimensions; Z (2) represents the output of the second-layer graph convolution; W (2) represents the weight matrix of the second-layer GCN; The second layer of the graph convolutional network GCN performs the following operations: where W (2) ∈R 128×64 ; Output the feature Z: Z = Z (2) , where Z ∈ R N×64 represents the encoding of the high-order co-occurrence relationship; Step 2-5-3, fusing visual and co-occurrence features through the attention mechanism: The multi-scale feature map visual feature v ∈ R of the detection head B×H×W×C , performing global max pooling and global average pooling on each channel to obtain the channel description vector v max ∈ R B×C , v avg ∈ R B×C , v max is the channel description vector obtained by taking the maximum value of each channel of the visual feature v through global max pooling, and v avg is the channel description vector obtained by taking the average value of each channel of the visual feature v through global average pooling; v is the multi-scale visual feature map output by the detection head; Concatenate the visual features v max and v avg into the channel description vector V of the visual features, then concatenate it with the feature Z, and generate the attention weight matrix M through a linear layer: M = Sigmoid(Linear(Concat(V, Z)), where M ∈ R B×C , Sigmoid is the activation function; Linear represents the linear transformation layer; Broadcast the attention weights to the original feature map size and multiply them element-wise with the visual features: Ffused = v⊙Broadcast(M), where Ffused ∈ ℝ B×C×H×W ; Ffused is the fused feature map; Broadcast is dimension broadcasting, which is used to expand the attention weight matrix M to the same spatial dimension as the visual feature V; Step 2-5-4, dynamic guidance in the inference stage: Take the preliminary prediction category set S = {s1, s2,..., sk} output by the detection head as the input; where sk represents the kth category label predicted by the detection head; According to the conditional probability matrix P and the top m categories with the highest co-occurrence probability in the preliminary prediction category set S: T = Top m (∑ sk∈S Ps,:); where T is the target category set for dynamic guidance, and Topm represents the operation of taking the top m maximum values; Ps,: represents the row corresponding to the category sk; Attention Enhancement: In the detection head, the confidence of the prediction box corresponding to class T is weighted: Confidence new = Confidence orig ·(1 + α·∑P S,t / h), where α is an intermediate parameter; Confidence orig is the confidence of the predicted bounding box in the original output of the detection head; Confidence new is the confidence dynamically adjusted according to the co-occurrence probability; h is the number of categories for dynamic guidance; ∑P S,t is the sum of the co-occurrence probabilities of the target category t and all categories in the preliminary prediction category set S.
8. The method according to claim 7, characterized in that, In Step 2-6, the adaptive feature fusion module performs the following steps: Step 2-6-1, extract multi-scale features through the YOLOv7 backbone network: Fbackbone = {F1, F2, F3}, where \(F_r\in\mathbb{R}\) Hr×Wr×Cr , \(H_1 = H / 8\), \(W_1 = W / 8\), \(C_1 = 256\); \(H_2 = H / 16\), \(W_2 = W / 16\), \(C_2 = 512\); \(H_3 = H / 32\), \(W_3 = W / 32\), \(C_3 = 1024\); \(F_{\text{backbone}}\) is a set of multi-scale feature maps extracted by the YOLOv7 backbone network; \(F_1\) is a shallow feature map, \(F_2\) is a middle-layer feature map, \(F_3\) is a deep feature map; \(F_r\) is the \(r\)-th layer feature map, which is a multi-scale feature map extracted by the YOLOv7 backbone network; \(H_r\) and \(W_r\) represent the height and width of the feature map \(F_r\) respectively; \(C_r\) is the number of channels of the feature map \(F_r\); H1 is the height of the input image after being downscaled by the backbone network, and W1 is the width of the input image after being downscaled by the backbone network; C1 is the number of channels of the shallow features, C2 is the number of channels of the middle features, and C3 is the number of channels of the deep features; Upsample the deep features and concatenate them with the shallow features: F fpn2 = Concat(Upsample(F3), F2), F fpn1 = Concat(Upsample(F fpn2 ), F1); Among them, F fpn2 is the second layer of fused features of the Feature Pyramid Network (FPN), and Upsample represents upsampling; F fpn1 is the first-layer fused feature of the Feature Pyramid Network (FPN); Step 2-6-2, Adaptive weight generation and feature alignment: Define a learnable weight parameter x ∈ R for each pair of fused features, initialized to a uniform distribution: x ∼ U(-0.1, 0.1); U(-0.1, 0.1) represents a uniform distribution; Dynamically update x through the semantic difference between the backbone and neck features: Δx = Conv 1x1 (Concat(F backbone , F fpn )) x new = x + η·Δx + μ·(x - x new ) where η represents the learning rate, μ is the momentum coefficient, and Δx is the update amount of the weight parameter; Conv 1x1 is a 1x1 convolutional layer; F backbone are the multi-scale features output by the backbone network, namely {F1, F2, F3}; F fpn are the multi-level features after the fusion of the Feature Pyramid Network FPN; x new are the updated weight parameters; Bilinearly upsample the low-resolution features to make the size of the low-resolution features consistent with that of the high-resolution features, and adjust the number of channels through 1x1 convolution to obtain the aligned feature map F aligned : F aligned = Conv 1x1 (F backbone ); Step 2-6-4, Dynamic feature fusion and post-processing: Perform batch normalization BN and SiLU activation on the parameter x: x norm = BN(x), x act = SiLU(x norm ); where x norm represents the weight parameter after batch normalization; x act represents the weight parameter after being processed by the SiLU activation function; Generate two complementary weight maps w1 and w2 by weighting with the Sigmoid function: w1 = Sigmoid(x act ) w2 = 1 - w1, Perform channel-wise weighting on the aligned feature maps: F fused = w1·F aligned + w2·F fpn ; Among them, F fused represents the feature map after dynamic weighted fusion, and F fpn represents the fused features output by the feature pyramid network; Generate a fused feature pyramid: F final ={F fused1 ,F fused2 ,F fused3}; Among which F final represents the set of feature pyramids after multi-level fusion; F fused1 is the highest-resolution fused feature; F fused2 is the medium-resolution fusion feature; F fused3 is the lowest-resolution fusion feature; Perform LeakyReLU activation and batch normalization on the fused features: F out = BN(LeakyReLU(F fused )) Among which F out represents the final feature map of the output.
9. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores program code. When the program code is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 8.
10. A storage medium, characterized in that, Stores a computer program or instructions. When the computer program or instructions run on a computer, they execute the steps of the method according to any one of claims 1 to 8.
Citation Information
Cited By
Fish abnormal behavior monitoring, description and diagnosis method
CN120783398A
Training method and system for multi-modal large model in automobile field
CN120875082A
Method and device for establishing defect detection model of lightweight power equipment, defect detection method and chip
CN121458684A
Underwater image enhancement method and system based on physical prior and target features
CN121810516A