Cyanobacteria Detection Method Based on Multi-Task Deep Learning and Adaptive Background Removal

By adopting multi-task deep learning and adaptive background removal methods in cyanobacteria detection, combined with graph convolutional network and self-attention mechanism, the problem of insufficient background interference processing capabilities in complex water environments is solved, and high-precision and robust cyanobacteria detection is achieved.

CN119904733BActive Publication Date: 2025-06-27ANHUI AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510387246.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-27
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing cyanobacterial detection methods lack background interference treatment capabilities and poor adaptability in complex water environments, and traditional methods are difficult to cope with dynamic changes in water environments, resulting in insufficient detection accuracy and robustness.

Method used

Using a method based on multi-task deep learning and adaptive background removal, the multi-modal adaptive graph attention network MAGAN and multi-task deep learning framework are constructed, combined with the graph convolution network and self-attention mechanism, the local features of the image are extracted and the degree of attention is adaptively adjusted to achieve joint optimization of background removal and cyanobacteria detection.

Benefits of technology

It improves the accuracy and robustness of cyanobacteria detection in complex water environments, enhances the adaptability and spatial perception of the model, effectively suppresses background noise, and reduces false detection and missed detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904733B_ABST
    Figure CN119904733B_ABST
Patent Text Reader

Abstract

The present invention is applicable to the field of cyanobacteria detection. Specifically, a cyanobacteria detection method based on multi-task deep learning and adaptive background removal is provided. In the detection method, a MAGAN model is constructed to extract local features through a graph convolutional network, and a self-attention mechanism is used to enhance global context awareness. Finally, the foreground and background are optimized through a background removal module to distinguish them, thereby realizing the adaptive background removal task. Then, a multi-task deep learning framework is adopted. By simultaneously learning the objectives of multiple tasks in the same model, the model can share knowledge from multiple tasks. The cyanobacteria detection task and the background removal task share some features and are collaboratively optimized to achieve background removal, providing support for cyanobacteria recognition. Cyanobacteria recognition provides reference information for background removal, enhancing the synergistic effect between tasks. The model can more accurately distinguish cyanobacteria from other water substances, reducing false detections and missed detections.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection by machine vision, and in particular to a cyanobacteria detection method based on multi-task deep learning and adaptive background removal. Background Art

[0002] The occurrence of cyanobacteria blooms poses a major threat to water quality, ecological environment, and even human health. Although traditional cyanobacteria detection methods, including microscopy, chemical analysis, and sensor technology, can provide relatively accurate detection results, they all have disadvantages such as long sampling cycles, large detection area limitations, and cumbersome operations. Therefore, the use of remote sensing images and image processing technology for automated cyanobacteria detection has become a hot topic of extensive research in recent years. The spectral reflectance characteristics of cyanobacteria in remote sensing data (such as chlorophyll a concentration, etc.) are unique, but the identification of cyanobacteria is challenging due to the presence of interferences such as algae, phytoplankton, and suspended matter in water bodies. In order to solve these problems, researchers have made significant progress by using intelligent algorithms such as deep learning, combined with image analysis, target detection, and background removal technologies.

[0003] Existing deep learning cyanobacteria identification methods can be roughly divided into two categories:

[0004] The first category: those based on image recognition mainly focus on the target detection, classification and segmentation of cyanobacteria through water images. The disadvantages of the methods based on image recognition are: on the one hand, different water environments and types of cyanobacteria may limit the generalization ability of the model; in new environments or when facing unprecedented types of cyanobacteria, the recognition accuracy of the model may decrease; on the other hand, due to the presence of various substances in water bodies, such as phytoplankton, suspended matter, etc., their spectral characteristics may be similar to those of cyanobacteria, which makes the existing single target detection methods prone to false detection or missed detection, especially when the spectral differences between cyanobacteria and other biological substances are not obvious.

[0005] The second category: remote sensing data-based methods use satellite, multispectral or hyperspectral remote sensing images to monitor and predict the dynamics of cyanobacteria on a large scale. The disadvantages of methods based on remote sensing data are: on the one hand, remote sensing data are usually high-dimensional and large-capacity, and processing and analyzing these data requires professional techniques and tools; in addition, the acquisition and processing of remote sensing data may be affected by factors such as satellite orbits and cloud cover, resulting in incompleteness and inaccuracy of the data; on the other hand, the spatial resolution of satellite images is limited, and the periodicity of image acquisition limits the timely monitoring of changes in cyanobacteria. Summary of the invention

[0006] The purpose of the embodiment of the present invention is to provide a cyanobacteria detection method based on multi-task deep learning and adaptive background removal, aiming to solve the technical problems that existing methods have insufficient ability to handle background interference and poor adaptability in complex water environments, and traditional methods such as detection based on optical sensors and simple image processing algorithms are difficult to cope with the dynamic changes of the water environment, resulting in insufficient detection accuracy and robustness.

[0007] To achieve the above object, the present invention provides the following technical solutions.

[0008] An embodiment of the present invention provides a cyanobacteria detection method based on multi-task deep learning and adaptive background removal. The cyanobacteria detection method includes the following steps:

[0009] Construct a detection model, where the detection model includes a multi-modal adaptive graph attention network MAGAN and a multi-task deep learning framework;

[0010] In the multi-modal adaptive graph attention network MAGAN, extract the local features of the image through a graph convolutional network GCN, combine the graph attention mechanism, and adaptively adjust and enhance the attention degree of each region in the image to obtain the features output after background removal;

[0011] Extract general features from the features output after background removal through a shared convolutional layer to generate a feature map; in the multi-task deep learning framework, use convolutional layers and fully connected layers to classify or segment cyanobacteria to generate a detection result, determine whether the image contains cyanobacteria, and locate the cyanobacteria region; and reconstruct the image background through convolutional layers and fully connected layers to generate an image after background removal; design a multi-task loss function, combine the cyanobacteria detection loss and the background removal task loss, and perform joint optimization through weighted combination;

[0012] Train and optimize the model parameters, use the water surface image as the input of the model, and output the detection result of cyanobacteria in the water surface image.

[0013] Further, the step of extracting the local features of the image through the graph convolutional network GCN includes:

[0014] For the input image, use the superpixel segmentation algorithm SLIC to divide the image into multiple regions with similar features, and obtain the vector F through multi-modal feature extraction i , F i = [C i , T i , S i , L i , where C i represents the color feature vector, T i represents the texture feature vector, S i represents the shape feature vector, and L i represents the illumination feature vector.

[0015] Further, combining with the graph attention mechanism, the steps of adaptively adjusting and enhancing the attention degree of each region in the image to obtain the features output after background removal include:

[0016] Using the graph attention mechanism to adaptively weight the image features, calculating the attention weights between nodes to dynamically adjust the attention degree of each region in the image;

[0017] The attention weights will perform weighted summation on the features of the nodes to obtain the updated node features;

[0018] The features of each node are updated by performing weighted summation with the features of neighboring nodes in the graph convolution operation;

[0019] By calculating the contrast loss between the foreground and the background to strengthen the differences between regions, the background is removed to obtain the output features.

[0020] Further, in the graph attention mechanism, the model calculates an attention weight α for each node ij , and the attention weight α ij represents the correlation between the node and other nodes;

[0021] The calculation formula of the attention mechanism is expressed as follows:

[0022] ;

[0023] where: Q i is the query vector of node i, K j is the key vector of node j, k represents the set of neighbor nodes of node i, and α ij is the attention weight of node i to node j; W Q represents the weight matrix of the query vector Q; T represents the transpose operation; represents the key vector of node k; by calculating to measure the similarity between node i and node j, introducing a non-linear transformation through the LeakyReLU activation function; the attention coefficients are normalized by softmax so that the sum of the attention weights of each node is 1;

[0024] Finally, the feature of node i will be obtained through the following weighted aggregation process: , where, h j is the feature of node j.

[0025] Further, in the graph convolutional network GCN, the feature vector of each superpixel region is updated in the graph convolutional layer by the pixel features of its neighborhood to obtain ; The update of the graph convolutional layer is as follows:

[0026] ;

[0027] Where, is the feature representation of node i at the l-th layer; N(i) represents the set of neighbor nodes of node i; A ij is the element of the adjacency matrix between nodes i and j, reflecting their similarity; W (l) is the learning weight matrix of the graph convolutional layer; σ is the ReLu activation function; c ij represents the normalization coefficient, where, , where, deg(i) represents the degree of node i; deg(j) represents the degree of node j.

[0028] Furthermore, the step of strengthening the difference between regions by calculating the contrast loss between the foreground and the background and removing the background to obtain the output features includes:

[0029] After being processed by the graph convolution and the self-attention mechanism, the obtained node features contain foreground information and background information ;

[0030] Introduce a background removal module to differentially process the extracted features. This module uses the contrast loss L contrastiv to distinguish the difference between cyanobacteria and the background, which is expressed as follows:

[0031] ;

[0032] Where, and respectively represent the features of the foreground and background regions, and N is the number of samples;

[0033] The output layer separates the foreground and the background, and uses a fully connected layer to map each pixel or region feature of the image to the class labels of the background and the foreground, which is expressed as: y i =softmax(W out h i ); where, h i represents the feature vector of node i, y i represents the classification result of node i, and the classification structure is foreground or background; W out is the weight matrix of the fully connected layer.

[0034] Furthermore, the step of extracting general features from the features output after background removal through a shared convolutional layer to generate a feature map includes:

[0035] Input the features output after background removal into a shared convolutional layer F shared, which is expressed as follows:

[0036] F shared (x) = ReLU(Conv(H MAGAN ));

[0037] Among them, F shared extracts general visual features, which are applicable to all subsequent tasks; the ReLU activation function is used to enhance non-linearity; each convolutional layer extracts different levels of features of the image; H MAGAN represents the node feature matrix obtained after processing the input image.

[0038] Furthermore, the step of using the convolutional layer and the fully connected layer to classify or segment cyanobacteria, generate a detection result, determine whether the image contains cyanobacteria, and locate the cyanobacteria area includes:

[0039] After all convolutional features are extracted, the feature map is flattened and passed to the fully connected layer for final cyanobacteria segmentation;

[0040] Using the fully connected layer F algae_fc to segment cyanobacteria, determine whether the image contains cyanobacteria, and accurately locate the cyanobacteria area, which is expressed as follows:

[0041] F algae (x) = ReLU(F algae_fc (flatten(F shared (x)))));

[0042] F algae_fc (x) = W i ·x + b i ;

[0043] Among them, i is the number of fully connected layers. When i takes 1, 2, and 3, the corresponding [W1, W2, W3] and [b1, b2, b3] are the weight matrix and bias term of the corresponding fully connected layer respectively;

[0044] The output layer is a Sigmoid layer, which is used for the segmentation task and outputs the probability of whether each pixel belongs to the cyanobacteria area.

[0045] Furthermore, the step of reconstructing the image background through the convolutional layer and the fully connected layer to generate an image after removing the background includes:

[0046] Transposed convolution is used to upsample the feature map and reconstruct the background; among them, the size of the input feature map F in is H′×W′×C in , and it is upsampled to H×W×C out through transposed convolution. The transposed convolution operation is expressed as follows: F out = DeConv(Fin , K deconv , stride, padding); where: F in is the input feature map with dimensions H′×W′×C in ; K deconv is the convolution kernel of the transposed convolution of size 3×3, used to learn the spatial relationship in the image; stride represents the stride; padding represents the padding; F out is the output feature map with dimensions H×W×C out .

[0047] Furthermore, the step of designing the multi-task loss function, combining the cyanobacteria detection loss and the background removal task loss, and jointly optimizing through weighted combination includes:

[0048] Design a multi-task loss function, combine the losses of the cyanobacteria detection and background removal tasks, and jointly optimize through weighted combination, where the cyanobacteria detection loss L algae uses the cross-entropy loss, expressed as:

[0049] ;

[0050] where, represents the actual algae value, is the predicted algae value;

[0051] The background removal loss L bg uses the mean squared error MSE loss, expressed as follows:

[0052] ;

[0053] where, y bgijk represents the corresponding true pixel value in the input image; y bg adjustedijk represents the pixel value of the image after background removal; H and W are the height and width of the output image respectively, and C is the number of channels of the output image; the total loss L total is expressed as follows: L total = λ algae L algae + λ bg L bg ;

[0054] where, λ algae and λ bg both represent weight coefficients, used to balance the losses of the cyanobacteria detection task and the background removal task.

[0055] Compared with the prior art, the beneficial effects of the cyanobacteria detection method based on multi-task deep learning and adaptive background removal of the present invention are:

[0056] First, the graph convolutional network of the MAGAN model architecture of the present invention can process the spatial information of images in a graph structure, convert the water body image into a graph structure, and extract local features by learning the relationships between nodes in the graph. This enables the model to have stronger spatial perception ability when dealing with the complex morphology of cyanobacteria and its interaction with the background;

[0057] Second, the self-attention mechanism of the MAGAN model architecture of the present invention further enhances the model's ability to accurately distinguish cyanobacteria from the background in complex environments by focusing on the global context information of each region in the image, can effectively suppress background noise, and improve the recognition accuracy of cyanobacteria;

[0058] Third, the present invention also adopts an adaptive background removal mechanism, which enables the model to dynamically adjust the background processing method according to the characteristics of different water environments, ensuring that the effect of background removal does not affect the accuracy of cyanobacteria detection. Through this adaptive adjustment, the background removal module can effectively avoid the problems of over-removal or insufficient removal of background interference in complex scenes.

[0059] Fourth, the present invention combines a multi-task deep learning framework, enabling the cyanobacteria detection task and the background removal task to share some features and be co-optimized, thereby further improving the overall performance. This multi-task co-optimization mechanism ensures the effective information transfer and complementarity between the cyanobacteria detection task and the background removal task, enhances the adaptability and robustness of the model, and enables the method to exhibit superior detection accuracy and reliability in complex and dynamic water environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention.

[0061] Figure 1 It is the architecture diagram of the multi-task deep learning framework in the cyanobacteria detection method based on multi-task deep learning and adaptive background removal of the present invention;

[0062] Figure 2 For Figure 1 It is the structural block diagram of the shared convolutional layer in

[0063] Figure 3 It is the architecture diagram of the multi-modal adaptive graph attention network MAGAN in the cyanobacteria detection method based on multi-task deep learning and adaptive background removal of the present invention;

[0064] Figure 4 It is the overall architecture diagram of the cyanobacteria detection method based on multi-task deep learning and adaptive background removal of the present invention;

[0065] Figure 5 This is the implementation flowchart of the cyanobacteria detection method based on multi-task deep learning and adaptive background removal of the present invention. Specific implementation manners

[0066] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0067] Aiming at the problems of insufficient background interference processing ability and poor adaptability of existing cyanobacteria detection methods in complex water environments, the present invention proposes a cyanobacteria detection method based on multi-task deep learning and adaptive background removal. In this method, first, a MAGAN model is constructed to extract local features through a graph convolutional network, and a self-attention mechanism is used to enhance global context awareness. Finally, the foreground and background distinction is optimized through a background removal module, thereby realizing the adaptive background removal task. Then, a multi-task deep learning framework is adopted. By learning the objectives of multiple tasks simultaneously in the same model, the model can share knowledge from multiple tasks. The cyanobacteria detection task and the background removal task share some features and are jointly optimized. The background removal provides support for cyanobacteria recognition, and the cyanobacteria recognition provides reference information for background removal, enhancing the synergy between tasks. The model can more accurately distinguish cyanobacteria from other water substances, reduce false detections and missed detections. At the same time, the adaptive background removal strategy enables the model to be adjusted according to the characteristics of different water area environments, effectively coping with detection tasks under various water conditions and providing a more flexible and reliable solution.

[0068] The following describes in detail the specific implementation of the cyanobacteria detection method based on multi-task deep learning and adaptive background removal of the present invention with reference to specific embodiments.

[0069] Please refer to Figures 1 - 5 , in an embodiment of the present invention, a cyanobacteria detection method based on multi-task deep learning and adaptive background removal is provided. The cyanobacteria detection method includes the following steps:

[0070] S101. Construct a detection model, where the detection model includes a multi-modal adaptive graph attention network MAGAN and a multi-task deep learning framework;

[0071] S102. In the multi-modal adaptive graph attention network MAGAN, extract local features of the image through a graph convolutional network GCN, and combine the graph attention mechanism to adaptively adjust and enhance the attention degree of each region in the image to obtain the features output after background removal;

[0072] S103. Extract general features from the features output after background removal through a shared convolutional layer to generate a feature map; in a multi-task deep learning framework, use convolutional layers and fully connected layers to classify or segment cyanobacteria, generate detection results, determine whether the image contains cyanobacteria, and locate the cyanobacteria region; and reconstruct the image background through convolutional layers and fully connected layers to generate an image after background removal; design a multi-task loss function, combine the cyanobacteria detection loss and the background removal task loss, and perform joint optimization through weighted combination.

[0073] S104. Train and optimize the model parameters, use the water surface image as the input of the model, and output the detection results of cyanobacteria in the water surface image.

[0074] Among them, before step S101, the present invention also includes the steps of data collection and preprocessing; specifically, collect the original water surface images, and perform data augmentation on the original image x by methods such as rotation, flipping, and scaling to increase the diversity of training data. Perform normalization processing on the image to enhance the contrast and clarity of the image, and obtain the enhanced data x'.

[0075] In one implementation, when the present invention collects the original water surface images, the water surface images are collected by drones, satellite remote sensing, or water body monitoring devices, and check whether the format of the input image is a supported format (such as JPEG, PNG, etc.), if not, perform format conversion. Adjust the image to a fixed size (such as 256x256 pixels) to ensure the consistency of subsequent processing.

[0076] Furthermore, during the preprocessing of the original image, for the original image x, the machine uses Gaussian filters and median filters to perform convolution operations on the image to remove Gaussian noise and salt-and-pepper noise in the image; perform histogram equalization on the image to enhance the contrast of the image; then normalize the image pixel values to the range of [0, 1] to obtain the enhanced data x', which is convenient for subsequent processing.

[0077] Furthermore, please refer to Figure 1 and Figure 2 , in step S102 of the present invention, the step of extracting local features of the image through the graph convolutional network GCN includes:

[0078] For the input image x', use the superpixel segmentation algorithm SLIC to divide the image into multiple regions with similar features, and obtain the vector F through multi-modal feature extraction i , F i = [C i , T i , S i , L i , C i represents the color feature vector, T i represents the texture feature vector, Si represents the shape feature vector, L i represents the illumination feature vector.

[0079] Furthermore, the step of adaptively adjusting and enhancing the attention degree of each region in the image by combining the graph attention mechanism to obtain the output features after background removal includes:

[0080] Using the graph attention mechanism to adaptively weight the image features and calculate the attention weight α between nodes ij to dynamically adjust the attention degree of each region in the image;

[0081] The attention weight α ij will perform weighted summation on the features of the nodes to obtain the updated node features ; the features of each node are updated by performing weighted summation with the features of neighboring nodes in the graph convolution operation to obtain ;

[0082] The background removal module reinforces the differences between regions by calculating the contrast loss L between the foreground and the background contrastive to remove the background and obtain the output features H MAGAN ;

[0083] Among them, for the input layer of the multi-modal adaptive graph attention network MAGAN, first, the image x′ is segmented using superpixel segmentation SLIC (Simple Linear Iterative Clustering), and superpixel regions with regular shapes and uniform sizes are generated by clustering pixel points. Each superpixel usually contains a relatively consistent region, which can reduce redundant information in subsequent calculations. For each superpixel region, the present invention extracts color, texture, shape, and illumination features and converts them into a vector F i , F i = [C i , T i , S i , L i , where C i represents the color feature vector, T i represents the texture feature vector, S i represents the shape feature vector, and L i represents the illumination feature vector.

[0084] In one implementation of the color feature C i , the color feature represents the color information of each superpixel region. For all pixels within each superpixel region, the average RGB color value of the pixels in the region is calculated as the color feature of the region. The color feature C i is represented as: , where I(p) is the color value of pixel p, and P i is the set of pixels within superpixel region i, and N i is the number of pixels within the superpixel region, and C i is the color feature vector of superpixel region i; the color feature C i is quantized and represented as C i = [R i , G i , B i .

[0085] In one implementation of the texture feature T i , the gray-level co-occurrence matrix (GLCM) extracts the texture features (contrast, energy, homogeneity, entropy) of the image by statistically analyzing the co-occurrence relationship of pixel pairs with different gray levels in the image, and converts the image from RGB (color image) to grayscale image; where: I gray = 0.2989×R + 0.5870×G + 0.1140×B, where R, G, and B are the pixel values of the red, green, and blue channels respectively; assuming the direction we choose is "horizontal", that is, the horizontal distance between two adjacent pixels is 1, a gray-level co-occurrence matrix P(i, j, d, θ) is generated, where i and j represent gray levels, d is the distance between pixel pairs (such as 1 pixel, 2 pixels), and θ is the direction of pixel pairs (for example, 0°, 45°, 90°, 135°);

[0086] Contrast measures the contrast of the image texture and calculates the degree of difference between matrix elements, expressed as:

[0087] ;

[0088] Energy is used to measure the uniformity of the texture in the image and reflects the smoothness of the image, expressed as: ;

[0089] Homogeneity is used to measure the smoothness of the image and reflects the consistency of the texture, expressed as: ;

[0090] Entropy is used to measure the complexity or information content of the image texture, expressed as:

[0091] ;

[0092] where the texture feature vector T i is expressed as:

[0093] T i = [Contrast i , Energyi , Homogeneity i , Entropy i 。

[0094] Furthermore, in one implementation of the shape feature, the contour extraction method (findContours in OpenCV) is used to connect the edges in the image into a closed contour; at this time, the contour information of each superpixel region is extracted; among them, the area of the contour can help distinguish regions of different sizes. As part of the shape feature, the area of the contour is expressed as follows: A i = Area(C i ); where C i represents the contour of the i-th superpixel region;

[0095] By calculating the bounding rectangle bounding box of the contour, the size and shape of the region can be obtained, which is expressed as follows: BoundingBox i = [x min , y min , x max , y max ; The perimeter can represent the edge length of the contour, which is expressed as follows: P i = Perimeter(C i );

[0096] Furthermore, the geometric invariant features of the contour are captured by Hu moments, such as rotation, scaling, and mirror invariance; Hu moments are usually calculated through contour points, which is expressed as follows: S i = [Hu1, Hu2,..., Hu7];

[0097] Through the above methods, we can extract multiple geometric features (such as area, perimeter, shape descriptors, etc.) from the contour of the superpixel region and form them into a feature vector S i , which is expressed as:

[0098] S i = [A i , P i , BoundingBox i , Hu i .

[0099] In one implementation of the illumination feature, by converting the RGB image into a grayscale image, the brightness information of each superpixel region is extracted;

[0100] The grayscale value L i is used as the illumination feature of the superpixel region, which is expressed as follows:

[0101] ; where gray(p) is the color value of pixel p, and P i is the set of pixels within the superpixel region, and N i is the number of pixels within the superpixel region.

[0102] The graph convolutional network of the MAGAN model architecture of the present invention can process the spatial information of an image in a graph structure, convert the water body image into a graph structure, and extract local features by learning the relationships between nodes in the graph. This enables the model to have stronger spatial perception ability when dealing with the complex morphology of cyanobacteria and its interaction with the background.

[0103] Furthermore, in the graph attention mechanism provided by the embodiments of the present invention, the model calculates an attention weight α for each node ij , and the attention weight α ij represents the correlation between a node and other nodes; the calculation formula of the attention mechanism is expressed as follows:

[0104] ;

[0105] where Q i is the query vector of node i, K j is the key vector of node j, k represents the set of neighbor nodes of node i, and α ij is the attention weight of node i to node j; W Q represents the weight matrix of the query vector Q; T represents the transpose operation; represents the key vector of node k, and the similarity between node i and node j is measured by calculating , and a non-linear transformation is introduced through the LeakyReLU activation function; the attention coefficient is normalized by softmax so that the sum of the attention weights of each node is 1;

[0106] Specifically, the query vector Q i is the transformation of the feature representation of node i, which is obtained by multiplying the feature of node i by a weight matrix W Q ; the query vector Q i is used to match with the features of other nodes in the self-attention mechanism to determine the degree of attention of node i to other nodes, where Q i =W Q h i , h i is the original feature of node i, and W Q is the learned weight matrix, usually a matrix of d×d′, where d is the dimension of the original feature and d′ is the dimension of the query vector;

[0107] The key vector K jis a transformation of the feature representation of node j, which is used together with the query vector to calculate the similarity between nodes. In the self-attention mechanism, node i will use the query vector Q i to match with the key vector K j of node j to calculate the attention degree, which is expressed as: K j =W k h j ;

[0108] These two matrices, the weight matrices W Q and W k are learned parameters used to map the features of nodes to the space of query vectors and key vectors. Through these weight matrices, the model can learn how to calculate attention between different nodes, so as to dynamically adjust the attention area according to the context. Among them, W Q ∈R d×d′ , W k ∈R d×d′ .

[0109] Furthermore, the calculation of queries and keys: For each node i, first calculate its query vector Q i and the key vector K j of each neighbor node j. Similarity calculation: Then, calculate the similarity i between the query vector Q j and the key vector K , which is usually completed through the dot product of vectors;

[0110] Finally, the feature of node i will be obtained through the following weighted aggregation process: ; where h j is the feature of node j. This aggregation process helps the model to focus more attention on important neighbor nodes, thus enhancing the focus on the cyanobacteria area.

[0111] Furthermore, in the graph convolutional network GCN, the feature vector of each superpixel region is updated in the graph convolutional layer through the pixel features of its neighborhood to obtain ; The update of the graph convolutional layer is as shown in the formula: ; where, is the feature representation of node i at the l-th layer; N(i) represents the set of neighbor nodes of node i; A ij is the element of the adjacency matrix between nodes i and j, reflecting their similarity; W (l) is the learned weight matrix of the graph convolutional layer; σ is the ReLu activation function; c ij represents the normalization coefficient, , deg(i) represents the degree of node i; deg(j) represents the degree of node j.

[0112] Further, in the embodiments of the present invention, the step of removing the background to obtain the output features by calculating the contrast loss between the foreground and the background to strengthen the difference between regions includes:

[0113] After being processed by graph convolution and self-attention mechanism, the obtained node features contain foreground information and background information ;

[0114] To perform adaptive background removal, the present invention introduces a background removal module to differentially process the extracted features. This module uses the contrast loss L contrastiv to distinguish the difference between cyanobacteria and the background, facilitating better distinction between foreground information and background information, as shown below:

[0115] ;

[0116] Among them, and respectively represent the features of the foreground and background regions, and N is the number of samples;

[0117] This loss function L contrastiv aims to enhance the model's ability to distinguish between the two by maximizing the difference between foreground and background features; represents the square of the Euclidean distance between the foreground and background feature vectors;

[0118] The output layer separates the foreground and the background, and uses a fully connected layer to map each pixel or region feature of the image to the class labels of the background and the foreground, expressed as: y i =softmax (W out h i ); where h i represents the feature vector of node i, y i represents the classification result of node i, and the classification structure is foreground or background; W out is the weight matrix of the fully connected layer.

[0119] Feature dimension of the feature vector h i : In image processing tasks, nodes usually correspond to pixels or superpixel regions in the image. The dimension of the feature vector h i is determined by the network architecture design and usually includes multiple feature channels such as color, texture, and shape; for example, the color feature may include the values of the red, green, and blue channels, the texture feature may include the texture information extracted by filters, and the shape feature may include the edge detection results, etc.;

[0120] Information fusion: In MAGAN, the node feature h iIt is obtained by fusing information from multiple modalities (such as color, texture, shape, etc.). This fusion enables the feature vector of each node to comprehensively represent its attributes and context information in the image.

[0121] The present invention adopts an adaptive background removal mechanism, which enables the model to dynamically adjust the background processing method according to the characteristics of different water environments, ensuring that the effect of background removal does not affect the accuracy of cyanobacteria detection. Through this adaptive adjustment, the background removal module can effectively avoid the problems of over-removal or insufficient removal of background interference in complex scenarios.

[0122] Further, please refer to Figure 1 and Figure 2 , in the embodiments of the present invention, the key to multi-task learning is to share features, allowing the model to jointly optimize in multiple related tasks;

[0123] Set the output feature of the MAGAN model to be H MAGAN , which is a matrix containing the features of image regions. Usually, the output features of MAGAN will be fed into a shared feature layer, and then the output of this layer will be passed to different task-specific output layers; specifically:

[0124] The features H MAGAN output after background removal are input into a shared convolutional layer F shared , expressed as: F shared (x) = ReLU(Conv(H MAGAN )), where H MAGAN represents the node feature matrix obtained after processing the input image, and the general visual features extracted by F shared are applicable to all subsequent tasks; the ReLU activation function is used to enhance non-linearity; each convolutional layer extracts different levels of features of the image, from low-level features (edges, textures) to high-level features (cyanobacteria shape and structure); a 3×3 convolutional kernel is used, and the standard convolutional kernel size can effectively capture local features in the image. As the convolutional layer deepens, the number of channels gradually increases from 64 to 512, enabling the extraction of more complex features; max pooling: a pooling operation is applied after each convolutional layer, using a 2×2 pooling window with a stride of 2 to further reduce the size of the feature map and reduce the computational amount.

[0125] The output feature H MAGAN is the node feature matrix obtained after processing the input image. These features represent the semantic information of each region in the image and can be used for subsequent tasks such as cyanobacteria detection and background removal. Specifically, H MAGANis an N×D matrix, where N is the number of regions in the image, i.e., nodes, and D is the feature dimension of each node. These features capture the spatial and semantic relationships between image regions, providing rich information for multi-task learning; when inputting H MAGAN into the subsequent multi-task deep learning framework, common features can be further extracted through shared convolutional layers and then separately passed to each task-specific branch (such as cyanobacteria detection and background removal) to achieve collaborative optimization.

[0126] Furthermore, the step of using convolutional layers and fully connected layers to classify or segment cyanobacteria, generate detection results, determine whether the image contains cyanobacteria, and locate the cyanobacteria region includes:

[0127] After all convolutional features are extracted, the feature map is flattened and passed to the fully connected layer for final cyanobacteria segmentation;

[0128] Using the fully connected layer F algae_fc to segment cyanobacteria, determine whether the image contains cyanobacteria, and accurately locate the cyanobacteria region, which is expressed as follows:

[0129] F algae (x) = ReLU(F algae_fc (flatten(F shared (x)))));

[0130] F algae_fc (x) = W i ·x + b i ;

[0131] where i is the number of fully connected layers, and when i takes 1, 2, and 3, the corresponding [W1, W2, W3] and [b1, b2, b3] are the weight matrices and bias terms of the corresponding fully connected layers respectively;

[0132] where F algae represents the overall function of the cyanobacteria detection branch, which is responsible for extracting features from the input image and classifying or segmenting cyanobacteria; specifically: F algae (x) represents the processing process of the cyanobacteria detection branch for the input image x, including feature extraction, flattening, fully connected layer processing, and activation function application, and finally generating the detection result of cyanobacteria; F algae_fc is the fully connected layer function in the cyanobacteria detection branch, which is responsible for converting the flattened feature map into the final output result;

[0133] Furthermore, the output layer is a Sigmoid layer for the segmentation task, which outputs the probability y of whether each pixel belongs to the cyanobacteria region algae predi .

[0134] The self-attention mechanism of the MAGAN model architecture of the present invention further enhances the model's ability to accurately distinguish cyanobacteria from the background in complex environments by focusing on the global context information of each region in the image, effectively suppressing background noise and improving the recognition accuracy of cyanobacteria.

[0135] Further, in the embodiment of the present invention, the step of reconstructing the image background through the convolutional layer and the fully connected layer to generate the image after removing the background includes:

[0136] The background removal task aims to remove the background from the image and retain the cyanobacteria region. For this purpose, the present invention uses transposed convolution to upsample the feature map and reconstruct the background; among them, the input feature map F in has a size of H′×W′×C in , and it is upsampled to H×W×C through transposed convolution out . The transposed convolution operation is expressed as follows: F out =DeConv(F in ,K deconv ,stride,padding);

[0137] Where: F in is the input feature map with a size of H′×W′×C in ; K deconv is the convolution kernel of the transposed convolution, with a size of 3×3, and this convolution kernel learns the spatial relationship in the image; stride is the stride, which determines the step size during each convolution, usually set to 2 or 4 to achieve upsampling; padding is the padding, which determines the processing method for the edges during the convolution operation; F out is the output feature map with a size of H×W×C out It has a higher spatial resolution; finally, is the image output of the background removal task.

[0138] The present invention combines a multi-task deep learning framework, enabling the cyanobacteria detection task and the background removal task to share some features and be jointly optimized, thereby further improving the overall performance. This multi-task collaborative optimization mechanism ensures effective information transfer and complementarity between the cyanobacteria detection task and the background removal task, enhancing the adaptability and robustness of the model, and making the method exhibit excellent detection accuracy and reliability in complex dynamic water environments.

[0139] Further, the step of designing the multi-task loss function, combining the cyanobacteria detection loss and the background removal task loss, and performing joint optimization through weighted combination includes:

[0140] Design the multi-task loss function L, combine the losses of the cyanobacteria detection and the background removal tasks, and perform joint optimization through weighted combination; the cyanobacteria detection loss Lalgae The cross - entropy loss is adopted and is expressed as follows:

[0141] ;

[0142] where represents the actual algae value, is the predicted algae value;

[0143] The background removal loss \(L\) bg uses the mean - squared error (MSE) loss and is expressed as follows:

[0144] ;

[0145] where \(y\) bgijk represents the corresponding true pixel value in the input image; \(y\) bg adjustedijk represents the pixel value of the image after background removal; \(H\) and \(W\) are respectively the height and width of the output image, and \(C\) is the number of channels of the output image; The total loss \(L\) total is expressed as follows: \(L\) total =\(\lambda\) algae \(L\) algae +\(\lambda\) bg \(L\) bg ; where \(\lambda\) algae and \(\lambda\) bg both represent weight coefficients used to balance the losses of the cyanobacteria detection task and the background removal task.

[0146] Furthermore, in step S104 provided by the present invention, during the process of training and optimizing the detection model, first, the parameters of the model are initialized. The learning rate \(\eta\) is taken as 0.001, the first - order momentum decay rate \(\beta_1\) is taken as 0.9, the second - order momentum decay rate \(\beta_2\) is taken as 0.999, and the numerical value \(\varepsilon\) is taken as \(10\) -8 ;

[0147] Bind the optimizer to the parameters \(\alpha\), \(W\) (l) , \(W\) out , \(W\) conv , \(W\) fc , \(b\) fc , \(\lambda\) algae , \(\lambda\) bg ; \(M = \{F\) shared , \(F\) algae , \(F\) bg , \(\alpha\}\); \(R=\{W\) (l) , \(W\) out , \(W\) conv , \(W\) fc , \(b\) fc , \(\lambda\) algae , \(\lambda\) bg \};

[0148] where \(W\) outDenote the weight matrix of the fully connected layer; W conv Denote the convolutional kernel weights of the shared convolutional layer; W fc Denote the weights of the fully connected layer of the cyanobacteria detection branch; b fc Denote the bias of the fully connected layer of the cyanobacteria detection branch; λ algae 、λ bg Denote the weight parameters in the loss function; F shared Denote the extracted general visual features; F algae Denote the fully connected layer for cyanobacteria classification or segmentation; F bg Denote the background removal result, which can be the predicted label of the background region or the image reconstruction output.

[0149] Train using the Adam optimizer:

[0150] , which means that in the Adam optimizer, using the total L total The process of updating the model parameter θ;

[0151] ; which means that in the Adam optimizer, using the contrastive loss L contrastive The process of updating the model parameter θ;

[0152] ; which means that in the Adam optimizer, using the cyanobacteria detection loss L algae The process of updating the model parameter θ;

[0153] Among them, θ represents the model parameter, and η represents the learning rate;

[0154] Training loop: Optimize the model parameters through the training data. Taking multi-task deep learning as an example:

[0155] ;

[0156] ;

[0157] ;

[0158] ;

[0159] ;

[0160] In Among them, Denote that: The general feature map F shared (x) output by the shared feature layer is input to the cyanobacteria segmentation branch F algae , and this branch decodes the features through the fully connected layer and outputs the probability that each pixel belongs to the cyanobacteria region; Representation: Prediction result of cyanobacteria segmentation branch is the actual algae value, is the predicted algae value, and the cross-entropy loss L is calculated algae for optimizing the classification ability of the model;

[0161] In , Representation: The output features of the cyanobacteria segmentation branch are passed to the background removal branch, upsampled by transposed convolution, and an image with the background removed is generated This process aims to reconstruct an image that only retains the cyanobacteria region; Representation: The mean squared error (MSE) loss L is calculated between the image after background removal and the true background removal result (i.e., the clean foreground image) bg for optimizing the accuracy of background removal;

[0162] In specifically, the following operations are performed:

[0163] Input: Current parameter θ, total loss L total and learning rate η;

[0164] Calculate the gradient: Calculate the gradient of the total loss L with respect to θ through backpropagation total ;

[0165] Update the parameter: Adjust θ according to the Adam formula to update it in the direction of decreasing loss;

[0166] Output: The updated parameter θ for the next round of training.

[0167] Specifically, in the embodiments of the present invention, the convolution kernel weights and biases of the shared convolutional layer are updated according to the rules of the gradient and the Adam optimizer; the weights and biases of the fully connected layer of the cyanobacteria detection branch are updated, and the Adam optimizer is applied to update the parameters; the weights and biases of the transposed convolution kernel of the background removal branch are updated by calculating the gradient and applying the Adam update step; the weight parameter λ in the loss function is adjusted according to the rules of the gradient and the Adam optimizer;

[0168] Input the preprocessed image x, perform adaptive background removal through the multi-modal adaptive graph attention network (MAGAN) to obtain the features after background removal, input the features after background removal into the multi-task deep learning framework, the shared convolutional layer extracts general features to generate a feature map, and the cyanobacteria detection branch uses convolutional layers and fully connected layers to classify or segment cyanobacteria to generate a detection result, represented as:

[0169] ;

[0170] The background removal branch reconstructs the image background through convolutional layers and fully connected layers, generating the image with the background removed, which is expressed as: ;

[0171] Combining the cyanobacteria detection loss and the background removal task loss, a multi-task loss function is obtained through weighted combination: .

[0172] Repeat the above processes of forward propagation, loss calculation, backpropagation, and parameter update until the performance of the model on the validation set reaches the expectation or meets certain convergence conditions.

[0173] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated examples described herein.

Claims

1. A cyanobacteria detection method based on multi-task deep learning and adaptive background removal, characterized in that: The cyanobacteria detection method comprises the following steps: Build a detection model, which includes a multimodal adaptive graph attention network MAGAN and a multi-task deep learning framework; In the multimodal adaptive graph attention network MAGAN, the local features of the image are extracted through the graph convolution network GCN, and combined with the graph attention mechanism, the attention level of each area in the image is adaptively adjusted and enhanced to obtain the output features after background removal; the graph attention mechanism is used to adaptively weight the image features, and the attention weights between nodes are calculated to dynamically adjust the attention level of each area in the image; the attention weights will weighted sum the features of the nodes to obtain the updated node features; the features of each node are updated by weighted summing with the features of the neighboring nodes in the graph convolution operation; the difference between the regions is strengthened by calculating the contrast loss between the foreground and the background, and the background is removed to obtain the output features; In the step of strengthening the difference between regions by calculating the contrast loss between the foreground and the background and removing the background to obtain the output features, after being processed by the graph convolution and self-attention mechanism, the obtained node features contain foreground information. and background information , a background removal module is introduced to differentiate the extracted features. The background removal module uses the contrast loss L contrastive To distinguish the difference between cyanobacteria and background; the output layer separates the foreground and background, and uses the fully connected layer to map each pixel or region feature of the image to the category label of the background and foreground; The features output after background removal are used to extract common features through a shared convolutional layer to generate a feature map; In the multi-task deep learning framework, convolutional layers and fully connected layers are used to classify or segment cyanobacteria, generate detection results, determine whether the image contains cyanobacteria, and locate the cyanobacteria area; and reconstruct the image background through convolutional layers and fully connected layers to generate an image with the background removed; Design a multi-task loss function, combine the cyanobacteria detection loss and the background removal task loss, and perform joint optimization through weighted combination; Train and optimize the model parameters, take the water surface image as the input of the model, and output the detection results of cyanobacteria in the water surface image.

2. The method for detecting cyanobacteria based on multi-task deep learning and adaptive background removal according to claim 1, characterized in that: The steps of extracting local features of an image through a graph convolutional network GCN include: For the input image, the superpixel partitioning algorithm SLIC is used to divide the image into multiple regions with similar features, and the vector F is obtained through multimodal feature extraction. i , F i =[C i ,T i ,S i ,L i ]; Among them, C i represents the color feature vector, T i represents the texture feature vector, S i represents the shape feature vector, L i Represents the illumination feature vector.

3. The method for detecting cyanobacteria based on multi-task deep learning and adaptive background removal according to claim 2, characterized in that: In the graph attention mechanism, the model calculates an attention weight α for each node ij , attention weight α ij Indicates the correlation between a node and other nodes; Among them, the calculation formula of the attention mechanism is expressed as follows: ; Where: Q i is the query vector of node i, K j is the key vector of node j, k represents the set of neighbor nodes of node i, α ij is the attention weight of node i to node j; W Q represents the weight matrix of the query vector Q; T represents the transpose operation; represents the key vector of node k; By calculation To measure the similarity between node i and node j, a nonlinear transformation is introduced through the LeakyReLU activation function; the attention coefficient is normalized by softmax so that the sum of the attention weights of each node is 1; the feature of node i The following weighted aggregation process will be used to obtain: , where h j is the feature of node j.

4. The method for detecting cyanobacteria based on multi-task deep learning and adaptive background removal according to claim 3, characterized in that: In the graph convolutional network GCN, the feature vector of each superpixel region is In the graph convolution layer, the pixel features of its neighborhood are updated to obtain ; The update of the graph convolutional layer is expressed as: ; in: is the feature representation of node i at layer l; N(i) represents the set of neighbor nodes of node i; A ij is the adjacency matrix element between node i and node j, reflecting their similarity; W (l) is the learning weight matrix of the graph convolutional layer; σ is the ReLu activation function; c ij represents the normalization coefficient, where: , where deg(i) represents the degree of node i; deg(j) represents the degree of node j.

5. The method for detecting cyanobacteria based on multi-task deep learning and adaptive background removal according to claim 4, characterized in that: Contrastive loss L contrastive It is expressed as follows: ; in, and Represent the features of the foreground and background areas respectively, and N is the number of samples; Use a fully connected layer to map each pixel or region feature of the image to the category label of background and foreground, expressed as: i =softmax(W out h i ), where h i represents the feature vector of node i; y i represents the classification result of node i, which is foreground or background; W out is the weight matrix of the fully connected layer.

6. The method for detecting cyanobacteria based on multi-task deep learning and adaptive background removal according to claim 5, characterized in that: The steps of extracting common features from the features output after background removal through a shared convolutional layer and generating a feature map include: The features output after background removal are input into a shared convolutional layer F shared , which is expressed as follows: F shared (x)=ReLU(Conv(H MAGAN )); Among them, F shared What is extracted are common visual features; H MAGAN Represents the node feature matrix obtained after processing the input image; Enhance nonlinearity through ReLU activation function; Each convolutional layer extracts features at different levels of the image.

7. The method for detecting cyanobacteria based on multi-task deep learning and adaptive background removal according to claim 6, characterized in that: The steps of using the convolutional layer and the fully connected layer to classify or segment the cyanobacteria, generate the detection results, determine whether the image contains cyanobacteria, and locate the cyanobacteria area include: After all convolutional features are extracted, the feature maps are flattened and passed to the fully connected layer for the final cyanobacteria segmentation. algae_fc Segment the cyanobacteria, determine whether the image contains cyanobacteria, and accurately locate the cyanobacteria area, as shown below: F algae (x)=ReLU(F algae_fc (flatten(F shared (x)))); F algae_fc (x)=W i ·x+b i ; Where i is the number of fully connected layers, and when i is 1, 2, or 3, the corresponding [W1, W2, W3] and [b1, b2, b3] are the weight matrix and bias term of the corresponding fully connected layer respectively; The output layer is a Sigmoid layer, which is used for segmentation tasks and outputs the probability of whether each pixel belongs to the cyanobacteria area.

8. The method for detecting cyanobacteria based on multi-task deep learning and adaptive background removal according to claim 7, characterized in that: The steps of reconstructing the image background through the convolution layer and the fully connected layer to generate the image after removing the background include: Transposed convolution is used to upsample the feature map and reconstruct the background; among them, the input feature map F in The dimensions are H′×W′×C in , upsample it to H×W×C by transposed convolution out , the transposed convolution operation is expressed as follows: F out =DeConv(F in ,K deconv ,stride,padding); Among them: F in The size is H'×W'×C in Input feature map of K deconv is a convolution kernel of a transposed convolution of size 3×3, which is used to learn the spatial relationship in the image; stride represents the stride; padding represents padding; F out The size is H×W×C out The output feature map of .

9. The method for detecting cyanobacteria based on multi-task deep learning and adaptive background removal according to claim 8, characterized in that: The steps of designing a multi-task loss function, combining the cyanobacteria detection loss and the background removal task loss, and performing joint optimization through weighted combination include: A multi-task loss function is designed, which combines the losses of cyanobacteria detection and background removal tasks and performs joint optimization through weighted combination. algae Using cross entropy loss, expressed as: ; in, Indicates the actual algae value, is the predicted algae value; Background removal loss L bg Using mean square error MSE loss, it is expressed as follows: ; Among them, y bgijk Represents the corresponding real pixel value in the input image; y bg adjustedijk Represents the pixel value of the image after background removal; H and W are the height and width of the output image, respectively, and C is the number of channels of the output image; the total loss L total It is expressed as follows: L total =λ algae L algae +λ bg L bg ; Among them, λ algae and λ bg Both represent weight coefficients, which are used to balance the losses of the cyanobacteria detection task and the background removal task.

Citation Information

Patent Citations

  • Unmanned aerial vehicle-mounted small target detection and positioning method and system under complex background

    CN112489032A

  • Method for intelligently predicting dynamic state of cyanobacteria population in reservoir

    CN117493942A