A method for identifying floating objects on the water surface based on semantic segmentation and image anomaly detection
The integration of semantic segmentation and image anomaly detection enhances water surface debris detection accuracy and adaptability, addressing the inefficiencies of traditional methods by automating the process and reducing human error.
Patent Information
- Application Number
- CN202310894099.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-07-20
AI Technical Summary
Traditional water surface floating object detection methods rely on manual monitoring, consume a lot of manpower and material resources, and are susceptible to environmental interference, making it difficult to achieve real-time and efficient water surface floating object recognition.
Using a method based on semantic segmentation and image abnormality detection, the semantic segmentation network and image abnormality detection network are constructed to identify floating objects on the water surface, including pre-processing, feature pyramid FPN network, surface segmentation network, abnormality detection network and image classification network, to realize automatic recognition of floating objects on the water surface.
It improves the comprehensiveness and accuracy of identification of floating objects on the water surface, reduces the cost of manpower and material resources, and can realize automated monitoring in complex environments, adapt to more scenarios, and reduce missed inspections.
Smart Images

Figure CN116824352B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water surface floating object image processing, and specifically, to a method for identifying water surface floating objects based on semantic segmentation and image anomaly detection. Background Art
[0002] The problem of water surface floating object pollution seriously affects human production and life. Real-time detection of water quality is an important link in water quality management and floating object pollution prevention. Due to the complexity of the water surface environment, water surface images have characteristics such as uneven illumination, being easily contaminated by noise, and being easily affected by weather, which makes the detection of water surface floating objects have certain particularity. Traditional object detection algorithms have limited feature extraction capabilities under the influence of external environments or noise interference.
[0003] The monitoring of the water surface environment mainly relies on arranging special personnel to manually view the monitoring screen. Although this method is simple, it requires a large amount of human and material resources; according to research, after staring at the monitoring screen for 22 minutes continuously, the human eye will ignore more than 95% of the moving information in the screen, and facing multiple monitoring screens for a long time is likely to cause fatigue of the monitoring personnel, making it difficult to respond promptly to some problems that occur on the water surface, especially under bad weather conditions. Summary of the Invention
[0004] The present invention is to solve the above-mentioned deficiencies existing in the prior art, and proposes a method for identifying water surface floating objects based on semantic segmentation and image anomaly detection, in order to solve the problem of identifying water surface floating objects in video monitoring scenarios. By constructing a semantic segmentation network and an image anomaly detection network, water surface floating objects in complex video monitoring scenarios can be obtained, thereby greatly improving the comprehensiveness and accuracy of floating object identification.
[0005] The present invention adopts the following technical solutions to achieve the above-mentioned invention purpose:
[0006] A method for identifying water surface floating objects based on semantic segmentation and image anomaly detection according to the present invention is characterized by including the following steps:
[0007] Step 1: Obtain a water surface floating object image data set and perform preprocessing of screening, standardization, and size adjustment in sequence to obtain a preprocessed water surface floating object image sample data set X = {X1,..., X n ,..., X N}, where X n represents the nth water surface floating object image, and N represents the total number of samples; the water surface area divided from X n is used as the label of X n , denoted as Y n ; and Y n∈C, where C represents the set of categories; annotate the water surface floating object image sample dataset X to obtain its corresponding mask mask set S = {S1,..., S n ,..., S N}; where S n represents the true category information of each pixel point in X n .
[0008] Step 2, build a water surface segmentation network based on the Feature Pyramid Network (FPN) to process the nth water surface floating object image X n and obtain the water surface image P n of the water surface floating object image X n .
[0009] Use Equation (4) to construct the loss function L α of the water surface segmentation network:
[0010] L α = L ce + λL ls (4)
[0011] In Equation (4), λ is the weight parameter; L ls represents the Lovasz-SoftMax loss; L ce represents the cross-entropy loss;
[0012] Step 3, build a water surface anomaly detection network based on a single classification algorithm to process the nth water surface image P n and obtain the floating object region image Z n of X n .
[0013] Use Equation (7) to construct the loss function L β of the water surface anomaly detection network:
[0014]
[0015] In Equation (7), μ is the weight parameter; L θ represents the encoding loss, represents the classification loss;
[0016] Step 4, build an image classification network based on ResNet, which sequentially includes a backbone network module and a classification module:
[0017] Step 4.1, the backbone network module adds an average pooling layer after the ResNet101 network;
[0018] The floating object region image Z n of X nInput the backbone network module, extract the feature maps output by the last three convolutional layers and perform feature fusion, and then obtain Z after dimensionality reduction processing by the average pooling layer n The feature vector T n ;
[0019] Step 4.2: The classification module is successively composed of a fully connected layer and a Softmax activation function, and input the feature vector T n into the classification module to obtain the nth water surface floating object image X n The predicted classification label for the mth category , m = 1, 2,..., M, where M is the total number of floating object categories, so as to obtain the category result of the floating object according to the corresponding relationship between the predicted classification label and the floating object category;
[0020] Step 4.3: Use Equation (8) to construct the loss function L γ :
[0021]
[0022] In Equation (8), J n,m represents the true category label of X n for the mth category;
[0023] Step 5: Based on the water surface floating object image sample dataset X = {X1,..., X n ,..., X N}, use the gradient descent method to train the water surface segmentation network, the water surface anomaly detection network and the image classification network, and calculate the loss functions L α , L β and L γ to update the network parameters until the loss function converges, so as to obtain the trained water surface segmentation model, water surface anomaly detection model and image classification model. Among them, the water surface segmentation model is used to segment the water surface floating object image collected in real time to obtain the water surface image; the anomaly detection model performs anomaly detection on the water surface image to obtain the floating object area, and then inputs it into the trained image classification model to obtain the predicted category of the floating object, and finally outputs the floating object area set and the category of the floating object.
[0024] The feature of the water surface floating object recognition method based on semantic segmentation and image anomaly detection according to the present invention is that the water surface segmentation network successively includes: a backbone network module, a pyramid pooling module PPM, a flow alignment module FAM and a refinement residual module RRB, and obtains the water surface image P according to the following steps n :
[0025] Step 2.1: The backbone network is based on the ResNet101 network and sequentially includes: a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer;
[0026] Input the nth water surface floating object image X n into the backbone network and sequentially process it through five convolutional layers. Each convolutional layer generates a corresponding feature map where, represents the water surface convolutional feature map generated by the ith convolutional layer;
[0027] Step 2.2: The pyramid pooling module PPM is composed of k pyramids with different scales. Perform k different scale pooling operations on the water surface convolutional feature map generated by the 5th convolutional layer to obtain k water surface feature maps with different sizes; then perform upsampling operations on the k water surface feature maps with different sizes respectively to restore them to the size of, and finally concatenate the k upsampled water surface feature maps in the channel dimension to obtain the nth low-resolution water surface pooled feature map
[0028] Step 2.3: The flow alignment module FAM upsamples the water surface pooled feature map to the same size as the water surface convolutional feature map generated by the (j + 1)th convolutional layer through bilinear interpolation, then concatenates the two in the channel dimension, and warps the concatenated feature map into the water surface convolutional feature map to generate the jth flow alignment feature map Thus, the flow alignment feature map set of the nth water surface pooled feature map is obtained. Concatenate the three flow alignment feature maps and use pointwise convolution to reduce the dimension to obtain a normalized flow alignment feature map
[0029] Step 2.4: Use Equation (1) to construct the cross-entropy loss L of the flow alignment module FAM ce :
[0030]
[0031] In Equation (1), represents the predicted value of the ith pixel in , p n,i represents the true value of the ith pixel in the mask mask S n , and N represents the number of pixels;
[0032] Step 2.5: The refined residual block RRB consists of e1 residual units, an integration adder, and a 1×1 dimensionality reduction convolutional layer. Each residual unit first uses a 1×1 ordinary convolutional layer to reduce the dimensionality of the input features. The reduced-dimensional features are divided into two branches. One branch is directly input into the adder, and the other branch is first passed through two cascaded 3×3 ordinary convolutional layers and then input into the adder to obtain integrated features.
[0033] Input the four convolutional layer feature maps of the water surface into the refined residual block RRB respectively. After passing through e1 residual units, integrated feature maps are obtained.
[0034] Input the four integrated feature maps into the integration adder, and then perform dimensionality reduction through the 1×1 dimensionality reduction convolutional layer to obtain a refined residual feature map. Then, perform pointwise convolution on the normalized flow alignment feature map and the refined residual feature map to obtain the water surface floating object image X n of the water surface image P n ;
[0035] Step 2.6: Use Equation (2) to construct the Lovasz-SoftMax loss L of the refined residual block RRB ls :
[0036]
[0037] In Equation (2), |C| represents the number of categories in the category set C, c represents any category in the category set C, is the Lovasz extension of ΔJ c , and ΔJ c is the Jaccard index of category c; m i (c) represents the vector of misclassifying the prediction of category c for the i-th pixel in S n , and there is:
[0038]
[0039] In Equation (3), represents the predicted value of the i-th pixel in the refined residual feature map .
[0040] The water surface anomaly detection network sequentially includes: an encoder module f θ , a texture enhancement module TEM, a pyramid texture feature extraction module PTFEM, and a classifier module and obtains the floating object region image Z of X n as followsn :
[0041] Step 3.1. The encoder module f θ is composed of r basic units, and each basic unit is successively composed of a two-dimensional convolution Conv2D and an activation function layer LeakyReLU;
[0042] Using h1 as the step size, the nth water surface image P n is divided into H blocks of size h2 to obtain a block set P n ={w n,1 ,..., w n,h ,..., w n,H}, where H is the total number of blocks of P n , and w n,h is the hth block in P n ;
[0043] Input w n,h into the encoder f θ for processing to extract the hth encoded feature A n,h ;
[0044] Step 3.2. Use Equation (5) to construct the loss L θ of the encoder f θ :
[0045]
[0046] In Equation (5), w n,h′ is the adjacent block of w n,h ;
[0047] The texture enhancement module TEM is composed of a global average pooling layer, a one-dimensional quantization coding QCO layer, and an MLP multi-layer perceptron;
[0048] After the hth encoded feature A n,h is processed by the global average pooling layer, the hth texture pooling feature g n,h is obtained. Then, calculate the cosine similarity between each pixel feature vector of the encoded feature A n,h and the texture pooling feature g n,h to generate a feature similarity matrix G n,h . The one-dimensional quantization coding QCO layer performs one-dimensional quantization coding on the feature similarity matrix G n,h to obtain the hth quantization coding matrix E n,h ; The hth quantization coding matrix E n,h generates the hth statistical feature D n,h after passing through the MLP multi-layer perceptron. The hth quantization coding matrix E n,h is then combined with the hth statistical feature Dn,h After multiplication, the h-th high-quality texture feature O is obtained n,h ;
[0049] Step 3.4: The pyramid texture feature extraction module PTFEM is composed of b1 parallel branches with different scales, and each parallel branch contains a texture feature extraction unit; the texture feature extraction unit is sequentially composed of b2 two-dimensional quantization coding 2d-QCO layers and b3 multi-layer perceptrons MLP;
[0050] The h-th high-quality texture feature O n,h and the h-th encoded feature A n,h After feature fusion, the h-th fused feature map K is generated n,h And input it into the pyramid texture feature extraction module PTFEM. First, divide the h-th fused feature map K n,h into b1 feature maps with different scales, and then input them into b1 parallel branches with different scales for processing respectively. Each branch extracts the corresponding texture representation feature map through its own texture feature extraction unit Among them, represents the texture representation feature map generated by the q-th branch; perform upsampling operations on the texture representation feature maps obtained from the b1 branches respectively to restore them to the size of the fused feature map K n,h Then, fuse the b1 upsampled feature maps to generate the h-th multi-scale texture feature F n,h , and finally fuse F n,h with A n,h to generate the h-th texture feature B n,h ;
[0051] Step 3.5: The classifier module is composed of x1 linear units, and each linear unit is sequentially composed of a linear layer and an activation function layer LeakyReLU;
[0052] The h-th texture feature B n,h is input into the classifier to obtain the anomaly score of the h-th block w n,h , so as to obtain the anomaly scores of all blocks; among them, the anomaly score of the pixel point i in the n-th water surface image P n is the mean value of the sum of the anomaly scores of all the cut blocks where the pixel point i is located;
[0053] Set the anomaly score threshold, and extract the anomaly regions in the anomaly scores of the pixel points in the n-th water surface image P n that are higher than the anomaly score threshold, so as to obtain the floating object region image Z n of X n ;
[0054] Step 3.6. Construct the classifier using Equation (6). Loss of
[0055]
[0056] In Equation (6), Cross_entropy represents cross entropy. represents a random small patch in the nth water surface image Pn. represents Any one of the small patches in the 8 surrounding directions centered on, where 1 ≤ p1, p2 ≤ H, and y is relative to azimuth label of, and y = 1, 2,..., 8.
[0057] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program for supporting the processor to execute the water surface floating object recognition method, and the processor is configured to execute the program stored in the memory.
[0058] A computer-readable storage medium according to the present invention, characterized in that a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the water surface floating object recognition method when run by a processor.
[0059] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0060] 1. The present invention uses a deep learning algorithm to improve the accuracy of floating object recognition, reduce human and material resources, and reduce labor costs. There are various types of floating object garbage, with diverse characteristics such as color and texture. Instead of identifying the categories of water surface floating objects one by one, it is better to directly regard all floating objects as an anomaly. No matter what kind of floating objects are on the water surface, as long as they are different from the normal water surface, they are anomalies. As long as there are floating objects on the water surface, it can be determined as an anomaly, thus realizing the automatic detection of water surface abnormal conditions. The semantic segmentation and image anomaly detection algorithms adopted by the present invention can automatically identify the floating object garbage on the water surface. Combined with a camera, the water surface conditions in multiple areas can be monitored simultaneously, thereby realizing more efficient management.
[0061] 2. Compared with the previous object detection methods, the proposed method for identifying water surface floating objects in the present invention has better accommodation for floating object categories, and is more comprehensive and accurate in identifying floating objects. If anomaly detection is directly used to identify water surface floating objects, the monitored area must be free of background interference, and the entire monitored area must be water surface, which is not practical. Therefore, the method of the present invention first uses a semantic segmentation network to perform preprocessing of water surface segmentation, separating the foreground part and the background of the water surface, so that the entire detection system can adapt to more scenarios, greatly improving the applicability and practicality of the method. By using the method of anomaly detection to identify floating object garbage, and then through image classification, various floating object garbage can be identified, effectively reducing the missed detection situation of floating object garbage identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 It is the overall flowchart for identifying water surface floating objects in the present invention;
[0063] Figure 2 It is the inference flowchart for identifying water surface floating objects in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] In this embodiment, a method for identifying water surface floating objects based on semantic segmentation and image anomaly detection mainly first uses a semantic segmentation network to segment the water surface and the background part in the water surface floating object image, obtaining the water surface image while eliminating the interference of the background part on the identification of floating objects. Then, an anomaly detection network is used to detect the floating object area in the water surface image, and an image classification network is used to identify the specific category of the floating object. As Figure 1 shown, the entire process can be specifically divided into the following steps:
[0065] Step 1. Obtain the water surface floating object image dataset and perform preprocessing of screening, standardization, and size adjustment in sequence to obtain the preprocessed water surface floating object image sample dataset X = {X1,..., X n ,..., X N}, where X n represents the nth water surface floating object image, and N represents the total number of samples; the water surface area divided from X n is used as the label of X n , denoted as Y n ; and Y n ∈C, where C represents the category set; the water surface floating object image sample dataset X is labeled to obtain its corresponding mask mask set S = {S1,..., S n ,..., S N}; among them, S n represents the true category information of each pixel point in X n ;
[0066] Step 2: Build a water surface segmentation network based on the Feature Pyramid Network (FPN), which successively includes: a backbone network module, a Pyramid Pooling Module (PPM), a Flow Alignment Module (FAM), and a Refinement Residual Block (RRB).
[0067] Step 2.1: The backbone network is based on the ResNet101 network and successively includes: a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer.
[0068] In this embodiment, the first convolutional layer is composed of a 7×7 convolutional layer, a Batch Normalization (BN) layer, a Rectified Linear Unit (ReLU) activation function layer, and a Max Pooling layer. The second convolutional layer is composed of three residual blocks. The third convolutional layer is composed of four residual blocks. The fourth convolutional layer is composed of twenty-three residual blocks. The fifth convolutional layer is composed of three residuals. Each residual block contains two 1×1 convolutional layers and one 3×3 convolutional layer.
[0069] Input the nth water surface floating object image X n into the backbone network and successively process it through five convolutional layers. Each convolutional layer generates a corresponding feature map. Among them, represents the water surface convolutional feature map generated by the i-th convolutional layer.
[0070] Step 2.2: The Pyramid Pooling Module (PPM) is composed of k different-scale pyramids. Perform k different-scale pooling operations on the water surface convolutional feature map generated by the 5th convolutional layer to obtain k different-sized water surface feature maps. Then, perform upsampling operations on the k different-sized water surface feature maps respectively to restore them to the size of. Finally, concatenate the k upsampled water surface feature maps in the channel dimension to obtain the nth low-resolution water surface pooled feature map In this embodiment, k is taken as 4, and the four different scales are 1×1, 2×2, 3×3, and 6×6 respectively.
[0071] Step 2.3: The Flow Alignment Module (FAM) upsamples the water surface pooled feature map to the same size as the water surface convolutional feature map generated by the (j + 1)-th convolutional layer through bilinear interpolation. Then, concatenate the two in the channel dimension and warp the concatenated feature map into the water surface convolutional feature map . In this embodiment, the warping operation is composed of two 3×3 convolutional layers to generate the j-th flow alignment feature map Thus, obtain the flow alignment feature map set of the nth water surface pooled feature map Concatenate three flow alignment feature maps Dimensionality reduction is performed by point - by - point convolution to obtain a normalized flow - aligned feature map
[0072] Step 2.4: Use Equation (1) to construct the cross - entropy loss \(L\) of the flow - alignment module FAM ce :
[0073]
[0074] In Equation (1), represents the predicted value of the \(i\) - th pixel in n,i and \(p\) represents the mask mask \(S\) n the true value of the \(i\) - th pixel in, and \(N\) represents the number of pixels;
[0075] The refined residual block RRB consists of \(e1\) residual units, an integration adder, and a \(1\times1\) dimensionality - reduction convolutional layer. Each residual unit first uses a \(1\times1\) ordinary convolutional layer to reduce the dimensionality of the input features. The reduced - dimensional features are divided into two branches. One branch is directly input into the adder, and the other branch first passes through two cascaded \(3\times3\) ordinary convolutional layers and then is input into the adder to obtain the integrated features. In this embodiment, \(e1\) is taken as 4;
[0076] The four - layer feature maps after convolution of the water surface are respectively input into the refined residual block RRB, and after passing through \(e1\) residual units, an integrated feature map is obtained
[0077] The four integrated feature maps are input into the integration adder, and after dimensionality reduction by a \(1\times1\) dimensionality - reduction convolutional layer, a refined residual feature map is obtained Then the normalized flow - aligned feature map and the refined residual feature map are subjected to point - by - point convolution to obtain a binary image of the water - surface segmentation result. Among them, the binary image means that the gray - scale values of the pixel points on the image are respectively set to 0 and 255, that is, pure black and pure white. Then, using image - processing technology, all pixel points with a gray - scale value of 255 in the obtained binary image are set to 1, and then multiplied pixel - by - pixel with the water - surface floating - object image \(X\) n to obtain the water - surface image \(P\) of the water - surface floating - object image \(X\) n ; n ;
[0078] Step 2.6: Use Equation (2) to construct the Lovasz - SoftMax loss \(L\) of RRB ls :
[0079]
[0080] In Equation (2), |C| represents the number of categories in the category set C, and c represents any category in the category set C. is the Lovasz extension of ΔJ c , and ΔJ c is the Jaccard index of category c; m i (c) represents the vector of mispredicting the category c of the i-th pixel in S n , and there is:
[0081]
[0082] In Equation (3), represents the predicted value of the i-th pixel in the refined residual feature map ;
[0083] Step 2.7. Use Equation (4) to construct the loss function L of the water surface segmentation network α :
[0084] L α = L ce + λL ls (4)
[0085] In Equation (4), λ is the weight parameter;
[0086] Step 3. Build a water surface anomaly detection network based on a single classification algorithm, which successively includes: an encoder module f θ , a texture enhancement module TEM, a pyramid texture feature extraction module PTFEM, and a classifier module
[0087] Step 3.1. The encoder module f θ is composed of r basic units, and each basic unit is successively composed of a two-dimensional convolution Conv2D and an activation function layer LeakyReLU; in this embodiment, r is taken as 8;
[0088] Taking h1 as the step size, divide the n-th water surface image P n into H blocks of size h2 to obtain a block set P n = {w n,1 ,..., w n,h ,..., w n,H}, where H is the total number of blocks of P n , and w n,h is the h-th block in P n ; in this embodiment, h1 is taken as 4, and the size of h2 is 32×32;
[0089] Input w n,h into the encoder f θ for processing to extract the h-th encoded feature An,h ;
[0090] Step 3.2. Construct the encoder f using Equation (5). θ The loss L of θ :
[0091]
[0092] In Equation (5), w n,h′ is the neighboring block of w n,h ; the neighboring blocks are the cut blocks in eight directions: up, down, left, right, upper left, upper right, lower left, and lower right.
[0093] Step 3.3. The texture enhancement module TEM consists of a global average pooling layer, a one-dimensional quantization coding QCO layer, and an MLP multi-layer perceptron.
[0094] After the h-th encoded feature A n,h is processed by the global average pooling layer, the h-th texture pooling feature g n,h is obtained. Then, calculate the cosine similarity between each pixel feature vector of the encoded feature A n,h and the texture pooling feature g n,h to generate the feature similarity matrix G n,h . The one-dimensional quantization coding QCO layer performs one-dimensional quantization coding on the feature similarity matrix G n,h to obtain the h-th quantization coding matrix E n,h ; the h-th quantization coding matrix E n,h generates the h-th statistical feature D n,h after passing through the MLP multi-layer perceptron. The h-th quantization coding matrix E n,h is then multiplied by the h-th statistical feature D n,h to obtain the h-th high-quality texture feature O n,h ;
[0095] Step 3.4. The pyramid texture feature extraction module PTFEM consists of b1 parallel branches with different scales. Each parallel branch contains a texture feature extraction unit; the texture feature extraction unit consists of b2 two-dimensional quantization coding 2d-QCO layers and b3 multi-layer perceptrons MLP in sequence; in this embodiment, b1 is taken as 4, b2 is taken as 1, b3 is taken as 1, and the four scales in PTFEM are 1, 2, 4, and 8 respectively, that is, the input feature is evenly divided into 1, 2, 4, and 8 small blocks.
[0096] After fusing the h-th high-quality texture feature O n,h and the h-th encoded feature A n,h , the h-th fused feature map K n,hAnd input it into the pyramid texture feature extraction module PTFEM, and first transform the hth fusion feature map K n,h After being divided into b1 feature maps of different scales, they are input into b1 parallel branches of different scales for processing. Each branch extracts the corresponding texture representation feature map through its own texture feature extraction unit. in, represents the texture representation feature map generated by the qth branch; the texture representation feature maps obtained by the b1th branch are upsampled respectively to restore to the fusion feature map K n,h The size of the h-th multi-scale texture feature F is generated by fusing the b1 upsampled feature maps. n,h Finally, F n,h With A n,h After fusion, the hth texture feature B is generated n,h ;
[0097] Step 3.5: Classifier module It is composed of x1 linear units, each of which is composed of a linear layer and an activation function layer LeakyReLU in sequence; in this embodiment, x1 is 4;
[0098] The h-th texture feature B n,h Input classifier and get the hth block w n,h The abnormality scores of all blocks are obtained by calculating the abnormality scores of all blocks; among them, the nth water surface image P n The anomaly score of pixel i in is the mean of the sum of anomaly scores of all slices where pixel i is located;
[0099] Set an abnormal score threshold. In this embodiment, the abnormal score threshold is set to 0.22. n The abnormal area with higher abnormal score than the abnormal score threshold is extracted from the abnormal score of the pixel point. According to the abnormal score threshold, the pixel points with higher abnormal score than the abnormal score threshold are set to 255, and the pixel points with lower abnormal score than the abnormal score threshold are set to 0 to obtain the outline of the abnormal area. Since the outline of the abnormal area is not necessarily a regular rectangular area, the image processing technology is used to intercept the minimum circumscribed rectangle of the abnormal area outline, so as to obtain X n The floating area image Z n ;
[0100] Step 3.6: Use formula (6) to build a classifier Loss
[0101]
[0102] In Equation (6), Cross_entropy represents cross entropy, represents a random small block in n the nth water surface image P represents any one of the small blocks in the 8 surrounding directions centered on , where 1 ≤ p1, p2 ≤ H, and y is the azimuth label relative to , and y = 1, 2,..., 8;
[0103] Step 3.7. Use Equation (7) to construct the loss function L of the water surface anomaly detection network β :
[0104]
[0105] In Equation (7), μ is the weight parameter;
[0106] Step 4. Build an image classification network based on ResNet, including a backbone network module and a classification module in sequence:
[0107] Step 4.1. The backbone network module adds an average pooling layer after the ResNet101 network;
[0108] X n 's floating object area image Z n is input into the backbone network module, and the feature maps output by the last three convolutional layers are extracted and feature fused. After dimensionality reduction by the average pooling layer, the feature vector T of n Z is obtained n ;
[0109] Step 4.2. The classification module consists of a fully connected layer and a Softmax activation function in sequence, and the feature vector T n is input into the classification module to obtain the predicted classification label for the mth category of the nth water surface floating object image X n m = 1, 2,..., M, where M is the total number of floating object categories. Thus, according to the correspondence between the predicted classification label and the floating object category, the floating object category result is obtained; in this embodiment, classification label 1 corresponds to the floating object category of bottle, 2 corresponds to duckweed,..., and M corresponds to others. m = 1, 2,..., M, where M is the total number of floating object categories. Thus, according to the correspondence between the predicted classification label and the floating object category, the floating object category result is obtained; in this embodiment, classification label 1 corresponds to the floating object category of bottle, 2 corresponds to duckweed,..., and M corresponds to others.
[0110] Step 4.3. Use Equation (8) to construct the loss function L of the image classification network γ :
[0111]
[0112] In Equation (8), J n,m represents X nThe true class label for the m-th category;
[0113] Step 5. Based on the water surface floating object image sample dataset X = {X1,..., X n ,..., X N}, use the gradient descent method to train the water surface segmentation network, the water surface anomaly detection network, and the image classification network, and calculate the loss functions L α , L β and L γ to update the network parameters until the loss functions converge, so as to obtain the trained water surface segmentation model, water surface anomaly detection model, and image classification model. Among them, the water surface segmentation model is used to segment the water surface floating object image collected in real time to obtain the water surface image; the anomaly detection model performs anomaly detection on the water surface image to obtain the floating object area, and then inputs it into the trained image classification model to obtain the predicted category of the floating object, and finally outputs the set of areas where the floating object is located and the category of the floating object. The specific recognition process is as Figure 2 shown, including:
[0114] The water surface floating object image passes through the water surface segmentation model to obtain a binary image of the water surface segmentation result, and then uses image processing technology to multiply the water surface floating object image by the binary image to obtain the water surface image. At this time, the background part of the water surface image is removed, and only the water surface and the floating objects in the water surface are retained;
[0115] Next, send the water surface image into the water surface anomaly detection model, set the anomaly score threshold, and determine whether there are floating objects in the water surface image according to the anomaly score threshold and the pixel anomaly score of the water surface image. If there are no floating objects on the water surface, the result that there are no floating objects on the water surface can be directly output; if there are floating objects on the water surface, the contour of the abnormal area is obtained from the relationship between the abnormal score of the water surface image pixel and the abnormal score threshold, and then the minimum circumscribed rectangle of the contour is obtained according to the abnormal area contour, and the minimum circumscribed rectangle of the contour is intercepted on the water surface image to obtain the water surface floating object area image;
[0116] Then send the water surface floating object area image into the image classification model to obtain the classification label value, and obtain the specific category of the floating object according to the correspondence between the classification label and the floating object category, and finally output the area where the floating object is located and the category of the floating object.
[0117] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0118] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium. When the computer program is run by the processor, it executes the steps of the above method.
Claims
1. A method for identifying floating objects on the water surface based on semantic segmentation and image anomaly detection, characterized in that, Including the following steps: Step 1: Obtain the water surface floating object image dataset and perform preprocessing of screening, standardization, and size adjustment in sequence to obtain the preprocessed water surface floating object image sample dataset X = {X1,..., X n ,..., X N}, where X n represents the nth water surface floating object image, and N represents the total number of samples; the water surface area divided from X n is used as the label of X n , denoted as Y n ; and Y n ∈ C, where C represents the category set; annotate the water surface floating object image sample dataset X to obtain its corresponding mask set S = {S1,..., S n ,..., S N}; among them, S n represents the true category information of each pixel point in X n . Step 2: Build a water surface segmentation network based on the Feature Pyramid Network (FPN) to process the nth water surface floating object image X n and obtain the water surface image P n of the water surface floating object image X n ; Construct the loss function \(L\) of the water surface segmentation network using Equation (4) α :[[-END]] L α = L ce + λL ls (4) In formula (4), λ is the weight parameter; L ls represents the Lovasz-SoftMax loss; L ce represents the cross-entropy loss; Step 3: Build a water surface anomaly detection network based on a single classification algorithm to process the nth water surface image P n and obtain the floating object area image Z n of X n ; Construct the loss function \(L\) of the water surface anomaly detection network using Equation (7). β : In formula (7), μ is the weight parameter; L θ represents the encoding loss, and represents the classification loss; Step 4, construct an image classification network based on ResNet, successively including a backbone network module and a classification module: Step 4.1, add an average pooling layer after the ResNet101 network in the backbone network module; X n Floating object area image Z n Input into the backbone network module, extract the feature maps output by the last three convolutional layers and perform feature fusion, and then obtain Z after dimensionality reduction by the average pooling layer n Feature vector T n ; Step 4.2: The classification module is successively composed of a fully connected layer and a Softmax activation function, and the feature vector T n is input into the classification module to obtain the nth water surface floating object image X n for the predicted classification label of the mth category where M is the total number of floating object categories. Thus, according to the correspondence between the predicted classification label and the floating object category, the category result of the floating object can be obtained; Step 4.3: Construct the loss function \(L\) of the image classification network using Equation (8) γ :[[]]END]] In formula (8), J n,m represents X n as the true class label of the m-th category; Step 5. Based on the water surface floating object image sample dataset X = {X1,..., X n ,..., X N}, use the gradient descent method to train the water surface segmentation network, the water surface anomaly detection network, and the image classification network, and calculate the loss functions L α , L β , and L γ to update the network parameters until the loss functions converge, so as to obtain the trained water surface segmentation model, water surface anomaly detection model, and image classification model. Among them, the water surface segmentation model is used to segment the water surface floating object image collected in real time to obtain the water surface image; the anomaly detection model performs anomaly detection on the water surface image to obtain the floating object area, and then inputs it into the trained image classification model to obtain the predicted category of the floating object, and finally outputs the set of areas where the floating object is located and the category of the floating object.
2. The method for identifying floating objects on the water surface based on semantic segmentation and image anomaly detection according to claim 1, wherein The water surface segmentation network sequentially includes: a backbone network module, a pyramid pooling module PPM, a flow alignment module FAM, and a refinement residual module RRB, and obtains a water surface image P according to the following steps n : Step 2.1, the backbone network is based on the ResNet101 network and successively includes: a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer; Input the nth floating object image X on the water surface n into the backbone network, and successively process it through five convolutional layers, and each convolutional layer generates a corresponding feature map respectively wherein represents the convolutional feature map of the water surface generated by the ith convolutional layer; Step 2.2: The pyramid pooling module PPM is composed of pyramids of k different scales. For the water surface convolution feature map generated by the 5th convolutional layer perform pooling operations of k different scales to obtain water surface feature maps of k different sizes; then perform upsampling operations on the k water surface feature maps of different sizes respectively to restore them to the size of, and finally concatenate the k upsampled water surface feature maps in the channel dimension to obtain the nth low-resolution water surface pooling feature map Step 2.
3. The flow alignment module FAM upsamples the water surface pooling feature map to the same size as the water surface convolution feature map generated by the (j + 1)-th convolutional layer using bilinear interpolation. Then, after concatenating the two in the channel dimension, the concatenated feature map is warped into the water surface convolution feature map to generate the j-th flow alignment feature map Thus, the flow alignment feature map set of the n-th water surface pooling feature map is obtained. Three flow alignment feature maps are dimension-reduced by pointwise convolution to obtain a normalized flow alignment feature map Step 2.
4. Construct the cross-entropy loss \(L\) of the flow alignment module FAM using Equation (1). ce : In formula (1), represents the predicted value of the i-th pixel in n,i the mask S n the true value of the i-th pixel in, and N represents the number of pixels; Step 2.5, the refined residual block RRB is composed of e1 residual units, an integration adder, and a 1×1 dimensionality reduction convolutional layer. Each residual unit first uses a 1×1 ordinary convolutional layer to reduce the dimensionality of the input features. The reduced-dimensional features are divided into two branches. One branch is directly input into the adder, and the other branch is first passed through two cascaded 3×3 ordinary convolutional layers and then input into the adder to obtain integrated features; The four-layer feature maps of the water surface convolution are respectively input into the refinement residual block RRB, and after passing through e1 residual units, an integrated feature map is obtained Integrate four feature maps Input them into the integration adder, and after dimension reduction by the 1×1 dimensionality reduction convolutional layer, obtain a refined residual feature map Then, perform pointwise convolution on the normalized flow alignment feature map and the refined residual feature map to obtain the water surface floating object image X n of the water surface image P n ; Step 2.
6. Construct the Lovasz-SoftMax loss \(L\) of the refined residual block RRB using Equation (2). ls : In formula (2), |C| represents the number of categories in the category set C, and c represents any category in the category set C. is the Lovasz extension of ΔJ c , where ΔJ c is the Jaccard index of category c; m i (c) represents the vector of mispredictions of the category c for the i-th pixel in S n , and there is: In formula (3), represents the predicted value of the \(i\)-th pixel in the refined residual feature map.
3. The method for identifying floating objects on the water surface based on semantic segmentation and image anomaly detection according to claim 2, wherein: The water surface anomaly detection network sequentially includes: an encoder module f θ , a texture enhancement module TEM, a pyramid texture feature extraction module PTFEM, and a classifier module and obtains the floating object area image Z n of X n : Step 3.1, the encoder module f θ is composed of r basic units, and each basic unit is composed of a two-dimensional convolution Conv2D and a LeakyReLU activation function layer in sequence; Taking h1 as the step size, divide the nth water surface image P n into H blocks of size h2 to obtain the block set P n ={w n,1 ,..., w n,h ,..., w n,H}, where H is the total number of blocks of P n , and w n,h is the hth block in P n ; Input w n,h into the encoder f θ for processing, and extract the h-th encoded feature A n,h ; Step 3.
2. Construct the encoder \(f\) using Equation (5) θ with a loss \(L\) θ as follows: In formula (5), w n,h′ is the adjacent block of w n,h ; Step 3.3, the texture enhancement module TEM is composed of a global average pooling layer, a one-dimensional quantization coding QCO layer, and an MLP multi-layer perceptron; For the h-th encoded feature A n,h After being processed by the global average pooling layer, the h-th texture pooling feature g is obtained n,h , and then the cosine similarity between each pixel feature vector of the encoded feature A n,h and the texture pooling feature g n,h is calculated to generate the feature similarity matrix G n,h . The one-dimensional quantization coding QCO layer performs one-dimensional quantization coding on the feature similarity matrix G n,h to obtain the h-th quantization coding matrix E n,h ; The h-th quantization coding matrix E n,h generates the h-th statistical feature D after passing through the MLP multi-layer perceptron n,h . The h-th quantization coding matrix E n,h is multiplied by the h-th statistical feature D n,h to obtain the h-th high-quality texture feature O n,h ; Step 3.4, the pyramid texture feature extraction module PTFEM is composed of b1 parallel branches with different scales, and each parallel branch contains a texture feature extraction unit; the texture feature extraction unit is successively composed of b2 two-dimensional quantization coding 2d-QCO layers and b3 multi-layer perceptrons MLP; Fuse the h-th high-quality texture feature O n,h with the h-th encoded feature A n,h to generate the h-th fused feature map K n,h and input it into the pyramid texture feature extraction module PTFEM. First, divide the h-th fused feature map K n,h into b1 feature maps of different scales, and then input them into b1 parallel branches of different scales for processing respectively. Each branch extracts the corresponding texture representation feature map through its own texture feature extraction unit wherein represents the texture representation feature map generated by the q-th branch; perform upsampling operations on the texture representation feature maps obtained from the b1 branches respectively to restore them to the size of the fused feature map K n,h and then fuse the b1 upsampled feature maps to generate the h-th multi-scale texture feature F n,h , and finally fuse F n,h with A n,h to generate the h-th texture feature B n,h ; Step 3.5, the classifier module is composed of x1 linear units, and each linear unit is successively composed of a linear layer and an activation function layer LeakyReLU; The h-th texture feature B n,h Input the classifier and get the hth block w n,h The abnormality scores of all blocks are obtained by calculating the abnormality scores of all blocks; among them, the nth water surface image P n The anomaly score of pixel i in is the mean of the sum of anomaly scores of all the slices where the pixel i is located; Set an abnormal score threshold value, and extract the abnormal region where the abnormal scores of the pixel points in the nth water surface image P n are higher than the abnormal score threshold value, so as to obtain the floating object region image Z n of X n ; Step 3.
6. Construct the classifier using Equation (6). loss In formula (6), Cross_entropy represents cross entropy, represents a random small block in the n n-th water surface image P represents any one of the eight small blocks in the eight directions around the center of , where 1 ≤ p1, p2 ≤ H, and y is the azimuth label relative to , and y = 1, 2,..., 8.
4. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor to execute the water surface floating object recognition method according to any one of claims 1-3, and the processor is configured to execute the program stored in the memory.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it executes the steps of the water surface floating object recognition method according to any one of claims 1-3.
Citation Information
Patent Citations
Low-quality underwater image fish target detection method
CN115410078A
High-speed rail overhead line system foreign matter detection method and system based on single classification and anomaly generation
CN115690730A