An offshore oil spill detection method, device, medium and equipment
By improving the U-Net network in offshore oil spill detection, using the ResNet50 network as an encoder and introducing a multi-head cross attention mechanism module, the problem of low accuracy of offshore oil spill detection in the prior art is solved, and more accurate feature extraction and identification of oil spill area is achieved.
Patent Information
- Application Number
- CN202411326514.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-09-23
AI Technical Summary
The prior art has a problem of low accuracy in offshore oil spill detection, especially the lack of accurate extraction of the characteristic features of isolated oil spill areas.
The improved U-Net network is built by replacing the encoder in the U-Net network with the ResNet50 network and adding a multi-head cross attention mechanism module to the input of the decoder deconvolution module. The network supplements the lost oil spill characteristics in the convolution operation through the splicing module, and captures long-distance dependencies through the multi-head cross-attention mechanism to enhance feature association.
The accuracy of offshore oil spill detection is improved, the loss of features of isolated oil spill areas during convolution is avoided, and the ability to identify and position oil spill areas is enhanced.
Smart Images

Figure CN119313920B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of marine monitoring, and particularly to a method, device, medium and equipment for detecting oil spills at sea. Background Art
[0002] An oil spill at sea usually refers to a maritime accident that causes marine pollution, such as a collision, grounding, or wave damage of a ship at sea, resulting in the leakage of fuel oil from a cargo ship. These leakage incidents usually occur in areas far from land, causing the oil to spread, disperse, and float on the ocean surface, causing extensive pollution and damage to marine life and ecosystems. Therefore, effective monitoring and estimation of the oil leakage volume are key steps in dealing with leakage incidents.
[0003] In the prior art, the U-Net network is a method for image segmentation, and can detect oil spills at sea by segmenting oil spill images. Its network structure consists of two parts: an encoder and a decoder, with feature transfer through skip connections in the middle. The encoder extracts high-level features of the image through a series of convolutional operations; the skip connections transfer the feature maps of the encoder to the decoder part to retain and utilize the detailed information, enabling the decoder to obtain fine-grained information from the encoder during the upsampling process; the decoder gradually restores the spatial resolution of the image through upsampling and uses the feature maps in the skip connections to assist in reconstruction.
[0004] However, after multiple convolutional operations of the encoder, the features of the larger oil spill areas in the oil spill image can be well extracted. However, the information of some isolated oil spill areas in the oil spill image may be lost, resulting in inaccurate oil spill detection results. Summary of the Invention
[0005] Based on this, in order to solve the technical problem of low accuracy of oil spill detection at sea in the prior art, the present invention provides a method, device, medium and equipment for detecting oil spills at sea.
[0006] The present invention provides a method for detecting oil spills at sea, including:
[0007] Replacing the encoder in the original U-Net network with a ResNet50 network to construct an improved encoder, where the ResNet50 network includes a splicing module for splicing the input and output of the convolutional operation; adding a multi-head cross-attention mechanism module to the input end of the deconvolution module of the decoder in the original U-Net network to construct an improved decoder; and constructing an improved U-Net network including the improved encoder, the improved decoder, and the skip connection of the original U-Net network;
[0008] Collecting oil spill images at sea to construct a dataset, and using the dataset to train the improved U-Net network to obtain an oil spill detection model;
[0009] Input the oil spill image to be detected into the oil spill detection model. Through the improved encoder, perform convolution operations on the oil spill image to be detected to reduce the size of the oil spill image to be detected. During the convolution operation, through the splicing module, splice the input of the convolution operation with the output of the convolution operation to supplement the loss of oil spill features caused by the reduction of the size of the oil spill image to be detected during the convolution operation, and obtain oil spill feature maps of different scales. Through the first decoder transposed convolution module in the improved decoder, perform transposed convolution operations on the oil spill feature map with the smallest scale to restore the size of the oil spill feature map. Before the remaining decoder transposed convolution modules perform transposed convolution operations, convey the oil spill feature map of the corresponding scale to the input end of the corresponding decoder transposed convolution module through skip connections to splice with the restored oil spill feature map output from the previous decoder transposed convolution module, and obtain the spliced feature map of the corresponding scale. And through the multi-head cross-attention mechanism module, capture the long-range dependencies in the spliced feature map to obtain the association between the oil spill features supplemented by the splicing module in the spliced feature map and other oil spill features in the spliced feature map, and weight the oil spill features supplemented by the residual connection module that are associated with other oil spill features in the spliced feature map to obtain the oil spill weighted feature map of the corresponding scale. The remaining decoder transposed convolution modules perform transposed convolution operations on the oil spill weighted feature map of the corresponding scale to obtain oil spill restored feature maps of different scales. Classify the pixels in the oil spill restored feature map with the largest size into oil spill pixels and non-oil spill pixels, and output the oil spill area in the oil spill image to be detected.
[0010] Further, the ResNet50 network includes 5 sequentially connected residual modules, and the standard convolution layers in each residual module are replaced with a convolution block and an identity block structure connected in sequence; the convolution block is used to adjust the channels and size of the oil spill feature map; the identity block is used to increase the depth of the network without changing the channels and size of the oil spill feature map to prevent the vanishing of gradients during the transmission of oil spill features.
[0011] Further, the conveying the oil spill feature map of the corresponding scale to the input end of the corresponding decoder transposed convolution module through skip connections to splice with the restored oil spill feature map output from the previous decoder transposed convolution module to obtain the spliced feature map of the corresponding scale specifically includes:
[0012] Splice the fourth oil spill feature map output by the fourth residual module in the ResNet50 network with the first restored oil spill feature map output by the first decoder transposed convolution module to obtain the first spliced feature map;
[0013] The third oil spill feature map output by the third residual module in the ResNet50 network and the second oil spill recovery feature map output by the second decoder transposed convolution module are concatenated to obtain a second concatenated feature map;
[0014] The second oil spill feature map output by the second residual module in the ResNet50 network and the third oil spill recovery feature map output by the third decoder transposed convolution module are concatenated to obtain a third concatenated feature map;
[0015] The first oil spill feature map output by the first residual module in the ResNet50 network and the fourth oil spill recovery feature map output by the fourth decoder transposed convolution module are concatenated to obtain a fourth concatenated feature map.
[0016] Further, adding a multi-head cross-attention mechanism module to the input end of the decoder transposed convolution module of the original U-Net network specifically includes:
[0017] Adding a first multi-head cross-attention mechanism module, a second multi-head cross-attention mechanism module, a third multi-head cross-attention mechanism module, and a fourth multi-head cross-attention mechanism module to the input ends of the second decoder transposed convolution module, the third decoder transposed convolution module, the fourth decoder transposed convolution module, and the fifth decoder transposed convolution module in the original U-Net network respectively:
[0018] The i-th multi-head cross-attention mechanism module is used to capture the long-range dependence relationship in the i-th concatenated feature map and output the i-th oil spill weighted feature map corresponding to the i-th concatenated feature map, where i = 1, 2, 3, 4; specifically including:
[0019] The i-th multi-head cross-attention mechanism module uses the oil spill feature map from the concatenation module in the i-th concatenated feature map as the query Q and key K in the attention mechanism, and uses the oil spill recovery feature map from the decoder transposed convolution module in the i-th concatenated feature map as the value V in the attention mechanism;
[0020] Based on the query Q and key K, calculate the attention weight matrix, multiply the attention weight matrix by the value V and then perform an upsampling operation to obtain a first feature matrix; perform positional encoding on the oil spill feature map from the residual module and then perform a dot product with the first feature matrix to obtain a first oil spill weighted feature map; perform positional encoding on the oil spill recovery feature map from the decoder transposed convolution module and then perform an upsampling operation to obtain a second oil spill weighted feature map;
[0021] Concatenate the first oil spill weighted feature map and the second oil spill weighted feature map to obtain the i-th oil spill weighted feature map output by the i-th multi-head cross-attention mechanism module.
[0022] The present invention provides an offshore oil spill detection device, including:
[0023] A model construction module, configured to construct an improved encoder by replacing the encoder in the original U-Net network with a ResNet50 network, where the ResNet50 network includes a splicing module for splicing the input of the convolution operation and the output of the convolution operation; construct an improved decoder by adding a multi-head cross-attention mechanism module to the input end of the deconvolution module of the original U-Net network decoder; and construct an improved U-Net network including the improved encoder, the improved decoder, and the skip connection of the original U-Net network.
[0024] A model training module, configured to collect offshore oil spill images to construct a dataset, and use the dataset to train the improved U-Net network to obtain an offshore oil spill detection model.
[0025] An oil spill detection module, configured to input the offshore oil spill image to be detected into the offshore oil spill detection model, perform convolution operations on the offshore oil spill image to be detected through the improved encoder to reduce the size of the offshore oil spill image to be detected. During the convolution operation, the input of the convolution operation and the output of the convolution operation are spliced through the splicing module to supplement the oil spill feature loss caused by the reduction of the size of the offshore oil spill image to be detected during the convolution operation, and obtain oil spill feature maps of different scales; perform deconvolution operations on the oil spill feature map with the smallest scale through the first decoder deconvolution module in the improved decoder to restore the size of the oil spill feature map; before performing deconvolution operations on the remaining decoder deconvolution modules, convey the oil spill feature map of the corresponding scale to the input end of the corresponding decoder deconvolution module through the skip connection to splice with the oil spill restoration feature map output from the previous decoder deconvolution module to obtain the spliced feature map of the corresponding scale; and capture the long-range dependence relationship in the spliced feature map through the multi-head cross-attention mechanism module to obtain the association between the oil spill features supplemented by the splicing module in the spliced feature map and other oil spill features in the spliced feature map, and weight the oil spill features supplemented by the residual connection module that are associated with other oil spill features in the spliced feature map to obtain the oil spill weighted feature map of the corresponding scale; the remaining decoder deconvolution modules perform deconvolution operations on the oil spill weighted feature map of the corresponding scale to obtain oil spill restoration feature maps of different scales; classify the pixels in the oil spill restoration feature map with the largest size into oil spill pixels and non-oil spill pixels, and output the oil spill area in the offshore oil spill image to be detected.
[0026] The present invention provides a computer-readable storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned offshore oil spill detection method is implemented.
[0027] The present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned offshore oil spill detection method is implemented.
[0028] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects:
[0029] In the offshore oil spill detection method provided by the present invention, the encoder is replaced with a ResNet50 network and a multi-head cross-attention mechanism module is introduced to improve the U-Net network; the ResNet50 module supplements the loss of oil spill features caused by the reduction of the size of the offshore oil spill image to be detected during the convolution operation by splicing the input of the convolution operation and the output of the convolution operation, avoiding the loss of features of isolated oil spill areas during the convolution process; after splicing the output of the ResNet50 network of the same scale with the output of the decoder's deconvolution and inputting it into the multi-head cross-attention mechanism module, since the multi-head cross-attention mechanism module can capture the long-range dependence relationships in the spliced feature map, the association between the oil spill features supplemented by the splicing module in the spliced feature map and other oil spill features in the spliced feature map can be obtained. Therefore, only the supplementary features that have a significant association with other oil spill features are given higher weights, rather than weighting all the supplementary features, transmitting the useful information in the supplementary features, and making the obtained information of isolated oil spill areas more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0031] Figure 1 is a schematic flow chart of an offshore oil spill detection method provided by the present invention;
[0032] Figure 2 is a schematic diagram of the U-Net network structure provided by the present invention;
[0033] Figure 3 is a schematic diagram of the Resnet50 convolutional block provided by the present invention;
[0034] Figure 4 is a schematic diagram of the Resnet50 identity block provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0036] Vessels, aircraft, satellites, etc. are widely used to identify and monitor oil spills. Traditional methods include aerial photography and on-site surveys, but they are costly. Some vessels are equipped with special radars to detect oil at sea, but their visibility and coverage are limited. Therefore, satellite observation has become one of the main systems for monitoring and controlling climate change. Remote sensing observations, especially those of SAR (Synthetic Aperture Radar) sensors, are favored due to their data collection capabilities. SAR sensors record backscatter signals of different frequencies and polarizations to form two-dimensional images. Since SAR sensors are not affected by sunlight and cloudy weather, have a wide coverage, and are less costly compared to aircraft equipped with SAR sensors, they are widely used in marine oil spill detection.
[0037] In this context, this embodiment provides a method for detecting marine oil spills. This method improves the original U-Net network by replacing the encoder part of the U-Net network with the residual blocks of Resnet50 and introducing a multi-head cross-attention mechanism in the residual connections of the decoder part to efficiently extract image information, and uses the improved U-Net network to segment SAR images into two categories: floating oil and background.
[0038] Figure 1 The flow of the method for detecting marine oil spills provided in this embodiment is shown. Specifically, in combination with Figure 1 The method for detecting marine oil spills will be described in detail as follows, which specifically includes the following steps:
[0039] S1: Replace the encoder in the original U-Net network with a ResNet50 network to construct an improved encoder. The ResNet50 network includes a splicing module for splicing the input of the convolution operation and the output of the convolution operation; add a multi-head cross-attention mechanism module to the input end of the deconvolution module of the decoder of the original U-Net network to construct an improved decoder; and construct an improved U-Net network including the improved encoder, the improved decoder, and the skip connection of the original U-Net network.
[0040] The U-Net network structure is as Figure 2 shown, and Figure 2The left part in it, that is, the encoder part, is replaced by the ResNet50 network. The ResNet50 network includes 5 sequentially connected residual modules; the decoder includes five decoder deconvolution modules connected in sequence; the outputs of the residual modules of the same scale are sent to the input ends of the corresponding decoder deconvolution modules through skip connections to splice the oil spill feature maps and oil spill recovery feature maps of the same scale, and the spliced feature maps are input into the multi-head cross-attention mechanism module for weighting operations, and the weighted spliced feature maps are input into the next decoder deconvolution module for further decoding.
[0041] In the ResNet50 network, the standard convolutional layer in each residual module is replaced by a sequentially connected convolutional block and identity block structure. The convolutional block is used to adjust the feature map channels and the size of the feature map to extract features of different scales; the identity block is used to increase the depth of the network without changing the size or number of channels of the feature map to prevent the vanishing gradient during the feature transfer process.
[0042] The network structure of the convolutional block is as Figure 3 shown, including an input layer, two parallel convolutional branches connected to the output end of the input layer at the same time, a splicing layer connected to the output ends of the two parallel convolutional branches at the same time, a ReLU activation function layer connected to the output end of the splicing layer, and an output layer; one of the convolutional branches includes a 1×1 convolutional layer, a batch normalization layer, a ReLU activation function layer, a 3×3 convolutional layer, a batch normalization layer, a ReLU activation function layer, a 1×1 convolutional layer, and a batch normalization layer connected in sequence; the other convolutional branch includes a 1×1 convolutional layer and a batch normalization layer connected in sequence.
[0043] The network structure of the identity block is as Figure 4 shown, including an input layer, a 1×1 convolutional layer, a batch normalization layer, a ReLU activation function layer, a 1×1 convolutional layer, a batch normalization layer, a ReLU activation function layer, a 1×1 convolutional layer, a batch normalization layer, a fusion layer, a ReLU activation function layer, and an output layer connected in sequence; among them, the output end of the input layer is connected to the input end of the fusion layer through a residual connection.
[0044] The working steps of the improved encoder are as follows:
[0045] The size of the input SAR image is 256x256x1, where 1 represents the number of channels. The image is input into the convolutional block of the first residual module, and the following operations are performed on one of the branches in the convolutional block:
[0046] First, perform a convolution operation using a 1x1 convolutional kernel and 16 convolutional layers with a stride of 1. The size of the output feature map is 256x256x16. Normalize the output feature map and apply the ReLU activation function. Then, perform a 3x3 convolution operation with the number of output channels set to 16 and padding set to "same" to keep the size unchanged. The size of the feature map generated in this step is still 256x256x16. Again, perform batch normalization and apply the ReLU activation function to the feature map. Next, perform another convolution operation using a 1x1 convolutional kernel and 16 convolutional layers and perform normalization. In another branch of the convolutional block, after performing a 1x1 convolution operation and normalization on the input image through a residual connection, concatenate it with the output of the previous branch. Apply the ReLU activation function to the concatenated result to obtain the output of the convolutional block, with the output being a feature of size 256x256x16.
[0047] Input the feature of size 256x256x16 into the identity block module and perform the following operations:
[0048] First, perform a convolution operation using a 1x1 convolutional kernel and 16 convolutional layers to output a feature map of size 256x256x16. Normalize the output feature map and apply the ReLU activation function. Then, perform a second 3x3 convolution operation with the number of output channels set to 16 and padding set to "same" to keep the size unchanged. The size of the feature map generated in this step is still 256x256x16. Again, perform batch normalization and apply the ReLU activation function to the feature map. After performing another convolution operation using a 1x1 convolutional kernel and 16 convolutional layers and performing normalization, concatenate it with the input feature through a residual connection. Apply the ReLU activation function to the concatenated result to obtain the output of the identity block, with the output being a feature of size 256x256x16. The convolutional block is placed in the first position to adjust the channels and the size of the feature map, and the identity block is used to deepen the network. After performing a 2x2 max pooling operation, obtain a feature map of size 128x128x16 pixels, which is the input to the second residual module.
[0049] Similarly, in the second residual module, the same steps as above are followed. First, it passes through a 1x1 convolution, normalization, and ReLU activation function in the convolutional block, and then a 3x3 convolution operation, normalization, and ReLU activation function are performed. Then, it passes through a 1x1 convolution and normalization operation. The difference from the steps of the first residual module is that the number of convolutional layers in the second residual module is 32, 32, and 32 respectively. Finally, after performing a 1x1 convolution operation on the input features of the second residual module, a residual connection is made with the output part, and a feature map with a size of 128x128x32 is output. Then, it passes through a 1x1 convolution, normalization, and ReLU activation function in the identity block, and then a 3x3 convolution operation, normalization, and ReLU activation function are performed, and then a 1x1 convolution and normalization operation. The number of convolutional layers in the identity block is 32, 32, and 32 respectively. Finally, a residual connection is made between the input features and the output part, and a feature map with a size of 128x128x32 is output. After a 2x2 max pooling operation, a feature map with a size of 64x64x32 is obtained, which serves as the input to the third residual module.
[0050] The processing method of the third residual module is the same as above. The number of convolutional layers in the convolutional block is replaced with 64, and the number of convolutional layers in the identity block is replaced with 64. After a max pooling operation, a feature map with a size of 32x32x64 is obtained, which serves as the input to the fourth residual module.
[0051] The processing method of the fourth residual module is the same as above. The number of convolutional layers in the convolutional block is replaced with 128, and the number of convolutional layers in the identity block is replaced with 128. After a max pooling operation, a feature map with a size of 16x16x128 is obtained, which serves as the input to the fifth residual module.
[0052] The processing method of the fifth residual module is the same as above. The number of convolutional layers in the convolutional block is replaced with 256, and the number of convolutional layers in the identity block is replaced with 256. Then, a bilinear upsampling operation is performed using the nearest neighbor interpolation method to obtain a feature map with a size of 32x32x128 pixels.
[0053] The upsampled feature map will be concatenated with the corresponding feature map of the encoder. After concatenation, the number of channels of the feature map will increase. During the concatenation process, a multi-head cross-attention mechanism module is added. The multi-head cross-attention mechanism module is used after the skip connection, mainly to combine the feature map with richer high-level semantics with the high-resolution map from the skip connection. The core idea of the multi-head cross-attention mechanism module is to filter out the irrelevant or noisy regions in the skip connection and highlight the relevant regions. The inputs of the multi-head cross-attention mechanism module are the result S (the output of each residual module) from the skip connection and the result Y (the output of each decoder transposed convolution module) processed by the previous layer of the decoder. Specifically:
[0054] The input of the first decoder deconvolution module is the fifth oil spill feature map output by the fifth residual module in the ResNet50 network, and the output of the first decoder deconvolution module is the first oil spill recovery feature map.
[0055] The fourth oil spill feature map output by the fourth residual module in the ResNet50 network and the first oil spill recovery feature map are concatenated to obtain the first concatenated feature map; the first concatenated feature map is weighted by the first multi-head cross-attention mechanism module to obtain the first oil spill weighted feature map; the input of the second decoder deconvolution module is the first oil spill weighted feature map, and the output of the second decoder deconvolution module is the second oil spill recovery feature map.
[0056] The third oil spill feature map output by the third residual module in the ResNet50 network and the second oil spill recovery feature map are concatenated to obtain the second concatenated feature map; the second concatenated feature map is weighted by the second multi-head cross-attention mechanism module to obtain the second oil spill weighted feature map; the input of the third decoder deconvolution module is the second oil spill weighted feature map, and the output of the third decoder deconvolution module is the third oil spill recovery feature map.
[0057] The second oil spill feature map output by the second residual module in the ResNet50 network and the third oil spill recovery feature map are concatenated to obtain the third concatenated feature map; the third concatenated feature map is weighted by the third multi-head cross-attention mechanism module to obtain the third oil spill weighted feature map; the input of the fourth decoder deconvolution module is the third oil spill weighted feature map, and the output of the fourth decoder deconvolution module is the fourth oil spill recovery feature map.
[0058] The first oil spill feature map output by the first residual module in the ResNet50 network and the fourth oil spill recovery feature map are concatenated to obtain the fourth concatenated feature map; the fourth concatenated feature map is weighted by the fourth multi-head cross-attention mechanism module to obtain the fourth oil spill weighted feature map; the input of the fifth decoder deconvolution module is the fourth oil spill weighted feature map, and the output of the fifth decoder deconvolution module is the fifth oil spill recovery feature map.
[0059] The inputs of the multi-head cross-attention mechanism module are the result S from the skip connection and the result Y processed by the previous layer respectively. The result after Y is embedded is used as K and Q in the attention mechanism, and the result output by S is used as V. That is, Q and K are taken from the current decoding feature, and V is taken from the corresponding encoding feature. The combination of features consists of two ways: one is to use Q and K to obtain the weight matrix of the current feature, multiply it by V to obtain the information of the important part in the segmentation, then upsample, and then take the position-encoded S for dot multiplication. The other is to upsample the position-encoded Y and then perform a convolution operation. Finally, the features obtained from the first and the second are concatenated together to get the new decoding feature. Specifically:
[0060] The $i$-th multi-head cross-attention mechanism module uses the features from the residual module in the $i$-th concatenated feature map as the query $Q$ and key $K$ in the attention mechanism, and uses the features from the decoder deconvolution module in the $i$-th concatenated feature map as the value $V$ in the attention mechanism; where $i = 1, 2, 3, 4$. Calculate the attention weight matrix based on the query $Q$ and key $K$, multiply the attention weight matrix by the value $V$ and then perform an upsampling operation to obtain the first feature matrix; perform positional encoding on the features from the residual module and then perform a dot product with the first feature matrix to obtain the first global feature map; perform positional encoding on the features from the decoder deconvolution module and then perform an upsampling operation to obtain the second global feature map; concatenate the first global feature map and the second global feature map to obtain the output of each multi-head cross-attention mechanism module.
[0061] The first decoder deconvolution module in the decoder uses two ordinary convolutional layers. Each convolutional layer includes 128 convolutional kernels, and the size of each convolutional kernel is $(3, 3)$, and the ReLU activation function is used to implement the decoding operation. The remaining decoder deconvolution modules are the same as above, all going through two ordinary convolutional layers, and the size of the convolutional kernels is $(3, 3)$, but the number of convolutional kernels becomes 64, 32, and 16 in sequence. The fifth decoder deconvolution module uses a $1\times1$ convolutional operation to obtain a feature map with a size of $256\times256\times1$ pixels, and this feature map is used as the final output of the improved U-Net network.
[0062] The above improved Unet network has the following improvements:
[0063] 1. Replace the ordinary convolutional layer with the convolutional block and identity block structures in ResNet50. By introducing residual connections, not only can the network be deepened to solve the gradient vanishing problem; but also concatenate the input of the convolutional operation with the output of the convolutional operation to supplement the spilled oil features lost during the convolutional operation, avoiding the loss of features in isolated spilled oil regions during the convolutional process. In the extraction of offshore spilled oil region features, it can better capture complex image features and more accurately identify spilled oil region features.
[0064] 2. The multi-head cross-attention mechanism is introduced into the skip connections of the U-Net network to interact the features of the encoder and the decoder, extract and fuse key features, and reduce the influence of noise. The multi-head cross-attention mechanism allows the model to focus on the most relevant features in the encoder during the decoding stage, enabling the model to more accurately identify and locate the oil spill area during the feature extraction process of the marine oil spill area. Since the multi-head cross-attention mechanism module can capture the long-range dependencies in the spliced feature maps, after splicing the output of the ResNet50 network of the same scale with the input of the decoder and inputting it into the multi-head cross-attention mechanism module, higher weights can be given only to those supplementary features that have a significant association with other oil spill features, rather than weighting all the features supplemented through residual connections, transmitting the useful information in the supplementary features and making the obtained isolated oil spill area information more accurate.
[0065] S2: Collect marine oil spill images to construct a dataset, and use the dataset to train the improved U-Net network to obtain a marine oil spill detection model.
[0066] First, a large number of datasets of oil spill events in different regions of the world are selected to train and evaluate the performance of the proposed framework. Each category is divided into a training set and a test set in a ratio of 7:3. And the SAR images in the dataset are geometrically corrected to obtain 700 SAR images for constructing the dataset. Then Python is used to process the SAR images in blocks: open the input digital elevation model file to obtain the spectral bands of red R, green G, and blue B; obtain the width (number of columns) and height (number of rows) of the image from the attributes of the red band, divide the width and height of the image by the pixel size, which is 256, to get the integer number of rows and columns to divide the image into blocks of a fixed size. Then calculate the remainders of the width and height of the image divided by the pixel size for processing incomplete pixel blocks. Then use a double loop to process the SAR images in blocks. The outer loop is responsible for traversing the rows of the image, and then the inner loop traverses the columns of each row, so that the image can be accessed and processed block by block without loading the entire image into memory at once, improving the efficiency of data processing. In the loop, first read a part of the image data from the red, green, and blue channels respectively to form an RGB image block. Then convert the RGB image block to a grayscale image by calling a function. Finally, check whether white pixels need to be added according to the conditions. If the conditions are met (there are remainders after segmentation of the height and width and the current image block is in the last row or the last column of the image), then adjust the size of the grayscale image block and add white pixels to the edge of the image block. Finally, multiply the pixel values of the processed grayscale image block by 512 and convert it to an unsigned 16-bit integer type. Save the processed image block data as a file in the ".TIFF" format. Use a function to splice the processed image blocks (or partitions) into a whole image.
[0067] Next, the improved U-Net network is trained using the dataset. The remote sensing images in the training set are preprocessed and converted into images with a size of 256x256x1 pixels, and then input into the CNN. During the model training process, the performance of the model is evaluated using accuracy, recall, and F1-score.
[0068] S3: Input the image of the oil spill to be detected into the oil spill detection model. The improved encoder is used to perform convolution operations on the image of the oil spill to be detected, so as to reduce the size of the image of the oil spill to be detected. During the convolution operation, the input of the convolution operation and the output of the convolution operation are concatenated through the concatenation module to supplement the loss of oil spill features caused by the reduction of the size of the image of the oil spill to be detected during the convolution operation, and oil spill feature maps of different scales are obtained; the first decoder deconvolution module in the improved decoder is used to perform deconvolution operations on the oil spill feature map with the smallest scale to restore the size of the oil spill feature map; before the remaining decoder deconvolution modules perform deconvolution operations, the oil spill feature maps of the corresponding scale are sent to the input end of the corresponding decoder deconvolution module through skip connections, so as to be concatenated with the restored oil spill feature map output from the previous decoder deconvolution module to obtain the concatenated feature map of the corresponding scale; and the long-range dependence relationship in the concatenated feature map is captured through the multi-head cross-attention mechanism module to obtain the association between the oil spill features supplemented by the concatenation module in the concatenated feature map and other oil spill features in the concatenated feature map, and the oil spill features supplemented by the residual connection module and associated with other oil spill features in the concatenated feature map are weighted to obtain the oil spill weighted feature map of the corresponding scale; the remaining decoder deconvolution modules perform deconvolution operations on the oil spill weighted feature map of the corresponding scale to obtain oil spill restored feature maps of different scales; the pixels in the oil spill restored feature map with the largest size are classified into oil spill pixels and non-oil spill pixels, and the oil spill area in the image of the oil spill to be detected is output.
[0069] The above is the oil spill detection method provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding oil spill detection device, including:
[0070] A model construction module, configured to construct an improved encoder by replacing the encoder in the original U-Net network with a ResNet50 network, where the ResNet50 network includes a concatenation module for concatenating the input of the convolution operation and the output of the convolution operation; construct an improved decoder by adding a multi-head cross-attention mechanism module to the input end of the decoder deconvolution module in the original U-Net network; and construct an improved U-Net network including the improved encoder, the improved decoder, and the skip connection of the original U-Net network.
[0071] A model training module, which is used to collect offshore oil spill images to construct a dataset, and use the dataset to train an improved U-Net network to obtain an offshore oil spill detection model.
[0072] An oil spill detection module, which is used to input the offshore oil spill image to be detected into the offshore oil spill detection model, and perform convolution operations on the offshore oil spill image to be detected through an improved encoder to reduce the size of the offshore oil spill image to be detected. During the convolution operation, the input of the convolution operation is spliced with the output of the convolution operation through a splicing module to supplement the loss of oil spill features caused by the reduction of the size of the offshore oil spill image to be detected during the convolution operation, and obtain oil spill feature maps of different scales; perform deconvolution operations on the oil spill feature map with the smallest scale through the first decoder deconvolution module in the improved decoder to restore the size of the oil spill feature map; before performing deconvolution operations on the remaining decoder deconvolution modules, send the oil spill feature map of the corresponding scale to the input end of the corresponding decoder deconvolution module through a skip connection to splice with the oil spill restoration feature map output from the previous decoder deconvolution module to obtain a spliced feature map of the corresponding scale; and capture the long-range dependence relationship in the spliced feature map through a multi-head cross-attention mechanism module to obtain the association between the oil spill features supplemented by the splicing module in the spliced feature map and other oil spill features in the spliced feature map, and weight the oil spill features supplemented by the residual connection module that are associated with other oil spill features in the spliced feature map to obtain an oil spill weighted feature map of the corresponding scale; the remaining decoder deconvolution modules perform deconvolution operations on the oil spill weighted feature map of the corresponding scale to obtain oil spill restoration feature maps of different scales; classify the pixels in the oil spill restoration feature map with the largest size into oil spill pixels and non-oil spill pixels, and output the oil spill area in the offshore oil spill image to be detected.
[0073] For the specific limitations of the offshore oil spill detection device, reference can be made to the limitations of the offshore oil spill detection method in the above text, which will not be elaborated here. Each module in the above offshore oil spill detection device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0074] The present invention also provides a computer-readable storage medium, which stores a computer program, and the computer program can be used to execute the above Figure 1 provided offshore oil spill detection method.
[0075] The present invention also provides a computer device structure. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 provided offshore oil spill detection method.
[0076] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include at least one of non-volatile and volatile memories. The non-volatile memory can include a read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. The volatile memory can include a random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as a static random access memory (SRAM) or a dynamic random access memory (DRAM), etc.
[0077] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded by the present invention.
Claims
1. A method for detecting oil spills at sea, characterized in that: include: An improved encoder is constructed by replacing the encoder in the original U-Net network with a ResNet50 network, wherein the ResNet50 network includes a splicing module for splicing the input of the convolution operation and the output of the convolution operation; an improved decoder is constructed by adding a multi-head cross attention mechanism module at the input end of the deconvolution module of the original U-Net network decoder; and an improved U-Net network including an improved encoder, an improved decoder and a jump connection of the original U-Net network is constructed; Collect offshore oil spill images to build a data set, use the data set to train the improved U-Net network, and obtain an offshore oil spill detection model; The offshore oil spill image to be detected is input into the offshore oil spill detection model, and the offshore oil spill image to be detected is convolved by improving the encoder to reduce the size of the offshore oil spill image to be detected. During the convolution operation, the input of the convolution operation is spliced with the output of the convolution operation by a splicing module to compensate for the loss of oil spill features caused by the size reduction of the offshore oil spill image to be detected during the convolution operation, so as to obtain oil spill feature maps of different scales; the first decoder deconvolution module in the improved decoder is used to perform a deconvolution operation on the minimum-scale oil spill feature map to restore the size of the oil spill feature map; Before the residual decoder deconvolution module performs deconvolution operation, the oil spill feature map of the corresponding scale is transmitted to the input end of the corresponding decoder deconvolution module through jump connection to be spliced with the oil spill recovery feature map output from the previous decoder deconvolution module to obtain the spliced feature map of the corresponding scale; and the long-distance dependency in the spliced feature map is captured through the multi-head cross attention mechanism module to obtain the association between the oil spill features supplemented by the splicing module in the splicing feature map and other oil spill features in the splicing feature map, and the oil spill features supplemented by the residual connection module that are associated with other oil spill features in the splicing feature map are weighted to obtain the oil spill weighted feature map of the corresponding scale; the residual decoder deconvolution module performs deconvolution operation on the oil spill weighted feature map of the corresponding scale to obtain oil spill recovery feature maps of different scales; the pixels in the oil spill recovery feature map of the largest size are classified into oil spill pixels and non-oil spill pixels, and the oil spill area in the marine oil spill image to be detected is output.
2. The method for detecting oil spill at sea as claimed in claim 1, characterized in that: The ResNet50 network includes 5 residual modules connected in sequence, and the standard convolution layer in each residual module is replaced by a convolution block and an identity block structure connected in sequence; the convolution block is used to adjust the channel and size of the oil spill feature map; the identity block is used to increase the depth of the network without changing the channel and size of the oil spill feature map to prevent the gradient from disappearing during the oil spill feature transmission process.
3. The method for detecting oil spill at sea as claimed in claim 2, characterized in that: The step of transmitting the oil spill feature map of the corresponding scale to the input end of the corresponding decoder deconvolution module through the jump connection to be spliced with the oil spill recovery feature map output from the previous decoder deconvolution module to obtain the spliced feature map of the corresponding scale specifically includes: The fourth oil spill feature map output by the fourth residual module in the ResNet50 network and the first oil spill recovery feature map output by the first decoder deconvolution module are spliced to obtain a first spliced feature map; The third oil spill feature map output by the third residual module in the ResNet50 network and the second oil spill recovery feature map output by the second decoder deconvolution module are spliced to obtain a second spliced feature map; The second oil spill feature map output by the second residual module in the ResNet50 network and the third oil spill recovery feature map output by the third decoder deconvolution module are spliced to obtain a third spliced feature map; The first oil spill feature map output by the first residual module in the ResNet50 network and the fourth oil spill recovery feature map output by the fourth decoder deconvolution module are spliced to obtain a fourth spliced feature map.
4. The method for detecting oil spill at sea as claimed in claim 3, characterized in that: The method adds a multi-head cross attention mechanism module to the input of the original U-Net network decoder deconvolution module, specifically including: In the original U-Net network, the first multi-head cross-attention mechanism module, the second multi-head cross-attention mechanism module, the third multi-head cross-attention mechanism module and the fourth multi-head cross-attention mechanism module are added to the input ends of the second decoder deconvolution module, the third decoder deconvolution module, the fourth decoder deconvolution module and the fifth decoder deconvolution module respectively: The i-th multi-head cross attention mechanism module is used to capture the long-range dependencies in the i-th concatenated feature map and output the i-th oil spill weighted feature map corresponding to the i-th concatenated feature map, i = 1, 2, 3, 4; specifically includes: The i-th multi-head cross attention mechanism module uses the oil spill feature map from the splicing module in the i-th splicing feature map as the query Q and key K in the attention mechanism, and uses the oil spill recovery feature map from the decoder deconvolution module in the i-th splicing feature map as the value V in the attention mechanism; The attention weight matrix is calculated based on the query Q and the key K, and the attention weight matrix is multiplied by the value V and then upsampled to obtain a first feature matrix; the oil spill feature map from the residual module is position-encoded and dot-multiplied with the first feature matrix to obtain a first oil spill weighted feature map; the oil spill recovery feature map from the decoder deconvolution module is position-encoded and upsampled to obtain a second oil spill weighted feature map; The first oil spill weighted feature map and the second oil spill weighted feature map are spliced to obtain the i-th oil spill weighted feature map output by the i-th multi-head cross attention mechanism module.
5. A marine oil spill detection device, characterized in that: include: A model building module, for building an improved encoder by replacing the encoder in the original U-Net network with a ResNet50 network, wherein the ResNet50 network includes a splicing module for splicing the input of the convolution operation and the output of the convolution operation; building an improved decoder by adding a multi-head cross attention mechanism module at the input end of the deconvolution module of the original U-Net network decoder; and building an improved U-Net network including an improved encoder, an improved decoder and a jump connection of the original U-Net network; The model training module is used to collect offshore oil spill images to build a data set, and use the data set to train the improved U-Net network to obtain an offshore oil spill detection model; The oil spill detection module is used to input the marine oil spill image to be detected into the marine oil spill detection model, and perform convolution operation on the marine oil spill image to be detected by improving the encoder to reduce the size of the marine oil spill image to be detected. During the convolution operation, the input of the convolution operation is spliced with the output of the convolution operation by the splicing module to compensate for the loss of oil spill features caused by the size reduction of the marine oil spill image to be detected during the convolution operation, so as to obtain oil spill feature maps of different scales; the first decoder deconvolution module in the improved decoder is used to perform deconvolution operation on the oil spill feature map of the smallest scale to restore the size of the oil spill feature map; Before the residual decoder deconvolution module performs deconvolution operation, the oil spill feature map of the corresponding scale is transmitted to the input end of the corresponding decoder deconvolution module through jump connection to be spliced with the oil spill recovery feature map output from the previous decoder deconvolution module to obtain the spliced feature map of the corresponding scale; and the long-distance dependency in the spliced feature map is captured through the multi-head cross attention mechanism module to obtain the association between the oil spill features supplemented by the splicing module in the splicing feature map and other oil spill features in the splicing feature map, and the oil spill features supplemented by the residual connection module that are associated with other oil spill features in the splicing feature map are weighted to obtain the oil spill weighted feature map of the corresponding scale; the residual decoder deconvolution module performs deconvolution operation on the oil spill weighted feature map of the corresponding scale to obtain oil spill recovery feature maps of different scales; the pixels in the oil spill recovery feature map of the largest size are classified into oil spill pixels and non-oil spill pixels, and the oil spill area in the marine oil spill image to be detected is output.
6. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
7. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Transformer oil leakage identification method and system
CN114757938A
SAR marine oil spill image segmentation method based on improved U-Net
CN117975006A