Underwater target detection method and device
Through the underwater object detection method of multi-level nested U-shaped network and frequency domain channel attention layer, the problems of underwater image edge blur and high-frequency information loss are solved, and high-precision significant object detection in complex underwater environments are achieved.
Patent Information
- Application Number
- CN202510536458.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-26
AI Technical Summary
Due to blurring edges and loss of high-frequency information in underwater images, existing object detection algorithms cannot accurately identify significant targets.
Using a multi-level nested U-shaped network structure, combined with the frequency domain channel attention layer and the significance graph generation module, the underwater image is detected through the pre-trained model, and the frequency domain channel attention layer is used to enhance feature extraction, and combined with the edge enhancement of weighted cross entropy loss function optimization model training.
Effectively capture and emphasize target objects in complex underwater environments, improve detection accuracy and robustness, and accurately identify significant targets in underwater images.
Smart Images

Figure CN120544016A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an underwater target detection method and device. Background Art
[0002] Underwater images are widely used in fields such as exploring seabed resources, monitoring the marine environment, and visually guiding underwater robots. These underwater images are crucial for studying underwater life, underwater archaeology, pipeline repair, and performing various tasks. However, due to the inherent complexity of the underwater environment and variable lighting conditions, underwater images may suffer from color degradation, reduced contrast, and blurred details. These problems greatly increase the difficulty of underwater salient object detection tasks. Specifically, they are reflected in the following aspects:
[0003] (1) Due to the absorption and scattering of light by water in underwater environments, there is a significant deviation in the color of the image. This color distortion will cause the boundary between the target object and the surrounding environment to become blurred, seriously affecting the accuracy of salient target detection.
[0004] (2) Due to the scattering and absorption effects of water, the contrast between the target object and the background is low, which makes it difficult for traditional target detection algorithms to effectively separate the target from the background.
[0005] (3) Due to the limited penetration distance of light in underwater environments, the details in the imaging (that is, high-frequency information) are blurred and lost, which is especially reflected in the edges and textures of the target objects, making it difficult for traditional target detection algorithms to accurately capture the characteristics of the target.
[0006] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0007] The embodiments of the present application provide an underwater target detection method and device to at least solve the technical problem that existing target detection algorithms are unable to accurately identify significant targets in underwater images due to the characteristics of edge blur and high-frequency information loss in underwater images.
[0008] According to one aspect of an embodiment of the present application, a method for underwater target detection is provided, comprising: acquiring a first underwater image; performing target detection on the first underwater image using a pretrained underwater target detection model to obtain a first saliency map corresponding to the first underwater image, and determining a first saliency target in the first underwater image based on the first saliency map, wherein the predicted pixel value of each pixel in the first saliency map is used to represent the probability that the pixel belongs to the first saliency target, and the underwater target detection model comprises: a multi-level nested U-shaped network and a saliency map generation module, and each level of the U-shaped network comprises: a shallow feature extraction module and a deep feature extraction module, and the shallow feature extraction module comprises at least: a frequency domain channel attention layer.
[0009] Optionally, the training process of the underwater target detection model includes: constructing an initial learning model including a shallow feature extraction module, a deep feature extraction module and a saliency map generation module; obtaining a training sample set and a sample label set, wherein the training sample set includes multiple second underwater images as training samples, and the sample label set includes a binary image corresponding to each second underwater image as a sample label, and the mask value of each pixel point in the binary image is used to characterize whether the pixel point belongs to the salient target area in the second underwater image; using the training sample set and the sample label set to iteratively train the initial learning model to obtain an underwater target detection model.
[0010] Optionally, the shallow feature extraction module in each level of the U-shaped network includes N levels of downsampling layers, and the first M levels of downsampling layers are U-shaped blocks, and the last NM levels of downsampling layers are residual U-shaped blocks of expanded versions. The first M levels of downsampling layers all include frequency domain channel attention layers. The shallow feature extraction module is used to extract shallow features of underwater images to obtain corresponding shallow feature maps, where M is less than N, M is greater than UM, and M and N are both positive integers; the deep feature extraction module in each level of the U-shaped network includes N-1 levels of upsampling layers, and the first NM-1 levels of upsampling layers are expanded versions. The residual U-shaped block of Zhang's version is used, and the remaining M-level upsampling layers are U-shaped blocks. The deep feature extraction module is used to fuse the feature map output by each upsampling layer with the feature map output by the corresponding level of downsampling layer through skip connection to obtain a deep feature map; the saliency map generation module includes at least: a convolution layer, an upsampling layer, a feature map fusion layer, and a sigmoid activation function layer. The saliency map generation module is used to fuse the feature map output by the N-th level downsampling layer and the feature map output by the N-1 level upsampling layer to obtain the corresponding saliency map.
[0011] Optionally, the expanded version of the residual U-shaped block includes: at least one expanded convolution layer, a residual connection layer, a sampling layer, and an output layer, wherein the expanded convolution layer is used to extract features from the feature map output by the previous sampling layer according to a preset expansion rate to obtain a corresponding feature map.
[0012] Optionally, the frequency domain channel attention layer includes: a compression module and an activation module, wherein the compression module is used to split the feature map output by the previous downsampling layer along the channel dimension to obtain multiple two-dimensional matrices, and perform discrete cosine transform on each two-dimensional matrix to obtain the corresponding two-dimensional frequency component, and then splice the two-dimensional frequency components corresponding to each two-dimensional matrix to obtain the corresponding compressed channel attention vector; the activation module includes a first fully connected layer, a Relu activation function layer, a second fully connected layer and a sigmoid activation function layer connected in sequence, wherein the first fully connected layer is used to perform dimensionality reduction processing on the channel attention vector output by the compression module; the second fully connected layer is used to perform dimensionality recovery processing on the channel attention vector processed by the Relu activation function layer; the output of the activation module is a feature map obtained by channel-by-channel multiplication of the channel attention vector output by the second fully connected layer and the feature map output by the previous downsampling layer.
[0013] Optionally, the initial learning model is iteratively trained using the training sample set and the sample label set to obtain an underwater target detection model, including: for each training batch in the iterative training process, each training sample of the training batch is input into the initial learning model to obtain a second saliency map corresponding to each training sample output by the initial learning model; the target loss function is constructed using the sample labels and the second saliency map corresponding to each training sample of the training batch, and the model parameters of the initial learning model are adjusted according to the target loss function.
[0014] Optionally, the target loss function is a weighted cross entropy loss function based on edge enhancement, and the expression of the target loss function is:
[0015]
[0016] Where N represents the total number of training samples in the training batch, H and W represent the height and width of the training samples respectively, and g n,h,w Represents the mask value of the binary image corresponding to the nth training sample at position (h, w), s n,h,w represents the predicted pixel value at position (h, w) of the second saliency map corresponding to the nth training sample, and s, g∈{0,1}, e n,h,w represents the value of the edge feature map at position (h, w) obtained by performing an average pooling operation on the binary image corresponding to the nth training sample using a sliding window of a preset size, and e n,h,w The expression can be written as:
[0017]
[0018] Among them, G n ={g n |0 <g n<1}∈R N×1×H×W Represents the binary image corresponding to the nth training sample, S(G n )={s n |0 n <1}∈R N×1×H×W represents the second saliency map corresponding to the nth training sample.
[0019] According to another aspect of an embodiment of the present application, an underwater target detection device is also provided, including: an acquisition module for acquiring a first underwater image; a detection module for performing target detection on the first underwater image using a pre-trained underwater target detection model, obtaining a first saliency map corresponding to the first underwater image, and determining a first saliency target in the first underwater image based on the first saliency map, wherein the predicted pixel value of each pixel point in the first saliency map is used to represent the probability that the pixel point belongs to the first saliency target, and the underwater target detection model includes: a multi-level nested U-shaped network and a saliency map generation module, and each level of the U-shaped network includes: a shallow feature extraction module and a deep feature extraction module, and the shallow feature extraction module includes at least: a frequency domain channel attention layer.
[0020] According to another aspect of an embodiment of the present application, a computer program product is further provided, the computer program product comprising: a computer program, wherein when the computer program is executed by a processor, the above-mentioned underwater target detection method is implemented.
[0021] According to another aspect of an embodiment of the present application, an electronic device is further provided, comprising: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned underwater target detection method through the computer program.
[0022] In an embodiment of the present application, a pre-trained underwater target detection model is used to perform target detection on a first underwater image to obtain a first saliency map corresponding to the first underwater image, and a first salient target within the first underwater image is determined based on the first saliency map. The underwater target detection model includes a multi-level nested U-shaped network and a saliency map generation module, and each level of the U-shaped network includes a shallow feature extraction module and a deep feature extraction module, each of which includes at least a frequency domain channel attention layer. The above-mentioned underwater target detection model can achieve the technical effect of effectively capturing and emphasizing target objects in complex underwater environments, thereby achieving the purpose of improving detection accuracy and robustness, and thereby solving the technical problem that existing target detection algorithms cannot accurately identify salient targets within underwater images due to the characteristics of blurred edges and loss of high-frequency information in underwater images. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0024] Figure 1 is a flow chart of an optional underwater target detection method according to an embodiment of the present application;
[0025] Figure 2 is a schematic structural diagram of an optional underwater target detection device according to an embodiment of the present application;
[0026] Figure 3 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0027] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0028] It should be noted that the terms "first", "second", etc. in the specification, claims, and drawings of the present application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.
[0029] In order to better understand the embodiments of the present application, some nouns or terms that appear in the description of the embodiments of the present application are first translated and explained as follows:
[0030] Salient Object Detection (SOD): It is used to distinguish the most obvious areas in an image. It is widely used in the field of computer vision, such as visual tracking, image retrieval, non-photo-level rendering, 4D saliency detection, and reference-free image quality assessment.
[0031] Upsampling: Also known as image enlargement or image interpolation, its primary purpose is to enlarge the original image. Common upsampling methods include nearest neighbor interpolation, bilinear interpolation, bicubic interpolation, deconvolution (also known as transposed convolution), and unpooling. Deconvolution and bilinear interpolation are commonly used in deep learning models because they effectively increase the size of feature maps while preserving or restoring the information in them.
[0032] Downsampling: Also known as downsampling, it reduces the size of an image. Its purpose is to make the image fit within the display area, thereby generating a thumbnail of the corresponding image. Therefore, downsampling can be understood as pooling.
[0033] Pooling: A downsampling operation. It generates a new feature map by defining a spatial neighborhood (usually a rectangular area) and performing statistical processing (such as taking the maximum value or average value) on the features within the spatial neighborhood. Therefore, the pooling operation usually follows the convolutional layer.
[0034] Convolution: It is a mathematical operation that slides a sliding window (also called a convolution kernel or filter) on the input image or feature map and calculates the weighted sum (including the bias term) of the elements in the window and the corresponding convolution kernel elements to generate a feature map.
[0035] U-Net: A convolutional neural network based on deep learning, primarily used for image segmentation tasks. It consists of two parts: an encoder (downsampling path) and a decoder (upsampling path), and is named U-Net because of its U-shaped structure. U-Net uses skip connections to combine feature maps from the encoder and decoder to preserve more spatial information and improve positioning accuracy.
[0036] Dilated convolution: Also known as atrous convolution or dilated convolution, it differs from normal convolution in the number of parameters within the neural network, except for the size of the convolution kernel. The only difference is the introduction of a hyperparameter called the dilation rate, which defines the spacing between values when the convolution kernel processes data. As a result, dilated convolution has a larger receptive field.
[0037] The Discrete Cosine Transform (DCT) is a transform related to the Fourier Transform. It is similar to the Discrete Fourier Transform, but uses only real numbers. The Discrete Cosine Transform is equivalent to a Discrete Fourier Transform of approximately twice its length, performed on a real even number. There are generally several forms of the Discrete Cosine Transform, the most commonly used being DCT-II (also known as the "inverse discrete cosine transform" or "inverse discrete cosine transform"). As a common image compression method, DCT-II converts images from the spatial domain to the frequency domain. In the frequency domain, the low-frequency components of an image (i.e., those with gentle changes) typically contain the majority of the image information, while the high-frequency components (i.e., those with drastic changes) contain less information. Therefore, image compression can be achieved by retaining the low-frequency components and discarding the high-frequency components.
[0038] Example 1
[0039] According to an embodiment of the present application, a method for underwater target detection is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0040] Figure 1 FIG. 1 is a flow chart of an underwater target detection method according to an embodiment of the present application, such as Figure 1 As shown, the method includes the following steps S101-S104, wherein:
[0041] Step S102: Acquire a first underwater image.
[0042] In the technical solution provided in step S102 above, the first underwater image can be any form of digital image captured or generated in an underwater environment. This includes, but is not limited to: natural underwater landscape images (such as images containing various marine organisms, plants, and other natural objects), underwater archaeological images (such as images containing underwater shipwreck sites and ancient relics), industrial inspection images (such as images captured during the inspection of underwater pipelines, bridges, docks, or other infrastructure), underwater sports images (such as sports images of divers, swimmers, and underwater athletes), etc. This embodiment of the application does not impose any specific restrictions on the type of the first underwater image.
[0043] In step S104 , a pre-trained underwater object detection model is used to perform object detection on the first underwater image to obtain a first salient map corresponding to the first underwater image, and a first salient object in the first underwater image is determined based on the first salient map.
[0044] In the technical solution provided in step S104 above, a pre-trained underwater object detection model is invoked to perform object detection on the first underwater image, thereby obtaining a first saliency map corresponding to the first underwater image. The predicted pixel value of each pixel in the first saliency map represents the probability that the pixel belongs to the first salient object. Furthermore, the first salient object in the first underwater image can be determined based on the first saliency map.
[0045] The underwater target detection model in the embodiment of the present application includes: a multi-level nested U-shaped network U n Net (n represents the level of the nested U-shaped structure and can theoretically be set to any positive integer) and a saliency map generation module. Each level of the U-shaped network includes: a shallow feature extraction module (for downsampling operations) and a deep feature extraction module (for upsampling operations). This design is mainly due to the fact that underwater images often face imaging quality degradation caused by water characteristics, such as color degradation, reduced contrast, blurred details, high noise, etc. Therefore, the underwater target detection model adopts a multi-level nested U-shaped network structure, which enables the model to gradually enrich the level and depth of features during the process of layer-by-layer feature extraction, thereby achieving a comprehensive and in-depth analysis of underwater images.
[0046] In addition, the shallow feature extraction module includes at least: a frequency domain channel attention layer. That is, the frequency domain channel attention layer is used to replace the original convolution layer in the shallow feature extraction module in order to more accurately focus on the key information of specific features of underwater images. At the same time, frequency domain analysis is introduced, so that the model can better understand complex phenomena such as light scattering and color distortion in underwater images, thereby improving the understanding and analysis of image content.
[0047] Based on the scheme defined by the above steps S102 to S104, it can be known that in the embodiment of the present application, by designing the feature extraction part of the model as a multi-level nested U-shaped network structure, the model performs feature learning and fusion at different scales. At the same time, the use of the frequency domain channel attention layer in the feature extraction part can effectively focus on and enhance the key information in the frequency domain in the initial processing stage and the later deepening stage of the underwater image, respectively, so that the model is more sensitive to the edge and texture features of the salient targets in underwater image processing, thereby providing more accurate predictions on complex boundaries; and then the salient map generation module uses a multi-stage output strategy to integrate feature information at different levels to generate the final salient map to determine the salient targets in the underwater image, thereby achieving the technical effect of effectively capturing and emphasizing the target objects in a complex underwater environment, and achieving the purpose of improving detection accuracy and robustness.
[0048] The following describes the various steps of the underwater target detection method in combination with a specific implementation process.
[0049] As an optional implementation, in the technical solution provided in step S104 above, the training process of the underwater target detection model may include:
[0050] Step S1: construct an initial learning model including a multi-level nested U-shaped network and a saliency map generation module.
[0051] Specifically, the initial learning model includes a multi-level nested U-shaped network and a saliency map generation module, and each level of the U-shaped network includes: a shallow feature extraction module and a deep feature extraction module, and the shallow feature extraction module and the deep feature extraction module include at least: a frequency domain channel attention layer.
[0052] Optionally, for the shallow feature extraction module in each level of the U-shaped network, the shallow feature extraction module includes N levels of downsampling layers, where:
[0053] The first M downsampling layers are U-shaped blocks, and all of them include a frequency domain channel attention layer. This design allows the underwater object detection model to focus on more important channel information in the frequency domain while performing convolution operations, thereby enhancing the expressiveness of features. It should be noted that the size of the U-shaped blocks used in the first M downsampling layers can be configured according to the spatial resolution of the input feature map to ensure that detailed information can be effectively captured at different scales. Therefore, for larger input feature maps, larger convolution kernels and frequency domain channel attention layers can be used to capture more details.
[0054] The last NM-level downsampling layers are dilated versions of the residual U-block. This is because the resolution of the feature map is already low in the following downsampling layers. If the sampling rate is further reduced, details that are crucial to the background information may be lost. The dilated version of the residual U-block can replace the traditional pooling and upsampling operations with dilated convolutions, thereby maintaining the resolution of the feature map unchanged.
[0055] The shallow feature extraction module in the above structure not only retains more spatial information when extracting shallow features from underwater images, but also expands the model's receptive field, enabling it to capture a wider range of contextual information. By avoiding pooling and upsampling operations, information loss is reduced.
[0056] The frequency domain channel attention layer is introduced in the first M-level downsampling process to deal with the loss of underwater image information caused by color distortion and its interference with the underwater salient object detection task. The frequency domain channel attention layer includes: a compression module and an activation module, where:
[0057] (1)Compression module.
[0058] In the channel attention mechanism, feature maps are compressed according to the channel dimension, converting each two-dimensional feature channel into a real number with a global receptive field, and the output dimension matches the number of input feature channels. This compression method can effectively capture information from the entire feature space and provide richer context for the network.
[0059] At present, the traditional channel attention mechanism mainly relies on global average pooling (GAP) to capture the global information of each channel, wherein GAP obtains a real number that can reflect the importance of the channel by averaging the feature maps of each channel. However, although the global average pooling (GAP) method can capture the global information of each channel, it may ignore the differences and correlations between different frequency components. In contrast, the discrete cosine transform (DCT) can better retain the features of more frequencies and extract more robust feature representations. To this end, the embodiment of the present application proposes the use of discrete cosine transform (DCT) to replace global average pooling (GAP) to achieve feature channel compression in the frequency domain.
[0060] This is because the expression of one-dimensional discrete cosine transform can be written as:
[0061]
[0062] Among them, f k Represents the value of the kth frequency component in the frequency domain after DCT transformation, x i Represents the value of the i-th sample point in the time domain (ie, spatial domain), L represents the total number of sample points, k represents the index of the frequency component, and i represents the index of the sample point.
[0063] Then when k=0, the above expression can be written as:
[0064]
[0065] The above f0 represents all sample points x i If f0 is divided by the total number of all sample points L, the average value of the sample points is obtained, that is:
[0066]
[0067] The above expression becomes the definition of global average pooling (GAP).
[0068] The expression of two-dimensional discrete cosine transform can be written as:
[0069]
[0070] sth∈{0,1,…,H-1},w∈{0,1,…,W-1}
[0071] in, Represents the values of the hth and wth height components in the frequency domain after two-dimensional DCT transformation. Represents the value of the sample point in the i-th row and j-th column of the spatial domain, H represents the total height of the sample points, W represents the total width of the sample points, h represents the index of the frequency component in the height direction, w represents the index of the frequency component in the width direction, i represents the row index of the sample point, and j represents the column index of the sample point.
[0072] Then when h=0, w=0, the above expression can be written as:
[0073]
[0074] above Represents all sample points The sum of Dividing it by the total number of all sample points H*W, we get the average value of the sample points, that is:
[0075]
[0076] The above expression becomes the definition of global average pooling (GAP).
[0077] Therefore, global average pooling (GAP) can be viewed as the DC component of the discrete cosine transform (DCT), and its result is proportional to the lowest frequency component of the two-dimensional DCT. Therefore, using GAP in the channel attention mechanism means that only the lowest frequency information is retained, while information at other frequencies is discarded. However, this omitted frequency domain information also contains useful information about the channel. In underwater image processing, the loss or blurring of high-frequency information is a common problem, so relying solely on GAP may limit the model's understanding and processing capabilities of image details.
[0078] Specifically, the compression module can be processed according to the following process:
[0079] Step 1: Split the feature map output by the previous downsampling layer along the channel dimension to obtain multiple two-dimensional matrices, where each matrix represents a specific feature channel in the feature map. Therefore, the multiple two-dimensional matrices obtained by segmentation can be expressed as: [X0, X1, ..., X n-1 ], and X i ∈R 1×H×W ,i∈{0,1,...,n-1}.
[0080] Step 2: Perform discrete cosine transform on each two-dimensional matrix to obtain the corresponding two-dimensional frequency components. i The corresponding two-dimensional frequency components can be expressed as:
[0081]
[0082] in, represents the compressed feature vector, C i Indicates the number of low-frequency components retained by the i-th two-dimensional matrix after DCT transformation, [u i ,v i ] is the same as the i-th two-dimensional matrix X i The index of the corresponding two-dimensional frequency component, Represents the i-th two-dimensional matrix X i The pixel value at position (h,w), Denotes the DCT-II transform at position (h,w) corresponding to the frequency index [u i ,v i ], and the basis function of DCT-II is:
[0083]
[0084] Step 3: Concatenate the two-dimensional frequency components corresponding to each two-dimensional matrix to obtain the corresponding compressed channel attention vector. Therefore, the expression of the compressed channel attention vector can be written as:
[0085] Freq=cat([Freq 0 ,Freq 1 ,...,Freq n-1 ])
[0086] (2) Activate the module.
[0087] In order to utilize the information extracted from the compression operation, we will now use an excitation operation to obtain channel dependencies. To achieve this goal, this excitation operation must meet two conditions: flexibility (in particular, it must be able to learn nonlinear interactions between channels) and the ability to learn non-mutually exclusive relationships, which allows enhancing multiple channels.
[0088] In order to effectively limit the complexity of the model and promote generalization ability, the embodiment of the present application introduces a parameterized gating mechanism, which consists of a bottleneck structure consisting of two fully connected (FC) layers. This design aims to finely control the flow of information through nonlinear transformation and feature selection.
[0089] Specifically, the activation module can be processed according to the following process:
[0090] Step 1: Use the first fully connected layer to reduce the dimension of the channel attention vector output by the compression module, thereby reducing the size of the parameter space and thus reducing the complexity and computational burden of the model.
[0091] Step 2: Use the Relu activation function layer to process the compressed connection after dimensionality reduction output of the first fully connected layer to enhance the model's ability to learn complex feature representations.
[0092] Step 3: Use the second fully connected layer to restore the dimension of the channel attention vector processed by the ReLU activation function layer to the same channel dimension as the original input, thereby ensuring that the output of the gating mechanism can interact or combine with the original input element by element, thereby retaining the necessary feature information.
[0093] Step 4: Use the sigmoid activation function layer to convert the channel attention vector output by the second fully connected layer into a weight or importance score, where the value of the weight or importance score is between 0 and 1, which indicates the extent to which each channel feature map should be retained or suppressed in the original input.
[0094] Therefore, the expression for the weight vector can be written as:
[0095] s=σ(W2δ(W1Freq))
[0096] Among them, σ represents the Sigmoid function, δ represents the ReLU function, W1 and W2 represent the weight parameters of the first fully connected layer and the second fully connected layer respectively, and Freq represents the compressed channel attention vector.
[0097] Step 5: Multiply the weight vector output by the second fully connected layer and the feature map output by the previous downsampling layer channel by channel according to the following formula (i.e., multiply the weight by the feature map element of each channel) to obtain the corresponding feature map:
[0098]
[0099] In the formula F scale (u c ,s c ) represents the weight scalar s c and feature map u c ∈R H×W Channel-by-channel multiplication between s c =g(Freq; w g ), and w g Represents the weight parameters of the two fully connected layers.
[0100] The frequency-domain channel attention layer adjusts the activation state of each channel feature map. This means that channels deemed more important to the task (i.e., with relatively high weights) are enhanced, while channels deemed less important or potentially noisy (i.e., with relatively low weights) are suppressed. This ensures that the output feature map retains more high-frequency information relevant to salient object detection while reducing irrelevant or redundant information. This allows subsequent processing steps (such as upsampling) to focus more on key features, thereby improving the accuracy and robustness of salient object detection.
[0101] Optionally, for the deep feature extraction module in each level of the U-shaped network, the deep feature extraction module includes N-1 levels of upsampling layers, where:
[0102] The first NM-1 upsampling layers also use a dilated version of the residual U-shaped block, similar to the NM-1 downsampling layers. This design enables the NM-1 upsampling layer to expand the receptive field without increasing computational complexity, thereby capturing more contextual information. By introducing dilated convolution, the NM-1 upsampling layer can more effectively utilize the spatial information of the feature map, improving the ability to locate salient objects.
[0103] The remaining M-level upsampling layers are U-shaped blocks.
[0104] The deep feature extraction module takes the output feature map of the previous upsampling layer and the output feature map of the symmetrical downsampling layer as input in each upsampling layer to fuse information at different levels across stages. This design enables the model to simultaneously utilize low-level and high-level features to more comprehensively understand the input image and generate more accurate saliency maps.
[0105] Among them, the above M is less than N, M is greater than UM, and M and N are both positive integers.
[0106] For example, N can be 6 and M can be 4. Encoder_1, Encoder_2, Encoder_3, and Encoder_4 of the shallow feature extraction module use the normal version of the U-type block, Encoder_5 and Encoder_6 use the dilated version of the residual U-type block, and introduce the frequency domain attention layer in the first four downsampling layers. The deep feature extraction module Dncoder_5 uses the dilated version of the residual U-type block, while Dncoder_4, Dncoder_3, Dncoder_2, and Dncoder_1 all use the normal version of the U-type block.
[0107] In addition, the expanded version of the residual U-shaped block includes: at least one dilated convolution layer, a residual connection layer, a sampling layer, and an output layer. The dilated convolution layer is used to extract features from the feature map output by the previous sampling layer according to a preset dilation rate to obtain the corresponding feature map. The layers include a downsampling layer (i.e., encoder layer) and an upsampling layer (i.e., decoder layer).
[0108] Optionally, the saliency map generation module utilizes a multi-stage output mechanism to fully utilize feature information at different levels to improve the accuracy of saliency detection. Specifically, the saliency map generation module fuses the feature maps output by the Nth downsampling layer and the feature maps output by the N-1th upsampling layer to obtain the corresponding saliency map. The saliency map generation module includes at least: a convolutional layer, an upsampling layer, a feature map fusion layer, and a sigmoid activation function layer. During the fusion process, the functions of each layer are as follows:
[0109] Step 1: Use the convolution layer to process the feature map output by the Nth downsampling layer and the feature map output by the N-1th upsampling layer to obtain multiple feature maps. The convolution kernel size of the convolution layer can be 3*3;
[0110] Step 2: Convert each feature map output in the first step through the sigmoid activation function layer to obtain the probability map corresponding to each feature map;
[0111] Step 3: Use the upsampling layer to perform upsampling operations on each feature map. The upsampling process ensures that in the subsequent fusion step, the feature maps output by different layers have the same spatial dimensions.
[0112] Step 4: Use the feature map fusion layer to fuse these feature maps. The fused feature maps not only contain low-level detail information, but also incorporate high-level semantic information. This cross-stage fusion method enables the model to better understand the image content and accurately distinguish between salient objects and non-salient areas.
[0113] Step 5: Use the convolution layer and sigmoid activation function layer to process the fused feature map to obtain the final saliency map. The convolution kernel size of this convolution layer is 1*1.
[0114] Step S2: Obtain a training sample set and a sample label set.
[0115] The training sample set includes multiple second underwater images as training samples, and the sample label set includes a binary image corresponding to each second underwater image as a sample label. The mask value of each pixel in the binary image is used to indicate whether the pixel belongs to a salient target region within the second underwater image. Generally, the pixel value of each pixel in the binary image is either 0 or 1, where 1 indicates that the corresponding pixel is a salient target and 0 indicates that the corresponding pixel is a background pixel.
[0116] It should be noted that the embodiments of the present application can directly use the public data set as the training sample set and sample label set; or, manually label the target area of the second underwater image used as the training sample to generate a corresponding binary image.
[0117] Step S3: Iteratively train the initial learning model using the training sample set and the sample label set to obtain an underwater target detection model.
[0118] In the technical solution provided in the above step S3, the iterative training process can be implemented through the following steps: for each training batch in the iterative training process, each training sample of the training batch is input into the initial learning model to obtain the second saliency map corresponding to each training sample output by the initial learning model; the target loss function is constructed using the sample labels and the second saliency map corresponding to each training sample of the training batch, and the model parameters of the initial learning model are adjusted according to the target loss function.
[0119] Since the task of salient object detection aims to completely and clearly highlight the salient areas in an image, this often involves a clear demarcation of boundary regions. However, Binary Cross Entropy (BCE), as a pixel-level loss function, performs error calculation and supervision on each pixel separately, which can distinguish salient from non-salient areas to a certain extent. However, due to its characteristic of processing each pixel independently, it often ignores the spatial relationship between pixels, which will lead to blurred or chaotic boundaries.
[0120] To this end, an embodiment of the present application proposes using an edge-enhanced binary cross entropy loss function (EBCE) as a target loss function to accurately optimize the edge areas of salient objects.
[0121] Optionally, the target loss function is an edge-enhanced binary cross entropy loss function (EBCE). The specific implementation is as follows: by performing average pooling on the mask values of each pixel in the binary image, an edge feature map that highlights edge and texture information is obtained. The average pooling operation helps extract edge features in the image. Using this edge feature map, the weights of corresponding positions in the model can be adjusted so that the model pays more attention to edge information during training. This can guide the model to learn more detailed features about the edge, thereby improving edge detection performance. The average pooled mask values are then used as weights to calculate the binary cross entropy loss with the prediction results output by the model. This way, the model will pay more attention to learning edge and texture information during training, thereby achieving accurate segmentation of salient target edges.
[0122] Therefore, the expression of the above objective loss function can be written as:
[0123]
[0124] Where N represents the total number of training samples in the training batch, H and W represent the height and width of the training samples respectively, and g n,h,w Represents the mask value of the binary image corresponding to the nth training sample at position (h, w), s n,h,w represents the predicted pixel value at position (h, w) of the second saliency map corresponding to the nth training sample, and s, g∈{0,1}, e n,h,w It represents the value of the edge feature map at position (h, w) obtained by performing an average pooling operation on the binary image corresponding to the nth training sample using a sliding window of a preset size (such as 5*5), and e n,h,w The expression can be written as:
[0125]
[0126] Among them, G n ={g n |0 <g n <1}∈R N×1×H×W Represents the binary image corresponding to the nth training sample, S(G n )={s n |0 n <1}∈R N×1×H×W represents the second saliency map corresponding to the nth training sample.
[0127] Therefore, by calling the above-mentioned underwater target detection model to analyze the first underwater image, the first salient target and the background in the image are effectively detected and segmented to obtain a first salient map, and the first salient target in the first underwater image is accurately identified based on the first salient map.
[0128] In summary, the underwater target detection method in the embodiments of the present application has the following technical advantages over traditional target detection methods:
[0129] (1) The underwater target detection model composed of a multi-level nested U-shaped network and a saliency map generation module in the present embodiment can effectively capture feature information at different scales and levels. Specifically, the nested U-shaped network improves the model's feature fusion capability and computational efficiency by combining skip connections and residual U-shaped blocks, avoiding the problem of a sharp increase in computational cost and memory consumption caused by an increase in the number of layers.
[0130] (2) The embodiment of the present application introduces a frequency domain channel attention module into the shallow feature extraction module and uses discrete cosine transform (DCT) instead of traditional global average pooling (GAP), which can not only capture global contrast information, but also effectively enhance the high-frequency information of local edges, thereby solving the problem of blurring and loss of high-frequency information of underwater images and improving the adaptability and detection accuracy of the model to complex underwater environments.
[0131] (3) The embodiment of the present application adopts a multi-stage output mechanism in the saliency map generation module, that is, multiple intermediate saliency maps are generated from the deepest layer of the encoder to each layer of the decoder, and by upsampling and fusing these saliency maps of different resolutions, the model can utilize information at different levels to generate more comprehensive and accurate detection results.
[0132] (4) The embodiment of the present application adopts a weighted cross entropy loss function based on edge enhancement, which, while retaining pixel-level supervision, particularly strengthens the loss calculation of boundary areas, improves the performance of the model in processing edge details, and thus helps to more accurately define the contours of significant targets, especially in the case of blurred boundaries in underwater images.
[0133] Example 2
[0134] According to an embodiment of the present application, an underwater target detection device for implementing the underwater target detection method in embodiment 1 is also provided. Figure 2 As shown, the underwater target detection device at least includes: an acquisition module 22 and a detection module 24, wherein:
[0135] An acquisition module 22 is configured to acquire a first underwater image;
[0136] The detection module 24 is used to perform target detection on the first underwater image using a pre-trained underwater target detection model, obtain a first saliency map corresponding to the first underwater image, and determine a first salient target in the first underwater image based on the first saliency map, wherein the predicted pixel value of each pixel in the first saliency map is used to represent the probability that the pixel belongs to the first salient target. The underwater target detection model includes: a multi-level nested U-shaped network and a saliency map generation module, and each level of the U-shaped network includes: a shallow feature extraction module and a deep feature extraction module, and the shallow feature extraction module includes at least: a frequency domain channel attention layer.
[0137] In addition, the underwater target detection state provided in the embodiment of the present application further includes a model training module 26, and the model training module 26 is used to train the underwater target detection model according to the following method, including:
[0138] Step S1: construct an initial learning model including a multi-level nested U-shaped network and a saliency map generation module.
[0139] Among them, the initial learning model includes a multi-level nested U-shaped network and a saliency map generation module, and each level of the U-shaped network includes: a shallow feature extraction module and a deep feature extraction module, and the shallow feature extraction module and the deep feature extraction module include at least: a frequency domain channel attention layer.
[0140] Optionally, for the shallow feature extraction module in each level of the U-shaped network, the shallow feature extraction module includes N levels of downsampling layers, where:
[0141] The first M downsampling layers are U-shaped blocks, and all of them include a frequency domain channel attention layer. This design allows the underwater object detection model to focus on more important channel information in the frequency domain while performing convolution operations, thereby enhancing the expressiveness of features. It should be noted that the size of the U-shaped blocks used in the first M downsampling layers can be configured according to the spatial resolution of the input feature map to ensure that detailed information can be effectively captured at different scales. Therefore, for larger input feature maps, larger convolution kernels and frequency domain channel attention layers can be used to capture more details.
[0142] The last NM-level downsampling layers are dilated versions of the residual U-block. This is because the resolution of the feature map is already low in the following downsampling layers. If the sampling rate is further reduced, details that are crucial to the background information may be lost. The dilated version of the residual U-block can replace the traditional pooling and upsampling operations with dilated convolutions, thereby maintaining the resolution of the feature map unchanged.
[0143] The shallow feature extraction module in the above structure not only retains more spatial information when extracting shallow features from underwater images, but also expands the model's receptive field, enabling it to capture a wider range of contextual information. By avoiding pooling and upsampling operations, information loss is reduced.
[0144] The frequency domain channel attention layer is introduced in the first M-level downsampling process to deal with the loss of underwater image information caused by color distortion and its interference with the underwater salient object detection task. The frequency domain channel attention layer includes: a compression module and an activation module, where:
[0145] (1)Compression module.
[0146] In the channel attention mechanism, feature maps are compressed according to the channel dimension, converting each two-dimensional feature channel into a real number with a global receptive field, and the output dimension matches the number of input feature channels. This compression method can effectively capture information from the entire feature space and provide richer context for the network.
[0147] At present, the traditional channel attention mechanism mainly relies on global average pooling (GAP) to capture the global information of each channel, wherein GAP obtains a real number that can reflect the importance of the channel by averaging the feature maps of each channel. However, although the global average pooling (GAP) method can capture the global information of each channel, it may ignore the differences and correlations between different frequency components. In contrast, the discrete cosine transform (DCT) can better retain the features of more frequencies and extract more robust feature representations. To this end, the embodiment of the present application proposes the use of discrete cosine transform (DCT) to replace global average pooling (GAP) to achieve feature channel compression in the frequency domain.
[0148] Specifically, the compression module can be processed according to the following process:
[0149] Step 1: Split the feature map output by the previous downsampling layer along the channel dimension to obtain multiple two-dimensional matrices, where each matrix represents a specific feature channel in the feature map. Therefore, the multiple two-dimensional matrices obtained by segmentation can be expressed as: [X0, X1, ..., X n-1 ], and X i ∈R 1×H×W ,i∈{0,1,...,n-1}.
[0150] Step 2: Perform discrete cosine transform on each two-dimensional matrix to obtain the corresponding two-dimensional frequency components. i The corresponding two-dimensional frequency components can be expressed as:
[0151]
[0152] in, represents the compressed feature vector, C i Indicates the number of low-frequency components retained by the i-th two-dimensional matrix after DCT transformation, [u i ,v i ] is the same as the i-th two-dimensional matrix X i The index of the corresponding two-dimensional frequency component, Represents the i-th two-dimensional matrix X i The pixel value at position (h,w), Denotes the DCT-II transform at position (h,w) corresponding to the frequency index [u i ,v i ], and the basis function of DCT-II is:
[0153]
[0154] Step 3: Concatenate the two-dimensional frequency components corresponding to each two-dimensional matrix to obtain the corresponding compressed channel attention vector. Therefore, the expression of the compressed channel attention vector can be written as:
[0155] Freq=cat([Freq 0 ,Freq 1 ,...,Freq n-1 ])
[0156] (2) Activate the module.
[0157] In order to utilize the information extracted from the compression operation, we will now use an excitation operation to obtain channel dependencies. To achieve this goal, this excitation operation must meet two conditions: flexibility (in particular, it must be able to learn nonlinear interactions between channels) and the ability to learn non-mutually exclusive relationships, which allows enhancing multiple channels.
[0158] In order to effectively limit the complexity of the model and promote generalization ability, the embodiment of the present application introduces a parameterized gating mechanism, which consists of a bottleneck structure consisting of two fully connected (FC) layers. This design aims to finely control the flow of information through nonlinear transformation and feature selection.
[0159] Specifically, the activation module can be processed according to the following process:
[0160] Step 1: Use the first fully connected layer to reduce the dimension of the channel attention vector output by the compression module, thereby reducing the size of the parameter space and thus reducing the complexity and computational burden of the model.
[0161] Step 2: Use the Relu activation function layer to process the compressed connection after dimensionality reduction output of the first fully connected layer to enhance the model's ability to learn complex feature representations.
[0162] Step 3: Use the second fully connected layer to restore the dimension of the channel attention vector processed by the ReLU activation function layer to the same channel dimension as the original input, thereby ensuring that the output of the gating mechanism can interact or combine with the original input element by element, thereby retaining the necessary feature information.
[0163] Step 4: Use the sigmoid activation function layer to convert the channel attention vector output by the second fully connected layer into a weight or importance score, where the value of the weight or importance score is between 0 and 1, which indicates the extent to which each channel feature map should be retained or suppressed in the original input.
[0164] Therefore, the expression for the weight vector can be written as:
[0165] s=σ(W2δ(W1Freq))
[0166] Among them, σ represents the Sigmoid function, δ represents the ReLU function, W1 and W2 represent the weight parameters of the first fully connected layer and the second fully connected layer respectively, and Freq represents the compressed channel attention vector.
[0167] Step 5: Multiply the weight vector output by the second fully connected layer and the feature map output by the previous downsampling layer channel by channel according to the following formula (i.e., multiply the weight by the feature map element of each channel) to obtain the corresponding feature map:
[0168]
[0169] In the formula F scale (u c ,s c ) represents the weight scalar s c and feature map u c ∈R H×W Channel-by-channel multiplication between s c =g(Freq; w g ), and w g Represents the weight parameters of the two fully connected layers.
[0170] The frequency-domain channel attention layer adjusts the activation state of each channel feature map. This means that channels deemed more important to the task (i.e., with relatively high weights) are enhanced, while channels deemed less important or potentially noisy (i.e., with relatively low weights) are suppressed. This ensures that the output feature map retains more high-frequency information relevant to salient object detection while reducing irrelevant or redundant information. This allows subsequent processing steps (such as upsampling) to focus more on key features, thereby improving the accuracy and robustness of salient object detection.
[0171] Optionally, for the deep feature extraction module in each level of the U-shaped network, the deep feature extraction module includes N-1 levels of upsampling layers, where:
[0172] The first NM-1 upsampling layers also use a dilated version of the residual U-shaped block, similar to the NM-1 downsampling layers. This design enables the NM-1 upsampling layer to expand the receptive field without increasing computational complexity, thereby capturing more contextual information. By introducing dilated convolution, the NM-1 upsampling layer can more effectively utilize the spatial information of the feature map, improving the ability to locate salient objects.
[0173] The remaining M-level upsampling layers are U-shaped blocks.
[0174] The deep feature extraction module takes the output feature map of the previous upsampling layer and the output feature map of the symmetrical downsampling layer as input in each upsampling layer to fuse information at different levels across stages. This design enables the model to simultaneously utilize low-level and high-level features to more comprehensively understand the input image and generate more accurate saliency maps.
[0175] Among them, the above M is less than N, M is greater than UM, and M and N are both positive integers.
[0176] For example, N can be 6 and M can be 4. Encoder_1, Encoder_2, Encoder_3, and Encoder_4 of the shallow feature extraction module use the normal version of the U-type block, Encoder_5 and Encoder_6 use the dilated version of the residual U-type block, and introduce the frequency domain attention layer in the first four downsampling layers. The deep feature extraction module Dncoder_5 uses the dilated version of the residual U-type block, while Dncoder_4, Dncoder_3, Dncoder_2, and Dncoder_1 all use the normal version of the U-type block.
[0177] In addition, the expanded version of the residual U-shaped block includes: at least one dilated convolution layer, a residual connection layer, a sampling layer, and an output layer. The dilated convolution layer is used to extract features from the feature map output by the previous sampling layer according to a preset dilation rate to obtain the corresponding feature map. The layers include a downsampling layer (i.e., encoder layer) and an upsampling layer (i.e., decoder layer).
[0178] Optionally, the saliency map generation module utilizes a multi-stage output mechanism to fully utilize feature information at different levels to improve the accuracy of saliency detection. Specifically, the saliency map generation module fuses the feature maps output by the Nth downsampling layer and the feature maps output by the N-1th upsampling layer to obtain the corresponding saliency map. The saliency map generation module includes at least: a convolutional layer, an upsampling layer, a feature map fusion layer, and a sigmoid activation function layer. During the fusion process, the functions of each layer are as follows:
[0179] Step 1: Use the convolution layer to process the feature map output by the Nth downsampling layer and the feature map output by the N-1th upsampling layer to obtain multiple feature maps. The convolution kernel size of the convolution layer can be 3*3;
[0180] Step 2: Convert each feature map output in the first step through the sigmoid activation function layer to obtain the probability map corresponding to each feature map;
[0181] Step 3: Use the upsampling layer to perform upsampling operations on each feature map. The upsampling process ensures that in the subsequent fusion step, the feature maps output by different layers have the same spatial dimensions.
[0182] Step 4: Use the feature map fusion layer to fuse these feature maps. The fused feature maps not only contain low-level detail information, but also incorporate high-level semantic information. This cross-stage fusion method enables the model to better understand the image content and accurately distinguish between salient objects and non-salient areas.
[0183] Step 5: Use the convolution layer and sigmoid activation function layer to process the fused feature map to obtain the final saliency map. The convolution kernel size of this convolution layer is 1*1.
[0184] Step S2: Obtain a training sample set and a sample label set.
[0185] The training sample set includes multiple second underwater images as training samples, and the sample label set includes a binary image corresponding to each second underwater image as a sample label. The mask value of each pixel in the binary image is used to indicate whether the pixel belongs to a salient target region within the second underwater image. Generally, the pixel value of each pixel in the binary image is either 0 or 1, where 1 indicates that the corresponding pixel is a salient target and 0 indicates that the corresponding pixel is a background pixel.
[0186] It should be noted that the embodiments of the present application can directly use the public data set as the training sample set and sample label set; or, manually label the target area of the second underwater image used as the training sample to generate a corresponding binary image.
[0187] Step S3: Iteratively train the initial learning model using the training sample set and the sample label set to obtain an underwater target detection model.
[0188] In the technical solution provided in the above step S3, the iterative training process can be implemented through the following steps: for each training batch in the iterative training process, each training sample of the training batch is input into the initial learning model to obtain the second saliency map corresponding to each training sample output by the initial learning model; the target loss function is constructed using the sample labels and the second saliency map corresponding to each training sample of the training batch, and the model parameters of the initial learning model are adjusted according to the target loss function.
[0189] Since the task of salient object detection aims to completely and clearly highlight the salient areas in an image, this often involves a clear demarcation of boundary regions. However, Binary Cross Entropy (BCE), as a pixel-level loss function, performs error calculation and supervision on each pixel separately, which can distinguish salient from non-salient areas to a certain extent. However, due to its characteristic of processing each pixel independently, it often ignores the spatial relationship between pixels, which will lead to blurred or chaotic boundaries.
[0190] To this end, an embodiment of the present application proposes using an edge-enhanced binary cross entropy loss function (EBCE) as a target loss function to accurately optimize the edge areas of salient objects.
[0191] Optionally, the target loss function is an edge-enhanced binary cross entropy loss function (EBCE). The specific implementation is as follows: by performing average pooling on the mask values of each pixel in the binary image, an edge feature map that highlights edge and texture information is obtained. The average pooling operation helps extract edge features in the image. Using this edge feature map, the weights of corresponding positions in the model can be adjusted so that the model pays more attention to edge information during training. This can guide the model to learn more detailed features about the edge, thereby improving edge detection performance. The average pooled mask values are then used as weights to calculate the binary cross entropy loss with the prediction results output by the model. This way, the model will pay more attention to learning edge and texture information during training, thereby achieving accurate segmentation of salient target edges.
[0192] Therefore, the expression of the above objective loss function can be written as:
[0193]
[0194] Where N represents the total number of training samples in the training batch, H and W represent the height and width of the training samples respectively, and g n,h,w Represents the mask value of the binary image corresponding to the nth training sample at position (h, w), s n,h,w represents the predicted pixel value at position (h, w) of the second saliency map corresponding to the nth training sample, and s, g∈{0,1}, e n,h,w It represents the value of the edge feature map at position (h, w) obtained by performing an average pooling operation on the binary image corresponding to the nth training sample using a sliding window of a preset size (such as 5*5), and e n,h,w The expression can be written as:
[0195]
[0196] Among them, G n ={g n |0 <g n <1}∈R N×1×H×W Represents the binary image corresponding to the nth training sample, S(G n )={s n |0 n <1}∈R N×1×H×W represents the second saliency map corresponding to the nth training sample.
[0197] Furthermore, the detection module 24 can call the underwater target detection model trained by the above-mentioned model training module 26 to analyze the first underwater image, so as to effectively detect and segment the first salient target and background in the image, obtain a first salient map, and accurately identify the first salient target in the first underwater image based on the first salient map.
[0198] It should be noted that each module in the underwater target detection device in the embodiment of the present application corresponds one-to-one to each implementation step of the underwater target detection method in Example 1. Since a detailed description has been given in Example 1, some details not reflected in this embodiment can be referred to Example 1 and will not be elaborated here.
[0199] Example 3
[0200] According to an embodiment of the present application, a computer program product is further provided, which includes a computer program, wherein when the computer program is executed by a processor, the underwater target detection method in Example 1 is implemented.
[0201] According to an embodiment of the present application, a non-volatile storage medium is also provided, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the underwater target detection method in Example 1 by running the computer program.
[0202] According to an embodiment of the present application, a processor is further provided, which is used to run a computer program, wherein the underwater target detection method in Example 1 is executed when the computer program is run.
[0203] According to an embodiment of the present application, an electronic device is also provided, which includes: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the underwater target detection method in Example 1 through the computer program.
[0204] Specifically, when the computer program is running, the following steps are executed: obtaining a first underwater image; performing target detection on the first underwater image using a pre-trained underwater target detection model to obtain a first saliency map corresponding to the first underwater image, and determining a first saliency map in the first underwater image based on the first saliency map, wherein the predicted pixel value of each pixel in the first saliency map is used to represent the probability that the pixel belongs to the first saliency map, and the underwater target detection model includes: a multi-level nested U-shaped network and a saliency map generation module, and each level of the U-shaped network includes: a shallow feature extraction module and a deep feature extraction module, and the shallow feature extraction module includes at least: a frequency domain channel attention layer.
[0205] As an optional implementation, the electronic device may be in the form of a mobile terminal, a computer terminal or a similar computing device. Figure 3 FIG1 shows a hardware structure block diagram of an electronic device for implementing an underwater target detection method. Figure 3 As shown, the electronic device 30 may include one or more (illustrated as 302a, 302b, ..., 302n in the figure) processors 302 (the processor 302 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 304 for storing data, and a transmission device 306 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 3 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 3 More or fewer components than shown, or with Figure 3 Different configurations shown.
[0206] It should be noted that the one or more processors 302 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the electronic device 30. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0207] Memory 304 can be used to store software programs and modules for application software, such as the program instructions / data storage device corresponding to the underwater target detection method in the embodiments of the present application. Processor 302 executes the software programs and modules stored in memory 304 to perform various functional applications and data processing, thereby implementing the aforementioned application vulnerability detection method. Memory 304 can include high-speed random access memory (RAM) and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, memory 304 may further include memory remotely located from processor 302, which can be connected to electronic device 30 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0208] The transmission device 306 is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by the communications provider of the electronic device 30. In one embodiment, the transmission device 306 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In one embodiment, the transmission device 306 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0209] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the electronic device 30 .
[0210] The serial numbers of the above embodiments are for description only and do not represent the advantages or disadvantages of the embodiments.
[0211] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0212] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0213] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs.
[0214] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0215] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program code.
[0216] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for underwater target detection, characterized in that: include: acquiring a first underwater image; Target detection is performed on the first underwater image using a pre-trained underwater target detection model to obtain a first saliency map corresponding to the first underwater image, and a first salient target in the first underwater image is determined based on the first saliency map, wherein the predicted pixel value of each pixel in the first saliency map is used to represent the probability that the pixel belongs to the first salient target. The underwater target detection model includes: a multi-level nested U-shaped network and a saliency map generation module, and each level of the U-shaped network includes: a shallow feature extraction module and a deep feature extraction module, and the shallow feature extraction module includes at least: a frequency domain channel attention layer.
2. The method according to claim 1, characterized in that The training process of the underwater target detection model includes: Constructing an initial learning model including the shallow feature extraction module, the deep feature extraction module and the saliency map generation module; Obtaining a training sample set and a sample label set, wherein the training sample set includes a plurality of second underwater images as training samples, the sample label set includes a binarized image corresponding to each second underwater image as a sample label, and a mask value of each pixel in the binarized image is used to indicate whether the pixel belongs to a salient target area in the second underwater image; The initial learning model is iteratively trained using the training sample set and the sample label set to obtain the underwater target detection model.
3. The method according to claim 2, characterized in that The shallow feature extraction module in each level of the U-shaped network includes N levels of downsampling layers, and the first M levels of downsampling layers are U-shaped blocks, and the last NM levels of downsampling layers are expanded versions of residual U-shaped blocks. The first M levels of downsampling layers all include the frequency domain channel attention layer. The shallow feature extraction module is used to extract shallow features of the underwater image to obtain a corresponding shallow feature map, wherein M is less than N, M is greater than UM, and M and N are both positive integers; The deep feature extraction module in each level of the U-shaped network includes N-1 levels of upsampling layers, and the first NM-1 levels of upsampling layers are expanded versions of residual U-shaped blocks, and the remaining M levels of upsampling layers are U-shaped blocks. The deep feature extraction module is used to fuse the feature map output by each level of upsampling layer with the feature map output by the corresponding level of downsampling layer through skip connections to obtain a deep feature map; The saliency map generation module includes at least: a convolution layer, an upsampling layer, a feature map fusion layer, and a sigmoid activation function layer. The saliency map generation module is used to perform feature fusion on the feature map output by the Nth level downsampling layer and the feature map output by the N-1th level upsampling layer to obtain the corresponding saliency map.
4. The method according to claim 3, characterized in that The expanded version of the residual U-shaped block includes: at least one expanded convolution layer, a residual connection layer, a sampling layer, and an output layer, wherein the expanded convolution layer is used to extract features from the feature map output by the previous sampling layer according to a preset expansion rate to obtain a corresponding feature map.
5. The method according to claim 3, characterized in that The frequency domain channel attention layer includes: a compression module and an activation module, wherein, The compression module is used to segment the feature map output by the previous downsampling layer along the channel dimension to obtain multiple two-dimensional matrices, and perform discrete cosine transform on each of the two-dimensional matrices to obtain corresponding two-dimensional frequency components, and then splice the two-dimensional frequency components corresponding to each of the two-dimensional matrices to obtain the corresponding compressed channel attention vector; The activation module includes a first fully connected layer, a Relu activation function layer, a second fully connected layer and a sigmoid activation function layer connected in sequence, wherein the first fully connected layer is used to perform dimensionality reduction processing on the channel attention vector output by the compression module; the second fully connected layer is used to perform dimensionality recovery processing on the channel attention vector processed by the Relu activation function layer; the output of the activation module is a feature map obtained by multiplying the channel attention vector output by the second fully connected layer and the feature map output by the previous downsampling layer channel by channel.
6. The method according to claim 2, characterized in that Iteratively training the initial learning model using the training sample set and the sample label set to obtain the underwater target detection model includes: For each training batch in the iterative training process, each training sample of the training batch is input into the initial learning model to obtain a second saliency map corresponding to each training sample output by the initial learning model; a target loss function is constructed using the sample labels corresponding to each training sample of the training batch and the second saliency map, and the model parameters of the initial learning model are adjusted according to the target loss function.
7. The method according to claim 6, characterized in that The target loss function is a weighted cross entropy loss function based on edge enhancement, and the expression of the target loss function is: Where N represents the total number of training samples in the training batch, H and W represent the height and width of the training samples respectively, and g n,h,w Represents the mask value of the binary image corresponding to the nth training sample at position (h, w), s n,h,w represents the predicted pixel value at position (h, w) of the second saliency map corresponding to the nth training sample, and s, g∈{0,1}, e n,h,w represents the value of the edge feature map at position (h, w) obtained by performing an average pooling operation on the binary image corresponding to the nth training sample using a sliding window of a preset size, and e n,h,w The expression can be written as: Among them, G n ={g n |0 <g n <1}∈R N×1×H×W Represents the binary image corresponding to the nth training sample, S(G n )={s n |0 n <1}∈R N×1×H×W represents the second saliency map corresponding to the nth training sample. 8. An underwater target detection device, characterized in that: include: An acquisition module, configured to acquire a first underwater image; A detection module is configured to perform target detection on the first underwater image using a pre-trained underwater target detection model, obtain a first saliency map corresponding to the first underwater image, and determine a first salient target within the first underwater image based on the first saliency map, wherein the predicted pixel value of each pixel in the first saliency map is used to represent the probability that the pixel belongs to the first salient target. The underwater target detection model includes: a multi-level nested U-shaped network and a saliency map generation module, and each level of the U-shaped network includes: a shallow feature extraction module and a deep feature extraction module, and the shallow feature extraction module includes at least: a frequency domain channel attention layer.
9. A computer program product, characterized in that include: A computer program, wherein when the computer program is executed by a processor, the underwater target detection method according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the underwater target detection method according to any one of claims 1 to 7 through the computer program.
Citation Information
Cited By
Underwater target detection method and device based on frequency domain guide feature enhancement
CN122156946A
An underwater target detection method and device based on frequency domain guided feature enhancement
CN122156946B