Underwater image salient target detection method and system based on conditional diffusion model

Through frequency-domain-spatial domain entanglement enhancement and stable time-step mask prediction technology, the problems of overconfident misprediction and boundary offset in traditional underwater salient target detection are solved, and high-precision and robust underwater salient target detection is achieved.

CN120823375APending Publication Date: 2025-10-21HAINAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511068049.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Traditional underwater salient target detection methods are prone to overconfident mispredictions and boundary shift problems in complex underwater environments. In addition, the intermediate iteration results of the conditional diffusion model are easily disturbed by noise and lack a stable iterative prediction mechanism.

Method used

Frequency domain-spatial domain entanglement enhancement and stable time-step mask prediction technology are adopted. The complementary features of RGB image and depth image are used to generate robust conditional features through Fourier domain-spatial domain entanglement enhancement. The stability of the iterative results is screened and weighted fused through the stable time-step mask prediction module to improve the accuracy and robustness of edge generation.

Benefits of technology

It effectively solves the problems of overconfident misprediction and boundary offset in underwater environments, improves the accuracy and boundary stability of underwater salient target detection, and generates high-precision and high-robustness underwater salient target segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823375A_ABST
    Figure CN120823375A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of salient target detection, and provides an underwater image salient target detection method and system based on a conditional diffusion model, and the method comprises the steps: carrying out the feature extraction of an RGB image and a depth image in each time step, and obtaining a plurality of RGB features and depth features with different resolutions; fourier domain perception enhancement is carried out on the RGB features and the depth features under the same resolution, and global Fourier optimization features are obtained; generating splicing features based on the RGB features and the depth features under the same resolution, and performing spatial domain perception enhancement on the splicing features to obtain spatial optimization features; fusing the global Fourier optimization features and the spatial optimization features to obtain fusion results, and fusing the fusion results under different resolutions to obtain condition features of each time step; and predicting the condition features of different time steps and the loading reference image to obtain prediction results of different time steps, and screening and aggregating based on the prediction results of different time steps to obtain an aggregation detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of salient target detection, and in particular relates to a method and system for detecting salient targets in underwater images based on a conditional diffusion model. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] With the rapid development of underwater robotics, marine resource exploration, and ecological monitoring, underwater salient object detection (USOD) has become a fundamental component of underwater visual perception. However, the underwater environment suffers from complex degradation phenomena such as light absorption, scattering, color distortion, and suspended particle noise, resulting in widespread issues such as low contrast, blurred edges, and texture loss in underwater images. Traditional USOD methods, based on an encoder-decoder architecture, utilize only spatial RGB-D features to directly predict salient masks from pixel-level probability maps. This is prone to overconfident mispredictions and boundary shifts due to underwater degradation. Furthermore, the Conditional Diffusion Model (CDM), while demonstrating advantages in image generation tasks such as progressive denoising and multi-step iteration, has not yet been introduced to the USOD field. Furthermore, its intermediate results are easily affected by the quality of the conditional-guided features, making it difficult to achieve its intended effect.

[0004] In other words, traditional RGB-D salient object detection methods rely solely on spatial features, leading to overconfident mispredictions and boundary shifts in complex underwater backgrounds. When directly applying diffusion models to USOD, poor quality conditional features can amplify intermediate errors. Diffusion models' intermediate iterations are susceptible to noise perturbations and lack a stable iterative prediction mechanism. Summary of the Invention

[0005] In order to solve the above problems, the present invention proposes a method and system for underwater image salient target detection based on the conditional diffusion model. The present invention comprises two main parts: Fourier domain-spatial domain entanglement enhancement and stable time-step mask prediction. The frequency domain-spatial domain entanglement enhancement generates robust conditional features by utilizing the dominant information of multimodal features in the spatial domain and the Fourier domain, which are used to guide the diffusion model to correctly identify salient targets in underwater images; then, the stable time-step mask prediction effectively guides the edge generation process by integrating a stable iterative process, so as to achieve more refined underwater salient target boundary generation.

[0006] According to some embodiments, a first solution of the present invention provides a method for detecting salient objects in underwater images based on a conditional diffusion model, which adopts the following technical solutions: The method for detecting salient objects in underwater images based on the conditional diffusion model includes: Preprocess the original underwater image to obtain RGB image, depth image and loading reference image; Perform feature extraction on the RGB image and depth image in each time step to obtain multiple RGB features and depth features with different resolutions; Perform Fourier domain perceptual enhancement on RGB features and depth features at the same resolution to obtain global Fourier optimized features; Generate splicing features based on RGB features and depth features at the same resolution, perform spatial domain perception enhancement on the splicing features, and obtain spatial optimization features; The global Fourier optimization feature and the spatial optimization feature are subtracted and then spatial attention operation is performed to obtain the salient feature. The fusion result is obtained based on the salient feature, global Fourier optimization and spatial optimization features. The fusion results at different resolutions are spliced ​​to obtain the conditional feature of each time step. The diffusion model is used to predict the conditional features and loading reference graphs at different time steps to obtain the prediction results at different time steps. The prediction results at different time steps are filtered and aggregated to obtain the aggregated detection results.

[0007] Furthermore, the RGB features and depth features at each resolution are enhanced in the Fourier domain to obtain global Fourier optimized features, specifically: Extract the phase and amplitude components corresponding to the RGB features and depth features at each resolution; Generate a spliced ​​phase component and a spliced ​​amplitude component based on the phase component and the amplitude component corresponding to the RGB feature and the depth feature; The spliced ​​phase component and the spliced ​​amplitude component are convolved and then inverse Fourier transformed to obtain the global Fourier optimization feature.

[0008] Furthermore, the phase and amplitude components of the RGB features and depth features at each resolution are extracted, specifically: Convolution processing is performed on the RGB features and depth features at each resolution to obtain RGB convolution features and depth convolution features; Perform Fourier transform on the RGB convolution features to obtain RGB phase components and RGB amplitude components; Perform Fourier transform on the deep convolution features to obtain the deep phase component and the deep amplitude component.

[0009] Furthermore, the splicing features are generated based on the RGB features and depth features at each resolution, and the splicing features are enhanced by spatial perception to obtain spatial optimization features, specifically: Extract the RGB features and depth features at each resolution and then splice the resulting features; The spliced ​​features are processed by three multi-scale convolutions simultaneously to obtain three core components of global self-attention, and the self-attention map is obtained after the global self-attention operation; Perform multi-scale convolution processing on the splicing features to obtain convolution splicing features; The self-attention map and the convolutional splicing features are concatenated and then convolved to obtain spatially optimized features.

[0010] Furthermore, the spatial attention operation is performed after subtracting the global Fourier optimization feature and the spatial optimization feature to obtain the salient feature. Based on the salient feature, the global Fourier optimization and the spatial optimization feature, the fusion result is obtained, specifically: The spatial attention operation is performed after subtracting the global Fourier optimized features from the spatial optimized features to obtain the salient features; Multiply the salient feature with the global Fourier optimization feature and then add it to the spatial optimization feature to obtain the fused optimization feature; After splicing the fusion optimization features and the spatial optimization features, a channel attention operation is performed to obtain the channel fusion optimization features; Add the channel fusion optimization features to the RGB features to obtain the fusion result.

[0011] Furthermore, the prediction results at different time steps are filtered and aggregated to obtain the aggregated detection results, specifically: The predicted change rate of the current time step is obtained by subtracting the predicted result of the current time step from the predicted result of the previous time step; All prediction change rates within the total iteration time step T are sorted from small to large, and T / 2 prediction results with low prediction change rates are selected for weighted average aggregation to obtain the aggregated detection result.

[0012] According to some embodiments, a second solution of the present invention provides a system for detecting salient objects in underwater images based on a conditional diffusion model, which employs the following technical solutions: The underwater image salient object detection system based on the conditional diffusion model includes: The image processing module is used to pre-process the original underwater image to obtain the RGB image, depth map and loaded reference image; The feature extraction module is used to extract features from the RGB image and depth image in each time step, and obtain multiple RGB features and depth features of different resolutions respectively; The Fourier domain perception enhancement module is used to perform Fourier domain perception enhancement on the RGB features and depth features at each resolution to obtain global Fourier optimized features; The spatial perception enhancement module is used to generate splicing features based on the RGB features and depth features at each resolution, and perform spatial perception enhancement on the splicing features to obtain spatially optimized features; The conditional feature generation module is used to perform spatial attention operations after subtracting the global Fourier optimization features from the spatial optimization features to obtain salient features. Based on the salient features, global Fourier optimization, and spatial optimization features, a fusion result is obtained. The fusion results at different resolutions are spliced ​​to obtain the conditional features of each time step. The prediction result fusion module is used to use the diffusion model to predict the conditional features and loading reference graphs at different time steps to obtain the prediction results at different time steps, and to filter and aggregate the prediction results at different time steps to obtain the aggregated detection results.

[0013] According to some embodiments, a third aspect of the present invention provides a computer-readable storage medium.

[0014] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for detecting salient targets in underwater images based on a conditional diffusion model as described in the first solution.

[0015] According to some embodiments, a fourth aspect of the present invention provides a computer device.

[0016] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the method for detecting salient targets in underwater images based on a conditional diffusion model as described in the first solution above are implemented.

[0017] According to some embodiments, a fifth aspect of the present invention provides a computer program product or computer program.

[0018] A computer program product or computer program includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in the underwater image salient target detection method based on the conditional diffusion model as described in the first solution above.

[0019] Compared with the prior art, the present invention has the following beneficial effects: The method of the present invention addresses two major issues in underwater environments: overconfident mispredictions and boundary shifts, through two main processes: frequency-domain-spatial domain entanglement enhancement and stable time-step mask prediction. On the one hand, Fourier-domain-spatial domain entanglement enhancement stems from the discovery of the Fourier spectrum of underwater scenes: RGB images and depth maps exhibit significant complementary characteristics in the Fourier domain. Specifically, the Fourier transform converts the image into the frequency domain and extracts amplitude and phase information. The RGB amplitude primarily represents the distribution of light and dark in the scene, helping to distinguish different objects from the background. The depth map phase focuses on structural and contour information, which helps distinguish the shape and boundaries of objects. By complementing these Fourier domain fusion features with the spatial domain, the model can more comprehensively understand underwater scene information, improve the quality of conditional feature generation, and address overconfident mispredictions. On the other hand, stable time-step mask prediction is based on mutual information theory, which states that each iteration step is related to the previous iteration step, and the ultimate goal of each iteration step is to obtain a correct result. The stable time-step mask prediction module calculates the IoU change rate between masks in adjacent time steps, automatically locks the five most stable iteration steps, and then performs weighted fusion on these stable masks to filter out early noise disturbances and enhance boundary consistency, ultimately outputting high-precision and robust underwater salient object segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0021] Figure 1 1 is a diagram showing the overall processing framework of the method for detecting salient objects in underwater images based on the conditional diffusion model in an embodiment of the present invention; Figure 2 1 is a processing framework diagram of Fourier domain-spatial domain entanglement enhancement in an embodiment of the present invention; Figure 3 1 is a processing framework diagram of Fourier domain perception enhancement and spatial domain perception enhancement in an embodiment of the present invention; Figure 4 This is a comparison chart of the effects of the method proposed in the embodiment of the present invention and other existing methods on different data sets. DETAILED DESCRIPTION

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0023] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0024] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0025] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.

[0026] Example 1 like Figure 1 As shown, this embodiment provides a method for detecting salient targets in underwater images based on a conditional diffusion model. This embodiment uses the method applied to a server as an example. It is understandable that the method can also be applied to a terminal, and can also be applied to a system including a terminal, a server, and a server, and is implemented through the interaction between the terminal and the server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application. In this embodiment, the method includes the following steps: Preprocess the original underwater image to obtain RGB image, depth image and loading reference image; Perform feature extraction on the RGB image and depth image in each time step to obtain multiple RGB features and depth features with different resolutions; Perform Fourier domain perceptual enhancement on RGB features and depth features at the same resolution to obtain global Fourier optimized features; Generate splicing features based on RGB features and depth features at the same resolution, perform spatial domain perception enhancement on the splicing features, and obtain spatial optimization features; The global Fourier optimization feature and the spatial optimization feature are subtracted and then spatial attention operation is performed to obtain the salient feature. The fusion result is obtained based on the salient feature, global Fourier optimization and spatial optimization features. The fusion results at different resolutions are spliced ​​to obtain the conditional feature of each time step. The diffusion model is used to predict the conditional features and loading reference graphs at different time steps to obtain the prediction results at different time steps. The prediction results at different time steps are filtered and aggregated to obtain the aggregated detection results.

[0027] like Figure 1 As shown, the method proposed in this embodiment includes two main parts: frequency-domain-spatial domain entanglement enhancement and stable time-step mask prediction. Frequency-domain-spatial domain entanglement enhancement uses the proposed Fourier-domain perception enhancement module and the spatial-domain perception module to reconstruct multimodal feature representations, improving the model's understanding of underwater scenes. A redundancy suppression attention strategy is used to enhance the underwater condition feature representation. The Fourier-domain perception enhancement module uses Fourier convolution to generate complementary features of different modalities in the Fourier domain; the spatial-domain perception module combines multi-scale convolution to extract local-global spatial features of different modalities. Subsequently, the redundancy suppression attention strategy uses spatial channel attention to highlight and fuse the complementary information of the two, effectively eliminating the interference caused by redundant information. Stable time-step mask prediction is based on mutual information theory. By calculating the stability of the rate of change of the multi-step iterative results of the diffusion model, the optimal intermediate mask is selected and weightedly fused to reduce the interference of noise perturbations on the prediction results, thereby generating a robust prediction result. In summary, this embodiment regards USOD as a mask generation task based on the conditional diffusion model, rather than a traditional detection task. The algorithm achieves complementary fusion of two modal information in dual domains through frequency-space entanglement enhancement, outputs robust conditional features with texture, edge and spectral consistency, and reduces the uncertainty generated by the iterative process through stable time-step mask prediction, thereby significantly improving the accuracy and boundary stability of underwater salient target detection.

[0028] Step S1: pre-process the original underwater image to obtain an RGB image, a depth image, and a loaded reference image; Step S2: Perform feature extraction on the RGB image and depth image in each time step to obtain multiple RGB features and depth features of different resolutions, specifically: Combined with each time step, the RGB image and depth image are input into the PVT model to extract features, and multiple RGB features and depth features of different resolutions are obtained; Among them, PVT-Pyramid Vision Transformer, pyramid vision transformer model.

[0029] Since the traditional USOD paradigm relies on pixel-level probabilities, which can lead to overconfident erroneous predictions and the accuracy of the conditional diffusion model is limited by the conditional guidance features, this embodiment uses the Fourier domain-spatial domain entanglement enhancement process to process multiple RGB features and depth features of different resolutions, and integrates the Fourier information of multiple RGB features and depth features of different resolutions as an auxiliary spatial information to provide a more comprehensive understanding of underwater scenes.

[0030] like Figure 2 As shown in the figure, the Fourier-spatial domain entanglement enhancement process combines the complementary properties of depth maps and RGB images in the frequency and spatial domains. It can utilize the characteristics of RGB images and depth maps in different domains to integrate information such as color, texture, edges, spectrum, amplitude, and energy, which is conducive to learning discriminative representations. At the same time, the redundant suppression attention strategy suppresses redundant information in the fusion process, reducing feature interference and generating more discriminative underwater scene condition features required by the diffusion model.

[0031] Through frequency domain analysis, it is found that in underwater scenes, RGB images and depth maps exhibit significant complementary features in the Fourier domain. The Fourier transform converts the image into the frequency domain and extracts amplitude and phase information. The RGB amplitude mainly represents the light and dark distribution of the scene, which helps to distinguish different objects and backgrounds. The phase of the depth map focuses on structure and contour information, which helps to distinguish the shape and boundary of objects. Based on this discovery, this embodiment proposes a Fourier domain-spatial domain entanglement module to collaboratively fuse dual modalities, using Fourier domain features as auxiliary branches of spatial domain features to reconstruct more discriminative multimodal feature representations.

[0032] Specifically, the multiple RGB features and depth features of different resolutions output by each layer of the PVT model will be enhanced through Fourier domain-spatial domain entanglement, including two branches: Fourier domain perception enhancement and spatial domain perception enhancement. The results of the two branches are then spliced ​​and fused to obtain conditional features, including: Step S3: Perform Fourier domain perceptual enhancement on the RGB features and depth features at the same resolution to obtain global Fourier optimized features, specifically: Extract the phase and amplitude components corresponding to the RGB features and depth features at each resolution, specifically: Convolution processing is performed on the RGB features and depth features at each resolution to obtain RGB convolution features and depth convolution features; Perform Fourier transform on the RGB convolution features to obtain RGB phase components and RGB amplitude components; Perform Fourier transform on the deep convolution feature to obtain the deep phase component and the deep amplitude component; Generate a spliced ​​phase component and a spliced ​​amplitude component based on the phase component and the amplitude component corresponding to the RGB feature and the depth feature; The spliced ​​phase component and the spliced ​​amplitude component are convolved and then inverse Fourier transformed to obtain the global Fourier optimization feature.

[0033] like Figure 3 As shown in Figure 2, in Fourier domain perception enhancement, the RGB features and depth features at different resolutions are outputted through Fourier transform to obtain the corresponding phase and amplitude components. The Fourier transform process can be expressed as:

[0034]

[0035] in, is the phase component, is the amplitude component, and It is The RGB features and depth features of the layer output.

[0036] The corresponding phase components and amplitude components are spliced ​​on the channels respectively, and the energy transformation and spatial structure between different modes are globally deconstructed through 1×1 convolution to help the network learn subtle changes in underwater scenes. Finally, the inverse Fourier transform is used to obtain the final global Fourier optimization feature. The process can be expressed as:

[0037] in, represents 1×1 convolution, Indicates the The RGB phase component and depth phase component channels of the layer are stitched together. Indicates the The RGB amplitude component and depth amplitude component channels of the layer are stitched together. represents the inverse Fourier transform, is the global Fourier optimization feature.

[0038] Step S4: Generate splicing features based on RGB features and depth features at the same resolution, perform spatial domain perception enhancement on the splicing features, and obtain spatial optimization features, specifically: Extract the RGB features and depth features at each resolution and then splice the resulting features; The spliced ​​features are processed by three multi-scale convolutions at the same time to obtain three core components of global self-attention , after the global self-attention operation, the self-attention map is obtained; Perform multi-scale convolution processing on the splicing features to obtain convolution splicing features; The self-attention map and the convolutional splicing features are concatenated and then convolved to obtain spatially optimized features.

[0039] like Figure 3 As shown in the process of spatial perception enhancement, the first Layer splicing features , the concatenated features will go through 1×1 convolution and a set of multi-scale convolutions to obtain the required self-attention operation , the self-attention map is obtained through the global self-attention operation After that, it can be superimposed with the local convolution result to reasonably capture the significant information in the image and adapt to the multi-scale requirements of underwater objects. The acquisition process is:

[0040] in, Indicates splicing by channel, represents the self-attention map, represents 3×3 and 5×5 multi-scale convolution, It is a spatial optimization feature.

[0041] Step S5: Perform spatial attention operation after subtracting the global Fourier optimization feature from the spatial optimization feature to obtain salient features. Based on the salient features, global Fourier optimization, and spatial optimization features, a fusion result is obtained. The fusion results at different resolutions are spliced ​​to obtain the conditional features of each time step, specifically: The spatial attention operation is performed after subtracting the global Fourier optimized features from the spatial optimized features to obtain the salient features; Multiply the salient feature with the global Fourier optimization feature and then add it to the spatial optimization feature to obtain the fused optimization feature; After splicing the fusion optimization features and the spatial optimization features, a channel attention operation is performed to obtain the channel fusion optimization features; Add the channel fusion optimization features to the RGB features to obtain the fusion result; The fusion results at different resolutions are spliced ​​together to obtain the conditional features of each time step.

[0042] The output results of the Fourier domain perception enhancement and spatial domain perception enhancement branches are entangled through a redundant suppression attention strategy to learn the information difference between Fourier features and spatial features. Complementary features are obtained by subtracting the two features, and spatial attention operations are used to highlight the complementary features. This removes the redundant information interference caused by the fusion process of the two and enhances feature discriminability. The specific process is as follows:

[0043]

[0044]

[0045] in, represents a 3×3 convolution, It represents the difference between Fourier optimization feature and spatial optimization feature. represents global maximum pooling, represents global average pooling, represents the salient features after using spatial attention, represents the input of channel attention, represents the Sigmoid activation function, represents multiplication, It represents a more discriminative fusion result after removing redundancy.

[0046] Finally, the more discriminative fusion results obtained by each layer will be adjusted through splicing and fusion to obtain the conditional features that need to be input into the diffusion model as a guide to enhance the ability of the diffusion model in downstream tasks.

[0047] Stable time-step mask prediction is based on mutual information theory. We recognize that the mask of each time step of the diffusion model contains useful information related to the target data. A stable rather than highly variable iterative process means that the result of the next step is based on the previous step, indicating that the information transfer is continuous. Therefore, after the rough features of the target are formed, the subsequent stable iterative process can effectively guide the edge generation process because the noise influence is relatively stable. Based on this, the proposed mask prediction module at a stable time step calculates the prediction change rate between each step. , choose the one with lower rate of change The predicted values ​​of the time steps are taken as the effective predicted values, and then these binary mask results are weighted averaged and voted to obtain the final aggregated result. , making the prediction more stable and precise. Selecting results with lower transformation rates can not only reduce abnormal results caused by noise mutations, but also take advantage of the guidance ability of stable results on boundaries.

[0048] During the diffusion model mask generation process, the model generates multiple intermediate results during the multi-step iterative denoising process. Based on mutual information theory, we recognize that the mask generated at each time step contains useful information about the target data. Different intermediate results have different representations of target boundaries and contours. Integrating these diverse intermediate results can provide richer information for final mask generation, thereby improving the quality of the segmentation results. The results of the early iterations are primarily used to generate rough features of the target. Once these rough features are formed, the noise influence is relatively stable, and the subsequent stable iterations can effectively guide the edge generation process, as shown in step S6.

[0049] Step S6: Use the diffusion model to predict the conditional features and loading reference graphs at different time steps to obtain prediction results at different time steps. Filter and aggregate the prediction results at different time steps to obtain aggregated detection results, specifically: The diffusion model is used to predict the conditional characteristics and loading reference graphs at different time steps to obtain the prediction results at different time steps; The predicted change rate of the current time step is obtained by subtracting the predicted result of the current time step from the predicted result of the previous time step; All prediction change rates within the total iteration time step T are sorted from small to large, and T / 2 prediction results with low prediction change rates are selected for weighted average aggregation to obtain the aggregated detection result.

[0050] Therefore, this embodiment proposes a process of mask prediction based on stable time steps, by calculating the predicted change rate between each time step. , sort the predicted change rates of all time steps from small to large, and select the ones with lower change rates The prediction results of the time steps are taken as the effective prediction values, and then the binary mask results of these effective prediction values ​​are weighted averaged to obtain the final aggregated detection results. , making the prediction more stable and precise. By choosing a lower transformation rate Time steps can not only reduce the outliers caused by noise mutations, but also take advantage of its ability to guide edges. The process of mask prediction based on stable time steps is as follows:

[0051]

[0052] in, is the total iteration time step, the default is 10, is the total iteration time step Half of the value, the default value is 5, For the The prediction results of the time step, It is The rate of change of the time step, is the aggregated detection result, Indicates the smallest rate of change The subscript of the prediction result.

[0053] In the overall processing framework of the method described in this embodiment, during the training process, weighted intersection over union (IoU) loss and weighted binary cross entropy (BCE) loss are used as loss functions. Therefore, the loss function It can be expressed as:

[0054] in, and represents the weighted intersection-over-union (IoU) loss and weighted binary cross entropy (BCE) loss, Load the reference image.

[0055] The method proposed in this embodiment is compared with other state-of-the-art methods (TC-USOD, Dual-SAM, SENET, etc.) on the USOD10K dataset and the USOD dataset. It is proven that the method proposed in this embodiment can achieve superior results on the USOD task. It can not only successfully distinguish the difference between objects and backgrounds, but also obtain accurate edge results, even in challenging areas such as Figure 4 The scenes of rows 2 and 3 are shown.

[0056] Example 2 This embodiment provides a system for detecting salient objects in underwater images based on a conditional diffusion model, including: The image processing module is used to pre-process the original underwater image to obtain the RGB image, depth map and loaded reference image; The feature extraction module is used to extract features from the RGB image and depth image in each time step, and obtain multiple RGB features and depth features of different resolutions respectively; The Fourier domain perception enhancement module is used to perform Fourier domain perception enhancement on the RGB features and depth features at each resolution to obtain global Fourier optimized features; The spatial perception enhancement module is used to generate splicing features based on the RGB features and depth features at each resolution, and perform spatial perception enhancement on the splicing features to obtain spatially optimized features; The conditional feature generation module is used to perform spatial attention operations after subtracting the global Fourier optimization features from the spatial optimization features to obtain salient features. Based on the salient features, global Fourier optimization, and spatial optimization features, a fusion result is obtained. The fusion results at different resolutions are spliced ​​to obtain the conditional features of each time step. The prediction result fusion module is used to use the diffusion model to predict the conditional features and loading reference graphs at different time steps to obtain the prediction results at different time steps, and to filter and aggregate the prediction results at different time steps to obtain the aggregated detection results.

[0057] The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiment 1. It should be noted that the above modules as part of the system can be executed in a computer system such as a set of computer executable instructions.

[0058] The descriptions of the various embodiments in the above embodiments have different focuses. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0059] The proposed system can be implemented in other ways. For example, the system embodiment described above is merely illustrative. For example, the above module division is only a logical function division. In actual implementation, other division methods may be used. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not implemented.

[0060] Example 3 This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the method for detecting salient objects in underwater images based on a conditional diffusion model as described in the first embodiment above are implemented.

[0061] Example 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the method for detecting salient targets in underwater images based on the conditional diffusion model as described in the first embodiment above are implemented.

[0062] Example 5 This embodiment provides a computer program product or computer program, including computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps of the underwater image salient target detection method based on the conditional diffusion model described in the above-mentioned embodiment 1.

[0063] Those skilled in the art will appreciate that embodiments of the present invention may provide methods, systems, or computer program products. Thus, the present invention may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage) containing computer-usable program code.

[0064] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0065] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0066] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0067] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0068] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.

Claims

1. A method for detecting salient objects in underwater images based on a conditional diffusion model, characterized in that: include: Preprocess the original underwater image to obtain RGB image, depth image and loading reference image; Perform feature extraction on the RGB image and depth image in each time step to obtain multiple RGB features and depth features with different resolutions; Perform Fourier domain perceptual enhancement on RGB features and depth features at the same resolution to obtain global Fourier optimized features; Generate splicing features based on RGB features and depth features at the same resolution, perform spatial domain perception enhancement on the splicing features, and obtain spatial optimization features; The global Fourier optimization feature and the spatial optimization feature are subtracted and then spatial attention operation is performed to obtain the salient feature. The fusion result is obtained based on the salient feature, global Fourier optimization and spatial optimization features. The fusion results at different resolutions are spliced ​​to obtain the conditional feature of each time step. The diffusion model is used to predict the conditional features and loading reference graphs at different time steps to obtain the prediction results at different time steps. The prediction results at different time steps are filtered and aggregated to obtain the aggregated detection results.

2. The method for detecting salient objects in underwater images based on a conditional diffusion model according to claim 1, wherein: The RGB features and depth features at each resolution are enhanced in the Fourier domain to obtain the global Fourier optimized features, specifically: Extract the phase and amplitude components corresponding to the RGB features and depth features at each resolution; Generate a spliced ​​phase component and a spliced ​​amplitude component based on the phase component and the amplitude component corresponding to the RGB feature and the depth feature; The spliced ​​phase component and the spliced ​​amplitude component are convolved and then inverse Fourier transformed to obtain the global Fourier optimization feature.

3. The method for detecting salient objects in underwater images based on a conditional diffusion model according to claim 2, wherein: Extract the phase and amplitude components of the RGB features and depth features at each resolution, specifically: Convolution processing is performed on the RGB features and depth features at each resolution to obtain RGB convolution features and depth convolution features; Perform Fourier transform on the RGB convolution features to obtain RGB phase components and RGB amplitude components; Perform Fourier transform on the deep convolution features to obtain the deep phase component and the deep amplitude component.

4. The method for detecting salient objects in underwater images based on a conditional diffusion model according to claim 1, wherein: The method generates splicing features based on the RGB features and depth features at each resolution, performs spatial domain perception enhancement on the splicing features, and obtains spatial optimization features, specifically: Extract the RGB features and depth features at each resolution and then splice the resulting features; The spliced ​​features are processed by three multi-scale convolutions simultaneously to obtain three core components of global self-attention, and the self-attention map is obtained after the global self-attention operation; Perform multi-scale convolution processing on the splicing features to obtain convolution splicing features; The self-attention map and the convolutional splicing features are concatenated and then convolved to obtain spatially optimized features.

5. The method for detecting salient objects in underwater images based on a conditional diffusion model according to claim 1, wherein: The spatial attention operation is performed after subtracting the global Fourier optimization feature and the spatial optimization feature to obtain the salient feature. Based on the salient feature, the global Fourier optimization and the spatial optimization feature, the fusion result is obtained, specifically: The spatial attention operation is performed after subtracting the global Fourier optimized features from the spatial optimized features to obtain the salient features; Multiply the salient feature with the global Fourier optimization feature and then add it to the spatial optimization feature to obtain the fused optimization feature; After splicing the fusion optimization features and the spatial optimization features, a channel attention operation is performed to obtain the channel fusion optimization features; Add the channel fusion optimization features to the RGB features to obtain the fusion result.

6. The method for detecting salient objects in underwater images based on a conditional diffusion model according to claim 1, wherein: Based on the prediction results of different time steps, the aggregated detection results are obtained by screening and aggregation, specifically: The predicted change rate of the current time step is obtained by subtracting the predicted result of the current time step from the predicted result of the previous time step; All prediction change rates within the total iteration time step T are sorted from small to large, and T / 2 prediction results with low prediction change rates are selected for weighted average aggregation to obtain the aggregated detection result.

7. Underwater image salient object detection system based on conditional diffusion model, characterized by: include: The image processing module is used to pre-process the original underwater image to obtain the RGB image, depth map and loaded reference image; The feature extraction module is used to extract features from the RGB image and depth image in each time step, and obtain multiple RGB features and depth features of different resolutions respectively; The Fourier domain perception enhancement module is used to perform Fourier domain perception enhancement on the RGB features and depth features at each resolution to obtain global Fourier optimized features; The spatial perception enhancement module is used to generate splicing features based on the RGB features and depth features at each resolution, and perform spatial perception enhancement on the splicing features to obtain spatially optimized features; The conditional feature generation module is used to perform spatial attention operations after subtracting the global Fourier optimization features from the spatial optimization features to obtain salient features. Based on the salient features, global Fourier optimization, and spatial optimization features, a fusion result is obtained. The fusion results at different resolutions are spliced ​​to obtain the conditional features of each time step. The prediction result fusion module is used to use the diffusion model to predict the conditional features and loading reference graphs at different time steps to obtain the prediction results at different time steps, and to filter and aggregate the prediction results at different time steps to obtain the aggregated detection results.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for detecting salient targets in underwater images based on a conditional diffusion model are implemented as described in any one of claims 1 to 6.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method for detecting salient targets in underwater images based on a conditional diffusion model are implemented as described in any one of claims 1 to 6.

10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the steps of the underwater image salient object detection method based on the conditional diffusion model according to any one of claims 1 to 6.

Citation Information

Cited By

  • RGB-D image saliency detection method based on frequency decoupling mode interaction

    CN121962587A