Marine oil spill segmentation method and system based on spatial spectrum frequency constraint diffusion model

By employing bi-branch frequency analysis and a spatial-spectral attention mechanism within a spatial-spectral frequency-constrained diffusion model, the challenge of automated identification of emulsified states in marine oil spill detection was solved. This enabled high-precision oil spill segmentation under different glare conditions, improving detection efficiency and accuracy.

CN121582272APending Publication Date: 2026-02-27FIRST INSTITUTE OF OCEANOGRAPHY MNR
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511776132.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing technologies for marine oil spill detection suffer from limitations such as reliance on human-driven spectral interpretation for emulsification state identification, inability to achieve end-to-end automated segmentation, and difficulty in effectively resolving complex spectral variations and subtle texture patterns under different solar glare conditions, resulting in insufficient detection accuracy and efficiency.

Method used

A method based on a spatial-spectral frequency-constrained diffusion model is adopted. By combining a dual-branch frequency resolver and a spatial-spectral attention mechanism with multispectral remote sensing images for iterative denoising, the automatic and accurate identification of oil spill segmentation masks can be achieved.

Benefits of technology

Precise pixel-level segmentation of unemulsified oil, emulsified oil, and positive-negative contrast spills was achieved under different solar glare conditions, improving detection accuracy and robustness while reducing process complexity and uncertainty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121582272A_ABST
    Figure CN121582272A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of remote sensing image processing and computer vision, and discloses a marine oil spill segmentation method and system based on a spatial spectrum frequency constraint diffusion model. Preprocessing the multispectral remote sensing image to obtain reflectivity data after Rayleigh correction; inputting the reflectivity data into a trained spatial-spectral frequency constraint diffusion model, and generating an oil spill segmentation mask through iterative denoising; wherein the spatial spectrum frequency constraint diffusion model is integrated with a double-branch frequency analyzer (DF-P) and a spatial spectrum attention condition coding module (SS-A). Through innovative architecture design, end-to-end oil spill accurate segmentation under the condition of cross-sun flare is realized, the segmentation performance is remarkably improved, and the core design thought also has important reference value for similar optical remote sensing ground feature segmentation tasks; and meanwhile, more powerful technical support can be provided for marine environment protection and oil spill emergency response.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of remote sensing image processing and computer vision, and particularly relates to a marine oil spill segmentation method and system based on a spatial-spectral frequency constraint diffusion model. BACKGROUND

[0002] Marine oil spill is the release phenomenon of liquid hydrocarbon substances in the marine environment, and its source covers natural leakage and human discharge. In the process of oil spill diffusion, its form is driven by comprehensive physical dynamic factors such as wind field, ocean current and wave, and accompanied by physical, chemical and biological processes such as evaporation, dissolution, emulsification and biodegradation, finally forming a complex pollution form with thickness gradient and emulsification heterogeneity. This not only poses a serious threat to the global marine ecosystem, but also has a profound and adverse impact on the sustainable development of fishery resources and coastal economy. With the continuous growth of global energy demand, marine oil exploration, exploitation and transportation activities are becoming more and more frequent, which objectively leads to a significant increase in the incidence of oil spills. In the face of sudden oil spill events, it is crucial to quickly and accurately identify and detect the scope of oil spill in order to develop effective emergency response strategies and optimize cleanup plans.

[0003] Remote sensing detection of marine surface oil can be achieved through various sensors, and the main means are synthetic aperture radar (SAR) and optical remote sensing. Oil film reduces Bragg scattering signals by suppressing sea surface capillary waves and short gravity waves, thus showing negative contrast of oil and water on SAR images under certain wind field conditions. However, due to its dependence on sea surface roughness, it is vulnerable to false positives from similar objects (such as algae blooms, low wind speed areas), and its ability to distinguish emulsification states is limited. Optical remote sensing can detect oil spills based on the difference in optical properties (absorption and scattering) between oil and water, and can further identify the emulsification state of oil spills. On the other hand, optical remote sensing can also detect oil films based on their smoothing effect on the sea surface with the aid of solar glint reflection. In this way, optical remote sensing can detect very thin oil films, and the presence of solar glint makes the oil slick in the optical image appear brighter (positive contrast) or darker (negative contrast) than the surrounding water. This phenomenon can be explained by observing the geometric shape and intensity of solar glint reflection. Studies have shown that the angle between the direction of specularly reflected sunlight and the direction of sensor detection (θ) is the key to determining the contrast of oil slicks in optical images. When θ is less than 90 degrees, the oil slick appears brighter than the surrounding water (positive contrast); when θ is greater than 90 degrees, the oil slick appears darker than the surrounding water (negative contrast). The ratio of the radiance of the oil slick to that of the clean water is the most effective indicator to evaluate the difference of the glint reflection between the oil slick and the clean water, and the critical angle of the contrast and inversion of the glint reflection between the oil slick and the clean water is located at 12° to 13°. Even if the property of the oil slick does not change dramatically in a short time, the optical response characteristics of the oil slick are significantly different in different sensor images due to different observation geometries. When the observation geometry is far from the mirror direction, the reflectivity of the oil slick is mainly affected by its own optical properties at this time; the emulsified oil is in positive contrast with the background water due to the strong backscattering of incident light, while the unemulsified oil has a lower reflectivity than the background water due to the strong absorption of incident light, showing negative contrast. However, when the observation geometry is closer to the mirror direction, the smoothness of the oil film to the sea surface is more likely to form a mirror reflection, resulting in strong solar glint reflection, which masks the inherent optical property characteristics of emulsified and unemulsified oil, making both of them show positive contrast characteristics. This complex contrast inversion phenomenon, as well as the optical response difference of different emulsification states under different glint conditions, constitutes the core challenge of oil spill detection by optical remote sensing.

[0004] In recent years, deep learning technology has brought a new paradigm to oil spill detection, significantly improving detection accuracy by extracting multi-scale spatial spectral features through models such as convolutional neural networks (CNN) and visual transformers (ViT). Numerous researchers have used various deep learning methods to detect oil spills, identify oil types, and retrieve oil thickness based on airborne hyperspectral images combined with ground-controlled experiments. Although these innovative technologies perform well with high-resolution unmanned aerial vehicle hyperspectral data, the response characteristics of oil spills in controlled experiments differ significantly from real marine environments, so the effectiveness of these algorithms still needs to be strictly tested with real scene data. At the same time, scholars have also analyzed the satellite optical characteristics of oil spills in real marine environments based on satellite multispectral remote sensing payloads and conducted detection algorithm research. For example, Du et al. proposed the CBF-CNN framework, which effectively addresses the class imbalance problem in multispectral datasets, significantly improving recognition rates in weak glint conditions. Jiang et al. innovatively integrated Landsat-8 / 9 multispectral and thermal infrared data to improve the accuracy of emulsification state identification in complex marine environments. However, the above research generally focuses on oil spill detection under weak solar glint reflection conditions, and the complexity of oil spill optical characteristics and segmentation challenges under strong solar glint reflection conditions remain insufficient.

[0005] In recent years, more and more scholars have begun to develop general oil spill detection algorithms that adapt to different solar glint reflection conditions. For example, Jiao et al. tried to combine deep learning with spectral spatial geometric features to detect positive and negative contrast oil spills, but the recognition of emulsified state still relies on human-driven spectral interpretation, and cannot achieve completely end-to-end automated segmentation. Wang et al. developed an oil spill optical remote sensing detection method based on the YOLOv5 framework, which can automatically identify oil spill areas under different solar glint reflection contrasts in multi-source satellite images, and can distinguish between emulsified and non-emulsified oil types. However, this model can only automatically output the target level rectangular detection box and class label, and the fine pixel-level segmentation mask still needs to rely on post-processing spectral quantization analysis, increasing the complexity and uncertainty of the process. In summary, under the condition of crossing solar glint reflection (positive / negative contrast difference), automatically generating pixel-level segmentation masks of different emulsified oil spills is still a technical bottleneck that needs to be broken through.

[0006] Although the encoder-decoder structure deep learning feature extraction algorithm commonly used in marine oil spill detection tasks performs well in specific scenarios, it mainly relies on static pre-defined operations (such as fixed convolution kernels or fixed mode attention mechanisms) for single feature extraction, and its receptive field or attention range is often limited, which makes it difficult to effectively analyze and model the complex, dynamic and nonlinear spectral changes and subtle texture patterns exhibited by oil spills of different emulsified states under different solar glint intensities. In contrast, denoising diffusion probabilistic models (DDPM) have shown impressive performance in image generation, denoising, restoration and super-resolution tasks, providing a promising alternative. The feature representation learned by diffusion models has been proven to significantly improve the performance of various downstream discriminant tasks, such as object detection, classification or semantic segmentation. With its unique probabilistic generation process, diffusion models gradually introduce Gaussian noise to simulate image degradation by forward diffusion, and then iteratively reconstruct the original target mask by reverse denoising, showing excellent feature extraction capability in image segmentation tasks. Compared with traditional single forward propagation feature extraction methods, the probabilistic iterative learning paradigm of diffusion models has significant advantages in feature extraction robustness, adaptability to complex and variable scenarios, and ability to capture subtle differences, making it particularly suitable for processing complex and variable remote sensing image features.

[0007] Through the above analysis, the problems and defects of the prior art are: (1) The method of combining deep learning with spectral spatial geometric features to detect positive and negative contrast oil spills still relies on human-driven spectral interpretation for emulsified state recognition, and cannot achieve completely end-to-end automated segmentation.

[0008] (2) Existing optical remote sensing detection methods for oil spills based on the YOLOv5 framework can only automatically output target-level rectangular detection boxes and category labels. Fine pixel-level segmentation masks still need to rely on post-processing spectral quantization analysis, which increases the complexity and uncertainty of the process.

[0009] (3) The encoder-decoder structure deep learning feature extraction algorithm commonly used in marine oil spill detection tasks mainly relies on static predefined operations for single feature extraction. Its receptive field or attention range is often limited, which makes it difficult to effectively analyze and model the complex, dynamic and nonlinear spectral changes and subtle texture patterns exhibited by oil spills under different solar flare intensities and emulsification states.

[0010] (4) Although the iterative optimization process of DDPM is powerful, the standard implementation lacks a dedicated mechanism to resolve the complex physical characteristics of solar flares or distinguish subtle spectral differences between emulsion states.

[0011] To address the limitations of the existing technologies, particularly the challenges of oil-water contrast reversal caused by solar glare interference and the difficulty of achieving pixel-level end-to-end emulsification state identification using existing methods, this invention aims to achieve high-precision automated detection of oil spills in complex marine environments through the synergistic effect of dual-branch frequency analysis and spatial-spectral attention mechanisms. Summary of the Invention

[0012] To overcome the problems existing in related technologies, the present invention discloses an embodiment of a marine oil spill segmentation method and system based on a spatial spectrum frequency-constrained diffusion model, the technical solution of which is as follows: This invention is implemented as follows: a marine oil spill segmentation method based on a spatial frequency-constrained diffusion model, comprising the following steps: S1, acquire multispectral remote sensing images; S2, preprocess the multispectral remote sensing image to obtain Rayleigh-corrected reflectance data; S3, the reflectivity data is input into the trained spatial frequency constraint diffusion model, and an oil spill segmentation mask is generated through iterative denoising; the spatial frequency constraint diffusion model adds noise to the segmentation mask in the forward diffusion process, and reconstructs the segmentation mask from the noise in the reverse denoising process. The reverse denoising process is implemented by a neural network, which includes: A segmentation encoder is used to extract features from a noisy mask. A conditional encoder is used to process the Rayleigh-corrected reflectivity data. The conditional encoder integrates a dual-branch frequency resolver DF-P and a spatial spectrum attention conditional coding module SS-A. Its output features are fused with the features of the segmentation encoder to guide the inverse denoising process. A segmentation decoder is configured to fuse the output features of the segmentation encoder and the conditional encoder, and to reconstruct the oil spill segmentation mask.

[0013] In step S1, the multispectral remote sensing image includes, but is not limited to, an optical remote sensing image carried by an airplane or a satellite.

[0014] In step S2, the calculation formula of the Rayleigh correction is as follows: ;

[0015] In the formula, is the corrected reflectivity, is the total radiance of the pixel, is the Rayleigh scattering radiance, is the solar incident irradiance, is the solar zenith angle; The preprocessed image, combined with the manually labeled oil spill true value mask, constitutes a data set for model training, verification and testing; In step S3, the forward diffusion process includes: adding Gaussian noise to the original data by steps, so that is gradually converted into a pure noise image ; at each time step , noise is added to the data with a predefined variance schedule , which is used to provide training data for the reverse denoising process and simulate the degradation of the image under noise disturbance; given the original image , the noise image at any time step is represented as: ;

[0016] In the formula, is the original image, is the noise image at time step , , is the predefined variance schedule at time step , is a Gaussian distribution, is an identity matrix; The reverse denoising process includes: training a neural network to gradually remove noise from the image with noise and finally recover the original data ; at each time step, the model learns to predict the image at the previous time step ; the denoising function of the model is represented as: ;

[0017] wherein, is the condition information, is the time step of the noisy image, is the image of the previous time step predicted by the model, is the time step, is the neural network parameter.

[0018] In step S3, the segmentation encoder is responsible for extracting multi-scale feature representations from the input image with noise ; composed of a series of residual blocks and down-sampling layers, for capturing spatial information at different scales; the initial convolutional layer maps the input channels to the initial feature dimension; the down-sampling path gradually increases the number of feature channels while reducing the spatial size of the feature map through a series of dimension multiplication factors; each down-sampling stage contains at least one residual block; The conditional encoder is used to independently process the condition information I of the original remote sensing image and perform deep feature extraction and enhancement; the conditional encoder integrates a dual-branch frequency parser DF-P and a spectral-spatial attention conditional encoding module SS-A; the initial convolutional layer maps the number of original image channels to the same initial feature dimension as the segmentation encoder; at each down-sampling stage, an attention module is followed by a customized conditional processing module, which integrates DF-P and SS-A; the feature map size of the conditional processing module is consistent with the feature map size of the current down-sampling stage, and the feature dimension matches the number of channels of the current stage; the enhanced context information output by the conditional encoder is effectively fused with the features of the segmentation encoder through the attention mechanism, thereby accurately guiding the denoising process; The segmentation decoder uses the features extracted by the segmentation encoder and the conditional encoder to gradually reconstruct a fine oil spill segmentation mask; the segmentation decoder is built using up-sampling and residual blocks, and receives multi-scale features from the encoder through skip connections; the up-sampling path is symmetrical to the down-sampling path, gradually increasing the spatial size of the feature map while reducing the number of feature channels; each up-sampling stage contains two residual blocks; the feature dimension of the skip connection is the sum of the input dimension and the conditional feature dimension; finally, a residual block and a convolutional layer map the processed features back to the output mask channel number.

[0019] In step S3, the dual-branch frequency parser DF-P decomposes the image features in the frequency domain, separating and processing the high-frequency and low-frequency information in the image; the high-frequency features correspond to the details, edges, and textures of the image, and the low-frequency features represent the overall structure, background distribution, and brightness changes of the image; The dual-branch frequency parser performs two-dimensional Fourier transform on the input feature map, converting the input feature map to the frequency domain; for a two-dimensional image , the Fourier transform is expressed as: ;

[0020] wherein, is the input image feature in spatial domain, is the output feature in frequency domain, is the image size, is the spatial coordinate, is the frequency coordinate, is the imaginary unit; the dual-branch frequency analyzer applies a learnable frequency mask to high-frequency and low-frequency components respectively; high-frequency processing: applying specific learnable weights to the high-frequency part of the Fourier-transformed feature map ; low-frequency processing: applying learnable weights to the low-frequency part of the Fourier-transformed feature map ; the feature map after frequency domain operation is represented as: = ; wherein, is the weighted feature map, is the unweighted feature map, represents the high-frequency or low-frequency mask; the dual-branch frequency analyzer generates a dynamic frequency attention map according to the real part splicing result of the Fourier transform of the noise image and the conditional image , thereby adaptively adjusting the weights of high-frequency and low-frequency; after frequency domain operation, the feature map is converted back to the spatial domain through inverse Fourier transform and integrated into subsequent processing; wherein, the inverse Fourier transform is represented as: ;

[0021] is the feature returned to the spatial domain after inverse transformation, is the feature after frequency domain operation.

[0022] In step S3, the spatial-spectral attention conditional encoding module SS-A optimizes the extraction process of spatial-spectral features by fusing spectral and spatial attention mechanisms; a convolution layer is used to extract basic features from the conditional image , wherein, is the real space, is the batch size, is the number of channels, is the image height, is the image width, is the feature dimension, is the number of categories; the base features are rearranged as ; Spectral attention mechanism: realized through a sequence module, including adaptive average pooling, convolution layer, ReLU activation and Sigmoid activation; for input features , the calculation formula of spectral attention weight is as follows: ;

[0023] In the formula, denotes global average pooling, is a convolution layer, ReLU is an activation function, and Sigmoid is a Sigmoid activation function; Spatial attention mechanism: realized through a convolution layer, ReLU activation and Sigmoid activation; for input features , the calculation formula of spatial attention weight is as follows: ;

[0024] Through joint modeling, the optimized spatial-spectral features are then input into the denoising process of the diffusion model as enhanced conditional information, and work together with the iterative optimization mechanism of the diffusion model.

[0025] In step S3, the neural network is trained by minimizing the difference between the predicted noise and the actual noise; the loss function adopts mean square error, and when the model target is to predict noise , the loss function is: ;

[0026] Wherein, is the difference between the predicted noise and the actual noise, is the actual added noise, is the noise predicted by the model.

[0027] In step S3, the oil spill segmentation mask includes one or more categories of unemulsified oil, emulsified oil, positive contrast oil spill and negative contrast oil spill, for realizing pixel-level segmentation under different solar flare conditions.

[0028] Another purpose of the present application is to provide a marine oil spill segmentation system based on a spatial-spectral frequency constraint diffusion model, which is used to realize the marine oil spill segmentation method based on the spatial-spectral frequency constraint diffusion model, and the system comprises: A data acquisition module is used to acquire a multispectral remote sensing image. a preprocessing module configured to preprocess the multispectral remote sensing image to obtain Rayleigh-corrected reflectance data; a segmentation processing module including a hyperspectral-frequency constrained diffusion model integrated with a dual-branch frequency parser DF-P and a hyperspectral attention conditional encoding module SS-A, configured to receive the Rayleigh-corrected reflectance data output by the preprocessing module and generate an oil spill segmentation mask through an iterative denoising process; The segmentation processing module includes: a segmentation encoding unit configured to extract multi-scale features from the noisy mask; a conditional encoding unit configured to process the Rayleigh-corrected reflectance data, the conditional encoding unit integrated with a dual-branch frequency parser DF-P and a hyperspectral attention conditional encoding module SS-A; a segmentation decoding unit configured to fuse the output features of the segmentation encoding unit and the conditional encoding unit to reconstruct the oil spill segmentation mask; a result output module configured to output or store the oil spill segmentation mask.

[0029] In combination with all the above technical solutions, the present application has the following beneficial effects: First, marine oil spills pose a serious threat to the ecological environment and economic development, so rapid and accurate identification is crucial. Although optical remote sensing has the advantage of wide-area monitoring, the complex optical characteristics caused by solar glint pose a challenge to traditional methods. To address the limitations of existing deep learning methods in identifying variable glint conditions and emulsification states, the present invention proposes a spatial-spectral attention conditional diffusion probability model (OSS-Diff). This model achieves breakthroughs through three innovations: First, it uses the iterative denoising capability of the diffusion probability model to accurately capture the high-brightness reflection in strong glint areas and the negative contrast features in weak light areas. Second, it designs a dual-branch frequency analyzer (DF-P) to separate and process high-frequency boundary features and low-frequency background features in the frequency domain, significantly improving positive and negative contrast oil spill identification and suppressing noise. Third, it develops a spatial-spectral attention module (SS-A) to adaptively fuse multispectral spatial-spectral information and enhance key features for emulsification state discrimination. Experimental results show that the OSS-Diff model achieves F1 scores of 0.907 and 0.862 for unemulsified oil and emulsified oil identification under weak glint conditions, and an F1 score of 0.918 for positive contrast oil spill detection under strong glint conditions. Compared with advanced methods such as Unet, DeepLabV3+, and TransUnet, the model's average performance improves by 8.3%. Ablation experiments confirm that DF-P and SS-A contribute 39.9% and 27.2% of the performance gain, respectively. The model successfully handles the positive and negative contrast reversal phenomenon in the dual-star observation scenario and effectively identifies the emulsification state. In the same image with different contrast oil spill regions, the model achieves F1 scores of 0.901 and 0.934 for negative contrast oil spills and F1 scores of 0.923 and 0.9505 for positive contrast oil spills, demonstrating strong adaptability to glint changes. Under interference environments such as uneven illumination, rough sea surface, background noise, clouds, and their shadows, the model maintains an average F1 value of 0.872 or higher, demonstrating excellent robustness. Therefore, the present invention provides an efficient and reliable new paradigm for accurate oil spill monitoring under complex solar glint conditions.

[0030] Second, the present invention proposes an innovative spatial-spectral attention conditional diffusion probability model (OSS-Diff) that successfully solves the oil spill segmentation problem in optical remote sensing images under complex solar glint conditions by integrating a dual-branch frequency analyzer (DF-P) and a spatial-spectral attention encoding module (SS-A). Experimental results show that this method has significantly improved performance in multiple key performance indicators, providing a new technical solution for marine oil spill monitoring.

[0031] Third, the frequency processing (FF-Parser) in the prior art only focuses on suppressing high-frequency noise, and such a "single-path" suppression scheme cannot cope with the complex physical problem caused by solar flare, that is, the existence of high-frequency flare noise and low-frequency background distortion. The "double-branch frequency parser" (DF-P) of the present application realizes the preservation of high-frequency flare edges and the adaptive reweighting of low-frequency background distortion through high-low frequency separation processing, which is the key to solving the above-mentioned physical problems. The automatic and high-precision oil spill segmentation capability provided by the present application can directly serve the national marine environmental monitoring, oil spill emergency response departments, and commercial satellite data service companies. By replacing the traditional high-cost and low-efficiency manual interpretation.

[0032] Fourth, the present application first combines the iterative diffusion probability model (DDPM) with the frequency domain (DF-P) and the space-spectrum domain (SS-A) constraints dedicated to remote sensing physical characteristics, and realizes end-to-end, automatic pixel-level segmentation of oil spill emulsification state under strong solar flare interference. BRIEF DESCRIPTION OF DRAWINGS

[0033] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure; Figure 1 is a flow chart of the marine oil spill segmentation method based on the space-spectrum frequency constraint diffusion model provided by the embodiments of the present application; Figure 2 is a whole architecture diagram of the OSS-Diff model provided by the embodiments of the present application; Figure 3 is a double-branch frequency parser structure diagram provided by the embodiments of the present application; Figure 4 is a space-spectrum attention conditional encoding module structure diagram provided by the embodiments of the present application; Figure 5 is an oil spill result diagram under different solar flare reflection conditions provided by the embodiments of the present application; wherein (a1)-(c1) are oil spill images under weak solar flare reflection conditions, (f1)-(j1) are oil spill images under strong solar flare reflection conditions, and (a2)-(j2) are oil spill detection results of the model proposed by the present application; Figure 6 is a F1-Score comparison diagram of the present application and other advanced algorithms. DETAILED DESCRIPTION

[0034] In order to make the above objectives, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application are described in detail below. In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the spirit of the present application, so the present application is not limited to the specific implementations disclosed below.

[0035] The present application first applies the diffusion probability model to the field of oil spill segmentation, and the iterative denoising mechanism can effectively capture the feature changes under different glint conditions. The dual-branch frequency parser proposed in the present application separates the frequency domain features, significantly improving the model's ability to identify positive and negative contrast oil spills. The spatial-spectral attention module designed in the present application realizes the adaptive fusion of multispectral information, enhancing the model's ability to distinguish emulsification state. The organic combination of these innovations enables the model to perform excellently in complex scenarios. In order to overcome the limitations of existing encoder-decoder structure oil spill detection methods, the present application proposes a novel "Oil Spill Segmentation with Diffusion model (OSS-Diff) based on spatial-spectral attention conditional diffusion probability model and dual-branch frequency parsing method". The core innovation of this method lies in fully utilizing the inherent advantages of diffusion model in complex and dynamic feature learning, and introducing two key modules: 1. Dual-branch Frequency Parser (DF-P): In the frequency domain, image features are decomposed and processed, effectively separating and finely controlling high-frequency (such as edges, textures) and low-frequency (such as background brightness distribution) information, enhancing the model's ability to identify positive / negative contrast oil spill boundaries and noise suppression.

[0036] 2. Spatial-spectral Attention conditioned encoder module (SS-A): adaptively fuses the spatial context and spectral band information of multispectral data, highlighting the reflectivity characteristics of oil spills under different wavebands, and strengthening the attention to local details (especially the key areas for emulsification state discrimination).

[0037] The present application overcomes the limitations of existing methods in oil spill segmentation under different solar glint conditions, and ultimately realizes the precise and automated pixel-level identification of positive contrast oil spills, negative contrast oil spills, and different emulsification states (unemulsified oil, emulsified oil), thereby providing more reliable and efficient technical support for intelligent monitoring and emergency response of marine oil spills.

[0038] In embodiment 1, as Figure 1As shown, the ocean oil spill segmentation method based on the hyperspectral frequency constraint diffusion model provided by the embodiment of the present application comprises the following steps: S1, acquiring a multispectral remote sensing image; S2, preprocessing the multispectral remote sensing image to obtain reflectivity data after Rayleigh correction; S3, inputting the reflectivity data into the trained hyperspectral frequency constraint diffusion model to generate an oil spill segmentation mask through iterative denoising; the hyperspectral frequency constraint diffusion model adds noise to the segmentation mask through a forward diffusion process, and reconstructs the segmentation mask from the noise through a reverse denoising process; Wherein, the reverse denoising process is realized by a neural network, which comprises: a segmentation encoder for extracting features from the noisy mask; a conditional encoder for processing the reflectivity data after Rayleigh correction, the conditional encoder integrating a double-branch frequency parser DF-P and a hyperspectral attention conditional encoding module SS-A; a segmentation decoder for fusing the output features of the segmentation encoder and the conditional encoder and reconstructing the oil spill segmentation mask.

[0039] As a preferred, the embodiment of the present application mainly uses the CZI data carried by HY-1C and HY-1D and the HY-3A CZI-II sensor data. HY-1C was launched on September 7, 2018, and HY-1D was launched on June 11, 2020. The two satellites work together, and their CZI sensors have a 950-kilometer wide scanning band and a 50-meter spatial resolution, achieving a global revisit frequency of two times in three days, providing high-frequency observation capability for dynamic monitoring of marine environment. HY-3A satellite was launched on November 16, 2023, carrying CZI2 load multispectral spatial resolution increased to 20m, side swing condition can see more than 1000km, coverage cycle is 3 days 1 times.

[0040] The present application takes HY-1C / D CZI and HY-3A CZI2 sensor data as the main data source, and combines Landsat-8 / 9 OLI and Sentinel-2 MSI data for application verification. These satellites all have similar band configurations as CZI. The band specifications of HY-1C / D CZI, HY-1E CZI-II and other satellite loads used are shown in Table 1.

[0041] Table 1 Band specifications of optical loads of satellites used

[0042] The construction of the oil spill dataset provided by the embodiment of the present application comprises: a. Data space distribution: The present application collects data of 495 oil spill areas, which covers different solar glint conditions, including 95 oil spill images under strong glint conditions and 400 oil spill images under weak glint conditions. The spatial distribution of these oil spill images is extensive, mainly concentrated in key marine areas such as the Persian Gulf, the Bay of Bengal, the South China Sea, the Strait of Malacca, and the Yellow Sea. This diverse geographical distribution and glint condition coverage ensures the representativeness of the dataset and the generalization ability of the model.

[0043] b. Data preprocessing: All remote sensing images in the present application are processed into Rayleigh-corrected reflectance data (R) ) to eliminate the influence of sensor errors and Rayleigh scattering effects on image quality. The Rayleigh correction calculation formula is as follows: ;

[0044] In the formula, is the corrected Rayleigh-corrected reflectance, is the total radiance of the pixel (L ), is the Rayleigh scattering radiance, is the solar incident irradiance, is the solar zenith angle; After preprocessing, the images combined with the manually labeled oil spill true value mask constitute the complete dataset used for model training, verification, and testing in the present application.

[0045] In order to clearly define the difference in glint reflection between oil spill and clean sea surface, the present application also calculates the angle between the direction of mirror reflection of sunlight in the oil spill area and the detection direction of the sensor (θ ), and the calculation formula is as follows: ;

[0046] In the formula, is the angle between the direction of mirror reflection of sunlight in the oil spill area and the detection direction of the sensor, is the satellite zenith angle, is the relative azimuth angle between the sun and the satellite.

[0047] The OSS-Diff model framework provided by the embodiment of the present application is based on the denoising diffusion probability model (DDPM), and its core lies in the iterative denoising process, that is, gradually recovering the original segmentation mask (Y ) from the pure noise image (X ). This process includes two stages: forward diffusion process and reverse denoising process.

[0048] Forward diffusion process: This is a Markov chain process, which involves progressively diffusion forward into the original data ( Add Gaussian noise to gradually transform it into a pure noise image. At each time step ( The noise will be scheduled with a predefined variance ( This process is fixed and non-learnable; its main purpose is to provide training data for the inverse denoising process and to simulate image degradation under noise perturbation. Given the original image... any time step Noisy images It can be represented as: ;

[0049] In the formula, .

[0050] Inverse denoising process: This is the core of the model's learning process, designed to train a neural network. From images with noise ( Noise is gradually removed in the process, and the original data is eventually recovered. At each time step, the model learns to predict the image from the previous time step. The denoising function of the model can be expressed as: ;

[0051] In the formula, This is conditional information.

[0052] Figure 2 The overall architecture of the OSS-Diff model consists of three main components: a segment encoder, a conditional encoder, and a segment decoder, which are connected via time steps ( t The sinusoidal position encoding is used as a guide.

[0053] Segment encoder: This module is responsible for processing noisy input images ( The model extracts multi-scale feature representations. Its structure is similar to the downsampling path of the standard U-Net, consisting of a series of residual blocks and downsampling layers to capture spatial information at different scales. The input image size is uniformly 256 pixels, the mask has 4 channels, and the input image also has 4 channels. The initial convolutional layer maps the input channels to the initial feature dimension, which defaults to 64. The downsampling path gradually increases the number of feature channels while decreasing the spatial size of the feature map through a series of dimension multiplication factors (defaulting to (1, 2, 4, 8)). Each downsampling stage contains at least one residual block, and the default number of normalized groups for the residual blocks is 8.

[0054] Condition encoder: This module processes the condition information I of the original remote sensing image independently and performs deep feature extraction and enhancement. Its structure is similar to the down-sampling path of the segmentation encoder, but its core is to integrate a dual-branch frequency parser (DF-P) and a spectral-spatial attention condition encoding module (SS-A). The initial convolutional layer maps the original image channel number (4 bands) to the same initial feature dimension (default 64) as the segmentation encoder, using a 7x7 convolution kernel and a padding of 3. At each down-sampling stage, an attention module is followed by a custom condition processing module that integrates DF-P and SS-A. The feature map size of the condition processing module is consistent with the feature map size of the current down-sampling stage, and the feature dimension also matches the channel number of the current stage. The enhanced context information output by the condition encoder is effectively fused with the features of the segmentation encoder through an attention mechanism, accurately guiding the denoising process and improving the accuracy of oil spill identification.

[0055] Segmentation decoder: This module uses the features extracted by the segmentation encoder and the condition encoder to gradually reconstruct a fine oil spill segmentation mask. The decoder uses up-sampling and residual blocks, and receives multi-scale features from the encoder through a skip connection to preserve spatial details and achieve pixel-level accurate segmentation. The up-sampling path is symmetrical to the down-sampling path, with the spatial size of the feature map gradually increasing while the number of feature channels gradually decreasing. Each up-sampling stage also contains two residual blocks. The feature dimension of the skip connection is the sum of the input dimension and the condition feature dimension. Finally, a residual block and a convolutional layer map the processed features back to the required output mask channel number.

[0056] The dual-branch frequency parser provided by the embodiment of the present application includes: in order to further enhance the feature extraction capability of the model under complex glare conditions, the present application introduces a dual-branch frequency parser (DF-P) for high-frequency and low-frequency separation, as shown in Figure 3 The parser decomposes image features in the frequency domain, aiming to effectively separate and process high-frequency and low-frequency information in the image. High-frequency features usually correspond to details, edges and textures (such as oil-water boundaries) of the image, while low-frequency features represent the overall structure, background distribution and brightness changes of the image.

[0057] DF-P achieves its function by operating on input features in the frequency domain. Specifically, it first performs two-dimensional Fourier transform (FFT2) on the input feature map, converting it to the frequency domain. For a two-dimensional image , its Fourier transform can be represented as: ;

[0058] where is the image size, is the frequency domain coordinate.

[0059] Then, the DF-P acts on the high and low frequency components respectively using learnable frequency masks.

[0060] High frequency processing: In strong highlight regions, the positive contrast spill often exhibits high-brightness edges, and the DF-P can highlight these high-frequency edge information. By applying specific learnable weights to the high-frequency part of the Fourier-transformed feature map , the model can enhance the information of fine structures such as oil-water boundaries.

[0061] Low frequency processing: In weak highlight regions, the negative contrast characteristics between emulsified oil and unemulsified oil may exhibit subtle brightness differences, and the DF-P can capture these low-frequency features more clearly. By applying learnable weights to the low-frequency part of the Fourier-transformed feature map , the model can better analyze the background reflection distribution and overall brightness changes.

[0062] The feature map after frequency domain operation can be represented as: = ; Wherein, represents the high-frequency or low-frequency mask.

[0063] The DF-P generates a dynamic frequency domain attention map according to the Fourier transform real part splicing result of the noise image and the conditional image , so as to adaptively adjust the weights of high and low frequencies, further improving the segmentation accuracy and robustness of the diffusion model under different highlight conditions. After frequency domain operation, the feature map is converted back to the spatial domain through inverse Fourier transform (IFFT2) and integrated into subsequent processing. The inverse Fourier transform can be represented as: ;

[0064] The spectral-spatial attention conditional encoding module provided by the embodiment of the application comprises: In order to further improve the performance of oil spill segmentation, the application designs a spectral-spatial attention conditional encoding module (SS-A), as shown in Figure 4 . The module optimizes the extraction process of spectral-spatial features by fusing spectral and spatial attention mechanisms. In multispectral remote sensing data, spectral information and spatial information are complementary, and effective combination of the two is crucial for accurate target identification. The module first extracts basic features from the conditional image , wherein is the feature dimension, is the number of categories. Subsequently, these features are rearranged as to process different categories of features independently.

[0065] Spectral attention mechanism: allows the model to adaptively weight the contribution of different wavebands to the oil spill features, highlighting the reflectance characteristics of the oil spill in specific wavebands under different sun glint conditions. It is implemented through a sequential module consisting of adaptive average pooling, convolutional layers, ReLU activation, and Sigmoid activation. For input features , the calculation of spectral attention weights can be represented as: ;

[0066] where represents global average pooling, is a convolutional layer, ReLU is an activation function, and Sigmoid is an activation function.

[0067] Spatial attention mechanism: enhances the ability to capture local details, enabling the model to more accurately identify oil-water boundaries and texture changes within the oil spill. It is implemented through convolutional layers, ReLU activation, and Sigmoid activation. For input features , the calculation of spatial attention weights can be represented as: ;

[0068] Through this joint modeling, the spectral-spatial embedding module can generate more discriminative spectral-spatial feature representations. These optimized spectral-spatial features are then input as enhanced conditional information into the denoising process of the diffusion model, working in conjunction with the iterative optimization mechanism of the diffusion model. This synergistic effect ensures that the model can dynamically adapt to complex sun glint scenarios, significantly enhancing feature discrimination and improving the accuracy and robustness of oil spill segmentation under different sun glint intensities and observation geometries.

[0069] The loss function provided by the embodiment of the present application includes: the present application trains the neural network by minimizing the difference between the predicted noise and the actual noise. The loss function usually adopts mean square error (MSE) to ensure that the model can accurately estimate and remove noise.

[0070] When the model goal is to predict noise , the loss function can be represented as: ;

[0071] where is the actual added noise, is the noise predicted by the model.

[0072] The embodiment 2 of the present application provides a marine oil spill segmentation system based on a hyperspectral frequency constraint diffusion model, which comprises: a data acquisition module configured to acquire a multispectral remote sensing image; a preprocessing module configured to preprocess the multispectral remote sensing image to obtain Rayleigh-corrected reflectance data; a segmentation processing module comprising a hyperspectral frequency constraint diffusion model which is trained and integrated with a dual-branch frequency parser DF-P and a hyperspectral attention conditional encoding module SS-A, and is configured to receive the Rayleigh-corrected reflectance data output by the preprocessing module and generate an oil spill segmentation mask through an iterative denoising process; The segmentation processing module comprises: a segmentation encoding unit configured to extract multi-scale features from a noisy mask; a conditional encoding unit configured to process the Rayleigh-corrected reflectance data, the conditional encoding unit being integrated with the dual-branch frequency parser DF-P and the hyperspectral attention conditional encoding module SS-A; a segmentation decoding unit configured to fuse the output features of the segmentation encoding unit and the conditional encoding unit to reconstruct the oil spill segmentation mask; a result output module configured to output or store the oil spill segmentation mask.

[0073] According to the embodiments of the present application, the present application further provides a computer device, which comprises at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor implements the steps in any of the above method embodiments when executing the computer program.

[0074] The present application further provides a computer readable storage medium storing a computer program, wherein the computer program is executable by a processor to implement the steps in any of the above method embodiments.

[0075] The present application further provides an information data processing terminal, which is configured to provide a user input interface to implement the steps in any of the above method embodiments when executed on an electronic device, and the information data processing terminal is not limited to a mobile phone, a computer, or a switch.

[0076] The present application further provides a server, which is configured to provide a user input interface to implement the steps in any of the above method embodiments when executed on an electronic device.

[0077] The present application further provides a computer program product, which, when executed on an electronic device, enables the electronic device to implement the steps in any of the above method embodiments.

[0078] The integrated unit, if implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on such understanding, the present application can implement all or part of the processes in the above-mentioned embodiment methods by a computer program to instruct related hardware to complete, and the computer program can be stored in a computer-readable storage medium. When the processor executes the computer program, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium at least includes any entity or device capable of carrying the computer program code to the photographing device / terminal equipment, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc.

[0079] To further demonstrate the positive effects of the above embodiments, the present application based on the above technical solutions carries out the following experiments.

[0080] 1. Experimental setup and evaluation metrics: To comprehensively and quantitatively evaluate the proposed oil spill detection method, the present application uses a variety of metrics widely recognized in the fields of computer science and remote sensing to measure the performance of the model and the accuracy of oil spill detection. These metrics have been proven to be highly applicable in similar evaluations.

[0081] Precision measures the proportion of actual positive examples among those identified as positive by the model, and its calculation formula is: ;

[0082] Recall measures the proportion of all actual positive examples that are correctly identified by the model, and its calculation formula is: ;

[0083] F1 score is the harmonic mean of precision and recall, providing a balanced classification performance metric, especially suitable for class-imbalanced datasets. Its calculation formula is: ;

[0084] Wherein, TP (true positive) represents the number of correctly classified oil spill pixels, TN (true negative) represents the number of correctly classified non-oil spill pixels, FN (false negative) represents the number of oil spill pixels that were incorrectly classified as non-oil spill pixels, and FP (false positive) represents the number of non-oil spill pixels that were incorrectly classified as oil spill pixels.

[0085] 2. Oil spill detection under different solar glare reflection conditions; To evaluate the performance of the oil spill detection model under different solar flare conditions, including its ability to segment the oil spill area and distinguish emulsion states, this invention selects representative oil spill images from weak solar flare (WSG) and strong solar flare (SSG) environments, conducts experimental analysis, and evaluates its visual consistency and quantification accuracy.

[0086] To verify the model's visual accuracy, this invention uses... Figure 5 Compare the input image, the human visual interpretation reference, and the model inference results. For example... Figure 5 As shown in Figures (a), (b), (c), (d), and (e), under weak solar glare conditions, the model's detection results are highly consistent with the human reference and effectively distinguish the emulsified state; Figure 5 As shown in Figures (f) and (j), the model results also match the reference height under strong solar flare conditions, successfully segmenting the oil spill area.

[0087] Table 2. Accuracy of oil spill detection under different glare reflection conditions

[0088] To quantify and evaluate model performance, this invention analyzes precision, recall, and F1-Score, with results shown in Table 2. Under weak solar flare (WSG) conditions, OS-N's F1 score is 0.942 (precision 0.952, recall 0.933), indicating high accuracy and completeness. In contrast, under strong solar flare (SSG) conditions, OS-P's F1 score is 0.918 (precision 0.963, recall 0.883), with higher precision than OS-N but lower recall, indicating that the model's accuracy remains stable under strong solar flare conditions, but coverage is slightly compromised. Under weak solar glare reflection conditions, the F1 score of NEOS was 0.907 (precision 0.924, recall 0.892), while the F1 score of EOS was 0.862 (precision 0.843, recall 0.884). The precision of EOS was lower than that of NEOS, but the recall was higher, indicating that the model's accuracy in distinguishing emulsion states was slightly reduced, but the integrity was maintained well.

[0089] In the experimental verification part, the present application conducts a systematic evaluation based on data sets from five global oil spill-prone regions. Under weak sunlight conditions, the model achieves F1 scores of 0.907 and 0.862 for unemulsified oil and emulsified oil identification, respectively. Under strong sunlight conditions, the F1 score for detecting positive contrast oil spills is 0.918. Compared with existing advanced methods such as Unet, DeepLabV3+, and TransUnet, the average comprehensive performance is improved by 8.3%, especially in emulsified oil identification.

[0090] The OSS-Diff model proposed by the present application achieves end-to-end precise oil spill segmentation across different solar flare conditions through innovative architectural design. This method not only significantly improves segmentation performance, but also has important reference value for similar optical remote sensing feature segmentation tasks. This method is expected to provide more powerful technical support for marine environmental protection and oil spill emergency response.

[0091] Overall, the experimental results show that the model achieves high-precision oil spill pollution range detection under different solar flare conditions and high-precision oil spill state differentiation under weak solar flare reflection conditions. These results verify that the model proposed by the present application has excellent oil spill detection capability under different solar flare reflection conditions.

[0092] 3. Precision comparison with other advanced algorithms; To comprehensively evaluate the oil spill detection performance of the OSS-Diff model and other advanced models, the present application analyzes the F1-Score of Unet, DeepLabV3+, TransUnet, and OSS-Diff in detecting oil spills (OS-N, OS-P) of different contrasts under different solar flare intensities and detecting oil spills (NEOS, EOS) of different emulsion states under weak flare conditions. The experimental results are shown in Figure 6 .

[0093] To analyze the difference in oil spill detection accuracy under weak and strong flare conditions, the present application compares the F1 values of OS-N and OS-P to evaluate the robustness of the algorithm to flare interference. The results show that OSS-Diff achieves F1 scores of 0.942 and 0.918 on OS-N and OS-P, respectively, leading Unet (0.887, 0.875), DeepLabV3+ (0.916, 0.887), and TransUnet (0.931, 0.895). Specifically, Unet has low accuracy due to insufficient deep feature extraction, DeepLabV3+ has insufficient edge optimization despite its multi-scale modeling benefits for complex scenes, and TransUnet has strong global dependency capture capability but its detailed accuracy is limited by the deep information over-reliance of the Transformer. This result shows that OSS-Diff can effectively handle oil spill segmentation challenges under different flare intensities.

[0094] In order to evaluate the ability of the oil spill emulsion state to distinguish in the weak glare environment, the present application tests the sensitivity of the algorithm to the change of the emulsion state by comparing the F1 value of NEOS and EOS. The results show that OSS-Diff achieves 0.907 and 0.862 F1 scores on NEOS and EOS, respectively, which is better than Unet (0.860, 0.517), DeepLabV3+ (0.891, 0.574) and TransUnet (0.912, 0.676). Among them, Unet is limited in detail preservation due to its simple structure, DeepLabV3+ is better in context modeling but still loses points for small targets, and the overall performance of TransUnet is strong but the precision is low due to the decline of edge segmentation in the emulsion state. This result shows that OSS-Diff has more advantages in emulsion state differentiation.

[0095] Overall, the experimental results show that OSS-Diff is consistently leading in all oil spill categories, indicating that it overcomes the limitations of traditional deep learning algorithms. The advantage is attributed to the iterative denoising characteristics of the diffusion model, as well as the key feature fusion of the Fourier double-branch frequency analyzer and the spectral attention conditional encoder, ultimately achieving more accurate and robust oil spill segmentation in complex marine environments.

[0096] The above describes only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any modification, equivalent replacement and improvement made by those skilled in the art within the technical range disclosed by the present application, as long as it is within the spirit and principles of the present application, should be covered within the protection scope of the present application.

Claims

1. A method for marine oil spill segmentation based on a spatial frequency-constrained diffusion model, characterized in that, The method includes the following steps: S1, acquire multispectral remote sensing images; S2, preprocess the multispectral remote sensing image to obtain Rayleigh-corrected reflectance data; S3, the reflectivity data is input into the trained spatial frequency constraint diffusion model, and an oil spill segmentation mask is generated through iterative denoising; the spatial frequency constraint diffusion model adds noise to the segmentation mask in the forward diffusion process, and reconstructs the segmentation mask from the noise in the reverse denoising process. The reverse denoising process is implemented by a neural network, which includes: A segmentation encoder is used to extract features from a noisy mask. A conditional encoder is used to process the Rayleigh-corrected reflectivity data. The conditional encoder integrates a dual-branch frequency resolver DF-P and a spatial spectrum attention conditional coding module SS-A. A segmentation decoder is used to fuse the output features of the segmentation encoder and the conditional encoder, and reconstruct the oil spill segmentation mask.

2. The marine oil spill segmentation method based on the spatial spectrum frequency-constrained diffusion model according to claim 1, characterized in that, In step S1, the multispectral remote sensing image comes from HY-1C, HY-1D, HY-3A, Landsat-8 / 9, or Sentinel-2 satellite sensors.

3. The marine oil spill segmentation method based on the spatial spectrum frequency-constrained diffusion model according to claim 1, characterized in that, In step S2, the Rayleigh correction is calculated using the following formula: ; In the formula, The reflectance is the Rayleigh-corrected value. For the total radiance of a pixel, Rayleigh scattering radiance, Solar incident irradiance, The solar zenith angle; The preprocessed images, combined with manually labeled oil spill ground truth masks, form a dataset for model training, validation, and testing.

4. The marine oil spill segmentation method based on the spatial spectrum frequency-constrained diffusion model according to claim 1, characterized in that, In step S3, the forward diffusion process includes: gradually moving the data towards the original data. Add Gaussian noise, Gradually transforming into a pure noise image ; at each time step Noise is scheduled according to a predefined variance. Added to the data to provide training data for the inverse denoising process and to simulate image degradation under noise perturbation; given the original image any time step Noisy images Represented as: ; In the formula, For the original image, For time step Noisy images, , In time step Predefined variance scheduling It follows a Gaussian distribution. It is the identity matrix; The reverse denoising process includes: training a neural network. From noisy images The noise was gradually removed, and the original data was eventually recovered. At each time step, the model learns to predict the image from the previous time step. The denoising function of the model is expressed as: ; In the formula, For conditional information, For time step Noisy images, The image predicted by the model at the previous time step. For time steps, These are the parameters of the neural network.

5. The marine oil spill segmentation method based on the spatial spectrum frequency-constrained diffusion model according to claim 1, characterized in that, In step S3, the segmentation encoder is responsible for processing the noisy input image. Extracting multi-scale feature representations from them; It consists of a series of residual blocks and downsampling layers to capture spatial information at different scales; the initial convolutional layer maps the input channels to the initial feature dimension; The downsampling path gradually increases the number of feature channels through a series of dimensionality multiplication factors, while reducing the spatial size of the feature map; each downsampling stage contains at least one residual block; The conditional encoder is used to independently process the conditional information I of the original remote sensing image and perform depth feature extraction and enhancement. The conditional encoder integrates the dual-branch frequency parser DF-P and the spatial spectral attention conditional coding module SS-A; the initial convolutional layer maps the original image channel number to the same initial feature dimension as the segmentation encoder. In each downsampling stage, a custom conditional processing module is connected after the attention module. This module integrates DF-P and SS-A. The feature map size of the conditional processing module is consistent with the feature map size of the current downsampling stage, and the feature dimension matches the number of channels in the current stage. The enhanced context information output by the conditional encoder is effectively fused with the features of the segmentation encoder through the attention mechanism, thereby accurately guiding the denoising process. The segmentation decoder uses features extracted by the segmentation encoder and the conditional encoder to progressively reconstruct a fine oil spill segmentation mask; the segmentation decoder is constructed using upsampling and residual blocks, and receives multi-scale features from the encoder through skip connections; The upsampling path is symmetrical to the downsampling path, gradually increasing the spatial size of the feature map while decreasing the number of feature channels; each upsampling stage contains two residual blocks; the feature dimension of the skip connection is the sum of the input dimension and the conditional feature dimension; Finally, a residual block and a convolutional layer map the processed features back to the number of output mask channels.

6. The marine oil spill segmentation method based on the spatial spectrum frequency-constrained diffusion model according to claim 1, characterized in that, In step S3, the dual-branch frequency resolver DF-P decomposes the image features in the frequency domain, separating and processing the high-frequency and low-frequency information in the image; the high-frequency features correspond to the details, edges and textures of the image, while the low-frequency features represent the overall structure, background distribution and brightness changes of the image. The dual-branch frequency analyzer performs a two-dimensional Fourier transform on the input feature map, converting it to the frequency domain; for a two-dimensional image... Fourier transform Represented as: ; In the formula, For the input image features in the spatial domain, For the output characteristics in the frequency domain, Image size, Spatial domain coordinates, For frequency domain coordinates, The imaginary unit; The dual-branch frequency resolver uses a learnable frequency domain mask to apply to the high-frequency and low-frequency components respectively; High-frequency processing: Applying specific learnable weights to the high-frequency components of the Fourier transform feature map. ; Low-frequency processing: Learnable weights are applied to the low-frequency components of the Fourier transform feature map. ; The feature map after frequency domain operation is represented as follows: = ; in, This is the weighted feature map. This is the unweighted feature map. Represents a high-frequency or low-frequency mask; The dual-branch frequency resolver is based on the noise image and conditional images The Fourier transform real part concatenation result generates a dynamic frequency domain attention map, thereby adaptively adjusting the weights of high and low frequencies; after frequency domain operations, the feature map is transformed back to the spatial domain through inverse Fourier transform and incorporated into subsequent processing. Among them, the inverse Fourier transform Represented as: ; The characteristics of the spatial domain returned after the inverse transformation are... These are the features after frequency domain operations.

7. The marine oil spill segmentation method based on the spatial spectrum frequency-constrained diffusion model according to claim 1, characterized in that, In step S3, the spatial-spectral attention conditional coding module SS-A optimizes the extraction process of spatial-spectral features by fusing spectral and spatial attention mechanisms; it extracts features from the conditional image through a convolutional layer. Extracting basic features ,in, For the real number space, For batch size, For the number of channels, Image height, Image width, It is the feature dimension. It is the number of categories; the basic features are rearranged as ; Spectral attention mechanism: Implemented through a sequence module, including adaptive average pooling, convolutional layers, ReLU activation, and sigmoid activation; for input features Spectral attention weights The calculation formula is: ; In the formula, Indicates global average pooling. This is a convolutional layer, with ReLU as the activation function. Use the Sigmoid activation function; Spatial attention mechanism: implemented through convolutional layers, ReLU activation, and sigmoid activation; for input features Spatial attention weights The calculation formula is: ; Through joint modeling, optimized spatial-spectral features Subsequently, as enhanced conditional information, it is input into the denoising process of the diffusion model, working synergistically with the iterative optimization mechanism of the diffusion model.

8. The marine oil spill segmentation method based on the spatial spectrum frequency-constrained diffusion model according to claim 1, characterized in that, In step S3, the neural network is trained by minimizing the difference between the predicted noise and the actual noise; The loss function uses mean squared error when the model objective is to predict noise. When the loss function is: ; in, To predict the difference between the noise and the actual noise, The actual noise added. It is noise in the model prediction.

9. The marine oil spill segmentation method based on the spatial spectrum frequency-constrained diffusion model according to claim 1, characterized in that, In step S3, the oil spill segmentation mask includes one or more categories of unemulsified oil, emulsified oil, positive contrast oil spill, and negative contrast oil spill, used to achieve pixel-level segmentation under different solar glare conditions.

10. A marine oil spill separation system based on a spatial frequency-constrained diffusion model, characterized in that, This system is used to implement the marine oil spill segmentation method based on the spatial spectrum frequency-constrained diffusion model as described in any one of claims 1-9, and the system includes: The data acquisition module is used to acquire multispectral remote sensing images; The preprocessing module is used to preprocess the multispectral remote sensing image to obtain Rayleigh-corrected reflectance data; The segmentation processing module includes a trained spatial frequency constrained diffusion model that integrates a dual-branch frequency resolver DF-P and a spatial spectrum attention conditional coding module SS-A. It is used to receive Rayleigh-corrected reflectivity data output by the preprocessing module and generate an oil spill segmentation mask through an iterative denoising process. The segmentation processing module includes: The segmentation coding unit is configured to extract multi-scale features from a noisy mask; The conditional coding unit is configured to process the Rayleigh-corrected reflectance data, and the conditional coding unit integrates a dual-branch frequency resolver DF-P and a spatial spectrum attention conditional coding module SS-A. The segmentation decoding unit is configured to fuse the output features of the segmentation coding unit and the conditional coding unit to reconstruct the oil spill segmentation mask; The result output module is used to output or store the oil spill segmentation mask.

Citation Information

Patent Citations

  • A sea surface oil spill optical remote sensing detection method based on flare reflection difference

    CN109284709A

  • High-resolution remote sensing image semantic segmentation method based on diffusion model

    CN120673054A

  • Marine spatio-temporal data interpolation method based on remote sensing condition information diffusion

    CN120950852A

  • method of mobile search for hydrocarbon deposits and bottom objects, detection of signs of the emergence of hazardous phenomena on the sea shelf

    RU2015112546A