SAR-to-optical image generation method and device for fire monitoring
By using the Denoising Diffusion Probability Model (DDPM) and a hierarchical structure guidance mechanism, the problems of optical remote sensing occlusion and GAN training instability in the generation of optical images from SAR are solved, generating stable and reliable optical images to meet the high-efficiency and high-precision requirements of fire monitoring.
Patent Information
- Application Number
- CN202511375729.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-01-09
AI Technical Summary
Existing SAR-to-optical image generation technologies for fire monitoring suffer from problems such as optical remote sensing being obscured by smoke and clouds, difficulty in intuitive interpretation of SAR images, instability in GAN training, high computational cost of diffusion models, and discontinuous block generation structures, thus failing to meet the needs of fire monitoring.
The Denoising Diffusion Probability Model (DDPM) combined with a hierarchical structure guidance mechanism is adopted. Optical features are introduced during the diffusion process through a cross-domain translator. The normalized burn index is calculated using near-infrared and short-wave infrared bands extracted from multispectral images to generate optical style images. By combining regional weighted reconstruction loss and global consistency loss, the spectral consistency and overall structural integrity of the fire area are ensured.
The generated optical images are stable and reliable under smoke and cloud cover conditions, and the fire area identification is refined, meeting the visualization needs of fire monitoring, reducing the computational burden, improving the continuity and accuracy of fire monitoring, and are suitable for large-scale fire scenarios.
Smart Images

Figure CN121304716A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of remote sensing image processing, and particularly relates to a SAR-to-optical image generation method and device for fire monitoring. BACKGROUND
[0002] Fire, as a common major disaster, has strong suddenness and fast spreading speed, which poses a serious threat to the ecological environment and human life and property safety. Efficient and accurate fire monitoring and post-disaster assessment are key links to reduce fire hazards, and remote sensing technology has become a core technical means in the field of fire monitoring due to its wide coverage, short observation period and no limitation by ground conditions. In current remote sensing monitoring technology, synthetic aperture radar (SAR) and optical remote sensing are two mainstream technologies, each of which has its own characteristics in fire monitoring scenarios, but also has technical limitations that cannot be broken through alone. At the same time, SAR-to-optical image generation technology still faces many problems to be solved in adapting to fire monitoring needs.
[0003] In related technologies, SAR remote sensing technology is based on microwave imaging principle and has all-weather and all-day working ability, which is not affected by weather and light conditions such as cloud, fog, smoke, day and night, etc. This core advantage makes it irreplaceable in fire monitoring, especially suitable for scenes where a large amount of smoke is blocked during fire occurrence and traditional optical remote sensing cannot penetrate, and it can stably obtain the ground structure information of the fire area to provide basic data support for fire location and fire range delineation. However, the imaging mechanism of SAR image also leads to significant limitations. On the one hand, SAR image presents the distribution of backscattering coefficients of ground objects, rather than the intuitive optical reflection characteristics of ground objects, and there is obvious speckle noise in the image, and ground object interpretation depends on professional knowledge, which is difficult to directly use for visual analysis and rapid judgment of fire monitoring. On the other hand, SAR image lacks the spectral information possessed by optical image, and cannot intuitively distinguish the burned and unburned areas of fire through spectral features such as vegetation greenness and soil brightness like optical image, which makes it difficult to provide sufficient information support in fine application scenarios such as fire disaster grade assessment and post-disaster ecological recovery monitoring.
[0004] In the related art, optical remote sensing technology captures visible light, near-infrared and other band reflection information of ground objects to generate optical images with strong intuitiveness and rich spectral information. In fire monitoring, optical images can quickly identify the burned area of the fire by the characteristics of the decrease in the near-infrared reflectance of vegetation and the increase in the short-wave infrared reflectance. At the same time, with the help of multispectral information, the intensity of the fire can be evaluated, and the types of the damaged vegetation can be counted, and other refined analyses can be performed. Optical images are commonly used data types in fire monitoring and post-disaster assessment. However, the limitations of optical remote sensing technology are also very prominent. The imaging quality of optical remote sensing technology is heavily dependent on weather and lighting conditions. Smoke and clouds that are common during a fire can cause serious obstruction to optical signals, resulting in data loss and blurring of optical images, and making it impossible to continuously and stably obtain information about the fire area. In addition, the revisit period of some high-resolution optical satellites is long, which makes it difficult to meet the real-time monitoring needs during the fire spreading process, and limits the application in the dynamic tracking of sudden fires.
[0005] To solve the problems of difficult interpretation of SAR images and easy obstruction of optical images, SAR-to-optical image generation technology uses artificial intelligence methods such as deep learning to establish a mapping relationship between SAR data and optical data, and converts difficult-to-interpret SAR images into intuitive optical-style images, thereby combining the all-weather advantage of SAR with the visualization and spectralization advantage of optical images, and providing a new technical path for fire monitoring. Although the current SAR-to-optical image generation technology has made some progress, it still faces the following technical difficulties when adapting to the specific scenario of fire monitoring:
[0006] First, existing generation methods mostly use general image translation frameworks such as generative adversarial networks (GAN), and do not introduce domain prior information for the fire monitoring scenario, which cannot guarantee the generation quality of the fire area. The generated optical images may have problems such as distortion of spectral features in the fire area (e.g., confusion between burned vegetation and unburned vegetation colors), and blurred boundaries, which makes it difficult to meet the needs of fire area identification and disaster assessment.
[0007] Second, traditional GAN models are prone to mode collapse and unstable training during the training process, resulting in defects such as artifacts and discontinuous structures in the generated images. Although the denoising diffusion probability model (DDPM) that has emerged in recent years can improve the stability of generation, the computational overhead of the full-process conditional constraints is large, and it is difficult to balance the generation efficiency and result quality when processing large-scale SAR images commonly used in fire monitoring.
[0008] Third, fire monitoring often involves a large area (such as forest fires, grassland fires), and it is necessary to process large-scale SAR images. The prior art usually adopts the method of generating and splicing in blocks, but the edge of the block is not connected and the global structure is broken in the blocking process, which leads to the problem that the generated large-scale optical image cannot fully reflect the spatial distribution characteristics of the fire area, affecting the accuracy of the overall disaster analysis. SUMMARY
[0009] To this end, the present application provides a SAR-to-optical image generation method and device for fire monitoring, which solves the problems in the prior art that optical remote sensing is blocked by smoke clouds, SAR images are difficult to interpret intuitively, GAN training is unstable and has artifacts, diffusion model calculation overhead is large, lacks domain prior, and the generated structure is discontinuous, which cannot meet the needs of fire monitoring.
[0010] In order to achieve the above purpose, the present application provides the following technical scheme: a SAR-to-optical image generation method for fire monitoring, comprising the following steps:
[0011] SAR image acquisition and preprocessing: acquiring synthetic aperture radar (SAR) satellite images in a fire monitoring scene, performing radiometric calibration and geometric correction on the original SAR satellite image data; then performing normalization and difference operation on the polarization channels of the SAR satellite image, and combining the operation results to form a pseudo-RGB image; the large-scale satellite image is cut into fixed-size image blocks, and a preset overlap strategy is used for processing;
[0012] Spectrum feature extraction and fire area guidance: acquiring multispectral image data of the target area, extracting near-infrared (NIR) and short-wave infrared (SWIR) bands from the multispectral image data; calculating the normalized burn index (NBR) based on the extracted near-infrared (NIR) and short-wave infrared (SWIR) bands, and using the normalized burn index (NBR) to distinguish the burned area and the unaffected area; performing threshold segmentation on the calculated NBR image to obtain a binary fire mask, and distinguishing the burned area and the background area through the fire mask during model training; in the model training stage, a regionally weighted reconstruction loss function is used to calculate the error between the generated optical image and the real optical image, and a global consistency loss is introduced to calculate the entire multispectral image;
[0013] Diffusion model generation and structure guidance: taking the denoising diffusion probability model (DDPM) as the image generation backbone, embedding a cross-domain translator in the diffusion model, and combining a hierarchical structure guidance mechanism to complete the generation of SAR satellite images to optical images; the cross-domain translator is introduced at a certain time in the diffusion process, and unconditional denoising and conditional denoising are respectively performed in the reverse sampling stage according to different time steps; the hierarchical structure guidance mechanism adopts a double-channel reconstruction strategy, respectively generates a structure guidance image through a global low-resolution path and generates a local reconstruction result through a local high-resolution path, defines a structure sensitive area and constructs a spatial adaptive weight function, and based on the spatial adaptive weight function, the structure guidance image and the local reconstruction result are weighted and fused to obtain the final image reconstruction result;
[0014] Output optical style image: according to the final image reconstruction result, for local areas, generate small image results to show the spectral consistency and texture details of the fire-burned area; in a scene with a size exceeding a certain scale, obtain a complete large-scale optical style image through splicing.
[0015] As an optimal solution of the SAR-to-optical image generation method for fire monitoring, in the SAR image acquisition and preprocessing step:
[0016] The SAR satellite image adopts C-band Sentinel-1 data source and L-band alos data source; wherein, the C-band Sentinel-1 data source includes two polarization channels of vertical emission / vertical reception VV and vertical emission / horizontal reception VH, and the L-band alos data source includes two polarization channels of horizontal emission / horizontal reception HH and horizontal emission / vertical reception HV;
[0017] When the polarization channels of the SAR satellite image are normalized and differenced, and the operation results are combined to form a pseudo-RGB image:
[0018] The pseudo-RGB image corresponding to the C-band SAR image, the red channel is represented by the absolute value difference of VV and VH, the green channel is represented by VH, and the blue channel is represented by VV;
[0019] The pseudo-RGB image corresponding to the L-band SAR image, the red channel is represented by the absolute value difference of HH and HV, the green channel is represented by HV, and the blue channel is represented by HH; the size of the cropped image block is set to 256x256 pixels, and the preset overlap strategy is 64-pixel overlap.
[0020] As an optimal solution of the SAR-to-optical image generation method for fire monitoring, in the spectral feature extraction and fire area guidance step, when calculating the normalized burn index NBR:
[0021] The result of subtracting the pixel value of the short-wave infrared band from the pixel value of the near-infrared band is divided by the result of adding the pixel value of the short-wave infrared band to the pixel value of the near-infrared band;
[0022] When threshold segmentation is performed on the NBR image, an empirical threshold is set in advance. When the NBR value is less than the empirical threshold, the corresponding pixel point is determined as a fire area, and otherwise, the corresponding pixel point is determined as a non-fire area, so as to obtain a binary fire mask with the same size as the original image. The pixel with a value of 1 in the mask represents a fire area, and the pixel with a value of 0 represents a non-fire area.
[0023] When the error between the generated optical image and the real optical image is calculated by using the region-weighted reconstruction loss function, the fire mask is multiplied element by element with the difference between the generated optical image and the real optical image, and then the square norm of the obtained result is calculated.
[0024] The spectral guidance loss function is obtained by weighting and summing the region-weighted reconstruction loss and the global consistency loss by using a preset weight parameter.
[0025] As an optimal solution of the SAR-to-optical image generation method for fire monitoring, in the diffusion model generation and structure guidance step:
[0026] The forward diffusion process of the denoising diffusion probability model DDPM follows the first-order Markov chain characteristic. The noisy image at each time is obtained by transferring the noisy image at the previous time through a set Gaussian distribution. The mean of the Gaussian distribution is related to the noisy image at the previous time and a preset variance scheduling parameter, and the variance is the preset variance scheduling parameter.
[0027] After a number of iterations, the image data gradually approaches a standard Gaussian distribution. The edge distribution of the noisy image at any time with respect to the initial clean image follows a Gaussian distribution. The mean of the Gaussian distribution is related to the cumulative product of the initial clean image and the multi-step variance scheduling parameter, and the variance is related to the result of 1 minus the cumulative product.
[0028] In the reverse denoising process, the posterior distribution of the noisy image at the previous time with respect to the noisy image at the current time and the initial clean image also follows a Gaussian distribution. The diffusion model predicts the noise added to the image to approximate the mean of the Gaussian distribution through the denoising network.
[0029] The target of the diffusion model training is to minimize the mean square error between the real noise and the predicted noise of the denoising network. The mean square error is obtained by performing expectation operation on the initial clean image, the added noise, and different time steps.
[0030] As an optimal solution of the SAR-to-optical image generation method for fire monitoring, in the diffusion model generation and structure guidance step:
[0031] The cross-domain translator receives a diffusion state and a pseudo- RGB image constructed by a SAR polarization channel at a set time, and generates a cross-domain feature after processing;
[0032] In the reverse sampling stage, when the time step is greater than the set time, the diffusion model completes the denoising inversion only through the diffusion network itself without relying on additional conditions;
[0033] When the time step is less than or equal to the set time, the diffusion model generates a cross-domain feature as an additional input for conditional denoising.
[0034] A regularization loss is added to the training target, which includes two parts: one is the mean square error of noise prediction to ensure the denoising accuracy under the condition of input, and the other is the square of the Frobenius norm of the parameter gradient of the cross-domain translator multiplied by a preset regularization coefficient to suppress overfitting of the diffusion model and ensure stable training.
[0035] As an optimal solution of the SAR-to-optical image generation method for fire monitoring, in the diffusion model generation and structure guiding step, when processing the global low-resolution path of the hierarchical structure guiding mechanism:
[0036] First, the input high-resolution image is down-sampled by a bilinear method to obtain a low-resolution image; then the low-resolution image is input into a backbone network combined with the cross-domain translator to generate a rough reconstruction result; finally, the rough reconstruction result is up-sampled by a bilinear method to obtain a structure guiding image;
[0037] When processing the local high-resolution path of the hierarchical structure guiding mechanism:
[0038] First, the original high-resolution image is divided into multiple overlapping image blocks; each image block is independently processed by inputting into the same backbone network to obtain corresponding local features;
[0039] Then, all local features are aggregated to form a local reconstruction result through a spatial splicing operation; a structure sensitive area is defined, and the structure sensitive area includes all pixel points with a horizontal or vertical Manhattan distance to the nearest image block edge less than or equal to a preset boundary threshold;
[0040] A spatially adaptive weight function is constructed based on the structure sensitive area to smoothly transition the weight in the neighborhood of the image block boundary; finally, the image reconstruction result is output by pixel-wise weighted fusion of the structure guiding image and the local reconstruction result through the spatially adaptive weight function, and the sum of the weighted proportion of the structure guiding image and the weighted proportion of the local reconstruction result is 1.
[0041] The application also provides a SAR-to-optical image generation device for fire monitoring, comprising:
[0042] The SAR image acquisition and preprocessing module is configured to acquire synthetic aperture radar (SAR) satellite images in a fire monitoring scene, perform radiometric calibration and geometric correction on original SAR satellite image data, normalize and perform differential operation on a polarization channel of the SAR satellite images, combine the operation results to form a pseudo RGB image, and crop a large-format satellite image into a fixed-size image block and process the image block using a preset overlap strategy.
[0043] The spectral feature extraction and fire region guiding module is configured to acquire multispectral image data of a target region, extract a near-infrared (NIR) band and a short-wave infrared (SWIR) band from the multispectral image data, calculate a normalized burn ratio (NBR) based on the extracted NIR band and SWIR band, distinguish a burned area and an unaffected area using the NBR, perform threshold segmentation on the calculated NBR image to obtain a binary fire mask, and distinguish the burned area and the background area in the model training process through the fire mask.
[0044] The diffusion model generation and structure guiding module is configured to use a denoising diffusion probability model (DDPM) as an image generation backbone, embed a cross-domain translator in the diffusion model, and complete generation of SAR satellite images to optical images in combination with a hierarchical structure guiding mechanism.
[0045] The optical style image output module is configured to generate a small image result for a local region based on the final image reconstruction result to show spectral consistency and texture details of a fire burned area, and obtain a complete large-scale optical style image through splicing in a scene exceeding a set scale.
[0046] As a preferred solution of the SAR-to-optical image generation device for fire monitoring, in the SAR image acquisition and preprocessing module:
[0047] The SAR satellite image adopts a C-band Sentinel-1 data source and an L-band alos data source; the C-band Sentinel-1 data source includes two polarization channels of vertical emission / vertical reception VV and vertical emission / horizontal reception VH, and the L-band alos data source includes two polarization channels of horizontal emission / horizontal reception HH and horizontal emission / vertical reception HV;
[0048] The pseudo RGB image corresponding to the C-band SAR image, the red channel is represented by the absolute value difference between VV and VH, the green channel is represented by VH, and the blue channel is represented by VV;
[0049] The pseudo RGB image corresponding to the L-band SAR image, the red channel is represented by the absolute value difference between HH and HV, the green channel is represented by HV, and the blue channel is represented by HH; the size of the cropped image block is set to 256*256 pixels, and the preset overlap strategy is 64-pixel overlap.
[0050] As an optimal scheme of the SAR-to-optical image generation device for fire monitoring, in the spectral feature extraction and fire region guiding module:
[0051] The result of subtracting the short-wave infrared band pixel value from the near-infrared band pixel value is divided by the result of adding the near-infrared band pixel value and the short-wave infrared band pixel value;
[0052] When performing threshold segmentation on the NBR image, an empirical threshold is preset, when the NBR value is less than the empirical threshold, the corresponding pixel point is determined as a fire region, otherwise, it is a non-fire region, so as to obtain a binary fire mask with the same size as the original image, the pixel with a value of 1 in the mask represents a fire region, and the pixel with a value of 0 represents a non-fire region;
[0053] When calculating the error between the generated optical image and the real optical image by using the region-weighted reconstruction loss function, the fire mask is multiplied element by element with the difference between the generated optical image and the real optical image, and then the square norm of the obtained result is calculated;
[0054] The spectral guiding loss function is obtained by weighting and summing the region-weighted reconstruction loss and the global consistency loss by using a preset weight parameter.
[0055] As an optimal scheme of the SAR-to-optical image generation device for fire monitoring, in the diffusion model generation and structure guiding module:
[0056] The forward diffusion process of the denoising diffusion probability model DDPM follows the first-order Markov chain characteristic, the noisy image at each time is obtained by transferring the noisy image at the previous time through a set Gaussian distribution, the mean of the Gaussian distribution is related to the noisy image at the previous time and a preset variance scheduling parameter, and the variance is the preset variance scheduling parameter.
[0057] After several steps of iteration, the image data gradually approaches a standard Gaussian distribution; the edge distribution of the noisy image at any moment with respect to the initial clean image obeys a Gaussian distribution, the mean value of the Gaussian distribution is related to the cumulative product of the initial clean image and the multi-step variance scheduling parameter, and the variance is related to the result of 1 minus the cumulative product;
[0058] In the reverse denoising process, the posterior distribution of the noisy image at the previous moment with respect to the noisy image at the current moment and the initial clean image also obeys a Gaussian distribution, and the diffusion model adds noise to the image through the denoising network to approximate the mean value of the Gaussian distribution;
[0059] The target of the diffusion model training is to minimize the mean square error between the real noise and the denoising network predicted noise, and the mean square error is obtained by expectation operation on the initial clean image, added noise and different time steps.
[0060] As an optimal solution of the SAR-to-optical image generation device for fire monitoring, the diffusion model generates and the structure guiding module includes:
[0061] The cross-domain translator receives the diffusion state and the pseudo-RGB image constructed by the SAR polarization channel at a set moment, and generates cross-domain features after processing;
[0062] In the reverse sampling stage, when the time step is greater than the set moment, the diffusion model does not rely on additional conditions and only completes denoising inversion through the diffusion network itself;
[0063] When the time step is less than or equal to the set moment, the diffusion model takes the generated cross-domain features as additional input and performs conditional denoising.
[0064] The regularization loss is added to the training target, which includes two parts, one part is the mean square error of noise prediction to ensure the denoising accuracy under the condition of input, and the other part is the square of the Frobenius norm of the cross-domain translator parameter gradient, multiplied by a preset regularization coefficient to suppress overfitting of the diffusion model and ensure stable training.
[0065] As an optimal solution of the SAR-to-optical image generation device for fire monitoring, the diffusion model generates and the structure guiding module includes:
[0066] The bilinear downsampling method is used for the input high-resolution image to obtain a low-resolution image; then the low-resolution image is input into the backbone network combined with the cross-domain translator to generate a rough reconstruction result; finally, the bilinear upsampling is performed on the rough reconstruction result to obtain a structure guiding image.
[0067] As an optimal solution of the SAR-to-optical image generation device for fire monitoring, the diffusion model generates and the structure guiding module includes:
[0068] The original high-resolution image is divided into a plurality of overlapping image blocks; each image block is independently processed by being input into the same backbone network respectively to obtain corresponding local features;
[0069] All the local features are aggregated by a spatial splicing operation to form a local reconstruction result; a structure-sensitive region is defined, and the structure-sensitive region includes all pixel points with a horizontal or vertical Manhattan distance to the edge of the nearest image block being less than or equal to a preset boundary threshold;
[0070] A spatial adaptive weight function is constructed based on the structure-sensitive region to make the weight in the neighborhood of the image block boundary smoothly transition; finally, the image reconstruction result is output by performing pixel-by-pixel weighted fusion on the structure guide image and the local reconstruction result through the spatial adaptive weight function, and the sum of the weighted proportion of the structure guide image and the weighted proportion of the local reconstruction result by the spatial adaptive weight function is 1.
[0071] The present application has the following advantages:
[0072] Firstly, on the one hand, the present application relies on the all-weather and anti-smoke / cloud sheltering characteristics of SAR images as the basic data source for generating optical images, ensuring that scene data can be stably obtained even if there is a large amount of smoke and cloud cover when a fire occurs; on the other hand, through SAR-to-optical image generation technology, SAR backscatter information that is difficult to intuitively interpret is converted into images with optical visual features, without relying on optical remote sensing data that is easily blocked, and clear and directly interpretable visual results can be provided for fire monitoring, greatly improving the continuity and reliability of fire monitoring under extreme weather conditions, and being particularly suitable for forest, grassland and other fire scenes prone to large-scale smoke blocking.
[0073] Secondly, by means of the near-infrared and short-wave infrared bands extracted from multispectral images, the NBR index is calculated and a fire mask is generated to clearly define the range of the fire area; then, through a regionally weighted reconstruction loss function, higher optimization weights are given to the pixels in the fire area to ensure that the spectral difference between the burned area and the unburned area in the generated optical image is clear; at the same time, the global consistency loss constraint is combined with the structure of the non-fire area to ensure that the entire image meets the needs of fine identification of the fire area and guarantees the spatial structure integrity of the non-fire area, thereby providing accurate data support for fine monitoring tasks such as fire intensity classification and burned area statistics.
[0074] Third, only introduce a lightweight cross-domain translator at the key time step of the diffusion process, inject SAR pseudo-RGB features, avoid the additional computational burden brought by full-process conditional constraints, and greatly reduce the model inference time while ensuring the fidelity of the generated image details. Compared with the traditional block generation method, this design can more efficiently process large-format SAR images commonly used in fire monitoring, without sacrificing image size or the number of blocks due to computational efficiency issues, meeting the rapid monitoring needs of large-area fires and gaining valuable time for fire spread trend analysis.
[0075] Fourth, the structure guide map generated by the global low-resolution path can encode the topological constraints of the scene level, ensuring the global structure coherence of large-scale images. The local high-resolution path addresses the problem of easy detail loss in image blocks, preserving the fine spectral and texture features of fire areas. Through the fusion of structure-sensitive regions and spatial adaptive weights, the weights of the image block edge neighborhood are smoothly transitioned, eliminating the obvious edge gaps or texture misplacement caused by block stitching, and generating a complete large-scale optical image that can clearly present local fire details and accurately reflect the global spatial distribution of fire areas, providing a comprehensive spatial reference for overall fire disaster assessment and rescue force deployment.
[0076] Fifth, it can generate small images and stitched large images respectively, which fits the actual application scenarios of fire monitoring, avoids the problem of generated results deviating from monitoring needs due to the lack of field adaptability of general methods, and makes the technical solution directly applicable to real-time fire monitoring, post-disaster rapid assessment, and other practical work, reducing the operation threshold and analysis difficulty of grassroots monitoring personnel. BRIEF DESCRIPTION OF DRAWINGS
[0077] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only exemplary, and for those skilled in the art, other drawings can be derived from the provided drawings without creative labor.
[0078] The structures, proportions, sizes, etc. shown in the specification are only used to cooperate with the content disclosed in the specification, to be understood and read by those skilled in the art, and do not define the limiting conditions for the implementation of the present application, so they do not have technical significance. Any modification of structure, change of proportion relationship or adjustment of size, without affecting the effect and purpose that the present application can produce, should still fall within the scope of the technical content disclosed by the present application.
[0079] Figure 1 The flowchart of the SAR-to-optical image generation method for fire monitoring provided in the embodiments of the present application;
[0080] Figure 2 The technical roadmap of the SAR-to-optical image generation method for fire monitoring provided in the embodiments of the present application is provided;
[0081] Figure 3 The SAR image preprocessing diagram in the SAR-to-optical image generation method for fire monitoring provided in the embodiments of the present application is provided;
[0082] Figure 4 The generation of the fire spectrum map and the true value map for comparison is provided in the embodiments of the present application;
[0083] Figure 5 The device architecture schematic diagram of the SAR-to-optical image generation method for fire monitoring provided in the embodiments of the present application is provided. DETAILED DESCRIPTION
[0084] The embodiments of the present application are described below by specific embodiments, and those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in the specification. Obviously, the described embodiments are part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0085] Embodiment 1
[0086] Referring to Figure 1 and Figure 2 , the embodiment 1 of the present application provides a SAR-to-optical image generation method for fire monitoring, comprising the following steps:
[0087] S1, SAR image acquisition and preprocessing: acquiring synthetic aperture radar (SAR) satellite images in a fire monitoring scene, performing radiometric calibration and geometric correction on the original SAR satellite image data; then performing normalization and difference operation on the polarization channels of the SAR satellite images, combining the operation results to form a pseudo-RGB image; cutting the large-format satellite images into fixed-size image blocks, and processing them with a preset overlap strategy;
[0088] S2, spectral feature extraction and fire region guidance: obtain multispectral image data of the target region, extract near-infrared band NIR and short-wave infrared band SWIR from the multispectral image data; calculate normalized burn index NBR based on the extracted near-infrared band NIR and short-wave infrared band SWIR, and use the normalized burn index NBR to distinguish the burned area and the unaffected area; threshold segmentation is performed on the calculated NBR image to obtain a binary fire mask, and the burned area and the background area are distinguished by the fire mask in the model training process; in the model training stage, a region weighted reconstruction loss function is used to calculate the error between the generated optical image and the real optical image, and a global consistency loss is introduced to calculate the entire multispectral image;
[0089] S3, diffusion model generation and structure guidance: taking the denoising diffusion probability model DDPM as the image generation backbone, embedding a cross-domain translator in the diffusion model, and combining a hierarchical structure guidance mechanism to complete the generation of SAR satellite images to optical images; the cross-domain translator is introduced at a set time in the diffusion process, and unconditional denoising and conditional denoising are respectively performed in the reverse sampling stage according to different time steps; the hierarchical structure guidance mechanism adopts a double-channel reconstruction strategy, respectively generates a structure guidance image through a global low-resolution path and generates a local reconstruction result through a local high-resolution path, defines a structure sensitive area and constructs a spatial adaptive weight function, and based on the spatial adaptive weight function, the structure guidance image and the local reconstruction result are weighted and fused to obtain the final image reconstruction result;
[0090] S4, output optical style image: according to the final image reconstruction result, for local areas, a small image result is generated to show the spectral consistency and texture details of the fire burned area; in a scene with a size exceeding a set scale, a complete large-scale optical style image is obtained through splicing.
[0091] In the SAR image acquisition and preprocessing step S1 in the embodiment:
[0092] The SAR satellite image uses C-band Sentinel-1 data source and L-band alos data source; wherein, the C-band Sentinel-1 data source includes two polarization channels of vertical emission / vertical reception VV and vertical emission / horizontal reception VH, and the L-band alos data source includes two polarization channels of horizontal emission / horizontal reception HH and horizontal emission / vertical reception HV;
[0093] When the polarization channels of the SAR satellite image are normalized and differenced, and the operation results are combined to form a pseudo-RGB image: for the pseudo-RGB image corresponding to the C-band SAR image, the red channel is represented by the absolute value difference of VV and VH, the green channel is represented by VH, and the blue channel is represented by VV; for the pseudo-RGB image corresponding to the L-band SAR image, the red channel is represented by the absolute value difference of HH and HV, the green channel is represented by HV, and the blue channel is represented by HH; the size of the cropped image block is set to 256x256 pixels, and the preset overlap strategy is 64-pixel overlap.
[0094] Specifically, the synthetic aperture radar (SAR) satellite image in the fire monitoring scene is obtained, mainly using C-band Sentinel-1 data sources, including vertical transmission / vertical reception (VV) and vertical transmission / horizontal reception (VH) two polarization channels, and L-band alos data sources, including horizontal transmission / horizontal reception (HH) and horizontal transmission / vertical reception (HV) two polarization channels. Due to the speckle noise and complex backscattering characteristics of SAR images, direct use for optical image generation will lead to interpretation difficulties, so it needs to be preprocessed. The specific steps include: first, the original SAR data is radiometrically calibrated and geometrically corrected to eliminate system errors and ensure the consistency of spatial coordinates; second, as shown in Figure 3 , the VV and VH polarization channels are normalized and differenced, and combined to form a pseudo-RGB image. For the C-band, |VV-VH| represents the red channel, VH represents the green channel, and VV represents the blue channel. For the L-band, |HH-HV| represents the red channel, HV represents the green channel, and HH represents the blue channel, thereby highlighting the polarization difference and enhancing the spatial texture features; finally, for large-area satellite images, they are cropped into fixed-size image blocks (256x256 pixels), and a certain overlap strategy is used, such as 64-pixel overlap, to ensure the integrity of local details and the diversity of training samples. Through the above preprocessing steps, the speckle noise problem in the SAR image can be effectively alleviated, and high-quality input data can be provided for subsequent spectral feature guidance and diffusion generation.
[0095] In this embodiment, in the spectral feature extraction and fire region guidance step S2, in order to ensure that the generated optical style image has higher spectral consistency and discrimination ability in the fire-burned area, a spectral feature extraction and fire region guidance mechanism is proposed. As shown in Figure 2 (a), the core idea is to use the Normalized Burn Ratio (NBR) widely used in fire monitoring as a spectral prior, which is introduced into the diffusion generation process to constrain the learning of the model in the fire-sensitive area.
[0096] wherein, when calculating the normalized burn ratio NBR:
[0097] The result of subtracting the short-wave infrared band pixel value from the near-infrared band pixel value is divided by the result of adding the near-infrared band pixel value and the short-wave infrared band pixel value, and the specific formula is as follows:
[0098]
[0099] In the formula, I NIR represents the near-infrared band pixel value, I SWIR represents the short-wave infrared band pixel value; the NBR index is highly sensitive to the influence of fire, and can effectively distinguish between burned areas and unaffected areas.
[0100] Wherein, when the NBR image is threshold segmented, an empirical threshold is set in advance, when the NBR value is less than the empirical threshold, the corresponding pixel point is determined as a fire area, otherwise it is a non-fire area, so as to obtain a binary fire mask with the same size as the original image, the pixel with a value of 1 in the mask represents a fire area, and the pixel with a value of 0 represents a non-fire area. When the NBR image is threshold segmented, the empirical threshold δ is set to obtain the binary fire mask M burn The formula is:
[0101] M burn = 1 (NBR < δ)
[0102] In the formula, M burn is a binary matrix with the same size as the original image, the pixel with a value of 1 represents a fire area, and the pixel with a value of 0 represents a non-fire area. Through the mask, the burned area and the background area can be distinguished in the training process.
[0103] Wherein, when the region weighted reconstruction loss function is used to calculate the error between the generated optical image and the real optical image, the fire mask is multiplied with the difference between the generated optical image and the real optical image element by element, and then the square norm of the obtained result is calculated; the spectral guided loss function is obtained by weighting and summing the region weighted reconstruction loss and the global consistency loss with the preset weight parameters.
[0104] Specifically, in the model training stage, the region weighted reconstruction loss function is used to calculate the error between the generated optical image and the real optical image. For the pixels in the fire area, a higher weight is given to ensure the spectral consistency; for the non-fire area, the global consistency constraint is maintained, so as to balance the local accuracy and the overall quality. Wherein, the region weighted reconstruction loss function is defined as:
[0105]
[0106] In the formula, represents the generated optical image, I optrepresents a reference optical image, represents an element-wise multiplication; the spectral guidance loss function constructed by combining the region weighted reconstruction loss function and the global consistency loss function is:
[0107] L SFG =λ1·L burn +λ2·L global
[0108] In the formula, λ1 and λ2 are weight parameters for balancing the local constraint of the burned area and the global consistency of the overall image, L global is the global consistency loss. Through the above spectral feature extraction and fire area guidance, the application can optimize the spectral recovery effect of the fire area in the generation process, ensure that the details of the burned area are more real and consistent, and at the same time maintain the naturalness and structural integrity of the overall image, providing reliable support for subsequent fire monitoring and post-disaster assessment.
[0109] In this embodiment, in the diffusion model generation and structure guidance step S3, the generated backbone adopts a denoising diffusion probability model (DDPM), and a cross-domain translator is embedded therein, and at the same time, hierarchical structure guidance is combined to improve the stability, structural continuity and detail fidelity of SAR to optical image translation in the fire scene. Unlike traditional generation models, the core idea of the denoising diffusion probability model is to perturb a clean image x0 into a pure Gaussian noise distribution p(x0) by gradually injecting Gaussian noise into it, and to learn how to reverse this process.
[0110] Wherein, the forward diffusion process of the denoising diffusion probability model DDPM follows the characteristics of a first-order Markov chain, and the noisy image at each time is obtained by transferring the noisy image at the previous time through a set Gaussian distribution, the mean of the Gaussian distribution is related to the noisy image at the previous time and the preset variance scheduling parameter, and the variance is the preset variance scheduling parameter;
[0111] After several iterations, the image data gradually approaches the standard Gaussian distribution; the edge distribution of the noisy image at any time with respect to the initial clean image obeys a Gaussian distribution, the mean of the Gaussian distribution is related to the cumulative product of the initial clean image and the multi-step variance scheduling parameter, and the variance is related to the result of 1 minus the cumulative product;
[0112] In the reverse denoising process, the posterior distribution of the noisy image at the previous time with respect to the noisy image at the current time and the initial clean image also obeys a Gaussian distribution, and the diffusion model predicts the noise added to the image to approximate the mean of the Gaussian distribution through the denoising network;
[0113] The target of the diffusion model training is to minimize the mean square error between the real noise and the denoising network predicted noise, and the mean square error is obtained by performing expectation operation on the initial clean image, the added noise and different time steps.
[0114] Specifically, in the forward diffusion process of the Denoising Diffusion Probability Model (DDPM), q(x) t |x t-1 The chain is defined as a first-order Markov chain, and its transition process is formulated as follows:
[0115]
[0116] In the formula, β t ∈(0,1) represents a predefined variance scheduling parameter, x t Let x represent the noisy image at time t. t-1 Let q(x) represent the noisy image at time t-1, and I represent the identity matrix. After T iterations, the data gradually becomes approximately Gaussian distributed, i.e., q(x) T The marginal distribution formula at any time t is: ≈ N(0, I);
[0117]
[0118] In the formula, x0 represents a clean image. This represents the cumulative product of the variance scheduling parameters during the t-step forward diffusion process; it is permissible not to calculate intermediate steps x1,…,x t-1 In this case, x at any time can be obtained directly from x0. t .
[0119] In the reverse denoising process, the distribution p θ (x t-1 |x t The parameterization is performed by a neural network, with the goal of progressively denoising noisy samples and restoring the original data distribution. This is based on the assumption that the posterior distribution q(x) is... t-1 |x t When x0 follows a Gaussian distribution, this distribution can be analytically written as:
[0120]
[0121] In the formula, This represents the mean of the Gaussian distribution. This represents the variance of the Gaussian distribution; the model uses a denoising network ε. θ (x t Approximate mean of t) That is, predicting the added noise; the training objective is to minimize the Kullback-Leibler divergence between the true posterior and the model's approximate posterior, which ultimately simplifies to the mean squared error loss, as shown in the formula:
[0122]
[0123] In the formula, Let ε represent the expectation operation, ε represent the Gaussian noise added to the clean image, and θ represent the parameters of the denoising network. This objective function ensures that the model can progressively recover clean data from noisy samples, thereby achieving the generation of high-quality images.
[0124] In this embodiment, in the diffusion model generation and structure guidance step S2, as follows: Figure 2 As shown in (b), to improve cross-modal translation efficiency, this invention introduces a lightweight translator T at a critical moment τ in the diffusion process. φ Conditions are injected only at this step, thus avoiding the additional computational overhead of relying on condition constraints throughout the diffusion inversion process.
[0125] The cross-domain translator receives the diffusion state and the pseudo-RGB image constructed by the SAR polarization channel at a set time, and generates cross-domain features after processing.
[0126] During the backsampling phase, when the time step is greater than the set time, the diffusion model completes the denoising inversion solely through the diffusion network itself without relying on additional conditions.
[0127] When the time step is less than or equal to the set time, the diffusion model uses the generated cross-domain features as additional input for conditional denoising.
[0128] A regularization loss is added to the training objective. The regularization loss consists of two parts: one part is the mean square error of noise prediction to ensure the denoising accuracy under conditional input, and the other part is the square of the Frobenius norm of the gradient of the cross-domain translator parameters, which is then multiplied by a preset regularization coefficient to suppress overfitting of the diffusion model and ensure training stability.
[0129] Specifically, at time τ, the cross-domain translator receives the current diffusion state. yτ The pseudo-RGB image constructed from the SARVV / VH polarization channels is used as input to generate cross-domain feature representations:
[0130]
[0131] In the formula, z τ Let φ represent the cross-domain features generated at time τ, and let φ represent the parameters of the cross-domain translator. This represents a pseudo-RGB image constructed from SAR polarization channels. During the backsampling stage, when time step t > τ, the model performs unconditional denoising, relying solely on the diffusion network itself to complete the inversion. When time step t ≤ τ, conditional denoising is performed, using cross-domain features z. τ As an additional input, to ensure that pseudo-RGB information provides effective constraints on the generation process at critical moments.
[0132] To maintain the stability of cross-domain conditional injection, a regularization loss is added to the training objective, with the following formula:
[0133]
[0134] wherein the first term ||ε-ε θ (y t ,t,z τ *)|| 2 is the mean square error of noise prediction, used to ensure the denoising accuracy under the condition input, denotes the cross-domain feature of the key time step, the second term is a regularization term of the translator, used to suppress overfitting and ensure stable training, and γ is a regularization coefficient, denotes the gradient of the cross-domain translator parameter, denotes the square of the Frobenius norm. Through this design, the application can realize stable cross-domain mapping based on pseudo RGB images while effectively reducing the computational burden, thereby improving the quality and efficiency of SAR to optical image generation in a fire scene.
[0135] In this embodiment, in the diffusion model generation and structure guiding step S3, as shown in (c) of FIG. 3, Figure 2 To solve the problem of structure discontinuity caused by block-based reconstruction in large-scale remote sensing images, as well as geometric misalignment and edge artifacts, etc., the application proposes a hierarchical structure guiding mechanism. This mechanism adopts a dual-channel reconstruction strategy, combining a global low-resolution path and a local high-resolution path, thereby maintaining spectral details while ensuring the integrity of spatial structure.
[0136] In the global low-resolution path of the hierarchical structure guiding mechanism:
[0137] First, the input high-resolution image is down-sampled using a bilinear method to obtain a low-resolution image; then the low-resolution image is input into a backbone network combined with a cross-domain translator to generate a rough reconstruction result; finally, the rough reconstruction result is up-sampled using a bilinear method to obtain a structure guiding image;
[0138] In the local high-resolution path of the hierarchical structure guiding mechanism:
[0139] First, the original high-resolution image is divided into multiple overlapping image blocks; each image block is independently processed by being input into the same backbone network to obtain corresponding local features;
[0140] Then, all local features are aggregated to form a local reconstruction result through a spatial concatenation operation; a structure-sensitive region is defined, and the structure-sensitive region includes all pixel points with a horizontal or vertical Manhattan distance to the edge of the nearest image block less than or equal to a preset boundary threshold;
[0141] The spatial adaptive weight function is constructed based on the structure sensitive region, so that the weights in the neighborhood of the image block boundary are smoothly transitioned; and the final image reconstruction result is output by pixel-by-pixel weighted fusion of the structure guided image and the local reconstruction result through the spatial adaptive weight function, and the sum of the weighted proportion of the structure guided image and the weighted proportion of the local reconstruction result by the spatial adaptive weight function is 1.
[0142] Specifically, the global low resolution path of the hierarchical structure guided mechanism performs low resolution image reconstruction on the input high resolution image I HR ∈R H×W×3 The bilinear down-sampling is performed to obtain a low resolution image, and the formula is as follows:
[0143] I LR bilinear HR
[0144] In the formula, I LR represents the low resolution image, D bilinear represents a bilinear down-sampling operator, H and W respectively represent the height and width of the high resolution image, and 3 represents the channel number of the image; the backbone network f generates a rough reconstruction result based on the low resolution input, and the formula is as follows:
[0145] G LR DMT LR
[0146] In the formula, G LR represents the rough reconstruction result, and f DMT represents the backbone network combined with the cross-domain translator; then the structure guided image is obtained by bilinear up-sampling, and the formula is as follows:
[0147] G guide bilinear bilinear HR
[0148] In the formula, G guide represents the structure guided image, and U bilinear represents a bilinear up-sampling operator. The guided image can encode the topological constraints at the scene level, but due to the reduction in resolution, there is a certain loss of high frequency texture information.
[0149] Meanwhile, the local high resolution path divides the original image I HR into K overlapping image blocks Each image block is independently processed through the same backbone network f to obtain a local feature representation, and is aggregated through a spatial splicing operation to obtain a local reconstruction result:
[0150]
[0151] where G local denotes the local reconstruction result, denotes the spatial stitching operation, P (k) denotes the k-th image patch. This way can preserve fine spectral features (such as spectral difference within the burned area), but is prone to structure mismatch due to independent processing of image patches. To alleviate this problem, the present application further introduces a structure-sensitive region Γ boundary , defined as follows:
[0152] Γ boundary = {p = (i, j) | min(d x (p), d y (p)) ≤ ω}
[0153] where p = (i, j) denotes the pixel coordinate in the image, i and j denote the row and column coordinates of the pixel, respectively, and ω is the boundary threshold, d x (p) denotes the horizontal Manhattan distance from pixel p to the nearest image patch edge, d y (p) denotes the vertical Manhattan distance from pixel p to the nearest image patch edge, and ω is the boundary threshold, d x , d y denote the Manhattan distance from the pixel to the nearest image patch edge. Based on this region, a spatially adaptive weight function a(p) is constructed, which allows smooth transition within the boundary neighborhood, thereby reducing artifacts and misalignment.
[0154] Finally, the present application obtains the final reconstruction output through weighted fusion:
[0155]
[0156] where denotes the final reconstruction output, and denotes the pixel-wise weighted operation. Through this fusion method, the present application effectively solves the structure discontinuity and artifact problem caused by block processing while maintaining spectral consistency and detail fidelity, thereby obtaining more complete, clear and reliable optical image reconstruction results in large-scale fire monitoring scenarios.
[0157] Referring to Figure 4 , after the preprocessing in step S1, the spectral feature guidance in step S2, the diffusion generation process guided by the cross-domain translator and the hierarchical structure in step S3, the present application outputs the optical style image in step S4. For local regions, small detailed results can be generated to show the spectral consistency and texture details of the burned area. In large-scale scenes, complete large-scale images are obtained through stitching, maintaining the continuity and spatial consistency of the global structure. The generated results not only intuitively approach the real optical image, but also can assist more accurate disaster identification and assessment in fire monitoring tasks.
[0158] The application scenarios of the present application are as follows:
[0159] Scenario one, real-time dynamic monitoring scenario of forest and grassland fire
[0160] In open areas such as forests and grasslands where large-scale fires are prone to occur, fires are often accompanied by a large amount of smoke. Traditional optical remote sensing satellites or unmanned aerial vehicles cannot obtain clear images due to smoke obstruction, making it difficult for monitoring personnel to grasp the fire situation in real time. The present application can generate clear optical style images based on SAR satellites such as C-band Sentinel-1, L-band alos, or SAR equipment carried by unmanned aerial vehicles to obtain all-weather SAR images. Even in thick smoke and cloudy weather, the generated images can intuitively distinguish between burning areas, burned areas, and unaffected areas, clearly showing the direction and speed of fire spread. Monitoring personnel can adjust the firefighting strategy in real time based on the images, such as deploying firefighting equipment to the front of the fire spread, and warning residents and important facilities in advance, effectively improving the efficiency of emergency disposal of forest and grassland fires.
[0161] Scenario two, fine monitoring scenario of urban and industrial park fires
[0162] Urban building groups and industrial parks have complex burning materials and hidden fire characteristics. Traditional monitoring methods are easily affected by building obstructions and local smoke, making it difficult to accurately locate the fire source and assess the damage range. The present application generates optical style images by processing high-resolution SAR images, which can clearly present details such as building appearance and road layout. Combined with the advantage of spectral consistency, it can distinguish the burning damage level of building exterior walls and the damage situation of equipment in the park. The color change of buildings in the image can be used to judge the burning damage level, and the texture features of road areas can be used to identify whether there are collapsed materials blocking the rescue channel. This image can provide accurate on-site environmental information for firefighters, assist in developing a breakthrough rescue plan, and provide intuitive evidence for subsequent fire cause investigation.
[0163] Scenario three, post-disaster rapid assessment and damage assessment scenario
[0164] After the fire is extinguished, quickly and accurately assessing the burned area, damaged ground object type and loss degree is the key to post-disaster reconstruction and insurance claims. Traditional post-disaster assessment relies on manual field investigation, which is time-consuming and labor-intensive and difficult to quickly cover large areas of fire. The present application can process SAR images after a fire and generate complete large-scale optical style images: on the one hand, thanks to the structural coherence advantage, it can present the spatial distribution of the entire fire area completely and accurately calculate the burned area, such as distinguishing the burned vegetation area and damaged building area by different spectral characteristics in the image; on the other hand, through the spectral fidelity characteristic, it can distinguish the damage level of the ground object, such as whether the vegetation is completely burned or the building is only externally blackened or the internal structure is collapsed. Based on the image, the assessment personnel can quickly prepare the post-disaster assessment report, which provides data support for the government to develop reconstruction planning and the insurance company to calculate the claim amount, greatly shortening the post-disaster assessment period.
[0165] Scene four, fire emergency monitoring under adverse weather conditions
[0166] Under adverse weather conditions such as heavy rain, heavy snow, thick fog, night, etc., traditional optical remote sensing equipment is completely disabled due to insufficient light and signal shielding, while SAR equipment can obtain images but interpretation is difficult, leading to a blind period of fire monitoring. The present application relies on the all-weather imaging capability of SAR images and its stable generation effect, which can continuously output optical style images under the above adverse conditions: for example, in a night forest fire, the generated image can clearly show the location and spread trend of the fire point, avoiding misjudgment of the fire due to insufficient light; when urban waterlogging accompanied by fire occurs in heavy rain, the image can distinguish between water accumulation and burning areas, preventing rescue personnel from misjudging the road conditions and getting into danger. Under this scenario, the present application can ensure the continuity of fire monitoring and provide uninterrupted data support for fire emergency response under extreme conditions.
[0167] Embodiment 2
[0168] Referring to Figure 5 , the present application embodiment 2 also provides a SAR to optical image generation device for fire monitoring, comprising:
[0169] The SAR image acquisition and preprocessing module 100 is used for acquiring synthetic aperture radar (SAR) satellite images in a fire monitoring scene, performing radiation calibration and geometric correction on the original SAR satellite image data; then performing normalization and difference operation on the polarization channels of the SAR satellite image, combining the operation results to form a pseudo-RGB image; and cutting the large-format satellite image into fixed-size image blocks and processing them with a preset overlap strategy;
[0170] The spectral feature extraction and fire region guiding module 200 is configured to acquire multispectral image data of a target region, extract a near-infrared waveband NIR and a short-wave infrared waveband SWIR from the multispectral image data, calculate a normalized burn index NBR based on the extracted near-infrared waveband NIR and short-wave infrared waveband SWIR, distinguish a burned region and an unaffected region by using the normalized burn index NBR, perform threshold segmentation on the calculated NBR image to obtain a binary fire mask, and distinguish the burned region and the background region by using the fire mask in the model training process.
[0171] The diffusion model generation and structure guiding module 300 is configured to use a denoising diffusion probability model DDPM as an image generation backbone, embed a cross-domain translator in the diffusion model, and complete generation of SAR satellite images to optical images by combining a hierarchical structure guiding mechanism; the cross-domain translator is introduced at a set time in the diffusion process, and unconditional denoising and conditional denoising are respectively performed in the reverse sampling stage according to different time steps; the hierarchical structure guiding mechanism adopts a double-channel reconstruction strategy, generates a structure guiding image through a global low-resolution path and generates a local reconstruction result through a local high-resolution path, defines a structure sensitive region and constructs a spatial adaptive weight function, and performs weighted fusion on the structure guiding image and the local reconstruction result based on the spatial adaptive weight function to obtain a final image reconstruction result.
[0172] The optical style image output module 400 is configured to generate a small image result to show spectral consistency and texture details of a fire burned region for a local region according to the final image reconstruction result, and obtain a complete large-scale optical style image by splicing in a scene with a size exceeding a set size.
[0173] In the embodiment, the SAR image acquisition and preprocessing module 100 is configured to:
[0174] The SAR satellite images use C-band Sentinel-1 data sources and L-band alos data sources; the C-band Sentinel-1 data sources include two polarization channels of vertical emission / vertical reception VV and vertical emission / horizontal reception VH, and the L-band alos data sources include two polarization channels of horizontal emission / horizontal reception HH and horizontal emission / vertical reception HV.
[0175] The pseudo RGB image corresponding to the C-band SAR image is represented by using an absolute value difference between VV and VH in a red channel, VH in a green channel, and VV in a blue channel.
[0176] The pseudo- RGB image corresponding to the L-band SAR image, the red channel is represented by the absolute value difference between HH and HV, the green channel is represented by HV, and the blue channel is represented by HH; the size of the cropped image block is set to 256*256 pixels, and the preset overlap strategy is 64-pixel overlap.
[0177] In this embodiment, the spectral feature extraction and fire region guiding module 200 includes:
[0178] The result of subtracting the pixel value of the short-wave infrared band from the pixel value of the near-infrared band is divided by the result of adding the pixel value of the near-infrared band to the pixel value of the short-wave infrared band;
[0179] When performing threshold segmentation on the NBR image, an empirical threshold is preset, and when the NBR value is less than the empirical threshold, the corresponding pixel point is determined as a fire region, otherwise it is a non-fire region, so as to obtain a binary fire mask with the same size as the original image, wherein the pixel with a value of 1 represents a fire region, and the pixel with a value of 0 represents a non-fire region;
[0180] When calculating the error between the generated optical image and the real optical image by using the region-weighted reconstruction loss function, the fire mask is multiplied element by element with the difference between the generated optical image and the real optical image, and then the square norm of the obtained result is calculated;
[0181] The spectral guiding loss function is obtained by weighting and summing the region-weighted reconstruction loss and the global consistency loss by using a preset weight parameter.
[0182] In this embodiment, the diffusion model generation and structure guiding module 300 includes:
[0183] The forward diffusion process of the denoising diffusion probability model DDPM follows the characteristics of a first-order Markov chain, and each time's noisy image is obtained by transferring the previous time's noisy image through a set Gaussian distribution, the mean of the Gaussian distribution is related to the previous time's noisy image and a preset variance scheduling parameter, and the variance is the preset variance scheduling parameter.
[0184] After a number of iterations, the image data gradually approaches a standard Gaussian distribution; the edge distribution of the noisy image at any time with respect to the initial clean image obeys a Gaussian distribution, the mean of the Gaussian distribution is related to the cumulative product of the initial clean image and a multi-step variance scheduling parameter, and the variance is related to the result of 1 minus the cumulative product;
[0185] In the reverse denoising process, the posterior distribution of the previous time's noisy image with respect to the current time's noisy image and the initial clean image also obeys a Gaussian distribution, and the diffusion model predicts the noise added to the image by the denoising network to approximate the mean of the Gaussian distribution;
[0186] The diffusion model is trained to minimize the mean square error between the real noise and the predicted noise of the denoising network, which is obtained by taking the expectation of the initial clean image, the added noise and different time steps.
[0187] In this embodiment, the diffusion model generates and the structure guiding module 300 in:
[0188] The cross-domain translator receives the diffusion state and the pseudo- RGB image constructed by the SAR polarization channel at a set time, and generates cross-domain features after processing;
[0189] In the reverse sampling stage, when the time step is greater than the set time, the diffusion model completes the denoising inversion only through the diffusion network itself without relying on additional conditions;
[0190] When the time step is less than or equal to the set time, the diffusion model generates the cross-domain features as additional input for conditional denoising.
[0191] The regularization loss is added to the training target, which includes two parts. One part is the mean square error of noise prediction to ensure the denoising accuracy under the condition of input, and the other part is the square of the Frobenius norm of the cross-domain translator parameter gradient multiplied by the preset regularization coefficient to suppress the overfitting of the diffusion model and ensure the stability of the training.
[0192] In this embodiment, the diffusion model generates and the structure guiding module 300 in:
[0193] The input high-resolution image is down-sampled by bilinear interpolation to obtain a low-resolution image; the low-resolution image is input into the backbone network combined with the cross-domain translator to generate a rough reconstruction result; and finally the rough reconstruction result is up-sampled by bilinear interpolation to obtain a structure guiding image.
[0194] As an SAR-to-optical image generation device for fire monitoring, the diffusion model generates and the structure guiding module in:
[0195] The original high-resolution image is divided into multiple overlapping image blocks; each image block is input into the same backbone network for independent processing to obtain corresponding local features;
[0196] All local features are aggregated to form a local reconstruction result by a spatial concatenation operation; a structure sensitive region is defined, which includes all pixel points with a horizontal or vertical Manhattan distance to the nearest image block edge less than or equal to a preset boundary threshold.
[0197] The spatial adaptive weight function is constructed based on the structure sensitive region, so that the weights in the neighborhood of the image block boundary are smoothly transitioned; and the final image reconstruction result is output by performing pixel-by-pixel weighted fusion on the structure guided image and the local reconstruction result through the spatial adaptive weight function, and the sum of the weighted proportion of the structure guided image and the weighted proportion of the local reconstruction result by the spatial adaptive weight function is 1.
[0198] It should be explained that the information interaction and execution process between the modules of the device are based on the same concept as the method embodiment in Embodiment 1 of the present application, and the technical effects brought by the method embodiment are the same as those of the method embodiment, and the specific content can be referred to the description in the foregoing method embodiment of the present application, which will not be repeated here.
[0199] Embodiment 3
[0200] Embodiment 3 of the present application provides a non-transitory computer readable storage medium, the computer readable storage medium stores a program code of a SAR-to-optical image generation method for fire monitoring, the program code includes instructions for executing the SAR-to-optical image generation method for fire monitoring of Embodiment 1 or any possible implementation manner thereof.
[0201] SAR-to-optical image generation method for fire monitoring.
[0202] The computer readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)) and the like.
[0203] Embodiment 4
[0204] Embodiment 4 of the present application provides an electronic device, comprising: a memory and a processor;
[0205] The processor and the memory complete mutual communication through a bus; the memory stores program instructions executable by the processor, and the processor calling the program instructions can execute the SAR-to-optical image generation method for fire monitoring of Embodiment 1 or any possible implementation manner thereof.
[0206] Specifically, the processor can be implemented by hardware or software, when implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which realizes by reading software codes stored in a memory, the memory can be integrated in the processor, or can be located outside the processor and exist independently.
[0207] In the embodiments described above, all or some of the modules / units can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or some of the modules / units can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded into and executed by a computer, all or some of the procedures or functions as described in the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable system. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) manner.
[0208] It is obvious that those skilled in the art should understand that the modules or steps of the present application described above can be implemented by a general computing system, which can be concentrated on a single computing system or distributed on a network composed of multiple computing systems, and optionally, they can be implemented by program codes executable by a computing system, so that they can be stored in a storage system and executed by a computing system, and in some cases, the steps shown or described can be executed in different order, or they can be made into individual integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module. Thus, the present application is not limited to any particular combination of hardware and software.
[0209] Although the present application has been described in detail by the above general description and specific embodiments, some modifications or improvements can be made on the basis of the present application, which is obvious to those skilled in the art. Therefore, these modifications or improvements made on the basis of not deviating from the spirit of the present application are within the scope of the present application.
Claims
1. A method for generating optical images from SAR for fire monitoring, characterized in that, Includes the following steps: SAR Image Acquisition and Preprocessing: Acquire synthetic aperture radar (SAR) satellite images in fire monitoring scenarios, perform radiometric calibration and geometric correction on the raw SAR satellite image data; then normalize and perform differential operations on the polarization channels of the SAR satellite images, combine the operation results to construct a pseudo-RGB image; crop large-format satellite images into fixed-size image blocks, and process them using a preset overlap strategy. Spectral feature extraction and fire area guidance: acquire multispectral image data of the target area, and extract near-infrared (NIR) and short-wave infrared (SWIR) bands from the multispectral image data; calculate the normalized burn index (NBR) based on the extracted NIR and SWIR bands, and use the NBR to distinguish between burned areas and unaffected areas; perform threshold segmentation on the calculated NBR image to obtain a binarized fire mask, and use the fire mask to distinguish between burned areas and background areas during model training; During the model training phase, a region-weighted reconstruction loss function is used to calculate the error between the generated optical image and the real optical image, and a global consistency loss is introduced to calculate the error for the entire multispectral image. Diffusion Model Generation and Structure Guidance: The Denoising Diffusion Probability Model (DDPM) is used as the backbone for image generation. A cross-domain translator is embedded in the diffusion model, and a hierarchical structure guidance mechanism is combined to generate optical images from SAR satellite images. The cross-domain translator is introduced at a set time during the diffusion process. During the backsampling stage, unconditional denoising and conditional denoising are performed according to different time steps. The hierarchical structure guidance mechanism adopts a dual-channel reconstruction strategy, generating a structure guidance map through a global low-resolution path and generating local reconstruction results through a local high-resolution path. A structure-sensitive region is defined and a spatial adaptive weight function is constructed. The structure guidance map and the local reconstruction results are weighted and fused based on the spatial adaptive weight function to obtain the final image reconstruction result. Output optical style images: Based on the final image reconstruction results, for local areas, small image results are generated to show the spectral consistency and texture details of the fire-damaged areas; for scenes exceeding a set scale, complete large-scale optical style images are obtained by stitching.
2. The SAR-to-optical image generation method for fire monitoring according to claim 1, characterized in that, In the SAR image acquisition and preprocessing steps: SAR satellite imagery uses C-band Sentinel-1 data source and L-band alos data source; the C-band Sentinel-1 data source includes two polarization channels: vertical transmit / vertical receive (VV) and vertical transmit / horizontal receive (VH), and the L-band alos data source includes two polarization channels: horizontal transmit / horizontal receive (HH) and horizontal transmit / vertical receive (HV). When normalizing and differentiating the polarization channels of SAR satellite images, and then combining the results to construct a pseudo-RGB image: The pseudo-RGB image corresponding to the C-band SAR image is represented by the absolute difference between VV and VH for the red channel, VH for the green channel, and VV for the blue channel. The pseudo-RGB image corresponding to the L-band SAR image is represented by the absolute difference between HH and HV for the red channel, HV for the green channel, and HH for the blue channel. The size of the cropped image block is set to 256×256 pixels, and the default overlap strategy is 64 pixels overlap.
3. The SAR-to-optical image generation method for fire monitoring according to claim 1, characterized in that, In the spectral feature extraction and fire zone guidance steps, when calculating the normalized burn index (NBR): The result of subtracting the short-wave infrared band pixel value from the near-infrared band pixel value is divided by the result of adding the near-infrared band pixel value to the short-wave infrared band pixel value. When performing thresholding on an NBR image, an empirical threshold is preset. When the NBR value is less than the empirical threshold, the corresponding pixel is determined to be a fire area, and otherwise it is a non-fire area. This is to obtain a binary fire mask with the same size as the original image. Pixels with a value of 1 in the mask represent fire areas, and pixels with a value of 0 represent non-fire areas. When calculating the error between the generated optical image and the real optical image using the region-weighted reconstruction loss function, the difference between the fire mask and the generated optical image and the real optical image is multiplied element by element, and then the square norm is calculated on the result. The spectral guided loss function is obtained by weighting the regional weighted reconstruction loss and the global consistency loss separately by using preset weight parameters and then summing them.
4. The SAR-to-optical image generation method for fire monitoring according to claim 1, characterized in that, In the diffusion model generation and structure guidance steps: The forward diffusion process of the Denoising Diffusion Probability Model (DDPM) follows the properties of a first-order Markov chain. The noisy image at each time step is obtained by transferring the noisy image at the previous time step through a set Gaussian distribution. The mean of the Gaussian distribution is related to the noisy image at the previous time step and the preset variance scheduling parameter, and the variance is the preset variance scheduling parameter. After several iterations, the image data gradually approaches a standard Gaussian distribution; The edge distribution of the noisy image relative to the initial clean image at any time follows a Gaussian distribution. The mean of the Gaussian distribution is related to the cumulative product of the initial clean image and the multi-step variance scheduling parameters, and the variance is related to 1 minus the result of this cumulative product. In the reverse denoising process, the posterior distribution of the noisy image at the previous time step relative to the noisy image at the current time step and the initial clean image also follows a Gaussian distribution. The diffusion model predicts the noise added to the image through the denoising network to approximate the mean of this Gaussian distribution. The goal of training the diffusion model is to minimize the mean square error between the real noise and the noise predicted by the denoising network. The mean square error is obtained by performing expectation calculations on the initial clean image, the added noise, and different time steps.
5. The SAR-to-optical image generation method for fire monitoring according to claim 1, characterized in that, In the diffusion model generation and structure guidance steps: The cross-domain translator receives the diffusion state and the pseudo-RGB image constructed by the SAR polarization channel at a set time, and generates cross-domain features after processing. During the backsampling phase, when the time step is greater than the set time, the diffusion model completes the denoising inversion solely through the diffusion network itself without relying on additional conditions. When the time step is less than or equal to the set time, the diffusion model uses the generated cross-domain features as additional input for conditional denoising. A regularization loss is added to the training objective. The regularization loss consists of two parts: one part is the mean square error of noise prediction to ensure the denoising accuracy under conditional input, and the other part is the square of the Frobenius norm of the gradient of the cross-domain translator parameters, which is then multiplied by a preset regularization coefficient to suppress overfitting of the diffusion model and ensure training stability.
6. The SAR-to-optical image generation method for fire monitoring according to claim 1, characterized in that, In the diffusion model generation and structure guidance steps, during the global low-resolution path processing of the hierarchical structure guidance mechanism: First, bilinear downsampling is applied to the input high-resolution image to obtain a low-resolution image; then, the low-resolution image is input into the backbone network combined with the cross-domain translator to generate a coarse reconstruction result; finally, bilinear upsampling is applied to the coarse reconstruction result to obtain the structure guide map. When processing local high-resolution paths in a hierarchical guidance mechanism: First, the original high-resolution image is divided into multiple overlapping image blocks; each image block is input into the same backbone network and processed independently to obtain the corresponding local features; Then, all local features are aggregated through spatial stitching to form a local reconstruction result; define a structure-sensitive region, which includes all pixels whose horizontal or vertical Manhattan distance to the edge of the nearest image block is less than or equal to a preset boundary threshold; A spatially adaptive weighting function is constructed based on the structurally sensitive region to ensure a smooth transition of weights within the neighborhood of image block boundaries. The final image reconstruction result is obtained by weighting and fusing the structural guidance map and the local reconstruction result pixel by pixel using the spatially adaptive weighting function. The sum of the weighting ratio of the spatially adaptive weighting function on the structural guidance map and the weighting ratio on the local reconstruction result is 1.
7. A SAR-to-optical image generation device for fire monitoring, characterized in that, include: The SAR image acquisition and preprocessing module is used to acquire synthetic aperture radar (SAR) satellite images in fire monitoring scenarios, perform radiometric calibration and geometric correction on the raw SAR satellite image data, normalize and differentially process the polarization channels of the SAR satellite images, combine the results to construct a pseudo-RGB image, and crop large-format satellite images into fixed-size image blocks and process them using a preset overlap strategy. The spectral feature extraction and fire area guidance module is used to acquire multispectral image data of the target area, extract near-infrared (NIR) and short-wave infrared (SWIR) bands from the multispectral image data, calculate the normalized burn index (NBR) based on the extracted NIR and SWIR bands, and use the NBR to distinguish between burned areas and unaffected areas. The calculated NBR image is then thresholded to obtain a binarized fire mask, which is used to distinguish between burned areas and background areas during model training. During the model training phase, a region-weighted reconstruction loss function is used to calculate the error between the generated optical image and the real optical image, and a global consistency loss is introduced to calculate the error for the entire multispectral image. The diffusion model generation and structure guidance module uses the Denoising Diffusion Probability Model (DDPM) as the image generation backbone, embeds a cross-domain translator in the diffusion model, and combines it with a hierarchical structure guidance mechanism to generate optical images from SAR satellite images. The cross-domain translator is introduced at a set time during the diffusion process, and unconditional denoising and conditional denoising are performed respectively according to different time steps during the backsampling stage. The hierarchical structure guidance mechanism adopts a dual-channel reconstruction strategy, generating a structure guidance map through a global low-resolution path and generating local reconstruction results through a local high-resolution path. It defines structure-sensitive regions and constructs a spatial adaptive weight function. Based on the spatial adaptive weight function, the structure guidance map and the local reconstruction results are weighted and fused to obtain the final image reconstruction result. The optical style image output module is used to generate small image results for local areas based on the final image reconstruction results to show the spectral consistency and texture details of the fire-damaged area; and to obtain complete large-scale optical style images by stitching together scenes that exceed the set scale.
8. The SAR-to-optical image generation device for fire monitoring according to claim 7, characterized in that, In the SAR image acquisition and preprocessing module: SAR satellite imagery uses C-band Sentinel-1 data source and L-band alos data source; the C-band Sentinel-1 data source includes two polarization channels: vertical transmit / vertical receive (VV) and vertical transmit / horizontal receive (VH), and the L-band alos data source includes two polarization channels: horizontal transmit / horizontal receive (HH) and horizontal transmit / vertical receive (HV). The pseudo-RGB image corresponding to the C-band SAR image is represented by the absolute difference between VV and VH for the red channel, VH for the green channel, and VV for the blue channel. The pseudo-RGB image corresponding to the L-band SAR image is represented by the absolute difference between HH and HV for the red channel, HV for the green channel, and HH for the blue channel. The size of the cropped image block is set to 256×256 pixels, and the default overlap strategy is 64 pixels overlap.
9. The SAR-to-optical image generation device for fire monitoring according to claim 7, characterized in that, In the spectral feature extraction and fire area guidance module: The result of subtracting the short-wave infrared band pixel value from the near-infrared band pixel value is divided by the result of adding the near-infrared band pixel value to the short-wave infrared band pixel value. When performing thresholding on an NBR image, an empirical threshold is preset. When the NBR value is less than the empirical threshold, the corresponding pixel is determined to be a fire area, and otherwise it is a non-fire area. This is to obtain a binary fire mask with the same size as the original image. Pixels with a value of 1 in the mask represent fire areas, and pixels with a value of 0 represent non-fire areas. When calculating the error between the generated optical image and the real optical image using the region-weighted reconstruction loss function, the difference between the fire mask and the generated optical image and the real optical image is multiplied element by element, and then the square norm is calculated on the result. The spectral guided loss function is obtained by weighting the regional weighted reconstruction loss and the global consistency loss separately by using preset weight parameters and then summing them.
10. The SAR-to-optical image generation device for fire monitoring according to claim 7, characterized in that, In the diffusion model generation and structure guidance module: The forward diffusion process of the Denoising Diffusion Probability Model (DDPM) follows the properties of a first-order Markov chain. The noisy image at each time step is obtained by transferring the noisy image at the previous time step through a set Gaussian distribution. The mean of the Gaussian distribution is related to the noisy image at the previous time step and the preset variance scheduling parameter, and the variance is the preset variance scheduling parameter. After several iterations, the image data gradually approaches a standard Gaussian distribution; The edge distribution of the noisy image relative to the initial clean image at any time follows a Gaussian distribution. The mean of the Gaussian distribution is related to the cumulative product of the initial clean image and the multi-step variance scheduling parameters, and the variance is related to 1 minus the result of this cumulative product. In the reverse denoising process, the posterior distribution of the noisy image at the previous time step relative to the noisy image at the current time step and the initial clean image also follows a Gaussian distribution. The diffusion model predicts the noise added to the image through the denoising network to approximate the mean of this Gaussian distribution. The goal of training the diffusion model is to minimize the mean square error between the real noise and the noise predicted by the denoising network. The mean square error is obtained by performing expectation calculations on the initial clean image, the added noise, and different time steps. In the diffusion model generation and structure guidance module: The cross-domain translator receives the diffusion state and the pseudo-RGB image constructed by the SAR polarization channel at a set time, and generates cross-domain features after processing. During the backsampling phase, when the time step is greater than the set time, the diffusion model completes the denoising inversion solely through the diffusion network itself without relying on additional conditions. When the time step is less than or equal to the set time, the diffusion model uses the generated cross-domain features as additional input for conditional denoising. Regularization loss is added to the training objective. The regularization loss consists of two parts: one part is the mean square error of noise prediction to ensure the denoising accuracy under conditional input, and the other part is the square of the Frobenius norm of the gradient of the cross-domain translator parameters, which is then multiplied by a preset regularization coefficient to suppress overfitting of the diffusion model and ensure training stability. In the diffusion model generation and structure guidance module: The high-resolution input image is downsampled using a bilinear method to obtain a low-resolution image. The low-resolution image is then input into a backbone network combined with a cross-domain translator to generate a coarse reconstruction result. Finally, the coarse reconstruction result is upsampled using a bilinear method to obtain a structure guide map. In the diffusion model generation and structure guidance module: The original high-resolution image is divided into multiple overlapping image patches; each image patch is input into the same backbone network and processed independently to obtain the corresponding local features; Then, all local features are aggregated through spatial stitching to form a local reconstruction result; define a structure-sensitive region, which includes all pixels whose horizontal or vertical Manhattan distance to the edge of the nearest image block is less than or equal to a preset boundary threshold; A spatially adaptive weighting function is constructed based on the structurally sensitive region to ensure a smooth transition of weights within the neighborhood of image block boundaries. The final image reconstruction result is obtained by weighting and fusing the structural guidance map and the local reconstruction result pixel by pixel using the spatially adaptive weighting function. The sum of the weighting ratio of the spatially adaptive weighting function on the structural guidance map and the weighting ratio on the local reconstruction result is 1.
Citation Information
Cited By
Remote sensing image enhancement method for forest fire monitoring
CN122115252A