SAR (Synthetic Aperture Radar) super-resolution and target joint recognition method based on physical constraint diffusion model

By combining generative diffusion models with the physical mechanisms of SAR imaging, an end-to-end joint recognition framework is constructed, which solves the problem of synergistic optimization between SAR image super-resolution and target recognition. The generated images have strong physical interpretability, high recognition accuracy, and low computational resource consumption.

CN121190306APending Publication Date: 2025-12-23SICHUAN JIUTIAN EMBODIED INTELLIGENT TECHNOLOGY CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511183906.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

In existing technologies, SAR image super-resolution and target recognition tasks are separated, resulting in poor error propagation and physical interpretability, making it impossible to achieve end-to-end collaborative optimization. This leads to problems with recognition accuracy and reliability, especially in applications with high precision requirements.

Method used

By combining generative diffusion models with the physical mechanisms of SAR imaging, and through physical consistency residuals and latent variable sharing mechanisms, an end-to-end joint identification framework is constructed to achieve synergistic optimization of super-resolution reconstruction and target identification.

Benefits of technology

The generated SAR image texture conforms to the laws of electromagnetic scattering, improving recognition accuracy and reducing computational resource consumption, making it suitable for real-time deployment on resource-constrained platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 4C8CE6A0-EADA-46E4-8D2F-A434E92F2EC6
    Figure 4C8CE6A0-EADA-46E4-8D2F-A434E92F2EC6
Patent Text Reader

Abstract

The invention relates to the technical field of signal processing and target recognition, and discloses an SAR super-resolution and target combined recognition method based on a physical constraint diffusion model, and the method comprises the steps: obtaining a low-resolution SAR image, and carrying out the preprocessing of the low-resolution SAR image; performing super-resolution reconstruction on the preprocessed low-resolution SAR image by using a generative diffusion model; coupling a target identification branch on the generative diffusion model, wherein the target identification branch and an intermediate layer of the generative diffusion model share latent variable features; and carrying out end-to-end joint optimization on the generative diffusion model and the target identification branch through a joint loss function comprising image reconstruction loss, physical constraint loss formed by physical consistency residual errors and target identification loss. According to the method, the reconstructed sub-pixel-level textures and details are ensured to be clear visually, the common problems of artifacts, detail distortion, inconsistent structures and the like in a traditional deep learning method are effectively inhibited, and the reliability of super-resolution images is greatly enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of signal processing and target recognition, and particularly relates to a SAR super-resolution and target joint recognition method based on a physical constraint diffusion model. BACKGROUND

[0002] Synthetic Aperture Radar (SAR) is an active microwave remote sensing system, which plays an indispensable role in key fields such as military reconnaissance, national mapping, environmental monitoring and disaster assessment due to its unique advantages of all-weather and all-day observation. With the complication of application scenarios, the industry has put forward increasingly stringent requirements for the spatial resolution of SAR images and the accuracy of target automatic recognition.

[0003] In the prior art, improving the nominal resolution of SAR images and performing target recognition usually follow two technical routes, but both routes have significant technical bottlenecks: 1. A phased and serial processing flow: This is the most common processing mode. First, a super-resolution (SR) algorithm is used to enhance the original low-resolution SAR image to generate a high-resolution image; then, the high-resolution image is input into an independently trained target recognition model (such as a classifier or detector based on a convolutional neural network CNN) for analysis. The fundamental defect of this flow is task fragmentation and error propagation. The goal of the super-resolution algorithm is to minimize the pixel reconstruction error or improve the visual quality, and the details generated by the algorithm, especially those generated by data-driven models (such as conventional CNNs or generative adversarial networks GANs), may not conform to the electromagnetic scattering physical laws of SAR, and may even produce artifacts or false textures. These non-real details will be received as noise or misleading information by the downstream target recognition model, thereby seriously affecting the accuracy and reliability of recognition. Since the objective functions of the two models are independent of each other, end-to-end collaborative optimization cannot be achieved, resulting in a far-from-optimal overall system performance.

[0004] 2. Simply relying on data-driven deep learning models: In recent years, deep learning models represented by CNNs and GANs have been widely used in SAR super-resolution tasks. However, these models are essentially "black box" models that learn a mapping relationship from low resolution to high resolution through massive data, but they often completely ignore the unique and strict coherent imaging mechanism of SAR. This results in an output that may be visually clearer, but has poor physical interpretability and cannot guarantee the accuracy of key physical features such as scattering centers and shadows, posing a significant risk in high-precision and high-reliability applications such as military target recognition.

[0005] 3. Preliminary application of generative diffusion model: As the latest high-performance generative model, the diffusion model (Diffusion Model) surpasses GAN in image generation quality and is beginning to be explored for image super-resolution tasks. It can generate high-quality details by simulating an inverse process of gradually denoising from noise to recover images. However, if it is directly applied to the SAR field, it still faces similar core problems as the previous methods: first, its generation process lacks physical constraints and cannot guarantee that the reconstruction result conforms to the scattering law of SAR; second, it is still a general image generation framework and has not been deeply coupled and optimized with the downstream target recognition task.

[0006] Therefore, how to break through the limitations of traditional staged processing, deeply integrate the physical prior knowledge of SAR imaging into advanced generative models, and build a new technical framework that can achieve collaborative gain and end-to-end integrated optimization of super-resolution reconstruction and target recognition tasks is a technical problem that needs to be solved in the current SAR intelligent processing field. SUMMARY

[0007] In order to solve the problems existing in the prior art, the present application aims to provide a method for realizing integrated processing of SAR image super-resolution reconstruction and target recognition task by using deep generative model, especially combining imaging physics principles with reversible diffusion network.

[0008] In order to achieve the above-mentioned application purposes, the technical solutions provided by the present application include: The SAR super-resolution and target joint recognition method based on the physically constrained diffusion model includes the following steps: Obtain a low-resolution SAR image and pre-process it; Use a generative diffusion model to perform super-resolution reconstruction on the pre-processed low-resolution SAR image, wherein the reconstruction process of the generative diffusion model is constrained by a physical consistency residual calculated based on the SAR imaging physical mechanism; Couple a target recognition branch on the generative diffusion model, which shares latent variable features with the intermediate layer of the generative diffusion model to output target recognition results; End-to-end joint optimize the generative diffusion model and the target recognition branch through a joint loss function including image reconstruction loss, physical constraint loss composed of the physical consistency residual, and target recognition loss.

[0009] Preferably, the calculation method of the physical consistency residual includes: Construct a SAR forward scattering propagation model for describing the mapping relationship between the physical structure of the target and the echo signal; inputting the intermediate reconstructed image generated by the generative diffusion model into the forward scattering propagation model to generate a predicted echo; calculating a difference between the predicted echo and a real observed echo to obtain the physical consistency residual.

[0010] Preferably, the method for super-resolution reconstruction of the preprocessed low-resolution SAR image by using a generative diffusion model comprises: mapping the physical consistency residual into a gating factor, and applying the gating factor to a condition control path of the generative diffusion network for dynamically adjusting the correction strength of the physically inconsistent area in the reconstruction process.

[0011] Preferably, the preprocessing comprises sequentially performing denoising, geometric correction and registration, and amplitude normalization and phase unwrapping processing on the low-resolution SAR image data.

[0012] Preferably, the method for super-resolution reconstruction of the preprocessed low-resolution SAR image by using a generative diffusion model further comprises: performing forward noise disturbance on the preprocessed low-resolution SAR image data to generate a series of noisy samples; and performing reverse denoising by a noise predictor to gradually recover a super-resolution image from the noisy samples.

[0013] Preferably, the end-to-end joint optimization further comprises: According to the gradient size of the image reconstruction loss, the physical constraint loss and the target recognition loss, the weight of each in the joint loss function is adaptively adjusted to balance the convergence speed of different tasks.

[0014] Preferably, when performing super-resolution reconstruction, the method further comprises the step of: dynamically adjusting or prematurely truncating the denoising iteration step number of the generative diffusion model according to the convergence of the physical consistency residual, so as to reduce the computational overhead while ensuring the reconstruction accuracy.

[0015] Preferably, the joint loss function further comprises a multi-scale perception loss for comparing the differences between the generated super-resolution image and the high-resolution ground truth image at multiple feature layers during the training process, so as to improve the texture detail quality of the reconstructed image.

[0016] Advantages 1、The electromagnetic scattering propagation equation of SAR imaging is used as an explicit physical constraint in the present application to correct and guide the generation process of the diffusion model in real time, so that the reconstructed sub-pixel level texture and details not only look clear in vision, but also conform to the real electromagnetic scattering law in physics, effectively suppressing the common problems such as artifacts, distorted details and inconsistent structures in traditional deep learning methods, and greatly enhancing the reliability of the super-resolution image.

[0017] 2. This invention deeply couples super-resolution reconstruction and target recognition into a single model architecture through a latent variable sharing mechanism. The recognition branch directly utilizes features rich in fine structure and texture information learned by the diffusion model during the reconstruction process, while the gradient of the recognition task is also backpropagated to optimize the reconstruction process, generating features more conducive to classification. This collaborative optimization mechanism breaks through the performance bottleneck of traditional sequential methods, significantly improving the model's target recognition accuracy and generalization ability under different resolutions and scenarios.

[0018] 3. To address the inherent problems of high computational cost and numerous iterations in diffusion models, this invention proposes a dynamic truncation diffusion step size control strategy based on physical consistency residuals. This strategy can adaptively terminate redundant iterative calculations early while ensuring image generation quality and physical consistency, significantly reducing computational resource consumption and providing feasibility for real-time deployment on resource-constrained platforms such as spaceborne and airborne systems. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the SAR super-resolution and target joint identification method based on a physical constraint diffusion model provided in a preferred embodiment of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be described in detail below. It should be understood that the specific embodiments described herein are only for explaining this invention and are not intended to limit this invention.

[0021] like Figure 1 As shown, this invention provides a joint SAR super-resolution and target recognition method based on a physically constrained diffusion model. Its core idea is to no longer treat super-resolution reconstruction and target recognition as two independent sequential tasks, but rather to construct an end-to-end unified deep learning framework guided in real-time by physical mechanisms. This framework uses a generative diffusion model as its backbone, introducing electromagnetic scattering physics as a strong constraint at each step of the model's image generation. Simultaneously, through a latent variable sharing mechanism, the model's recognition and reconstruction branches co-evolve, ultimately improving the physical realism of image reconstruction, detail fidelity, and target recognition accuracy through a jointly optimized objective function.

[0022] S1. Acquire low-resolution SAR images and perform preprocessing.

[0023] Those skilled in the art will understand that the purpose of the preprocessing is to eliminate noise, correct geometric distortions, and standardize the data, providing high-quality, format-consistent input for subsequent modeling. In some preferred embodiments, the preprocessing includes the following specific operations: First, low-resolution amplitude or phase map data output from the SAR imaging system is acquired. This can include single-polarization, fully polarized complex images, or multi-channel interferometric SAR data. Next, denoising processing is performed on the raw data, for example, using a Lee adaptive filter to effectively suppress system noise and multiplicative speckle noise while preserving target edge and texture details to the maximum extent. Then, geometric correction and registration are performed. Auxiliary data such as orbital parameters are used for geometric correction of the images. For multi-temporal or multi-channel data, precise registration is performed using methods such as phase correlation to ensure pixel-level alignment of all images in physical space. After correction, the images are cropped to a uniform size, and the amplitude and phase components are processed separately: amplitude is normalized (e.g., min-max scaling), and phase is unwrapped (e.g., using the Itoh algorithm) to eliminate phase blur. Finally, the processed amplitude and phase components are merged by channel to form a standardized multi-channel tensor, which serves as the input to subsequent models.

[0024] S2. Super-resolution reconstruction of the preprocessed low-resolution SAR image is performed using a generative diffusion model, wherein the reconstruction process of the generative diffusion model is constrained by a physical consistency residual calculated based on the SAR imaging physical mechanism.

[0025] This step aims to transform the physical laws of SAR imaging into computable constraints, which can be used to guide and correct the generation process of deep learning models in real time. Specifically, the step of using physically consistent residual constraints to reconstruct the diffusion model includes the following sub-steps: S21. Construct a SAR forward scattering propagation model. This model, based on electromagnetic wave scattering theory, can accurately describe the mathematical mapping relationship between a given ground target scene (i.e., the super-resolution image being generated by the model) and the radar echo signal received by the SAR sensor. In one embodiment, this forward model can be implemented by a system transfer function pre-calculated in the frequency domain, which is uniquely determined by physical parameters such as radar wavelength and imaging geometry.

[0026] S22. In each step of the inverse denoising process using the diffusion model, the physical consistency residual is calculated using the forward model. Specifically, the intermediate reconstructed image currently generated by the diffusion model is input into the constructed forward scattering propagation model to obtain a "predicted echo". Then, this predicted echo is compared with the actually observed radar echo (e.g., the difference between the two is calculated), and the difference is the "physical consistency residual". The magnitude of this residual intuitively quantifies the extent to which the currently generated image deviates from the true physical scattering law.

[0027] S23. The residual is used to constrain the reconstruction process. This invention provides various constraint methods. In a preferred embodiment, the constraint is embodied in two aspects: First, the L2 norm of the residual can be used as a physical constraint loss term, directly added to the overall training objective of the model (i.e., the joint loss function). Second, to achieve finer dynamic control, the physically consistent residual can be mapped to a gating factor. For example, the ratio of the residual to the true echo energy (i.e., the residual ratio ρ) can be used as a scalar gating signal. Then, this gating factor is applied to the conditional control path (e.g., the conditional normalization layer) of the generative diffusion network. When the residual is large (i.e., the physical inconsistency is significant), the gating factor will correspondingly increase the modulation intensity of the relevant physical correction channels in the network, thereby dynamically adjusting the correction intensity for physically inconsistent regions and guiding the generation process to converge in a physically more correct direction.

[0028] S3. A target recognition branch is coupled onto the generative diffusion model. The target recognition branch shares latent variable features with the intermediate layer of the generative diffusion model to output the target recognition result.

[0029] The core of this step is to build a unified network architecture that can perform super-resolution reconstruction and object recognition simultaneously, and to train it jointly end-to-end.

[0030] The steps for super-resolution reconstruction using a generative diffusion model follow the basic principles of diffusion models. Specifically, they include: First, forward noise perturbation is applied to the input low-resolution SAR image data. This involves gradually and controllably adding Gaussian noise to the image within a preset number of steps T until it becomes a purely noise distribution. This process allows the model to learn the distribution patterns of the noise. Then, during training and inference, inverse denoising is performed using a deep neural network (often called a noise predictor, such as the U-Net architecture). Starting from a random noise sample, the network predicts and removes noise components at each step, gradually and iteratively recovering a clear super-resolution image.

[0031] A key aspect of this step is coupling an object recognition branch onto the generative diffusion model. This recognition branch does not exist independently but shares latent variable features with the intermediate layers of the diffusion model. In other words, during the inverse denoising process, the deeper layers of the diffusion model generate a series of feature maps (i.e., latent variables) containing rich spatial structure and texture information. These feature maps are used both for subsequent image reconstruction and as input to the object recognition branch. Subsequently, the object recognition branch processes these shared latent variable features, for example through additional convolutional layers and attention modules, to output object recognition results (such as object category, confidence level, etc.).

[0032] S4. Perform end-to-end joint optimization of the generative diffusion model and the target recognition branch using a joint loss function that includes image reconstruction loss, physical constraint loss consisting of the physical consistency residual, and target recognition loss.

[0033] This step aims to achieve synergistic gains between the super-resolution reconstruction and object recognition tasks. The joint loss function comprises at least three parts: 1. Image reconstruction loss (such as L2 loss) used to measure the pixel difference between the reconstructed image and the real high-resolution image; 2. The aforementioned physical constraint loss consisting of physical consistency residuals; 3. Target recognition loss (such as cross-entropy loss) used to measure the difference between the recognition result and the true label.

[0034] In a preferred embodiment, to further improve the visual perception quality of the generated image, the joint loss function may also include a multi-scale perceptual loss. This loss compares the generated image with the real image at multiple different feature layers through a pre-trained discriminator network, thereby better preserving the texture details and structural information of the image.

[0035] Furthermore, to achieve more efficient joint optimization, this invention proposes a preferred scheme: during training, the weights of the image reconstruction loss, physical constraint loss, and object recognition loss in the joint loss function are adaptively adjusted based on their respective gradient magnitudes. This mechanism can automatically balance the convergence speeds of different tasks, dynamically allocating more training resources to tasks with slower convergence or smaller gradients, thereby avoiding a single task dominating the training process and achieving more stable and efficient collaborative optimization.

[0036] In a preferred embodiment, the super-resolution reconstruction further includes the step of dynamically adjusting or prematurely truncating the number of denoising iterations of the generative diffusion model based on the convergence of the physical consistency residual, so as to reduce computational overhead while ensuring reconstruction accuracy.

[0037] Specifically, in each iteration of the reverse denoising process, the system monitors the changes in the physical consistency residual ratio ρ and pixel reconstruction error in real time. When these metrics stabilize within a preset low threshold for several consecutive steps, the system determines that the reconstruction has sufficiently converged and immediately terminates the remaining denoising iterations, directly outputting the current result. This mechanism can significantly reduce unnecessary computational steps while ensuring the final image quality and physical consistency, thereby greatly reducing the average inference latency and computational overhead.

[0038] The method proposed in this invention can be applied to distributed remote sensing networks composed of spaceborne processing units and ground stations. After receiving real low-resolution SAR data, it undergoes preprocessing and, under the aforementioned dynamic step-size control mechanism, can efficiently and synchronously output high-resolution SAR images, along with information such as target category, confidence level, and location contained therein, for downstream tasks such as military reconnaissance, disaster monitoring, and urban change detection.

[0039] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A SAR super-resolution and target joint identification method based on a physically constrained diffusion model, characterized in that, Including the following steps: Acquire low-resolution SAR images and perform preprocessing; Super-resolution reconstruction of the preprocessed low-resolution SAR image is performed using a generative diffusion model, wherein the reconstruction process of the generative diffusion model is constrained by a physical consistency residual calculated based on the physical mechanism of SAR imaging. A target recognition branch is coupled to the generative diffusion model. The target recognition branch shares latent variable features with the intermediate layer of the generative diffusion model to output the target recognition result. The generative diffusion model and the target recognition branch are jointly optimized end-to-end using a joint loss function that includes image reconstruction loss, physical constraint loss consisting of the physical consistency residuals, and target recognition loss.

2. The SAR super-resolution and target joint identification method based on a physically constrained diffusion model as described in claim 1, characterized in that, The method for calculating the physical consistency residual includes: Construct a SAR forward scattering propagation model to describe the mapping relationship between the target's physical structure and the echo signal; The intermediate reconstructed image generated by the generative diffusion model is input into the forward scattering propagation model to generate the predicted echo; The difference between the predicted echo and the actual observed echo is calculated to obtain the physical consistency residual.

3. The SAR super-resolution and target joint identification method based on a physically constrained diffusion model as described in claim 1 or 2, characterized in that, The method for super-resolution reconstruction of the preprocessed low-resolution SAR image using a generative diffusion model includes: The physical consistency residual is mapped to a gating factor, and the gating factor is applied to the conditional control path of the generative diffusion network to dynamically adjust the correction intensity of the physically inconsistent regions during the reconstruction process.

4. The SAR super-resolution and target joint identification method based on a physically constrained diffusion model as described in claim 1, characterized in that, The preprocessing includes: sequentially performing denoising, geometric correction and registration, amplitude normalization and phase unwrapping on the low-resolution SAR image data.

5. The SAR super-resolution and target joint identification method based on a physically constrained diffusion model as described in claim 1, characterized in that, The method for super-resolution reconstruction of the preprocessed low-resolution SAR image using a generative diffusion model further includes: performing forward noise perturbation on the preprocessed low-resolution SAR image data to generate a series of noisy samples; and performing inverse denoising through a noise predictor to gradually recover the super-resolution image from the noisy samples.

6. The SAR super-resolution and target joint identification method based on a physically constrained diffusion model as described in claim 1, characterized in that, The end-to-end joint optimization also includes: The weights of the image reconstruction loss, physical constraint loss, and target recognition loss in the joint loss function are adaptively adjusted based on their respective gradient magnitudes to balance the convergence speed of different tasks.

7. The SAR super-resolution and target joint identification method based on a physically constrained diffusion model as described in claim 1, characterized in that, The super-resolution reconstruction also includes the following steps: dynamically adjusting or prematurely truncating the number of denoising iterations of the generative diffusion model based on the convergence of the physical consistency residual, so as to reduce computational overhead while ensuring reconstruction accuracy.

8. The SAR super-resolution and target joint identification method based on a physically constrained diffusion model as described in claim 1, characterized in that: The joint loss function also includes a multi-scale perceptual loss, which is used to compare the differences between the generated super-resolution image and the high-resolution ground truth image at multiple feature layers during training, in order to improve the texture detail quality of the reconstructed image.

Citation Information

Cited By

  • Camera image acquisition parameter verification method and system based on big data analysis

    CN121582094A