Electromagnetic inverse scattering imaging method of diffusion model based on multi-modal fusion features
By combining a diffusion model based on multimodal fusion features with a large language model, the problem of insufficient generalization ability of electromagnetic imaging methods when the initial guess is inappropriate and the data modes are missing is solved, and high-precision and stable electromagnetic backscattering imaging is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU DIANZI UNIV
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-21
AI Technical Summary
Existing electromagnetic imaging methods tend to converge to local minima when the initial guess is inappropriate, especially for strong scatterers with high contrast or large electrical size. Furthermore, they lack generalization ability when data modes are missing or the measurement scene changes, resulting in low imaging efficiency and poor accuracy.
A diffusion model based on multimodal fusion features is adopted, combined with a large language model. By introducing multimodal information such as BP plot, current distribution and electromagnetic scattering field data for feature extraction and fusion, a diffusion model is established for electromagnetic backscattering imaging, which enhances the model's generalization ability under different conditions.
It significantly improves the reconstruction accuracy and stability of electromagnetic backscattering imaging, possesses inherent robustness to conditional inputs, and can automatically rely on other conditional information to complete high-quality reconstruction when some conditions are missing, thus improving the practicality and adaptability of the method in real and complex scenarios.
Smart Images

Figure CN121904533A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of electromagnetic imaging and computational imaging, and particularly to an electromagnetic inverse scattering imaging method based on a diffusion model with multimodal feature embedding. This method integrates a deep generative model and a large language model to achieve highly universal and generalizable two-dimensional imaging in a variety of complex and unseen environments.
[0002] Electromagnetic inverse scattering imaging (ESI) reconstructs high-quality images of objects using the interaction of electromagnetic waves and is widely used in security inspection, underground detection, medical imaging, and radar imaging. Traditional electromagnetic imaging methods, such as those based on the Lippmann-Schwinger integral equation, often converge to local minima when the initial guess is inappropriate, especially for strong scatterers with high contrast or large electrical size. Furthermore, their enormous computational cost limits their practical applicability.
[0003] In recent years, deep learning technology has provided new solutions to the aforementioned problems. This method can learn effective feature representations from large amounts of sample data, achieving rapid target reconstruction and significantly improving imaging efficiency. However, existing deep learning-based inverse scattering imaging methods typically use the scattered field or its simple transformations as network input, such as backpropagation (BP) images or other forms of pre-generated images. This fails to fully exploit the information inherent in the scattered field data itself and lacks mechanisms for deep feature extraction and adaptive fusion of different modal data (such as scattered field, current distribution, and semantic priors). This leads to problems such as overfitting and poor generalization ability if the initial data (such as BP images) is of low quality, and makes accurate target reconstruction difficult when measurement data is limited. In summary, existing methods often lack sufficient generalization ability and universality when facing practical challenges such as missing data modalities, changes in measurement scenarios, or unknown target attributes.
[0004] Diffusion models, as an emerging generative model, have demonstrated powerful learning and generative capabilities in fields such as image processing, speech recognition, and 3D reconstruction in recent years. However, when directly applied to electromagnetic inverse scattering tasks, they often struggle to achieve high-quality and high-generalization-rate generation results due to their difficulty in effectively extracting and adapting multimodal features.
[0005] Therefore, there is an urgent need to develop a reconstruction method that can effectively integrate multimodal electromagnetic features. This method can adapt to most conditions as input, and while maintaining the high accuracy of the generated model, it can further improve the universality and generalization performance of electromagnetic backscattering imaging. Summary of the Invention
[0006] This invention provides an electromagnetic backscattering imaging method based on a diffusion model with multimodal fusion features. It combines a diffusion model with a large language model and applies it to electromagnetic backscattering imaging tasks to achieve multi-scale generation and reconstruction of electromagnetic fields. This method introduces multimodal data, such as BP maps, current distributions, and electromagnetic scattering field data, as conditional inputs into the diffusion model, enabling the model to fully integrate observational information during the generation process, significantly improving the reconstruction accuracy and stability of the image. Simultaneously, the large language model introduces prior linguistic information, further enhancing the model's generalization ability under different conditions. Notably, the method described in this invention possesses inherent robustness to conditional inputs. During the training phase, the model learns the complementarity and intrinsic correlation between different modalities through a multimodal feature fusion mechanism. Therefore, during the inference phase, no structural modifications or parameter adjustments are required, allowing for flexible handling of situations where one or more conditional data (such as BP maps, current distributions, and textual descriptions) are missing. When some conditions are missing, the model can automatically rely on the remaining available conditional information (such as using only scattering field data) to complete effective reconstruction, significantly improving the method's practicality and adaptability in real-world complex scenarios.
[0007] The electromagnetic backscattering imaging method based on a diffusion model with multimodal feature embedding comprises the following steps:
[0008] Step (1): Under a specific two-dimensional electromagnetic backscattering imaging environment, delineate the region of interest containing the scatterer under test, and deploy a transmitting antenna and a receiving antenna array around the region of interest. Collect the scattering field data of the scatterer under test in free space through the antenna array, and combine it with the incident field information to construct the input-output sample pair of the model.
[0009] Step (2): Before the diffusion model training stage, features from the BP reconstructed image, induced current distribution, scattering field data, and textual semantic description are introduced through a multimodal feature fusion mechanism. Each modal data is input into its corresponding feature extraction network for encoding, and then adaptively fused using a cross-modal attention mechanism to generate fused features with both physical and semantic consistency. These features guide the diffusion model to more accurately recover the target dielectric constant distribution during denoising, improving reconstruction accuracy and generalization ability.
[0010] Step (3): Establish a diffusion model in two-dimensional space. A forward diffusion process is formed by gradually adding Gaussian noise to the samples. In the reverse process, a neural network guided by fusion features is used to predict the noise distribution and gradually denoise to restore the original samples, thereby realizing the two-dimensional reconstruction and generation of the dielectric constant of the target region. The generalization performance of the model is verified by testing across datasets.
[0011] The sample is the target dielectric constant.
[0012] Step (1) describes the deployment of transmitting and receiving antenna arrays in a specific two-dimensional electromagnetic backscattering imaging environment to acquire the scattering field data of the scattering object under test, as follows:
[0013] The two-dimensional electromagnetic backscattering imaging region is divided into two areas: the region of interest (ROI) and the measurement region. Against a free-space background, a rectangular region is selected as the ROI to accelerate large matrix operations using the Fast Fourier Transform (FFT). The measurement region is located outside the ROI, with transmitting and receiving antennas uniformly deployed. A known incident field is applied to the ROI. The scattered field data of the region of interest were measured. .
[0014] During the simulation modeling phase, scattering field data Calculated based on the electromagnetic scattering equation, its expression is as follows:
[0015] (Formula 1)
[0016] in, This represents the mapping operator from the induced current in the region of interest to the scattered field in the measurement region. This represents the mapping operator between the induced current and the internal scattered field within the region of interest. Indicates the incident field within the region of interest. The dielectric constant is represented by a two-dimensional distribution in the region of interest.
[0017] Based on this, the scattered field data obtained through calculation Together with the corresponding dielectric constant distribution ε, they constitute training sample pairs. ,in As input to the diffusion model, ε serves as the output label for the diffusion model, guiding the learning of the diffusion model and the reconstruction of the two-dimensional electromagnetic field.
[0018] Optionally, the training sample pair further includes one or more of the following modal data: a BP reconstructed image obtained based on the scattered field data using a backpropagation algorithm, the induced current distribution J within the region of interest, and a textual semantic description L of the scattering object to be imaged. The scattered field data, the backpropagation BP reconstructed image, the induced current distribution, the textual semantic description, and the ground truth value of the medium parameter distribution together constitute a multimodal training sample pair. The data sources for the training sample pair include simulation data and measured data.
[0019] Step (2) introduces fusion features during the training phase of the diffusion model, as follows:
[0020] Scattering field data The BP reconstructed image, induced current distribution J, and textual semantic description L are respectively fed into their corresponding feature extractors for encoding, then fused through a cross-modal attention mechanism, and finally adapted by convolution to be combined with noisy samples. Consistent size fusion features The text semantic description L is extracted using a pre-trained large language model to better utilize the prior knowledge of the large language model and the relationships between semantics.
[0021] Step (3) involves establishing a diffusion model in two-dimensional space to achieve two-dimensional reconstruction and generation of the dielectric constant of the target region, as detailed below:
[0022] A diffusion model is established in two-dimensional space to learn the scattering field data. The fusion features obtained from BP reconstructed image, induced current distribution J, and textual semantic description L The implicit mapping relationship between dielectric constant distribution and the process involves two stages: forward diffusion and reverse denoising.
[0023] 3-1. Forward diffusion process:
[0024] The forward process is a progressively noisy Markov process used to gradually perturb the original sample ε0 (i.e., the true dielectric constant distribution) into Gaussian noise. Its expression is:
[0025] , (Formula 2)
[0026] in, This represents the noisy sample at step t. The decay coefficient corresponding to time step t, The sample noise follows a standard normal distribution. After T-step diffusion, the sample distribution gradually approximates a standard Gaussian distribution.
[0027] Through this process, the model can learn the mapping from ordered dielectric distribution to random noise, laying a statistical foundation for subsequent reverse generation.
[0028] 3-2. Reverse denoising process:
[0029] The reverse process uses a neural network to progressively remove noise, recovering the target dielectric constant distribution from the Gaussian distribution samples. The probability distribution generated in the reverse process is shown below. Represented as:
[0030] , (Formula 3)
[0031] in, and These represent the mean and variance predicted by the neural network, respectively. The fusion features of the observation data corresponding to the target are used to guide the model to learn a feature distribution consistent with physical observations during the generation process.
[0032] In each step of reverse generation, the model predicts the noise term of the current sample. And calculate the denoised samples based on the prediction results:
[0033] , (Formula 4)
[0034] By repeatedly performing this denoising process, the model eventually generates a two-dimensional image representation that approximates the true dielectric constant distribution ε0.
[0035] This invention employs a diffusion model constructed based on a two-dimensional U-Net architecture. The network consists of an encoder, a bottleneck layer, and a decoder. The encoder extracts multi-scale spatial features through downsampling, while the decoder gradually restores spatial resolution and preserves detailed information through upsampling combined with skip connections. To enhance the modeling capability of electromagnetic data, this fused feature is appended to noisy samples at each denoising step, guiding the diffusion model to gradually recover the target dielectric constant distribution. This fully utilizes the observation field information and multi-scale structural features, significantly improving the accuracy, stability, and generalization performance of imaging.
[0036] Compared with the prior art, the significant advantages of this invention lie in its superior generalization performance and universality:
[0037] Robustness to missing input data: During the training phase, the model learns the inherent complementarity between different modalities through a multimodal feature fusion mechanism. Therefore, during the inference (deployment) phase, when one or more conditional data (such as BP plots, current distributions, and text descriptions) are partially missing or of poor quality, the model can automatically rely on the remaining valid information to complete high-quality reconstruction without adjusting the model structure or parameters, greatly improving its practicality in real-world complex scenarios.
[0038] Cross-dataset generalization ability: As demonstrated in the experiments of the specific implementation, a model trained on a single dataset (such as single-person pose data) can be directly applied to another dataset with a different distribution (such as two-person pose data) and maintain excellent performance. This proves that what this invention learns is the essential characteristic of the electromagnetic inverse scattering problem, rather than overfitting to a specific training set. Attached Figure Description
[0039] Figure 1 This is a schematic diagram of the specific process of the method of the present invention.
[0040] Figure 2 This is a schematic diagram of the arrangement of the receiving and transmitting antenna arrays in the method of the present invention.
[0041] Figure 3This is a schematic diagram of the training and generation process of the network framework constructed in the method of the present invention.
[0042] Figure 4 This is a schematic diagram of the reconstructed two-dimensional scene result of the present invention. Detailed Implementation
[0043] The present invention will now be described in further detail with reference to the accompanying drawings.
[0044] like Figure 1 As shown, this invention proposes an electromagnetic inverse scattering imaging method based on a diffusion model with multimodal feature embedding. The method consists of three main steps: acquisition of multimodal electromagnetic feature data and sample construction, extraction and fusion of multimodal conditional features, and reconstruction and generation of conditional diffusion models.
[0045] Step (1), Multimodal electromagnetic feature data acquisition and sample construction: In the two-dimensional electromagnetic backscattering imaging scenario, a transmitting and receiving antenna array is set up around the target area to collect the scattered field observation data generated by the target; Combine the incident field information and the corresponding true value of the medium parameter distribution to construct a training sample pair containing at least one modal condition among the scattered field observation data, backpropagation reconstructed image, induced current distribution and text semantic description.
[0046] Step (2), Extraction and fusion of multimodal conditional features: The multimodal conditional data in the training sample pairs are respectively input into the corresponding feature extraction networks for encoding, and adaptive fusion is performed through cross-modal attention mechanism to generate a unified fusion feature vector;
[0047] Step (3), Reconstruction and generation of conditional diffusion model: Construct a diffusion generation model based on the fused feature vector; the model adds noise to the true value of medium parameters through the forward diffusion process, and in the reverse denoising process, it uses the fused feature vector to guide the denoising neural network to predict noise, and gradually reconstructs the medium parameter distribution image of the target area.
[0048] Step (1) involves setting up transmitting and receiving antennas around the region of interest to acquire the target's scattered field data, as detailed below:
[0049] 1-1. In the implementation of this invention, to construct the dataset for the two-dimensional electromagnetic backscattering imaging task, 10,000 two-dimensional digital images from the publicly available MNIST dataset were selected as the base samples. First, the digital shape regions in the original images were regarded as targets, and random relative permittivity values were assigned to them, typically between 1 and 2, to simulate weak scatterers. Simultaneously, several dots were randomly superimposed on the background region and assigned random permittivity values to increase scene complexity, enabling the model to learn more diverse permittivity distribution characteristics.
[0050] 1-2. To verify the generalization ability of this invention in complex real-world testing scenarios, a publicly available human posture electromagnetic scattering dataset was used for testing. This dataset, first published by Hongrui Zhang et al. in "Semantic regularization of electromagnetic inverse problems," contains three subsets: "single-person random," "single-person intermediate," and "multi-person intermediate," totaling 30,900 samples, covering a variety of human postures and electromagnetic scattering environments from simple to complex. This dataset was collected in a complex indoor environment, where target morphology, relative position, and electromagnetic interaction with the environment are highly variable, providing a challenging benchmark for evaluating the generalization ability of this method in out-of-distribution scenarios.
[0051] The fusion feature extraction model described in step (2) is as follows:
[0052] 2-1. For scattering field data This invention designs a dedicated Scattering Field Encoder. This encoder first processes the original complex scattering field matrix through a multi-layer one-dimensional convolutional network to extract its local frequency domain features. Then, adaptive average pooling is used to unify the feature length to 16, and a linear projection layer maps the feature dimension to the hidden dimension of a Transformer. Finally, sinusoidal position encoding is introduced, and global dependency modeling is performed through a multi-layer Transformer encoder, outputting a depth feature representation with a dimension of 256, i.e., the scattering field depth feature.
[0053] 2-2. For BP image reconstruction, this invention employs a deep convolutional coding network (BPCNN) to process two-dimensional image data. The network contains multiple convolutional blocks, each consisting of a double convolutional layer, batch normalization, ReLU activation, and 2×2 max pooling, progressively downsampling the feature map to (256, 4, 4). After unifying the output size through adaptive pooling, the features are flattened and projected onto a 256-dimensional feature space, i.e., the BP image depth features; effectively preserving the spatial structure information of the image, such as edges and textures.
[0054] 2-3. For the induced current distribution (J), this invention designs an encoding scheme (CurrentDensityEncoder) that emphasizes both dimensionality reduction and feature extraction to address its high-dimensional characteristics. First, a multi-layer fully connected network is used to compress the high-dimensional spatial points to 64 dimensions, significantly reducing computational complexity. Then, the dimensionality-reduced feature matrix is reconstructed into a two-dimensional feature map, and spatial-frequency joint features are further extracted through convolutional pooling operations, finally outputting a compact 256-dimensional representation, i.e., the depth feature of the current distribution.
[0055] 2-4. For text semantic features (text_indices), this invention constructs a text encoder based on a pre-trained language model. The (20, 768) dimensional text embedded by BERT is first projected onto a 256-dimensional feature space, and after adding positional encoding, semantic dependencies are captured through a multi-layer Transformer encoder. A masked weighted average pooling strategy is adopted to dynamically aggregate sequence features according to the effective length of the text, finally generating a 256-dimensional semantic representation aligned with the electromagnetic feature dimension, i.e., the text deep feature.
[0056] 2-5. The modal features are deeply fused using a cross-modal attention fusion module. The scattering field depth features, BP image depth features, current distribution depth features, and optional text depth features are mapped to the same-dimensional feature space through a linear projection layer. The correlation weights between the modal features are calculated using a cross-modal multi-head attention mechanism, and adaptive weighted fusion is performed based on these weights. The fused features are then further refined by a multi-layer global Transformer encoder, ultimately outputting a 256-dimensional globally fused feature containing multimodal information, providing enhanced conditional input for the subsequent diffusion model.
[0057] Step (3) involves establishing a diffusion model on a two-dimensional scene, as detailed below:
[0058] During the training phase, a real two-dimensional dielectric constant distribution image is randomly selected from the training sample set. and in time step Randomly sample one diffusion step within the range Gaussian noise was gradually added to the real samples according to Equation 2 of the diffusion process to generate noisy samples. This invention achieves inverse denoising and image reconstruction based on U-Net. At the model input end, noisy samples are... Corresponding time step code and the fusion characteristics of the target Simultaneously input into the network, where the fused features are directly added... Above. The network outputs a predicted noise term after forward propagation. The denoising and reconstruction are then performed according to Formula 4 of the reverse diffusion process. This process evolves gradually from high-noise samples to low-noise samples, ultimately reaching a state where... The distribution of dielectric constant is recovered to match the true dielectric constant. Approximate reconstruction results.
[0059] During the training phase, the loss function is calculated in each iteration based on the mean square error between the predicted noise and the actual noise. The network parameters are updated using the backpropagation algorithm. As the iteration progresses, the network continuously optimizes its denoising capabilities, enabling it to accurately recover the two-dimensional dielectric constant characteristics under different noise levels until the model converges.
[0060] Table 1 Comparison of Reconstruction Methods
[0061]
[0062] To verify the generalization performance of the model, this invention uses single-person test data for model training. After training, this invention uses two-person test data as the detection standard. The generation effect and accuracy of the two-person model were tested on the final model, and tested on some traditional methods using the same training method. The generation comparison results are shown in Table 1. As shown in Table 1, the "Semantic Regularization" method is an electromagnetic inverse scattering imaging method based on semantic regularization, which is used as a baseline for comparison. Experimental results show that in the highly challenging cross-dataset generalization test (the model is trained on single-person data and tested on two-person data), the baseline method "Semantic Regularization" has a generalization success rate of 0%, indicating that it is severely overfitting to the training data distribution. In contrast, the method of this invention achieves a generalization success rate of 70%, which is significantly better than the baseline method, fully demonstrating the strong generalization ability of the proposed framework when dealing with out-of-distribution data.
Claims
1. An electromagnetic backscattering imaging method based on a diffusion model with fusion features, characterized in that... Includes the following steps: Step (1), Multimodal electromagnetic feature data acquisition and sample construction: Under a specific two-dimensional electromagnetic backscattering imaging environment, the region of interest containing the scattering object to be tested is delineated, and a transmitting antenna and receiving antenna array are set up around the region of interest; The scattered field data of the scattering body under test in free space is collected by an antenna array. Combined with the incident field information and the true value of the corresponding medium parameter distribution, the input-output sample pair of the model is constructed. Step (2), Extraction and fusion of multimodal conditional features: The multimodal conditional data in the training sample pairs are respectively input into the corresponding feature extraction network for encoding, and adaptive fusion is performed through cross-modal attention mechanism to generate fusion features with physical and semantic consistency, which are used to guide the diffusion model to more accurately recover the target dielectric constant distribution in the denoising process; Step (3) Reconstruction and generation of conditional diffusion model: Construct a diffusion model with fusion features as conditions; the model adds noise to the true value of medium parameters through forward diffusion process, and in the reverse denoising process, the fusion features are used to guide the denoising neural network to predict noise, and gradually reconstruct the medium parameter distribution image of the region of interest.
2. The electromagnetic backscattering imaging method based on a diffusion model with multimodal fusion features according to claim 1, characterized in that, Step (1) is as follows: The two-dimensional electromagnetic backscattering imaging region is divided into two regions: the region of interest (ROI) and the measurement region. Against a free-space background, a rectangular region is selected as the ROI. The measurement region is located outside the ROI, and transmitting and receiving antennas are uniformly deployed there. A known incident field is applied to the ROI. The scattered field data of the target area were measured. ; During the simulation modeling phase, scattering field data Calculated based on the electromagnetic scattering equation, its expression is as follows: (Official 1) in, This represents the mapping operator from the induced current in the region of interest to the scattered field in the measurement region. This represents the mapping operator between the induced current and the internal scattered field within the region of interest. Indicates the incident field within the region of interest. The dielectric constant is distributed in two dimensions in the region of interest. The scattered field data obtained through calculation Together with the corresponding dielectric constant distribution ε, they constitute training sample pairs. ,in As input to the diffusion model, ε serves as the output label for the diffusion model, guiding the learning of the diffusion model and the reconstruction of the two-dimensional electromagnetic field.
3. The electromagnetic backscattering imaging method based on a diffusion model with multimodal fusion features according to claim 1, characterized in that, The training sample pair includes one or more of the following modal data: a backpropagation reconstructed image obtained from the scattered field observation data through the backpropagation algorithm, the induced current distribution in the region of interest, and a textual semantic description; the scattered field observation data, the backpropagation reconstructed image, the induced current distribution, the textual semantic description, and the ground truth value of the medium parameter distribution together constitute a multimodal training sample pair.
4. The electromagnetic backscattering imaging method based on a diffusion model with multimodal fusion features according to claim 1, characterized in that, The data sources for the training sample pairs include simulation data and measured data.
5. The electromagnetic backscattering imaging method based on a diffusion model with fusion features according to claim 1, characterized in that, Step (2) specifically includes: 2-1. Scattered field feature encoding: The scattered field data is input into a one-dimensional convolutional coding network to extract its local frequency domain features. The scattered field depth features are obtained by global modeling through positional encoding and multi-layer Transformer encoder. 2-2. Backpropagation Reconstructed Image Feature Encoding: The backpropagation reconstructed image is input into a two-dimensional convolutional coding network, and its spatial structure features are extracted through multi-layer convolution and pooling operations to obtain the depth features of the BP image; 2-3. Current distribution feature encoding: The induced current distribution data is first reduced in dimensionality through a multi-layer fully connected network, and then its features are extracted through a convolutional coding network to obtain the current distribution depth features; 2-4. Text semantic feature encoding: When the text semantic description exists, the text semantic features are extracted through a pre-trained large language model and its encoder to obtain the text deep features; 2-5. Cross-modal feature fusion: The scattering field depth features, BP image depth features, current distribution depth features, and optional text depth features are mapped to the same-dimensional feature space through a linear projection layer; the correlation weights between each modality feature are calculated using a cross-modal multi-head attention mechanism, and adaptive weighted fusion is performed based on the weights to generate a unified fused feature.
6. The two-dimensional electromagnetic inverse scattering imaging method based on a diffusion model with multimodal fusion features according to claim 5, characterized in that... Step (3) is as follows: A diffusion generation model is established in two-dimensional space to learn the implicit mapping relationship between multimodal fusion features and dielectric constant distribution; the process includes two stages: forward diffusion and reverse denoising. 3-1. Forward diffusion process: The forward process is a progressively noisy Markov process used to improve the quality of the original samples. The gradual perturbation is in the form of Gaussian noise, and its expression is: , (Official 2) in, Indicates the first Noisy samples of the step, To keep pace with time The corresponding attenuation coefficient, For random noise that follows a standard normal distribution; after After the first diffusion step, the sample distribution gradually approximates the standard Gaussian distribution; 3-2. Reverse denoising process: The reverse process gradually removes noise through a neural network, recovering the target dielectric constant distribution from the Gaussian distributed samples. The probability distribution generated in the reverse process is expressed as: (Official 3) in, and These represent the mean and variance predicted by the neural network, respectively. The fusion features of the observation data corresponding to the target are used to guide the model to learn a feature distribution consistent with physical observations during the generation process; In each step of reverse generation, the model predicts the noise term of the current sample. And calculate the denoised samples based on the prediction results: , (Official 4) By repeatedly performing this denoising process, the model eventually generates a two-dimensional image representation that approximates the true dielectric constant distribution ε0.
7. The two-dimensional electromagnetic inverse scattering imaging method based on a diffusion model with multimodal fusion features according to claim 5, characterized in that, The denoising neural network adopts a U-Net architecture, including an encoder, a bottleneck layer, and a decoder. The encoder is used to downsample and extract multi-scale features, and the decoder is used to upsample and combine skip connections to recover spatial details. The fused features are used as conditional inputs in each layer of the denoising neural network to guide the denoising process. During the training phase, the loss function is calculated based on the mean square error between the predicted noise and the actual noise in each iteration, and the network parameters are updated through the backpropagation algorithm. As the iterations continue, the network continuously optimizes its denoising capabilities, enabling it to accurately recover the two-dimensional dielectric constant characteristics under different noise levels until the model converges.