An agricultural remote sensing image defogging method and system based on a multi-scale hole map network and frequency domain physical perception, and an electronic device

By combining multi-scale dilated convolution and dynamic graph networks with frequency domain physical sensing, the problem of large-scale fog occlusion in remote sensing images was solved, achieving more efficient defogging and image quality improvement.

CN122115265APending Publication Date: 2026-05-29NORTHEAST INST OF GEOGRAPHY & AGRIECOLOGY C A S

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NORTHEAST INST OF GEOGRAPHY & AGRIECOLOGY C A S
Filing Date
2026-03-10
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies in remote sensing image processing suffer from problems such as limited receptive field, difficulty in dealing with large-scale thick fog obscuration, lack of global context reasoning, and insufficient utilization of physical mechanisms, resulting in poor image interpretation and analysis effects.

Method used

We employ multi-scale dilated convolution to expand the receptive field, combine it with dynamic graph networks to capture global dependencies, and perform refined reconstruction through frequency domain physical perception. This leads to the construction of a dehazing method based on multi-scale dilated graph networks and frequency domain physical perception, which utilizes implicit fog layer attention maps to adjust feature fusion and frequency domain enhancement.

Benefits of technology

It significantly expands the receptive field, solves the problem of texture breakage caused by large-area fog obscuration, improves the color fidelity and texture clarity of images, and achieves a more efficient dehazing effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115265A_ABST
    Figure CN122115265A_ABST
Patent Text Reader

Abstract

The application provides an agricultural remote sensing image defogging method and system based on a multi-scale hollow map network and a frequency domain physical perception and an electronic device, and comprises the following steps: acquiring a paired foggy and non-foggy agricultural remote sensing image dataset in a preset area, and constructing an end-to-end remote sensing image dataset without a mask; based on the remote sensing image dataset, a U-Net type encoder-decoder structure with a multi-scale hollow perception module, a dynamic graph neural network module and a physical guided feature extractor is adopted to construct a defogging model; wherein the implicit fog layer attention map obtained by the physical guided feature extractor is used to generate a gate weight and adjust the frequency domain enhancement intensity, and cross-module collaborative defogging is realized. Based on the defogging model, the foggy agricultural remote sensing image to be processed is defogged to obtain a defogged image. The application can effectively remove thick fog and shadows without fog layer labeling, and can maintain the continuity and realism of the ground texture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing image processing and computer vision technology, specifically relating to a method, system and electronic equipment for dehazing agricultural remote sensing images based on multi-scale hole map networks and frequency domain physical perception. Background Technology

[0002] Optical remote sensing imagery is of great value in fields such as Earth observation, environmental monitoring, and military reconnaissance. However, due to atmospheric conditions, remote sensing images are often affected by atmospheric scattering, such as haze, fog, and thick fog, leading to the loss of ground feature information and seriously affecting subsequent image interpretation and analysis.

[0003] Currently, the mainstream technical solutions in this field mainly cover several categories: First, methods based on physical models, such as dark channel priors. These methods rely on ideal atmospheric scattering models, assuming that the fog layer is a linear fog-like obstruction. When facing thick fog with a transmittance close to 0, division operations are prone to numerical instability and color distortion. Second, methods based on traditional convolutional neural networks (CNNs). Although they can achieve dehazing through learning, the local receptive field of convolution operations is limited, making it difficult to use information from distant fog-free areas in the image to repair large areas of fog obstruction, often leading to structural problems such as "river breaks" and "road disappearances". Third, methods based on generative adversarial networks (GANs). Although they can generate realistic textures, they are prone to producing "illusionary" content and lack physical interpretability and structural fidelity.

[0004] In summary, existing technologies have three major drawbacks: First, the receptive field is limited, making it difficult to cope with large-scale thick fog obstruction; second, there is a lack of global context reasoning, making it impossible to effectively utilize non-local information to repair texture breaks; and third, the physical mechanisms are not fully utilized, and the purely data-driven model has poor robustness under complex atmospheric conditions. Summary of the Invention

[0005] This invention addresses the shortcomings and deficiencies of existing technologies by providing a method, system, and electronic device for dehazing agricultural remote sensing images based on multi-scale dilated graph networks and frequency domain physical perception. This method employs multi-scale dilated convolution to expand the receptive field, utilizes dynamic graph networks to capture global dependencies, and combines frequency domain physical priors for refined reconstruction. The implicit fog layer interest map serves as a unified scheduling signal, simultaneously constraining the information injection range of dynamic graph inference and adjusting the frequency domain enhancement intensity, thus distinguishing it from the simple module stacking of existing technologies.

[0006] To achieve the above objectives, the present invention provides the following solution: A method for dehazing agricultural remote sensing images based on multi-scale hole map networks and frequency domain physical sensing includes: Acquire paired agricultural remote sensing image datasets with and without fog within a preset area, and construct a maskless end-to-end remote sensing image dataset. Based on the remote sensing image dataset, a dehazing model is constructed using a U-Net encoder-decoder structure that incorporates a multi-scale hole perception module, a dynamic graph neural network module, and a physically guided feature extractor. Based on the aforementioned defogging model, the foggy agricultural remote sensing image to be processed is defogging to obtain a defogging image.

[0007] Preferably, the defogging model includes: A multi-scale hole sensing module, located at the front end of the encoder, includes three convolutional branches for extracting multi-scale features of remote sensing images. The encoder is used to downsample the multi-scale features to obtain a deep semantic feature map; The first dynamic graph neural network module, embedded in the bottleneck layer of the U-Net architecture, is used to transform the deep semantic feature map into a topological graph structure to obtain a fused feature map. A decoder is used to upsample the fused feature map and obtain a decoded feature map using a second dynamic graph neural network module embedded in the intermediate layer; A physical-guided feature extractor, located at the end of the decoder, is used to regress implicit physical parameters based on the decoded feature map. The implicit physical parameters include at least a transmittance map, an atmospheric light map, and an implicit fog layer interest map. Based on the implicit fog layer interest map, a gating weight is generated to adjust the feature fusion intensity of the dynamic graph neural network module. At the same time, affine feature modulation is performed in the spatial domain and high-low frequency separation processing is performed in the frequency domain to generate a fog-free image.

[0008] Preferably, both the first dynamic graph neural network module and the second dynamic graph neural network module include: Dynamic layer construction is used to divide the input feature map into non-overlapping blocks, map them into node vectors, calculate the cosine similarity between nodes, select the nodes with the highest cosine similarity as the neighbor set, and obtain the graph structure; wherein, each node vector corresponds to a block region in the input feature map, and the node vectors updated by graph attention are backfilled and rearranged according to the original block position to form a graph network output feature map of the same size as the input feature map. The graph attention aggregation layer is used to update node features using attention coefficients based on the graph structure to obtain graph network branch features; An adaptive gated fusion layer is used to perform element-wise gated weighted fusion of the initial convolutional features input to this module and the graph network branch features after aligning them in the same spatial position to obtain fused features. The initial convolutional features participate in the fusion as identity mapping branches that preserve local texture. The gate weights are obtained by mapping the interest graph of the implicit fog layer and are used to limit the range of graph network information injection and reduce artifacts in clear areas.

[0009] Preferably, the method for generating fog-free images includes: Spatial domain fog-aware modulation is applied to the decoded feature map to obtain the modulation features; The frequency components of the decoded feature map are separated using a low-pass filter operator; By fusing the modulation features and the frequency components, a frequency domain enhancement feature is obtained; wherein the frequency components include low-frequency components and high-frequency components, and the high-frequency gain coefficient is obtained by mapping the transmittance map or the implicit fog layer interest map and clipped to a specified range; Based on the frequency domain enhancement features, the fog-free image is obtained using convolutional layers and activation functions.

[0010] Preferably, the method for obtaining the modulation features includes: Convolutional layers are used to regress and predict potential transmittance maps, atmospheric light maps, and implicit fog interest maps from decoded feature maps; The implicit fog concern graph is normalized and cropped to 0 to 1 to obtain the gating weight, which is used to characterize the feature injection intensity of the fog-occluded region. Based on the predicted transmittance map and atmospheric light map, affine coefficients are generated; wherein the affine coefficients include scaling coefficients and translation coefficients, and the scaling coefficients and translation coefficients are aligned with the decoded feature map in terms of spatial size and number of channels; Based on the gating weights and the affine coefficients, spatial domain fog-aware modulation is performed on the decoded feature map; wherein the decoded feature map is normalized and then scaled and translated to obtain intermediate modulation features, and the intermediate modulation features and the original decoded feature map are fused element-wise by the gating weights to obtain the modulation features.

[0011] Preferably, the difference between the generated fog-free image and the real fog-free image is calculated by using a hybrid loss function, and the network parameters of the defogging model are updated by backpropagation; wherein the hybrid loss function includes pixel domain loss, structural domain loss and perceptual domain loss.

[0012] This invention also provides an agricultural remote sensing image dehazing system based on multi-scale hole map networks and frequency domain physical sensing, for implementing the method, comprising: The data acquisition module is used to acquire paired agricultural remote sensing image datasets with and without fog within a preset area, and to construct a maskless end-to-end remote sensing image dataset. The dehazing model construction module is used to construct a dehazing model based on the remote sensing image dataset, using a U-Net encoder-decoder structure that incorporates a multi-scale hole perception module, a dynamic graph neural network module, and a physically guided feature extractor. The dehazing image generation module is used to dehaze the foggy agricultural remote sensing image to be processed based on the dehazing model, and obtain the dehazing image.

[0013] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the computer program is executed by the processor to perform the aforementioned method for dehazing agricultural remote sensing images based on multi-scale hole map networks and frequency domain physical sensing.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) Multi-scale perception and maskless strategy: This invention significantly expands the effective receptive field without increasing the computational burden through the three-way parallel dilated convolution of the MSDilate module, enabling it to capture large-scale fog distribution features. At the same time, it abandons the dependence on manual masking, achieving fully end-to-end adaptive dehazing and reducing data acquisition costs.

[0015] (2) Non-local texture restoration capability: To address the texture breakage problem caused by large-area thick fog occlusion, this invention utilizes dual-level dynamic graph inference (Dual GNN). By dynamically constructing a k-nearest neighbor graph, the occluded area can "see" similar fog-free textures in the distance. Combined with CloudGate adaptive gating, global restoration information is introduced only in the fog area, effectively solving the problems of "broken river" and "broken road".

[0016] (3) Spatial-Frequency Joint Physical Enhancement: Instead of using simple linear subtraction, this invention proposes a feature modulation mechanism based on implicit physical parameters. The CAF module restores contrast in the spatial domain using affine transformation, while the FGVS module separates and suppresses low-frequency fog components and enhances high-frequency details in the frequency domain. This physically guided design results in images that outperform pure data-driven models in both color fidelity and texture clarity. Attached Figure Description

[0017] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of the overall process of the agricultural remote sensing image dehazing algorithm based on multi-scale hole map network and frequency domain physical perception in this embodiment. Figure 2 This is a structural framework diagram of the deep network model in this embodiment; Figure 3 This is a schematic diagram of the CAF module in this embodiment; Figure 4 This is a schematic diagram of the FGVS module in this embodiment; Figure 5This is a dehazing effect diagram of the algorithm of the present invention in this embodiment; Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Attached image description: 1010, Processor; 1020, Memory; 1030, Input / Output Interface; 1040, Communication Interface; 1050, Bus. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] Example 1: like Figure 1 As shown, a method for dehazing agricultural remote sensing images based on multi-scale hole map networks and frequency domain physical sensing includes: S1: Obtain paired agricultural remote sensing image datasets with and without fog within a preset area, and construct a maskless end-to-end remote sensing image dataset. This embodiment uses publicly available remote sensing dehazing datasets such as RRSHID. Unlike existing technologies, this invention does not load fog mask files at all during the training phase, and only uses foggy images (Input) and fog-free reference images (Target).

[0023] Specifically, we acquire paired remote sensing image datasets with and without fog. Without using manual masking, we preprocess the datasets only through normalization and geometric noise data augmentation to construct maskless end-to-end training data as input to the model.

[0024] In step S1, data preprocessing specifically includes: Normalization processing: Normalization of the pixel values ​​of the acquired remote sensing images The linear mapping to the interval [-1, 1] is given by the following formula: in, In this embodiment .

[0025] Geometry and noise enhancement: During the training phase, random cropping, random 90-degree rotation, horizontal flipping, vertical flipping, and Gaussian noise injection (variance 5.0-20.0) are applied simultaneously to pairs of foggy and fog-free images to improve the robustness of the model to different fog morphologies and sensor thermal noise. In order to maintain the spectral characteristics of remote sensing images, color jitter is not used in this embodiment.

[0026] S2: Based on remote sensing image datasets, a U-Net encoder-decoder structure is constructed by introducing a multi-scale hole perception module, a dynamic graph neural network module, and a physically guided feature extractor.

[0027] A further implementation method is that the defogging model includes: The multi-scale hole sensing module (MSDilateBlock), located at the front end of the encoder, includes three convolutional branches for extracting multi-scale features from remote sensing images; such as... Figure 2 As shown, the specific processing procedure is as follows: Input feature X passes through three branches in parallel, output feature The calculation formula is: in, The dilation rate is expressed as of Convolution operation, This indicates a channel splicing operation. Indicates the use of feature fusion Convolutional layer. Maps the number of channels back to the base channel (192 in this example).

[0028] The encoder is used to downsample multi-scale features to obtain a deep semantic feature map; the encoder contains successive downsampling modules, each consisting of a step size of 2. It consists of convolutional layers and residual blocks. Deep semantic features are extracted progressively. After each downsampling, the feature map size is halved and the number of channels is doubled, until the bottleneck layer is reached.

[0029] The first dynamic graph neural network module, embedded in the bottleneck layer of the U-Net architecture, is used to transform deep semantic feature maps into a topological graph structure to obtain a fused feature map. The second dynamic graph neural network module is embedded in the intermediate layer of the decoder. A further implementation wherein both the first and second dynamic graph neural network modules include: Dynamic layer construction is used to divide the input feature map into non-overlapping blocks, map them to node vectors, and calculate the cosine similarity between nodes. The nodes with the highest cosine similarity are selected as a neighbor set to obtain the graph structure. Each node vector corresponds to a block region in the input feature map. After graph attention updates, the node vectors are backfilled and rearranged according to their original block positions to form a graph network output feature map of the same size as the input feature map. Specifically, the node... With nodes Cosine similarity between The formula is: Select The largest Each node is used as a neighbor set. In this embodiment, the following is selected: The largest Each node is used as a neighbor set. .

[0030] The graph attention aggregation layer is used to update node features and obtain graph network branch features based on the graph structure and attention coefficients; specifically, it utilizes attention coefficients... Update node features : in, For learnable weight matrix, For attention vectors, Indicates splicing, This is the residual scaling factor. In this embodiment, the number of attention heads is set to 12, and the node feature dimension is 512 to ensure that the model has sufficient semantic representation capabilities.

[0031] An adaptive gated fusion layer is used to perform element-wise gated weighted fusion of the initial convolutional features input to this module and the graph network branch features after aligning them in the same spatial position to obtain fused features. The initial convolutional features participate in the fusion as identity mapping branches that preserve local texture. The gate weights are obtained by mapping the interest graph of the implicit fog layer and are used to limit the range of graph network information injection and reduce artifacts in clear areas.

[0032] Computational fusion features : in, It is a gated convolutional layer. It is the Sigmoid activation function. For element-wise multiplication, This represents the initial convolutional feature map input to this module; The feature map is obtained by inverse mapping after updating the features of the graph network nodes. The two are then aligned in size and channel before participating in gated fusion; the gate weights This is used to control the intensity of information injection into the graph network, thereby enhancing the structural recovery capability of fog-covered areas while maintaining clear regional details.

[0033] The decoder upsamples the fused feature map and uses a second dynamic graph neural network module embedded in the intermediate layer to obtain the decoded feature map; the decoder contains successive upsampling modules, each consisting of a step size of 2. It consists of deconvolution and residual blocks. The dynamic graph neural network module is deployed in the bottleneck layer of the network and the intermediate layer after the first upsampling of the decoder, forming a two-level global inference architecture to capture long-distance texture dependencies at different resolution scales.

[0034] A physical-guided feature extractor, located at the end of the decoder, is used to regress implicit physical parameters based on the decoded feature map. The implicit physical parameters include at least a transmittance map, an atmospheric light map, and an implicit fog layer interest map. Based on the implicit fog layer interest map, a gating weight is generated to adjust the feature fusion intensity of the dynamic graph neural network module. At the same time, affine feature modulation is performed in the spatial domain and high-low frequency separation processing is performed in the frequency domain to generate a fog-free image.

[0035] A further implementation method includes generating a fog-free image by: Spatial domain fog-aware modulation is applied to the decoded feature map to obtain modulation features. A further implementation method for obtaining modulation features includes: using convolutional layers to regress and predict potential transmittance maps, atmospheric light maps, and implicit fog interest maps from the decoded feature map; calculating adaptive weighting coefficients based on the implicit fog interest map; wherein the implicit fog interest map is normalized and cropped to 0 to 1 to obtain gating weights, which are used to characterize the feature injection intensity in the fog-occluded region. Affine coefficients are generated based on the predicted transmittance map and atmospheric light map; wherein the affine coefficients include scaling and translation coefficients, and the scaling and translation coefficients are aligned with the decoded feature map in terms of spatial size and number of channels.

[0036] Based on the gating weights and the affine coefficients, spatial domain fog-aware modulation is performed on the decoded feature map; wherein the decoded feature map is normalized and then scaled and translated to obtain intermediate modulation features, and the intermediate modulation features and the original decoded feature map are fused element-wise by the gating weights to obtain the modulation features.

[0037] Specifically, using Convolutional layers directly regress from decoder features to predict the underlying transmittance map. Atmospheric light map and implicit fog layer attention map This process is entirely based on self-supervised learning and requires no external truth mask supervision.

[0038] like Figure 3 As shown, spatial domain fog-aware modulation (CAF): the spatial domain fog-aware modulation module decodes feature maps. With implicit physical parameter input As input. The implicit physical parameter input. The image is obtained by stitching together the transmittance map and the atmospheric light map, and then input into a parameter mapping network to generate scaling factors. and Translation coefficient For the input feature map Normalization was performed to obtain the above. With scaling factor Element-wise multiplication yields the scaling feature, which is then combined with the translation coefficient. Element-wise addition is performed to obtain the fused feature map. For the fused feature map conduct Convolution yields transform features and the input feature map The output feature map is obtained by performing residual summation. The residual addition is used to introduce physically guided modulation information while preserving the original structural information; based on the predicted physical parameters and Generate affine coefficients , The characteristic modulation formula is: in, For normalization operations, For implicit fog-based attention graphs The calculated adaptive weighting coefficients.

[0039] like Figure 4 As shown, a low-pass filter operator is used to separate the frequency components of the decoded feature map; specifically, frequency domain separation processing (FGVS) utilizes a low-pass filter operator. Separate frequency components, i.e. ; By fusing modulation features and frequency components, frequency domain enhancement features are obtained. The frequency components include low-frequency and high-frequency components, and the high-frequency gain coefficient is obtained by mapping the transmittance map or the implicit fog layer interest map and cropping it to a specified range; Frequency domain enhancement and reconstruction: in, For dehazing modulation based on physical parameters, This is the high-frequency gain coefficient. (In this embodiment, it is set to 2.0).

[0040] Based on frequency domain enhancement features, fog-free images are obtained using convolutional layers and activation functions. Specifically, the final result is achieved through... Convolutional layers and the Tanh activation function directly output fog-free images.

[0041] A further implementation involves calculating the difference between the generated fog-free image and the real fog-free image using a hybrid loss function, and then backpropagating to update the network parameters of the defogging model; wherein the hybrid loss function includes pixel domain loss, structural domain loss and perceptual domain loss.

[0042] Hybrid loss function The expression is as follows: in, , , The weighting coefficients used to balance the contributions of each loss term are positive real numbers and are not required to sum to 1. In this embodiment, we take... , , Domain loss Defined as: Pixel domain loss Defined as: In the formula, These represent the mean and variance of the image patch, respectively. For covariance, It is a constant. These are generated images and real images, respectively.

[0043] The definition of is: in, This represents the feature map obtained in the l-th layer after the input image is fed into the pre-trained feature extraction network. and These represent the height and width of the feature map, respectively. This represents the set of layers used to compute perceived differences.

[0044] This embodiment does not introduce Lab color loss or physical consistency loss based on explicit imaging model residuals. The physical guidance is achieved through fog-sensing modulation and frequency domain separation enhancement within the network.

[0045] S3: Based on the defogging model, the foggy agricultural remote sensing image to be processed is defogging to obtain a defogging image.

[0046] Example 2 This invention also provides an agricultural remote sensing image dehazing system based on multi-scale hole map networks and frequency domain physical sensing, and a method for implementing this system, comprising: The data acquisition module is used to acquire paired agricultural remote sensing image datasets with and without fog within a preset area, constructing a maskless end-to-end remote sensing image dataset; the experimental hardware environment includes: GPU: NVIDIA RTX 3080 Ti (12GB) 1; CPU: Intel(R) Xeon(R) CPU E5-2686 v4 @ 2.30GHz; Memory: 28GB; Hard drive: 30 GB system disk, 60 GB data disk.

[0047] The software environment for the experiment includes: Operating system: Ubuntu 5.19.0-50-generic; Programming language: Python 3.10; Deep learning framework: PyTorch 2.1.2; CUDA version: 12.1.

[0048] The dehazing model building module is used to construct a dehazing model based on a remote sensing image dataset, using a U-Net encoder-decoder structure that incorporates a multi-scale hole perception module, a dynamic graph neural network module, and a physically guided feature extractor. Training parameter configuration: Optimizer: AdamW optimizer is used; Momentum parameter: .9, ; Learning rate strategy: Initial learning rate set to The learning rate is dynamically adjusted using a cosine annealing strategy, with a minimum learning rate set to... .

[0049] Training cycle: The total number of training epochs is set to 300.

[0050] Batch Size: Set to 4. Loss Weights: and The weights are all set to 5.0 to balance pixel precision and perceived quality.

[0051] The effect diagram obtained by verifying the invention through inputting test set data is shown below. Figure 5 As shown. The first column is the foggy image, the second column is the label, and the third column is the fog-free image generated by this invention.

[0052] The dehazing image generation module is used to dehaze foggy agricultural remote sensing images based on a dehazing model to obtain dehazing images.

[0053] Example 3 Based on the same inventive concept, corresponding to any of the above embodiments, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the agricultural remote sensing image dehazing method based on multi-scale hole map network and frequency domain physical sensing as described in any of the above embodiments.

[0054] Figure 6 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0055] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0056] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0057] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0058] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB (Universal Serial Bus), network cable, etc.) or wireless means (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).

[0059] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0060] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0061] The system described in the above embodiments is used to implement the corresponding method in any of the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0062] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for dehazing agricultural remote sensing images based on multi-scale hole map networks and frequency domain physical sensing, characterized in that, include: Acquire paired agricultural remote sensing image datasets with and without fog within a preset area, and construct a maskless end-to-end remote sensing image dataset. Based on the remote sensing image dataset, a dehazing model is constructed using a U-Net encoder-decoder structure that incorporates a multi-scale hole perception module, a dynamic graph neural network module, and a physically guided feature extractor. Based on the aforementioned defogging model, the foggy agricultural remote sensing image to be processed is defogging to obtain a defogging image.

2. The method according to claim 1, characterized in that, The defogging model includes: A multi-scale hole sensing module, located at the front end of the encoder, includes three convolutional branches for extracting multi-scale features of remote sensing images. The encoder is used to downsample the multi-scale features to obtain a deep semantic feature map; The first dynamic graph neural network module, embedded in the bottleneck layer of the U-Net architecture, is used to transform the deep semantic feature map into a topological graph structure to obtain a fused feature map. A decoder is used to upsample the fused feature map and obtain a decoded feature map using a second dynamic graph neural network module embedded in the intermediate layer; A physical-guided feature extractor, located at the end of the decoder, is used to regress implicit physical parameters based on the decoded feature map. The implicit physical parameters include at least a transmittance map, an atmospheric light map, and an implicit fog layer interest map. Based on the implicit fog layer interest map, a gating weight is generated to adjust the feature fusion intensity of the dynamic graph neural network module. At the same time, affine feature modulation is performed in the spatial domain and high-low frequency separation processing is performed in the frequency domain to generate a fog-free image.

3. The method according to claim 2, characterized in that, Both the first dynamic graph neural network module and the second dynamic graph neural network module include: Dynamic layer construction is used to divide the input feature map into non-overlapping blocks, map them into node vectors, calculate the cosine similarity between nodes, select the nodes with the highest cosine similarity as the neighbor set, and obtain the graph structure; wherein, each node vector corresponds to a block region in the input feature map, and the node vectors updated by graph attention are backfilled and rearranged according to the original block position to form a graph network output feature map of the same size as the input feature map. The graph attention aggregation layer is used to update node features using attention coefficients based on the graph structure to obtain graph network branch features; An adaptive gated fusion layer is used to perform element-wise gated weighted fusion of the initial convolutional features input to this module and the graph network branch features after aligning them in the same spatial position to obtain fused features. The initial convolutional features participate in the fusion as identity mapping branches that preserve local texture. The gate weights are obtained by mapping the interest graph of the implicit fog layer and are used to limit the range of graph network information injection and reduce artifacts in clear areas.

4. The method according to claim 2, characterized in that, Methods for generating fog-free images include: Spatial domain fog-aware modulation is applied to the decoded feature map to obtain the modulation features; The frequency components of the decoded feature map are separated using a low-pass filter operator; By fusing the modulation features and the frequency components, a frequency domain enhancement feature is obtained; wherein the frequency components include low-frequency components and high-frequency components, and the high-frequency gain coefficient is obtained by mapping the transmittance map or the implicit fog layer interest map and clipped to a specified range; Based on the frequency domain enhancement features, the fog-free image is obtained using convolutional layers and activation functions.

5. The method according to claim 4, characterized in that, The method for obtaining the modulation features includes: Convolutional layers are used to regress and predict potential transmittance maps, atmospheric light maps, and implicit fog interest maps from decoded feature maps; The implicit fog concern graph is normalized and cropped to 0 to 1 to obtain the gating weight, which is used to characterize the feature injection intensity of the fog-occluded region. Based on the predicted transmittance map and atmospheric light map, affine coefficients are generated; wherein the affine coefficients include scaling coefficients and translation coefficients, and the scaling coefficients and translation coefficients are aligned with the decoded feature map in terms of spatial size and number of channels; Based on the gating weights and the affine coefficients, spatial domain fog-aware modulation is performed on the decoded feature map; wherein the decoded feature map is normalized and then scaled and translated to obtain intermediate modulation features, and the intermediate modulation features and the original decoded feature map are fused element-wise by the gating weights to obtain the modulation features.

6. The method according to claim 4, characterized in that, The difference between the generated fog-free image and the real fog-free image is calculated by a hybrid loss function, and the network parameters of the defogging model are updated by backpropagation; wherein the hybrid loss function includes pixel domain loss, structural domain loss and perceptual domain loss.

7. A dehazing system for agricultural remote sensing images based on multi-scale hole map networks and frequency domain physical sensing, used to implement the method described in any one of claims 1-6, characterized in that, include: The data acquisition module is used to acquire paired agricultural remote sensing image datasets with and without fog within a preset area, and to construct a maskless end-to-end remote sensing image dataset. The dehazing model construction module is used to construct a dehazing model based on the remote sensing image dataset, using a U-Net encoder-decoder structure that incorporates a multi-scale hole perception module, a dynamic graph neural network module, and a physically guided feature extractor. The dehazing image generation module is used to dehaze the foggy agricultural remote sensing image to be processed based on the dehazing model, and obtain the dehazing image.

8. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the computer program is executed by the processor according to any one of claims 1-6, the method for dehazing agricultural remote sensing images based on multi-scale hole map networks and frequency domain physical sensing.