A remote sensing fusion method based on frequency domain reconstruction and cross-modal channel exchange

By employing frequency domain reconstruction and cross-modal channel exchange, the contradiction between high precision and low computing power in multimodal remote sensing data fusion is resolved, achieving efficient and real-time multimodal remote sensing data fusion, improving feature extraction and classification accuracy, and making it suitable for national land surveys and urban governance.

CN121214149BActive Publication Date: 2026-03-27湖南工商大学
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies face challenges in processing multimodal remote sensing data, including the contradiction between high accuracy and low computing power, the inefficiency of multi-stage processes, the difficulty of multimodal data fusion, and the mismatch between real-time requirements and existing technologies. These limitations restrict the practical application of remote sensing technology.

Method used

A method based on frequency domain reconstruction and cross-modal channel exchange is adopted, including a two-stage preprocessing of adaptive noise reduction and calibration, format and dimension conversion, a three-stage iterative cyclic feature extraction, and a fusion process of an adaptive cross-modal hierarchical adversarial fusion network, to achieve efficient fusion of hyperspectral images and radar images.

Benefits of technology

It achieves high-precision, real-time, and highly adaptable multimodal remote sensing data fusion, improves feature extraction and classification accuracy, meets low computing power requirements, and is suitable for high-precision data processing in land surveys and urban governance scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121214149B_ABST
    Figure CN121214149B_ABST
Patent Text Reader

Abstract

The application relates to a remote sensing fusion method based on frequency domain reconstruction and cross-modal channel exchange, comprising the following steps: performing format and dimension conversion on a hyperspectral image and a radar image after two-stage preprocessing; performing three-stage iterative cycle feature extraction on the converted hyperspectral image tensor and the radar image tensor to obtain hyperspectral features and radar features; each stage of feature extraction comprises the following steps: respectively performing adaptive frequency domain transformation on the hyperspectral features and the radar features, calculating the correlation between the hyperspectral image tensor and the radar image tensor, and respectively performing feature modulation on the transformed hyperspectral features and radar features based on the correlation; the modulated hyperspectral features and radar features are subjected to a residual-Transformer model to output the corresponding hyperspectral features and radar features of each stage; and the hyperspectral features and radar features output in the third stage are fused through an adaptive cross-modal hierarchical adversarial fusion network to obtain final fusion features.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of remote sensing image processing, in particular to a remote sensing fusion method based on frequency domain reconstruction and cross-modal channel exchange. BACKGROUND

[0002] In the field of remote sensing image processing today, with the wide application of multi-source data such as hyperspectral images (HSI), synthetic aperture radar (SAR) images and laser radar (LiDAR) images, the existing technology has the problems of contradiction between high precision and low computing power, inefficiency of multi-stage process, difficulty of multi-modal data fusion, and mismatch between real-time requirements and existing technology when processing multi-modal remote sensing data. These problems seriously limit the promotion and use of remote sensing technology in practical applications. Therefore, there is an urgent need for an end-to-end multi-modal fusion method that can achieve high precision and real-time processing under low computing power conditions to meet the fine and timely requirements in various scenarios. SUMMARY

[0003] Therefore, it is necessary to provide a remote sensing fusion method based on frequency domain reconstruction and cross-modal channel exchange, comprising:

[0004] S1: obtaining hyperspectral images and radar images, and performing two-stage preprocessing of adaptive noise reduction and calibration on each modal image, and performing format and dimension conversion on the hyperspectral images and radar images after two-stage preprocessing;

[0005] S2: performing three-stage iterative cycle feature extraction on the converted hyperspectral image tensor and radar image tensor to obtain hyperspectral features and radar features; each stage of feature extraction includes: performing adaptive frequency domain transformation on the hyperspectral features and radar features respectively, calculating the correlation between the hyperspectral image tensor and the radar image tensor, and based on the correlation, modulating the transformed hyperspectral features and radar features respectively, and outputting the corresponding hyperspectral features and radar features of each stage through the residual-Transformer model; the first stage uses the converted hyperspectral image tensor and radar image tensor, and the second and third stages use the output of the previous stage;

[0006] S3: fusing the hyperspectral features and radar features output in the third stage through an adaptive cross-modal hierarchical adversarial fusion network to obtain the final fusion features.

[0007] Beneficial effects: the method provides an efficient, accurate, real-time and adaptive solution for remote sensing fusion in the field of remote sensing image processing. BRIEF DESCRIPTION OF DRAWINGS

[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0009] Figure 1 The flow chart of the remote sensing fusion method based on frequency domain reconstruction and cross-modal channel exchange in the embodiments of the present application. DETAILED DESCRIPTION

[0010] In order to make the above-mentioned purposes, features and advantages of the present application more apparent and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings. In the following description, a large number of specific details are set forth in order to provide a sufficient understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the spirit of the present application, so the present application is not limited to the specific embodiments disclosed below.

[0011] In addition, the terms "first", "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise explicitly specified.

[0012] As shown in Figure 1 The present embodiment provides a remote sensing fusion method based on frequency domain reconstruction and cross-modal channel exchange, comprising:

[0013] S1: Obtain hyperspectral images and radar images, and perform two-stage preprocessing of adaptive noise reduction and calibration on each modal image, and perform format and dimension conversion on the hyperspectral images and radar images after two-stage preprocessing.

[0014] In the present embodiment, the radar images include synthetic aperture radar images and laser radar images.

[0015] Specifically, the two-stage preprocessing of the hyperspectral images comprises:

[0016] The hyperspectral images are decomposed into low-frequency structure components and high-frequency noise components by two-dimensional wavelet transform;

[0017] The information entropy of each band in the hyperspectral images is calculated, and each information entropy is normalized to obtain the corresponding band weight;

[0018] The high-frequency noise components are weighted based on the wave band weights corresponding to the high-frequency noise components, the low-frequency structure components and the weighted high-frequency noise components are added, and then the added result is inversely transformed through a wavelet to reconstruct a denoised hyperspectral image;

[0019] A ground truth sample is obtained, the ground truth sample is a hyperspectral image used for calibration, and a first average spectrum of any one ground object category in each wave band of the ground truth sample is calculated;

[0020] A second average spectrum of the corresponding ground object category in each wave band of the hyperspectral image is calculated;

[0021] The denoised hyperspectral image is calibrated based on the first average spectrum and the second average spectrum to obtain a calibrated hyperspectral image.

[0022] The two-stage preprocessing of the radar image includes:

[0023] For a synthetic aperture radar image, the filter window size is dynamically adjusted according to the scattering coefficients of each polarization channel in the synthetic aperture radar image, the filter window is used to filter the synthetic aperture radar image of the corresponding polarization channel, and the calculation formula is:

[0024] ;

[0025] wherein, represents the synthetic aperture radar image of the pth polarization channel after filtering, represents the variance of the scattering coefficient of the pth polarization channel, represents the noise variance obtained by dark area estimation on the synthetic aperture radar image, represents the synthetic aperture radar image of the pth polarization channel, represents the mean value of the synthetic aperture radar image of the pth polarization channel;

[0026] After filtering each polarization channel of the synthetic aperture radar image, a filtered synthetic aperture radar image is obtained;

[0027] For a laser radar image, the mean value of pixels in the laser radar image is calculated and the standard deviation Pixels in the laser radar image that are beyond the range of are determined as abnormal points, otherwise they are determined as normal points; the abnormal points are inversely distance weighted interpolated and completed based on the normal points in the neighborhood of the abnormal points, and the calculation formula is:

[0028] ;

[0029] wherein, represents the jth abnormal point in the laser radar image after completion, represents the coordinates of the abnormal point, K represents the number of normal points in the neighborhood of the abnormal point, represents the Euclidean distance between the abnormal point and the kth normal point in its neighborhood, represents the normal points in the neighborhood of the abnormal point in the laser radar image, represents the kth normal point coordinates in the neighborhood of the abnormal point;

[0030] After each abnormal point in the laser radar image is completed, a completed laser radar image is obtained;

[0031] Based on the calibrated hyperspectral image, the normalized cross-correlation value of each image block in the filtered synthetic aperture radar image or the completed laser radar image and the corresponding image block in the calibrated hyperspectral image is calculated, the offset value when the normalized cross-correlation value is maximum is selected, and each image block in the filtered synthetic aperture radar image or the completed laser radar image is calibrated based on the offset value, to obtain a calibrated synthetic aperture radar image or a laser radar image. The calculation formula is:

[0032] ;

[0033] wherein, represents the calibrated image block in the filtered synthetic aperture radar image or the completed laser radar image, represents the coordinates of the calibrated image block, represents the filtered synthetic aperture radar image or the completed laser radar image, represents the coordinates of the image block before calibration, represents the offset value when the normalized cross-correlation value is maximum.

[0034] Further, the format and dimension conversion includes:

[0035] The two-stage preprocessed hyperspectral image and radar image are respectively converted into tensor format;

[0036] The hyperspectral image and radar image in tensor format are cropped or interpolated according to the preset spatial dimension requirement, and the spatial dimensions of the hyperspectral image and radar image in tensor format are unified;

[0037] The hyperspectral image and radar image in tensor format with unified spatial dimensions are respectively normalized to obtain a hyperspectral image tensor and a radar image tensor.

[0038] S2: Perform a three-stage iterative feature extraction on the transformed hyperspectral image tensor and radar image tensor to obtain hyperspectral features and radar features. Each stage of feature extraction includes: performing adaptive frequency domain transformation on the hyperspectral features and radar features respectively, calculating the correlation between the hyperspectral image tensor and the radar image tensor, and performing feature modulation on the transformed hyperspectral features and radar features based on the correlation. The modulated hyperspectral features and radar features are then passed through a residual-transformer model to output the corresponding hyperspectral features and radar features for each stage. The first stage uses the transformed hyperspectral image tensor and radar image tensor, while the second and third stages use the output of the previous stage.

[0039] Specifically, feature extraction at any stage includes:

[0040] Step 1: Calculate the local variance of the hyperspectral image tensor or the hyperspectral features output from the previous stage using a 3×3 sliding window. The calculation formula is as follows:

[0041] ;

[0042] in, Represents local variance. Indicates the position of elements in hyperspectral or radar features. Indicates the position of the sliding window. This represents the hyperspectral image tensor or the hyperspectral features output from the previous stage. Indicates the inside of the sliding window The mean;

[0043] Adjusting the number of FFT subbands based on local variance and variance threshold is expressed as follows:

[0044] ;

[0045] in, Indicates in Number of FFT subbands at a given location Indicates the variance threshold;

[0046] Based on the adjusted number of FFT sub-band divisions, perform a regional FFT on the hyperspectral image tensor or the hyperspectral features output from the previous stage, decomposing them into the first real part and the first imaginary part.

[0047] Step 2: Evaluate the noise level of each region in the radar image tensor or the radar features output from the previous stage based on the local mean and standard deviation (calculated using the signal-to-noise ratio formula). For regions with noise levels greater than or equal to the noise threshold, first perform Gaussian smoothing, and then perform regional FFT on the Gaussian smoothed regions according to the fixed number of sub-bands (6 sub-bands), decomposing them into the second real part and the second imaginary part.

[0048] Step 3: Calculate the correlation between the hyperspectral image tensor or the hyperspectral feature output by the last stage and the radar image tensor or the radar feature output by the last stage, and the calculation formula is:

[0049] ;

[0050] wherein, represents the correlation between the hyperspectral image tensor or the hyperspectral feature output by the last stage and the radar image tensor or the radar feature output by the last stage at the position, represents the element at the position in ; represents the radar image tensor or the radar feature output by the last stage, represents the element at the position in ; represents a regularization coefficient, represents a module;

[0051] The weights of the first real part, the first imaginary part, the second real part and the second imaginary part are calculated based on the correlation, and the calculation formula includes:

[0052] ;

[0053] ;

[0054] ;

[0055] ;

[0056] wherein, represents the weight of the first real part, represents the weight of the first imaginary part, represents the weight of the second real part, represents the weight of the second imaginary part, represents a sigmoid function, represents a global average pooling, represents a global maximum pooling, represents a 1x1 convolution, represents a splicing operation, represents a maximum value, represents the first real part, represents the first imaginary part, represents the second real part, represents the second imaginary part;

[0057] The weight of the first real part, the first imaginary part, the second real part, and the second imaginary part is scaled based on a training round factor. The weight scaling calculation formula of the first real part is:

[0058] ;

[0059] ;

[0060] ;

[0061] wherein, represents the weight of the scaled first real part, represents a scaling factor, represents a training round factor;

[0062] Step 4: The corresponding real part or imaginary part is modulated based on the scaled weight (the scaled weight is multiplied with the corresponding real part or imaginary part element by element), and IFFT is performed on the modulated first real part and first imaginary part to obtain hyperspectral frequency domain enhanced features, and IFFT is performed on the modulated second real part and second imaginary part to obtain radar image frequency domain enhanced features;

[0063] Step 5: The hyperspectral frequency domain enhanced features and the radar image frequency domain enhanced features are respectively input into a residual-Transformer model to output the corresponding hyperspectral features and radar features of the current stage.

[0064] Further, in the second stage and the third stage, the variance threshold is gradually reduced (0.1 in the first stage, 0.08 in the second stage), the correlation degree calculation formula is updated based on the scaled weight of the previous stage, and the model parameters in the residual-Transformer model are modified.

[0065] In this embodiment, the residual-Transformer model includes a multi-layer perception and a multi-head Transformer. The hyperspectral frequency domain enhanced features and the radar image frequency domain enhanced features are respectively input into two layers of multi-layer perceptions to obtain corresponding first features and second features. The first features or the second features are preliminarily fused through the multi-head Transformer, and the preliminary fusion results are residual fused with the corresponding hyperspectral frequency domain enhanced features and radar image frequency domain enhanced features through 1x1 convolution to output the corresponding hyperspectral features and radar features of the current stage.

[0066] S3: The hyperspectral features and the radar features output in the third stage are fused through the adaptive cross-modal hierarchical adversarial fusion network to obtain the final fusion features.

[0067] Specifically, the fusion process includes:

[0068] Step 1: Calculate the first attention map of hyperspectral features to radar features in the third stage output, the calculation formula is:

[0069] ;

[0070] wherein, represents the first attention map of hyperspectral features to radar features, represents the element position in the hyperspectral features or radar features, represents the element in the hyperspectral features in the third stage output, represents the element in the radar features in the third stage output, represents the transpose, represents the regularization coefficient;

[0071] The second attention map of radar features to hyperspectral features in the third stage output, the calculation formula is:

[0072] ;

[0073] wherein, represents the second attention map of radar features to hyperspectral features;

[0074] The radar features in the third stage output are weighted based on the first attention map, and the hyperspectral features in the third stage output are weighted based on the second attention map;

[0075] Step 2: Divide the weighted hyperspectral features and radar features respectively by using sliding windows of different sizes, and respectively obtain first local feature blocks and second local feature blocks; Calculate the correlation between the first local feature blocks and the second local feature blocks, the calculation formula is:

[0076] ;

[0077] wherein, represents the correlation between the ith first local feature block and the jth second local feature block, represents the variance, represents the covariance, represents the ith first local feature block, represents the jth second local feature block;

[0078] Step 3: Weight each correlation based on the corresponding channel importance weight to obtain a weighted correlation, the calculation formula is:

[0079] ;

[0080] wherein, ​channel importance weight of the c-th local feature block representing the feature of the m-th modality, m is 1 or 2, equal to 1 when the modality is hyperspectral image, equal to 2 when the modality is radar image, number representing the correlation degree, correlation degree between the c-th first local feature block and the k-th second local feature block, weighted correlation degree of the c-th local feature block representing the hyperspectral feature or the radar feature;

[0081] Step 4: Weighted sum the first local feature block and the second local feature block based on the weighted correlation degree, to obtain the first local feature block or the second local feature block after channel exchange, the calculation formula is:

[0082] ;

[0083] wherein, the local feature block after channel exchange in the feature of the m-th modality, the local feature block in the feature of the m-th modality, the local feature block in the feature of the other modality, sigmoid function, learning parameter,

[0084] Step 5: Fuse the first local feature block or the second local feature block after channel exchange corresponding to each sliding window, and restore to the original spatial dimension, to obtain the local exchanged hyperspectral feature or radar feature, the calculation formula is:

[0085] ;

[0086] wherein, the radar feature of the m-th modality after local exchange, fusion operation, the local feature block after channel exchange in the feature of the m-th modality in the s-th sliding window;

[0087] Step 6: Remodel the local exchanged hyperspectral feature and radar feature into sequence format, and input into the lightweight state space model, to obtain the hyperspectral remodeling feature and the radar remodeling feature respectively;

[0088] Step 7: Generate the conditional guided feature based on the hyperspectral remodeling feature and the radar remodeling feature, the calculation formula is:

[0089] ;

[0090] wherein, C represents the conditional guided feature, global average pooling, ReLU activation function, denotes layer normalization, denotes convolution operation, denotes concatenation operation, denotes hyperspectral remodeling feature, radar remodeling feature;

[0091] Step 8: up-sampling the conditional guided feature based on bilinear interpolation up-sampling to restore to the original spatial dimension to obtain the final fusion feature.

[0092] The embodiment also includes: calculating the mean square error loss and the structural similarity loss based on the expected feature and the final fusion feature, and weighting and summing to obtain the total loss, and optimizing the feature extraction process and the fusion process by minimizing the total loss.

[0093] The remote sensing fusion method based on frequency domain reconstruction and cross-modal channel exchange provided by the embodiment has the following beneficial effects:

[0094] The fusion method realizes efficient fusion of multi-source remote sensing data such as hyperspectral (HSI), synthetic aperture radar (SAR), or laser radar (LiDAR) through a series of innovative technical means. The method realizes high-precision feature extraction and classification accuracy improvement, meets low-power demand, high efficiency, real-time performance, and strong adaptability, and the final fusion feature has stronger expression ability and classification performance.

[0095] The fusion method can be applied to land survey and resource investigation scenarios. Through the fusion of HSI, SAR, LiDAR and other multi-source remote sensing data, the downstream model can more accurately identify the land use type of the forest area, provide high-precision classification results for land survey, and more accurately assess the reserves and distribution of mineral resources, providing strong technical support for resource investigation. It can also be applied to urban management and planning scenarios. Through the fusion of these multi-modal data, the downstream model can quickly and accurately identify newly built illegal buildings, provide timely information support for urban management, and also generate high-precision urban topographic maps and land use maps to help planning personnel understand the spatial structure and land use status of the city, thereby making more reasonable urban planning.

[0096] The technical features of the above-described embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described, but as long as the combinations of the technical features do not exist, they should be considered as the scope of the present disclosure.

[0097] The above embodiments only express several implementation ways of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation to the patent scope of the application. It should be pointed out that for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, which all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A remote sensing fusion method based on frequency domain reconstruction and cross-modal channel exchange, characterized in that, Comprise: S1: obtain hyperspectral images, radar images, and perform adaptive noise reduction and calibration on the images of each modality for two-stage preprocessing, and perform format and dimension conversion on the hyperspectral images and radar images after two-stage preprocessing; S2: perform three-stage iterative cycle feature extraction on the converted hyperspectral image tensor and radar image tensor to obtain hyperspectral features and radar features; each stage of feature extraction includes: performing adaptive frequency domain transformation on the hyperspectral features and radar features, respectively, calculating the correlation between the hyperspectral image tensor and the radar image tensor, and based on the correlation, modulating the transformed hyperspectral features and radar features, respectively, passing the modulated hyperspectral features and radar features through a residual-Transformer model, and outputting the corresponding hyperspectral features and radar features of each stage; the first stage uses the converted hyperspectral image tensor and radar image tensor, and the second and third stages use the output of the previous stage; S3: fuse the hyperspectral features and radar features output by the third stage through an adaptive cross-modality hierarchical adversarial fusion network to obtain the final fused features; the fusion process includes: Step 1: calculate the first attention map from the hyperspectral features output by the third stage to the radar features, the calculation formula is: ; wherein, represents a first attention map from hyperspectral features to radar features, represents an element position in the hyperspectral features or radar features, represents an element in the hyperspectral features output by the third stage, represents an element in the radar features output by the third stage, represents an element in the hyperspectral features output by the third stage, represents an element in the radar features output by the third stage, represents a regularization coefficient; The second attention map from the radar features output by the third stage to the hyperspectral features, the calculation formula is: ; wherein, represents a second attention map of radar features to hyperspectral features; Weight the radar features output by the third stage based on the first attention map, and weight the hyperspectral features output by the third stage based on the second attention map; Step 2: divide the weighted hyperspectral features and radar features using sliding windows of different sizes to obtain first local feature blocks and second local feature blocks, respectively; calculate the correlation between the first local feature blocks and the second local feature blocks, the calculation formula is: ; wherein, denotes a correlation between the i-th first local feature block and the j-th second local feature block, denotes a variance, denotes a covariance, denotes the i-th first local feature block, denotes the j-th second local feature block; Step 3: weight each correlation based on the corresponding channel importance weight to obtain the weighted correlation, the calculation formula is: ; wherein, a channel importance weight of a c-th local feature block representing a characteristic of a m-th modality, m is equal to 1 or 2, equal to 1 when the modality is hyperspectral image, equal to 2 when the modality is radar image, a number representing a correlation degree, a correlation degree between a c-th first local feature block and a k-th second local feature block, a weighted correlation degree of a c-th local feature block representing a hyperspectral characteristic or a radar characteristic; Step 4: weight and sum the first local feature blocks and the second local feature blocks based on the weighted correlation to obtain the first local feature blocks or the second local feature blocks after channel exchange, the calculation formula is: ; wherein, denotes a local feature block in the channel exchanged feature of the m-th modality, denotes a local feature block in the feature of the m-th modality, denotes a local feature block in the feature of another modality, denotes a sigmoid function, denotes a learning parameter; Step 5: fuse the first local feature blocks or the second local feature blocks corresponding to each sliding window after channel exchange and restore them to the original spatial dimension to obtain the hyperspectral features or radar features after local exchange, the calculation formula is: ; wherein, represents a local feature block of the m-th modality after channel swapping, represents a fusion operation, represents a local feature block of the m-th modality in the s-th sliding window after channel swapping; Step 6: reshape the hyperspectral features and radar features after local exchange into a sequence format and input them into a lightweight state space model to obtain hyperspectral reshaped features and radar reshaped features, respectively; Step 7: generate conditional guided features based on the hyperspectral reshaped features and radar reshaped features, the calculation formula is: ; wherein C denotes a conditional guidance feature, denotes a global average pooling, denotes a ReLU activation function, denotes a layer normalization, denotes a convolution operation, denotes a concatenation operation, denotes a hyperspectral reshaping feature, a radar reshaping feature; Step 8: upsample the conditional guided features based on bilinear interpolation to restore them to the original spatial dimension to obtain the final fused features.

2. The remote sensing fusion method based on frequency domain reconstruction and cross-modality channel exchange according to claim 1, characterized in that, The radar images include synthetic aperture radar images and laser radar images.

3. The method of claim 2, wherein, The two-stage preprocessing of the hyperspectral images includes: Decompose the hyperspectral images into low-frequency structure components and high-frequency noise components through two-dimensional wavelet transform; Calculate the information entropy of each band in the hyperspectral image, and normalize each information entropy to obtain the corresponding band weight; The high-frequency noise components are weighted based on the wave band weights corresponding to the high-frequency noise components, the low-frequency structure components and the weighted high-frequency noise components are added, and then the added high-frequency noise components are reconstructed through inverse wavelet transform to obtain the denoised hyperspectral image; A ground truth sample is obtained, the ground truth sample is a hyperspectral image used for calibration, and a first average spectrum of any one ground object category in each wave band of the ground truth sample is calculated; A second average spectrum of the corresponding ground object category in each wave band of the hyperspectral image is calculated. The denoised hyperspectral image is calibrated based on the first average spectrum and the second average spectrum to obtain a calibrated hyperspectral image.

4. The remote sensing fusion method based on frequency domain reconstruction and cross-modal channel exchange according to claim 3, characterized in that, The two-stage preprocessing of the radar image includes: For a synthetic aperture radar image, the filter window size is dynamically adjusted according to the scattering coefficients of each polarization channel in the synthetic aperture radar image, and the filter window is used to filter the synthetic aperture radar image of the corresponding polarization channel, and the calculation formula is: ; wherein, denotes the synthetic aperture radar image of the p-th polarization channel after filtering, denotes the variance of the scattering coefficient of the p-th polarization channel, denotes the noise variance resulting from the dark region estimation of the synthetic aperture radar image, denotes the synthetic aperture radar image of the p-th polarization channel, denotes the mean value of the synthetic aperture radar image of the p-th polarization channel; After filtering each polarization channel of the synthetic aperture radar image, a filtered synthetic aperture radar image is obtained. For the laser radar image, the mean value of pixels in the laser radar image is calculated and the standard deviation Pixels in the laser radar image beyond the range of are determined as abnormal points, and otherwise as normal points; the abnormal points are interpolated and completed by inverse distance weighted interpolation based on the normal points in the neighborhood of the abnormal points, and the calculation formula is: ; wherein, represents the jthoutlier point in the completed laser radar image, represents the outlier point coordinate, and K represents the number of normal points in the neighborhood of the outlier point, represents the Euclidean distance between the outlier point and the kthnormal point in its neighborhood, represents the normal points in the neighborhood of the outlier point in the laser radar image, represents the kthnormal point coordinate in the neighborhood of the outlier point; After each abnormal point in the laser radar image is completed, a completed laser radar image is obtained. Based on the calibrated hyperspectral image, the normalized cross-correlation value of each image block in the filtered synthetic aperture radar image or the completed laser radar image and the corresponding image block in the calibrated hyperspectral image is calculated, the offset value when the normalized cross-correlation value is maximum is selected, and the coordinates of each image block in the filtered synthetic aperture radar image or the completed laser radar image are calibrated based on the offset value to obtain a calibrated synthetic aperture radar image or laser radar image, and the calculation formula is: ; wherein, denotes a calibrated image patch of the filtered synthetic aperture radar image or the completed lidar image, denotes a calibrated image patch coordinate, denotes a filtered synthetic aperture radar image or a completed lidar image, denotes a pre-calibrated image patch coordinate, denotes a shift at which the normalized cross-correlation value is maximum.

5. The method of claim 1, wherein, The format and dimension conversion includes: The hyperspectral image and the radar image after the two-stage preprocessing are converted into a tensor format respectively; The hyperspectral image and the radar image in the tensor format are cropped or interpolated according to a preset spatial dimension requirement, and the spatial dimensions of the hyperspectral image and the radar image in the tensor format are unified; The hyperspectral image and the radar image in the tensor format with unified spatial dimensions are normalized respectively to obtain a hyperspectral image tensor and a radar image tensor respectively.

6. The method of claim 1, wherein, In S2, any one stage feature extraction includes: Step 1: The local variance of the hyperspectral image tensor or the hyperspectral feature output in the last stage is calculated by using a 3*3 sliding window, the FFT sub-band division number is adjusted according to the local variance and the variance threshold, and the hyperspectral image tensor or the hyperspectral feature output in the last stage is executed by region-based FFT according to the adjusted FFT sub-band division number to be decomposed into a first real part and a first imaginary part; Step 2: The noise level of each region in the radar image tensor or the radar feature output in the last stage is evaluated based on the local mean and the standard deviation, the region with a noise level greater than or equal to a noise threshold is first executed by Gaussian smoothing, and then the region after Gaussian smoothing is executed by region-based FFT according to a fixed sub-band division number to be decomposed into a second real part and a second imaginary part; Step 3: Calculate the correlation between the hyperspectral image tensor or the hyperspectral feature output by the previous stage and the radar image tensor or the radar feature output by the previous stage; based on the correlation, calculate the weights of the first real part, the first imaginary part, the second real part and the second imaginary part respectively, and scale the weights of the first real part, the first imaginary part, the second real part and the second imaginary part based on the training round factor respectively; Step 4: Based on the scaled weights, modulate the corresponding real part or imaginary part, perform IFFT on the modulated first real part and first imaginary part to obtain the hyperspectral frequency domain enhanced feature, and perform IFFT on the modulated second real part and second imaginary part to obtain the radar image frequency domain enhanced feature; Step 5: The hyperspectral frequency domain enhanced feature and the radar image frequency domain enhanced feature are respectively input into the residual-Transformer model to output the corresponding hyperspectral feature and radar feature of the current stage.

7. The method of claim 6, wherein, In the second stage and the third stage, gradually reduce the variance threshold, update the correlation calculation formula based on the scaled weights of the previous stage, and modify the model parameters in the residual-Transformer model.

8. The method of claim 6, wherein, The residual-Transformer model includes a multi-layer perceptron and a multi-head Transformer. The hyperspectral frequency domain enhanced feature and the radar image frequency domain enhanced feature are respectively input into the multi-layer perceptron to obtain corresponding first features and second features. The first features or the second features are input into the multi-head Transformer for preliminary fusion, and the preliminary fusion results are residual fused with the corresponding hyperspectral frequency domain enhanced features and radar image frequency domain enhanced features through 1x1 convolution to output the corresponding hyperspectral features and radar features of the current stage.

9. The method of claim 1, wherein, It also includes: Calculate the mean square error loss and the structural similarity loss based on the expected feature and the final fused feature, and weight the sum to obtain the total loss. The feature extraction process and the fusion process are optimized by minimizing the total loss.

Citation Information

Patent Citations

  • Hyperspectral and laser radar joint classification method of category perception fusion network

    CN118859230A