Multi-source environment interpretation method and device based on linear frequency modulation analysis

The linear frequency modulation analysis method is used to perform feature decomposition and cross-modal fusion of hyperspectral images and lidar data, solving the domain offset problem in multi-source remote sensing data classification and achieving efficient multi-source environment interpretation in small sample scenarios. It is suitable for complex airborne remote sensing scenarios.

CN120708089APending Publication Date: 2025-09-26BEIJING INST OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510861759.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

There is a domain shift problem in the classification of multi-source remote sensing data, which leads to a decrease in the generalization performance of the model in practical applications. In particular, it is difficult to effectively integrate multimodal information and amplify scene differences in the case of small samples.

Method used

The linear frequency modulation analysis method is used to decompose hyperspectral images and lidar data into amplitude spectrum and phase spectrum. Weight mapping and modulation are performed through convolutional networks to generate multi-source joint features. Time-varying frequency domain modeling is performed through linear frequency modulation basis functions, and cross-modal interaction and feature fusion are combined with windowed Transformer to generate multi-scale shared features.

Benefits of technology

It effectively alleviates multi-source distribution differences, enhances feature robustness, suppresses noise, solves feature drift caused by seasonal changes and environmental changes, achieves efficient cross-regional object classification, reduces dependence on target domain labeled data, and maintains high generalization performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708089A_ABST
    Figure CN120708089A_ABST
Patent Text Reader

Abstract

The invention provides a multi-source environment interpretation method and device based on linear frequency modulation analysis, and belongs to the field of space-based remote sensing intelligent process.The method comprises the steps that a hyperspectral image and laser radar digital surface model data of a target area are obtained, local neighborhood windows with pixels as the centers are extracted respectively, and the local neighborhood windows are obtained; generating a spatial-spectral data cube and elevation data; performing fractional domain feature extraction on the spatial spectrum data cube and the elevation data in sequence, and obtaining reconstruction features based on linear frequency modulation basis function mapping and cross-modal fusion; and identifying and classifying the reconstructed features to obtain a target area ground feature classification result. According to the method, through cascade design of fractional domain transformation, frequency modulation modeling and cross-modal attention, multi-source distribution alignment, dynamic feature enhancement and efficient fusion and generalization are realized, and an interpretable and low-dependence solution is provided for multi-source cross-regional ground feature classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of air-based remote sensing intelligent processing, and in particular relates to a multi-source environment interpretation method, device, equipment and medium based on linear frequency modulation analysis. Background Art

[0002] The rapid development of multi-sensor remote sensing technology across large spatial scales has brought unprecedented Earth observation capabilities. This, coupled with my country's strategic needs for aerospace sensing in the new era, places higher demands on the processing and analysis of aerospace-based multi-sensor remote sensing data. The establishment of a multi-spectral, full-element Earth observation system can directly promote the fusion and application of hyperspectral, lidar, and other data. Natural resource management requires the establishment of an "air, land, and sea" monitoring network, while ecological and environmental protection also requires the development of coordinated aerospace remote sensing technologies to achieve accurate inversion of vegetation cover.

[0003] However, multi-source remote sensing data classification faces a core challenge: domain shift. This is the distribution discrepancy between the training data (source domain) and the test data (target domain). This discrepancy severely impacts the generalization performance of models in practical applications. This includes: 1) spectral shift: variations in imaging conditions lead to different spectra for the same object; 2) spatial texture variation: differences in sensor resolution lead to inconsistent features; 3) sensor variability: different platform band settings and radiometric characteristics; and 4) temporal dynamics: seasonal changes and environmental evolution cause feature drift. Directly applying existing models across regions can result in accuracy drops of over 20%. Multi-source data fusion exacerbates this complexity, necessitating the design of specialized fusion modules.

[0004] To address this issue, existing methods are mainly divided into 1) statistical feature transformations: such as Maximum Mean Difference (MMD) and Correlation Alignment (CORAL), which reduce domain differences by aligning distribution statistics; 2) geometric feature transformations: including subspace alignment and optimal transfer; and 3) deep domain adaptation: adversarial training, federated learning, and self-supervised learning. Despite progress, existing methods still face the following shortcomings: the complex relationships between modalities in multi-source data make it difficult to fully extract ground feature features. This also makes it difficult to fully exploit complementary information in feature-level or decision-level fusion, and scene differences increase with the number of modalities. Therefore, how to design a deep neural network method suitable for multi-source environment interpretation based on linear frequency modulation analysis, and even how to reduce modal and regional differences in small sample sizes, is a key issue that needs to be addressed by technicians in the field of airborne remote sensing intelligent processing. Summary of the Invention

[0005] In order to solve the above problems, the present invention provides a multi-source environment interpretation method, apparatus, device and medium based on linear frequency modulation analysis.

[0006] In order to achieve the above object, the present invention provides the following technical solutions: A multi-source environment interpretation method based on linear frequency modulation analysis, the method comprising: Obtain hyperspectral images and lidar digital surface model data of the target area, and extract local neighborhood windows centered on pixels to generate spatial spectrum data cubes and elevation data; Decomposing the spatial spectrum data cube and the elevation data into an amplitude spectrum and a phase spectrum respectively; performing weight mapping on the phase spectrum through a convolutional network, generating a modulation weight matrix and multiplying the matrix element-by-element with the amplitude spectrum to output a multi-source joint feature; Performing time-varying frequency domain modeling on the multi-source joint features through linear frequency modulation basis functions to generate multi-transformation domain enhanced features; The multi-transform domain enhanced features are divided into non-overlapping blocks, and the non-overlapping blocks are filtered through a window mask to obtain visible blocks; the window Transformer is used to perform cross-modal interaction on the visible blocks to generate multi-scale shared features; based on the multi-scale shared features, a query vector, a key vector and a value vector are generated, the query vector and the key vector are matched through a cross-modal attention mechanism, and the value vector is weightedly fused based on the matching results to output the reconstructed features; the reconstructed features are identified and classified to obtain the classification results of the target area objects.

[0007] Optionally, the multi-source joint feature includes a hyperspectral fusion feature and a lidar fusion feature, and the phase spectrum is weighted mapped by a convolutional network to generate a modulation weight matrix and multiply the matrix element-by-element with the amplitude spectrum. The output multi-source joint feature includes: Performing fractional Fourier transform on the empty spectrum data cube to decompose it into the first magnitude spectrum and the first phase spectrum; Performing fractional Fourier transform on the elevation data to decompose it into a second magnitude spectrum and a second phase spectrum; Input the first phase spectrum and the second phase spectrum into the two-dimensional convolution layer respectively, and output a single-channel feature map; Normalize the single-channel feature map into a first modulation weight matrix and a second modulation weight matrix through a Sigmoid function; The first modulation weight matrix is ​​multiplied element-by-element with the first amplitude spectrum to generate a hyperspectral fusion feature, and the second modulation weight matrix is ​​multiplied element-by-element with the second amplitude spectrum to generate a lidar fusion feature.

[0008] Optionally, performing time-varying frequency domain modeling on the multi-source joint features by using a linear frequency modulation basis function to generate multi-transformation domain enhanced features includes: The frequency modulation basis function is defined as: ; in, x is the multi-source joint feature of the input, represents the Chirp frequency modulation coefficient, W Prepresents the nonlinear branching parameter; Decomposing the two multi-source joint features into linear combinations of frequency modulation basis functions respectively to generate a first multi-transform domain component and a second multi-transform domain component corresponding to the hyperspectral fusion feature and the lidar fusion feature; The transform domain components are reconstructed into spatial domain features through fractional Fourier transform to obtain corresponding first multi-transform domain enhanced features and second multi-transform domain enhanced features.

[0009] Optionally, the multi-scale shared features include a first multi-scale shared feature corresponding to the first multi-transform domain enhancement feature and a second multi-scale shared feature corresponding to the second multi-transform domain enhancement feature; a query vector, a key vector, and a value vector are generated based on the multi-scale shared features, the query vector and the key vector are matched through a cross-modal attention mechanism, and the value vector is weightedly fused based on the matching result, and the output reconstructed features include: Extracting a query vector from the first multi-scale shared feature to obtain a hyperspectral query feature; generating a key vector and a value vector from the second multi-scale shared feature through frequency modulation projection; the key vector is the lidar frequency modulation key; Calculate the frequency domain similarity between the hyperspectral query feature and the lidar frequency modulation key to generate cross-modal attention weights; The weights are weightedly fused with the value vector and connected with the first multi-scale shared feature residual to obtain the reconstructed feature.

[0010] A multi-source environment interpretation device based on linear frequency modulation analysis, the device comprising: The acquisition module is used to obtain the hyperspectral image and lidar digital surface model data of the target area, extract the local neighborhood window centered on the pixel, and generate the spatial spectrum data cube and elevation data; A construction module is used to construct a multi-source fusion model and train the multi-source fusion model to obtain an environment interpretation model; the environment interpretation model includes a score domain feature extraction module, a frequency modulation analysis module, and a cross-modal attention network; The fractional domain feature extraction module is used to decompose the spatial spectrum data cube and elevation data into an amplitude spectrum and a phase spectrum; perform weight mapping on the phase spectrum through a convolutional network, generate a modulation weight matrix, and multiply it element-by-element with the amplitude spectrum to output a multi-source joint feature; The frequency modulation analysis module is used to perform time-varying frequency domain modeling on the multi-source joint features through linear frequency modulation basis functions to generate multi-transformation domain enhanced features; The cross-modal attention network is used to divide the multi-transform domain enhanced features into non-overlapping blocks, and filter the non-overlapping blocks through a window mask to obtain visible blocks; use the window Transformer to perform cross-modal interaction on the visible blocks to generate multi-scale shared features; generate a query vector, a key vector and a value vector based on the multi-scale shared features, match the query vector and the key vector through a cross-modal attention mechanism, and perform weighted fusion on the value vector based on the matching results to output reconstructed features; identify and classify the reconstructed features to obtain the classification results of the target area objects.

[0011] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the multi-source environment interpretation method based on linear frequency modulation analysis is implemented.

[0012] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the multi-source environment interpretation method based on linear frequency modulation analysis is implemented.

[0013] The multi-source environment interpretation method based on linear frequency modulation analysis provided by the present invention has the following beneficial effects: First, the spatial spectrum and elevation data are decomposed into amplitude and phase spectra. The amplitude spectrum is dynamically adjusted through weight mapping of the phase spectrum, effectively aligning differences in multi-source distributions and enhancing feature robustness. This mitigates the problem of different spectra for the same object due to imaging conditions or sensor variations. Secondly, linear frequency modulation basis functions are used to model the joint features in the time-varying frequency domain, generating multi-transform domain enhanced features that can characterize temporal dynamics and suppress irrelevant noise, thereby addressing feature drift caused by seasonal changes and environmental evolution. Finally, visible blocks are filtered using a window mask, and cross-modal interaction is combined with the Transformer to generate multi-scale shared features. This effectively fuses spatial texture features from different sensors. Through query, key, and value matching and weighting, collaborative fusion at the feature and decision levels is achieved, avoiding the amplification of scene differences caused by the addition of modalities in traditional methods. This method reduces the model's reliance on labeled data in the target domain through frequency domain enhancement and cross-modal interaction, maintaining high generalization performance even in small sample sizes. In summary, this method achieves multi-source distribution alignment, dynamic feature enhancement, and efficient fusion and generalization through the cascade design of "score domain transformation → frequency modulation modeling → cross-modal attention", providing an interpretable and low-dependency solution for multi-source cross-regional object classification, which is particularly suitable for complex scenes of airborne remote sensing. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] To more clearly illustrate the embodiments of the present invention and its design, the following briefly introduces the drawings required for this embodiment. The drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be derived from these drawings without inventive effort.

[0015] Figure 1 The present invention is a flowchart of a multi-source environment interpretation method based on linear frequency modulation analysis according to an exemplary embodiment of the present invention.

[0016] Figure 2 Schematic diagram of a fractional domain feature extraction module and a frequency modulation analysis module provided according to an exemplary embodiment of the present invention; wherein (a) is a schematic diagram of the fractional domain multi-source feature extraction module; (b) is a schematic diagram of the frequency modulation analysis module.

[0017] Figure 3 A schematic diagram of a cross-modal attention network provided by the present invention according to an exemplary embodiment.

[0018] Figure 4 This is a block diagram of a multi-source environment interpretation device based on linear frequency modulation analysis according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0019] In order to enable those skilled in the art to better understand the technical solution of the present invention and to be able to implement it, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and are not intended to limit the scope of protection of the present invention.

[0020] The technical solutions provided by various embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0021] First, the present invention provides a multi-source environment interpretation method based on linear frequency modulation analysis, specifically Figure 1 As shown, the following steps are included: S101, obtaining hyperspectral images and lidar digital surface model data of the target area, extracting local neighborhood windows centered on pixels, and generating spatial spectrum data cubes and elevation data.

[0022] For example, for hyperspectral images, we extract an n×n×C cube centered around each pixel, where n = 5 or 7 and C is the number of spectral channels. For lidar data, we extract an n×n neighborhood window of the same spatial size, preserving the elevation channel information. Through these operations, we can obtain both spatial spectral data cubes and elevation data.

[0023] Next, the spatial-spectral data cube and elevation data need to be processed to obtain the terrain characteristics of the target area. In this step, a multi-source fusion model can be constructed and trained to obtain an environmental interpretation model. Based on this environmental interpretation model, the spatial-spectral data cube and elevation data are processed.

[0024] Among them, the environmental interpretation model includes a score domain feature extraction module, a frequency modulation analysis module and a cross-modal attention network.

[0025] In one embodiment, a fractional domain feature extraction module is first constructed, which includes the following: Figure 2 The fractional-domain space-frequency-phase perception block (a) performs a fractional Fourier transform on the input, converting it to the fractional domain of space-frequency mixture. The transform domain amplitude and phase are then fed into two 2D convolutional layers. Subsequently, a phase-aware function combines the frequency domain amplitude and phase information to generate a multi-source joint feature.

[0026] Then build the frequency modulation analysis module, which is based on the Chirp-based Fourier analysis method, such as Figure 2 As shown in (b), the multi-source joint feature is modeled in the time-varying frequency domain through the frequency modulation basis function to generate multi-transformation domain enhanced features.

[0027] Finally, a cross-modal attention network is constructed, which includes a window mask mechanism, a modality-sharing encoder (i.e., a Chirp-based multi-scale attention encoder), a Chirp-based multi-scale attention decoder, and a feature fusion and classifier. Specifically, two fractional domain signal extraction and analysis modules are used as basic units, and combined with nonlinear activation layers, a cross-modal deep attention mechanism network is designed, such as Figure 3 shown.

[0028] The constructed multi-source fusion model also needs to be trained. Specifically, training samples can be obtained, which include sample hyperspectral images and sample lidar digital surface model data of the same area and the real ground object classification results corresponding to the area; the sample hyperspectral images and sample lidar digital surface model data are input into the multi-source fusion model to obtain the output predicted ground object classification results; the multi-source fusion model is trained with the goal of minimizing the deviation between the real ground object classification results and the predicted ground object classification results.

[0029] For example, suppose the multi-source remote sensing data in the labeled source region is , and the unlabeled target area data is .in, and represent hyperspectral images and lidar data, respectively; and is the spatial size, and is the number of channels. is the known source region label space used for training (including class), and the test label space of the target area unknown. and Denote the source region and target region respectively. The feature fusion and classifier loss are defined as After completing the multimodal feature extraction and fusion, a domain adaptation module based on dual discriminators is designed to overcome the domain shift problem in cross-domain classification of multi-source remote sensing data. The discriminator loss is , combined with the aforementioned classifier loss and data reconstruction loss , the optimization objective function of the entire network is .

[0030] Then, based on the constructed source region training set, the network is trained and optimized end-to-end, so that it can effectively realize the interpretation of multi-source remote sensing cross-regional complex environments with low network parameters and small samples.

[0031] Finally, the target area test set is input into the trained classification model to obtain the corresponding predicted label. Compare and evaluate the performance of the proposed multi-source remote sensing cross-regional complex environment interpretation model.

[0032] S102, decomposing the spatial spectrum data cube and the elevation data into an amplitude spectrum and a phase spectrum respectively; performing weight mapping on the phase spectrum through a convolutional network, generating a modulation weight matrix and multiplying the matrix element-by-element with the amplitude spectrum, and outputting a multi-source joint feature.

[0033] Specifically, the spatial spectrum data cube and elevation data are subjected to fractional Fourier transform respectively, and decomposed into amplitude spectrum and phase spectrum; the phase spectrum is weighted mapped through a convolutional network to generate a modulation weight matrix and multiply it element-by-element with the amplitude spectrum to output multi-source joint features.

[0034] In this step, the spatial spectrum data cube and elevation data are input into the environmental interpretation model, and the spatial spectrum data cube and elevation data are respectively subjected to fractional Fourier transform through the fractional domain feature extraction module to decompose them into amplitude spectrum and phase spectrum; the phase spectrum is weighted mapped through the convolutional network to generate a modulation weight matrix and multiply it element-by-element with the amplitude spectrum to output multi-source joint features.

[0035] In one embodiment, after the spatial spectrum data cube and elevation data are input into the fractional domain feature extraction module, a fractional Fourier transform is performed on the spatial spectrum data cube corresponding to the hyperspectral data to decompose it into a first amplitude spectrum and a first phase spectrum; a fractional Fourier transform is performed on the elevation data corresponding to the lidar to decompose it into a second amplitude spectrum and a second phase spectrum; the first phase spectrum and the second phase spectrum are respectively input into a two-dimensional convolution layer to output a single-channel feature map; the single-channel feature map is normalized into a first modulation weight matrix and a second modulation weight matrix through a Sigmoid function; the first modulation weight matrix is ​​element-wise multiplied by the first amplitude spectrum to generate a hyperspectral fusion feature, and the second modulation weight matrix and the second amplitude spectrum are element-wise multiplied to generate a lidar fusion feature.

[0036] Taking the hyperspectral image data modality as an example, the input signal , the number of spectral channels is C, and the complex domain eigendecomposition after fractional Fourier transform is: (1) Among them, the amplitude spectrum is: (L2 norm), phase spectrum: (The main value is ), | is the Hadamard product (element-wise multiplication), j Is an imaginary unit.

[0037] Then, for the phase information, a phase weight generation module is designed to map the phase information into phase modulation weights through a convolutional network: (2) in, is the phase modulation weight, , is a learnable parameter, is the Sigmoid function, which constrains the weights to be within (0,1).

[0038] Assume cascade convolution operation , , then the amplitude and phase multi-domain fusion characteristics of the output hyperspectral data are: (3) The optimal modulation weight after normalization is: (4) Among them, ||·|| represents the matrix -norm, T 1 and T 2 are the cascade features in formula (3).

[0039] Similarly, Figure 2The fractional domain elevation-frequency-phase fusion feature of the lidar modality in (a) is defined as .

[0040] S103 , performing time-varying frequency domain modeling on the multi-source joint feature through a linear frequency modulation basis function to generate a multi-transformation domain enhanced feature.

[0041] After obtaining the multi-source joint features, the extracted multi-source joint features are analyzed using the Chirp-based Fourier analysis method of the frequency modulation analysis module. Specifically, a frequency modulation basis function is first defined, and the two multi-source joint features are linearly combined with the frequency modulation basis function to generate the first and second frequency domain components corresponding to the hyperspectral fusion features and the lidar fusion features. The frequency domain components are then reconstructed into spatial domain features through an inverse fractional Fourier transform, resulting in the corresponding first and second multi-transform domain enhancement features.

[0042] The multi-source joint features are modeled in the time-varying frequency domain through the linear frequency modulation basis function of the frequency modulation analysis module to generate multi-transformation domain enhanced features; the multi-transformation domain enhanced features are divided into non-overlapping blocks through the cross-modal attention network, and the non-overlapping blocks are filtered through the window mask to obtain visible blocks; the window Transformer is used to perform cross-modal interaction on the visible blocks to generate multi-scale shared features; based on the multi-scale shared features, query vectors, key vectors and value vectors are generated, the query vectors and key vectors are matched through the cross-modal attention mechanism, and the value vectors are weightedly fused based on the matching results to output reconstructed features; the reconstructed features are identified and classified to obtain the classification results of the target area objects.

[0043] for Figure 2 (b) Chirp-based Fourier analysis module. For any square integrable signal function , there exists a linear combination of frequency-modulated basis functions:

[0044] (5) in, represents the weight of the cosine component, represents the Chirp coefficient, k represents the index of the component, represents the weight of the sine component, is the trainable parameter matrix.

[0045] This makes the combination -norm converges to FM basis function composition Overcomplete framework.

[0046] In the proposed window attention mechanism, the local features of each window can be regarded as time-frequency signals. Chirp-FAN uses the frequency modulation basis function:

[0047] (6) The features within the local window are modeled in the time-varying frequency domain to generate multi-transform domain enhanced features.

[0048] Input to the Chirp-Fourier Analysis Layer (Chirp-FAN) Through learnable parameters and Perform nonlinear activation and map it into a combination of frequency modulated basis functions: (7) Among them, x is the multi-source joint feature of the input, and is the nonlinear branching parameter, is the activation function, is the Chirp frequency modulation coefficient.

[0049] S104, performing nonlinear activation and weighted fusion on the multi-transformation domain enhancement features based on the hyperspectral image and the lidar digital surface model data, outputting the reconstructed features, identifying and classifying the reconstructed features, and obtaining the ground object classification results.

[0050] Specifically, the multi-transform domain enhanced features are divided into non-overlapping blocks through the cross-modal attention network, and the non-overlapping blocks are filtered through the window mask to obtain visible blocks; the window Transformer is used to perform cross-modal interaction on the visible blocks to generate multi-scale shared features; based on the multi-scale shared features, the query vector, key vector and value vector are generated, the query vector and key vector are matched through the cross-modal attention mechanism, and the value vector is weightedly fused based on the matching results to output the reconstructed features; the reconstructed features are identified and classified to obtain the classification results of the target area objects.

[0051] Based on the above steps, the multi-transform domain enhancement feature includes a first multi-transform domain enhancement feature and a second multi-transform domain enhancement feature, corresponding to the hyperspectral image and lidar digital surface model data, respectively. Therefore, the multi-scale shared feature includes a first multi-scale shared feature corresponding to the first multi-transform domain enhancement feature and a second multi-scale shared feature corresponding to the second multi-transform domain enhancement feature. A query vector is then extracted from the first multi-scale shared feature to obtain a hyperspectral query feature. A key vector and a value vector are generated from the second multi-scale shared feature through frequency modulation projection; the key vector is the lidar frequency modulation key. The frequency domain similarity between the hyperspectral query feature and the lidar frequency modulation key is calculated to generate a cross-modal attention weight. This weight is then weightedly fused with the value vector and concatenated with the residual of the first multi-scale shared feature to obtain a reconstructed feature.

[0052] Specifically, first, the two fractional domain feature extraction and analysis units of formulas (3) and (6) are used as the basic structure, and the spatial enhancement features are divided into block formats through block embedding to generate a set of non-overlapping blocks. Then, the channel dimension is transformed from Map to Next, a window mask mechanism is introduced to design a multimodal mask learning strategy. The window mask mechanism is used to filter visible feature blocks and generate mask features that only contain a small part of the visible blocks.

[0053] Then, to model the relationship between modalities, a modality-sharing encoder is constructed to enhance feature consistency through cross-modal interaction. Assume that the specific features of hyperspectral and lidar modalities are and , the shared encoding process is:

[0054] (8) Where SwinT represents windowed Transformer, and [ , ] represents matrix concatenation. The encoded multimodal and multiscale features of the source region are and , the target area characteristics are and .

[0055] The decoder contains a shared decoder and two modality decoders for cross-modal reconstruction, such as Figure 3 As shown in the lower part. First, the shared decoder: alternately use two Chirp-based analysis modules and two upsampling to restore the spatial resolution. Let the decoder be , then the preliminary reconstruction features can be expressed as:

[0056] (9) Assume the encoder output is , the upsampling process utilizes the frequency modulation interpolation characteristics of Chirp-FAN. Secondly, modality decoder: Taking hyperspectral multi-domain feature reconstruction as an example: 1) Self-attention: First, the preliminary reconstruction features of the hyperspectral modality are Projected into query through three fully connected layers ,key Sum , learn the intra-modal frequency domain prior and obtain self-attention , improving the spectral continuity modeling capability. 2) Cross-attention interaction: Subsequently, in order to integrate multimodal complementarity to accurately reconstruct the hyperspectral modality, and After cascading, the mapping is to the inter-modal key, and the geometric features of the lidar are mapped to the cross-modal key through frequency modulation projection. , and hyperspectral query The frequency domain matching degree determines the cross-modal fusion attention .

[0057] Fusion and and with the value Multiply and reconstruct the hyperspectral modality , based on multi-domain features, reconstruct data It can be expressed as: (10) in, Generated by the frequency modulation weight layer of Chirp-FAN to achieve frequency domain cross-modulation, IFrFT stands for inverse fractional Fourier transform, FrFT stands for fractional Fourier transform, Indicates point addition, Represents dot product.

[0058] Finally, the multi-modal and multi-scale features of the source region after fusion encoding are , the target domain features are , align channels and spatial dimensions. Then, construct multimodal feature fusion and classifier to predict the final classification map of SD .

[0059] This method first decomposes spatial and elevation data into amplitude and phase spectra. The amplitude spectrum is dynamically adjusted through weight mapping of the phase spectrum, effectively aligning differences in multi-source distributions and enhancing feature robustness. This mitigates the problem of identical-object, heterogeneous spectra due to imaging conditions or sensor variations. Secondly, linear frequency modulation basis functions are used to model the joint features in the time-varying frequency domain, generating multi-transform domain enhanced features that characterize temporal dynamics and suppress irrelevant noise, thereby addressing feature drift caused by seasonal changes and environmental evolution. Finally, visible blocks are filtered using a window mask, and a Transformer is combined for cross-modal interaction to generate multi-scale shared features. This effectively fuses spatial texture features from different sensors. Through query, key, and value matching and weighting, a collaborative fusion of feature and decision levels is achieved, avoiding the amplification of scene differences that occurs with the addition of modalities in traditional methods. This method, through frequency domain enhancement and cross-modal interaction, reduces the model's reliance on labeled data in the target domain, maintaining high generalization performance even in small sample sizes. In summary, this method achieves multi-source distribution alignment, dynamic feature enhancement, and efficient fusion and generalization through the cascade design of "score domain transformation → frequency modulation modeling → cross-modal attention", providing an interpretable and low-dependency solution for multi-source cross-regional object classification, which is particularly suitable for complex airborne remote sensing scenarios.

[0060] Based on the above steps, the present invention also provides another embodiment.

[0061] This embodiment selects a public airborne multi-sensor source cross-city remote sensing image dataset to conduct a simulation experiment on the multi-source remote sensing cross-regional complex environment interpretation method proposed in the present invention to verify its feasibility and effectiveness.

[0062] This dataset contains two urban scenes, Nashua and Hanover, USA. The Nashua dataset provides 450×1050 pixel hyperspectral and lidar data at a spatial resolution of 1 meter; the Hanover dataset provides 700×620 pixel hyperspectral and lidar data. The hyperspectral data contains 114 spectral channels (0.42–0.95 μm) and seven ground object categories for experimental research and comparative analysis. Fifty labeled samples from each ground object category in the Nashua dataset were randomly selected for model training, while samples from the Hanover dataset were used for testing and validation. In addition, this implementation selected three novel hyperspectral image classification algorithms developed in the past three years for comparison: HighDAN (Remote Sens. Environ. 2023), MLUDA (IEEE Trans. Geosci. Remote Sens. 2024), and SCLUDA (IEEE Trans. Geosci. Remote Sens. 2023).

[0063] Table 1 Classification results of different classification algorithms on the Nashua-Hanover dataset (%) Table 1 shows the classification results (%) of the algorithm of the present invention and the above-mentioned comparison methods on the data set, including the classification accuracy of each class, overall classification accuracy (OA), average classification accuracy (AA), and Kappa coefficient. The results shown are the average values ​​of 10 random experiments.

[0064] According to Table 1, the proposed model improves the OA index by approximately 5% compared to current state-of-the-art models in terms of cross-city object classification performance. Experimental results demonstrate the effectiveness of the proposed multi-source remote sensing cross-regional complex environment interpretation method in improving the interpretation performance of small-sample cross-regional multi-source remote sensing data.

[0065] Aiming at the field of airborne remote sensing intelligent processing, this paper proposes a method for interpreting multi-source remote sensing cross-regional complex environments. Through network structure derivation, experimental results, and comparative analysis, the feasibility and effectiveness of the algorithm proposed in this paper in improving the interpretation performance of small-sample cross-regional multi-source remote sensing data are demonstrated, thereby providing a solution reference for the intelligent processing of remote sensing images in scenarios where airborne platforms observe regional changes and sample scarcity.

[0066] Secondly, the present invention also provides a multi-source environment interpretation device based on linear frequency modulation analysis, such as Figure 4 Shown, including: The acquisition module 201 is used to acquire the hyperspectral image and laser radar digital surface model data of the target area, extract the local neighborhood window centered on the pixel, and generate the spatial spectrum data cube and elevation data.

[0067] The construction module 202 is used to construct a multi-source fusion model and train the multi-source fusion model to obtain an environment interpretation model; the environment interpretation model includes a score domain feature extraction module, a frequency modulation analysis module and a cross-modal attention network.

[0068] The fractional domain feature extraction module 203 is used to decompose the spatial spectrum data cube and elevation data into an amplitude spectrum and a phase spectrum; perform weight mapping on the phase spectrum through a convolutional network, generate a modulation weight matrix, and multiply the matrix element-by-element with the amplitude spectrum to output a multi-source joint feature.

[0069] The frequency modulation analysis module 204 is configured to perform time-varying frequency domain modeling on the multi-source joint features using linear frequency modulation basis functions to generate multi-transformation domain enhanced features.

[0070] The cross-modal attention network 205 is used to divide the multi-transformation domain enhanced features into non-overlapping blocks, and filter the non-overlapping blocks through a window mask to obtain visible blocks; use the window Transformer to perform cross-modal interaction on the visible blocks to generate multi-scale shared features; generate a query vector, a key vector and a value vector based on the multi-scale shared features, match the query vector and the key vector through a cross-modal attention mechanism, and perform weighted fusion on the value vector based on the matching results to output reconstructed features; identify and classify the reconstructed features to obtain the classification results of the target area objects.

[0071] Using this device, spatial and elevation data are first decomposed into amplitude and phase spectra. The amplitude spectrum is dynamically adjusted through weight mapping of the phase spectrum, effectively aligning differences in multi-source distributions and enhancing feature robustness. This mitigates the problem of identical-object, heterogeneous spectra caused by imaging conditions or sensor variations. Secondly, linear frequency modulation basis functions are used to model the joint features in the time-varying frequency domain, generating multi-transform domain enhanced features that characterize temporal dynamics and suppress irrelevant noise, thereby addressing feature drift caused by seasonal changes and environmental evolution. Finally, visible blocks are filtered using a window mask, and cross-modal interaction is combined with the Transformer to generate multi-scale shared features. This effectively fuses spatial texture features from different sensors. Through query, key, and value matching and weighting, collaborative fusion at the feature and decision levels is achieved, avoiding the amplification of scene differences that occurs with the addition of modalities in traditional methods. This method, through frequency domain enhancement and cross-modal interaction, reduces the model's reliance on labeled data in the target domain, maintaining high generalization performance even in small sample sizes. In summary, this method achieves multi-source distribution alignment, dynamic feature enhancement, and efficient fusion and generalization through the cascade design of "score domain transformation → frequency modulation modeling → cross-modal attention", providing an interpretable and low-dependency solution for multi-source cross-regional object classification, which is particularly suitable for complex scenes of airborne remote sensing.

[0072] The present invention also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 The steps of a multi-source environment interpretation method based on linear frequency modulation analysis are provided.

[0073] The present invention also provides a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The steps of a multi-source environment interpretation method based on linear frequency modulation analysis are provided.

[0074] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0075] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0076] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0077] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0078] It should be noted that the specific embodiments described above can enable those skilled in the art to more fully understand the present invention, but do not limit the present invention in any way. Therefore, although this specification has described the present invention in detail, those skilled in the art should understand that the present invention can still be modified or replaced with equivalents; and all technical solutions and improvements that do not depart from the spirit and scope of the present invention are included in the scope of protection of the patent for the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. A multi-source environment interpretation method based on linear frequency modulation analysis, characterized in that: The method comprises: Obtain hyperspectral images and lidar digital surface model data of the target area, and extract local neighborhood windows centered on pixels to generate spatial spectrum data cubes and elevation data; Decomposing the spatial spectrum data cube and the elevation data into an amplitude spectrum and a phase spectrum respectively; performing weight mapping on the phase spectrum through a convolutional network, generating a modulation weight matrix and multiplying the matrix element-by-element with the amplitude spectrum to output a multi-source joint feature; Performing time-varying frequency domain modeling on the multi-source joint features through linear frequency modulation basis functions to generate multi-transformation domain enhanced features; The multi-transform domain enhanced features are divided into non-overlapping blocks, and the non-overlapping blocks are filtered through a window mask to obtain visible blocks; the window Transformer is used to perform cross-modal interaction on the visible blocks to generate multi-scale shared features; based on the multi-scale shared features, a query vector, a key vector and a value vector are generated, the query vector and the key vector are matched through a cross-modal attention mechanism, and the value vector is weightedly fused based on the matching results to output the reconstructed features; the reconstructed features are identified and classified to obtain the classification results of the target area objects.

2. The multi-source environment interpretation method based on linear frequency modulation analysis according to claim 1, characterized in that: The multi-source joint features include hyperspectral fusion features and lidar fusion features. The phase spectrum is weighted mapped by a convolutional network to generate a modulation weight matrix and multiply it element-by-element with the amplitude spectrum. The output multi-source joint features include: Performing fractional Fourier transform on the empty spectrum data cube to decompose it into the first magnitude spectrum and the first phase spectrum; Performing fractional Fourier transform on the elevation data to decompose it into a second magnitude spectrum and a second phase spectrum; Input the first phase spectrum and the second phase spectrum into the two-dimensional convolution layer respectively, and output a single-channel feature map; Normalize the single-channel feature map into a first modulation weight matrix and a second modulation weight matrix through a Sigmoid function; The first modulation weight matrix is ​​multiplied element-by-element with the first amplitude spectrum to generate a hyperspectral fusion feature, and the second modulation weight matrix is ​​multiplied element-by-element with the second amplitude spectrum to generate a lidar fusion feature.

3. The multi-source environment interpretation method based on linear frequency modulation analysis according to claim 2, characterized in that: Performing time-varying frequency domain modeling on the multi-source joint features by using a linear frequency modulation basis function to generate multi-transformation domain enhanced features includes: The frequency modulation basis function is defined as: ; in, x is the multi-source joint feature of the input, represents the Chirp frequency modulation coefficient, W P represents the nonlinear branching parameter; Decomposing the two multi-source joint features into linear combinations of frequency modulation basis functions respectively to generate a first multi-transform domain component and a second multi-transform domain component corresponding to the hyperspectral fusion feature and the lidar fusion feature; The transform domain components are reconstructed into spatial domain features through fractional Fourier transform to obtain corresponding first multi-transform domain enhanced features and second multi-transform domain enhanced features.

4. The multi-source environment interpretation method based on linear frequency modulation analysis according to claim 3, characterized in that: The multi-scale shared features include a first multi-scale shared feature corresponding to the first multi-transform domain enhancement feature and a second multi-scale shared feature corresponding to the second multi-transform domain enhancement feature; Generate query vectors, key vectors, and value vectors based on multi-scale shared features. Match the query vectors and key vectors through a cross-modal attention mechanism, and perform weighted fusion on the value vectors based on the matching results. The output reconstructed features include: Extracting a query vector from the first multi-scale shared feature to obtain a hyperspectral query feature; generating a key vector and a value vector from the second multi-scale shared feature through frequency modulation projection; the key vector is the lidar frequency modulation key; Calculate the frequency domain similarity between the hyperspectral query feature and the lidar frequency modulation key to generate cross-modal attention weights; The weights are weightedly fused with the value vector and connected with the first multi-scale shared feature residual to obtain the reconstructed feature.

5. A multi-source environment interpretation device based on linear frequency modulation analysis, characterized in that: The device comprises: The acquisition module is used to obtain the hyperspectral image and lidar digital surface model data of the target area, extract the local neighborhood window centered on the pixel, and generate the spatial spectrum data cube and elevation data; A construction module is used to construct a multi-source fusion model and train the multi-source fusion model to obtain an environment interpretation model; the environment interpretation model includes a score domain feature extraction module, a frequency modulation analysis module, and a cross-modal attention network; The fractional domain feature extraction module is used to decompose the spatial spectrum data cube and elevation data into an amplitude spectrum and a phase spectrum; perform weight mapping on the phase spectrum through a convolutional network, generate a modulation weight matrix, and multiply it element-by-element with the amplitude spectrum to output a multi-source joint feature; The frequency modulation analysis module is used to perform time-varying frequency domain modeling on the multi-source joint features through linear frequency modulation basis functions to generate multi-transformation domain enhanced features; The cross-modal attention network is used to divide the multi-transform domain enhanced features into non-overlapping blocks, and filter the non-overlapping blocks through a window mask to obtain visible blocks; use the window Transformer to perform cross-modal interaction on the visible blocks to generate multi-scale shared features; generate a query vector, a key vector and a value vector based on the multi-scale shared features, match the query vector and the key vector through a cross-modal attention mechanism, and perform weighted fusion on the value vector based on the matching results to output reconstructed features; identify and classify the reconstructed features to obtain the classification results of the target area objects.

6. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

7. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 4 when executing the program.

Citation Information

Cited By

  • Hyperspectral cross-domain wetland mapping method and system of fractional order Fourier RWKV network

    CN122024074A

  • Hyperspectral cross-domain wetland mapping method and system based on fractional fourier rwkv network

    CN122024074B

  • Source remote sensing data fusion and intelligent interpretation method and system for digital agriculture

    CN122388614A