HSI and ranging data domain classification method and system based on domain expansion and feature decoupling

By fusing HSI and LiDAR data and using frequency domain amplitude perturbation adversarial domain extension technology, the problems of single data modality, insufficient domain offset simulation, and susceptibility of features to noise interference in cross-domain classification of hyperspectral images and LiDAR data are solved, achieving high-precision and low-cost cross-domain classification.

CN121788922APending Publication Date: 2026-04-03XIDIAN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies for cross-domain classification of hyperspectral images and LiDAR data suffer from problems such as single data modality, insufficient simulation of domain offset, susceptibility of mixed features to source domain noise interference, and insufficient deployment flexibility, resulting in low classification accuracy and high application costs.

Method used

By fusing HSI and LiDAR data, combining frequency domain amplitude perturbation and adversarial domain extension techniques, and through multi-stream feature extraction and explicit decoupling, feature weights are adaptively adjusted to achieve cross-domain classification.

Benefits of technology

It improves cross-domain classification accuracy, enhances the model's adaptability to unknown target domains, reduces application costs, and expands the coverage of source domain distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121788922A_ABST
    Figure CN121788922A_ABST
Patent Text Reader

Abstract

The invention discloses an HSI and ranging data cross-domain classification method and system based on domain expansion and feature decoupling, and aims to solve the problems of single mode, insufficient domain offset simulation and weak generalization ability in the prior art. Preprocessing the HSI and LiDAR data to obtain a source domain sample, and generating an expansion domain sample through frequency domain amplitude disturbance and adversarial domain expansion; extracting basic features by using a multi-stream convolutional neural network, and separating domain specific features and cross-modal and cross-domain domain invariant features; and adaptively fusing the three types of features through an attention network, and inputting the preprocessed target domain data into a trained model to obtain a classification result. Through multi-modal fusion, mixed domain expansion and feature decoupling, the cross-domain classification precision and generalization ability are improved, fine adjustment of target domain data is not needed, application is flexible, and the method is suitable for multi-scene remote sensing ground feature classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing technology, specifically relating to a method and system for cross-domain classification of hyperspectral image (HSI) light detection and ranking (LiDAR) data based on domain extension and feature decoupling in the field of multimodal remote sensing data cross-domain classification technology. This invention can be used to classify different types of land features in hyperspectral images in various fields such as agricultural yield estimation, environmental monitoring, and mineral exploration. Background Technology

[0002] In the field of remote sensing Earth observation, hyperspectral images (HSI) can capture the fine spectral features of ground objects due to their extremely high spectral resolution. However, the imaging process of hyperspectral images is highly susceptible to external environmental factors (such as light intensity, atmospheric transmission characteristics, and solar altitude angle) and sensor characteristics (such as noise level and spectral response function), leading to significant statistical distribution differences between hyperspectral data collected at different times and locations. This phenomenon is known as "domain shift." Traditional supervised classification methods are usually based on the assumption that "training data (source domain) and test data (target domain) follow the same probability distribution." When this assumption is not valid (i.e., domain shift exists), directly applying a model trained in the source domain to the target domain will result in a sharp decline in classification performance.

[0003] Nanjing University of Science and Technology disclosed a method for panchromatic sharpening of multispectral remote sensing images based on a frequency domain decomposition network in its patent application, "A Panchromatic Sharpening Method for Multispectral Remote Sensing Images Based on Frequency Domain Decomposition Network" (Application No. 202510123236.2, Publication No. CN 119579847 A). This method constructs a frequency domain decomposition and domain generalization network, which includes three core modules: a dynamic convolutional low-frequency extraction module, a high-frequency preservation module, and a spectral enhancement module. It also proposes a domain generalization training strategy, employing a generative adversarial network (GAN) structure and introducing an adversarial loss function. By intentionally misassigning domain labels, the generator is guided to achieve better cross-domain generalization capabilities under the supervision of the discriminator. This method effectively captures information at different scales, improves the quality of fused images, and achieves accurate fusion of invisible satellite data. However, this method still has the following shortcomings: First, the data modalities are limited to multispectral and panchromatic images, lacking multimodal information complementarity for elevation and geometric structure information, making it difficult to solve the core problem of "different objects with the same spectrum" and "different spectra for the same object" in land cover classification; Second, the domain offset simulation relies solely on training strategies for domain label misassignment and adversarial loss, without incorporating frequency domain perturbations based on the physical mechanisms of remote sensing imaging, thus failing to accurately reproduce complex nonlinear spectral distortions caused by environmental factors such as illumination and atmosphere, and having a limited source domain distribution coverage; Third, it cannot separate domain-specific information that varies with the domain from domain-invariant information that reflects the essence of land cover, making the model prone to fitting source domain-specific noise and limiting its cross-domain generalization ability; Fourth, the core objective is panchromatic image sharpening (fusion), rather than land cover classification, and it does not optimize feature discriminativeness for classification needs, making it unable to directly adapt to cross-domain land cover classification scenarios.

[0004] Harbin University of Science and Technology disclosed a collaborative classification system for hyperspectral images and LiDAR data using frequency domain feature learning combined with CNN in its patent application, "Cooperative Classification System for Hyperspectral Images and LiDAR Data Based on Frequency Domain Feature Learning and CNN" (Application No. 202510431731.X, Publication No. CN 120279330 A). This system decomposes the low-frequency spectral features and high-frequency spatial details of hyperspectral data using wavelet transform, and employs a hybrid convolutional architecture to achieve hierarchical feature extraction. Simultaneously, it designs a dynamic channel weighting mechanism to fuse multi-scale elevation features from LiDAR, and enhances the accuracy of terrain feature detail extraction through a residual attention mechanism. Finally, it constructs a frequency-spatial dual-stream feature interaction module to achieve cross-modal complementary feature alignment and adaptive fusion. This system improves cross-modal fusion efficiency and enhances the ability to acquire high-frequency detail information. However, the system still has the following shortcomings: First, it only decomposes existing data features through wavelet transform, without simulating potential domain shifts in cross-domain scenarios, resulting in insufficient coverage of source domain data distribution and poor adaptability when facing unknown target domains at different times and locations. Second, the extracted features are only a mixture of domain-specific and domain-invariant information, making the model susceptible to noise interference from the source domain environment and limiting its cross-domain generalization performance. Third, the modal fusion uses fixed logic for dynamic channel weighting and dual-stream interaction, which cannot dynamically allocate the weights of LiDAR features and domain-invariant features according to the degree of interference with HSI spectral information (such as in scenarios with poor lighting), making it difficult to achieve accurate fusion that "plays to its strengths and avoids its weaknesses." Fourth, deployment requires data adaptation or model fine-tuning for the target domain, resulting in high application costs and insufficient flexibility. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of the existing technologies by providing a cross-domain classification method and system for hyperspectral and LiDAR data based on domain extension and feature decoupling. This aims to solve the problems of existing methods having a single data modality, making it difficult to handle complex scenarios such as "different objects with the same spectrum" or "different spectra with the same object"; failing to accurately reproduce complex nonlinear domain shifts in remote sensing imaging, resulting in insufficient source domain coverage; extracting mixed features being susceptible to interference from specific noise in the source domain; and relying on target domain data (including unlabeled data), leading to insufficient deployment flexibility and high application costs.

[0006] The technical approach to achieving the objectives of this invention is as follows: Addressing the issue of single data modality, this invention employs HSI and LiDAR data fusion technology. HSI provides fine spectral information, while LiDAR provides elevation and geometric structure information. Furthermore, LiDAR data is insensitive to changes in shadow and illumination, creating a natural complementarity between the two. This compensates for the limitations of single-modal information discrimination, thereby resolving the low classification accuracy issues caused by "different objects with the same spectrum" or "different spectra for the same object." Addressing the problem of insufficient domain shift simulation, this invention employs a hybrid domain expansion technique combining frequency domain and adversarial methods. Frequency domain amplitude perturbation can simulate spectral style shifts caused by changes in physical environments such as illumination and atmosphere. Adversarial sample generation can uncover model weaknesses and cover extreme domain shift scenarios. This dual strategy significantly expands the source domain distribution range, thereby solving the problem that existing simple data augmentation techniques cannot restore complex domain shifts. To address the issue of insufficient feature decoupling, this invention employs multi-stream feature extraction and explicit decoupling techniques. A dedicated extractor separates domain-specific information from domain-invariant information, and then modal attention fusion technology is combined to adaptively adjust the weights of each feature. This avoids domain-specific noise interference while fully leveraging the advantages of multimodal collaboration, thus solving the problem of weak generalization ability of mixed features. To address the issue of insufficient deployment flexibility, this invention adopts a pure domain generalization design approach. It optimizes only the source domain data through domain expansion and feature decoupling, requiring no data from the target domain, thus overcoming the limitations of existing methods that rely on target domain data and have restricted application scenarios.

[0007] The implementation steps of the method of the present invention include the following:

[0008] Step 1: Preprocess the hyperspectral image HSI and LiDAR data of the target region to obtain source region samples;

[0009] Step 2: A hybrid strategy combining frequency domain amplitude perturbation and adversarial domain extension is adopted to perform frequency domain amplitude perturbation on the preprocessed HSI and LiDAR data to generate frequency domain enhanced samples. Adversarial samples are generated based on the frequency domain enhanced samples. The frequency domain enhanced samples and adversarial samples are then weighted and fused at a preset ratio to obtain extended domain samples.

[0010] Step 3: Use a convolutional neural network to extract the basic features from the source domain samples and the extended domain samples respectively; use the domain-specific extractor and the domain-invariant extractor in the convolutional neural network to extract the domain-specific features corresponding to the source domain samples and the extended domain samples, as well as the cross-modal and cross-domain domain-invariant features.

[0011] Step 4: Input the extracted domain-specific features and domain-invariant features into the attention network to generate adaptive attention weights. Then, perform weighted fusion of the HSI domain-specific features, LiDAR domain-specific features, and domain-invariant features to obtain a fused feature vector.

[0012] Step 5: Input the fused feature vector into the classification head of the convolutional neural network to output the probability distribution of land cover categories; use a multi-loss joint optimization strategy to iteratively update the parameters of the cross-domain classification model to obtain the trained cross-domain classification model.

[0013] Step 6: Input the HSI data and LiDAR data of the target domain to be classified, which have been preprocessed in the same way as in Step 1, into the trained cross-domain classification model and output the land cover classification results.

[0014] Furthermore, the step of performing frequency domain amplitude perturbation processing on the preprocessed HSI and LiDAR data to generate frequency domain enhanced samples is as follows:

[0015] The first step is to perform two-dimensional discrete cosine transform (DCT) on the preprocessed HSI slice samples and LiDAR slice samples respectively, so as to transform the data from the spatial domain to the frequency domain and obtain the frequency domain matrix.

[0016] The second step involves using a Gaussian low-pass filter with a cutoff frequency of 0.3 to extract low-frequency components containing domain-specific information such as illumination and style from the frequency domain matrix, while retaining high-frequency components containing semantic information such as ground texture and structure.

[0017] The third step is to multiply the amplitude spectrum of the extracted low-frequency components by a random scaling factor α, which is uniformly sampled in the interval [0.7, 1.3] while keeping the phase spectrum unchanged; and to apply a weak perturbation of ±5% to the retained high-frequency components.

[0018] The fourth step is to perform inverse two-dimensional discrete cosine transform (IDCT) on the frequency domain data after the above processing to reconstruct the data back into the spatial domain.

[0019] The fifth step involves introducing a reconstruction coefficient of 0.85 to fuse the reconstructed data across different bands, generating frequency-enhanced samples.

[0020] Furthermore, the steps for generating adversarial examples are as follows:

[0021] The first step is to use the projected gradient descent algorithm as the adversarial example generation algorithm;

[0022] The second step is to set the algorithm parameters: perturbation step size α = 0.01, number of iterations T = 10, and perturbation boundary [0,1].

[0023] The third step is to generate adversarial examples for the current iteration according to the following formula:

[0024] ;

[0025] in, This represents the adversarial example in the (t+1)th iteration. This represents the projection operation within the [0,1] perturbation boundary. Let represent the adversarial example in the t-th iteration. Represents a symbolic function. This represents the gradient used to calculate the loss function with respect to the input samples; Let f(x) represent the cross-entropy loss function, f(x) represent the feature extraction and classification mapping function of the cross-domain classification model in the current training phase, and y represent the true land cover category label of the sample.

[0026] The cross-entropy loss function for: ;in, Let the i-th element in the one-hot encoded vector representing the true land cover category label be the element that determines whether the sample belongs to the i-th class. , Let C represent the probability value of the i-th type of land cover predicted by the cross-domain classification model, and let C represent the total number of land cover categories. This represents a logarithmic operation with base 10.

[0027] Fourth step: Repeat the above steps and use the adversarial sample obtained after T iterations as the final generated adversarial sample.

[0028] Furthermore, the preset ratio is determined in the following way: a pre-trained basic feature extraction network with the same structure as the cross-domain classification model is used to extract feature vectors of source domain data and a small amount of target domain reference data, and the maximum mean difference between the two is calculated; the weights ω1 of the frequency domain augmented samples and ω2 of the adversarial samples are set to satisfy ω1+ω2=1, and dynamically adjusted according to the result of the maximum mean difference: if the result of the maximum mean difference shows that the style difference is mainly concentrated in the low-frequency spectral distribution, then ω1∈[0.6,0.8] and ω2∈[0.2,0.4] are set; if the result of the maximum mean difference shows that the target domain contains complex sensor noise or abnormal environment, then ω1∈[0.3,0.5] and ω2∈[0.5,0.7].

[0029] Furthermore, the convolutional neural network is composed of four functionally differentiated feature extractor modules cascaded together: a shared-weight CNN basic feature extractor, an HSI domain-specific extractor, a LiDAR domain-specific extractor, and a domain-invariant extractor. Specifically, the shared-weight CNN basic feature extractor sequentially performs the following signal transformation operations on the preprocessed HSI and LiDAR samples, outputting a basic feature signal with a dimension of 2×2×128: ​​① Conv2d → ReLU activation → BatchNorm2d normalization; where the parameters of Conv2d are set to: 64 output channels, 3×3 kernel size, stride 1, and padding 1; ② MaxPool2d pooling, where the parameters of MaxPool2d are set to: 2×2 pooling kernel size and stride 2; ③ Conv2d → ReLU activation → BatchNorm2d normalization; where the parameters of Conv2d are set to: 128 output channels, 3×3 kernel size, stride 1, and padding 1; ④ MaxPool2d pooling; the parameters of MaxPool2d are set to: pooling kernel size 2×2, stride 2; the HSI domain-specific extractor concatenates and fuses the basic feature signals of HSI source domain samples and HSI extended domain samples output by the shared weight CNN basic feature extractor, and performs the following signal transformation operations in sequence to form an HSI domain-specific feature signal with an output dimension of 2×2×256: ① Conv2d→ReLU activation; where the parameters of Conv2d are set to: number of output channels 64, convolution kernel size 3×3, padding 1; ② Conv2d; the parameters of Conv2d are set to: number of output channels 32, convolution kernel size 1×1; the LiDAR domain-specific extractor concatenates and fuses the basic feature signals of LiDAR source domain samples and LiDAR extended domain samples output by the shared weight CNN basic feature extractor, and performs the following signal transformation operations in sequence to form a LiDAR domain-specific feature signal with a dimension of 2×2×256. The output signal of domain-specific features; adopts the same two-layer CNN network signal transformation structure as the HSI domain-specific extractor: ① Conv2d→ReLU activation; where the parameters of Conv2d are set as follows: 64 output channels, 3×3 kernel size, and 1 padding; ② Conv2d; the parameters of Conv2d are set as follows: 32 output channels and 1×1 kernel size;The domain-invariant extractor concatenates and fuses the basic feature signals of four types of samples—HSI source domain, LiDAR source domain, HSI extended domain, and LiDAR extended domain—output by the shared-weight CNN basic feature extractor, forming a cross-modal, cross-domain domain-invariant feature output signal with a dimension of 2×2×512. Signal processing employs a three-layer shared-weight CNN network to perform feature signal transformation, with the following steps: ① Conv2d → ReLU activation; where Conv2d parameters are set to: 128 output channels, 3×3 kernel size, and 1 padding; ② Conv2d → ReLU activation; where Conv2d parameters are set to: 64 output channels, 3×3 kernel size, and 1 padding; ③ Conv2d; where Conv2d parameters are set to: 64 output channels, 1×1 kernel size.

[0030] Further, the steps for extracting domain-specific features for each modality are as follows: First, using a shared-weight CNN basic feature extractor, feature encoding is performed on the original HSI samples and LiDAR samples in the source domain, as well as the HSI extended domain samples and LiDAR extended domain samples in the extended domain, outputting a 128-dimensional shallow basic feature map with a dimension of 2×2×128; Second, the encoded features of the original HSI samples and the encoded features of the extended HSI samples are concatenated to obtain concatenated features with a dimension of 2×2×256, which serve as the input to the HSI domain-specific extractor; Third, the encoded features of the original LiDAR samples and the encoded features of the extended LiDAR samples are concatenated to obtain concatenated features with a dimension of 2×2×256, which serve as the input to the LiDAR domain-specific extractor; Fourth, the input features of the HSI domain-specific extractor are fed into a preset two-layer CNN network, and the network calculates and outputs the HSI domain-specific features; Fifth, the LiDAR... The input features of the domain-specific extractor are fed into a two-layer CNN network with the same structure as the HSI domain-specific extractor, and the network calculates and outputs LiDAR domain-specific features.

[0031] Furthermore, the cross-modal and cross-domain domain-invariant features are obtained in the following way: the feature signals of four types of samples output by the shared-weight CNN basic feature extractor—HSI source domain, LiDAR source domain, HSI extended domain, and LiDAR extended domain—are cascaded and fused, and then the feature signals are obtained after feature transformation by a three-layer shared-weight CNN network. After processing by the above network, the feature signals are not affected by domain offset factors such as light intensity, atmospheric transmission characteristics, solar elevation angle, sensor noise level, and spectral response function, and can stably characterize the inherent essential properties of ground objects.

[0032] Furthermore, the steps of the multi-loss joint optimization strategy for iteratively updating the parameters of the cross-domain classification model are as follows:

[0033] The first step is to use the Adam optimizer, setting the initial learning rate to 0.0001, the weight decay factor to 1e-5, and the batch size to 32;

[0034] The second step involves inputting the fused feature vectors of the source domain samples and the extended domain samples into the cross-domain classification model, iteratively updating the model parameters until the model's total loss function converges. The convergence condition is that the fluctuation range of the total loss function value is less than 1e-4 within 10 consecutive iterations, resulting in a trained cross-domain classification model. The total loss function is as follows: ;in, Represents the cross-entropy classification loss; The domain difference loss is represented by the maximum mean difference loss, which is used to calculate the difference in feature distributions between the source domain and the extended domain. The feature consistency loss is represented by the mean squared error loss, which constrains the semantic consistency between domain-specific features and domain-invariant features. and The loss balance coefficients are set to 0.5 and 0.3 respectively.

[0035] The cross-domain classification model consists of six cascaded functional modules: a shared-weight CNN basic feature extractor, an HSI domain-specific extractor, a LiDAR domain-specific extractor, a domain-invariant extractor, an attention network module, and a classification head.

[0036] The structure and parameters of the four functional modules—shared weight CNN basic feature extractor, HSI domain-specific extractor, LiDAR domain-specific extractor, and domain-invariant extractor—are the same as the structure and parameters of the convolutional neural network described in claim 5.

[0037] The attention network module consists of a cascaded global average pooling layer and a fully connected layer. The input of the global average pooling layer is a concatenated feature of HSI domain-specific features, LiDAR domain-specific features, and domain-invariant features, and the output is a dimension of 1×1×(32+32+64)=1×1×128. The fully connected layer has an input dimension of 128 and an output dimension of 3 → a Sigmoid activation function, with the outputs corresponding to the three attention weights of HSI domain-specific features, LiDAR domain-specific features, and domain-invariant features, respectively.

[0038] The classification head consists of a fully connected layer and a cascaded Softmax activation function layer; the input dimension of the fully connected layer is 32 + 32 + 64 = 128, and the output dimension is equal to the total number of land cover categories C.

[0039] The present invention provides a cross-domain classification system for HSI and ranging data based on domain extension and feature decoupling, comprising the following modules: wherein,

[0040] Data preprocessing module: used to preprocess the hyperspectral image HSI and LiDAR data of the target region to obtain source domain samples;

[0041] Hybrid Domain Extension Module: Includes a frequency domain amplitude perturbation submodule and an adversarial domain extension submodule. The frequency domain amplitude perturbation submodule is used to perform frequency domain amplitude perturbation processing on the preprocessed HSI and LiDAR data to generate frequency domain enhanced samples. The adversarial domain extension submodule is used to generate adversarial samples based on the frequency domain enhanced samples. The frequency domain enhanced samples and adversarial samples are weighted and fused to obtain extended domain samples.

[0042] Multi-stream feature extraction and decoupling module: Composed of a shared-weight CNN basic feature extractor, an HSI domain-specific extractor, a LiDAR domain-specific extractor, a domain-invariant extractor, and a dual-loss constraint unit, it realizes the explicit separation of basic feature encoding, domain-specific information, and domain-invariant information, while ensuring the completeness and discriminative power of the features;

[0043] Modal attention fusion module: includes a global average pooling layer and a fully connected attention network, used to adaptively generate attention weights for various features, realize dynamic weighted fusion of domain-specific features and domain-invariant features, and highlight the role of high-contribution features;

[0044] The classification prediction and optimization module consists of a classification head and a multi-loss joint optimization unit. It maps the fused features to the probability distribution of land cover categories and updates the model parameters iteratively through multi-loss collaborative optimization to improve the model's classification accuracy and cross-domain generalization ability.

[0045] Compared with the prior art, the present invention has the following advantages.

[0046] First, the method of this invention introduces LiDAR modal data and utilizes its elevation / geometric information, which is insensitive to changes in shadow and illumination, to complement the spectral information of HSI. This overcomes the shortcomings of insufficient complementarity of single modal information in existing technologies and effectively solves the classification problems of "different objects with the same spectrum" and "different spectra of the same object". As a result, the overall accuracy (OA) of cross-domain classification of this invention is improved by 5%-15%, which is significantly better than the single hyperspectral domain generalization method.

[0047] Second, the method of this invention employs a hybrid strategy combining frequency domain amplitude perturbation and adversarial domain extension, overcoming the shortcomings of existing technologies such as simple domain migration simulation and insufficient source domain coverage. This significantly enhances the adaptability of the model to unknown target domains. It can simulate spectral style changes caused by physical environments such as illumination and atmosphere, and also cover extreme domain migration scenarios where the model is prone to errors, greatly expanding the range of source domain distribution.

[0048] Third, the present invention constructs a feature decoupling and modal attention fusion mechanism, which overcomes the shortcomings of existing technologies such as the susceptibility of mixed features to source domain noise pollution and low feature utilization. It can effectively separate domain-specific information from domain-invariant information, adaptively adjust the weights of each feature, give full play to the advantages of multimodal collaboration, and improve the Kappa coefficient of the model by more than 8% on average in the unknown target domain.

[0049] Fourth, the module of this invention adopts a pure domain generalization implementation method, which overcomes the shortcomings of existing systems that rely on target domain data and have poor deployment flexibility. Training can be completed using only source domain data, without the need for target domain data preprocessing or model fine-tuning, thus reducing application costs. At the same time, the core module has good scalability, can be adapted to different types of hyperspectral and LiDAR data, and can also be extended to cross-domain classification tasks of other multimodal remote sensing data, making it more widely applicable. Attached Figure Description

[0050] Figure 1 This is a flowchart of the present invention;

[0051] Figure 2 This is a flowchart of the domain extension module of the present invention. Detailed Implementation

[0052] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0053] Referring to FIG1, the implementation steps of an embodiment of the method of the present invention will be further described.

[0054] Step 1, data preprocessing.

[0055] Data Reading: Professional geographic data processing tools are used to read HSI and LiDAR data, and corresponding land cover category labels are extracted simultaneously to ensure a one-to-one correspondence between data and labels, providing a foundation for subsequent supervised training.

[0056] Normalization: The Min-Max normalization method is used to map HSI data and LiDAR data to the [0,1] interval respectively. The core purpose is to eliminate the interference caused by the difference in the scale of data of different modalities, so that the model can learn the features of the two types of data fairly and improve the training stability.

[0057] Data tiling: The normalized HSI and LiDAR data are tiled in 9×9 pixel sizes with a step size of 3 pixels. The core of this parameter setting is to preserve the spatial context information of ground features by utilizing the overlapping areas of adjacent tiles, avoiding feature loss due to tile fragmentation, and ensuring that the model can learn the complete spatial relationships of ground features.

[0058] Step 2, hybrid domain extension.

[0059] Referring to Figure 2, the specific operation process of generating extended domain samples through a dual strategy of frequency domain amplitude perturbation and adversarial domain extension is further described.

[0060] Frequency domain perturbation processing: Two-dimensional discrete cosine transforms are performed on the HSI and LiDAR slice samples of the training set to convert the data from the spatial domain to the frequency domain, achieving frequency domain separation of information; a Gaussian low-pass filter with a cutoff frequency of 0.3 is constructed to filter out low-frequency components containing domain-specific information such as illumination and style, while retaining high-frequency components reflecting core semantic information such as ground texture and structure; random scaling adjustment (scaling range 0.7-1.3) is applied to the amplitude spectrum of the low-frequency components to maintain the stability of the data structure while keeping the phase spectrum unchanged; a small perturbation is applied to the high-frequency components to enhance feature diversity; the processed frequency domain data is inversely transformed back to the spatial domain, and a reconstruction coefficient of 0.85 is introduced to fuse the data of each band, generating frequency domain enhanced samples that balance realism and domain offset simulation effects.

[0061] Adversarial example generation: Adversarial examples are generated using the projective gradient descent algorithm. A small perturbation step size and a reasonable number of iterations are set, while the perturbation boundary is limited to be consistent with the range of normalized data. The core logic is to calculate the gradient of the loss function with respect to the input sample based on the difference between the model prediction results and the true labels of the samples in the current training stage, and gradually apply small perturbations along the gradient direction. After multiple iterations, adversarial examples that can expose the weaknesses of the model are generated, thereby covering extreme domain offset scenarios.

[0062] Hybrid sample generation: Frequency domain augmented samples and adversarial samples are weighted and fused at a ratio of 0.9:0.1. This ratio was determined through experiments and can take into account both the physical realism of frequency domain augmented samples and the generalization enhancement effect of adversarial samples, ultimately forming extended domain samples to expand the distribution range of source domain data.

[0063] Step 3: Implementation of multi-stream feature extraction and explicit decoupling.

[0064] Basic Feature Extraction: A CNN with shared weights is used as the basic feature extractor. Through two consecutive operations of "convolution-activation-normalization-pooling", the spatial dimension of the data is gradually compressed and the abstraction of the features is improved. The first operation realizes the initial feature extraction and dimensionality improvement, and the second operation further enhances the feature expression capability. The final output is a 128-dimensional shallow feature map that can reflect the basic attributes of the data, providing high-quality input for subsequent feature decoupling.

[0065] Domain-Specific Feature Extraction: Dedicated extraction networks with identical structures are designed for both HSI and LiDAR modalities, each employing a two-layer convolutional structure. The basic features of the original HSI samples and the extended domain samples are concatenated and input into the HSI domain-specific extractor to enhance the domain-specific information specific to the HSI modality. Similarly, the basic features of the LiDAR data are concatenated and input into the dedicated extractor to output LiDAR domain-specific features. The structures of the two extractors are kept consistent to ensure matching of output feature dimensions, preparing for subsequent fusion.

[0066] Domain-invariant feature extraction: A shared weight network composed of three convolutional layers is used to cascade and fuse the basic features of HSI original data, LiDAR original data, HSI extended samples, and LiDAR extended samples. Through multi-layer convolution, common features that are not affected by domain offset factors such as illumination, atmosphere, and sensor noise are gradually screened out, highlighting the inherent essential attributes of ground objects and providing stable feature support for cross-domain classification.

[0067] Dual Loss Constraints: Reconstruction Loss: A decoder network is constructed to reconstruct the original input data using extracted HSI domain-specific features, LiDAR domain-specific features, and domain-invariant features. The mean squared error is used to calculate the reconstruction error as the loss. The core purpose is to verify the completeness of the information of the three types of features through data reconstruction, ensuring that no key information is lost during the feature extraction process. Supervised Contrast Loss: A temperature parameter is set to adjust the similarity distribution in the feature space. By bringing the features of similar samples closer together and increasing the features of dissimilar samples further apart, the discriminative ability of the features is improved, allowing the model to more clearly distinguish different land cover categories.

[0068] Step 4: Implement modal attention fusion.

[0069] Attention network construction: First, global average pooling is performed on HSI domain-specific features, LiDAR domain-specific features, and domain-invariant features respectively to compress spatial dimension features into one-dimensional vectors and eliminate spatial redundancy information; then, the importance weights of different features in the classification task are learned through a two-layer fully connected network; finally, the weights are mapped to the [0,1] interval through an activation function to achieve adaptive focusing on high contribution features.

[0070] Feature weighted fusion: Based on the learned attention weights, the three types of features are linearly weighted and summed. The core logic is to allow the model to automatically adjust the influence of each feature in different scenarios. For example, in scenarios with severe lighting interference, the weights of LiDAR domain-specific features and domain-invariant features are automatically increased. In scenarios with clear spectral information, the role of HSI domain-specific features is preserved, achieving precise fusion that "maximizes strengths and minimizes weaknesses". The weighted feature map is flattened to obtain a one-dimensional fused feature vector for classification.

[0071] Step 5: Classification prediction and model optimization.

[0072] Classification head construction: The fused one-dimensional feature vector is mapped to the output dimension corresponding to the total number of land cover categories through a fully connected network. Then, the output is converted into the probability distribution of each category through the Softmax activation function, which intuitively reflects the model's prediction confidence for each land cover category.

[0073] Multi-loss joint optimization: The total loss function consists of classification loss, reconstruction loss, and supervised comparison loss. By setting reasonable weight coefficients, the influence of each loss is balanced. The classification loss ensures the classification accuracy of the model, the reconstruction loss ensures the completeness of features, and the supervised comparison loss improves the discriminative power of features. The Adam optimizer is used, and appropriate initial learning rate, weight decay coefficient, and batch size are set to gradually adjust the model parameters during training. An early stopping strategy is adopted, and training is stopped when the overall accuracy of the validation set does not improve for several consecutive iterations to avoid model overfitting and ensure the model's generalization ability in unknown target domains.

[0074] The cross-domain classification system of this invention comprises five core modules, which work together to achieve accurate cross-domain classification of multimodal remote sensing data. The specific modules and their functions are as follows:

[0075] Data preprocessing module: Used to acquire hyperspectral images (HSI) and LiDAR data of the target area, and eliminate differences in data units through normalization, slicing and other operations, retain spatial context information, and provide standardized input data for subsequent processing.

[0076] Hybrid Domain Extension Module: Includes a frequency domain amplitude perturbation submodule and an adversarial domain extension submodule. It generates extended domain samples through a dual strategy to simulate complex nonlinear domain offsets and expand the distribution range of source domain data.

[0077] Multi-stream feature extraction and decoupling module: Composed of a shared-weight CNN basic feature extractor, an HSI domain-specific extractor, a LiDAR domain-specific extractor, a domain-invariant extractor, and a dual-loss constraint unit, it realizes the explicit separation of basic feature encoding, domain-specific information, and domain-invariant information, while ensuring the completeness and discriminative power of the features.

[0078] Modal attention fusion module: includes a global average pooling layer and a fully connected attention network, used to adaptively generate attention weights for various features, realize dynamic weighted fusion of domain-specific features and domain-invariant features, and highlight the role of high-contribution features.

[0079] The classification prediction and optimization module consists of a classification head and a multi-loss joint optimization unit. It maps the fused features to the probability distribution of land cover categories and updates the model parameters iteratively through multi-loss collaborative optimization to improve the model's classification accuracy and cross-domain generalization ability.

[0080] The effectiveness of this invention can be further demonstrated through the following simulation.

[0081] I. Simulation Experiment Conditions.

[0082] 1. Hardware Platform The hardware configuration for the simulation experiment of this invention is as follows:

[0083] Processor: Intel Core i7-7820X CPU @ 3.60GHz;

[0084] Graphics card: NVIDIA GeForce RTX 2080 GPU (11GB VRAM);

[0085] Memory: 32GB DDR4 2666MHz;

[0086] Storage device: 1TB SSD solid-state drive;

[0087] 2. Software platform.

[0088] Operating system: Ubuntu 20.04 LTS;

[0089] Deep learning framework: TensorFlow 2.8.0 + Keras 2.8.0;

[0090] Programming language: Python 3.8.10;

[0091] Dependencies: NumPy 1.21.6, SciPy 1.7.3, Scikit-learn 1.0.2, OpenCV 4.5.5, GDAL 3.4.3;

[0092] Data processing tool: Matplotlib 3.5.3 (for result visualization).

[0093] 3. Experimental dataset.

[0094] The simulation experiments of this invention use three publicly available multimodal remote sensing datasets with typical domain migration characteristics, covering three core scenarios: cross-time series, cross-region, and cross-sensor. The dataset details are as follows:

[0095] Houston dataset (across time-series scenarios): The source domain (Houston 2013) contains 2530 samples, and the target domain (Houston 2018) contains 53200 samples, covering 7 land cover categories (healthy grassland, stressed grassland, trees, water bodies, residential areas, uninhabited areas, and roads). The HSI data has 48 bands and a spatial resolution of 2.5m.

[0096] The Na-Ha dataset (cross-regional scenario) contains 67,076 samples in the source domain (Nashua) and 38,457 samples in the target domain (Hanover), covering 7 land cover types (trees, grassland, water bodies, white roofs, gray roofs, roads, and bare soil). The HSI data has 114 bands and a spatial resolution of 1m.

[0097] The Trento-MUUFL dataset (across sensor / regional scenes) contains 10,487 samples in the source domain (Trento) and 36,623 samples in the target domain (MUUFL), with a total of 3 core land features (trees, buildings, and roads), 63 HSI bands, and a spatial resolution of 3m.

[0098] The source domain is used as the training set, while the target domain is used only as the test set and does not participate in the training. The data preprocessing method is as follows: noise bands are removed from HSI data and radiometric correction is performed. LiDAR data is uniformly converted to the Digital Surface Model (DSM) format. Both types of data are normalized to the [0,1] interval using Min-Max normalization. Finally, the data is slicing in 9×9 pixel size (step size 3 pixels) to preserve spatial context information.

[0099] Model training parameters: Batch size = 32, initial learning rate = 0.0001, weight decay coefficient = 1e-5, total number of iterations = 200 epochs, early stopping strategy (stop if the overall accuracy OA of the validation set does not improve after 15 consecutive epochs); loss function weights: λ1 (reconstruction loss) = 0.1, λ2 (supervised contrastive loss) = 0.05.

[0100] II. Simulation Content and Result Analysis.

[0101] The simulation experiment of this invention is a cross-domain classification performance comparison experiment (to verify the superiority of the overall scheme).

[0102] The simulation experiment of this invention uses the method of this invention and four existing technologies (D3Net, LLURnet, CLDA, HSI-CNN) to perform cross-domain classification tests on the target domains of three datasets.

[0103] Existing technology 1: The deep domain diversification network proposed in the paper "D3Net: Deep Domain Diversification Network for Cross-SceneHyperspectral Image Classification", abbreviated as "D3Net", uses GAN to generate diverse domain samples to improve generalization ability.

[0104] Existing technology 2: The paper "LLURnet: Low-Light Underwater Image Restoration with a Unified Network" proposes a remote sensing classification method adapted to a low-light image restoration network, abbreviated as "LLURnet", which focuses on feature enhancement and noise suppression.

[0105] Existing technology 3: The paper "CLDA: Cross-Domain Label Alignment for Hyperspectral ImageClassification" proposes a cross-domain label alignment method, abbreviated as "CLDA", which achieves domain adaptation through label mapping.

[0106] Existing technology 4: The single-modal HSI classification benchmark method uses CNN to extract spectral-spatial features without fusing LiDAR data, and is referred to as "HSI-CNN".

[0107] 2. Results and analysis of the simulation experiment.

[0108] To verify the effectiveness of simulation experiment 1 of this invention, cross-domain classification tests were conducted on the target domains of three datasets using the method of this invention and four existing technologies (D3Net, LLURnet, CLDA, and HSI-CNN). Standard evaluation metrics in the field of remote sensing classification, including overall accuracy (OA), average accuracy (AA), and Kappa coefficient, were used. The classification accuracy and OA, AA, and Kappa coefficients for each category were calculated based on the confusion matrix, as shown in Table 1.

[0109] The detailed explanations of the evaluation indicators (OA, AA, Kappa) are as follows:

[0110] ;

[0111] ;

[0112] ;

[0113] ;

[0114] Where C represents the total number of land cover categories. The overall classification accuracy (OA) is the percentage of correct classifications. This represents the expected accuracy for random classification.

[0115] Table 1. Overview of quantitative classification results for each method on the three datasets (unit: %)

[0116]

[0117] As shown in Table 1, the method of the present invention outperforms the four existing technologies in all three types of datasets in terms of OA, AA, and Kappa coefficients, and has significant advantages in the following aspects:

[0118] In the cross-temporal scenario (Houston): OA improves by 7.44 percentage points compared to the best existing technology 2 (D3Net) and by 18.12 percentage points compared to the single-modal benchmark (HSI-CNN).

[0119] Cross-regional scenario (Na-Ha): OA is improved by 4.57 percentage points compared with the best existing technology 5 (HSI-CNN), and the Kappa coefficient is improved by 4.94 percentage points.

[0120] Cross-sensor scenario (Trento-MUUFL): OA is improved by 5.90 percentage points compared to the best existing technology 2 (D3Net), and AA is improved by 6.42 percentage points.

[0121] The method of this invention significantly improves classification accuracy in categories with obvious "different objects with the same spectrum" characteristics (such as trees and grasslands, buildings and roads), with an average improvement of 8%-15%, proving that the multimodal fusion of hyperspectral and LiDAR and the feature decoupling mechanism effectively make up for the lack of discrimination ability of single modal information.

[0122] Existing technology 5. HSI-CNN achieves slightly higher OA on the Na-Ha dataset than some domain generalization methods, but its performance drops sharply in cross-temporal and cross-sensor scenarios, demonstrating the necessity of multimodal fusion for complex domain offset scenarios.

[0123] Simulation experiments show that:

[0124] This invention, through a collaborative design of hyperspectral and LiDAR multimodal fusion, hybrid domain extension, explicit feature decoupling, and modal attention fusion, achieves significantly better classification performance than existing mainstream domain generalization and remote sensing classification methods in three typical domain migration scenarios: cross-time series, cross-region, and cross-sensor. The overall accuracy is improved by 4.57%-7.44 percentage points compared with the best comparison method, effectively solving the problems of single modality, insufficient domain migration simulation, and weak generalization ability in existing technologies.

[0125] This invention requires no data (including unlabeled data) in the target domain. It can achieve efficient cross-domain generalization by optimizing source domain data alone. During deployment, there is no need to preprocess the target domain data or fine-tune the model. It has low application cost and high flexibility, and is suitable for practical remote sensing application scenarios such as agricultural yield estimation, environmental monitoring, land survey, and fine classification of vegetation.

Claims

1. A cross-domain classification method for HSI and ranging data based on domain extension and feature decoupling, characterized in that, The implementation steps of the classification method are as follows: Step 1: Preprocess the hyperspectral image HSI and LiDAR data of the target region to obtain source region samples; Step 2: A hybrid strategy combining frequency domain amplitude perturbation and adversarial domain extension is adopted to perform frequency domain amplitude perturbation on the preprocessed HSI and LiDAR data to generate frequency domain enhanced samples. Adversarial samples are generated based on the frequency domain enhanced samples. The frequency domain enhanced samples and adversarial samples are then weighted and fused at a preset ratio to obtain extended domain samples. Step 3: Use a convolutional neural network to extract the basic features from the source domain samples and the extended domain samples respectively; By using domain-specific extractors and domain-invariant extractors in convolutional neural networks, we can extract the domain-specific features corresponding to the source domain samples and the extended domain samples, as well as the domain-invariant features across modalities and domains. Step 4: Input the extracted domain-specific features and domain-invariant features into the attention network to generate adaptive attention weights. Then, perform weighted fusion of the HSI domain-specific features, LiDAR domain-specific features, and domain-invariant features to obtain a fused feature vector. Step 5: Input the fused feature vector into the classification head of the convolutional neural network to output the probability distribution of land cover categories; use a multi-loss joint optimization strategy to iteratively update the parameters of the cross-domain classification model to obtain the trained cross-domain classification model. Step 6: Input the HSI data and LiDAR data of the target domain to be classified, which have been preprocessed in the same way as in Step 1, into the trained cross-domain classification model and output the land cover classification results.

2. The cross-domain classification method according to claim 1, characterized in that, The steps in step 2 for generating frequency-enhanced samples by performing frequency domain amplitude perturbation on the preprocessed HSI and LiDAR data are as follows: The first step is to perform two-dimensional discrete cosine transform (DCT) on the preprocessed HSI slice samples and LiDAR slice samples respectively, so as to transform the data from the spatial domain to the frequency domain and obtain the frequency domain matrix. The second step involves using a Gaussian low-pass filter with a cutoff frequency of 0.3 to extract low-frequency components containing domain-specific information such as illumination and style from the frequency domain matrix, while retaining high-frequency components containing semantic information such as ground texture and structure. The third step is to multiply the amplitude spectrum of the extracted low-frequency components by a random scaling factor α, where α is uniformly sampled in the interval [0.7, 1.3] to keep the phase spectrum unchanged. Apply a weak perturbation of ±5% to the retained high-frequency components; The fourth step is to perform inverse two-dimensional discrete cosine transform (IDCT) on the frequency domain data after the above processing to reconstruct the data back into the spatial domain. The fifth step involves introducing a reconstruction coefficient of 0.85 to fuse the reconstructed data across different bands, generating frequency-enhanced samples.

3. The cross-domain classification method according to claim 1, characterized in that, The steps for generating adversarial examples in step 2 are as follows: The first step is to use the projected gradient descent algorithm as the adversarial example generation algorithm; The second step is to set the algorithm parameters: perturbation step size α = 0.01, number of iterations T = 10, and perturbation boundary [0,1]. The third step is to generate adversarial examples for the current iteration according to the following formula: ; in, This represents the adversarial example in the (t+1)th iteration. This represents the projection operation within the [0,1] perturbation boundary. Let represent the adversarial example in the t-th iteration. Represents a symbolic function. This represents the gradient used to calculate the loss function with respect to the input samples; Let f(x) represent the cross-entropy loss function, f(x) represent the feature extraction and classification mapping function of the cross-domain classification model in the current training phase, and y represent the true land cover category label of the sample. The cross-entropy loss function for: ;in, Let the i-th element in the one-hot encoded vector representing the true land cover category label be the element that determines whether the sample belongs to the i-th class. , Let C represent the probability value of the i-th type of land cover predicted by the cross-domain classification model, and let C represent the total number of land cover categories. This represents a logarithmic operation with base 10. Fourth step: Repeat the above steps and use the adversarial sample obtained after T iterations as the final generated adversarial sample.

4. The cross-domain classification method according to claim 1, characterized in that, The preset ratio mentioned in step 2 is determined in the following way: a pre-trained basic feature extraction network with the same structure as the cross-domain classification model is used to extract feature vectors of source domain data and a small amount of target domain reference data, and the maximum mean difference between the two is calculated; the weights ω1 of the frequency domain augmented samples and ω2 of the adversarial samples are set to satisfy ω1+ω2=1, and dynamically adjusted according to the maximum mean difference result: if the maximum mean difference result shows that the style difference is mainly concentrated in the low frequency spectral distribution, then ω1∈[0.6,0.8] and ω2∈[0.2,0.4] are set; if the maximum mean difference result shows that the target domain contains complex sensor noise or abnormal environment, then ω1∈[0.3,0.5] and ω2∈[0.5,0.7].

5. The cross-domain classification method according to claim 1, characterized in that, The convolutional neural network described in step 3 consists of four functionally differentiated feature extractor modules cascaded together: a shared-weight CNN basic feature extractor, an HSI domain-specific extractor, a LiDAR domain-specific extractor, and a domain-invariant extractor. Specifically, the shared-weight CNN basic feature extractor sequentially performs the following signal transformation operations on the preprocessed HSI and LiDAR samples, outputting a basic feature signal with a dimension of 2×2×128: ​​① Conv2d → ReLU activation → BatchNorm2d normalization; where the parameters of Conv2d are set to: 64 output channels, 3×3 kernel size, stride 1, and padding 1; ② MaxPool2d pooling, where the parameters of MaxPool2d are set to: 2×2 pooling kernel size, stride 2; ③ Conv2d → ReLU activation → BatchNorm2d normalization; where the parameters of Conv2d are set to: 128 output channels, 3×3 kernel size, stride 1, and padding 1; ④ MaxPool2d pooling; the parameters of MaxPool2d are set to: pooling kernel size 2×2, stride 2; the HSI domain-specific extractor concatenates and fuses the basic feature signals of HSI source domain samples and HSI extended domain samples output by the shared weight CNN basic feature extractor, and performs the following signal transformation operations in sequence to form an HSI domain-specific feature signal with an output dimension of 2×2×256: ① Conv2d→ReLU activation; where the parameters of Conv2d are set to: number of output channels 64, convolution kernel size 3×3, padding 1; ② Conv2d; the parameters of Conv2d are set to: number of output channels 32, convolution kernel size 1×1; the LiDAR domain-specific extractor concatenates and fuses the basic feature signals of LiDAR source domain samples and LiDAR extended domain samples output by the shared weight CNN basic feature extractor, and performs the following signal transformation operations in sequence to form a LiDAR domain-specific feature signal with a dimension of 2×2×256. The output signal of domain-specific features; adopts the same two-layer CNN network signal transformation structure as the HSI domain-specific extractor: ① Conv2d → ReLU activation; where the parameters of Conv2d are set as follows: 64 output channels, 3×3 kernel size, and 1 padding; ② Conv2d; the parameters of Conv2d are set as follows: 32 output channels and 1×1 kernel size;The domain-invariant extractor concatenates and fuses the basic feature signals of four types of samples—HSI source domain, LiDAR source domain, HSI extended domain, and LiDAR extended domain—output by the shared-weight CNN basic feature extractor to form a cross-modal, cross-domain domain-invariant feature output signal with a dimension of 2×2×512. Signal processing employs a three-layer shared-weight CNN network to perform feature signal transformation, with the following steps: ① Conv2d → ReLU activation; where Conv2d parameters are set to: 128 output channels, 3×3 kernel size, and 1 padding; ② Conv2d → ReLU activation; where Conv2d parameters are set to: 64 output channels, 3×3 kernel size, and 1 padding; ③ Conv2d; where Conv2d parameters are set to: 64 output channels, 1×1 kernel size.

6. The cross-domain classification method according to claim 1, characterized in that, The steps for extracting domain-specific features for each modality in step 3 are as follows: First, using a shared-weight CNN basic feature extractor, feature encoding is performed on the original HSI samples and LiDAR samples in the source domain, as well as the HSI extended domain samples and LiDAR extended domain samples in the extended domain, outputting a 128-dimensional shallow basic feature map with a dimension of 2×2×128; Second, the encoded features of the original HSI samples and the encoded features of the extended HSI samples are concatenated to obtain concatenated features with a dimension of 2×2×256, which serve as the input to the HSI domain-specific extractor; Third, the encoded features of the original LiDAR samples and the encoded features of the extended LiDAR samples are concatenated to obtain concatenated features with a dimension of 2×2×256, which serve as the input to the LiDAR domain-specific extractor; Fourth, the input features of the HSI domain-specific extractor are fed into a pre-defined two-layer CNN network, and the network calculates and outputs the HSI domain-specific features. The fifth step involves feeding the input features of the LiDAR domain-specific extractor into a two-layer CNN network with the same structure as the HSI domain-specific extractor, and then using the network to calculate and output LiDAR domain-specific features.

7. The cross-domain classification method according to claim 1, characterized in that, The cross-modal and cross-domain domain-invariant features mentioned in step 3 are obtained in the following way: the feature signals of the four types of samples output by the shared weight CNN basic feature extractor in the HSI source domain, LiDAR source domain, HSI extended domain and LiDAR extended domain are concatenated and fused, and the feature signals are obtained after feature transformation is performed by a three-layer shared weight CNN network. After being processed by the aforementioned network, the characteristic signal is unaffected by domain offset factors such as light intensity, atmospheric transmission characteristics, solar elevation angle, sensor noise level, and spectral response function, and can stably characterize the inherent essential properties of ground objects.

8. The cross-domain classification method according to claim 1, characterized in that, The steps for iteratively updating the parameters of the cross-domain classification model using the multi-loss joint optimization strategy described in step 5 are as follows: The first step is to use the Adam optimizer, setting the initial learning rate to 0.0001, the weight decay factor to 1e-5, and the batch size to 32; The second step involves inputting the fused feature vectors of the source domain samples and the extended domain samples into the cross-domain classification model, iteratively updating the model parameters until the model's total loss function converges. The convergence condition is that the fluctuation range of the total loss function value is less than 1e-4 within 10 consecutive iterations, resulting in a trained cross-domain classification model. The total loss function is as follows: ;in, Represents the cross-entropy classification loss; The domain difference loss is represented by the maximum mean difference loss, which is used to calculate the difference in feature distributions between the source domain and the extended domain. The feature consistency loss is represented by the mean squared error loss, which constrains the semantic consistency between domain-specific features and domain-invariant features. and The loss balance coefficients are set to 0.5 and 0.3 respectively.

9. The cross-domain classification method according to claim 8, characterized in that, The cross-domain classification model consists of six cascaded functional modules: a shared-weight CNN basic feature extractor, an HSI domain-specific extractor, a LiDAR domain-specific extractor, a domain-invariant extractor, an attention network module, and a classification head. The structure and parameters of the four functional modules—shared weight CNN basic feature extractor, HSI domain-specific extractor, LiDAR domain-specific extractor, and domain-invariant extractor—are the same as the structure and parameters of the convolutional neural network described in claim 5. The attention network module consists of a cascaded global average pooling layer and a fully connected layer. The input of the global average pooling layer is a concatenated feature of HSI domain-specific features, LiDAR domain-specific features, and domain-invariant features, and the output is a dimension of 1×1×(32+32+64)=1×1×128. The fully connected layer has an input dimension of 128 and an output dimension of 3 → a Sigmoid activation function, with the outputs corresponding to the three attention weights of HSI domain-specific features, LiDAR domain-specific features, and domain-invariant features, respectively. The classification head consists of a fully connected layer and a cascaded Softmax activation function layer; the input dimension of the fully connected layer is 32 + 32 + 64 = 128, and the output dimension is equal to the total number of land cover categories C.

10. A cross-domain classification system for HSI and ranging data based on domain extension and feature decoupling, characterized in that, This is implemented based on the hyperspectral image classification method according to any one of claims 1-9; Includes the following modules: (wherein,) Data preprocessing module: used to preprocess the hyperspectral image HSI and LiDAR data of the target region to obtain source domain samples; Hybrid Domain Extension Module: Includes a frequency domain amplitude perturbation submodule and an adversarial domain extension submodule. The frequency domain amplitude perturbation submodule is used to perform frequency domain amplitude perturbation processing on the preprocessed HSI and LiDAR data to generate frequency domain enhanced samples. The adversarial domain extension submodule is used to generate adversarial samples based on the frequency domain enhanced samples. The frequency domain enhanced samples and adversarial samples are weighted and fused to obtain extended domain samples. Multi-stream feature extraction and decoupling module: Composed of a shared-weight CNN basic feature extractor, an HSI domain-specific extractor, a LiDAR domain-specific extractor, a domain-invariant extractor, and a dual-loss constraint unit, it realizes the explicit separation of basic feature encoding, domain-specific information, and domain-invariant information, while ensuring the completeness and discriminative power of the features; Modal attention fusion module: includes a global average pooling layer and a fully connected attention network, used to adaptively generate attention weights for various features, realize dynamic weighted fusion of domain-specific features and domain-invariant features, and highlight the role of high-contribution features; The classification prediction and optimization module consists of a classification head and a multi-loss joint optimization unit. It maps the fused features to the probability distribution of land cover categories and updates the model parameters iteratively through multi-loss collaborative optimization to improve the model's classification accuracy and cross-domain generalization ability.

Citation Information

Patent Citations

  • Multispectral remote sensing image panchromatic sharpening method based on frequency domain decomposition network

    CN119579847A

  • Hyperspectral image and LiDAR data collaborative classification system based on combination of frequency domain feature learning and CNN

    CN120279330A

Cited By

  • A method for classifying features in a visual image

    CN122493148A