Remote sensing image segmentation method based on hyperspectral coring and spatial-temporal feature fusion
Through hyperspectral nucleation and spatiotemporal feature fusion technology, the problem of insufficient spectral and spatial features in remote sensing image segmentation is solved, efficient multimodal data fusion is achieved, and the accuracy and efficiency of remote sensing image segmentation is improved.
Patent Information
- Application Number
- CN202510587319.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The existing remote sensing image segmentation method cannot fully utilize the complementary characteristics of hyperspectral images and synthetic aperture radar images, resulting in insufficient coordination of spectral-space-texture multi-dimensional features and low segmentation accuracy.
The method of fusion of hyperspectral nucleation and spatiotemporal features is adopted to improve the spectral feature utilization of hyperspectral images through spectral correlation coefficient calculation and radial basis function kernel transformation. Combined with multi-scale convolution and spatiotemporal feature fusion technology, the spatial information of SAR images is fully extracted to achieve effective fusion of multimodal data.
It improves the accuracy and computing efficiency of remote sensing image segmentation, especially in complex scenarios, which can more accurately segment different objects, which are suitable for large-scale remote sensing data processing.
Smart Images

Figure CN120495663A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and more particularly to a remote sensing image segmentation method based on hyperspectral kernelization and spatiotemporal feature fusion. Background Art
[0002] With the rapid development of remote sensing technology, it has been widely used in fields such as environmental monitoring, land use, and military reconnaissance. Hyperspectral remote sensing imagery and synthetic aperture radar (SAR) imagery, as important data sources for surface observation, each possess unique characteristics. Hyperspectral imagery contains rich spectral information and can accurately distinguish different substances, but its spatial resolution is low and it is susceptible to atmospheric interference. SAR imagery can obtain surface information under all-weather and all-time conditions and has strong penetration capabilities. However, because it primarily relies on backscattering characteristics, the extraction of texture and edge features is relatively complex.
[0003] However, current remote sensing image segmentation methods either rely on traditional segmentation algorithms based on a single data source or employ simple feature-level or decision-level fusion methods. Single hyperspectral image processing methods (such as spectral angle matching and support vector machines) are insufficiently expressive of spatial features, and SAR image segmentation methods (such as region growing and level set processing) are susceptible to interference from speckle noise. Existing fusion methods, such as band stacking or probability map fusion, fail to fully account for the complementary mechanism between the continuous spectrum characteristics of hyperspectral images and the geometric scattering properties of SAR, resulting in insufficient synergy between spectral, spatial, and texture multidimensional features.
[0004] Therefore, how to effectively integrate the advantages of hyperspectral images and SAR images to achieve more accurate remote sensing image segmentation is an important research direction in the current field of remote sensing image processing. Summary of the Invention
[0005] In view of this, the present invention provides a remote sensing image segmentation method based on hyperspectral kernelization and spatiotemporal feature fusion, aiming to make full use of the spectral and spatial information of multimodal remote sensing data and improve the segmentation accuracy in complex scenes.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] A remote sensing image segmentation method based on hyperspectral kernelization and spatiotemporal feature fusion, comprising:
[0008] The initial hyperspectral image features are obtained, converted into a sequence, and the spectral correlation coefficient is calculated and determined. The spectral correlation coefficient is converted into the form of a radial basis function kernel and then deserialized to obtain the hyperspectral kernelization feature. The hyperspectral kernelization feature is added to the initial hyperspectral feature to obtain the hyperspectral image feature.
[0009] Perform multi-scale feature extraction on synthetic aperture radar images and obtain radar image features after splicing and fusion;
[0010] The spatiotemporal feature fusion of hyperspectral image features and radar image features is performed to obtain a segmented image.
[0011] In an optional embodiment, calculating the spectral correlation coefficient includes:
[0012] The hyperspectral image features converted into sequences are projected into three different spaces: query, key, and value, and the spectral correlation coefficient is determined according to the following formula;
[0013]
[0014] Where q and k represent the query vector and key vector of a spectrum in the hyperspectral image feature sequence, respectively. and They represent the query vector mean and key value vector mean of all spectra in the hyperspectral image feature sequence, respectively.
[0015] In an optional embodiment, after the spectral correlation coefficient is converted into a radial basis function kernel form, it is first multiplied with the value vector and then the deserial operation is performed.
[0016] In an optional embodiment, the spectral correlation coefficient is converted into the form of a radial basis function kernel according to the following formula:
[0017]
[0018] Where σ is an adjustable temperature parameter used to control the smoothness of the kernel function, S(q, k) represents the spectral correlation coefficient, and K S Represents the kernel function, which is the inner product of two feature maps.
[0019] In an optional embodiment, multi-layer hyperspectral kernelization feature extraction is performed on the initial hyperspectral image features, and the features are added together with the initial hyperspectral image features through skip connections to obtain hyperspectral image features.
[0020] In an optional embodiment, after performing multi-scale feature extraction, splicing and fusion on the synthetic aperture radar image, the radar image features are obtained by sequentially passing through a convolution layer, a flattening layer, a fully connected layer and an activation function layer.
[0021] In an optional embodiment, performing spatiotemporal feature fusion on the hyperspectral image features and the radar image features includes:
[0022] Perform average pooling and maximum pooling on hyperspectral image features and radar image features respectively;
[0023] Obtain channel weights based on average pooling features, and obtain spatial weights based on maximum pooling features;
[0024] The channel weights and spatial weights corresponding to the hyperspectral image features and radar image features are summed, and the hyperspectral image features and radar image features are weighted fused respectively.
[0025] In an optional embodiment, a composite function of a weighted sum of cross entropy loss, Dice loss, and Focal loss is used as the final loss function.
[0026] The present invention discloses a remote sensing image segmentation method based on hyperspectral kernelization and spatiotemporal feature fusion. Compared with the existing technology, it not only gives full play to the spectral advantages of hyperspectral images and the spatial advantages of SAR images, but also effectively solves the problem of multimodal data fusion of remote sensing images through hyperspectral kernelization and spatiotemporal feature fusion technology.
[0027] Experimental results show that this method outperforms existing methods on multiple remote sensing datasets and can achieve more accurate remote sensing image segmentation in complex terrain environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0029] Figure 1 A schematic diagram of the structure of the hyperspectral nucleation module provided by the present invention;
[0030] Figure 2 A schematic diagram of the multi-scale convolution module structure provided by the present invention;
[0031] Figure 3 This is a schematic diagram of the structure of the spatiotemporal feature fusion module provided by the present invention. DETAILED DESCRIPTION
[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0033] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0034] In this embodiment, a remote sensing image segmentation method based on hyperspectral kernelization and spatiotemporal feature fusion is disclosed to exploit the multimodal characteristics of remote sensing images. This method is used to fully extract the spatial and spectral information of the image and improve segmentation accuracy in complex scenes. The method mainly includes three parts: hyperspectral kernelization and multi-scale convolution feature extraction, and spatiotemporal feature fusion. The steps are as follows:
[0035] The initial hyperspectral image features are obtained, converted into a sequence, and the spectral correlation coefficient is calculated and determined. The spectral correlation coefficient is converted into the form of a radial basis function kernel and deserialized to obtain a hyperspectral kernel feature. The hyperspectral kernel feature is added to the initial hyperspectral feature to obtain the hyperspectral image feature.
[0036] Perform multi-scale feature extraction on synthetic aperture radar images and obtain radar image features after splicing and fusion;
[0037] The spatiotemporal feature fusion of hyperspectral image features and radar image features is performed to obtain a segmented image.
[0038] In this embodiment:
[0039] First, hyperspectral images usually contain rich spectral information, which can provide important prior knowledge for target detection and classification. In order to more effectively utilize the spectral correlation of hyperspectral data, this paper proposes a hyperspectral kernelization scheme.
[0040] In this embodiment, the hyperspectral image is first serialized to convert the three-dimensional tensor into a one-dimensional sequence, and the spectral correlation is calculated using a self-attention mechanism. Unlike the traditional dot product similarity calculation, the present invention introduces a spectral correlation coefficient to enhance the ability to utilize information from different bands. Subsequently, the spectral correlation coefficient is converted into the form of a radial basis function (RBF) kernel using a kernelization technique, thereby improving computational efficiency and reducing computational complexity. Finally, the characteristic structure is restored through a deserialization operation, so that the information of the hyperspectral image is fully integrated in the spatial and spectral dimensions.
[0041] In one embodiment, the present application constructs a hyperspectral kernelization module to improve the ability to utilize information from different bands, enhance the spatial-spectral fusion effect, and improve the accuracy of remote sensing image segmentation.
[0042] Its structure refers to Figure 1Specifically, spectral correlation calculation, kernelization transformation and serialization / deserialization operations are used to extract the key features of hyperspectral images, and effective fusion of cross-band information is achieved through a multi-level structure.
[0043] In an exemplary embodiment, the steps are as follows:
[0044] 1.1 Initial hyperspectral feature extraction
[0045] First, the hyperspectral image is input into a 3×3 convolutional layer for preliminary feature extraction to obtain the initial hyperspectral feature map. Where H and W represent the height and width of the image, respectively, and C is the number of spectral channels. The main purpose of this step is to extract low-level spatial and spectral features through convolution operations, laying the foundation for subsequent deep feature processing.
[0046] 1.2 Constructing a hyperspectral nucleation module
[0047] 1) Feature serialization: First, the upsampling layer is used to process the image to improve its feature resolution. Then, through the serialization operation, the three-dimensional tensor Convert to a one-dimensional sequence where N HW =H×W is the length of the sequence, and C is the number of spectral channels. This step flattens the spectral information of each pixel into a one-dimensional sequence, allowing the subsequent attention mechanism to operate in a one-dimensional feature space;
[0048] 2) Self-attention mechanism and spectral correlation:
[0049] Next, similar to the self-attention mechanism in Transformer, the input features Projected into three different spaces: query (Q), key (K), and value (V), the following is obtained through linear transformation:
[0050]
[0051] Among them, W Q , W K , W V is a learnable weight matrix that maps input features to query, key, and value spaces;
[0052] 3) Calculation of spectral correlation coefficient:
[0053] The traditional self-attention mechanism calculates the similarity between the query and the key through the dot product (such as cosine similarity), while this paper introduces the spectral correlation coefficient to replace the traditional similarity metric. The calculation formula of the spectral correlation coefficient is:
[0054]
[0055] Where q and k represent the query vector and key vector of a spectrum in the hyperspectral image feature sequence, respectively. and They represent the query vector mean and key value vector mean of all spectra in the hyperspectral image feature sequence, respectively.
[0056] The square of the spectral correlation coefficient ensures non-negativity and effectively reflects the correlation between spectral curves. This method is insensitive to amplitude changes in spectral curves (such as shadows or occlusions), and can therefore provide prior knowledge of spectral information, enhancing the model's adaptability to remote sensing data.
[0057] 4) Nuclearization
[0058] In order to further reduce the computational complexity, the kernelization technique is used to process the spectral correlation coefficient. Specifically, the SCC is converted into a radial basis function (RBF) kernel:
[0059]
[0060] Where σ is an adjustable temperature parameter used to control the smoothness of the kernel function, S(q, k) represents the spectral correlation coefficient, and K S Represents the kernel function, which is the inner product of two feature maps and is expressed as:
[0061] K S (q,k)=<ψ(q),ψ(k)〉
[0062] Among them, ψ(·) is the characteristic mapping function, which can be approximately calculated by Taylor expansion.
[0063] After the spectral correlation coefficient is converted into the form of radial basis function kernel, it is multiplied with the value vector and then the deserial operation is performed.
[0064] 5) Deserialization operation
[0065] After the hyperspectral kernelization module is processed, the output feature is still a sequence In order to restore it to image format, a deserialization operation is required, which is to rearrange the sequence into the shape of the original image so that it can be used for subsequent processing steps.
[0066] In order to further optimize the above technical solution, the present invention performs multi-layer hyperspectral kernelization feature extraction on the initial hyperspectral image features, that is, further extracts features by applying N hyperspectral kernelization modules. The extracted features are connected to the initial hyperspectral features through skip connections. The fused features are then further processed using a 3×3 convolutional layer to obtain the feature representation φ of the hyperspectral image. HSI , providing rich multimodal features for subsequent remote sensing image segmentation tasks.
[0067] Through this process, the present invention effectively extracts spatial and spectral information from hyperspectral images, providing powerful feature support for remote sensing image segmentation.
[0068] Part II:
[0069] For SAR images, since they lack explicit spectral information and mainly rely on backscattering characteristics for target recognition, this embodiment uses a multi-scale convolution module for feature extraction. It also improves the expressiveness of SAR images through feature splicing and convolution fusion, enabling it to enhance the edge, texture and structural features of SAR images in multi-scale space, making it more suitable for remote sensing image segmentation tasks.
[0070] In one embodiment, in the feature extraction process of synthetic aperture radar (SAR) images, in order to fully capture information at different scales, a multi-scale convolution module is constructed, the structure of which is as follows: Figure 2 As shown in Figure 2, by combining multiple convolution kernel sizes, the expressive power of features is enhanced.
[0071] In an exemplary embodiment, the feature extraction process is as follows:
[0072] 2.1 Multi-scale Convolution Extraction: This paper uses three different convolution kernel sizes (3×3, 5×5, and 7×7) to perform parallel feature extraction on SAR images to capture spatial information at different scales. Self.scale1, self.scale2, and self.scale3 correspond to three different convolution kernel sizes, respectively, and are used to extract local texture, edge information, and larger-scale structural features.
[0073] 2.2 Multi-scale feature splicing: For the features extracted by different convolution kernels mentioned above, this paper uses feature splicing (torch.cat) to splice the feature maps of each scale along the channel dimension to form a fused feature representation with multi-scale information, so as to enhance the model's perception of the target area;
[0074] 2.3 Channel Fusion and Dimensionality Reduction: To further fuse features of different scales and reduce the channel dimension to reduce the amount of computation, this paper introduces a 3×3 convolutional layer (self.conv1x1) to perform channel transformation on the concatenated features and uses a flattening operation to enhance key information while reducing redundant features.
[0075] 2.4 Full Connection and Activation: After convolution feature extraction and fusion, the present invention further introduces a fully connected layer and activation function to enhance the nonlinear expression capability of features and improve the discriminative ability of the model. Finally, the feature distribution is further optimized through 1×1 convolution to provide high-quality feature representation for subsequent segmentation tasks. The radar image feature φ is finally obtained. SAR .
[0076] Part III:
[0077] After extracting the features of hyperspectral and SAR images, this embodiment further constructs a spatiotemporal feature fusion module to fully utilize the spatiotemporal relationships of these data and improve the performance of remote sensing image segmentation. This module first utilizes a channel attention mechanism to calculate the importance of different channels through global pooling and one-dimensional convolution, and then uses a Softmax function for normalization to dynamically adjust the weight of each channel. Furthermore, a spatial attention mechanism is used to perform global pooling and two-dimensional convolution on the feature maps to calculate the weights in the spatial dimension. Softmax normalization is also used to ensure the rationality of feature fusion.
[0078] Finally, through a multi-level fusion strategy, remote sensing image information from different time points and sources is fully combined to improve the accuracy of the segmentation results.
[0079] In one embodiment, the fusion process refers to Figure 3 ;φ HSI and φ SAR These feature maps represent hyperspectral image features and radar image features, respectively. These feature maps contain remote sensing image information from different time points or sources. The spatiotemporal feature fusion module can fully utilize the spatiotemporal relationship between these two different features to improve remote sensing image segmentation, especially when dealing with complex land feature classifications.
[0080] In an exemplary embodiment, the specific process includes:
[0081] 3.1 Building a channel attention mechanism
[0082] 3.11 Global Pooling: Input Feature Map φ HSI and φ SAR Perform global average pooling (Avg Pooling) to aggregate information in the channel dimension. The specific operations are:
[0083] φ C =Concat(Avg(φ HSI ),Avg(φ SAR ))
[0084] Among them, φ CRepresents the channel features obtained by pooling operation, and Concat represents splicing in the channel dimension.
[0085] 3.12 One-dimensional convolution: Next, use a one-dimensional convolution layer (to calculate φ HSI and φ SAR Importance in the channel dimension. These convolutional layers will be used to learn weights in the channel dimension, thereby dynamically adjusting the contribution of each channel to the final feature:
[0086] W c1 =Conv1(φ C ),W c2 =Conv2(φ C )
[0087] Among them, Conv represents the convolution layer. These convolution operations can help the model adaptively adjust the weights in subsequent steps by learning the contributions of different channels;
[0088] 3.13Softmax Normalization: Use the Softmax function to normalize the channel weights to ensure that the sum of the weights is 1:
[0089]
[0090] Among them, e is the base of the exponential function, and the normalized weights help the model focus on more important channels during the fusion process.
[0091] 3.2 Building a Spatial Attention Mechanism
[0092] 3.21 Spatial Pooling: Perform global pooling operations in the spatial dimension to calculate the feature map φ HSI and φ SAR Importance of space:
[0093] φ S =Concat(Max(φ HSI ),Max(φ SAR ))
[0094] φ S It can capture important information from different source images in spatial dimensions and prepare for subsequent spatial weighting.
[0095] 3.22 2D Convolution: Next, a 2D convolution layer is used to calculate the weights in the spatial dimension. The purpose of this operation is to further strengthen the fusion of spatial information and improve the consistency of feature maps from different sources in the spatial dimension:
[0096] W s1 =Conv2D1(φ S ),W s2 =Conv2D2(φS )
[0097] This operation is used to capture the weight relationship in the spatial dimension and further enhance the fusion of spatial information.
[0098] 3.23Softmax Normalization: Normalize spatial weights:
[0099]
[0100] Through this normalization, the model can effectively focus on more important areas in the spatial dimension.
[0101] 3.3 Feature Fusion. Fusion of bi-temporal weights, weighted summation of channel weights and spatial weights, to obtain the final feature weights:
[0102] W′ f1 =W′ c1 +W′ s1 , W′ f2 =W′ c2 +W′ s2
[0103] Use weighted summation to transform the feature map φ HSI and φ SAR After fusion, we finally get the fused feature map:
[0104] φ + =(W′ f1 ×φ HSI )+(W′ f2 ×φ SAR )
[0105] φ + The fused feature map contains the feature maps from the two input maps φ HSI and φ SAR The important information is ready to be passed to the downstream task module for further processing.
[0106] Further, the feature map is converted into the segmentation result, i.e. φ + The rich information from different sensors and time points is passed to a convolutional layer to generate the final segmentation result. The core purpose of this step is to convert the fused multimodal features into a final pixel-level segmentation map through the convolution operation, thereby achieving accurate segmentation of different ground object categories in remote sensing images.
[0107] In order to ensure that the segmentation result can output a probability value at each pixel position, this application uses the Sigmoid activation function to process the convolution layer output in this step. The Sigmoid function can map the output of each pixel to between 0 and 1, indicating the probability that the pixel belongs to a certain class. In this way, the final segmentation result will be a probability map, where each pixel corresponds to a probability value of a class.
[0108] In an optional embodiment, a cross-entropy loss function is used for training. The cross-entropy loss function can effectively measure the difference between the probability distribution of the model output and the actual label, thereby guiding the model to optimize parameters. Specifically, the loss function is as follows:
[0109]
[0110] in, is the predicted probability value, y i is the true label, N y is the total number of pixels in the image. By minimizing this loss function, the model can continuously adjust parameters to improve segmentation accuracy.
[0111] Dice loss is used to measure the overlap between the model prediction results and the true label, which is especially effective for unbalanced categories. The formula is expressed as:
[0112]
[0113] This loss can give greater weight to samples that are difficult to classify and reduce the influence of simple samples.
[0114] further;
[0115]
[0116] in, is the predicted probability of the target class, α is the balancing factor, and γ is the focusing parameter.
[0117] In this embodiment, the above loss functions are weighted and summed to obtain a composite loss function:
[0118]
[0119] Among them, λ1, λ2 and λ3 are weight coefficients that can be adjusted according to the needs of the task.
[0120] In this example, after calculating the loss using the cross-entropy loss function, the model parameters are updated using the backpropagation algorithm. During the optimization process, the Adam optimizer is used to accelerate convergence, and the learning rate is adjusted to avoid overfitting. As training progresses, the model continuously adjusts its weights to minimize the loss function, ultimately achieving a model capable of accurately performing remote sensing image segmentation. Through these steps, the final segmentation result will be able to accurately separate different ground object categories from remote sensing images and provide high-quality segmented images for downstream tasks.
[0121] Compared with the prior art, the present invention has the following beneficial effects:
[0122] 1) The hyperspectral kernelization process improves the efficiency of spectral information utilization. Through spectral correlation calculation, self-attention mechanism, and kernelization transformation, the cross-band information fusion capability of hyperspectral imagery is enhanced. This allows the model to utilize spectral prior knowledge while reducing computational complexity and improving segmentation accuracy.
[0123] 2) The application of multi-scale convolution enhances the feature representation capability of SAR images. Convolution kernels of different scales are used to extract edge, texture, and structural features of SAR images in parallel. Feature concatenation and channel fusion are then used to optimize the representation capability, making it suitable for remote sensing scenes with different resolutions and complex backgrounds.
[0124] 3) Spatiotemporal feature fusion enhances the collaborative analysis capabilities of multi-temporal and multi-sensor data. Combining channel-attention and spatial-attention mechanisms, this module adaptively adjusts the feature weights of remote sensing data from different sources, enhancing key information and reducing redundancy, thereby improving target segmentation in complex environments.
[0125] 4) High computational efficiency. This application is suitable for large-scale remote sensing data processing. By reducing the complexity of attention calculations through kernelization technology and optimizing computational overhead by combining lightweight one-dimensional / two-dimensional convolution operations, this invention can efficiently process large-scale remote sensing data and meet practical application needs.
[0126] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0127] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A remote sensing image segmentation method based on hyperspectral kernelization and spatiotemporal feature fusion, characterized in that: The initial hyperspectral image features are obtained, converted into a sequence, and the spectral correlation coefficient is calculated and determined. The spectral correlation coefficient is converted into the form of a radial basis function kernel and deserialized to obtain a hyperspectral kernel feature. The hyperspectral kernel feature is added to the initial hyperspectral feature to obtain the hyperspectral image feature. Perform multi-scale feature extraction on synthetic aperture radar images and obtain radar image features after splicing and fusion; The spatiotemporal feature fusion of hyperspectral image features and radar image features is performed to obtain a segmented image.
2. The remote sensing image segmentation method according to claim 1, wherein: Calculate spectral correlation coefficients, including: The hyperspectral image features converted into sequences are projected into three different spaces: query, key, and value, and the spectral correlation coefficient is determined according to the following formula; Where q and k represent the query vector and key vector of a spectrum in the hyperspectral image feature sequence, respectively. and They represent the query vector mean and key value vector mean of all spectra in the hyperspectral image feature sequence, respectively.
3. The remote sensing image segmentation method according to claim 2, characterized in that: After the spectral correlation coefficient is converted into the form of radial basis function kernel, it is multiplied with the value vector and then the deserial operation is performed.
4. The remote sensing image segmentation method according to claim 1, wherein: The spectral correlation coefficient is converted into the form of radial basis function kernel according to the following formula: Where σ is an adjustable temperature parameter used to control the smoothness of the kernel function, S(q, k) represents the spectral correlation coefficient, and K S Represents the kernel function, which is the inner product of two feature maps.
5. The remote sensing image segmentation method according to claim 1, wherein: The initial hyperspectral image features are subjected to multi-layer hyperspectral kernelization feature extraction and are summed with the initial hyperspectral image features through jump connections to obtain the hyperspectral image features.
6. The remote sensing image segmentation method according to claim 1, wherein: After multi-scale feature extraction and splicing of synthetic aperture radar images, the radar image features are obtained through the convolution layer, flattening layer, fully connected layer and activation function layer in sequence.
7. The remote sensing image segmentation method according to claim 1, wherein: Perform spatiotemporal feature fusion of hyperspectral image features and radar image features, including: Perform average pooling and maximum pooling on hyperspectral image features and radar image features respectively; Obtain channel weights based on average pooling features, and obtain spatial weights based on maximum pooling features; The channel weights and spatial weights corresponding to the hyperspectral image features and radar image features are summed, and the hyperspectral image features and radar image features are weighted fused respectively.
8. The remote sensing image segmentation method according to claim 1, wherein: The final loss function is a composite function of the weighted sum of cross entropy loss, Dice loss and Focal loss.
Citation Information
Patent Citations
Hyperspectral-remote-sensing-image classification technology combining spectral, spatial and hierarchical information
CN108427913A
Rapid soil classification method based on visible near infrared spectrum and multi-target fusion
CN108827909A
Hyperspectral image space spectrum classification method and device considering spectral importance
CN110796163A
Hyperspectral remote sensing image classification method and related equipment
CN117422917A
Physical guidance generative confrontation hyperspectral super-resolution method based on non-registration
CN118212127A