Remote sensing image segmentation method based on hyperspectral kernelization and spatio-temporal feature fusion

By using hyperspectral kernelization and spatiotemporal feature fusion techniques, the problem of insufficient multimodal data fusion in remote sensing image segmentation is solved, improving the accuracy and efficiency of remote sensing image segmentation, especially enabling more accurate segmentation in complex terrain environments.

CN120495663BActive Publication Date: 2025-11-07耕宇牧星(北京)空间科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510587319.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-11-07
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

Existing remote sensing image segmentation methods cannot effectively integrate the advantages of hyperspectral imagery and synthetic aperture radar imagery, resulting in insufficient synergy of spectral-spatial-texture multidimensional features and low segmentation accuracy.

Method used

A method combining hyperspectral kernelization and spatiotemporal feature fusion is adopted. Hyperspectral image features are processed by calculating spectral correlation coefficients and radial basis function kernelization. Combined with multi-scale convolution and spatiotemporal feature fusion techniques, the feature representation capability of SAR images is improved, and multimodal data fusion of hyperspectral and SAR images is achieved.

Benefits of technology

It improves the accuracy and efficiency of remote sensing image segmentation, especially enabling more precise segmentation in complex terrain environments, and is suitable for large-scale remote sensing data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495663B_ABST
    Figure CN120495663B_ABST
Patent Text Reader

Abstract

The application discloses a remote sensing image segmentation method based on hyperspectral kernelization and space-time feature fusion, and belongs to the field of remote sensing image processing. The steps comprise the following: firstly, initial hyperspectral image features are acquired, the spectral correlation coefficient is calculated after being converted into a sequence, the spectral correlation coefficient is converted into the form of a radial basis function kernel and is subjected to inverse sequence operation to obtain hyperspectral kernelization features; the hyperspectral kernelization features and the initial hyperspectral features are added to obtain hyperspectral image features; then, multi-scale feature extraction is carried out on synthetic aperture radar image to obtain radar image features; finally, space-time feature fusion is carried out on the hyperspectral image features and the radar image features to obtain a segmentation image. The application can fully utilize the spectral and spatial information of multi-modal remote sensing data and improve the segmentation precision in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of remote sensing image processing, and more particularly to a remote sensing image segmentation method based on hyperspectral kernelization and spatio-temporal feature fusion. BACKGROUND

[0002] With the rapid development of remote sensing technology, it has been widely used in environmental monitoring, land use, military reconnaissance and other fields. As important data sources for ground observation, hyperspectral remote sensing images and synthetic aperture radar (SAR) images have their own unique characteristics. Hyperspectral images contain rich spectral information and can accurately distinguish different substances, but their spatial resolution is low and they are easily affected by atmospheric interference. SAR images can obtain ground information under all-weather and all-time conditions and have strong penetration ability, but due to their dependence on backscattering characteristics, the extraction of texture and edge features is more complex.

[0003] However, current remote sensing image segmentation methods either use traditional segmentation algorithms with a single data source or use simple feature-level or decision-level fusion methods. Single hyperspectral image processing methods (such as spectral angle matching and support vector machines) have insufficient spatial feature expression ability, and SAR image segmentation methods (such as region growing and level sets) are affected by speckle noise. Existing fusion methods, such as band stacking or probability map fusion, cannot fully consider the complementary mechanism between the continuous spectral characteristics of hyperspectral images and the geometric scattering characteristics of SAR images, resulting in insufficient coordination of spectral-spatial-texture multi-dimensional features.

[0004] Therefore, how to effectively fuse the advantages of hyperspectral images and SAR images to achieve more accurate remote sensing image segmentation is an important research direction in the field of remote sensing image processing. SUMMARY

[0005] Therefore, the present application provides a remote sensing image segmentation method based on hyperspectral kernelization and spatio-temporal feature fusion, aiming to fully utilize the spectral and spatial information of multi-modal remote sensing data and improve the segmentation accuracy in complex scenes.

[0006] To achieve the above purpose, the present application adopts the following technical solutions:

[0007] A remote sensing image segmentation method based on hyperspectral kernelization and spatio-temporal feature fusion, comprising:

[0008] Obtain the initial hyperspectral image features, convert them into a sequence, calculate the spectral correlation coefficient, convert the spectral correlation coefficient into the form of a radial basis function kernel, and perform a reverse sequence operation to obtain the hyperspectral kernelization features; add the hyperspectral kernelization features and the initial hyperspectral features to obtain the hyperspectral image features;

[0009] The multi-scale feature extraction is performed on the synthetic aperture radar image, and the radar image feature is obtained after splicing and fusing.

[0010] The spatial and temporal feature fusion is performed on the hyperspectral image feature and the radar image feature, and a segmentation image is obtained.

[0011] In an optional implementation, the spectral correlation coefficient is calculated, including:

[0012] The hyperspectral image feature converted into a sequence is projected into three different spaces of query, key and value, and the spectral correlation coefficient is determined according to the following formula:

[0013]

[0014] In the formula, q and k respectively represent a query vector and a key vector of a certain spectrum in the hyperspectral image feature sequence, and respectively represent the mean of the query vectors and the mean of the key value vectors of all spectra in the hyperspectral image feature sequence.

[0015] In an optional implementation, after the spectral correlation coefficient is converted into the form of a radial basis function kernel, the spectral correlation coefficient is multiplied by the value vector and then subjected to the inverse sequence operation.

[0016] In an optional implementation, the spectral correlation coefficient is converted into the form of a radial basis function kernel according to the following formula:

[0017]

[0018] In the formula, sigma is an adjustable temperature parameter for controlling the smoothness of the kernel function, S(q, k) represents the spectral correlation coefficient, and K S represents the kernel function, which is the inner product of two feature mappings.

[0019] In an optional implementation, multi-layer hyperspectral kernelized feature extraction is performed on the initial hyperspectral image feature, and the initial hyperspectral image feature is added together through a skip connection to obtain the hyperspectral image feature.

[0020] In an optional implementation, after the multi-scale feature extraction is performed on the synthetic aperture radar image, splicing and fusing are performed, and then the radar image feature is obtained through a convolution layer, a flattening layer, a fully connected layer and an activation function layer in sequence.

[0021] In an optional implementation, the spatial and temporal feature fusion is performed on the hyperspectral image feature and the radar image feature, including:

[0022] The average pooling and the maximum pooling are respectively performed on the hyperspectral image feature and the radar image feature;

[0023] The channel weight is obtained based on the average pooling feature, and the spatial weight is obtained based on the maximum pooling feature;

[0024] The channel weight and the spatial weight corresponding to the hyperspectral image feature and the radar image feature are summed, and the hyperspectral image feature and the radar image feature are weighted and fused respectively.

[0025] In an optional embodiment, a composite function of weighted sum of cross-entropy loss, Dice loss and Focal loss is used as the final loss function.

[0026] The application discloses a remote sensing image segmentation method based on hyperspectral kernelization and spatiotemporal feature fusion.

[0027] The experimental results show that the method is better than the existing method on multiple remote sensing data sets, and can realize more accurate remote sensing image segmentation in a complex ground object environment. BRIEF DESCRIPTION OF DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.

[0029] Figure 1 The hyperspectral kernelization module structure schematic diagram provided by the present application is shown in the figure;

[0030] Figure 2 The multi-scale convolution module structure schematic diagram provided by the present application is shown in the figure;

[0031] Figure 3 The spatiotemporal feature fusion module structure schematic diagram provided by the present application is shown in the figure. DETAILED DESCRIPTION

[0032] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0033] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without the specific details set forth in this description, that the present application can be practiced with other systems, and that the present application can be practiced using different techniques. Therefore, the present application is not limited to the embodiments described herein but rather the scope of the present application is to be given by the appended claims and their equivalents.

[0034] The embodiment of the present application aims at the multi-modal characteristics of remote sensing images, discloses a remote sensing image segmentation method based on hyperspectral kernelization and spatio-temporal feature fusion, which is used for fully extracting the spatial and spectral information of the image and improving the segmentation accuracy in complex scenes. The method mainly includes three parts of hyperspectral kernelization and multi-scale convolution feature extraction, and spatio-temporal feature fusion. The steps are as follows:

[0035] An initial hyperspectral image feature is obtained, the spectral correlation coefficient is calculated after being converted into a sequence, the spectral correlation coefficient is converted into the form of a radial basis function kernel and is subjected to a reverse sequence operation to obtain a hyperspectral kernelized feature; and the hyperspectral kernelized feature and the initial hyperspectral feature are added to obtain a hyperspectral image feature;

[0036] Multi-scale feature extraction is performed on a synthetic aperture radar image to obtain a radar image feature after splicing and fusion;

[0037] Spatio-temporal feature fusion is performed on the hyperspectral image feature and the radar image feature to obtain a segmentation image.

[0038] In the embodiment, the first part is that a hyperspectral image usually contains rich spectral information, which can provide important prior knowledge for target detection and classification. In order to more effectively utilize the spectral correlation of hyperspectral data, the present application proposes a hyperspectral kernelization scheme.

[0039] The first part is that a hyperspectral image usually contains rich spectral information, which can provide important prior knowledge for target detection and classification. In order to more effectively utilize the spectral correlation of hyperspectral data, the present application proposes a hyperspectral kernelization scheme.

[0040] In the embodiment, first, a hyperspectral image is subjected to a serialization operation, a three-dimensional tensor is converted into a one-dimensional sequence, and a self-attention mechanism is used to calculate spectral correlation. Different from the traditional point product similarity calculation, the present application introduces a spectral correlation coefficient to enhance the utilization ability of different band information. Subsequently, the spectral correlation coefficient is converted into the form of a radial basis function (RBF) kernel by using kernelization technology, so as to improve the calculation efficiency and reduce the calculation complexity. Finally, the feature structure is restored through a reverse serialization operation, so that the information of the hyperspectral image is fully fused in the spatial and spectral dimensions.

[0041] In one embodiment, the present application constructs a hyperspectral kernelization module, aiming at improving the utilization ability of different band information, enhancing the spatial-spectral fusion effect, and improving the accuracy of remote sensing image segmentation.

[0042] The structure thereof is referred to Figure 1Specifically, the key features of the hyperspectral image are extracted by using spectral correlation calculation, kernelization transformation, and serialization / deserialization operation, and effective fusion of cross-band information is realized through a multi-level structure.

[0043] In an exemplary embodiment, the steps are as follows:

[0044] 1.1 Initial hyperspectral feature extraction

[0045] First, the hyperspectral image is input into a 3x3 convolution layer for preliminary feature extraction to obtain the initial hyperspectral feature map where H and W represent the height and width of the image, respectively, and C is the number of spectral channels. The main purpose of this step is to extract low-level spatial and spectral features through convolution operation to lay the foundation for subsequent deep feature processing.

[0046] 1.2 Construction of hyperspectral kernelization module

[0047] 1) Feature serialization: First, an up-sampling layer is used to process the image to improve its feature resolution. Then, through serialization operation, the three-dimensional tensor is converted into a one-dimensional sequence where N HW = H x W is the length of the sequence, and C is the number of spectral channels. This step flattens the spectral information of each pixel into a one-dimensional sequence, allowing the subsequent attention mechanism to operate in a one-dimensional feature space;

[0048] 2) Self-attention mechanism and spectral correlation:

[0049] Next, similar to the self-attention mechanism in Transformer, the input feature is projected into three different spaces: query (Q), key (K), and value (V), through linear transformation to obtain:

[0050]

[0051] where W Q , W K , and W V are learnable weight matrices used to map the input feature to the query, key, and value spaces;

[0052] 3) Spectral correlation coefficient calculation:

[0053] The traditional self-attention mechanism calculates the similarity between the query and the key through dot product (such as cosine similarity), while the present invention introduces a spectral correlation coefficient to replace the traditional similarity measure. The calculation formula of the spectral correlation coefficient is:

[0054]

[0055] In the formula, q and k respectively represent a query vector and a key vector of a certain spectrum in a hyperspectral image feature sequence, And The query vector mean and the key value vector mean of all spectra in the hyperspectral image feature sequence are respectively represented.

[0056] The square of the spectral correlation coefficient ensures non-negativity and can effectively reflect the correlation between spectral curves. The method is not sensitive to the amplitude change (such as shadow or occlusion) of the spectral curve, so it can provide prior knowledge of spectral information and enhance the adaptability of the model to remote sensing data.

[0057] 4) Kernelization

[0058] In order to further reduce the computational complexity, the kernel technology is used to process the spectral correlation coefficient. Specifically, the SCC is converted into a radial basis function (RBF) kernel form:

[0059]

[0060] In the formula, σ is an adjustable temperature parameter for controlling the smoothness of the kernel function, S(q, k) represents the spectral correlation coefficient, K S Indicates the kernel function, which is the inner product of two feature mappings, and the formula is represented as:

[0061] K S (q,k)=<ψ(q),ψ(k)〉

[0062] Where ψ(·) is a feature mapping function, which can be approximated by Taylor expansion.

[0063] After converting the spectral correlation coefficient into a radial basis function kernel form, it is multiplied by the value vector and then subjected to a reverse sequence operation.

[0064] 5) Reverse sequence operation

[0065] After processing by the hyperspectral kernel module, the output feature is still a sequence In order to restore it to an image format, a reverse sequence operation is needed, that is, the sequence is rearranged into the shape of the original image for subsequent processing steps.

[0066] To further optimize the above technical solution, the present application performs multi-layer hyperspectral kernel feature extraction on the initial hyperspectral image features, that is, further extracts features by applying N hyperspectral kernel modules. The extracted features are connected to the initial hyperspectral features Addition is performed to retain more original feature information. Subsequently, a 3x3 convolution layer is applied to further process the fused features, and finally the feature representation φ of the hyperspectral image is obtained HSI , providing rich multi-modal features for subsequent remote sensing image segmentation tasks.

[0067] Through this process, the spatial and spectral information in the hyperspectral image is effectively extracted, providing strong feature support for remote sensing image segmentation.

[0068] Second part:

[0069] For SAR images, since they lack explicit spectral information and rely mainly on backscattering characteristics for target recognition, this embodiment uses a multi-scale convolution module for feature extraction, and through feature splicing and convolution fusion, the expression ability of SAR images is improved, so that they can enhance the edge, texture and structural features of SAR images in multi-scale space, making them more suitable for remote sensing image segmentation tasks.

[0070] In an embodiment, in the feature extraction process of synthetic aperture radar (SAR) images, a multi-scale convolution module is constructed to fully capture information at different scales, as shown in Figure 2 , which enhances the expression ability of features by combining multiple convolution kernel sizes.

[0071] In an exemplary embodiment, the feature extraction process is as follows:

[0072] 2.1 Multi-scale convolution extraction; this invention uses three different sizes of convolution kernels, 3x3, 5x5 and 7x7, to perform parallel feature extraction on SAR images to capture spatial information at different scales. Among them, self.scale1, self.scale2 and self.scale3 correspond to three different sizes of convolution kernels and are used to extract local texture, edge information and larger range of structural features;

[0073] 2.2 Multi-scale feature splicing; for the features extracted by the above different convolution kernels, this invention uses the feature splicing (torch.cat) operation to splice the feature maps at each scale along the channel dimension to form a fused feature representation with multi-scale information, in order to enhance the model's perception ability of the target area;

[0074] 2.3 Channel fusion and dimension reduction; in order to further fuse features at different scales and reduce the channel dimension to reduce the amount of calculation, this invention introduces a 3x3 convolution layer (self.conv1x1) to perform channel transformation on the spliced features, and uses a flattening operation to strengthen key information while reducing redundant features;

[0075] 2.4 Full connection and activation; after convolution feature extraction and fusion, the present application further introduces a full connection layer and an activation function to enhance the non-linear expression ability of the features and improve the discrimination ability of the model. Finally, the feature distribution is further optimized through 1x1 convolution to provide high-quality feature representation for subsequent segmentation tasks. Finally, the radar image feature φ SAR .

[0076] Part III:

[0077] After extracting the features of hyperspectral images and SAR images, the embodiment further constructs a spatio-temporal feature fusion module to fully utilize the spatio-temporal relationship of these data and improve the effect of remote sensing image segmentation. The module first uses a channel attention mechanism to calculate the importance of different channels through global pooling and one-dimensional convolution, and uses a Softmax function for normalization, thereby dynamically adjusting the weight of each channel. In addition, a spatial attention mechanism is used to perform global pooling and two-dimensional convolution on the feature map to calculate the weight in the spatial dimension, and the Softmax normalization is also used to ensure the rationality of feature fusion.

[0078] Finally, through a multi-level fusion strategy, the accuracy of the segmentation result is improved by fully combining remote sensing image information from different time points and different sources.

[0079] In one embodiment, the fusion process refers to Figure 3 ;φ HSI and φ SAR represent hyperspectral image features and radar image features, respectively, and these feature maps contain remote sensing image information from different time points or different sources. Through the spatio-temporal feature fusion module, the spatio-temporal relationship of the two different features can be fully utilized to improve the effect of remote sensing image segmentation, especially when dealing with complex ground object classes.

[0080] In an exemplary embodiment, the specific process includes:

[0081] 3.1 Building a channel attention mechanism

[0082] 3.11 Global pooling: Global average pooling (Avg Pooling) is performed on the input feature maps φ HSI and φ SAR to aggregate information in the channel dimension. The specific operation is:

[0083] φ C = Concat(Avg(φ HSI ), Avg(φ SAR ))

[0084] where φ CThis represents the channel features obtained through pooling operations, and Concat represents concatenation along the channel dimension.

[0085] 3.12 One-dimensional Convolution: Next, a one-dimensional convolutional layer is used to calculate φ separately. HSI and φ SAR The importance of each channel in the channel dimension. These convolutional layers will be used to learn the weights in the channel dimension, thereby dynamically adjusting the contribution of each channel to the final feature:

[0086] W c1 =Conv1(φ C ),W c2 =Conv2(φ C )

[0087] Here, Conv represents a convolutional layer. These convolutional operations learn the contributions of different channels, which helps the model adaptively adjust the weights in subsequent steps.

[0088] 3.13 Softmax Normalization: The Softmax function is used to normalize the channel weights, ensuring that the sum of the weights is 1.

[0089]

[0090] Here, e is the base of the exponential function, and the normalized weights help the model focus on more important channels during the fusion process.

[0091] 3.2 Constructing a Spatial Attention Mechanism

[0092] 3.21 Spatial Pooling: Performs global pooling operations in the spatial dimension to compute the feature map φ. HSI and φ SAR Spatial importance:

[0093] φ S =Concat(Max(φ) HSI ),Max(φ SAR ))

[0094] φ S It can capture important information in the spatial dimension from images from different sources, preparing for subsequent spatial weighting.

[0095] 3.22 Two-Dimensional Convolution: Next, two-dimensional convolutional layers are used to compute weights in the spatial dimension. The purpose of this operation is to further enhance the fusion of spatial information and improve the consistency of feature maps from different sources in the spatial dimension.

[0096] W s1 =Conv2D1(φ S ),W s2 =Conv2D2(φS )

[0097] This operation is used to capture the weight relationship in the spatial dimension, further strengthening the fusion of spatial information.

[0098] 3.23 Softmax normalization: normalize the spatial weights:

[0099]

[0100] Through this normalization, the model can effectively focus on the more important areas in the spatial dimension.

[0101] 3.3 Feature fusion. Fuse the dual-phase weights, and weight sum the channel weights and spatial weights to get the final feature weights:

[0102] W′ f1 = W′ c1 + W′ s1 , W′ f2 = W′ c2 + W′ s2

[0103] The feature maps φ HSI and φ SAR are fused using the weighted sum method, and the final fused feature map is obtained:

[0104] φ + = (W′ f1 × φ HSI ) + (W′ f2 × φ SAR )

[0105] φ + The fused feature map contains important information from the two input feature maps φ HSI and φ SAR , and is ready to be passed to the downstream task module for further processing.

[0106] Further, the feature map is converted into a segmentation result, i.e. φ + contains rich information from different sensors and time points, which is passed to a convolutional layer to generate the final segmentation result. The core purpose of this step is to convert the fused multi-modal features into the final pixel-level segmentation map through convolutional operation, so as to realize accurate segmentation of different ground object categories in remote sensing images.

[0107] To ensure that the segmentation result can output a probability value at each pixel position, the present application uses a Sigmoid activation function to process the convolution layer output in this step. The Sigmoid function can map the output of each pixel to between 0 and 1, representing the probability of the pixel belonging to a certain class. In this way, the final segmentation result will be a probability map, with each pixel point corresponding to a probability value of a class

[0108] In an optional embodiment, the cross-entropy loss function is used for training. The cross-entropy loss function can effectively measure the difference between the probability distribution of the model output and the actual label, thereby guiding the model to optimize the parameters. Specifically, the loss function is as follows:

[0109]

[0110] wherein, is the predicted probability value, y i is the true label, N y is the total number of pixels in the image. By minimizing this loss function, the model can continuously adjust the parameters to improve the segmentation accuracy.

[0111] The Dice loss is used to measure the overlap between the model's prediction and the true label, and is particularly effective for unbalanced classes. The formula is as follows:

[0112]

[0113] This loss can give greater weight to difficult-to-classify samples and reduce the influence of simple samples.

[0114] Further;

[0115]

[0116] wherein, is the predicted probability of the target class, α is the balance factor, and γ is the focus parameter.

[0117] In this embodiment, the above loss functions are weighted and summed to obtain a composite loss function:

[0118]

[0119] wherein, λ1, λ2 and λ3 are weight coefficients, which can be adjusted according to the needs of the task.

[0120] In this embodiment, after calculating the loss using the cross-entropy loss function, the parameters of the model are updated through the backpropagation algorithm. During the optimization process, the Adam optimizer is used to accelerate convergence, and the learning rate is adjusted to avoid overfitting. As the training progresses, the model continuously adjusts its weights to minimize the loss function, ultimately obtaining a model that can accurately perform remote sensing image segmentation. Through these steps, the final segmentation result will be able to accurately separate different ground object categories from the remote sensing image and provide high-quality segmentation images for downstream tasks.

[0121] Compared with the prior art, the beneficial effects of the present application include:

[0122] 1) The hyperspectral kernelization process improves the utilization efficiency of spectral information. Through spectral correlation calculation, self-attention mechanism and kernel transformation, the cross-band information fusion ability of hyperspectral images is enhanced, which reduces the computational complexity while utilizing spectral prior knowledge, and improves the segmentation accuracy.

[0123] 2) The application of multi-scale convolution enhances the feature expression ability of SAR images. Different scale convolution kernels are used to extract edge, texture and structure features of SAR images in parallel, and the representation ability is optimized through feature splicing and channel fusion, so that it can adapt to remote sensing scenes with different resolutions and complex backgrounds.

[0124] 3) The spatio-temporal feature fusion improves the collaborative analysis ability of multi-temporal and multi-sensor data. Combined with channel attention mechanism and spatial attention mechanism, this module can adaptively adjust the feature weights of different source remote sensing data, strengthen key information and reduce redundancy, and improve the target segmentation effect in complex environments.

[0125] 4) High computational efficiency, the present application is suitable for large-scale remote sensing data processing. Through kernel technology to reduce the complexity of attention calculation, combined with lightweight one-dimensional / two-dimensional convolution operation to optimize the calculation overhead, so that the present application can efficiently process large-scale remote sensing data and meet the actual application requirements.

[0126] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the related parts can be referred to the method part.

[0127] The foregoing description of the disclosed embodiments enables a person skilled in the art to make or use the application. Modifications of these embodiments will occur to persons of skill in the art, and that the appended claims are intended to cover all such modifications that do not depart from the true spirit and scope of the application. Therefore, the application is not limited to the embodiments shown but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A remote sensing image segmentation method based on hyperspectral kernelization and spatio-temporal feature fusion, characterized in that, an initial hyperspectral image feature is obtained, a spectral correlation coefficient is calculated after being converted into a sequence, the spectral correlation coefficient is converted into a radial basis function kernel form and an inverse sequence operation is performed to obtain a hyperspectral kernelized feature; the hyperspectral kernelized feature and the initial hyperspectral feature are added to obtain a hyperspectral image feature; the spectral correlation coefficient is calculated, including: the hyperspectral image feature converted into a sequence is projected into three different spaces of query, key and value, and the spectral correlation coefficient is determined according to the following formula; In the formula, q and k respectively represent a query vector and a key vector of a certain spectrum in a hyperspectral image feature sequence, and respectively represent the mean of the query vector and the mean of the key value vector of all spectra in the hyperspectral image feature sequence. and the spectral correlation coefficient is converted into a radial basis function kernel form according to the following formula: where σ is a temperature parameter that can be adjusted to control the degree of smoothing of the kernel function, S(q, k) represents the spectral correlation coefficient, K S represents the kernel function; multi-scale feature extraction is performed on a synthetic aperture radar image, and radar image features are obtained after splicing and fusion; spatio-temporal feature fusion is performed on the hyperspectral image features and the radar image features to obtain a segmentation image.

2. The remote sensing image segmentation method of claim 1, wherein, After the spectral correlation coefficient is converted into a radial basis function kernel form, it is multiplied by the value vector and then subjected to an inverse sequence operation.

3. The remote sensing image segmentation method of claim 1, wherein, Multi-layer hyperspectral kernelized feature extraction is performed on the initial hyperspectral image feature, and the initial hyperspectral image feature is added through a skip connection to obtain the hyperspectral image feature.

4. The remote sensing image segmentation method of claim 1, wherein, After multi-scale feature extraction and splicing fusion of a synthetic aperture radar image, radar image features are obtained through a convolution layer, a flattening layer, a fully connected layer and an activation function layer in turn.

5. The method of claim 1, wherein, Spatio-temporal feature fusion is performed on the hyperspectral image features and the radar image features, including: average pooling and maximum pooling are performed on the hyperspectral image features and the radar image features, respectively; channel weights are obtained based on the average pooled features, and spatial weights are obtained based on the maximum pooled features; the channel weights and the spatial weights corresponding to the hyperspectral image features and the radar image features are added, and the hyperspectral image features and the radar image features are weighted and fused, respectively.

6. The remote sensing image segmentation method of claim 1, wherein, A composite function of weighted sum of cross-entropy loss, Dice loss and Focal loss is used as the final loss function to measure the difference between the probability distribution of the model output and the actual label, so as to guide the model to optimize the parameters. With the training, the model will continuously adjust its weights to minimize the loss function, and finally obtain a model that can accurately perform remote sensing image segmentation.

Citation Information

Patent Citations

  • Hyperspectral-remote-sensing-image classification technology combining spectral, spatial and hierarchical information

    CN108427913A

  • Rapid soil classification method based on visible near infrared spectrum and multi-target fusion

    CN108827909A