Remote sensing image ground feature classification method, device, equipment and product
By transforming the remote sensing image from the spatial domain to the frequency domain, extracting and fusing the spatial domain and frequency domain features, and using multimodal features for geographic classification, the problem of insufficient geographic classification accuracy in remote sensing images in the prior art is solved, and higher classification accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510713150.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-30
AI Technical Summary
In the prior art, the accuracy of the classification of land objects in remote sensing images is insufficient, especially in complex scenarios, it is difficult to effectively distinguish land objects from noise.
By transforming the remote sensing image from the spatial domain to the frequency domain, the spatial domain feature map and the frequency domain feature map are extracted, and fusion is performed, and multimodal features are used for geographic classification. Specific methods include feature extraction using residual networks and two-layer routing attention networks, and feature fusion through dynamic exponential sliding averages.
Improve the accuracy, robustness and stability of the classification of land objects in remote sensing images, especially in complex scenarios, where land objects and noise can be more accurately identified.
Smart Images

Figure CN120219864A_ABST
Abstract
Description
Technical Field
[0001] This application is applied to the field of remote sensing image technology, and particularly relates to a method, device, equipment and product for classifying ground objects in remote sensing images. Background Art
[0002] With the rapid development of remote sensing technology, remote sensing images have been widely used in the fields of land resource monitoring, environmental assessment, disaster warning, etc., including classifying ground objects based on remote sensing images.
[0003] In practical applications, remote sensing images are usually acquired using multispectral sensors or hyperspectral sensors. These sensors can capture the spectral characteristics of ground objects in different bands, making the remote sensing images contain rich information on the characteristics of ground objects. At the same time, remote sensing images also have the characteristics of high complexity. For example, different ground objects may exhibit similar spectral characteristics, and the boundaries of different ground objects are blurred.
[0004] Facing complex remote sensing images, the accuracy of classifying ground objects in remote sensing images still needs to be improved. Summary of the Invention
[0005] To solve the above problems, this application proposes a method, device, equipment and product for classifying ground objects in remote sensing images, which can improve the accuracy of classifying ground objects in remote sensing maps.
[0006] The first aspect of this application provides a method for classifying ground objects in a remote sensing map, including: obtaining a target remote sensing image; transforming the target remote sensing image from the spatial domain to the frequency domain to obtain a frequency-domain image; extracting features from the target remote sensing image to obtain a spatial-domain feature map; extracting features from the frequency-domain image to obtain a frequency-domain feature map; fusing the spatial-domain feature map and the frequency-domain feature map to obtain a fused feature map; classifying the fused feature map to obtain the ground object classification result of the target remote sensing image.
[0007] In a possible implementation, feature extraction is performed on the target remote sensing image to obtain a spatial domain feature map; feature extraction is performed on the frequency domain image to obtain a frequency domain feature map; the spatial domain feature map and the frequency domain feature map are fused to obtain a fused feature map; and the fused feature map is classified to obtain the ground object classification result of the target remote sensing image, including: inputting the target remote sensing image and the frequency domain image into a classification model; performing feature extraction on the target remote sensing image through a first feature extraction network in the classification model to obtain the spatial domain feature map; performing feature extraction on the frequency domain image through a second feature extraction network in the classification model to obtain the frequency domain feature map; fusing the spatial domain feature map and the frequency domain feature map through a feature fusion network in the classification model to obtain the fused feature map; and classifying the fused feature map through a classification network in the classification model to obtain the ground object classification result of the target remote sensing image.
[0008] In a possible implementation, the first feature extraction network is a residual network, and the residual network includes a convolution module and a plurality of cascaded residual blocks. The performing feature extraction on the target remote sensing image through the first feature extraction network in the classification model to obtain the spatial domain feature map includes: performing feature extraction on the target remote sensing image through the convolution module to obtain a convolution feature map of the target remote sensing image; and performing multiple downsamplings on the convolution feature map through the plurality of cascaded residual blocks to obtain the spatial domain feature map.
[0009] In a possible implementation, the second feature extraction network is a double-layer routing attention network, and the double-layer routing attention network includes a convolution module and an attention module. The attention mechanism of the attention module adopts a double-layer routing attention mechanism. The performing feature extraction on the frequency domain image through the second feature extraction network in the classification model to obtain the frequency domain feature map includes: performing feature extraction on the frequency domain image through the convolution module to obtain a convolution feature map of the frequency domain image; performing cross-region feature interaction on the convolution feature map through the attention module based on the double-layer routing attention mechanism to obtain a first attention feature map; and obtaining the frequency domain feature map according to the first attention feature map.
[0010] In a possible implementation, the convolution module includes a convolutional layer and a feature concatenation layer. The position encoding matrix in the feature concatenation layer is a learning parameter during the training process. By using the convolution module to perform feature extraction on the frequency-domain image to obtain the convolutional feature map of the frequency-domain image, the method includes: dividing the frequency-domain image into blocks to obtain a plurality of non-overlapping image blocks; performing feature extraction on each of the plurality of non-overlapping image blocks through the convolutional layer to obtain the feature maps corresponding to the plurality of non-overlapping image blocks respectively; in the feature concatenation layer, sorting the feature maps corresponding to the plurality of non-overlapping image blocks according to the position order of the plurality of non-overlapping image blocks, and adding position information to the sorted feature maps according to the position encoding matrix to obtain the convolutional feature map.
[0011] In a possible implementation, the attention module includes an attention layer. By using the attention module to perform cross-region feature interaction on the convolutional feature map based on a two-layer routing attention mechanism to obtain a first attention feature map, the method includes: dividing the convolutional feature map into windows to obtain a plurality of feature windows; determining the two-layer routing between the plurality of feature windows according to the feature similarity between the plurality of feature windows in pairs, where the two-layer routing indicates the associated feature windows corresponding to the plurality of feature windows respectively in the dimension of feature similarity; determining the query vectors corresponding to the plurality of feature windows respectively, the key vectors corresponding to the plurality of feature windows respectively, and the value vectors corresponding to the plurality of feature windows respectively according to the two-layer routing and the feature representations corresponding to the plurality of feature windows on the convolutional feature map; in the attention layer, determining the attention weight matrix according to the query vectors corresponding to the plurality of feature windows respectively, the key vectors corresponding to the plurality of feature windows respectively, the value vectors corresponding to the plurality of feature windows respectively, and the weight parameters of the attention layer, and adding the attention weight matrix to the convolutional feature map to obtain the first attention feature map.
[0012] In a possible implementation, the attention module further includes layer normalization and a feed-forward neural network. After using the attention module to perform cross-region feature interaction on the convolutional feature map based on a two-layer routing attention mechanism to obtain a first attention feature map, the method further includes: performing preliminary processing on the first attention feature map through the layer normalization; enhancing the preliminarily processed attention feature map through the feed-forward neural network.
[0013] In a possible implementation, obtaining the frequency-domain feature map based on the first attention feature map includes: adjusting the dimension of the first attention feature map; through the attention module, performing cross-region feature interaction on the first attention feature map with adjusted dimension based on the double-layer routing attention mechanism to obtain a second attention feature map; and obtaining the frequency-domain feature map according to the second attention feature map.
[0014] In a possible implementation, fusing the spatial-domain feature map and the frequency-domain feature map through the feature fusion network in the classification model to obtain the fused feature map includes: in the feature fusion network, using dynamic exponential moving average to fuse the spatial-domain feature map and the frequency-domain feature map to obtain the fused feature map, where the smoothing factor in the exponential moving average is related to the training parameters of the feature fusion network.
[0015] A second aspect of the present application provides a device for classifying ground objects in a remote sensing map, including: an acquisition unit for acquiring a target remote sensing image; a spatial-domain frequency-domain transformation unit for transforming the target remote sensing image from the spatial domain to the frequency domain to obtain a frequency-domain image; a spatial-domain feature extraction unit for extracting features from the target remote sensing image to obtain a spatial-domain feature map; a frequency-domain feature extraction unit for extracting features from the frequency-domain image to obtain a frequency-domain feature map; a feature fusion unit for fusing the spatial-domain feature map and the frequency-domain feature map to obtain a fused feature map; and a ground object classification unit for classifying the fused feature map to obtain the ground object classification result of the target remote sensing image.
[0016] A third aspect of the present application provides an electronic device, including a memory and a processor; the memory is connected to the processor for storing a program; the processor is configured to implement the method for classifying ground objects in a remote sensing map as described in the first aspect of the present application or any possible implementation manner of the first aspect of the present application by running the program in the memory.
[0017] A fourth aspect of the present application provides a chip, including a processor and a data interface, and the processor reads and runs a program stored on a memory through the data interface to execute the method for classifying ground objects in a remote sensing map as described in the first aspect of the present application or any possible implementation manner of the first aspect of the present application.
[0018] A fifth aspect of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method for classifying ground objects in a remote sensing map as described in the first aspect of the present application or any possible implementation manner of the first aspect of the present application.
[0019] A sixth aspect of the present application provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the method for classifying ground objects in a remote sensing map as described in the first aspect of the present application or any possible implementation manner of the first aspect of the present application.
[0020] According to a method, device, equipment and product for classifying ground objects in a remote sensing map proposed by the present application, in the process of classifying ground objects in a target remote sensing image, the target remote sensing image is converted from the spatial domain to the frequency domain to obtain a frequency domain image; feature extraction is respectively performed on the target remote sensing image and the frequency domain image to obtain a spatial domain feature map and a frequency domain feature map. The spatial domain feature map reflects the detailed features of the target remote sensing image, such as color, texture, etc., and the frequency domain feature map reflects the distribution of different frequency components in the target remote sensing image, which can better distinguish the noise and effective information of the target remote sensing image; by classifying the fused feature map obtained by fusing the spatial domain feature map and the frequency domain feature map, the ground object classification result of the target remote sensing image is determined. Compared with the ground object classification of remote sensing images based on single-modal features, the present application can extract richer multi-modal features from complex remote sensing images, combine the advantages of the spatial domain feature map and the frequency domain feature map, and effectively improve the accuracy of ground object classification of remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0022] Figure 1 It is a schematic diagram of the implementation environment related to the embodiments of the present application.
[0023] Figure 2 It is a flowchart of the method for classifying ground objects in a remote sensing map provided by the embodiments of the present application Figure 1 。
[0024] Figure 3 It is a flowchart of the method for classifying ground objects in a remote sensing map provided by the embodiments of the present application Figure 2 。
[0025] Figure 4 It is an example diagram of classifying ground objects in a remote sensing image through a classification model.
[0026] Figure 5 It is a schematic structural diagram of the device for classifying ground objects in a remote sensing map provided by the embodiments of the present application.
[0027] Figure 6Schematic structural diagram of an electronic device provided according to an embodiment of the present application. Detailed implementation manners
[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0029] Sensors for collecting remote sensing images (such as multispectral sensors, hyperspectral sensors) can capture the spectral characteristics of ground objects in different bands (for example, the chlorophyll content and water content of crops have unique spectral characteristics in different bands). However, different ground objects may have similar spectral characteristics (for example, different crops have similar spectral characteristics at the same growth stage. Taking wheat and corn as examples, the reflectance of wheat and corn in some bands may be very close), and the boundaries of different ground objects may be relatively blurred (for example, agricultural land surfaces include various ground object types such as bare land, roads, and water bodies, and the boundaries between some different ground object types in these ground object types are blurred). Especially in low-resolution remote sensing images, there may be mixed pixels (the spectral characteristics of mixed pixels are a mixture of the spectral characteristics of different types of ground objects). This increases the difficulty of ground object classification in remote sensing images and affects the accuracy of ground object classification in remote sensing images. A more refined classification method is needed for ground object classification in remote sensing images. Among them, the ground object classification of remote sensing images can also be called remote sensing image classification, ground object classification based on remote sensing images, and ground object recognition based on remote sensing images, which refers to extracting ground object category information from remote sensing images.
[0030] In related technologies, the ground object classification methods for remote sensing images include ground object classification methods for remote sensing images based on traditional machine learning and ground object classification methods for remote sensing images based on deep learning.
[0031] Remote sensing image ground object classification methods based on traditional machine learning include supervised classification methods and unsupervised classification methods. In the supervised classification method, a sufficient number of representative remote sensing images are manually selected for each type of typical ground object as training samples; based on the training samples, machine learning algorithms such as maximum likelihood classification (MLC), support vector machine (SVM), and random forest (RF) are used to train the classifier to learn the ground object features, and the parameters of the classifier are adjusted during the training process to optimize the classification performance of the classifier. The classification effect on complex remote sensing images is limited. In the unsupervised classification method, a clustering algorithm is used to cluster the pixel spectra of the remote sensing image, and the clustering center is continuously iteratively optimized during the clustering process, and finally multiple clusters are obtained, and the ground object types corresponding to each cluster are manually interpreted. It can be seen that the remote sensing image ground object classification method based on traditional machine learning relies on feature engineering and domain knowledge, requires combining expert experience to optimize the classification process, has a limited classification effect on complex remote sensing images, and is difficult to capture the detailed information and context information in high-resolution remote sensing images.
[0032] The specific process of the remote sensing image ground object classification method based on deep learning may include: First, obtain the remote sensing image and label the ground object types of the remote sensing image; then, preprocess and data augment the remote sensing image. The preprocessing includes cutting the large-size remote sensing image into small sizes, and the data augmentation includes rotating, flipping, cropping, adding noise, color jitter, etc. to the cut image; then, construct a deep neural network and adapt the deep neural network; train the deep neural network through an optimization strategy and a loss function; put the trained deep neural network into the actual business system to achieve the ground object classification of the remote sensing image. The deep neural network is a powerful framework that can directly learn the image expression from a large amount of image data. Compared with the remote sensing image ground object classification method based on traditional machine learning, the remote sensing image ground object classification method based on deep learning applies the deep neural network to remote sensing images containing a large amount of unknown information, improving the accuracy of ground object classification for remote sensing images. However, the remote sensing image classification method based on deep learning focuses on single spatial domain features, resulting in insufficient ground object classification ability of the deep neural network for remote sensing images.
[0033] The remote sensing image classification method based on deep learning mainly has the following deficiencies: First, it is relatively sensitive to illumination changes and scale differences, and the classification effect of ground objects in complex scenarios (such as scenarios where the ground area is obscured by clouds and fog, scenarios where the spectral features of different types of ground objects have small differences, and scenarios where the spectral features of the same type of ground objects have large differences) decreases, especially the classification stability in the scenario of cloud and fog occlusion is poor; Second, remote sensing images contain complex backgrounds and noises. The deep neural network that focuses on single spatial domain features ignores high-frequency detail information, making it difficult to distinguish foreground targets and background information in high-noise images. When facing background interference, the classification performance of the model decreases, resulting in insufficient robustness of the deep neural network for classifying ground objects in remote sensing images under different scenarios; Third, if the deep neural network focuses on single frequency domain features, it is difficult to retain accurate spatial position information, resulting in blurred boundaries of different types of ground objects. It can be seen that focusing on a single feature domain leads to a decrease in the classification ability of the deep neural network, including a decrease in stability, robustness, and accuracy.
[0034] In view of this, the embodiments of the present application propose a method, device, equipment and product for classifying ground objects in a remote sensing map. Through spatial domain-frequency domain transformation, the remote sensing image is transformed into a frequency domain image, and feature extraction is performed on the remote sensing image and the frequency domain image respectively to obtain a spatial domain feature map and a frequency domain feature map; the spatial domain feature map and the frequency domain feature map are fused, and ground object classification of the remote sensing image is performed based on the fused features. Thus, the complementary enhancement of image features across the spatial domain and the frequency domain is realized, the problem of insufficient ground object classification ability of remote sensing images based on a single feature domain is solved, and the stability, robustness and accuracy of ground object classification of remote sensing images are effectively improved, especially the accuracy of classifying remote sensing images (or complex remote sensing images) in complex scenarios.
[0035] Exemplary implementation environment Please refer to Figure 1 , Figure 1 FIG. is a schematic diagram of the implementation environment according to the embodiments of the present application. In the implementation environment involved in the present application, it includes a processing device 110. On the processing device 110, the method for classifying ground objects in a remote sensing map provided by the embodiments of the present application can be used to classify the input remote sensing map. Optionally, a classification model can be deployed on the processing device 110 to classify the input remote sensing map through the classification model.
[0036] Among them, the processing device 110 can be a terminal device (such as a mobile phone, a computer, a smart wearable device, a smart pen, a vehicle-mounted terminal, etc.) or a server. Figure 1 Taking the processing device 110 as a server as an example.
[0037] Exemplary method Please refer to Figure 2, in an exemplary embodiment, a method for classifying ground objects in a remote sensing map is provided. The method for classifying ground objects in the remote sensing map includes the following steps: S201, obtain a target remote sensing image.
[0038] In this embodiment, the target remote sensing image refers to the remote sensing image for which ground object classification is currently performed. The target remote sensing image can be obtained from a sensor (multispectral sensor or hyperspectral sensor); alternatively, the target remote sensing image sent by other devices (such as the user's terminal device) can be obtained; or, the target remote sensing image can be obtained from a task list.
[0039] S202, transform the target remote sensing image from the spatial domain to the frequency domain to obtain a frequency-domain image.
[0040] In this embodiment, the target remote sensing image belongs to a spatial-domain image. The signal transformation method can be used to transform the target remote sensing image from the spatial domain to the frequency domain to obtain the frequency-domain image of the target remote sensing image.
[0041] In one example, the Fourier transform can be used to transform the target remote sensing image from the spatial domain to the frequency domain, that is, from the spatial-domain signal to the frequency-domain signal, to obtain a frequency-domain image.
[0042] Further, the fast Fourier transform (FFT) is used to transform the target remote sensing image from the spatial domain to the frequency domain to obtain a frequency-domain image, so as to improve the efficiency of transforming the target remote sensing image from the spatial domain to the frequency domain and save the computing resources in this transformation process.
[0043] S203, perform feature extraction on the target remote sensing image to obtain a spatial-domain feature map.
[0044] Among them, the spatial-domain feature map refers to the features extracted in the pixel space of the target remote sensing image, and the information such as the gray value of the pixel point and the position of the pixel point is concerned.
[0045] In this embodiment, the image feature extraction method can be used to perform feature extraction on the target remote sensing image to obtain the image features of the target remote sensing image, that is, the spatial-domain feature map.
[0046] S204, perform feature extraction on the frequency-domain image to obtain a frequency-domain feature map.
[0047] Among them, the frequency-domain feature map refers to the features extracted from the frequency components of the frequency-domain image, and the information such as the amplitude, phase, and distribution of the frequency components is concerned.
[0048] In this embodiment, the image feature extraction method can be used to perform feature extraction on the frequency-domain image to obtain the image features of the frequency-domain image, that is, the frequency-domain feature map.
[0049] For the land cover classification of remote sensing images, the frequency domain feature map has the following advantages: On the one hand, in the frequency domain feature map, there are significant differences between the frequency components corresponding to the foreground of the image and those corresponding to the background of the image. In remote sensing images with low contrast or high noise, the frequency domain feature map can still clearly show the true image features. Therefore, by using the frequency domain image, the separation of foreground objects and background interference can be achieved, and the recognition ability of land cover areas in remote sensing images with complex backgrounds (in remote sensing images with complex backgrounds, the land cover image area may be covered by background information), low-contrast remote sensing images, and high-noise remote sensing images can be improved; On the other hand, for frequency domain images with illumination changes or scale changes, the frequency domain features are invariant. For example, the consistency of the frequency domain energy distribution of vegetation in different seasons is relatively high. Therefore, based on the frequency domain image, the stability of land cover classification of remote sensing images can be improved; On the other hand, the phase features corresponding to land covers with similar spectral features (such as asphalt roads and shadow water surfaces) are different. Therefore, based on the frequency domain feature map, the classification ability of land cover types with small inter-class differences can be improved, and the classification accuracy in scenarios with small inter-class differences can be improved.
[0050] S205, fuse the spatial domain feature map and the frequency domain feature map to obtain a fused feature map.
[0051] In this embodiment, a multi-modal feature map fusion method is used to fuse the spatial domain feature map and the frequency domain feature map to obtain a fused feature map.
[0052] In one example, the spatial domain feature map and the frequency domain feature map can be stitched together to obtain a fused feature map; alternatively, the spatial feature map and the frequency domain feature map can be weighted to obtain a fused feature map.
[0053] In another example, a feature dynamic fusion mechanism (such as an attention mechanism) can be used to dynamically fuse the spatial domain feature map and the frequency domain feature map to obtain a fused feature map, so as to improve the fusion effect of multi-modal feature maps and further improve the land cover classification accuracy of the target remote sensing image.
[0054] S206, classify the fused feature map to obtain the land cover classification result of the target remote sensing image.
[0055] In this embodiment, the fused feature map is used as the final feature information of the target remote sensing image. By classifying the fused feature map, the land cover classification result of the target remote sensing image is obtained. The land cover classification result may include the land cover area in the target remote sensing image and the land cover type corresponding to the land cover area.
[0056] In the embodiments of the present application, based on the fusion features of the spatial domain feature map and the frequency domain feature map, ground object classification of remote sensing images is performed. The spatial domain feature map reflects the visual features and spatial location features of ground objects in the remote sensing image. There are significant differences between the frequency components of the foreground and background of the remote sensing image in the frequency domain feature map, realizing the complementary enhancement of the spatial domain feature map and the frequency domain feature map, providing more comprehensive remote sensing image features for the ground object classification of remote sensing images, improving the ground object classification ability of remote sensing images, and especially improving the ground object classification accuracy of remote sensing images in complex scenes.
[0057] In some embodiments, the ground object classification of the target remote sensing image can be achieved through a classification model. Corresponding embodiments will be given below with reference to the accompanying drawings.
[0058] Please refer to Figure 3 , in another exemplary embodiment, a method for ground object classification of a remote sensing map is provided. The method for ground object classification of the remote sensing map includes the following steps: S301, obtain a target remote sensing image.
[0059] S302, transform the target remote sensing image from the spatial domain to the frequency domain to obtain a frequency domain image.
[0060] Among them, the implementation principles and technical effects of S301~S302 can refer to the foregoing embodiments and will not be elaborated herein.
[0061] In a possible implementation manner, S302 may include: performing a Fourier transform on the target remote sensing image to obtain a spectrogram; performing centering and / or enhancement on the spectrogram to obtain a preprocessed spectrogram; and copying the preprocessed spectrogram multiple times to obtain a frequency domain image.
[0062] In this implementation manner, by performing a Fourier transform on the target remote sensing image, the target remote sensing image is converted from the spatial domain to the frequency domain to obtain a spectrogram. In the spectrogram, the low-frequency components correspond to the slowly changing parts of the target remote sensing image and contain most of the effective information of the target remote sensing image, such as the general shape and spectral characteristics of the ground objects in the target remote sensing image; the high-frequency components correspond to the rapidly changing parts of the target remote sensing image, such as object edges, image details, and image noise; the low-frequency components are usually distributed at the edges of the spectrogram (such as the left region of the spectrogram). In order to extract more effective information, the spectrogram can be centered, that is, the low-frequency part of the spectrogram is moved to the center position of the spectrogram; the spectrogram can be enhanced, especially the high-frequency features related to object edges and image details, to obtain richer high-frequency features from the spectrogram. After the preprocessing operations of centering or enhancement, the preprocessed spectrogram can be copied multiple times to increase the number of image channels so that the number of channels of the frequency domain image meets the requirements of the classification model for the number of channels of the input image.
[0063] Optionally, the formula for performing Fourier transform on the target remote sensing image is expressed as: ; where M and N respectively represent the number of rows and columns of pixel points in the target remote sensing image, represents the pixel value at the image coordinates (x, y) in the target remote sensing image, represents the signal metric value (which may include the amplitude and phase of signal vibration) at the image coordinates (u, v) in the spectrogram, is the imaginary unit, =-1, is the negative exponential base, and this negative exponential base represents orthogonal basis functions of different frequencies and is used to decompose signals in the spatial domain.
[0064] Optionally, the formula for centering the spectrogram is expressed as: ; where, represents the centered spectrogram.
[0065] Optionally, after centering the spectrogram to obtain the centered spectrogram, the centered spectrogram is enhanced, and the formula for enhancing the centered spectrogram is expressed as: ; where, represents taking the absolute value, represents the enhanced spectrogram. By means of logarithmic transformation, the signal metric value of the spectrogram is adjusted to a numerical interval more suitable for processing and observation, highlighting the frequency components with small and meaningful signal metric values and enhancing the contrast between different frequency components.
[0066] Optionally, the input image of the classification model is a three-channel image, and the target remote sensing image is a three-channel image; the preprocessed spectrogram is copied multiple times, including: copying the preprocessed spectrogram three times to obtain a frequency-domain image with three channels, so that the number of channels of the frequency-domain image is the same as that of the target remote sensing image, meeting the requirement of the classification model for the number of channels of the input image.
[0067] S303, input the target remote sensing image and the frequency-domain image into the classification model.
[0068] In this embodiment, the target remote sensing image and the frequency-domain image can be preprocessed respectively, and the preprocessed target remote sensing image and the preprocessed frequency-domain image are respectively input into the classification model.
[0069] In a possible implementation, preprocessing is performed on the target remote sensing image and the frequency-domain image respectively, including: performing a normalization operation on the target remote sensing image and the frequency-domain image respectively, to obtain the normalized target remote sensing image and the normalized frequency-domain image. After that, the normalized target remote sensing image and the normalized frequency-domain image can be input into the classification model respectively.
[0070] In this implementation, the formula for performing the normalization operation on the target remote sensing image can be expressed as: ; where represents the target remote sensing image, and the image size of the target remote sensing image is, for example, 256×256×3. represents the normalized target remote sensing image, represents the mean value of the pixel values in the target remote sensing image, represents the standard deviation of the pixel values in the target remote sensing image.
[0071] Among them, the formula for performing the normalization operation on the frequency-domain image can refer to the formula for performing the normalization operation on the target remote sensing image, and will not be elaborated here.
[0072] S304. Through the first feature extraction network in the classification model, feature extraction is performed on the target remote sensing image to obtain a spatial-domain feature map.
[0073] Among them, the classification model can adopt a deep neural network. The classification model includes a feature extraction network for remote sensing images, a feature extraction network for frequency-domain images, a feature fusion network for fusing the spatial-domain feature map and the frequency-domain feature map, and a classification network for ground object classification.
[0074] In the classification model, the feature extraction network for remote sensing images and the feature extraction network for frequency-domain images can be different feature extraction networks, so as to design and train the corresponding feature extraction networks respectively according to the differences between the spatial-domain image and the frequency-domain image, and improve the accuracy of spatial-domain feature and frequency-domain feature extraction. For the convenience of distinction, the feature extraction network for remote sensing images is called the first feature extraction network, and the feature extraction network for frequency-domain images is called the second feature extraction network.
[0075] In this embodiment, the first feature extraction network includes multiple network layers. After inputting the target remote sensing image into the classification model, through the multiple network layers in the first feature extraction network, feature extraction is performed on the target remote sensing image to obtain a spatial-domain feature map.
[0076] In a possible implementation, the first feature extraction network is a residual network. During the process of extracting features from the target remote sensing image, the residual network can retain the detailed information of the target remote sensing image, has strong image detail detection ability, and can more accurately capture the local texture features and local shape features of the target remote sensing image in the spatial domain.
[0077] Among them, the residual network includes a convolution module and multiple cascaded residual blocks.
[0078] Based on the fact that the first feature extraction network is a residual network, and the residual network includes a convolution module and multiple cascaded residual blocks, S304 may include: S3041, through the convolution module, extract features from the target remote sensing image to obtain a convolution feature map of the target remote sensing image; S3042, through multiple cascaded residual blocks, perform multiple times of downsampling on the convolution feature map to obtain a spatial domain feature map. Thus, first perform preliminary feature extraction through the convolution module, and then perform multiple times of downsampling through multiple cascaded residual blocks to extract more in-depth local detail features and obtain a rich and accurate spatial domain feature map.
[0079] In S3041, the convolution module may include a convolution layer, and the convolution layer in the convolution module can be used to extract features from the target remote sensing image to obtain a convolution feature map of the target remote sensing image.
[0080] Optionally, the convolution module may further include a batch normalization (BN) layer and an activation function. In the convolution module, perform convolution processing on the target remote sensing image through the convolution layer to obtain a feature map output by the convolution layer; input the feature map output by the convolution layer into the BN layer, and perform normalization processing on the feature map through the BN layer to obtain a feature map output by the BN layer; input the feature map output by the BN layer into the activation function, and after the processing of the activation function, obtain a convolution feature map of the target remote sensing image.
[0081] Introducing the BN layer and the activation function outside the convolution layer can bring many advantages. For example, the BN layer can improve the convergence speed of the classification model during the training process, prevent gradient problems (gradient explosion, gradient disappearance), and prevent overfitting. Another example is that the activation function can introduce a non-linear relationship, enhancing the learning ability of the classification model for the spatial domain features of the target remote sensing image.
[0082] Exemplarily, the convolution processing of the target remote sensing image through the convolution layer can be expressed as: ; Among them, represents the feature map output by the convolution layer, represents the convolution kernel parameters in the convolution layer, represents the bias parameter in the convolution layer, and Conv represents the convolution layer.
[0083] Exemplarily, the normalization processing of the feature map through the BN layer can be expressed as: ; where represents the feature map output by the BN layer, and BN represents the BN layer.
[0084] Furthermore, the activation function adopts the rectified linear unit (ReLU) activation function, and the feature map output by the BN layer is processed through the ReLU activation function.
[0085] Exemplarily, the processing of the feature map output by the BN layer through the ReLU activation function can be expressed as: ; The ReLU activation function can be expressed as: ; where a is a constant, for example, a is 0. In the ReLU activation function, for x with a value greater than a, the value of x is output, and for x with a value less than or equal to a, a is output. represents the feature map output by the ReLU activation function, that is, the convolutional feature map output by the convolutional module.
[0086] Exemplarily, in the convolutional module, the convolutional kernel size of the convolutional layer is 7×7, the stride adopted by the convolutional layer (i.e., the stride size when the convolutional kernel slides on the target remote sensing image) is 2, and the padding value adopted by the convolutional layer is 3. After feature extraction of the target remote sensing image with a scale size of 256×256×3 through this convolutional module, a convolutional feature map with a scale size of 128×128×64 can be obtained.
[0087] In S3042, in multiple cascaded residual blocks, the structure of each residual block is the same. The convolutional feature map can be input into the first residual block of multiple cascaded residual blocks. Starting from the first residual block, the convolutional feature map is downsampled, and through multiple cascaded residual blocks, the convolutional feature map is downsampled multiple times to obtain a spatial domain feature map.
[0088] Optionally, the residual network further includes a max pooling layer. The max pooling layer is located between the convolutional module and the first residual block of multiple cascaded residual blocks. After obtaining the convolutional feature map, the convolutional feature map can be input into the max pooling layer for max pooling operation to obtain the feature map output by the max pooling layer, so as to reduce the spatial size of the feature map; through multiple cascaded residual blocks, the feature map output by the max pooling layer is downsampled multiple times to obtain a spatial domain feature map.
[0089] Exemplarily, the pooling operation of the max pooling layer can be expressed as: ; where represents the feature map output by the max pooling layer, and MaxPool represents the max pooling layer.
[0090] Exemplarily, the size of the max pooling layer is 3×3, the stride adopted by the max pooling layer is 2, and the padding value adopted is 1. After performing the max pooling operation on the convolutional feature map with a size of 128×128×64, a feature map with a size of 64×64×64 output by the max pooling layer is obtained.
[0091] Optionally, performing multiple times of downsampling on the convolutional feature map through multiple cascaded residual blocks may include: in the i-th residual block, first, through the first convolutional layer in the i-th residual block, perform feature extraction on the feature map input to the i-th residual block to obtain the feature map output by the first convolutional layer, where i is greater than or equal to 1; input the feature map output by the first convolutional layer into the first BN layer in the i-th residual block, and perform normalization processing on the feature map input to the first BN layer through the first BN layer to obtain the feature map output by the first BN layer; input the feature map output by the first BN layer into the first activation function in the i-th residual block to obtain the feature map output by the first activation function; input the feature map output by the first activation function into the second convolutional layer in the i-th residual block, and perform feature extraction on the feature map input to the second convolutional layer through the second convolutional layer to obtain the feature map output by the second convolutional layer; input the feature map output by the second convolutional layer into the second BN layer in the i-th residual block, and perform normalization processing on the feature map input to the second BN layer through the second BN layer to obtain the feature map output by the second BN layer; input the feature map input to the i-th residual block and the feature map output by the second BN layer into the second activation function in the i-th residual block to obtain the feature map output by the second activation function, and the feature map output by the second activation function is the feature map output by the i-th residual block. Thus, through multiple convolutional layers, multiple BN layers, and multiple activation functions, multiple times of sampling of the feature map are realized to extract more and more accurate local detail features in the target remote sensing image.
[0092] Exemplarily, taking the last residual block as an example, its processing flow can be expressed as: ; ; ; Among them, Conv1, BN1, ReLU1, Conv2, BN2, and ReLU2 respectively represent the first convolutional layer, the first BN layer, the first activation function, the second convolutional layer, the second BN layer, and the second activation function in the last residual block. represents the feature map output by the first activation function. represents the feature map output by the second BN layer. represents the feature map output by the last residual block. The feature map output by the last residual block is the spatial domain feature map.
[0093] Exemplarily, the residual network includes 16 cascaded residual blocks, each residual block contains 2 convolutional layers with a size of 3×3, and performs 8-fold downsampling on the input feature map. Finally, a spatial domain feature map with a size of 3×3×512 can be obtained.
[0094] S305, through the second feature extraction network in the classification model, extracts features from the frequency domain image to obtain a frequency domain feature map.
[0095] In this embodiment, the second feature extraction network includes multiple network layers. After inputting the frequency domain image into the classification model, through the multiple network layers in the second feature extraction network, features are extracted from the frequency domain image to obtain a frequency domain feature map. Based on the frequency domain feature map, the classification model can improve the recognition ability of the ground object area in remote sensing images with complex backgrounds, low contrast, high noise, and small differences between ground object classes.
[0096] In a possible implementation manner, the second feature extraction network is a two-layer routing attention network. The two-layer routing attention network includes a convolutional module and an attention module. The attention mechanism of the attention module adopts a two-layer routing attention (Bi-Level Routing Attention, BRA) mechanism.
[0097] The two-layer routing attention network refers to a Vision Transformer with Bi-Level Routing Attention (BiFormer) network that combines two-layer routing attention. During the process of extracting features from the frequency domain image through the two-layer routing attention network, the two-layer routing attention mechanism can be used to model the long-range spatial dependence relationship between frequency domain features (especially between global frequency domain features and local frequency domain features), solving the problem that it is difficult for deep neural networks to model the long-range spatial dependence relationship between features. At the same time, compared with the traditional attention mechanism, the two-layer routing attention mechanism has higher computational efficiency and lower computational complexity, solving the problem that the computational complexity is too high after introducing the attention mechanism in the deep neural network, resulting in difficult training of the deep neural network.
[0098] Based on the second feature extraction network being a double - layer routing attention network, the double - layer routing attention network includes a convolution module and an attention module adopting a double - layer routing attention mechanism. S305 may include: S3051, through the convolution module, extract features from the frequency - domain image to obtain a convolution feature map of the frequency - domain image; S3052, through the attention module, perform cross - regional feature interaction on the convolution feature map based on the double - layer routing attention mechanism to obtain a first attention feature map; S3053, based on the first attention feature map, obtain the frequency - domain feature map. Thus, during the frequency - domain feature extraction process, using the double - layer routing attention mechanism in the double - layer routing attention network, establish long - distance spatial dependence relationships between frequency - domain features, extract more accurate frequency - domain features (such as periodic texture features of large - scale farmland), and at the same time, based on the attention mechanism, focus on the key features contained in the high - frequency components of the frequency - domain image (such as road edge features contained in the high - frequency components).
[0099] In S3051, the convolution module may include a convolution layer, and the convolution layer in the convolution module can be used to extract features from the frequency - domain image to obtain a convolution feature map of the frequency - domain image.
[0100] Optionally, the convolution module includes a convolution layer and a feature splicing layer, and the position encoding matrix in the feature splicing layer is a learning parameter during the training process. S3051 includes: dividing the frequency - domain image into blocks to obtain a plurality of non - overlapping image blocks; through the convolution layer, respectively extract features from the plurality of non - overlapping image blocks to obtain feature maps corresponding to the plurality of non - overlapping image blocks respectively; in the feature splicing layer, sort the feature maps corresponding to the plurality of non - overlapping image blocks according to the position order of the plurality of non - overlapping image blocks, and add position information to the sorted feature maps according to the position encoding matrix to obtain a convolution feature map.
[0101] In this optional method, the frequency-domain image can be segmented into multiple non-overlapping image patches through image segmentation; the multiple non-overlapping image patches are respectively input into the convolutional layer, and in the convolutional layer, the multiple non-overlapping image patches can be encoded into embedding vectors through linear projection to obtain feature maps corresponding to the multiple non-overlapping image patches respectively. Thus, by segmenting the frequency-domain image and extracting features from the segmented image patches respectively, the local features of the frequency-domain image can be accurately extracted. Then, in the feature splicing layer, the feature maps corresponding to the multiple non-overlapping image patches can be unfolded in the spatial dimension according to the position order of the multiple non-overlapping image patches to realize the spatial sorting of the feature maps corresponding to the multiple non-overlapping image patches, so that the feature maps corresponding to the multiple non-overlapping image patches are unfolded into an image sequence, and this image sequence forms a feature map; the position encoding matrix contains the position encodings corresponding to the multiple non-overlapping image patches respectively, and position encoding is added to the sorted feature map according to the position encoding matrix to enhance the second feature extraction network's perception of position information, solving the problem that it is difficult for a deep neural network to retain accurate spatial position information when focusing on frequency-domain features.
[0102] During the training process of the classification model, the position encoding matrix in the feature splicing layer is adjusted to improve the accuracy of the position encoding matrix and enable the second feature extraction network to perceive accurate position information.
[0103] Exemplarily, the processing process of the frequency-domain image in the convolutional module can be expressed as: ; ; where Conv represents the convolutional layer, represents the frequency-domain image. In this formula, can be first divided into multiple non-overlapping image patches; represents the feature maps corresponding to the multiple non-overlapping image patches respectively; Flatten() represents unfolding into a sequence according to the spatial dimension, represents the position encoding matrix, represents the convolutional feature map obtained after the frequency-domain image is processed by the convolutional layer and the feature splicing layer.
[0104] Exemplarily, the size of the frequency-domain image is 256×256×3. The frequency-domain image is segmented into non-overlapping blocks of size 4×4, and 64×64 non-overlapping blocks can be obtained. Through a convolutional layer with a convolutional kernel size of 4×4, these non-overlapping blocks are converted into embedding vectors in a linear projection manner to obtain the feature maps corresponding to these non-overlapping blocks respectively. Then, these feature maps are unfolded into a sequence in the spatial dimension and the corresponding position encoding is added to obtain a convolutional feature map with a size of 64×64×64.
[0105] In S3052, in the attention module, based on the double-layer routing attention mechanism, the query vector, key vector, and value vector of the local features in the convolutional feature map of the frequency-domain image can be determined, and feature processing is performed based on the query vector, key vector, and value vector of the local features to achieve cross-region feature interaction based on the double-layer routing attention mechanism, and the first attention feature map is obtained.
[0106] Optionally, the attention module includes an attention layer, and S3052 includes: dividing the convolutional feature map of the frequency-domain image into multiple feature windows; determining the double-layer routing between the multiple feature windows according to the feature similarity between the multiple feature windows pairwise, where the double-layer routing indicates the associated feature windows corresponding to the multiple feature windows respectively in the dimension of feature similarity; determining the query vector corresponding to each of the multiple feature windows, the key vector corresponding to each of the multiple feature windows, and the value vector corresponding to each of the multiple feature windows according to the double-layer routing and the feature representations corresponding to the multiple feature windows on the convolutional feature map; in the attention layer, determining the attention weight matrix according to the query vector corresponding to each of the multiple feature windows, the key vector corresponding to each of the multiple feature windows, the value vector corresponding to each of the multiple feature windows, and the weight parameters of the attention layer, and adding the attention weight matrix to the convolutional feature map to obtain the first attention feature map. It can be seen that in the double-layer routing attention mechanism, through the local feature windows and the attention vectors of the local feature windows, the focusing and capturing of high-frequency details (such as road cracks and building outlines) are realized, avoiding the redundancy of global calculations; by determining the correlation between the feature windows based on the feature similarity between the feature windows, global sparse attention is established to achieve the attention to key regions in the low-frequency components (such as large areas of farmland and water areas). In this way, the calculation of the cross-region dependence relationship between the local frequency-domain features is realized, and the calculation complexity is reduced, enabling the second feature extraction network to focus on the boundaries and contours of ground objects, and improving the accuracy, robustness, and stability of the classification model for classifying ground objects in remote sensing images in complex scenarios.
[0107] In this optional approach, after obtaining multiple feature windows, the similarities between the feature representations corresponding to the multiple feature windows on the convolutional feature map can be calculated to obtain the feature similarities between every two of the multiple feature windows. For the first feature window, the top K feature windows can be selected from the second feature windows as the associated feature windows corresponding to the first feature window in the order of decreasing feature similarity between the first feature window and the second feature windows. The first feature window is any one of the multiple feature windows, and the second feature windows are the remaining feature windows among the multiple feature windows except the first feature window. In this way, a two-layer routing between every two of the multiple feature windows is obtained. Then, for each feature window, the query vector, key vector, and value vector corresponding to the feature window can be calculated according to the features covered by the feature window on the convolutional feature map and the features covered by the associated feature windows of the feature window. In this way, the query vectors corresponding to the multiple feature windows, the key vectors corresponding to the multiple feature windows, and the value vectors corresponding to the multiple feature windows are obtained. Finally, according to the query vectors corresponding to the multiple feature windows, the key vectors corresponding to the multiple feature windows, the value vectors corresponding to the multiple feature windows, and the weight parameters of the attention layer, the attention weight matrix is determined, and the attention weight matrix is added to the convolutional feature map to obtain the first attention feature map.
[0108] Further, the attention module further includes layer normalization (LN) and a feed-forward neural network. After obtaining the first attention feature map, it further includes: preliminarily processing the first attention feature map through layer normalization; enhancing the preliminarily processed attention feature map through a feed-forward neural network (abbreviated as FFN, composed of fully connected layers). Thus, the problem of gradient explosion caused by the dot product operation in the attention mechanism is alleviated through layer normalization, and the feature representation of the attention feature map is optimized through the feed-forward neural network.
[0109] Exemplarily, the above processing process can be expressed as: ; ; ; ; ; Among them, and respectively represent the feature representation of the i-th feature window and the feature representation of the j-th feature window, Sim() represents calculating the feature similarity between the feature windows, represents and The feature similarity between; TopK() represents selecting the top k feature windows with the highest feature similarity, represents the associated feature window corresponding to the i-th feature window; , , respectively represent the query vector, key vector, and value vector corresponding to the i-th feature window, Softmax represents the normalized exponential function, and d represents the dimension of the query vector, represents the attention weight matrix of the i-th feature window; LN represents layer normalization, represents the first attention feature map after layer normalization processing, represents the first attention feature map enhanced by a feed-forward neural network.
[0110] Exemplarily, a convolutional feature map with a size of 64×64×64 is divided into feature windows with a size of 8×8, and 64 feature windows can be obtained. The above attention operations are performed on the 64 feature windows respectively, and finally, a first attention feature map with a size of 64×64×64 can be obtained.
[0111] In S3053, the first attention feature map can be determined as the frequency domain feature map, or the first attention feature map can be further processed to obtain the frequency domain feature map. In particular, cross-region feature interaction based on a double-layer routing attention mechanism can be performed on the first attention feature map to obtain a second attention feature map, and cross-region feature interaction based on a double-layer routing attention mechanism can also be performed on the second attention feature map to obtain a third attention feature map. Through repeated attention operations, deep cross-region feature interaction is established to improve the extraction effect of the frequency domain features.
[0112] Optionally, S3053 includes: adjusting the dimension of the first attention feature map; through an attention module, performing cross-region feature interaction based on a double-layer routing attention mechanism on the first attention feature map after dimension adjustment to obtain a second attention feature map; obtaining the frequency domain feature map according to the second attention feature map. Obtaining the frequency domain feature map according to the second attention feature map can include: determining the second attention feature map as the frequency domain feature map; or adjusting the dimension of the second attention; through an attention module, performing cross-region feature interaction based on a double-layer routing attention mechanism on the second attention feature map after dimension adjustment to obtain a third attention feature map, and adjusting the dimension of the third attention feature map to obtain the frequency domain feature map. In this way, through repeated dimension adjustment and attention operations, deep cross-region feature interaction is performed to improve the extraction effect of the frequency domain features.
[0113] In this optional method, multiple adjacent feature blocks on the same channel dimension in the first attention feature map can be merged, and the number of channels of the first attention feature map can be adjusted to achieve dimensional adjustment of the first attention feature map. The dimensional adjustment of the second attention feature map and the cross-region feature interaction of the second feature map can refer to the description of the first attention feature map above and will not be elaborated here.
[0114] Optionally, the frequency domain feature map and the spatial domain feature map have the same size to facilitate the fusion of the frequency domain feature map and the spatial domain feature map.
[0115] Exemplarily, for the first attention feature map with a size of 64×64×64 , adjacent 2×2 feature blocks are merged on the same channel dimension, making the size of the first attention feature map become 32*32*64; then, the number of channels of the first attention feature map is adjusted to 256, making the size of the first attention feature map become 32*32*256; then, through linear transformation, the number of channels of the first attention feature map is projected to 128 to reduce the output of redundant information, and the final size of the first attention feature map is 32*32*128. Cross-region feature interaction based on the double-layer routing attention mechanism is performed on the dimension-adjusted first attention feature map to obtain the second attention feature map , and then a dimensional adjustment operation is performed on the second attention feature map to obtain a second attention feature map with a size of 16*16*256 ; cross-region feature interaction based on the double-layer routing attention mechanism is performed on the dimension-adjusted second attention feature map to obtain the third attention feature map , and then a dimensional adjustment operation is performed on the third attention feature map to obtain a frequency domain feature map with a size of 8*8*512 .
[0116] S306. Through the feature fusion network in the classification model, the spatial domain feature map and the frequency domain feature map are fused to obtain a fused feature map.
[0117] In this embodiment, if the frequency-domain features and spatial features are simply concatenated, it will be impossible to establish deep associations between cross-domain features, and there is a lack of targeted modeling of high-order semantics in the frequency domain. To solve this problem, a feature fusion network is used to fuse the spatial-domain feature map and the frequency-domain feature map. During the training process, the feature fusion network can dynamically learn the deep associations between cross-domain features and the high-order semantics in the frequency domain, realize the dynamic fusion of cross-domain features, improve the feature fusion effect, and further improve the classification accuracy of remote sensing images based on the fused features for ground object classification.
[0118] In a possible implementation, S306 includes: In the feature fusion network, dynamic exponential moving average (EMA) is used to fuse the spatial-domain feature map and the frequency-domain feature map to obtain a fused feature map. The smoothing factor in the exponential moving average is related to the training parameters of the feature fusion network. During the training process, the training parameters of the feature fusion network can be dynamically adjusted. The change in the weight parameters causes the change in the smoothing factor in the exponential moving average, so that the feature fusion network can dynamically adjust the feature fusion effect during the training process, and further improve the fusion effect of the feature fusion network on cross-domain features in the actual application process.
[0119] Optionally, the smoothing factor in the dynamic exponential moving average is expressed as: ; ; where represents the dynamic fusion coefficient, that is, the smoothing factor in the exponential moving average; t represents the training steps of the classification model, represents the spatial-domain feature map extracted during the training process, represents the frequency-domain feature map extracted during the training process; w and b are the training parameters (i.e., learnable parameters) of the feature fusion network, where w is the weight parameter used to control the adjustment rate of the feature fusion network during the training process, and b is the bias parameter used to control the offset of the adjustment of the feature fusion network during the training process; Sigmoid represents the activation function used to map the variable value input to the activation function to the numerical interval from 0 to 1. It can be seen that during the training process, changes with the training steps t (for example, it biases towards the frequency-domain feature map during the initial few training processes to realize the dynamic update of; it biases towards the spatial-domain feature map during the later training process), and is a fixed value during the model verification and testing phases.
[0120] S307, through the classification network in the classification model, classify the fused feature map to obtain the ground object classification result of the target remote sensing image.
[0121] In this embodiment, the fused feature map is input into the classification network. In the classification network, the fused feature map can be mapped to the category space to obtain scores corresponding to multiple ground object types respectively. The scores corresponding to multiple ground object types are converted into probability values corresponding to multiple ground object types respectively. The ground object classification result can be determined according to the probability values corresponding to multiple ground object types respectively. For example, the ground object type with the largest probability value is determined as the ground object type corresponding to the remote sensing image.
[0122] Optionally, the classification network includes a global pooling layer, a fully connected layer, and a softmax function. First, the global pooling layer can be used to compress the spatial dimension of the fused feature map and retain the channel information of the fused feature map to obtain the feature map output by the global pooling layer. Then, the fully connected layer is used to map the feature map output by the global pooling layer to the category space to obtain scores corresponding to multiple ground object types respectively. After that, through the softmax function, the scores corresponding to multiple ground object types are converted into probabilities corresponding to multiple ground object types respectively. The formula for this process can be expressed as: ; ; ; ; where GlobalAvgPool represents the global pooling layer, represents the fused feature map, represents the feature map output by the global pooling layer; W and b represent the training parameters of the fully connected layer, Z represents the scores output by the fully connected layer; Softmax represents the softmax function, P represents the probability distribution, and the probabilities corresponding to multiple ground object types in the probability distribution, represents the probability corresponding to the first ground object type, represents the probability corresponding to the Cth ground object type.
[0123] In the embodiment of the present application, a classification model is used to implement the ground object classification of remote sensing images based on multi-modal features. In particular, in the classification model, the extraction effect of spatial domain features is improved through the residual network, the extraction effect of frequency domain features is improved through the double-layer routing attention network, and the fusion effect of spatial domain features and frequency domain features is improved through dynamic exponential moving average, which improves the accuracy, robustness, and stability of the classification model for ground object classification of remote sensing images from multiple aspects.
[0124] Figure 4 is an example diagram of ground object classification of remote sensing images by the classification model. As Figure 4As shown in the figure, the remote sensing image is input into the residual network, and the residual network is used to extract features from the remote sensing image to obtain a spatial domain feature map; the remote sensing image is subjected to FFT transformation to obtain a frequency domain image, and the frequency domain image is input into the double-layer routing attention network, and the double-layer routing attention network is used to extract features from the frequency domain image to obtain a frequency domain feature map; the spatial domain feature map and the frequency domain feature map are subjected to feature fusion based on exponential moving average to obtain a fusion feature map; finally, the fusion feature map is input into the classification layer for classification to obtain the ground object type corresponding to the remote sensing image.
[0125] Next, an embodiment of the training process of the classification model is provided.
[0126] In some embodiments, the training process of the classification model may include: obtaining a training data set, where the training data set includes training samples and sample labels of the training samples, the training samples are remote sensing images, and the sample labels of the training samples are the true ground object classification results of the remote sensing images; transforming the training samples from the spatial domain to the frequency domain to obtain the frequency domain images corresponding to the training samples; inputting the training samples and the frequency domain images corresponding to the training samples into the classification model; through the first feature extraction network in the classification model, extracting features from the training samples to obtain a spatial domain feature map; through the second feature extraction network in the classification model, extracting features from the frequency domain images corresponding to the training samples to obtain a frequency domain feature map; through the feature fusion network in the classification model, fusing the spatial domain feature map and the frequency domain feature map to obtain a fusion feature map; through the classification network in the classification model, classifying the fusion feature map to obtain the predicted ground object classification result of the training samples, and determining the loss value according to the predicted ground object classification result of the training samples and the sample labels of the training samples (i.e., the true ground object classification results of the training samples); according to the loss value, using an optimization algorithm to adjust the parameters of the classification model.
[0127] In this embodiment, in the process of adjusting the parameters of the classification model using an optimization algorithm according to the loss value, the parameters of the classification model can be adjusted by the backpropagation method. After multiple iterative trainings, the classification model tends to converge, and finally a classification model with better classification effect is obtained.
[0128] It should be noted that the classification process of the classification model for the training samples can refer to the classification process of the classification model for the target remote sensing image, which will not be elaborated here.
[0129] Optionally, obtain a first data set, perform data augmentation on the first data set to obtain a second data set; divide the second data set into a training data set, a validation test set, and a test data set. Among them, the first data set is a publicly available data set, and the first data set includes remote sensing images of multiple ground object types (such as agriculture, aircraft, baseball fields, beaches, buildings, woods, etc.); performing data augmentation on the first data set may include performing one or more operations of rotating, translating, and inverting the images in the first data set to improve the richness of the data in the data set.
[0130] Optionally, the optimization algorithm uses the Adaptive Moment Estimation (Adam) algorithm to improve the optimization effect. During the training process, the hyperparameters of the Adam algorithm can be initialized, the number of iterations and the batch size (i.e., the number of samples) used in each training can be initialized. In order to fully train the classification model to learn multi-modal features, a relatively large number of iterations can be set. Among them, the hyperparameters of the Adam algorithm include the learning rate and the weight decay coefficient.
[0131] Exemplarily, initialize the learning rate to 0.03, the weight decay coefficient to 0.0005, the batch size to 64, and the number of iterations to 10,000.
[0132] Optionally, the predicted ground object classification result of the training sample includes the probability values that the training sample belongs to multiple ground object types respectively, and the sample label of the training sample indicates the actual ground object type to which the training sample belongs. Based on this, the loss function of the classification model is expressed as: ; Among them, represents the probability value that the training sample belongs to the i-th type of ground object in the predicted ground object classification result; if the sample label of the training sample indicates that the training sample actually belongs to the i-th type of ground object, then y i = 1, otherwise y i = 0. C represents the number of ground object types.
[0133] Exemplary device Correspondingly, an embodiment of the present application further provides a ground object classification device for a remote sensing map.
[0134] Please refer to Figure 5 , in an exemplary embodiment, a ground object classification device 500 for a remote sensing map is provided. The ground object classification device 500 for a remote sensing map includes: an acquisition unit 501, a spatial domain frequency domain transformation unit 502, a spatial domain feature extraction unit 503, a frequency domain feature extraction unit 504, a feature fusion unit 505, and a ground object classification unit 506. Among them: An acquisition unit 501 for acquiring a target remote sensing image; a spatial-frequency domain transformation unit 502 for transforming the target remote sensing image from the spatial domain to the frequency domain to obtain a frequency-domain image; a spatial-domain feature extraction unit 503 for extracting features from the target remote sensing image to obtain a spatial-domain feature map; a frequency-domain feature extraction unit 504 for extracting features from the frequency-domain image to obtain a frequency-domain feature map; a feature fusion unit 505 for fusing the spatial-domain feature map and the frequency-domain feature map to obtain a fused feature map; a ground object classification unit 506 for classifying the fused feature map to obtain a ground object classification result of the target remote sensing image.
[0135] In a possible implementation manner, the spatial-domain feature extraction unit 503 is specifically configured to: input the target remote sensing image and the frequency-domain image into a classification model; extract features from the target remote sensing image through a first feature extraction network in the classification model to obtain a spatial-domain feature map. The frequency-domain feature extraction unit 504 is specifically configured to: extract features from the frequency-domain image through a second feature extraction network in the classification model to obtain a frequency-domain feature map. The feature fusion unit 505 is specifically configured to: fuse the spatial-domain feature map and the frequency-domain feature map through a feature fusion network in the classification model to obtain a fused feature map. The ground object classification unit 506 is specifically configured to: classify the fused feature map through a classification network in the classification model to obtain a ground object classification result of the target remote sensing image.
[0136] In a possible implementation manner, the first feature extraction network is a residual network, and the residual network includes a convolution module and a plurality of cascaded residual blocks. The spatial-domain feature extraction unit 503 is specifically configured to: extract features from the target remote sensing image through the convolution module to obtain a convolution feature map of the target remote sensing image; perform multiple downsamplings on the convolution feature map through the plurality of cascaded residual blocks to obtain a spatial-domain feature map.
[0137] In a possible implementation manner, the second feature extraction network is a double-layer routing attention network, and the double-layer routing attention network includes a convolution module and an attention module. The attention mechanism of the attention module adopts a double-layer routing attention mechanism. The frequency-domain feature extraction unit 504 is specifically configured to: extract features from the frequency-domain image through the convolution module to obtain a convolution feature map of the frequency-domain image; perform cross-region feature interaction on the convolution feature map through the attention module based on the double-layer routing attention mechanism to obtain a first attention feature map; obtain a frequency-domain feature map according to the first attention feature map.
[0138] In a possible implementation, the convolution module includes a convolutional layer and a feature concatenation layer. The position encoding matrix in the feature concatenation layer is a learning parameter during the training process. The frequency domain feature extraction unit 504 is specifically configured to: divide the frequency domain image into blocks to obtain a plurality of non-overlapping image blocks; perform feature extraction on the plurality of non-overlapping image blocks respectively through the convolutional layer to obtain feature maps corresponding to the plurality of non-overlapping image blocks respectively; in the feature concatenation layer, sort the feature maps corresponding to the plurality of non-overlapping image blocks according to the position order of the plurality of non-overlapping image blocks, and add position information to the sorted feature maps according to the position encoding matrix to obtain a convolutional feature map.
[0139] In a possible implementation, the attention module includes an attention layer. The frequency domain feature extraction unit 504 is specifically configured to: divide the convolutional feature map into windows to obtain a plurality of feature windows; determine a two-layer routing between the plurality of feature windows according to the feature similarity between the plurality of feature windows in pairs, where the two-layer routing indicates the associated feature windows corresponding to the plurality of feature windows respectively in the dimension of feature similarity; determine query vectors corresponding to the plurality of feature windows respectively, key vectors corresponding to the plurality of feature windows respectively, and value vectors corresponding to the plurality of feature windows respectively according to the two-layer routing and the feature representations corresponding to the plurality of feature windows on the convolutional feature map; in the attention layer, determine an attention weight matrix according to the query vectors corresponding to the plurality of feature windows respectively, the key vectors corresponding to the plurality of feature windows respectively, the value vectors corresponding to the plurality of feature windows respectively, and the weight parameters of the attention layer, and add the attention weight matrix to the convolutional feature map to obtain a first attention feature map.
[0140] In a possible implementation, the attention module further includes layer normalization and a feed-forward neural network. The frequency domain feature extraction unit 504 is further configured to: perform preliminary processing on the first attention feature map through layer normalization; enhance the preliminarily processed attention feature map through the feed-forward neural network.
[0141] In a possible implementation, the frequency domain feature extraction unit 504 is specifically configured to: adjust the dimension of the first attention feature map; perform cross-region feature interaction on the dimension-adjusted first attention feature map through the attention module based on the two-layer routing attention mechanism to obtain a second attention feature map; obtain a frequency domain feature map according to the second attention feature map.
[0142] In a possible implementation, the feature fusion unit 505 is specifically configured to: in the feature fusion network, fuse the spatial domain feature map and the frequency domain feature map by using dynamic exponential moving average to obtain a fused feature map, where the smoothing factor in the exponential moving average is related to the training parameters of the feature fusion network.
[0143] The ground object classification device 500 of the remote sensing map provided in this embodiment belongs to the same inventive concept as the ground object classification method of the remote sensing map provided in the foregoing embodiments of the present application, and can execute the ground object classification method of the remote sensing map provided in any of the foregoing embodiments of the present application, and has corresponding functional modules and beneficial effects for executing the ground object classification method of the remote sensing map. For technical details not described in detail in this embodiment, reference may be made to the specific processing content of the ground object classification method of the remote sensing map provided in the foregoing embodiments of the present application, which will not be elaborated here.
[0144] The functions implemented by each unit in the above device can be implemented by the same or different processors respectively, which is not limited in the embodiments of the present application.
[0145] It should be understood that the units in the above device can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or the functions of each unit of the device. The processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory inside the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of a hardware circuit, and the functions of some or all of the units can be implemented through the design of the hardware circuit. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are implemented through the design of the logical relationship of the components in the circuit. Again, for example, in another implementation, the hardware circuit can be implemented by a PLD. Taking an FPGA as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured through a configuration file, so as to implement the functions of some or all of the above units. All units of the above device can be all implemented in the form of a processor calling software, or all implemented in the form of a hardware circuit, or some implemented in the form of a processor calling software, and the remaining part implemented in the form of a hardware circuit.
[0146] In the embodiments of the present application, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and running capabilities, such as a CPU, a microprocessor, a GPU, or a DSP, etc. In another implementation, the processor can implement certain functions through the logical relationship of the hardware circuit, and the logical relationship of the hardware circuit is fixed or can be reconstructed. For example, the processor is a hardware circuit implemented by an ASIC or a PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a kind of ASIC, such as an NPU, a TPU, a DPU, etc.
[0147] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method. For example: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0148] In addition, each unit in the above device can be integrated in whole or in part, or can be independently implemented. In one implementation, these units are integrated together and implemented in the form of an SOC. The SOC can include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The types of the at least one processor can be different. For example, it includes CPU and FPGA, CPU and artificial intelligence processor, CPU and GPU, etc.
[0149] Exemplary electronic device Another embodiment of the present application also proposes an electronic device. Refer to Figure 6 As shown, the electronic device may include: a memory 600 and a processor 610; wherein, the memory 600 is connected to the processor 610 and is used for storing programs; the processor 610 is used for implementing the ground object classification method of the remote sensing map disclosed in any of the above embodiments by running the programs stored in the memory 600.
[0150] Specifically, the above electronic device may further include: a bus, a communication interface 620, an input device 630, and an output device 640.
[0151] The processor 610, the memory 600, the communication interface 620, the input device 630, and the output device 640 are interconnected through the bus. Among them: The bus may include a path for transmitting information between various components of the computer system.
[0152] The processor 610 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the solution of the present application. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0153] The processor 610 may include a main processor and may also include a baseband chip, a modem, etc.
[0154] The program for implementing the technical solution of this application is stored in the memory 600, and the operating system and other key services can also be stored. Specifically, the program can include program code, and the program code includes computer operation instructions. More specifically, the memory 600 can include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, etc.
[0155] The input device 630 can include devices for receiving data and information input by the user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.
[0156] The output device 640 can include devices for allowing information to be output to the user, such as a display screen, a printer, a speaker, etc.
[0157] The communication interface 620 can include any device of a transceiver type to communicate with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0158] The processor 610 executes the program stored in the memory 600 and calls other devices, and can be used to implement each step of any one of the ground object classification methods of the remote sensing map provided in the above embodiments of this application.
[0159] An embodiment of this application also proposes a chip, which includes a processor and a data interface. The processor reads and runs the program stored on the memory through the data interface to execute any one of the ground object classification methods of the remote sensing map provided in the above embodiments. For the specific processing process and its beneficial effects, reference can be made to the embodiment introduction of the ground object classification method of the remote sensing map above.
[0160] Exemplary computer program product and storage medium In addition to the above methods and devices, an embodiment of this application can also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the ground object classification method of the remote sensing map according to various embodiments of this application described in any of the above embodiments of this specification.
[0161] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0162] In addition, an embodiment of the present application can also be a storage medium on which a computer program is stored, and the computer program is executed by a processor to perform the steps in the method for classifying ground objects in a remote sensing map according to various embodiments of the present application described in any of the above embodiments of this specification.
[0163] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0164] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0165] The steps in the methods of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.
[0166] The modules and sub-modules in the devices and terminals in the embodiments of the present application can be combined, divided, and deleted according to actual needs.
[0167] In several embodiments provided by the present application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.
[0168] The modules or sub-modules described as separate components may or may not be physically separated. The components as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place, or they can be distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0169] In addition, each functional module or sub-module in various embodiments of the present application can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above-mentioned integrated modules or sub-modules can be implemented in the form of hardware or in the form of software functional modules or sub-modules.
[0170] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0171] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software units executed by a processor, or a combination of the two. The software units can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0172] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0173] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for land cover classification of remote sensing images, characterized in that, Including: Obtain a target remote sensing image; Transform the target remote sensing image from the spatial domain to the frequency domain to obtain a frequency domain image; Extract features from the target remote sensing image to obtain a spatial domain feature map; Extract features from the frequency domain image to obtain a frequency domain feature map; Fuse the spatial domain feature map and the frequency domain feature map to obtain a fused feature map; Classify the fused feature map to obtain the ground object classification result of the target remote sensing image.
2. The method for classifying ground objects in a remote sensing image according to claim 1, wherein The extracting features from the target remote sensing image to obtain a spatial domain feature map; extracting features from the frequency domain image to obtain a frequency domain feature map; fusing the spatial domain feature map and the frequency domain feature map to obtain a fused feature map; And, classifying the fused feature map to obtain the ground object classification result of the target remote sensing image, including: Input the target remote sensing image and the frequency domain image into a classification model; Extract features from the target remote sensing image through the first feature extraction network in the classification model to obtain the spatial domain feature map; Extract features from the frequency domain image through the second feature extraction network in the classification model to obtain the frequency domain feature map; Fuse the spatial domain feature map and the frequency domain feature map through the feature fusion network in the classification model to obtain the fused feature map; Classify the fused feature map through the classification network in the classification model to obtain the ground object classification result of the target remote sensing image.
3. The method for classifying ground objects in a remote sensing image according to claim 2, wherein, The first feature extraction network is a residual network, and the residual network includes a convolution module and a plurality of cascaded residual blocks. The extracting features from the target remote sensing image through the first feature extraction network in the classification model to obtain the spatial domain feature map includes: Extract features from the target remote sensing image through the convolution module to obtain a convolution feature map of the target remote sensing image; Perform multiple times of downsampling on the convolution feature map through the plurality of cascaded residual blocks to obtain the spatial domain feature map.
4. The method for classifying ground objects in a remote sensing image according to claim 2, characterized in that The second feature extraction network is a double-layer routing attention network. The double-layer routing attention network includes a convolution module and an attention module. The attention mechanism of the attention module adopts a double-layer routing attention mechanism. The extracting features from the frequency domain image through the second feature extraction network in the classification model to obtain the frequency domain feature map includes: Extract features from the frequency domain image through the convolution module to obtain a convolution feature map of the frequency domain image; Perform cross-region feature interaction on the convolution feature map based on the double-layer routing attention mechanism through the attention module to obtain a first attention feature map; Obtain the frequency domain feature map according to the first attention feature map.
5. The method for classifying ground objects in a remote sensing image according to claim 4, characterized in that, The convolution module includes a convolution layer and a feature splicing layer. The position encoding matrix in the feature splicing layer is a learning parameter during the training process. The extracting features from the frequency domain image through the convolution module to obtain the convolution feature map of the frequency domain image includes: Divide the frequency domain image into blocks to obtain a plurality of non-overlapping image blocks; Through the convolutional layer, feature extraction is respectively performed on the multiple non-overlapping image patches to obtain feature maps corresponding to the multiple non-overlapping image patches respectively; In the feature splicing layer, according to the position order of the multiple non-overlapping image patches, the feature maps corresponding to the multiple non-overlapping image patches are sorted, and position information is added to the sorted feature maps according to the position encoding matrix to obtain the convolutional feature map.
6. The method for classifying ground objects in a remote sensing image according to claim 4, wherein, The attention module includes an attention layer. Through the attention module, cross-region feature interaction based on a two-layer routing attention mechanism is performed on the convolutional feature map to obtain a first attention feature map, including: Perform window division on the convolutional feature map to obtain a plurality of feature windows; According to the feature similarity between every two of the plurality of feature windows, determine the two-layer routing between every two of the plurality of feature windows, where the two-layer routing indicates the associated feature windows corresponding to the plurality of feature windows respectively in the dimension of feature similarity; According to the two-layer routing and the feature representations corresponding to the plurality of feature windows on the convolutional feature map, determine the query vectors corresponding to the plurality of feature windows respectively, the key vectors corresponding to the plurality of feature windows respectively, and the value vectors corresponding to the plurality of feature windows respectively; In the attention layer, according to the query vectors corresponding to the plurality of feature windows respectively, the key vectors corresponding to the plurality of feature windows respectively, the value vectors corresponding to the plurality of feature windows respectively, and the weight parameters of the attention layer, determine the attention weight matrix, and add the attention weight matrix to the convolutional feature map to obtain the first attention feature map.
7. The method for classifying ground objects in a remote sensing image according to claim 6, characterized in that, The attention module further includes layer normalization and a feed-forward neural network. After obtaining the first attention feature map by performing cross-region feature interaction based on a two-layer routing attention mechanism on the convolutional feature map through the attention module, it further includes: Perform preliminary processing on the first attention feature map through the layer normalization; Enhance the attentional feature map after preliminary processing through the feed-forward neural network.
8. The method for classifying ground objects in a remote sensing image according to claim 4, characterized in that, The obtaining of the frequency domain feature map according to the first attention feature map includes: Perform dimension adjustment on the first attention feature map; Through the attention module, perform cross-region feature interaction based on a two-layer routing attention mechanism on the first attention feature map after dimension adjustment to obtain a second attention feature map; Obtain the frequency domain feature map according to the second attention feature map.
9. The method for classifying ground objects in a remote sensing image according to any one of claims 2 to 8, characterized in that, The obtaining of the fusion feature map by fusing the spatial domain feature map and the frequency domain feature map through the feature fusion network in the classification model includes: In the feature fusion network, adopt dynamic exponential moving average to fuse the spatial domain feature map and the frequency domain feature map to obtain the fusion feature map, where the smoothing factor in the exponential moving average is related to the training parameters of the feature fusion network.
10. An electronic device, characterized in that, Includes a memory and a processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the method for classifying ground objects in remote sensing images as described in any one of claims 1 to 9 by running the program in the memory.
11. A computer program product, characterized in that, It includes a computer program which, when executed by a processor, implements the method for classifying ground objects in remote sensing images as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Feature extraction method and device and electronic system
CN112883983A
Hyperspectral image classification method based on multidirectional dynamic routing
CN118674997A
Hyperspectral image classification method
CN119169399A
Multi-modal remote sensing image classification method and device, electronic equipment and medium
CN119600330A
Remote sensing image directed target detection method based on double-domain feature fusion
CN119919819A