Remote Sensing Scene Classification Method and Device Based on Visual Marker Clustering and Discarding
Through the clustering and discarding methods of visual markers, the remote sensing image classification model is optimized, and the problem of insufficient computing efficiency and accuracy in the existing technology is solved, and efficient and accurate remote sensing image classification is achieved.
Patent Information
- Application Number
- CN202510542050.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing remote sensing image classification model has shortcomings in taking into account both computing efficiency and accuracy, and it is impossible to achieve high efficiency and high accuracy at the same time.
The clustering and discarding method based on visual markers is adopted, and the clustering and discarding of visual markers is performed by constructing a remote sensing image classification model network architecture, using multi-scale feature extraction and fusion module, feature enhancement aggregation module and Transformer module, clustering and discarding of visual markers, including the calculation of the mean value in the Li group, the evaluation of the spatial distance of the Li group manifold and the importance of visual markers, and the attention mechanism is optimized to selectively discard unimportant visual markers.
It improves the calculation efficiency and accuracy of remote sensing image classification, effectively captures the geometric shape and semantic features of the scene, reduces the complexity of the model, and improves the classification performance.
Smart Images

Figure CN120071026B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing scene classification, and particularly relates to a remote sensing scene classification method and device based on clustering and discarding of visual tokens. Background Art
[0002] High-resolution remote sensing images (HRRSIs) can cover a wider scene range and finely present the texture structure, geometric shape, and spatial layout in the scene. Remote sensing scene classification (RSSC) based on HRRSIs is widely applied in fields such as urban scene analysis. However, due to the existence of scenes with similar geometric structures and spatial layouts in high-resolution remote sensing images, and at the same time containing different socioeconomic characteristics, these complexities pose great challenges to scene classification.
[0003] With the rapid development of deep learning technology, convolutional neural networks (CNNs) have achieved remarkable results in tasks such as image classification, segmentation, and object detection. CNNs have gradually become one of the mainstream methods for remote sensing image classification. However, these models have limitations in capturing global context information and long-range dependency relationships, especially when dealing with HRRSIs with complex backgrounds. To solve these problems, scholars have introduced Transformer into the remote sensing scene classification model. Using the hierarchical attention mechanism of Swin Transformer to process multi-scale features further improves the performance of remote sensing image classification.
[0004] Although the above models have improved in classification accuracy, some Vision Transformer models divide the image into several uniform and fixed-size blocks and generate corresponding visual tokens (Tokens) for them. The model treats the visual tokens of all regions equally, which results in a high time complexity of the model. Therefore, the existing models cannot simultaneously achieve high computational efficiency and high accuracy when performing remote sensing scene classification. Summary of the Invention
[0005] In view of this, an object of the present invention is to provide a remote sensing scene classification method and device based on clustering and discarding of visual tokens, aiming to solve the problem in the prior art that high computational efficiency and high accuracy cannot be achieved simultaneously when performing remote sensing image classification.
[0006] An object of the present invention is to provide a remote sensing scene classification method based on clustering and discarding of visual tokens, the method comprising:
[0007] Constructing a network architecture of a remote sensing image classification model;
[0008] Training the network architecture of the remote sensing image classification model to obtain a pre-trained remote sensing image classification model;
[0009] Obtain the remote sensing image to be classified, and input the remote sensing image to be classified into the pre-trained remote sensing image classification model to obtain the final classification result;
[0010] Among them, the network architecture of the remote sensing image classification model at least includes a multi-scale feature extraction and fusion module, a feature enhancement aggregation module, and a Transformer module connected in sequence. The Transformer module is used to cluster the visual tokens into fixed clusters when receiving visual tokens, and then merge the visual tokens within the same cluster into one visual token, so as to discard the corresponding visual tokens according to the importance degree of the visual tokens.
[0011] Furthermore, in the above remote sensing scene classification method based on clustering and discarding of visual tokens, the step of clustering the visual tokens into fixed clusters includes:
[0012] Calculate the local density ρ of each visual token based on the inner mean value in the Lie group of the received visual tokens:
[0013] ;
[0014] Among them, C represents the visual token, represents the inner mean value in the Lie group of the adjacent visual tokens of the i-th visual token, and represent the features of the i-th and j-th visual tokens respectively;
[0015] Determine the target visual token closest to the target as the clustering center according to the local density of the visual token and the distance in the Lie group manifold space;
[0016] And assign the other visual tokens to the closest clustering center according to the distance in the Lie group manifold space to form fixed clusters;
[0017] Among them, the calculation formula of the distance in the Lie group manifold space is:
[0018] ;
[0019] Among them, and represent the features of the i-th and j-th visual tokens respectively, represents the local density of the i-th visual token, ρ j represents the local density of the j-th visual token.
[0020] Furthermore, in the above remote sensing scene classification method based on clustering and discarding of visual tokens, the formula for merging the visual tokens within the same cluster into one visual token is:
[0021] ;
[0022] Among them, represents the cluster formed by the i-th visual marker, and respectively represent the features of the j-th visual marker and the importance degree of the corresponding in-group mean in the Lie group, m i represents the features of the visual marker after the clusters formed by the i-th visual marker are merged.
[0023] Furthermore, for the above remote sensing scene classification method based on clustering and discarding of visual markers, among them, the calculation formula for the importance degree of visual markers is as follows:
[0024] ;
[0025] Among them, represents the importance degree of the visual marker. The numerator in the formula represents the attention feature map, which is the click of the query vector obtained from the classified visual marker and the Keys value matrix, represents the classified visual marker, d =L / H, where L represents the vector length into which each visual marker is embedded, H represents the number of heads in the multi-head attention mechanism in the Transformer module, K is a fixed value set by the discard rate of the visual marker, and T represents the matrix transpose operation.
[0026] Furthermore, for the above remote sensing scene classification method based on clustering and discarding of visual markers, among them, the multi-scale feature extraction and fusion module is divided into shallow feature extraction and deep feature extraction;
[0027] The process of shallow feature extraction is as follows:
[0028] ;
[0029] Among them, represents the spatial position of the target object in the image, represents the color space information of the image, is used to extract the local gradient direction of the image, 3D-GMK(x,y) represents the spatial-spectral feature, is used to describe the spatial correlation between pixels, calculates the shortest distance between different target objects and is used to characterize the spatial relationship between different objects. T represents the matrix transpose operation;
[0030] The process of deep feature extraction is as follows:
[0031] Perform standard convolution operations on the received feature map to capture local spatial features;
[0032] After channel compression of the feature map obtained by standard convolution operation using 1×1 convolution, through adaptive pooling processing, the feature map is adjusted to a unified size;
[0033] The feature map processed by adaptive pooling is subjected to batch normalization and then the PeLK operation. The process of the PeLK operation is as follows:
[0034] ;
[0035] Among them, a represents the position of the feature map in the height (or row) direction, b represents the position of the feature map in the width (or column) direction, z(c,d) is the weight of the convolution kernel, w(c,d) is the position embedding, F is the input feature, r w represents the row, and c and d represent coordinates respectively;
[0036] The features extracted at different scales through shallow feature extraction and high-level feature extraction are used as Queries, and the results obtained by convolving the features extracted at different scales through shallow feature extraction and high-level feature extraction with a dilation rate of 1 are used as Keys and Values;
[0037] And after normalizing Keys and Queries and then fusing them, the result of the fusion is dot-product operated with Values, and a residual connection is made with the result obtained by convolving the features extracted at different scales through shallow feature extraction and high-level feature extraction with a dilation rate of 1 to obtain the multi-scale features that the final multi-scale feature extraction and fusion module needs to output.
[0038] Furthermore, in the above remote sensing scene classification method based on visual marker clustering and discarding, among them, the feature enhancement aggregation module includes parallel spatial feature enhancement aggregation and channel feature enhancement aggregation;
[0039] The process of spatial feature enhancement aggregation is:
[0040] The input feature map is subjected to batch normalization to standardize the feature map;
[0041] The feature map after batch normalization is processed through an adaptive pooling layer;
[0042] The pooled feature map is processed through a 1×1 convolutional layer;
[0043] The Lie group activation function is introduced to activate the feature map after convolution processing, and then the feature map is flattened. The features of each channel are assigned corresponding weights, so as to convert the multi-dimensional features into one-dimensional vectors;
[0044] The process of channel feature enhancement aggregation is:
[0045] Process the input feature map through global average pooling:
[0046] ;
[0047] Among them, represents the global average pooling result of the c-th channel, where H and W are the height and width of the feature map respectively, represents the pixel value of the input feature map on the c-th channel, and i and j represent the row and column values of the feature matrix respectively;
[0048] Extract and fuse the inter-channel features of the global average pooling result through an independent 1×1 convolutional layer;
[0049] Perform an addition operation on the features of all channels to complete the fusion, then perform non-linear processing through a convolutional layer and a SeLU activation function, and perform channel compression through a 1×1 convolution again to obtain the final fused feature map: ;
[0050] Among them, c represents the channel, is the fused feature map after fusion, represents the sum of the convolution results of all channels.
[0051] Furthermore, for the above remote sensing scene classification method based on visual token clustering and discarding, among them, the method further includes:
[0052] When the Transformer module receives the input features, perform convolutions with different dilation rates, and fuse the convolution results with different dilation rates as the Query, and use the convolution result with a dilation rate of 1 as the Keys and Values;
[0053] First perform BN operation on Keys and Query and then fuse them, perform a dot product operation on the fused result and Values, and perform a residual connection on the obtained result and the convolution result with a dilation rate of 1;
[0054] Then, the result obtained after passing through the BN operation, the DepConv operation, and the Lie group activation function in sequence is subjected to a second residual connection with the result after the previous residual connection
[0055] Another object of the present invention is to provide a remote sensing scene classification device based on visual token clustering and discarding, and the device includes:
[0056] A construction module for constructing the network architecture of the remote sensing image classification model;
[0057] The training module is used to train the remote sensing image classification model network architecture to obtain a pre-trained remote sensing image classification model;
[0058] The classification module is used to obtain the remote sensing image to be classified, and input the remote sensing image to be classified into the pre-trained remote sensing image classification model to obtain the final classification result;
[0059] Among them, the remote sensing image classification model network architecture includes at least a multi-scale feature extraction and fusion module, a feature enhancement aggregation module, and a Transformer module which are connected in sequence. The Transformer module is used to cluster the visual markers into fixed clusters after receiving the visual markers, and then merge the visual markers in the same cluster into one visual marker, so as to discard the corresponding visual markers according to their importance.
[0060] Another object of the present invention is to provide a readable storage medium having a computer program stored thereon, wherein the program implements the steps of the above method when executed by a processor.
[0061] Another object of the present invention is to provide an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of the above method when executing the program.
[0062] The present invention clusters and merges the visual markers formed by pixels obtained by feature-level fusion by using the Transformer module, and then selectively discards unimportant visual markers, ignoring unimportant parts and focusing only on important parts, thereby achieving efficient and accurate scene classification. This solves the problem in the prior art that it is impossible to simultaneously take into account both computational efficiency and high accuracy when classifying remote sensing images.
[0063] In addition, the present invention has at least the following beneficial effects:
[0064] 1. The multi-scale feature extraction and fusion module is divided into two stages: shallow feature extraction and high-level feature extraction. By mapping samples to the Lie group manifold space, shallow features such as geometric shapes and colors in HRRSI can be effectively captured. At the same time, combined with the mechanism based on human cognition, it simulates the selective attention of humans to different levels of features when processing visual information, that is, in the shallow stage, key geometric and texture information is extracted first, and in the high-level stage, more attention is paid to overall semantics and contextual relationships, so as to generate a more complete and semantic feature map for classification;
[0065] 2. The feature enhancement aggregation module processes spatial features and channel features in parallel, which enhances the model's focus on important areas while reducing information loss. It improves the ability to capture complex scene features while maintaining high computational efficiency.
[0066] 3. The introduction of PeLK operation, Lie group machine learning, etc. effectively expands the receptive field and enhances the ability to extract multi-scale features by simulating the peripheral vision processing method in human cognition. While reducing the model complexity, it significantly improves the computational efficiency and classification performance. Brief Description of the Drawings
[0067] Figure 1 It is a flowchart of remote sensing scene classification based on visual marker clustering and discarding provided by an embodiment of the present invention;
[0068] Figure 2 It is a structural block diagram of a remote sensing scene classification device based on visual marker clustering and discarding in the third embodiment of the present invention.
[0069] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. Specific Embodiments
[0070] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0071] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0072] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0073] The following will combine specific embodiments and drawings to detail how to improve the computational efficiency on the premise of ensuring the accuracy of remote sensing image classification.
[0074] Embodiment 1
[0075] Please refer to Figure 1 , which shows the remote sensing scene classification method based on visual marker clustering and discarding in the first embodiment of the present invention. The method includes steps S10 to S12.
[0076] Step S10, construct a network architecture for a remote sensing image classification model.
[0077] Among them, a new network architecture for a remote sensing image classification model is constructed. The network architecture for the remote sensing image classification model at least includes a multi-scale feature extraction and fusion module, a feature enhancement and aggregation module, and a Transformer module connected in sequence. The Transformer module is used to cluster visual tokens into fixed clusters when receiving them, and then merge the visual tokens within the same cluster into one visual token, so as to discard the corresponding visual tokens according to the importance degree of the visual tokens.
[0078] Specifically, the multi-scale feature extraction and fusion module aims to effectively capture shallow features such as geometry and color in high-resolution remote sensing images (HRRSIs) by mapping samples to the Lie group manifold space, and pays attention to the overall semantics and context relationships to generate a more complete and semantic feature map for classification. The feature enhancement and aggregation module enhances the model's attention to important regions by parallel processing spatial features and channel features, while reducing information loss. While improving the ability to capture features in complex scenes, this module maintains high computational efficiency. Finally, the Transformer module uses a preset clustering method to cluster and merge the received visual tokens, and then selectively discards unimportant visual tokens (such as backgrounds) to achieve efficient and accurate scene classification.
[0079] More specifically, in this embodiment, the implementation process and principle of the Transformer module are introduced in detail. Among the existing technologies, some Vision Transformer models uniformly divide an image into blocks of a fixed size and then generate corresponding visual tokens, resulting in the model treating all regions "equally". However, in people's perception, not all regions are equally important. Especially in scenes that are difficult to distinguish, people pay more attention to important details and ignore irrelevant features. To solve this problem, the embodiment of the present invention proposes a model of the Transformer module based on people's perception. The pixels obtained by feature-level fusion are used as the original visual tokens. After the Transformer module optimizes the attention mechanism, new visual tokens are generated, and a preset clustering method is used to cluster and merge them, and then unimportant visual tokens (such as backgrounds) are selectively discarded to achieve efficient and accurate scene classification.
[0080] First, after receiving the features output by the previous module, in order to learn features of different scales in different scenarios, parallel dilated convolutions are used for feature extraction. Using parallel dilated convolutions can effectively increase the receptive field while reducing the number of parameters in the model. The convolution results with different dilation rates of the features received from the previous module are fused and used as the Query. The convolution result with a dilation rate of 1 of the features received from the previous module is used as the Keys and Values. The Keys and Query are first subjected to BN operations and then fused. The fused result is dot-producted with the Values, and the obtained result is residual-connected with the convolution result with a dilation rate of 1. This operation is mainly to prevent gradient disappearance, reduce the training error of the model, and improve the generalization ability of the model.
[0081] The pixels in the feature map after the above operations are fewer than the original visual tokens, which can effectively reduce the computational complexity of the model. Next, first through the BN operation, the dependencies between channels are captured to maintain the stability and expressive ability of the feature map. In order to reduce the number of parameters in the model, DepConv operations are used to capture information such as local features of the scene, and Lie group activation functions are adopted. Finally, the obtained result is secondarily residual-connected with the result of the previous residual connection to enhance the representational ability of the model and learn effective features faster, improving the model's ability to process HRRSI.
[0082] Then, after the above optimized attention mechanism, new visual tokens are generated, and clustering and merging are performed for discarding. Among them, the visual tokens are grouped into a fixed cluster, and then the tokens within the same cluster are merged into one visual token. During the visual token clustering process, for a series of obtained visual tokens, first calculate the local density ρ of each visual token based on the within-group mean in its Lie group:
[0083] ;
[0084] where C represents the visual token, represents the within-group mean of the i-th visual token, and represent the features of the i-th and j-th visual tokens respectively;
[0085] To further screen important visual tokens, the Lie group manifold space distance is used to measure the target visual token closest to the target, and then clustering is performed by calculating the minimum distance between the target visual token and other visual tokens. Specifically, consider combining the distance and local density metrics to calculate the weight value of each visual token × . A higher weight value indicates a higher probability of becoming the center of a cluster. From the perspective of human cognition, by selecting the visual marker with the highest weight value (i.e., obtaining the target visual marker) as the center of the cluster, and then assigning other visual markers to the nearest cluster center according to the distance of the visual markers; among them, the calculation formula of the Lie group manifold space distance is:
[0086] ;
[0087] Among them, and respectively represent the features of the i-th and j-th visual markers, represents the local density of the i-th visual marker, ρ j represents the local density of the j-th visual marker.
[0088] Furthermore, for the merging of visual marker features, the traditional method is to directly calculate the average value of the features of the visual markers in the cluster. However, in visual tasks centered on human cognition, for visual markers with similar semantics, they are not equally important.
[0089] Therefore, in order to simulate the process of human cognition, the within-Lie-group mean is introduced to represent the importance degree of each visual marker. The within-Lie-group mean can effectively represent the most essential features of the scene, and a real symmetric matrix is used to represent the eigenvalues. This method can effectively reduce the memory space occupied by features and facilitate calculation. Calculate the within-Lie-group mean of each visual marker under the guidance of the importance degree of the visual marker.
[0090] Finally, the merged visual marker is:
[0091] ;
[0092] Among them, represents the cluster formed by the i-th visual marker, and respectively represent the feature of the j-th visual marker and the importance degree of its corresponding within-Lie-group mean, m i represents the feature of the merged visual marker of the cluster formed by the i-th visual marker.
[0093] The merged visual marker is used as the input of the improved Transformer module and as the Query, while the original visual marker is used as the Keys and Values. In order to simulate the process of human cognition, it is necessary to give play to the role and contribution of important visual markers. Therefore, the importance degree of visual markers is added to the attention mechanism of the aforementioned Transformer module, and the calculation is as follows:
[0094] ;
[0095] Among them, represents the number of channels of the Queries. Q and K respectively represent the query Query and the key Key in the attention mechanism, and V represents the value Value. represents the importance degree of the mean within the Lie group, and T represents the matrix transpose operation. This attention mechanism converts the human cognitive process into computable weights and can mimic the key region features of the scene that a person's visual perception focuses on.
[0096] In addition, on the basis of clustering and merging visual tokens, self-attention weights are used to calculate the importance degree of visual tokens, and unimportant visual tokens are discarded. For example, by calculating the importance degree of visual tokens, visual tokens with an importance degree lower than a preset threshold are discarded, which greatly reduces the number of unimportant visual tokens during the model training process and effectively improves the computational performance without reducing the model accuracy.
[0097] Among them, the importance degree of Tokens is calculated as follows:
[0098] ;
[0099] Among them, represents the importance degree of visual tokens, represents the classification visual tokens. The numerator in the above represents the attention feature map, which is the dot product of the query vector obtained from the classification visual tokens and the Keys value matrix. d = L / H, where L represents the length of each visual token embedded into a vector of length L, and H represents the number of heads in the multi-head attention mechanism.
[0100] Since the visual tokens in the value matrix are used for prediction classification according to linear combination, it is used to represent the importance degree of visual tokens. In actual calculation, according to to retain the top K visual tokens, and the value of K is set by the discard rate of visual tokens, and pre-trained classification visual tokens are used to calculate the importance degree. This operation does not require the gradient to be continuously passed to the starting point of the model, thereby reducing the memory occupancy and improving the computational performance of the model.
[0101] Step S11, train the remote sensing image classification model network architecture to obtain a pre-trained remote sensing image classification model.
[0102] Among them, train the remote sensing image classification model network architecture to obtain a remote sensing image classification model with stable performance and capable of accurate classification.
[0103] Step S12, obtain the remote sensing image to be classified, and input the remote sensing image to be classified into the pre-trained remote sensing image classification model to obtain the final classification result.
[0104] Specifically, the remotely sensed image to be classified is obtained. Since the remotely sensed image classification model has mastered the classification rules of remotely sensed images, the remotely sensed image to be classified can accurately output the final classification result.
[0105] In summary, the present invention realizes efficient and accurate scene classification by clustering and then merging the visual tokens formed by the pixels obtained through feature-level fusion using the Transformer module, and then selectively discarding unimportant visual tokens, ignoring the unimportant parts and only caring about the important parts. It solves the problem in the prior art that it is impossible to achieve both high computational efficiency and high accuracy when classifying remotely sensed images.
[0106] Embodiment 2
[0107] This embodiment also proposes a remotely sensed scene classification method based on clustering and discarding of visual tokens. The difference between the remotely sensed scene classification method based on clustering and discarding of visual tokens in this embodiment and the remotely sensed scene classification method based on clustering and discarding of visual tokens in Embodiment 1 is as follows:
[0108] The multi-scale feature extraction and fusion module is divided into two stages: shallow feature extraction and deep feature extraction. The process of shallow feature extraction is as follows:
[0109]
[0110] Among them, represents the spatial position of the target object in the image, represents the color space information of the image, which can effectively capture the color features of the image under brightness and color changes, such as blue ocean or green forest scenes, is used to extract the local gradient direction of the image, mainly used to describe the shape and edge information of the object, and can enhance the model's ability to detect targets. 3D-GMK(x,y) represents the spatial-spectral feature, which contains 13 directions and can obtain an enhanced feature map by combining without increasing parameters and computational complexity, is used to describe the spatial correlation between pixels and can capture the texture information of the image, such as roughness, contrast, and directionality, The minimum spanning tree (MST) algorithm is also introduced to calculate the shortest distance between different target objects, which is used to characterize the spatial relationship between different objects;
[0111] However, it is not comprehensive to classify scenes only relying on shallow features. Therefore, high-level features of the scene are further extracted by using high-level feature extraction. Traditional high-level feature extraction methods mainly rely on convolutional operations and extract semantic features of the scene by continuously deepening the depth of the network. However, as the depth of the network increases, the number of parameters and computational complexity of the convolutional neural network also increase. Especially when dealing with large-scale HRRSI, it is easy to cause excessive consumption of computing resources. This not only limits the practical application of the model but also increases the training cost. Therefore, the process of high-level feature extraction in the embodiments of the present invention is as follows:
[0112] Perform standard convolution operations on the received feature map to capture local spatial features;
[0113] Use 1×1 convolution on the feature map after standard convolution operations to compress the channels, reduce the number of parameters, extract more effective features, and then perform adaptive pooling processing to adjust the feature map to a unified size;
[0114] After the feature map processed by adaptive pooling is batch-normalized, the batch normalization operation can effectively improve the stability of training and accelerate the convergence process of the model. In order to learn features of different scales in the scene, after batch normalization, the PeLK operation is performed to replace the traditional convolution. This operation greatly expands the effective receptive field (ERF) while reducing the model parameter complexity, thus more effectively simulating the way the human eye perceives global and local information. Specifically, the process of the PeLK operation is as follows:
[0115] ;
[0116] Among them, a represents the position of the feature map in the height (or row) direction, b represents the position of the feature map in the width (or column) direction, z(c,d) is the weight of the convolution kernel, w(c,d) is the position embedding, F is the input feature, r w represents the row, and c and d represent coordinates respectively;
[0117] Finally, perform the fusion of multi-scale features. Use the features extracted at different scales after shallow feature extraction and high-level feature extraction as the Query, and use the results obtained by convolving the features extracted at different scales with a dilation rate of 1 as the Keys and Values;
[0118] Normalize the Keys and Query and then fuse them. Perform a dot product operation on the fused result and the Values, and perform a residual connection with the result obtained by convolving the extracted features with a dilation rate of 1 to obtain the multi-scale features that the final multi-scale feature extraction and fusion module needs to output;
[0119] Among them, the features extracted at different scales are fused to form a multi-scale feature representation. The fusion process is to use the features extracted at different scales as Queries, and the results obtained by convolving the features extracted at different scales with a dilation rate of 1 as Keys and Values. After normalizing the Keys and Queries, they are fused. The fused result is dot-producted with the Values, and a residual connection is made with the result obtained by convolving with a dilation rate of 1 to obtain the final multi-scale features. This operation is mainly used to prevent gradient disappearance, reduce the model training error, and improve the generalization ability of the model. After the feature map is processed by the above operations, the number of visual markers is reduced compared to the original visual markers, which can effectively reduce the computational complexity of the model.
[0120] In addition, the feature enhancement aggregation module includes parallel spatial feature enhancement aggregation and channel feature enhancement aggregation; by processing spatial features and channel features in parallel, the feature enhancement aggregation module enhances the model's attention to important regions while reducing information loss. While improving the ability to capture features in complex scenes, this module maintains a high computational efficiency;
[0121] Among them, the process of spatial feature enhancement aggregation is as follows:
[0122] First, the input feature map is processed by batch normalization to standardize the feature map; BN can not only stabilize the learning process of the network, but also effectively accelerate the convergence speed of the model. By reducing the internal covariate shift, BN can reduce the oscillation during model training, thereby accelerating convergence and improving the stability of the model;
[0123] The feature map after batch normalization is processed by an adaptive pooling layer; this layer dynamically adjusts the size of the feature map according to the input size to ensure that the output feature map has a fixed size, so as to meet the requirements of subsequent network layers;
[0124] The pooled feature map is processed by a 1×1 convolutional layer; to reduce the dimension and extract relevant features;
[0125] The Lie group activation function is introduced to activate the feature map after convolution processing. Subsequently, the feature map is flattened, and the features of each channel are assigned corresponding weights, thereby converting the multi-dimensional features into a one-dimensional vector; compared with the traditional sigmoid activation function, the Lie group activation function can better adapt to matrix samples and vector samples. Subsequently, the features are flattened, and the features of each channel are assigned corresponding weights, thereby strengthening the model's attention to important features. The flattening operation converts multi-dimensional features into a one-dimensional vector, simplifies the processing of subsequent fully connected layers, reduces the computational complexity, provides a more compact feature representation, and is suitable for classification or regression tasks.
[0126] In traditional feature extraction methods, single feature processing is usually adopted. Such a processing method results in insufficient information in a certain dimension and insufficient feature extraction. Therefore, in the embodiments of the present invention, channel feature enhancement aggregation is added as a supplement to the above-mentioned spatial feature enhancement aggregation;
[0127] Specifically, the process of channel feature enhancement aggregation is as follows:
[0128] The input feature map is processed by global average pooling to obtain the global feature description of each channel. The purpose of global average pooling is to generate a representative value for each channel, thereby capturing global information. The formula is as follows:
[0129] ;
[0130] Among them, represents the global average pooling result of the c-th channel, H and W are the height and width of the feature map respectively, represents the pixel value of the input feature map on the c-th channel;
[0131] The global average pooling results are passed through an independent 1×1 convolutional layer for inter-channel feature extraction and fusion; these global average pooling results are passed through an independent 1×1 convolutional layer for inter-channel feature extraction and fusion. The purpose of the 1x1 convolution is to remap the feature information of different channels and reduce the computational complexity by reducing the redundant information between channels. At the same time, the 1x1 convolution retains the global correlation between channels and enhances the feature expression ability through a non-linear activation function;
[0132] Finally, the features of all channels are added to complete the fusion and then passed through a convolutional layer and a SeLU activation function for non-linear processing, and then passed through a 1×1 convolution for channel compression to obtain the final fused feature map: ;
[0133] Among them, c represents the channel, is the fused feature map after fusion, represents the sum of the convolution results of all channels.
[0134] Embodiment III
[0135] Please refer to Figure 2 , which shows the remote sensing scene classification device based on visual marker clustering and discarding proposed in the third embodiment of the present invention. The device includes:
[0136] A construction module 100, configured to construct a network architecture of a remote sensing image classification model;
[0137] A training module 200 for training a remote sensing image classification model network architecture to obtain a pre-trained remote sensing image classification model;
[0138] A classification module 300 for obtaining a remote sensing image to be classified and inputting the remote sensing image to be classified into the pre-trained remote sensing image classification model to obtain a final classification result;
[0139] Wherein, the remote sensing image classification model network architecture at least includes a multi-scale feature extraction and fusion module, a feature enhancement aggregation module, and a Transformer module connected in sequence. The Transformer module is used to cluster visual tokens into fixed clusters when receiving visual tokens, and then merge the visual tokens within the same cluster into one visual token, so as to discard the corresponding visual tokens according to the importance of the visual tokens.
[0140] The functions or operation steps implemented when the above modules are executed are substantially the same as those in the above method embodiments, and will not be described in detail here.
[0141] Embodiment 4
[0142] On the other hand, the present invention also provides a readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in any one of the above Embodiments 1 to 2 are implemented.
[0143] Embodiment 5
[0144] On the other hand, the present invention also provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, the steps of the method described in any one of the above Embodiments 1 to 2 are implemented.
[0145] The technical features of each of the above embodiments can be combined arbitrarily. For the sake of brevity of description, all possible combinations of the technical features in the above embodiments are not described. However, as long as the combinations of these technical features do not conflict, they should be considered to be within the scope described in this specification.
[0146] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable storage medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0147] More specific examples (a non-exhaustive list) of computer-readable storage media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable storage medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0148] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0149] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0150] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.
Claims
1. A remote sensing scene classification method based on visual marker clustering and discarding, characterized in that, The method includes: Constructing a remote sensing image classification model network architecture; Training the remote sensing image classification model network architecture to obtain a pre-trained remote sensing image classification model; Obtaining a remote sensing image to be classified and inputting the remote sensing image to be classified into the pre-trained remote sensing image classification model to obtain a final classification result; Among them, the remote sensing image classification model network architecture at least includes a multi-scale feature extraction and fusion module, a feature enhancement aggregation module, and a Transformer module connected in sequence. The Transformer module is used to cluster visual tokens into fixed clusters when receiving visual tokens, and then merge the visual tokens within the same cluster into one visual token, so as to discard the corresponding visual tokens according to the importance degree of the visual tokens; The step of clustering visual tokens into fixed clusters includes: Calculating the local density ρ of each visual token based on the in-group mean of the received visual tokens in the Lie group; ; where C represents a visual marker, represents the Lie group inner mean of adjacent visual markers of the i-th visual marker, and represent the features of the i-th and j-th visual markers, respectively; Determining the target visual token closest to the target as the clustering center according to the local density of the visual token and the Lie group manifold space distance; And assigning other visual tokens to the nearest clustering center according to the Lie group manifold space distance to form fixed clusters; Among them, the calculation formula of the Lie group manifold space distance is: ; Among them, and represent the features of the \(i\)-th and \(j\)-th visual markers respectively, represents the local density of the \(i\)-th visual marker, \(\rho\) j represents the local density of the \(j\)-th visual marker; The formula for merging visual tokens within the same cluster into one visual token is: ; Among them, represents the cluster formed by the i-th visual marker clustering, and respectively represent the feature of the j-th visual marker and the importance degree of the mean within the corresponding Lie group, m i represents the feature of the visual marker after the clusters formed by the i-th visual marker clustering are merged; The calculation formula for the importance degree of visual tokens is as follows: ; Among them, represents the importance degree of the visual marker. The numerator in the formula represents the attention feature map, which is the dot product of the query vector obtained from classifying the visual marker and the Keys value matrix. represents classifying the visual marker. d =L / H, where L represents the vector into which each visual marker is embedded with a length of L, H represents the number of heads in the multi-head attention mechanism in the Transformer module, K is a fixed value set by the dropout rate of the visual marker, and T represents the matrix device operation.
2. The remote sensing scene classification method based on visual marker clustering and discarding according to claim 1, wherein The multi-scale feature extraction and fusion module is divided into shallow feature extraction and deep feature extraction; The process of shallow feature extraction is: ; Among them, represents the spatial position of the target object in the image, represents the color space information of the image, is used to extract the local gradient direction of the image, and 3D - GMK(x,y) represents the spatial - spectrum feature, is used to describe the spatial correlation between pixels, calculates the shortest distance between different target objects and is used to characterize the spatial relationship between different objects. T represents the matrix transpose operation; The process of deep feature extraction is: Performing standard convolution operations on the received feature map to capture local spatial features; After using 1×1 convolution for channel compression of the feature map after standard convolution operations, and then performing adaptive pooling processing to adjust the feature map to a unified size; Performing PeLK operation on the feature map after adaptive pooling processing after batch normalization. The process of PeLK operation is as follows: ; Among them, a represents the position of the feature map in the height (or row) direction, b represents the position of the feature map in the width (or column) direction, z(c,d) is the weight of the convolutional kernel, w(c,d) is the positional embedding, F is the input feature, r w represents the row, and c and d represent coordinates respectively; Taking the features extracted at different scales after shallow feature extraction and deep feature extraction as Queries, and taking the results obtained by using convolution with a dilation rate of 1 on the features extracted at different scales after shallow feature extraction and deep feature extraction as Keys and Values; And normalizing Keys and Queries and then fusing them. The fused result is dot-product operated with Values, and residual connection is performed with the results obtained by using convolution with a dilation rate of 1 on the features extracted at different scales after shallow feature extraction and deep feature extraction to obtain the multi-scale features that the multi-scale feature extraction and fusion module needs to output.
3. The remote sensing scene classification method based on visual marker clustering and discarding according to claim 2, characterized in that, The feature enhancement aggregation module includes parallel spatial feature enhancement aggregation and channel feature enhancement aggregation; The process of spatial feature enhancement aggregation is: The input feature map is processed by batch normalization to standardize the feature map; The feature map after batch normalization is processed through an adaptive pooling layer; The pooled feature map is processed through a 1×1 convolutional layer; The Lie group activation function is introduced to activate the feature map after convolution processing, and then the feature map is flattened. The features of each channel are assigned corresponding weights, so as to convert the multi-dimensional features into a one-dimensional vector; The process of channel feature enhancement and aggregation is as follows: The input feature map is processed through global average pooling: ; Among them, represents the global average pooling result of the c-th channel, where H and W are the height and width of the feature map respectively, represents the pixel value of the input feature map on the c-th channel, and i and j respectively represent the row and column values of the feature matrix; The global average pooling result is processed through an independent 1×1 convolutional layer to extract and fuse the features between channels; The features of all channels are subjected to an addition operation to complete the fusion, followed by non-linear processing through a convolutional layer and a SeLU activation function, and then channel compression is performed through a 1×1 convolution to obtain the final fused feature map: ; Among them, c represents the channel, is the fused feature map after fusion, represents the sum of the convolution results of all channels.
4. The remote sensing scene classification method based on vision marker clustering and discarding according to claim 3, characterized in that The method further includes: When the Transformer module receives the input features, convolutions with different dilation rates are performed, and the convolution results with different dilation rates are fused and used as Query, and the convolution result with a dilation rate of 1 is used as Keys and Values; Keys and Query are first subjected to BN operation and then fused. The fused result is subjected to a dot product operation with Values, and the obtained result is subjected to a residual connection with the convolution result with a dilation rate of 1; Subsequently, the results obtained after passing through the BN operation, DepConv operation, and Lie group activation function in sequence are subjected to a second residual connection with the result after the previous residual connection.
5. A remote sensing scene classification device based on visual marker clustering and discarding, which is used to implement the remote sensing scene classification method based on visual marker clustering and discarding described in any one of claims 1 to 4, characterized in that, The device includes: A construction module for constructing the network architecture of the remote sensing image classification model; A training module for training the network architecture of the remote sensing image classification model to obtain a pre-trained remote sensing image classification model; A classification module for obtaining the remote sensing image to be classified and inputting the remote sensing image to be classified into the pre-trained remote sensing image classification model to obtain the final classification result; Among them, the network architecture of the remote sensing image classification model at least includes a multi-scale feature extraction and fusion module, a feature enhancement and aggregation module, and a Transformer module connected in sequence. The Transformer module is used to cluster the visual tokens into fixed clusters when receiving the visual tokens, and then merge the visual tokens within the same cluster into one visual token, so as to discard the corresponding visual tokens according to the importance of the visual tokens.
6. A readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 4.
7. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Remote sensing scene classification method and device based on Lie group spatial features
CN117152547A