Multi-label remote sensing scene classification method and system based on state space model
By using state space model, linear multi-scale interaction enhancement model and linear semantic relationship in multi-label remote sensing scene classification, combining semantic and visual features, the problem of inaccurate classification in the existing technology is solved, and more efficient and accurate multi-label remote sensing image classification is achieved.
Patent Information
- Application Number
- CN202510111204.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-16
AI Technical Summary
The existing multi-label remote sensing scene classification method is difficult to ensure the exact correspondence between the extracted feature vectors and the actual geographic categories, resulting in inaccurate classification.
A multi-label remote sensing scene classification method based on state space model is adopted to establish a model through a linear multi-scale interaction enhancement model and a linear semantic relationship, and a pre-trained language model Bert extracts semantic embedding features and a convolutional neural network to extract visual features to establish the connection between different surface coverings in multi-label remote sensing images.
It significantly improves the classification accuracy of the model for multi-label remote sensing images, enhances the precise correspondence between semantic feature vectors and actual geographic categories, and reduces the computational complexity and hardware cost.
Smart Images

Figure CN120014463A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image label classification, and in particular relates to a multi-label remote sensing scene classification method and system based on a state space model. Background Art
[0002] Remote sensing images have rich spatial and spectral information and can provide valuable data on objects on the earth's surface, land use, and environmental changes. Since scenes in remote sensing images usually contain multiple types of objects, traditional single-label classification methods often have difficulty accurately describing the diversity and complexity of remote sensing images. Therefore, multi-label classification methods have received extensive attention in remote sensing image analysis, and their goal is to automatically identify and classify multiple types of object coverage targets from remote sensing images. In multi-label remote sensing scene classification, each remote sensing image contains multiple labels, representing different types of objects in the image. Compared with single-label classification tasks, multi-label classification requires the simultaneous prediction of the presence or absence of multiple labels in the image, which places higher requirements on the learning and generalization capabilities of the model.
[0003] At present, the methods of multi-label remote sensing scene classification can be divided into two types: feature enhancement methods and semantic association methods. Feature enhancement methods improve the model's ability to learn remote sensing image features by integrating advanced technologies such as attention mechanisms, convolutional neural networks, and multi-scale feature extraction. With the help of these technologies, the model can more accurately extract detailed information in the image, thereby effectively distinguishing different ground objects. The semantic association method focuses on the relationship between labels and uses the intrinsic connection between ground object categories to optimize the classification effect. For example, in remote sensing images, docks often appear at the same time as water bodies, while desert areas contain almost no water bodies. These spatial and semantic connections provide the model with useful contextual information, enabling the model to more reasonably predict the multi-label content in the image.
[0004] Existing feature enhancement methods (multi-scale feature information fusion methods) usually directly merge information of different scales by simple splicing or weighted averaging, ignoring the deep relationship between features of each scale and the dynamic changes of contextual information. To address this problem, the present invention proposes a linear multi-scale interactive enhancement model, which aims to establish global connections within the same scale based on local feature maps, and effectively capture the contextual relationship between different scales. This solves the problem of different scales of coverage of different objects in multi-label remote sensing images.
[0005] Existing semantic association methods usually rely on convolutional classifiers before establishing semantic connections. However, the limitation of this method is that it is difficult to ensure the accurate correspondence between the extracted feature vectors and the actual object categories. To solve this problem, the present invention introduces a pre-trained large-scale language model, which uses its powerful language understanding ability to capture text information related to different object coverage targets in remote sensing images from text data. By effectively fusing this text information with the features of remote sensing images, the model can establish a close association between images and labels at the semantic level.
[0006] Existing methods usually rely on Transformer and its variants in feature enhancement and semantic relationship modeling. Although these methods perform well in modeling global dependencies and capturing semantic information, their computational and training costs are high. In order to solve these problems, the present invention introduces a state space model to achieve modeling of sequence data at a lower computational cost, thereby significantly reducing computational complexity and hardware costs while capturing global information and establishing semantic associations, providing a more practical solution for tasks such as multi-label remote sensing image classification. Summary of the invention
[0007] The purpose of the present invention is to overcome the problem that the existing remote sensing scene classification methods are difficult to ensure the accurate correspondence between the extracted feature vectors and the actual object categories, resulting in inaccurate remote sensing scene classification, and to provide a multi-label remote sensing scene classification method and system based on a state space model.
[0008] In order to achieve the above object, the present invention adopts the following technical solution:
[0009] A multi-label remote sensing scene classification method based on a state space model comprises the following steps:
[0010] Acquire multi-label remote sensing images;
[0011] Extracting semantic embedding features from multi-label remote sensing images;
[0012] Extract visual features from multi-label remote sensing images;
[0013] Construct a linear multi-scale interaction enhancement model, perform multi-scale scanning on the extracted visual features based on the linear multi-scale interaction enhancement model, extract and output the long-distance dependencies represented by the visual features;
[0014] Construct a linear semantic relationship model based on which, under the guidance of the long-range dependency relationship represented by visual features, the relationship between different land cover objects in multi-label remote sensing images is established in combination with the extracted semantic embedding features;
[0015] Based on the relationship between different land covers, various land covers in multi-label remote sensing images are classified using a binary classifier.
[0016] In the step of extracting semantic embedding information in the multi-label remote sensing image, the semantic embedding information in the multi-label remote sensing image is extracted by the pre-trained language model Bert to generate a semantic vector. The process is defined as:
[0017] SE=[se1,se2,...,se c ]=[Φ bert (l1),Φ bert (l2),...,Φ bert (l c )],
[0018] where Φ bert (·) represents the text encoding process of Bert, Indicates dimension d l l j Semantic feature vector.
[0019] In the step of extracting visual features from multi-label remote sensing images, a convolutional neural network is used to extract and output as the extracted visual features, where H, W, and C represent the height, width, and number of channels, respectively.
[0020] In the step of constructing a linear multi-scale interaction enhancement model, performing multi-scale scanning on the extracted visual features based on the linear multi-scale interaction enhancement model, and extracting and outputting the long-distance dependency relationship represented by the visual features, the specific working process of the constructed linear multi-scale interaction enhancement model is as follows:
[0021] Receiving the extracted visual features as input features;
[0022] Regularize the input features;
[0023] Perform linear transformation on the input features after regularization to obtain input features with reduced dimensions;
[0024] Perform convolution operations on the reduced-dimensional input features to extract local spatial information in the image and generate feature maps;
[0025] Based on the multi-scale state space model, the generated feature map is downsampled to extract feature embeddings of different scales. The bidirectional structure is used to transfer and update the feature embeddings of different scales to obtain enhanced feature embeddings of different scales. The enhanced feature embeddings of different scales are added together to obtain the long-distance dependency relationship of visual feature representation.
[0026] The long-distance dependencies represented by the obtained visual features are output.
[0027] The downsampling operation is performed on the generated feature map to extract the feature embedding of different scales as follows:
[0028] P i = Downsample(P i-1 ),1<i≤N s ,
[0029] Where Downsample(·) is the downsampling operation, N s represents the total number of scales, Representation scale 1<i≤N s The feature embedding at i =(H×W) / 2(i-1).
[0030] The bidirectional structure is used to transfer and update the feature embeddings of different scales to obtain enhanced feature embeddings of different scales. The process of transferring and updating the feature embedding of the i-th scale is expressed as follows:
[0031]
[0032] Where P i+1 and P i-1 Can be viewed as high-scale and low-scale information, and represents the initial hidden state space of the forward and backward processes, Avgpool(·) represents the average pooling operation, which compresses the three-dimensional feature map into a two-dimensional feature vector suitable for the state space model, and W ms-s6-f , represents the learnable parameters in the forward process,
[0033] W ms-s6-b , The corresponding parameters represent the learnable parameters in the backward process, Represents enhanced feature embedding;
[0034] The enhanced feature embedding at the i-th scale is:
[0035] The linear semantic relationship establishment model is constructed based on the linear semantic relationship establishment model. Under the guidance of the long-distance dependency relationship represented by the visual features, the extracted semantic embedding features are combined to establish the connection between different surface covers in the multi-label remote sensing image. The specific working process of the constructed linear semantic relationship establishment model is as follows:
[0036] Receiving the extracted semantic embedding features as input features;
[0037] Regularize the input semantic embedding features;
[0038] Perform linear transformation on the regularized semantic embedding features to obtain semantic embedding features with reduced dimensions;
[0039] Perform a 1D convolution operation on the reduced-dimensional semantic embedding features to obtain formatted semantic embedding features;
[0040] Based on the feature-guided semantic relationship state space model, under the guidance of the long-distance dependency relationship represented by the visual features, a bidirectional structure is used to transfer and update the formatted semantic embedding features, and the forward semantic embedding features are added to the backward semantic embedding features to obtain the enhanced semantic embedding features.
[0041] The association between different land cover objects in multi-label remote sensing images is established based on the local features of the enhanced semantic embedding features.
[0042] The feature-guided semantic relationship state space model adopts a bidirectional structure to transfer and update the formatted semantic embedding features under the guidance of the long-distance dependency relationship represented by the visual features, and adds the forward semantic embedding features to the backward semantic embedding features to obtain the enhanced semantic embedding features. The specific working process is as follows:
[0043]
[0044] where Φ linear (·) represents linear embedding, and denote the initial hidden state space of the forward and backward processes of the feature-guided semantic relation state space model, W fgs-s6-f , and represents the learnable parameters in the forward process, W fgs-s6-b , and Represents the learnable parameters in the backward process, se` k Indicates k The semantic embedding obtained after the regularization layer, linear embedding layer and 1D convolution layer, ese k Refers to the enhanced semantic embedding;
[0045] The output of the feature-guided semantic relation state space model is expressed as
[0046] In the step of classifying various land covers in the multi-label remote sensing image by a binary classifier based on the relationship between different land covers, the specific representation of the binary classifier is as follows:
[0047]
[0048] where csj is the category score of the jth class, Sigmoid(·) represents the sigmoid activation function, and Φ linear (·) denotes a linear learnable layer.
[0049] A multi-label remote sensing scene classification system based on a state space model, comprising:
[0050] An acquisition module, used to acquire multi-label remote sensing images;
[0051] Semantic embedding extraction module, used to extract semantic embedding features in multi-label remote sensing images;
[0052] Visual feature extraction module, used to extract visual features from multi-label remote sensing images;
[0053] A linear multi-scale interaction enhancement model construction module is used to construct a linear multi-scale interaction enhancement model, perform multi-scale scanning on the extracted visual features based on the linear multi-scale interaction enhancement model, and extract and output the long-distance dependency relationship represented by the visual features;
[0054] The linear semantic relationship model building module is used to build a linear semantic relationship model. Based on the linear semantic relationship model, under the guidance of the long-distance dependency relationship represented by the visual features, the relationship between different surface covers in the multi-label remote sensing image is established in combination with the extracted semantic embedding features;
[0055] The classification module is used to classify various land cover in multi-label remote sensing images through binary classifiers based on the relationship between different land cover.
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] The present invention provides a multi-label remote sensing scene classification method and system based on a state space model. The method comprises the following steps: obtaining a multi-label remote sensing image; extracting semantic embedding features in the multi-label remote sensing image; extracting visual features in the multi-label remote sensing image; constructing a linear multi-scale interactive enhancement model, performing multi-scale scanning on the extracted visual features based on the linear multi-scale interactive enhancement model, extracting and outputting the long-distance dependency relationship represented by the visual features; constructing a linear semantic relationship establishment model, establishing a model based on the linear semantic relationship, under the guidance of the long-distance dependency relationship represented by the visual features, combining the extracted semantic embedding features to establish the connection between different surface covers in the multi-label remote sensing image; based on the connection between different surface covers, classifying various surface covers in the multi-label remote sensing image through a binary classifier. The linear multi-scale interactive enhancement model and the linear semantic relationship establishment model developed by the present invention fully consider the characteristics of remote sensing images, and the linear multi-scale interactive enhancement model enhances the representation of visual features by mining long-distance context information of different scales. In this way, the model can fully explore land cover targets of various scales from remote sensing images. The linear semantic relationship establishment model aims to establish semantic associations between different labels through the guidance of visual information. It can improve the ability of the efficient network based on the state space model to understand the complex content of remote sensing images by revealing the semantic relationship between different types of land cover, and improve the precise correspondence between the semantic feature vector and the actual ground object category. Through the collaboration between the two modules, the present invention can learn local information, global context, multi-scale clues and semantic associations from remote sensing images, improve the accuracy of model classification, which is very useful for multi-label remote sensing scene classification tasks.
[0058] Furthermore, based on the multi-scale state-space model, by deeply mining local features, multi-scale information, long-distance dependencies, and semantic relationships between different objects, the efficient network based on the state-space model can fully explore the complex content in remote sensing images and accurately determine the existence or non-existence of various categories.
[0059] Furthermore, since the linear multi-scale interaction enhancement module and the linear semantic relationship establishment module are both very usefully developed based on the state space model sense scene classification task, their computational complexity is linear, which makes the present invention more advantageous in terms of lightweight. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a flow chart of the present invention;
[0061] Figure 2 It is the overall structure diagram of the present invention;
[0062] Figure 3 This is a schematic diagram of the linear multi-scale interactive enhancement module proposed in the present invention;
[0063] Figure 4 A schematic diagram of a linear semantic relationship establishment module is proposed for the present invention. DETAILED DESCRIPTION
[0064] In order to further understand the content of the present invention, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention and are not intended to limit it.
[0065] Example 1
[0066] like Figure 1 As shown, a multi-label remote sensing scene classification method based on a state space model includes the following steps:
[0067] S1: Acquire multi-label remote sensing images;
[0068] S2: Extracting semantic embedding features from multi-label remote sensing images;
[0069] S3: Extracting visual features from multi-label remote sensing images;
[0070] S4: Construct a linear multi-scale interaction enhancement model, perform multi-scale scanning on the extracted visual features based on the linear multi-scale interaction enhancement model, extract and output the long-distance dependency relationship represented by the visual features;
[0071] S5: Construct a linear semantic relationship model. Based on the linear semantic relationship model, under the guidance of the long-range dependency relationship represented by visual features, the extracted semantic embedding features are combined to establish the relationship between different surface covers in multi-label remote sensing images;
[0072] S6: Based on the relationship between different land cover objects, a binary classifier is used to classify various land cover objects in multi-label remote sensing images.
[0073] Specifically, in S2, semantic embedding information is extracted from the candidate tag set (i.e., tag image) through a pre-trained language model to establish an accurate semantic association relationship; the premise of establishing an accurate semantic association relationship is to generate an accurate tag vector. To this end, the present invention uses a large-scale pre-trained language model Bert based on a tag candidate set L = {l1,l2,...,l c}, generate the semantic vector SE, the process can be defined as:
[0074] SE=[se1,se2,...,se c ]=[Φ bert (l1),Φ bert (l2),...,Φ bert (l c )],
[0075] where Φ bert (·) represents the text encoding process of Bert, Indicates dimension d l l j Semantic feature vector. It is worth noting that Bert only encodes text information during training and does not participate in the back propagation process. The purpose of this step is to reduce the training cost. Different from the adaptive semantic information mining commonly used in existing multi-label remote sensing scene classification methods, the label embedding obtained in the present invention combines the prior knowledge of large-scale language models, laying the foundation for further establishing semantic connections.
[0076] Specifically, in S3, the present invention uses a classic convolutional neural network ResNet18 to extract visual features from remote sensing images. ResNet18 is an excellent backbone for exploring the content of remote sensing images. It involves four residual modules that gradually extract local convolutional features, and deeper modules capture more comprehensive information to represent complex remote sensing scenes. The present invention selects the output of the fourth residual module as the extracted visual features, where H, W, and C represent the height, width, and number of channels, respectively.
[0077] Specifically, in S4, the present invention designs a linear multi-scale interactive enhancement model to cope with the challenges brought by the limited receptive field of view of ResNet18. Its main goal is to mine global information from remote sensing images of different scales and integrate these long-distance dependencies across multiple scales. In this way, the model can fully discover land cover of different scales and types in remote sensing images.
[0078] The structure of the linear multi-scale interaction enhancement model is as follows Figure 3 As shown in a. It contains a linear embedding layer, a deep convolution layer, a multi-scale state space model and another linear embedding layer. The linear embedding layer is intended to enhance the learnability of the model, the deep convolution layer focuses on further exploring local spatial information, and the multi-scale state space model layer is used to establish long-distance dependencies, which is the core of the linear multi-scale interaction enhancement model of the present invention. In addition, a regularization layer is applied before each linear embedding layer, a residual connection is used to prevent information loss, and a gated multi-layer perceptron (composed of a linear embedding layer and a SiLU activation function) is used to suppress irrelevant information.
[0079] Furthermore, the specific working process of the constructed linear multi-scale interactive enhancement model is as follows:
[0080] S41: receiving the extracted visual features as input features;
[0081] S42: Regularize the input features;
[0082] S43: performing linear transformation on the input features after regularization processing to obtain input features with reduced dimensions;
[0083] S44: performing a convolution operation on the input features with reduced dimensions, extracting local spatial information in the image, and generating a feature map;
[0084] S45: Based on the multi-scale state space model, down-sampling operation is performed on the generated feature map to extract feature embeddings of different scales. The bidirectional structure is used to transfer and update the feature embeddings of different scales to obtain enhanced feature embeddings of different scales. The enhanced feature embeddings of different scales are added together to obtain the long-distance dependency relationship of visual feature representation.
[0085] S46: Output the long-distance dependency relationship represented by the obtained visual features.
[0086] In S45, the information transmission method of the multi-scale state space model is as follows Figure 3 As shown in b. Considering the different sizes of land cover in remote sensing scenes, modeling long-range dependencies at a single scale is not enough to capture comprehensive global information. Therefore, the multi-scale state space model first extracts feature embeddings of different scales by:
[0087] P i = Downsample(P i-1 ),1<i≤N s ,
[0088] where Downsample(·) is the downsampling operation (i.e., 3D reshaping, 2×2 2D convolution, and flattening), N s represents the total number of scales, Representation scale 1<i≤N s The feature embedding at i =(H×W) / 2(i-1), and It is obtained by flattening F after passing through the regularization layer, linear layer, and deep convolutional layer.
[0089] The multi-scale state space model adopts a bidirectional structure, which is in line with the semantic understanding habits. For example, "ships" usually appear near "ports", and the appearance of "ports" is often accompanied by the appearance of "ships". In the forward process, high-scale embedded information is passed from top to bottom to the low-scale hidden state space, while in the backward process, low-scale clues are passed from bottom to top to the high-scale hidden state space. The long-distance dependencies between pixels at different positions are hidden in the state space. Specifically, the entire feature update process for the i-th scale can be expressed as,
[0090]
[0091] Where P i+1 and P i-1 Can be viewed as high-scale and low-scale information, and represents the initial hidden state space of the forward and backward processes, Avgpool(·) represents the average pooling operation, which compresses the three-dimensional feature map into a two-dimensional feature vector suitable for the state space model, and W ms-s6-f , represents the learnable parameters in the forward process, W ms-s6-b , The corresponding parameters represent the learnable parameters in the backward process, represents the enhanced feature embedding. Therefore, the above formula can be used to obtain the enhanced feature embedding at the i-th scale:
[0092] Use bilinear interpolation to unify the block sizes of different scales and add them together to get the output EF of the multi-scale state space model. This step can be expressed as,
[0093]
[0094] Where Upsample(·) represents the upsampling operation. With the help of the multi-scale state space model, our linear multi-scale interaction enhancement module achieves a comprehensive understanding of the long-range contextual relationships in remote sensing scene images. The global knowledge obtained not only fully reflects the connection between different locations within the same scale, but also includes the relationship between different scales. At the same time, the computational complexity of the linear multi-scale interaction enhancement module is O(2N s n), when N s When is a constant, the computational complexity is expressed as O(n), where n is the dimension of the input sequence. Therefore, the linear multi-scale interaction enhancement module maintains a linear characteristic in computational complexity.
[0095] In general, the linear multi-scale interaction enhancement model can effectively enhance the input feature F. It contains rich local convolution information, global information, and multi-scale information, and can highly represent the content contained in the input image.
[0096] Specifically, in S5, establishing semantic connections is crucial to reflecting the symbiotic relationship between land covers and improving the classification performance of multi-label remote sensing scene classification models. In addition, these semantic associations should be closely related to the visual content in the remote sensing image so that semantic embedding can effectively represent the different land covers contained therein. In order to meet the above requirements, the present invention models the relationship between various land covers under the guidance of visual features.
[0097] The framework for modeling linear semantic relationships is as follows: Figure 4 As shown in Figure 2. Similar to the linear multi-scale interaction enhancement model, it consists of a linear layer, a 1-D convolutional layer, a feature-guided semantic relationship state space model, and another linear layer. Regularization layers, residual connections, and gated multilayer perceptrons are also used. Here, we use a 1-D convolutional layer to adapt to the format of semantic embeddings and exploit a feature-guided semantic relationship state space model to mine the connections between different land cover targets.
[0098] Furthermore, the specific working process of the constructed linear semantic relationship model is as follows:
[0099] S51: receiving the extracted semantic embedding features as input features;
[0100] S52: Regularize the input semantic embedding features;
[0101] S53: performing linear transformation on the semantic embedding features after regularization processing to obtain semantic embedding features with reduced dimensions;
[0102] S54: performing a 1-dimensional convolution operation on the semantic embedding feature with reduced dimension to obtain a formatted semantic embedding feature;
[0103] S55: Based on the feature-guided semantic relationship state space model, under the guidance of the long-distance dependency relationship represented by the visual features, a bidirectional structure is used to transfer and update the formatted semantic embedding features, and the forward semantic embedding features are added to the backward semantic embedding features to obtain the enhanced semantic embedding features;
[0104] S56: Establishing the relationship between different land cover in multi-label remote sensing images based on the local features of enhanced semantic embedding features.
[0105] In S55, considering the non-causal nature of semantic information, the feature-guided semantic relationship state space model also analyzes the input semantic embedding features in a bidirectional manner to ensure that the linear semantic relationship establishment module thoroughly understands the relationship between different semantic categories. In order to ensure the relevance between semantic information and remote sensing images, we introduce the long-distance dependency represented by visual features as prior knowledge into the hidden state space. Under the guidance of visual information, the connection between different land covers can be gradually established. The above process can be expressed as follows,
[0106]
[0107] where Φ linear (·) represents linear embedding, and denote the initial hidden state space of the forward and backward processes of the feature-guided semantic relation state space model, W fgs-s6-f , and represents the learnable parameters in the forward process, W fgs-s6-b , and Represents the learnable parameters in the backward process, se` k Indicates k The semantic embedding obtained after the regularization layer, linear embedding layer and 1D convolution layer, ese k refers to the enhanced semantic embedding. Finally, the output of the feature-guided semantic relation state space model is expressed as In the above formula, Avgpool(Φ linear (F pyramid )) is the initial signal of the forward and backward processes. In the forward process, the initial hidden state space Encodes visual information. Since the state space model has a strong memory capacity, The visual information contained in can affect the subsequent hidden feature space 1≤k≤c is the formation of semantic associations. Similarly, the initial state space of the backward process It also contains visual information, making the backward hidden state space The semantic connections in 1≤k≤c can be established under the guidance of visual features. In this way, the semantic connections established by the linear semantic relationship establishment module can fully capture the relationship between different land covers in the input remote sensing image. In addition, similar to the linear multi-scale interactive enhancement module, the computational complexity of the linear semantic relationship establishment module is O(2n), which can be expressed as O(n), still maintaining the characteristics of linear complexity.
[0108] In summary, the linear semantic relationship building module enables the efficient network based on the state-space model to combine the content of remote sensing images and efficiently establish connections between semantics, thereby accurately reflecting the distribution and relationship of different surface covers in remote sensing images. By integrating global semantic context information through visual features, the linear semantic relationship building module provides a more comprehensive and accurate semantic understanding. This multi-level semantic association significantly enhances the classification performance and interpretation ability of the efficient network based on the state-space model in complex remote sensing image scenes.
[0109] Specifically, in S6, the semantic embedding output by the linear semantic relationship building module It contains rich local, global and multi-scale visual information, as well as highly representative semantic association information. These embeddings accurately represent the various land covers and their relationships in the input remote sensing images. Based on these embeddings, we first use multiple binary classifiers to predict the final category score CS = [cs1,cs2,...,cs c ],Right now,
[0110]
[0111] where csj is the category score of the jth class, Sigmoid(·) represents the sigmoid activation function, and Φ linear (·) denotes a linear learnable layer. Then, we adopt an asymmetric loss to optimize the model, which is defined as,
[0112]
[0113] Through the high-quality semantic embedding output by the linear semantic relationship building module, combined with multiple binary classifiers and an asymmetric loss function, the model is able to accurately predict the category scores of various land covers in remote sensing images and optimize its performance to suit specific task requirements.
[0114] Example 2
[0115] like Figure 2 As shown, a multi-label remote sensing scene classification system based on a state space model includes:
[0116] An acquisition module, used to acquire multi-label remote sensing images;
[0117] Semantic embedding extraction module, used to extract semantic embedding features in multi-label remote sensing images;
[0118] Visual feature extraction module, used to extract visual features from multi-label remote sensing images;
[0119] A linear multi-scale interaction enhancement model construction module is used to construct a linear multi-scale interaction enhancement model, perform multi-scale scanning on the extracted visual features based on the linear multi-scale interaction enhancement model, and extract and output the long-distance dependency relationship represented by the visual features;
[0120] The linear semantic relationship model building module is used to build a linear semantic relationship model. Based on the linear semantic relationship model, under the guidance of the long-distance dependency relationship represented by the visual features, the relationship between different surface covers in the multi-label remote sensing image is established in combination with the extracted semantic embedding features;
[0121] The classification module is used to classify various land cover in multi-label remote sensing images through binary classifiers based on the relationship between different land cover.
[0122] The present invention is mainly applied to multi-label remote sensing scene classification tasks, which is a basic task in remote sensing technology, and can automatically mark image categories, thereby liberating human resources. The present invention plays an important role in land resource utilization, urban planning, and environmental protection. In addition, the linear multi-scale interaction enhancement module and the linear semantic relationship establishment module proposed in the present invention have the characteristics of plug-and-play, can be integrated into many networks, and applied to different tasks. For example, the linear multi-scale interaction enhancement module can be used in hyperspectral classification, change detection, and semantic segmentation tasks, and enhances the feature representation ability of the model at a linear cost. The linear semantic relationship establishment module can be integrated into text retrieval tasks and image-text generation tasks to mine the correlation between different semantic information and enhance the interaction between the model's visual features and text features.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A multi-label remote sensing scene classification method based on a state space model, characterized in that: The steps include: Acquire multi-label remote sensing images; Extracting semantic embedding features from multi-label remote sensing images; Extract visual features from multi-label remote sensing images; Construct a linear multi-scale interaction enhancement model, perform multi-scale scanning on the extracted visual features based on the linear multi-scale interaction enhancement model, extract and output the long-distance dependencies represented by the visual features; Construct a linear semantic relationship model based on which, under the guidance of the long-range dependency relationship represented by visual features, the relationship between different land cover objects in multi-label remote sensing images is established in combination with the extracted semantic embedding features; Based on the relationship between different land covers, various land covers in multi-label remote sensing images are classified using a binary classifier.
2. According to claim 1, a multi-label remote sensing scene classification method based on a state space model is characterized in that: In the step of extracting semantic embedding information in the multi-label remote sensing image, the semantic embedding information in the multi-label remote sensing image is extracted by the pre-trained language model Bert to generate a semantic vector. The process is defined as: SE=[se1,se2,...,se c ]=[Φ bert (l1),Φ bert (l2),...,Φ bert (l c )], where Φ bert (·) represents the text encoding process of Bert, Indicates dimension d l l j Semantic feature vector.
3. The multi-label remote sensing scene classification method based on state space model according to claim 1 is characterized in that: In the step of extracting visual features from multi-label remote sensing images, a convolutional neural network is used to extract and output as the extracted visual features, where H, W, and C represent the height, width, and number of channels, respectively.
4. The multi-label remote sensing scene classification method based on state space model according to claim 1 is characterized in that: In the step of constructing a linear multi-scale interaction enhancement model, performing multi-scale scanning on the extracted visual features based on the linear multi-scale interaction enhancement model, and extracting and outputting the long-distance dependency relationship represented by the visual features, the specific working process of the constructed linear multi-scale interaction enhancement model is as follows: Receiving the extracted visual features as input features; Regularize the input features; Perform linear transformation on the input features after regularization to obtain input features with reduced dimensions; Perform convolution operations on the reduced-dimensional input features to extract local spatial information in the image and generate feature maps; Based on the multi-scale state space model, the generated feature map is downsampled to extract feature embeddings of different scales. The bidirectional structure is used to transfer and update the feature embeddings of different scales to obtain enhanced feature embeddings of different scales. The enhanced feature embeddings of different scales are added together to obtain the long-distance dependency relationship of visual feature representation. The long-distance dependencies represented by the obtained visual features are output.
5. The multi-label remote sensing scene classification method based on state space model according to claim 3 or 4, characterized in that: The downsampling operation is performed on the generated feature map to extract the feature embedding of different scales as follows: P i =Downsample(P i-1 ),1<i≤N s , Where Downsample(·) is the downsampling operation, N s represents the total number of scales, Representation scale 1<i≤N s The feature embedding at i =(H×W) / 2(i-1).
6. The multi-label remote sensing scene classification method based on state space model according to claim 5 is characterized in that: The bidirectional structure is used to transfer and update the feature embeddings of different scales to obtain enhanced feature embeddings of different scales. The process of transferring and updating the feature embedding of the i-th scale is expressed as follows: Where P i+1 and P i-1 Can be viewed as high-scale and low-scale information, and represents the initial hidden state space of the forward and backward processes, Avgpool(·) represents the average pooling operation, which compresses the three-dimensional feature map into a two-dimensional feature vector suitable for the state space model, and W ms-s6-f , represents the learnable parameters in the forward process, W ms-s6-b , The corresponding parameters represent the learnable parameters in the backward process, Represents enhanced feature embedding; The enhanced feature embedding at the i-th scale is:
7. The multi-label remote sensing scene classification method based on state space model according to claim 1 is characterized in that: The linear semantic relationship establishment model is constructed based on the linear semantic relationship establishment model. Under the guidance of the long-distance dependency relationship represented by the visual features, the extracted semantic embedding features are combined to establish the connection between different surface covers in the multi-label remote sensing image. The specific working process of the constructed linear semantic relationship establishment model is as follows: Receiving the extracted semantic embedding features as input features; Regularize the input semantic embedding features; Perform linear transformation on the regularized semantic embedding features to obtain semantic embedding features with reduced dimensions; Perform a 1D convolution operation on the reduced-dimensional semantic embedding features to obtain formatted semantic embedding features; Based on the feature-guided semantic relationship state space model, under the guidance of the long-distance dependency relationship represented by the visual features, a bidirectional structure is used to transfer and update the formatted semantic embedding features, and the forward semantic embedding features are added to the backward semantic embedding features to obtain the enhanced semantic embedding features. The association between different land cover objects in multi-label remote sensing images is established based on the local features of the enhanced semantic embedding features.
8. The multi-label remote sensing scene classification method based on state space model according to claim 7 is characterized in that: The feature-guided semantic relationship state space model adopts a bidirectional structure to transfer and update the formatted semantic embedding features under the guidance of the long-distance dependency relationship represented by the visual features, and adds the forward semantic embedding features to the backward semantic embedding features to obtain the enhanced semantic embedding features. The specific working process is as follows: where Φ linear (·) represents linear embedding, and denote the initial hidden state space of the forward and backward processes of the feature-guided semantic relation state space model, W fgs-s6-f , and represents the learnable parameters in the forward process, W fgs-s6-b , and represents the learnable parameters in the backward process, Indicates k The semantic embedding obtained after the regularization layer, linear embedding layer and 1D convolution layer, ese k Refers to the enhanced semantic embedding; The output of the feature-guided semantic relation state space model is expressed as 9. The multi-label remote sensing scene classification method based on state space model according to claim 1, characterized in that: In the step of classifying various land covers in the multi-label remote sensing image by a binary classifier based on the relationship between different land covers, the specific representation of the binary classifier is as follows: where csj is the category score of the jth class, Sigmoid(·) represents the sigmoid activation function, and Φ linear (·) denotes a linear learnable layer.
10. A multi-label remote sensing scene classification system based on a state space model, characterized in that: include: An acquisition module, used to acquire multi-label remote sensing images; Semantic embedding extraction module, used to extract semantic embedding features in multi-label remote sensing images; Visual feature extraction module, used to extract visual features from multi-label remote sensing images; A linear multi-scale interaction enhancement model construction module is used to construct a linear multi-scale interaction enhancement model, perform multi-scale scanning on the extracted visual features based on the linear multi-scale interaction enhancement model, and extract and output the long-distance dependency relationship represented by the visual features; The linear semantic relationship model building module is used to build a linear semantic relationship model. Based on the linear semantic relationship model, under the guidance of the long-distance dependency relationship represented by the visual features, the relationship between different surface covers in the multi-label remote sensing image is established in combination with the extracted semantic embedding features; The classification module is used to classify various land cover in multi-label remote sensing images through binary classifiers based on the relationship between different land cover.
Citation Information
Cited By
Geographic multi-factor classification and fusion modeling method based on state space model
CN120336444A