Multi-label remote sensing scene image classification method and related device
By adopting a feature fusion network based on semantic assistance in the multi-label remote sensing scene image classification, the problems of inaccuracy of semantic information and low credibility in the prior art are solved, and high-precision multi-label classification is achieved.
Patent Information
- Application Number
- CN202411994499.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-27
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
Due to the limitation of the supervision signal and the randomness of the deep learning process, the existing multi-label remote sensing scene image classification method may have inaccuracy in the semantic information extracted from the remote sensing image, which in turn reduces the credibility of the established semantic relationship and seriously affects the classification performance.
A multi-label remote sensing scene image classification method is adopted to classify through a pre-constructed feature fusion network based on semantic assistance. The network includes a dual-scale feature extraction module, a local semantic enhancement module, a cross-scale interactive attention module and a classification module, and introduces supervised semantic loss and semantic-based contrast learning loss to enhance the reliability and discrimination of semantic information.
By effectively integrating feature learning and semantic relationship modeling, the model's understanding of complex content of remote sensing images is improved, and the accuracy and discrimination ability of multi-label classification are significantly improved, ensuring the accuracy of classification results.
Smart Images

Figure CN119942186A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image processing, and in particular relates to a multi-label remote sensing scene image classification method and related devices. Background Art
[0002] With the rapid development of satellite technology, the content and resolution of remote sensing scene images have been significantly improved. With the improvement of image resolution, the same remote sensing scene image often contains multiple land cover targets, such as docks, vehicles and buildings. In this context, a single label cannot fully express the complex content of the remote sensing scene. Therefore, the classification task of multi-label remote sensing scene images that comprehensively characterize the land cover targets in the image through multiple semantic labels is attracting more and more researchers' attention.
[0003] At present, multi-label remote sensing scene image classification methods are mainly divided into two categories: feature enhancement-based methods and semantic association-based methods; among them, feature enhancement-based methods mainly combine advanced technologies such as attention mechanism and multi-scale information retrieval to enhance the feature learning ability of the model, so as to more effectively discover different targets in the image; semantic association-based methods use the inherent correlation between semantic labels to enhance the model's ability to understand the scene; for example, docks usually coexist with water, while deserts are usually not related to water bodies. The introduction of semantic relationships can help the model to perform multi-label classification more accurately and improve its interpretation; however, in the existing multi-label remote sensing scene image classification methods, due to the limited supervision signals and the randomness of the deep learning process, the semantic information extracted from the remote sensing images may be inaccurate, which in turn reduces the credibility of the established semantic relationships and seriously affects the classification performance. Summary of the invention
[0004] In view of the technical problems existing in the prior art, the present invention provides a multi-label remote sensing scene image classification method and related devices to solve the technical problem that the existing multi-label remote sensing scene image classification method may have inaccuracy in the semantic information extracted from the remote sensing image due to the limited supervision signal and the randomness of the deep learning process, thereby reducing the credibility of the established semantic relationship and seriously affecting the classification performance.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is:
[0006] The present invention provides a multi-label remote sensing scene image classification method, comprising:
[0007] Acquire remote sensing images to be classified;
[0008] Inputting the remote sensing image to be classified into a pre-built semantic-assisted feature fusion network for classification, and obtaining a multi-label remote sensing scene image classification result of the remote sensing image to be classified;
[0009] The pre-built semantic-assisted feature fusion network includes a dual-scale feature extraction module, a local semantic enhancement module, a cross-scale interactive attention module and a classification module;
[0010] The dual-scale feature extraction module is used to extract features from the remote sensing image to be classified, so as to capture low-level features and high-level features in the remote sensing image to be classified and obtain a feature map. and feature map
[0011] The local semantic enhancement module is used to enhance the feature map and feature map Semantic information is extracted respectively to obtain feature maps and feature map Wherein, the local semantic enhancement module introduces supervised semantic loss and semantic-based contrastive learning loss;
[0012] The cross-scale interactive attention module is used to and feature map Perform information exchange to obtain classification features and categorical features
[0013] The classification module is used to classify and categorical features Classification results of different scales are generated, and the classification results of different sizes are integrated based on the classification loss to obtain the final score as the multi-label remote sensing scene image classification result of the remote sensing image to be classified.
[0014] Furthermore, the backbone of the dual-scale feature extraction module is a ResNet18 with four residual layers; This is the feature map output by the last layer of ResNet18, which contains four residual layers; feature map It is the output feature map of the penultimate layer of ResNet18, which contains four residual layers.
[0015] Furthermore, the local semantic enhancement module includes a first lightweight semantic enhancement unit Transformer1 and a second lightweight semantic enhancement unit Transformer2 in parallel;
[0016] Among them, the first lightweight semantic enhancement unit Transformer1 is used to transform the feature map Extract semantic information and obtain feature maps The second lightweight semantic enhancement unit Transformer2 is used to Extract semantic information and obtain feature maps
[0017] Furthermore, the first lightweight semantic enhancement unit Transformer1 and the second lightweight semantic enhancement unit Transformer2 have the same structure, both of which include a supervised semantic discriminator, a contrastive learning module, an attention-based feature update mechanism module, a regularization module and a multi-layer convolution block;
[0018] A supervised semantic discriminator is used to extract or feature map Extract category-aware semantic information separately to obtain class-specific semantic information vectors
[0019] Contrastive learning module, used for semantic-based contrastive learning loss, which is derived from feature maps The class-specific semantic information vector extracted from the feature map The class-specific semantic information vector extracted from the image is extended to the semantic level through contrastive learning to obtain highly discriminative semantic information.
[0020] Attention-based feature update mechanism module for building feature maps or feature map The relationship between the feature pixels and the semantic information with high discriminability is used to obtain the enhanced feature map X e ; Among them, feature pixels are used as a medium to establish the connection between semantics;
[0021] The regularization module is used to enhance the feature map X e Perform regularization processing to obtain a regularized feature map;
[0022] Multi-layer convolution blocks are used to mine local information on regularized feature maps to obtain feature maps or feature map
[0023] Furthermore, the supervised semantic discriminator includes two parallel discriminative branches;
[0024] In the first discriminant branch, in the feature map or feature map Apply the sigmoid function and 1×1 convolution layer to generate the class attention map For feature maps or feature map and class attention map Each channel of is element-wise multiplied, followed by global average pooling to obtain the class-specific semantic information vector
[0025] In the second discriminant branch, the feature map or feature map After global average pooling and 1×1 convolution layer, the class activation vector is obtained Among them, the class activation vector Used to generate class attention maps based on supervised semantic loss process.
[0026] Furthermore, supervise the semantic loss as follows:
[0027]
[0028]
[0029] in, is the supervised semantic loss; BCE(·) is the binary cross entropy loss function; The feature map The i-th element in the class activation vector; y i is the jth label in the label vector; r i 2 The feature map For the i-th element in the class activation vector; N is the total number of elements in the class activation vector; r is the class activation vector; y is the label vector; σ(·) is the activation function; r j is the jth element in the class activation vector; c is the number of categories.
[0030] Furthermore, the semantic-based contrastive learning loss is as follows:
[0031]
[0032]
[0033] {P(a)=SD j \{a}}
[0034] SD={SD1,...,SD c}
[0035]
[0036]
[0037] in, is the semantic-based contrastive learning loss; a is the anchor sample with semantic label j in the single-label semantic dataset SD constructed based on the labels of remote sensing scenes; |P(a)| is the total number of positive sample sets of a; P(a) is the positive sample set of a; is the semantic-based contrastive learning loss corresponding to the anchor sample with semantic label j in SD; d(*) is the cosine similarity; p∈P(a) is the relevant positive sample; β is the temperature coefficient; n∈N(a) is the relevant negative sample; SD j is the sub-dataset of category j; is the indicator function; is the j-th class sample in the i-th element in the dual-scale semantic vector; DSF i is the i-th element in the dual-scale semantic vector, which is obtained by converting The class-specific semantic information vector extracted from the feature map The class-specific semantic information vector extracted from is concatenated.
[0038] Furthermore, the classification loss is specifically:
[0039]
[0040] in, is the classification loss: FS i For the final score.
[0041] The present invention also provides a multi-label remote sensing scene image classification system, comprising:
[0042] An image acquisition module is used to acquire remote sensing images to be classified;
[0043] A classification module is used to input the remote sensing image to be classified into a pre-built semantic-assisted feature fusion network for classification, so as to obtain a multi-label remote sensing scene image classification result of the remote sensing image to be classified;
[0044] The pre-built semantic-assisted feature fusion network includes a dual-scale feature extraction module, a local semantic enhancement module, a cross-scale interactive attention module and a classification module;
[0045] The dual-scale feature extraction module is used to extract features from the remote sensing image to be classified, so as to capture low-level features and high-level features in the remote sensing image to be classified and obtain a feature map. and feature map
[0046] The local semantic enhancement module is used to enhance the feature map and feature map Semantic information is extracted respectively to obtain feature maps and feature map Wherein, the local semantic enhancement module introduces supervised semantic loss and semantic-based contrastive learning loss;
[0047] The cross-scale interactive attention module is used to and feature map Perform information exchange to obtain classification features and categorical features
[0048] The classification module is used to classify and categorical features Classification results of different scales are generated, and the classification results of different sizes are integrated based on the classification loss to obtain the final score as the multi-label remote sensing scene image classification result of the remote sensing image to be classified.
[0049] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the multi-label remote sensing scene image classification method are implemented.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] The multi-label remote sensing scene image classification method provided by the present invention can fully understand the complex content in the remote sensing image by mining multi-scale clues, exploring diverse land cover targets, analyzing the relationship between them and enhancing feature representation, thereby realizing multi-label classification based on a semantically assisted feature fusion network; specifically, feature extraction is performed on the remote sensing image to be classified through a dual-scale feature extraction module to capture low-level features and high-level features in the remote sensing image to be classified, aiming to further mine semantic association information based on multi-scale high-level features to enhance the information expression ability of high-level features and realize the effective fusion of feature learning and semantic modeling; supervised semantic loss and semantic-based contrastive learning loss are introduced in a local semantic enhancement module, and a deep supervision mechanism is introduced, so that the network model can adaptively extract more accurate semantic information from the feature map, thereby improving the reliability of semantic relationships; wherein, contrastive learning is extended to the semantic level, and the distinguishability of semantic information is enhanced to improve the recognition ability of the model for unique targets in complex scenes and the overall classification performance; the present invention can effectively aggregate similar features and simultaneously pull out heterogeneous features, thereby significantly improving the model's discrimination ability and perception ability for fine-grained features, thereby ensuring the accuracy of the classification results.
[0052] The multi-label remote sensing scene image classification system and computer-readable storage medium provided by the present invention have all the advantages of the above-mentioned multi-label remote sensing scene image classification method. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0054] Figure 1 A flowchart of the multi-label remote sensing scene image classification method provided in Example 1;
[0055] Figure 2 It is a structural block diagram of the semantic-assisted feature fusion network pre-built in Example 1;
[0056] Figure 3 is a structural block diagram of the lightweight semantic enhancement unit Transforme in Example 1; wherein, Figure 3 (a) is the overall structural diagram of the lightweight semantic enhancement unit Transformer. Figure 3 (b) is the structural diagram of the supervised semantic discriminator; Figure 3 (c) is a structural diagram of a multi-layer convolutional block;
[0057] Figure 4 : is a structural block diagram of the cross-scale interactive attention module in Example 1;
[0058] Figure 5 This is a structural block diagram of the multi-label remote sensing scene image classification system provided in Example 1. DETAILED DESCRIPTION
[0059] In order to make the technical problems, technical solutions and beneficial effects solved by this application clearer, the technical solutions in the embodiments of this application will be described clearly and completely in combination with the drawings in the embodiments of this application; obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0060] Example 1
[0061] As attached Figure 1 As shown, this embodiment 1 provides a multi-label remote sensing scene image classification method, comprising the following steps:
[0062] Step 1: Obtain the remote sensing image to be classified. It should be noted that the present application does not limit the means or method for obtaining the remote sensing image.
[0063] Step 2: input the remote sensing image to be classified into a pre-built semantic-assisted feature fusion network for classification, and obtain a multi-label remote sensing scene image classification result of the remote sensing image to be classified.
[0064] As attached Figure 2 As shown, the pre-constructed semantic-assisted feature fusion network includes a dual-scale feature extraction module, a local semantic enhancement module, a cross-scale interactive attention module and a classification module.
[0065] The dual-scale feature extraction module is used to extract features from the remote sensing image to be classified, so as to capture low-level features and high-level features in the remote sensing image to be classified and obtain a feature map. and feature map The local semantic enhancement module is used to enhance the feature map and feature map Semantic information is extracted respectively to obtain feature maps and feature map The local semantic enhancement module introduces supervised semantic loss and semantic-based contrastive learning loss. The cross-scale interactive attention module is used to and feature map Perform information exchange to obtain classification features and categorical features The classification module is used to classify and categorical features Classification results of different scales are generated, and the classification results of different sizes are integrated based on the classification loss to obtain the final score as the multi-label remote sensing scene image classification result of the remote sensing image to be classified.
[0066] In this embodiment 1, the dual-scale feature extraction module is intended to extract multi-scale features from the remote sensing image to be classified; since remote sensing images usually contain land cover targets with the same semantics but different scales, it is difficult for single-scale features to fully represent complex content; at the same time, considering that the multi-scale features learned by the hierarchical convolutional neural network (CNN) contain rich information, that is, shallow-scale features focus on low-level information such as texture and color, while deep-scale features involve high-level context clues; the dual-scale feature extraction module uses ResNet18 containing four residual layers as the backbone to capture low-level and high-level features in the input remote sensing image; since the feature maps output by the first two layers of ResNet18 containing four residual layers contain redundant information, the dual-scale feature extraction module uses the outputs of the last two layers; specifically, the feature map output by the last layer of ResNet18 containing four residual layers is marked as the feature map The output feature map of the penultimate layer of ResNet18, which contains four residual layers, is marked as the feature map Feature map This is the feature map output by the last layer of ResNet18, which contains four residual layers. is the output feature map of the penultimate layer of ResNet18 containing four residual layers; H1, W1 and C1 are the feature maps The height, width and number of channels of H2, W2 and C2 are feature maps respectively. The height, width, and number of channels of the
[0067] In this embodiment 1, the local semantic enhancement module includes a first lightweight semantic enhancement unit Transformer1 and a second lightweight semantic enhancement unit Transformer2 in parallel, as shown in the attached Figure 2 As shown; the first lightweight semantic enhancement unit Transformer1 is used to Extract semantic information and obtain feature maps The second lightweight semantic enhancement unit Transformer2 is used to Extract semantic information and obtain feature maps Among them, the local semantic enhancement module explicitly establishes the relationship between the pixel area of the remote sensing image and various semantics, such as Figure 2 The solid line in the figure indicates that the local semantic enhancement module implicitly establishes the connection between different semantics, as shown in the attached figure. Figure 2 The dotted line in .
[0068] The first lightweight semantic enhancement unit Transformer1 and the second lightweight semantic enhancement unit Transformer2 have the same structure, both of which include a supervised semantic discriminator, a contrastive learning module, an attention-based feature update mechanism module, a regularization module and a multi-layer convolution block; it should be noted that the weights between the first lightweight semantic enhancement unit Transformer1 and the second lightweight semantic enhancement unit Transformer2 are not shared.
[0069] A supervised semantic discriminator is used to extract or feature map Extract category-aware semantic information separately to obtain class-specific semantic information vectors Contrastive learning module, used for semantic-based contrastive learning loss, which is derived from feature maps The class-specific semantic information vector extracted from the feature map The class-specific semantic information vector extracted from the image is extended to the semantic level by contrastive learning to obtain highly discriminative semantic information; the attention-based feature update mechanism module is used to establish the feature map or feature map The relationship between the feature pixels and the semantic information with high discriminability is used to obtain the enhanced feature map X e ; Among them, feature pixels are used as a medium to establish the connection between semantics; the regularization module is used to enhance the feature map X e Regularization is performed to obtain a regularized feature map; a multi-layer convolution block is used to perform local information mining on the regularized feature map to obtain a feature map or feature map
[0070] The supervised semantic discriminator consists of two parallel discriminative branches; in the first discriminative branch, or feature map Apply the sigmoid function and 1×1 convolution layer to generate the class attention map For feature maps or feature map and class attention map Each channel of is element-wise multiplied, followed by global average pooling to obtain the class-specific semantic information vector In the second discriminant branch, the feature map or feature map After global average pooling and 1×1 convolution layer, the class activation vector is obtained Among them, the class activation vector Used to generate class attention maps based on supervised semantic loss process.
[0071] The following takes one of the lightweight semantic enhancement units Transformer (LSET) as an example, and unifies the input data of the lightweight semantic enhancement unit into a local feature map. The first lightweight semantic enhancement unit Transformer1 and the second lightweight semantic enhancement unit Transformer2 are described in detail as follows:
[0072] As attached Figure 3 As shown in (a), the lightweight semantic enhancement unit Transformer includes a supervised semantic discriminator (SSFD), a contrastive learning module, an attention-based feature update mechanism module (AFUM) and a multi-layer convolutional block (MCB); SSFD is used to extract the local feature map from the local feature map. In order to accurately extract category-aware semantic information, ADUM is used to establish the relationship between feature pixels and semantics. At the same time, feature pixels are used as a medium to indirectly establish the connection between various semantics. MCB is used to further utilize local information.
[0073] As attached Figure 3 As shown in (b), SSFD consists of two parallel discriminative branches.
[0074] In the first discriminant branch, first in the local feature map Apply the sigmoid function and 1×1 convolution layer to generate the class attention map Among them, the class attention map The i-th one in represents the local feature map The contribution score of the i-th category; then, the local feature map and class attention map Element multiplication is performed on each channel of ; finally, global average pooling (GAP) is performed to construct a unique semantic information vector for each class and obtain a class-specific semantic information vector
[0075] It should be noted that the process of the first discrimination branch is formulated as follows:
[0076] CAM=σ(conv 1×1 (X))=[CS 1 ,...,CS c ]
[0077]
[0078] Among them, conv 1×1 (·) is 1×1 convolution; σ(·) is the sigmoid function; is element multiplication; GAP(·) is global average pooling; Contribute score to the jth class; is the semantic information vector of the jth class; is the class-specific semantic information vector.
[0079] In the second discriminant branch, the local feature map After global average pooling and 1×1 convolution layer, the class activation vector is obtained Among them, the class activation vector Used to generate class attention maps based on supervised semantic loss process.
[0080] After acquiring the semantics, the connections between the acquired semantics are modeled; the traditional means adopts the application of the sequence relationship method on the semantic information vector to realize the relationship, but it ignores the information hidden in the feature pixels, thereby limiting the contribution of the semantic relationship to the scene understanding; in order to overcome the above shortcomings, in this embodiment 1, AFUM is used to establish semantic relationships with the help of feature pixels, so that in the process of establishing the connection, the global semantics and the information in the local feature pixels can be considered at the same time, ensuring that the generated semantic association is good and reliable for interpreting the remote sensing scene.
[0081] Specifically, the AFUM is applied to the local feature map and highly discriminative semantic information to directly explore the relationship between feature pixels and semantics; wherein the formulation of the AFUM is as follows:
[0082]
[0083] Among them, φ(·) is the softmax function; W V , W Q and W K are three learnable projection matrices; X i,j is the local feature map The i-th row and j-th column element of i ' ,j To enhance the feature map X e The element in the i-th row and j-th column of ; is the scaling factor; M i,j is the weight matrix, including X i,j Correlation with various semantics.
[0084] It should be noted that when the enhanced feature map X is obtained e When , since the relationship between feature pixels and semantics has been determined, the connection between various semantics can also be obtained indirectly.
[0085] Next, a multi-layer convolutional block is used to mine local information on the regularized feature map. In MCB, it consists of two pairs of consecutive 3×3 convolutional layers and normalization layers, as shown in the attached figure. Figure 3 c; In addition, a residual structure is used to reduce information loss; After MCB, the output of LSET can be obtained
[0086] Based on the above examples, the feature map and feature map After two LSETs, it can be mapped to the feature map or feature map It contains rich local, pixel-wise semantics and semantic-semantic clues.
[0087] In this embodiment 1, considering that features of different scales contain different clues, a cross-scale interactive attention module (CIAM) is introduced to enhance clues of different scales, and the information of different scales is enhanced and supplemented through interactive attention maps; specifically, the valuable knowledge of different scales is interacted with each other to fully explore the content of remote sensing images at different scales.
[0088] As attached Figure 4 As shown in FIG, the working process of the cross-scale interactive attention module is as follows:
[0089] First, the feature map or feature map Simultaneously perform maximum pooling and average pooling to reduce the channel dimension and obtain feature maps or feature map
[0090] Next, the feature map or feature map Perform two 3×3 convolutions and use a sigmoid function to obtain the attention scores of feature map pixels at each scale in the spatial dimension. and
[0091] Afterwards, the attention score After the nearest neighbor interpolation of the upsampling operation, the attention score Add, then add to the feature map Multiply to get the classification features At the same time, the attention score After the average pooling of the downsampling operation, the attention score Add, then add to the feature map Multiply to get the classification features
[0092] The formulation of the cross-scale interactive attention module is as follows:
[0093]
[0094] Among them, AP(·) is the average pooling of the downsampling operation; NNI(·) is the nearest neighbor interpolation of the upsampling operation.
[0095] It should be noted that the cross-scale interactive attention module realizes effective information interaction between the two scale features, which promotes the collaboration between scales to refocus on the areas that were ignored during the forward propagation process; at the same time, it can deepen the understanding of the feature content and help reduce the impact of redundant information.
[0096] In this embodiment 1, the classification module includes two classifiers for classifying based on the classification features. and categorical features Generate classification results; as shown in the attached Figure 2 As shown in the figure, each classifier includes a 1×1 convolution layer and a global average pooling layer; the number of output channels of the convolution layer is equal to the number of categories c; finally, the results of the two classifiers are added and activated through the sigmoid function to obtain the final score of each class
[0097] Since the correctness of semantics is crucial for multi-label scene understanding, it is directly related to the generated class attention map CAM; therefore, this embodiment 1 introduces supervised semantic loss in the supervised semantic discriminator to ensure that the SSFD in LSET can achieve the correct semantics corresponding to the class activation vector.
[0098] Assume that the remote sensing scenes in a batch have been mapped to and Then, by reducing the class activation vector and the labels {y1,…,y N} to improve the accuracy of the class attention map CAM.
[0099] Specifically, the binary cross entropy (BCE) loss function is used, and the supervised semantic loss is expressed as:
[0100]
[0101]
[0102] in, is the supervised semantic loss; BCE(·) is the binary cross entropy loss function; The feature map The i-th element in the class activation vector; y i is the jth label in the label vector; The feature map For the i-th element in the class activation vector; N is the total number of elements in the class activation vector; r is the class activation vector; y is the label vector; σ(·) is the activation function; r j is the jth element in the class activation vector; c is the number of categories.
[0103] It should be noted that supervised semantic loss can push each channel of the feature map to a specific category, thereby obtaining a class activation map that can accurately highlight the interest area of each category to ensure the certainty of the extracted semantics and lay a solid foundation for establishing clear semantic connections.
[0104] In order to overcome the problem of low intra-class similarity and high inter-class similarity of land cover in remote sensing scenes, and to improve the model's ability to distinguish low intra-class similarity and high inter-class similarity of land cover in remote sensing scenes, this embodiment 1 proposes a semantic-based contrastive learning loss under the framework of contrastive learning, and introduces it into the local semantic enhancement module; because traditional contrastive learning usually adjusts sample distances at the image level; however, multiple semantics in the same image limit its development in multi-label remote sensing scene classification; therefore, the semantic-based contrastive learning loss is based on the class-specific semantic information vector extracted from SSFD to extend contrastive learning to the semantic level; specifically, by treating information vectors with the same semantics as positive pairs and information vectors with different semantics as negative pairs, the gap between positive / negative pairs is narrowed / widened to improve the discrimination of the learned semantics.
[0105] For semantic-based contrastive learning loss, an example is given below:
[0106] First, for a batch of remote sensing scenes D, assume that its semantic and The SSFD of two LSETs has been extracted; considering the rich scale clues, the two sets of semantics are connected to generate a dual-scale semantic vector {DSF1,...,DSF N};in, DSF i is the i-th element in the dual-scale semantic vector, which is obtained by converting The class-specific semantic information vector extracted from the feature map The class-specific semantic information vector extracted from is concatenated.
[0107] Then, a single-label semantic dataset SD is constructed based on the labels of the remote sensing scenes in the batch. The formula for constructing a single-label semantic dataset SD based on the labels of the remote sensing scenes is as follows:
[0108] SD={SD1,...,SD c}
[0109]
[0110] Among them, SD j is a sub-dataset of category j, that is, the semantics in the sub-dataset is marked as j; l ζ ∈{0,1} is the indicator function. If the condition ζ is satisfied, then l ζ =1; otherwise, l ζ =0;dsf i j is the j-th class sample in the i-th element in the dual-scale semantic vector.
[0111] Considering the lack of certain categories in remote sensing images, the extracted semantics may tend to be background; in order to alleviate the above problem, in this embodiment 1, the semantics in the remote sensing image are selected only according to the labels, and then the semantic contrast learning loss is used to minimize / maximize the distance between the semantics belonging to the same / different sub-datasets.
[0112] Specifically, the definition of semantic-based contrastive learning loss is as follows:
[0113]
[0114]
[0115] {P(a)=SD j \{a}}
[0116] {N(a)=SD\SD j}
[0117] in, is the semantic-based contrastive learning loss; a is the anchor sample with semantic label j in the single-label semantic dataset SD constructed based on the labels of remote sensing scenes; |P(a)| is the total number of positive sample sets of a; P(a) is the positive sample set of a; is the semantic-based contrastive learning loss corresponding to the anchor sample with semantic label j in SD; d(*) is the cosine similarity; p∈P(a) is the relevant positive sample; β is the temperature coefficient; n∈N(a) is the relevant negative sample.
[0118] It should be noted that the semantic-based contrastive learning loss is a semantic-level loss function, which can help the model capture highly discriminative semantic information, so as to better understand the differences and associations between various semantic categories, so that the model can have a deeper understanding of the characteristics and distribution of different land covers in remote sensing images, thereby improving the accuracy of the model in identifying different land covers.
[0119] In this embodiment 1, the BCE loss function is used to construct the classification loss to narrow the gap between the result of the classifier and the label y; wherein the formula of the classification loss is as follows:
[0120]
[0121] in, is the classification loss: FS i For the final score.
[0122] It is worth emphasizing that in this embodiment 1, supervised semantic loss, semantic-based contrastive learning loss and classification loss are introduced; wherein, supervised semantic loss and semantic-based contrastive learning loss ensure the reliability and discrimination of semantic information extracted by the local semantic enhancement module; classification loss guides the classification module to complete the multi-label remote sensing scene classification task; therefore, the overall loss function of the semantic-assisted feature fusion network is defined as:
[0123]
[0124] in, is the overall loss function of the semantic-assisted feature fusion network; λ is a hyperparameter; classification loss and supervised semantic loss For classification, the semantic-based contrastive learning loss For contrastive learning; by using semantic-based contrastive learning loss Add hyperparameters To balance different items; in the classification task, the three loss functions work together to ensure that the semantic-assisted feature fusion network can correctly extract semantic information, improve the model's ability to discriminate different land covers, and complete the multi-label remote sensing scene classification task.
[0125] The multi-label remote sensing scene image classification method described in this embodiment 1 can fully understand the complex content in the remote sensing image by mining multi-scale clues, exploring diverse land cover targets, analyzing the relationship between targets and enhancing feature representation, thereby realizing multi-label classification. Among them, the local semantic enhancement module in the feature fusion network based on semantic assistance can extract semantic feature vectors specific to the land cover target from the feature map, implicitly establish the connection between different semantics, and reintegrate them into the local convolutional features. In addition, the cross-scale interactive attention module in the feature fusion network based on semantic assistance is used to promote information interaction between features of different scales. By complementing each other's regions, the cross-scale interactive attention module reduces the information imbalance caused by scale changes and mines more valuable information.
[0126] In this embodiment 1, under the deep supervision and contrastive learning paradigm, supervised semantic loss and semantic-based contrastive learning loss are applied to the local semantic enhancement module to help the semantic-assisted feature fusion network obtain accurate and discriminative semantic information from remote sensing images; and the supervised semantic loss and semantic-based contrastive learning loss are combined with the classification loss, so that the semantic-assisted feature fusion network can be properly trained and obtain accurate multi-label classification results.
[0127] Example 2
[0128] As attached Figure 5 As shown, this embodiment 2 provides a multi-label remote sensing scene image classification system, including an image acquisition module and a classification module.
[0129] The image acquisition module is used to acquire the remote sensing image to be classified; the classification module is used to input the remote sensing image to be classified into a pre-built feature fusion network based on semantic assistance for classification, so as to obtain the multi-label remote sensing scene image classification result of the remote sensing image to be classified.
[0130] In this embodiment 2, the pre-constructed semantic-assisted feature fusion network includes a dual-scale feature extraction module, a local semantic enhancement module, a cross-scale interactive attention module and a classification module.
[0131] The dual-scale feature extraction module is used to extract features from the remote sensing image to be classified, so as to capture low-level features and high-level features in the remote sensing image to be classified and obtain a feature map. and feature map The local semantic enhancement module is used to enhance the feature map and feature map Semantic information is extracted respectively to obtain feature maps and feature map The local semantic enhancement module introduces supervised semantic loss and semantic-based contrastive learning loss; the cross-scale interactive attention module is used to and feature map Perform information exchange to obtain classification features and categorical features The classification module is used to classify and categorical features Classification results of different scales are generated, and the classification results of different sizes are integrated based on the classification loss to obtain the final score as the multi-label remote sensing scene image classification result of the remote sensing image to be classified.
[0132] It should be noted that the specific structure and principle of the pre-built semantic-assisted feature fusion network are described in detail in the corresponding content of the above-mentioned embodiment 1, and will not be repeated here.
[0133] Example 3
[0134] This embodiment 3 also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the multi-label remote sensing scene image classification method are implemented.
[0135] If the module / unit integrated in the multi-label remote sensing scene image classification method device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0136] Based on this understanding, the present embodiment 3 implements all or part of the processes in the above-mentioned multi-label remote sensing scene image classification method, and can also be completed by instructing related hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, the steps of the above-mentioned multi-label remote sensing scene image classification method can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc.
[0137] The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0138] It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media does not include electrical carrier signals and telecommunication signals.
[0139] The multi-label remote sensing scene image classification method provided by the present invention combines feature learning with semantic relationship modeling learning compared with the existing multi-label remote sensing scene classification method, and proposes two core modules: a local semantic enhancement module and a cross-scale interactive attention module; wherein the local semantic enhancement module extracts significant semantic information in the image and establishes associations between different semantics by deep mining of local convolutional feature maps; it not only helps the model understand the complex content of the remote sensing image from a semantic level, but also enhances the modeling ability of the intrinsic connection between multiple labels; the cross-scale interactive attention module introduces a selective mechanism at the feature level, captures key information related to high-level features in shallow features, suppresses interference from redundant and irrelevant features, and realizes cross-scale interactive fusion of feature information, thereby generating a more comprehensive and more representative feature description; finally, the features used by the model for classification fully integrate local information, global information and semantic association information, providing a reliable basis for high-precision classification.
[0140] Secondly, thanks to the excellent performance of the local semantic enhancement module in accurately mining semantic information, the present invention fully considers the characteristics of remote sensing images with high intra-class similarity and low inter-class similarity, and extends the idea of contrastive learning to the semantic level to further enhance the model's ability to distinguish features of different categories; by optimizing the feature distribution in the semantic space, it can effectively aggregate similar features and distance heterogeneous features, thereby significantly improving the model's discrimination ability and perception of fine-grained features.
[0141] In the present invention, the multi-label remote sensing scene classification task is mainly aimed at, aiming to realize the automatic category labeling of remote sensing images, significantly reduce the cost and workload of manual labeling, and improve the classification efficiency and accuracy; by effectively analyzing and identifying the complex scene information in the remote sensing image, it can provide efficient and reliable technical support for the fields of land resource utilization, urban planning and environmental protection. In addition, the local semantic enhancement module has significant flexibility and plug-and-play advantages, which can significantly improve the model's ability to express features, show excellent performance in the multi-label remote sensing scene classification task, and can be seamlessly adapted to other similar computer vision tasks; for example, in the classification of hyperspectral images, the local semantic enhancement module can better capture and process the complex information in the spectral data; in the change detection task, it can more accurately identify the area where changes occur in the image; in the semantic segmentation task, the local semantic enhancement module can improve the fine recognition ability of the target boundary and details.
[0142] The above embodiment is only one of the implementation methods that can realize the technical solution of the present invention. The scope of protection claimed by the present invention is not limited only to this embodiment, but also includes changes, replacements and other implementation methods that can be easily thought of by any technician familiar with the technical field within the technical scope disclosed by the present invention.
Claims
1. A multi-label remote sensing scene image classification method, characterized in that: include: Acquire remote sensing images to be classified; Inputting the remote sensing image to be classified into a pre-built semantic-assisted feature fusion network for classification, and obtaining a multi-label remote sensing scene image classification result of the remote sensing image to be classified; The pre-built semantic-assisted feature fusion network includes a dual-scale feature extraction module, a local semantic enhancement module, a cross-scale interactive attention module and a classification module; The dual-scale feature extraction module is used to extract features from the remote sensing image to be classified, so as to capture low-level features and high-level features in the remote sensing image to be classified and obtain a feature map. and feature map The local semantic enhancement module is used to enhance the feature map and feature map Separately extract semantic information to obtain feature maps and feature map Wherein, the local semantic enhancement module introduces supervised semantic loss and semantic-based contrastive learning loss; The cross-scale interactive attention module is used to and feature map Perform information exchange to obtain classification features and categorical features The classification module is used to classify and categorical features Classification results of different scales are generated, and the classification results of different sizes are integrated based on the classification loss to obtain the final score as the multi-label remote sensing scene image classification result of the remote sensing image to be classified.
2. A multi-label remote sensing scene image classification method according to claim 1, characterized in that: The backbone of the dual-scale feature extraction module is ResNet18 with four residual layers; This is the feature map output by the last layer of ResNet18, which contains four residual layers; feature map It is the output feature map of the penultimate layer of ResNet18, which contains four residual layers.
3. A multi-label remote sensing scene image classification method according to claim 1, characterized in that: The local semantic enhancement module includes a first lightweight semantic enhancement unit Transformer1 and a second lightweight semantic enhancement unit Transformer2 in parallel; Among them, the first lightweight semantic enhancement unit Transformer1 is used to transform the feature map Extract semantic information and obtain feature maps The second lightweight semantic enhancement unit Transformer2 is used to Extract semantic information and obtain feature maps 4. A multi-label remote sensing scene image classification method according to claim 3, characterized in that: The first lightweight semantic enhancement unit Transformer1 and the second lightweight semantic enhancement unit Transformer2 have the same structure, both of which include a supervised semantic discriminator, a contrastive learning module, an attention-based feature update mechanism module, a regularization module, and a multi-layer convolution block; A supervised semantic discriminator is used to extract or feature map Extract category-aware semantic information separately to obtain class-specific semantic information vectors Contrastive learning module, used for semantic-based contrastive learning loss, which is derived from feature maps The class-specific semantic information vector extracted from the feature map The class-specific semantic information vector extracted from the image is extended to the semantic level through contrastive learning to obtain highly discriminative semantic information. Attention-based feature update mechanism module for building feature maps or feature map The relationship between the feature pixels and the semantic information with high discriminability is used to obtain the enhanced feature map X e ; Among them, feature pixels are used as a medium to establish the connection between semantics; The regularization module is used to enhance the feature map X e Perform regularization processing to obtain a regularized feature map; Multi-layer convolution blocks are used to mine local information on regularized feature maps to obtain feature maps or feature map 5. A multi-label remote sensing scene image classification method according to claim 4, characterized in that: The supervised semantic discriminator consists of two parallel discriminative branches; In the first discriminant branch, in the feature map or feature map Apply the sigmoid function and 1×1 convolution layer to generate the class attention map For feature maps or feature map and class attention map Each channel of is element-wise multiplied, followed by global average pooling to obtain the class-specific semantic information vector In the second discriminant branch, the feature map or feature map After global average pooling and 1×1 convolution layer, the class activation vector is obtained Among them, the class activation vector Used to generate class attention maps based on supervised semantic loss process.
6. A multi-label remote sensing scene image classification method according to claim 5, characterized in that: Supervised semantic loss, as follows: in, is the supervised semantic loss; BCE(·) is the binary cross entropy loss function; The feature map The i-th element in the class activation vector; y i is the jth label in the label vector; The feature map For the i-th element in the class activation vector; N is the total number of elements in the class activation vector; r is the class activation vector; y is the label vector; σ(·) is the activation function; r j is the jth element in the class activation vector; c is the number of categories.
7. A multi-label remote sensing scene image classification method according to claim 4, characterized in that: Semantic-based contrastive learning loss, as follows: {P(a)=SD j \{to}} SD={SD1,...,SD c } in, is the semantic-based contrastive learning loss; a is the anchor sample with semantic label j in the single-label semantic dataset SD constructed based on the labels of remote sensing scenes; |P(a)| is the total number of positive sample sets of a; P(a) is the positive sample set of a; is the semantic-based contrastive learning loss corresponding to the anchor sample with semantic label j in SD; d(*) is the cosine similarity; p∈P(a) is the relevant positive sample; β is the temperature coefficient; n∈N(a) is the relevant negative sample; SD j is the sub-dataset of category j; is the indicator function; is the j-th sample in the i-th element of the dual-scale semantic vector; DSF i is the i-th element in the dual-scale semantic vector, which is obtained by converting The class-specific semantic information vector extracted from the feature map The class-specific semantic information vector extracted from is concatenated.
8. The multi-label remote sensing scene image classification method according to claim 1, characterized in that: The classification loss is specifically: in, is the classification loss: FS i For the final score.
9. A multi-label remote sensing scene image classification system, characterized in that: include: An image acquisition module, used for acquiring remote sensing images to be classified; A classification module is used to input the remote sensing image to be classified into a pre-built semantic-assisted feature fusion network for classification, so as to obtain a multi-label remote sensing scene image classification result of the remote sensing image to be classified; The pre-built semantic-assisted feature fusion network includes a dual-scale feature extraction module, a local semantic enhancement module, a cross-scale interactive attention module and a classification module; The dual-scale feature extraction module is used to extract features from the remote sensing image to be classified, so as to capture low-level features and high-level features in the remote sensing image to be classified and obtain a feature map. and feature map The local semantic enhancement module is used to enhance the feature map and feature map Semantic information is extracted respectively to obtain feature maps and feature map Wherein, the local semantic enhancement module introduces supervised semantic loss and semantic-based contrastive learning loss; The cross-scale interactive attention module is used to and feature map Perform information exchange to obtain classification features and categorical features The classification module is used to classify and categorical features Classification results of different scales are generated, and the classification results of different sizes are integrated based on the classification loss to obtain the final score as the multi-label remote sensing scene image classification result of the remote sensing image to be classified.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the multi-label remote sensing scene image classification method as described in any one of claims 1 to 7 are implemented.