Semi-supervised multispectral remote sensing image scene classification method, device, equipment and medium
By adopting a dual-branch network structure and spectral attention module in the multispectral remote sensing image scene classification, the problem of insufficient accuracy of pseudo-labels is solved, and higher model classification performance and information mining effects are achieved.
Patent Information
- Application Number
- CN202510482823.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing semi-supervised learning technology has the problem of insufficient pseudo-label accuracy in multispectral remote sensing image scene classification, resulting in limited model performance.
The dual-branch network structure is adopted, combining the spatial feature branch network and the spectral feature branch network, and the pseudo-label fusion of weak enhancement and strong enhancement processing can improve the accuracy of the pseudo-label, and enhance the spectral information extraction ability through the spectral attention module.
The pseudo-label accuracy of multi-spectral remote sensing images is improved, the classification performance of the model is improved, and the spatial and spectral information of the image is fully explored.
Smart Images

Figure CN120014370A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image scene classification, and in particular to a semi-supervised multispectral remote sensing image scene classification method, device, equipment and medium. Background Art
[0002] Existing semi-supervised learning techniques aim to improve the training efficiency and performance of models by combining a small amount of labeled data with a large amount of unlabeled data. In image classification tasks, semi-supervised methods are mainly divided into two categories: one is confidence-based pseudo-label generation, which guides model training by screening high-confidence pseudo-labels; the other is based on consistency regularization, which constrains the model to keep consistent predictions for the same data sample by applying different perturbations (such as augmentation operations) to the input data.
[0003] In recent years, semi-supervised learning has also been widely used in the field of remote sensing, especially in the scene classification task of multispectral remote sensing images. This type of technology is used to alleviate the problem of scarce labeled data. On the one hand, the data of multispectral remote sensing images is complex and non-uniform, and the consistency regularization method is not suitable for the scene classification task of multispectral remote sensing images. On the other hand, most of the existing scene classification models are mainly designed for RGB images, lacking the spectral information of multispectral remote sensing images, and the generated pseudo-labels are not accurate enough, which limits the performance of the model to a certain extent. Summary of the invention
[0004] The present invention provides a semi-supervised multispectral remote sensing image scene classification method, device, equipment and medium, which can fully mine the spectral information of multispectral remote sensing images, improve the accuracy of pseudo labels, and enhance the classification performance of the model.
[0005] The present invention provides a semi-supervised multispectral remote sensing image scene classification method, comprising: Acquire an unlabeled multispectral remote sensing image, and perform weak enhancement processing and strong enhancement processing on the unlabeled multispectral remote sensing image to obtain a weakly enhanced remote sensing image and a strongly enhanced remote sensing image; The weakly enhanced remote sensing image and the strongly enhanced remote sensing image are respectively input into a dual-branch network structure pre-constructed by a spatial feature branch network and a spectral feature branch network, wherein the spatial feature branch network is used to extract the spatial features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image, respectively, and obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spatial features according to the spatial features; the spectral feature branch network is used to extract the spectral features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image according to the spectral attention, and obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spectral features according to the spectral features; The weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features are fused to obtain a pseudo label of the unlabeled multispectral remote sensing image, and an unsupervised loss function is obtained according to the pseudo label and the strong enhancement prediction results corresponding to the spatial features and the spectral features, respectively. The unsupervised loss function is used to determine the total loss function; According to the total loss function, determine whether the dual-branch network structure has completed training, if so, use the dual-branch network structure as a scene classification model, if not, adjust the parameters of the dual-branch network structure, and return to the step of obtaining unlabeled multispectral remote sensing images until the training is completed; Based on the scene classification model, scene classification is performed on the target multispectral remote sensing image to obtain a target scene classification result.
[0006] As an embodiment, the spatial feature branch network is obtained based on a deep residual network, and the spectral feature branch network is obtained by adding a spectral attention module to each network layer of the deep residual network; The spectral attention module is used to determine the query parameters, key parameters and value parameters of the input image feature map based on the linear layer, perform dimensionality transformation on the query parameters, the key parameters and the value parameters, transmit the query parameters and the key parameters after dimensionality transformation to the normalization layer, determine the attention score between different spectral bands of the image feature map based on the normalization layer, perform calculations on the attention score and the value parameters after dimensionality transformation and transmit them to the output linear layer, obtain the output image feature map based on the output linear layer, perform dimensionality transformation on the output image feature map and output it to the next network layer.
[0007] As an embodiment, the calculation formula of the query parameter, the key parameter and the value parameter is as follows:
[0008] in, represents the query parameter, represents the key parameter, represents the value parameter, , , Represent the parameter matrices of the three linear layers, represents the input image feature map, , R represents the real number field, , , Respectively represent the number of channels, width and height of the input image feature map; Correspondingly, dimension transformation is performed on the query parameter, the key parameter, and the value parameter, including: The query parameter, the key parameter and the value parameter are respectively Dimension conversion to Dimension.
[0009] As an embodiment, the weak enhancement prediction result corresponding to the spatial feature and the weak enhancement prediction result corresponding to the spectral feature are fused to obtain a pseudo label of the unlabeled multispectral remote sensing image, including: Based on the entropy weighted algorithm, the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features are fused to obtain a pseudo label of the unlabeled multispectral remote sensing image.
[0010] As an embodiment, the entropy weighted algorithm is used to fuse the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features to obtain the pseudo label of the unlabeled multispectral remote sensing image, including: Determine the entropy of the weak enhancement prediction result corresponding to the spatial feature and the weak enhancement prediction result corresponding to the spectral feature, and determine the confidence of the spatial feature branch network and the spectral feature branch network according to the entropy; Determining the weights of the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features according to the respective confidences of the spatial feature branch network and the spectral feature branch network; According to the weights, the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features are weightedly fused to obtain a pseudo label of the unlabeled multispectral remote sensing image.
[0011] As an embodiment, obtaining an unsupervised loss function according to the pseudo-label and the strong enhancement prediction results corresponding to the spatial feature and the spectral feature respectively includes: The pseudo-label and the strong enhancement prediction results corresponding to the spatial feature and the spectral feature are substituted into the cross entropy loss function to obtain an unsupervised loss function.
[0012] As an embodiment, it also includes: Acquire annotated multispectral remote sensing images; Inputting the annotated multispectral remote sensing image into the dual-branch network structure to obtain a prediction result corresponding to the annotated multispectral remote sensing image; Obtaining a supervised loss function according to the labels corresponding to the annotated multispectral remote sensing image and the prediction results; The sum of the unsupervised loss function and the supervised loss function is used as the total loss function.
[0013] The present invention also provides a semi-supervised multispectral remote sensing image scene classification device, comprising: An acquisition module is used to acquire an unlabeled multispectral remote sensing image, and respectively perform weak enhancement processing and strong enhancement processing on the unlabeled multispectral remote sensing image to obtain a weakly enhanced remote sensing image and a strongly enhanced remote sensing image; A prediction module, used to input the weakly enhanced remote sensing image and the strongly enhanced remote sensing image into a dual-branch network structure pre-constructed by a spatial feature branch network and a spectral feature branch network, respectively, wherein the spatial feature branch network is used to extract the spatial features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image, respectively, and obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spatial features according to the spatial features; the spectral feature branch network is used to extract the spectral features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image, respectively, based on spectral attention, and obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spectral features according to the spectral features; A determination module is used to fuse the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features to obtain a pseudo label of the unlabeled multispectral remote sensing image, and obtain an unsupervised loss function according to the pseudo label and the strong enhancement prediction results corresponding to the spatial features and the spectral features, wherein the unsupervised loss function is used to determine the total loss function; A training module is used to determine whether the dual-branch network structure has completed training according to the total loss function. If so, the dual-branch network structure is used as a scene classification model. If not, the parameters of the dual-branch network structure are adjusted and the step of obtaining unlabeled multispectral remote sensing images is returned until the training is completed. The classification module is used to perform scene classification on the target multispectral remote sensing image based on the scene classification model to obtain a target scene classification result.
[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the semi-supervised multispectral remote sensing image scene classification method as described above is implemented.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the semi-supervised multispectral remote sensing image scene classification methods described above.
[0016] The semi-supervised multispectral remote sensing image scene classification method, device, equipment and medium provided by the present invention have a dual-branch network structure that fully utilizes the spatial and spectral information of the multispectral remote sensing image and the complementary information between different bands. The spectral feature branch network introduces spectral attention to further enhance the spectral information extraction capability and obtain pseudo-labels with higher accuracy, thereby improving the classification performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 This is one of the flow charts of the semi-supervised multispectral remote sensing image scene classification method provided by the present invention.
[0019] Figure 2 It is a structural schematic diagram of the spectral attention module provided by the present invention.
[0020] Figure 3 This is the second flow chart of the semi-supervised multispectral remote sensing image scene classification method provided by the present invention.
[0021] Figure 4 It is a structural schematic diagram of the semi-supervised multispectral remote sensing image scene classification device provided by the present invention.
[0022] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0024] Figure 1 This is one of the flow charts of the semi-supervised multispectral remote sensing image scene classification method provided by the present invention, such as Figure 1 As shown, the present invention provides a semi-supervised multispectral remote sensing image scene classification method, including steps S100-S500, wherein steps S100-S400 are a model training stage, and step S500 is a model application stage.
[0025] Step S100: acquiring an unlabeled multispectral remote sensing image, and performing weak enhancement processing and strong enhancement processing on the unlabeled multispectral remote sensing image to obtain a weakly enhanced remote sensing image and a strongly enhanced remote sensing image.
[0026] Step S200, respectively input the weakly enhanced remote sensing image and the strongly enhanced remote sensing image into a dual-branch network structure pre-constructed by a spatial feature branch network and a spectral feature branch network, wherein the spatial feature branch network is used to respectively extract the spatial features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image, and obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spatial features based on the spatial features; the spectral feature branch network is used to respectively extract the spectral features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image based on spectral attention, and obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spectral features based on the spectral features.
[0027] Step S300, the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features are merged to obtain a pseudo-label of the unlabeled multispectral remote sensing image, and an unsupervised loss function is obtained according to the pseudo-label and the strong enhancement prediction results corresponding to the spatial features and the spectral features respectively. The unsupervised loss function is used to determine the total loss function.
[0028] Step S400, based on the total loss function, determine whether the dual-branch network structure has completed training. If so, use the dual-branch network structure as a scene classification model. If not, adjust the parameters of the dual-branch network structure and return to the step of obtaining unlabeled multispectral remote sensing images until the training is completed.
[0029] Step S500: performing scene classification on the target multispectral remote sensing image based on the scene classification model to obtain a target scene classification result.
[0030] In the model training stage, labeled multispectral remote sensing images and unlabeled multispectral remote sensing images can be obtained as data sets. The data sets are divided into training set, test set and validation set according to a certain ratio. The training set, test set and validation set are used to train, test and validate the pre-built dual-branch network structure respectively to obtain a scene classification model.
[0031] Labeled multispectral remote sensing images can be directly input into the dual-branch network structure for feature extraction and result prediction. Unlabeled multispectral remote sensing images need to be processed by weak enhancement and strong enhancement respectively. Weak enhancement usually refers to a small adjustment to the image to improve the visual quality of the image, but it will not significantly change the content or structure of the image. Common weak enhancement methods include: brightness and contrast adjustment, filtering and color correction. Strong enhancement is a significant conversion and modification of the image, which is often used to highlight certain features or information, and may change the overall structure or content of the image. Common methods include: geometric transformation, feature extraction, image synthesis, etc.
[0032] The present invention utilizes the two branches of the dual-branch network structure to perform feature extraction and prediction on weakly enhanced remote sensing images and strongly enhanced remote sensing images respectively, and then uses the weakly enhanced prediction results of the two branches to obtain pseudo labels, thereby realizing the complementarity of spatial features and spectral features. At the same time, the spectral feature branch network performs feature extraction based on spectral attention, enhances the extraction of spectral information, fully mines the deep features of multispectral remote sensing images, and further improves the accuracy of pseudo labels of unlabeled multispectral remote sensing images.
[0033] Correspondingly, the accuracy of the unsupervised loss function obtained by the more accurate high-quality pseudo-label is also higher. On the basis that the data set is composed of labeled multispectral remote sensing images and unlabeled multispectral remote sensing images, the loss function of the model includes a supervised loss function and an unsupervised loss function, which can more accurately evaluate the gap between the model prediction value and the true value, and determine whether the prediction error of the dual-branch network structure reaches a preset value. If so, it can be determined that the dual-branch network structure has completed training, otherwise, it is determined whether the preset number of training iterations has been reached. If so, it is determined that the dual-branch network structure has completed training, otherwise, the parameters for determining the dual-branch network structure are adjusted, and new samples are randomly selected from the training set. Steps S100-S400 are repeated to implement iterative training of the dual-branch network structure.
[0034] After completing the model training, in response to the user's request, the target multispectral remote sensing image to be classified input by the user is obtained, and the target multispectral remote sensing image is input into the scene classification model. The two branches of the scene classification model respectively extract the spatial features and spectral features of the target multispectral remote sensing image and predict them to obtain the target scene classification result.
[0035] It can be understood that the dual-branch network structure of the present invention makes full use of the complementary information between different bands and fully mines the spectral information in multispectral images. The spectral feature branch network introduces spectral attention to mine the spectral information of multispectral remote sensing images and obtain pseudo-labels with higher accuracy, thereby improving the classification performance of the model.
[0036] Based on the above embodiment, as an optional embodiment, the spatial feature branch network is obtained based on a deep residual network, and the spectral feature branch network is obtained by adding a spectral attention module to each network layer of the deep residual network.
[0037] The spectral attention module is used to determine the query parameters, key parameters and value parameters of the input image feature map based on the linear layer, perform dimensionality transformation on the query parameters, the key parameters and the value parameters, transmit the query parameters and the key parameters after dimensionality transformation to the normalization layer, determine the attention score between different spectral bands of the image feature map based on the normalization layer, perform calculations on the attention score and the value parameters after dimensionality transformation and transmit them to the output linear layer, obtain the output image feature map based on the output linear layer, perform dimensionality transformation on the output image feature map and output it to the next network layer.
[0038] In order to ensure the complementarity of the spatial feature branch network and the spectral feature branch network, the spatial feature branch network and the spectral feature branch network are obtained based on the same feature extraction network. The difference is that in order to enhance the feature extraction capability of the spectral feature branch network so as to mine deep spectral information, the spectral feature branch network introduces a spectral attention module in the feature extraction network. Specifically, in an embodiment of the present invention, the original deep residual network ResNet50 is used as the spatial feature branch network, a spectral attention module is introduced in each network layer of the deep residual network ResNet50, and the modified ResNet50 is used as the spectral feature branch network, which further enhances its expertise in capturing spectral features.
[0039] like Figure 2 As shown, as an optional embodiment, the calculation formula of the query parameter, the key parameter and the value parameter is as follows:
[0040] in, represents the query parameter, represents the key parameter, represents the value parameter, , , Represent the parameter matrices of the three linear layers, represents the input image feature map, , R represents the real number field, , , Respectively represent the number of channels, width and height of the input image feature map. That is, for a given image feature map , R represents the field of real numbers, First, three linear layers are passed , and get, , , Represents the parameter matrix corresponding to Q, K, and V.
[0041] Correspondingly, the query parameter, the key parameter and the value parameter are dimensionally transformed, including: respectively transforming the query parameter, the key parameter and the value parameter from Dimension conversion to Dimension. and By transforming the dimension of , the relationship between different bands can be modeled.
[0042] The calculation formula of the attention score is as follows:
[0043] in, Represents the attention scores between different bands, Indicates the query parameter after dimension transformation. Represents the key parameter after dimension transformation, and Softmax represents the Softmax function.
[0044] The calculation formula for the output image feature map is as follows:
[0045] in, represents the output image feature map, represents the parameter matrix of the output linear layer, Represents the value parameter after dimension transformation. The output linear layer is used to linearly change the attention score.
[0046] In order to make the output image feature map It can be input to the next network layer and its dimension needs to be adjusted back to .
[0047] It can be understood that in order to effectively process multispectral images and make full use of unlabeled data, the present invention adopts a dual-branch network structure, which is used to extract spatial and spectral features respectively. The spectral attention module is used in the channel dimension, which can fully mine the spectral information in the unlabeled data and realize the full utilization of the unlabeled data.
[0048] like Figure 3 As shown, based on the above embodiment, as an optional embodiment, the pseudo label of the unlabeled multispectral remote sensing image is obtained according to the weak enhancement prediction result corresponding to the spatial feature and the weak enhancement prediction result corresponding to the spectral feature, including: Based on the entropy weighted algorithm, the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features are fused to obtain a pseudo label of the unlabeled multispectral remote sensing image.
[0049] In order to improve the quality of pseudo labels, the present invention adopts an entropy weighted fusion method to fuse the prediction results of the two branches on the weakly enhanced image, and the score corresponding to the largest category in the fusion result is greater than the threshold. The prediction results are retained as the final pseudo labels.
[0050] As an optional embodiment, the entropy-weighted algorithm is used to fuse the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features to obtain a pseudo label for the unlabeled multispectral remote sensing image, including steps S310 to S330.
[0051] Step S310, determining the entropy of the weak enhancement prediction result corresponding to the spatial feature and the weak enhancement prediction result corresponding to the spectral feature, and determining the confidence of the spatial feature branch network and the spectral feature branch network according to the entropy.
[0052] The confidence of each of the two branches, the spatial feature branch network and the spectral feature branch network, is determined by the entropy of their predictions. The common calculation formula for entropy is as follows:
[0053] in, represents the entropy of a branch for the prediction result z, Represents the prediction result z belongs to the category i The probability of and Is the network for category i, j Output logits.
[0054] Confidence is the inverse of entropy, and the public calculation formula for confidence is as follows:
[0055] in, Indicates the confidence of a branch for the prediction result z.
[0056] Substitute the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features into the above formula to obtain the confidence of the spatial feature branch network: and the confidence of the spectral feature branch network .
[0057] Step S320: determining the weights of the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features according to the confidences of the spatial feature branch network and the spectral feature branch network.
[0058] The calculation formula of the weight of the weak enhancement prediction result corresponding to the spatial feature is as follows:
[0059] in, Represents the weight of the weakly enhanced prediction result corresponding to the spatial feature.
[0060] The calculation formula of the weight of the weak enhancement prediction result corresponding to the spectral feature is as follows:
[0061] in, Represents the weight of the weak enhancement prediction result corresponding to the spectral feature.
[0062] Step S330: performing weighted fusion on the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features according to the weights to obtain a pseudo label of the unlabeled multispectral remote sensing image.
[0063] The calculation formula of the pseudo label is as follows:
[0064] in, represents the pseudo-label, which is a weighted combination of the prediction results of the two branches. and These are the weakly enhanced prediction results obtained by the two branches respectively.
[0065] It can be understood that the present invention obtains high-quality pseudo-labels by fusing the prediction results of the spatial feature branch network and the spectral feature branch network in an entropy-weighted manner. The obtained pseudo-labels simultaneously fuse the spatial information and the spectral information, and the quality of the pseudo-labels is higher. At the same time, the spectral information in the unlabeled data is fully mined, thereby achieving full utilization of the unlabeled data.
[0066] Based on the above embodiment, as an optional embodiment, the unsupervised loss function is obtained according to the pseudo-label and the strong enhancement prediction results corresponding to the spatial feature and the spectral feature, including: The pseudo-label and the strong enhancement prediction results corresponding to the spatial feature and the spectral feature are substituted into the cross entropy loss function to obtain an unsupervised loss function.
[0067] The unsupervised loss function is obtained based on the prediction results of the pseudo-label and the dual-branch network structure for the strongly enhanced remote sensing image. The calculation formula of the unsupervised loss function is as follows:
[0068]
[0069]
[0070] in, represents the unsupervised loss function corresponding to the spatial feature branch network, represents the unsupervised loss function corresponding to the spectral feature branch network, represents the batch size, represents the prediction result of the spatial feature branch network for the strongly enhanced remote sensing image, represents the prediction result of the spectral feature branch network for the strongly enhanced remote sensing image, represents the unsupervised loss function, represents cross-entropy, Represents unlabeled multispectral remote sensing images Perform strong enhancement processing.
[0071] It can be understood that the present invention uses pseudo labels to implement training supervision of strongly enhanced remote sensing images, achieves complementarity between spatial features and spectral features, and is conducive to improving classification performance.
[0072] Based on the above embodiment, as an optional embodiment, the semi-supervised multispectral remote sensing image scene classification method provided by the present invention further includes the following steps.
[0073] Acquire annotated multispectral remote sensing image. This step can be performed simultaneously with step S100.
[0074] The annotated multispectral remote sensing image is input into the dual-branch network structure to obtain a prediction result corresponding to the annotated multispectral remote sensing image. This step can be performed simultaneously with step S200.
[0075] According to the labels corresponding to the annotated multispectral remote sensing images and the prediction results, a supervised loss function is obtained. This step can be performed simultaneously with step S300.
[0076] The sum of the unsupervised loss function and the supervised loss function is used as the total loss function.
[0077] The supervised loss function is calculated based on the prediction of the two branches for the weakly enhanced image, which can be expressed as the following formula:
[0078]
[0079]
[0080] in, Represents input The corresponding true label, represents the supervised loss function corresponding to the spatial feature branch network, represents the supervised loss function corresponding to the spectral feature branch network, represents the batch size, represents the prediction result of the spatial feature branch network for the weakly enhanced remote sensing image, represents the prediction result of the spectral feature branch network for the weakly enhanced remote sensing image, represents the supervised loss function, Represents unlabeled multispectral remote sensing images Perform weak enhancement processing.
[0081] The total loss function is calculated as follows:
[0082] in, It is a hyperparameter used to balance the proportion of supervised and unsupervised losses.
[0083] Next, the hardware and software environment of the experiment, the dataset used for the experiment, the experimental settings, and the experimental evaluation metrics are introduced in detail, and the experimental results are compared with those of previous methods.
[0084] (1) Experimental environment
[0085] The detailed information of the environment configuration is shown in Table 1.
[0086] Table 1 Experimental environment configuration
[0087] (2) Experimental dataset
[0088] The method of the present invention is verified on multispectral remote sensing image datasets EuroSAT and SEN12MS.
[0089] (3) Experimental setup.
[0090] The present invention conducts six experimental settings on each data set, namely, using only 5 labeled data, 10 labeled data, 50 labeled data, 100 labeled data, 200 labeled data, and 300 labeled data for each category. During the training process, the learning rate, optimizer, , and batch size are set to 1×10 -4 , Adam, 0.9, 0.999, and 4. Pseudo label threshold is set to 0.8. The total number of training steps is .
[0091] (4) Experimental results
[0092] Table 2 Experimental results of EuroSAT dataset
[0093] PseudoLabel, MixMatch, FixMatch, FlexMatch, FreeMatch, and MSMatch are all existing scene classification methods. As can be seen from Table 2, the classification accuracy of the present invention is the highest under multiple label numbers. Only MSMatch is slightly higher than the present invention when the number of labels is 2000. However, MSMatch migrates the FixMatch method to multispectral data, but its backbone network only uses the Efficient network, and is not specially designed and optimized for the multi-channel characteristics of multispectral images, and does not fully mine the spectral information.
[0094] Table 3 SEN12MS experimental results
[0095] It can be seen from Table 3 that no matter what the number of labels is, the accuracy of the present invention is the highest.
[0096] In summary, the present invention designs a dual-branch network structure according to the characteristics of multispectral images, which is used to extract spatial information and spectral information from multispectral images respectively, wherein the spatial feature extraction branch uses the ResNet50 network, and the spectral feature extraction branch enhances the extraction of spectral information by introducing a spectral attention module in ResNet50. In addition, before the image is input, two strong and weak enhancements are performed, and the weak enhancement prediction results are fused to obtain high-quality pseudo labels, and the network is supervised to predict the strong enhanced image, thereby fully mining the information of the unlabeled image. The existing methods usually use only one network branch to obtain pseudo labels, while the present invention obtains high-quality pseudo labels by fusing the prediction results of the spatial and spectral branches based on entropy weighting. The obtained pseudo-labels simultaneously integrate spatial information and spectral information, and the quality of the pseudo-labels is higher than that of other methods. At the same time, the spectral information in the unlabeled data is fully mined, and the unlabeled data is fully utilized. Therefore, the present invention provides a new semi-supervised multispectral remote sensing image scene classification method. In addition, the attention mechanism adopted by the present invention is also different from most methods. The present invention uses it in the channel dimension, which is different from the existing attention mechanism implementation method.
[0097] The semi-supervised multispectral remote sensing image scene classification device provided by the present invention is described below. The semi-supervised multispectral remote sensing image scene classification device described below and the semi-supervised multispectral remote sensing image scene classification method described above can be referenced to each other.
[0098] Figure 4 : is a schematic diagram of the structure of the semi-supervised multi-spectral remote sensing image scene classification device provided by the present invention, such as Figure 4 As shown, the present invention also provides a semi-supervised multispectral remote sensing image scene classification device, including the following modules.
[0099] The acquisition module 410 is used to acquire an unlabeled multispectral remote sensing image, and respectively perform weak enhancement processing and strong enhancement processing on the unlabeled multispectral remote sensing image to obtain a weakly enhanced remote sensing image and a strongly enhanced remote sensing image; The prediction module 420 is used to input the weakly enhanced remote sensing image and the strongly enhanced remote sensing image into a dual-branch network structure pre-constructed by a spatial feature branch network and a spectral feature branch network, respectively. The spatial feature branch network is used to extract the spatial features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image, respectively, and obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spatial features according to the spatial features; the spectral feature branch network is used to extract the spectral features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image, respectively, based on spectral attention, and obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spectral features according to the spectral features; A determination module 430 is used to fuse the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features to obtain a pseudo label of the unlabeled multispectral remote sensing image, and obtain an unsupervised loss function according to the pseudo label and the strong enhancement prediction results corresponding to the spatial features and the spectral features, wherein the unsupervised loss function is used to determine the total loss function; The training module 440 is used to determine whether the dual-branch network structure has completed training according to the total loss function. If so, the dual-branch network structure is used as a scene classification model. If not, the parameters of the dual-branch network structure are adjusted, and the step of obtaining unlabeled multispectral remote sensing images is returned to be executed until the training is completed. The classification module 450 is used to perform scene classification on the target multispectral remote sensing image based on the scene classification model to obtain a target scene classification result.
[0100] As an embodiment, the spatial feature branch network is obtained based on a deep residual network, and the spectral feature branch network is obtained by adding a spectral attention module to each network layer of the deep residual network; The spectral attention module is used to determine the query parameters, key parameters and value parameters of the input image feature map based on the linear layer, perform dimensionality transformation on the query parameters, the key parameters and the value parameters, transmit the query parameters and the key parameters after dimensionality transformation to the normalization layer, determine the attention score between different spectral bands of the image feature map based on the normalization layer, perform calculations on the attention score and the value parameters after dimensionality transformation and transmit them to the output linear layer, obtain the output image feature map based on the output linear layer, perform dimensionality transformation on the output image feature map and output it to the next network layer.
[0101] As an embodiment, the calculation formula of the query parameter, the key parameter and the value parameter is as follows:
[0102] in, represents the query parameter, represents the key parameter, represents the value parameter, , , Represent the parameter matrices of the three linear layers, represents the input image feature map, , R represents the real number field, , , Respectively represent the number of channels, width and height of the input image feature map; Correspondingly, dimension transformation is performed on the query parameter, the key parameter, and the value parameter, including: The query parameter, the key parameter and the value parameter are respectively Dimension conversion to Dimension.
[0103] As an embodiment, the determining module 430 is further configured to: Based on the entropy weighted algorithm, the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features are fused to obtain a pseudo label of the unlabeled multispectral remote sensing image.
[0104] As an embodiment, the determining module 430 is further configured to: Determine the entropy of the weak enhancement prediction result corresponding to the spatial feature and the weak enhancement prediction result corresponding to the spectral feature, and determine the confidence of the spatial feature branch network and the spectral feature branch network according to the entropy; Determining the weights of the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features according to the respective confidences of the spatial feature branch network and the spectral feature branch network; According to the weights, the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features are weightedly fused to obtain a pseudo label of the unlabeled multispectral remote sensing image.
[0105] As an embodiment, the determining module 430 is further configured to: The pseudo-label and the strong enhancement prediction results corresponding to the spatial feature and the spectral feature are substituted into the cross entropy loss function to obtain an unsupervised loss function.
[0106] As an embodiment, it also includes: Acquire annotated multispectral remote sensing images; Inputting the annotated multispectral remote sensing image into the dual-branch network structure to obtain a prediction result corresponding to the annotated multispectral remote sensing image; Obtaining a supervised loss function according to the labels corresponding to the annotated multispectral remote sensing image and the prediction results; The sum of the unsupervised loss function and the supervised loss function is used as the total loss function. The semi-supervised multi-spectral remote sensing image scene classification device provided by the present invention is used to execute the semi-supervised multi-spectral remote sensing image scene classification method described in any of the above embodiments, and has the technical effect corresponding to the semi-supervised multi-spectral remote sensing image scene classification method, which will not be repeated.
[0107] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5As shown, the electronic device may include: a processor (processor) 510, a communication interface (Communications Interface) 520, a memory (memory) 530 and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call the logic instructions in the memory 530 to execute the semi-supervised multispectral remote sensing image scene classification method, which includes: obtaining an unlabeled multispectral remote sensing image, performing weak enhancement processing and strong enhancement processing on the unlabeled multispectral remote sensing image, respectively, to obtain a weakly enhanced remote sensing image and a strongly enhanced remote sensing image; respectively inputting the weakly enhanced remote sensing image and the strongly enhanced remote sensing image into a double-branch network structure pre-constructed by a spatial feature branch network and a spectral feature branch network, wherein the spatial feature branch network is used to respectively extract the spatial features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image, and obtain a strong enhancement prediction result and a weak enhancement prediction result corresponding to the spatial feature based on the spatial feature; the spectral feature branch network is used to respectively extract the spectral features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image based on spectral attention. And according to the spectral features, a strong enhancement prediction result and a weak enhancement prediction result corresponding to the spectral features are obtained; the weak enhancement prediction result corresponding to the spatial features and the weak enhancement prediction result corresponding to the spectral features are fused to obtain a pseudo-label of the unlabeled multispectral remote sensing image, and an unsupervised loss function is obtained according to the pseudo-label and the strong enhancement prediction results corresponding to the spatial features and the spectral features, and the unsupervised loss function is used to determine the total loss function; according to the total loss function, it is judged whether the dual-branch network structure has completed training, if so, the dual-branch network structure is used as a scene classification model, if not, the parameters of the dual-branch network structure are adjusted, and the step of obtaining the unlabeled multispectral remote sensing image is returned to execute until the training is completed; based on the scene classification model, the target multispectral remote sensing image is scene classified to obtain a target scene classification result.
[0108] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0109] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the semi-supervised multispectral remote sensing image scene classification method provided by the above methods, and the method includes: obtaining an unlabeled multispectral remote sensing image, performing weak enhancement processing and strong enhancement processing on the unlabeled multispectral remote sensing image, respectively, to obtain a weakly enhanced remote sensing image and a strongly enhanced remote sensing image; respectively inputting the weakly enhanced remote sensing image and the strongly enhanced remote sensing image into a dual-branch network structure pre-constructed by a spatial feature branch network and a spectral feature branch network, the spatial feature branch network is used to respectively extract the spatial features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image, and according to the spatial features, obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spatial features; the spectral feature branch network is used to extract the spatial features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image based on the spectral feature; the spectral feature branch network is used to extract the spatial features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image based on the spectral feature; the spectral feature branch network is used to extract the spatial features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image based on the spectral feature; the spectral feature branch network is used to extract the spatial features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image based on the spectral feature. The invention aims to extract the spectral features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image respectively, and obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spectral features according to the spectral features; fuse the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features to obtain the pseudo-labels of the unlabeled multispectral remote sensing image, and obtain the unsupervised loss function according to the pseudo-labels and the strong enhancement prediction results corresponding to the spatial features and the spectral features, and the unsupervised loss function is used to determine the total loss function; according to the total loss function, determine whether the dual-branch network structure has completed training, if so, use the dual-branch network structure as a scene classification model, if not, adjust the parameters of the dual-branch network structure, return to the step of obtaining the unlabeled multispectral remote sensing image, until the training is completed; perform scene classification on the target multispectral remote sensing image based on the scene classification model to obtain the target scene classification result.
[0110] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the semi-supervised multispectral remote sensing image scene classification method provided by the above-mentioned methods, the method comprising: obtaining an unlabeled multispectral remote sensing image, performing weak enhancement processing and strong enhancement processing on the unlabeled multispectral remote sensing image, respectively, to obtain a weakly enhanced remote sensing image and a strongly enhanced remote sensing image; respectively inputting the weakly enhanced remote sensing image and the strongly enhanced remote sensing image into a dual-branch network structure pre-constructed by a spatial feature branch network and a spectral feature branch network, the spatial feature branch network being used to respectively extract the spatial features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image, and according to the spatial features, obtaining a strong enhancement prediction result and a weak enhancement prediction result corresponding to the spatial features; the spectral feature branch network being used to respectively extract the weakly enhanced remote sensing image based on spectral attention. and the spectral features of the strong enhanced remote sensing image, and according to the spectral features, obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spectral features; fuse the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features to obtain the pseudo-label of the unlabeled multispectral remote sensing image, and according to the pseudo-label and the strong enhancement prediction results corresponding to the spatial features and the spectral features, obtain an unsupervised loss function, and the unsupervised loss function is used to determine the total loss function; according to the total loss function, determine whether the dual-branch network structure has completed training, if so, use the dual-branch network structure as a scene classification model, if not, adjust the parameters of the dual-branch network structure, return to execute the step of obtaining the unlabeled multispectral remote sensing image until the training is completed; perform scene classification on the target multispectral remote sensing image based on the scene classification model to obtain the target scene classification result.
[0111] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0112] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A semi-supervised multispectral remote sensing image scene classification method, characterized in that: include: Acquire an unlabeled multispectral remote sensing image, and perform weak enhancement processing and strong enhancement processing on the unlabeled multispectral remote sensing image to obtain a weakly enhanced remote sensing image and a strongly enhanced remote sensing image; The weakly enhanced remote sensing image and the strongly enhanced remote sensing image are respectively input into a dual-branch network structure pre-constructed by a spatial feature branch network and a spectral feature branch network, wherein the spatial feature branch network is used to extract the spatial features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image, respectively, and obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spatial features according to the spatial features; the spectral feature branch network is used to extract the spectral features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image according to the spectral attention, and obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spectral features according to the spectral features; The weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features are fused to obtain a pseudo label of the unlabeled multispectral remote sensing image, and an unsupervised loss function is obtained according to the pseudo label and the strong enhancement prediction results corresponding to the spatial features and the spectral features, respectively. The unsupervised loss function is used to determine the total loss function; According to the total loss function, determine whether the dual-branch network structure has completed training, if so, use the dual-branch network structure as a scene classification model, if not, adjust the parameters of the dual-branch network structure, and return to the step of obtaining unlabeled multispectral remote sensing images until the training is completed; The target multispectral remote sensing image is subjected to scene classification based on the scene classification model to obtain a target scene classification result.
2. The semi-supervised multispectral remote sensing image scene classification method according to claim 1, characterized in that: The spatial feature branch network is obtained based on a deep residual network, and the spectral feature branch network is obtained by adding a spectral attention module to each network layer of the deep residual network; The spectral attention module is used to determine the query parameters, key parameters and value parameters of the input image feature map based on the linear layer, perform dimensionality transformation on the query parameters, the key parameters and the value parameters, transmit the query parameters and the key parameters after dimensionality transformation to the normalization layer, determine the attention score between different spectral bands of the image feature map based on the normalization layer, perform calculations on the attention score and the value parameters after dimensionality transformation and transmit them to the output linear layer, obtain the output image feature map based on the output linear layer, perform dimensionality transformation on the output image feature map and output it to the next network layer.
3. The semi-supervised multispectral remote sensing image scene classification method according to claim 2, characterized in that: The calculation formulas for the query parameter, the key parameter, and the value parameter are as follows: ; in, represents the query parameter, represents the key parameter, represents the value parameter, , , Represent the parameter matrices of the three linear layers, represents the input image feature map, , R represents the real number field, , , Respectively represent the number of channels, width and height of the input image feature map; Correspondingly, dimension transformation is performed on the query parameter, the key parameter, and the value parameter, including: The query parameter, the key parameter and the value parameter are respectively Dimension conversion to Dimension.
4. The semi-supervised multispectral remote sensing image scene classification method according to any one of claims 1 to 3, characterized in that: The step of fusing the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features to obtain a pseudo label of the unlabeled multispectral remote sensing image includes: Based on the entropy weighted algorithm, the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features are fused to obtain a pseudo label of the unlabeled multispectral remote sensing image.
5. The semi-supervised multispectral remote sensing image scene classification method according to claim 4, characterized in that: The method of fusing the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features based on the entropy weighted algorithm to obtain the pseudo label of the unlabeled multispectral remote sensing image includes: Determine the entropy of the weak enhancement prediction result corresponding to the spatial feature and the weak enhancement prediction result corresponding to the spectral feature, and determine the confidence of the spatial feature branch network and the spectral feature branch network according to the entropy; Determining the weights of the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features according to the respective confidences of the spatial feature branch network and the spectral feature branch network; According to the weights, the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features are weightedly fused to obtain a pseudo label of the unlabeled multispectral remote sensing image.
6. The semi-supervised multispectral remote sensing image scene classification method according to claim 1, characterized in that: The unsupervised loss function is obtained according to the pseudo-label and the strong enhancement prediction results corresponding to the spatial feature and the spectral feature, including: The pseudo-label and the strong enhancement prediction results corresponding to the spatial feature and the spectral feature are substituted into the cross entropy loss function to obtain an unsupervised loss function.
7. The semi-supervised multispectral remote sensing image scene classification method according to claim 1, characterized in that: Also includes: Acquire annotated multispectral remote sensing images; Inputting the annotated multispectral remote sensing image into the dual-branch network structure to obtain a prediction result corresponding to the annotated multispectral remote sensing image; Obtaining a supervised loss function according to the labels corresponding to the annotated multispectral remote sensing image and the prediction results; The sum of the unsupervised loss function and the supervised loss function is used as the total loss function.
8. A semi-supervised multispectral remote sensing image scene classification device, characterized in that: include: An acquisition module is used to acquire an unlabeled multispectral remote sensing image, and respectively perform weak enhancement processing and strong enhancement processing on the unlabeled multispectral remote sensing image to obtain a weakly enhanced remote sensing image and a strongly enhanced remote sensing image; A prediction module, used to input the weakly enhanced remote sensing image and the strongly enhanced remote sensing image into a dual-branch network structure pre-constructed by a spatial feature branch network and a spectral feature branch network, respectively, wherein the spatial feature branch network is used to extract the spatial features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image, respectively, and obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spatial features according to the spatial features; the spectral feature branch network is used to extract the spectral features of the weakly enhanced remote sensing image and the strongly enhanced remote sensing image, respectively, based on spectral attention, and obtain the strong enhancement prediction results and the weak enhancement prediction results corresponding to the spectral features according to the spectral features; A determination module is used to fuse the weak enhancement prediction results corresponding to the spatial features and the weak enhancement prediction results corresponding to the spectral features to obtain a pseudo label of the unlabeled multispectral remote sensing image, and obtain an unsupervised loss function according to the pseudo label and the strong enhancement prediction results corresponding to the spatial features and the spectral features, wherein the unsupervised loss function is used to determine the total loss function; A training module is used to determine whether the dual-branch network structure has completed training according to the total loss function. If so, the dual-branch network structure is used as a scene classification model. If not, the parameters of the dual-branch network structure are adjusted and the step of obtaining unlabeled multispectral remote sensing images is returned until the training is completed. The classification module is used to perform scene classification on the target multispectral remote sensing image based on the scene classification model to obtain a target scene classification result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the semi-supervised multispectral remote sensing image scene classification method as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the semi-supervised multispectral remote sensing image scene classification method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Semi-supervised remote sensing target detection method and device based on detector decoupling
CN118570454A
Multi-mode-based unsupervised domain adaptive hyperspectral image classification method
CN119169451A
Semi-supervised spectral image classification method based on auxiliary task
CN119600349A
Hyperspectral remote sensing image semi-supervised classification method, apparatus, and device, and storage medium
WO2023000160A1
Cited By
Hyperspectral signal prediction method based on hyperspectral remote sensing image and pseudo label guidance
CN121834477A