Remote sensing image multi-label data classification method and device

By introducing feature fusion and guidance enhancement modules in the ResNest-50 network, the problem of failing to fully utilize image context relationships and multi-scale features in the prior art is solved, and a higher accuracy of multi-label remote sensing image classification is achieved.

CN120198710APending Publication Date: 2025-06-24PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510173914.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When using the ResNest network, the existing multi-label remote sensing image classification method fails to fully utilize the image context relationship and multi-scale features, resulting in low classification accuracy.

Method used

A multi-semantic label image data automatic classification network is designed, and the ResNest-50 network is used to combine the "up from top to bottom" feature fusion focus module, the "down from bottom to top" feature guidance enhancement module and the multi-scale feature fusion module to extract and fuse high, medium and low-level features to improve classification accuracy.

Benefits of technology

Through multi-scale feature fusion, feature redundancy and invalid information are avoided, and richer feature information is obtained, which significantly improves the classification accuracy of multi-label remote sensing images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198710A_ABST
    Figure CN120198710A_ABST
Patent Text Reader

Abstract

The invention provides a remote sensing image multi-label data classification method and device, and the method comprises the steps: obtaining a multi-semantic-label remote sensing image, dividing the multi-semantic-label remote sensing image into a training data set and a test data set according to a proportion, and carrying out the preprocessing; determining upper, middle and lower convolution features corresponding to the current training data by using the configured ResNest-50 network; utilizing a configured'top-down 'feature fusion focusing module to determine fusion features of the middle-layer and lower-layer convolution features; utilizing a configured'bottom-up 'feature guide enhancement module to determine enhanced features of the middle-layer and upper-layer fusion features; fusing the lower-layer fusion features, the middle-layer enhancement features and the upper-layer enhancement features by using a configured multi-scale feature fusion module to obtain classification features; and inputting the classification features into a classification training and label prediction unit, and determining an optimal label prediction model of the multi-label remote sensing image. According to the invention, the classification effect of the multi-semantic tag remote sensing image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular, to a method and device for classifying multi-label data of remote sensing images. Background Art

[0002] The multi-label data classification technology of remote sensing images is one of the important methods for earth observation using satellite images. During the process of satellite photographing surface coverings, due to the complex content and diverse structures of the ground scenes themselves, and the influence of adverse conditions such as lighting and cloud cover during the photographing process, the remote sensing scene image data is complex and contains noise. Currently, when using machine learning methods to train multi-semantic label remote sensing images and perform scene classification, there are often two problems that affect the classification performance. One is that there are certain feature differences in remote sensing images of the same semantic category in multi-semantic label images; the other is that there are certain feature similarities in scene images of different semantic categories in multi-semantic label images. Therefore, how to effectively utilize remote sensing images and accurately classify remote sensing scene images with multi-semantic labels has become a highly challenging topic.

[0003] In the methods of classifying multi-label remote sensing data using machine learning, currently popular deep convolutional neural networks include VGGNet-16, VGGNet-19, ResNet-50, GoogLeNet, etc. These networks usually contain millions or even tens of millions of parameters. If trained from scratch, a super-large-scale multi-semantic label remote sensing data sample is required. However, most of the data used for classifying multi-label remote sensing images is limited in quantity, and it is difficult to obtain a sufficient number of training samples to meet the needs of deep network training. Therefore, the models obtained by classifying multi-label remote sensing data using such networks have low accuracy in practical applications.

[0004] Patent CN115661551A discloses a wheat plant classification method based on an attention residual network. By adding a Split-Attention module during the feature extraction stage of the ResNest network, the network can better capture global and local details, improving the classification effect of wheat plants at the fine-grained level. However, this method does not fully utilize the relationship between image contexts and does not fuse and utilize the extracted multi-scale features.

[0005] Patent CN115620068A discloses a rock lithology classification method based on deep learning. An lithology classification model based on rock sample images is obtained by iteratively training the constructed ResNest-50 network. This method only utilizes the features of the last layer of the neural network and does not fuse and utilize multi-scale features.

[0006] Patent CN115546896A discloses a method for automatically classifying patients' lip language based on the ResNest network, but it also does not utilize the feature information obtained from other convolutional layers. The above methods prove the usability of the ResNest network in image classification, but none of them fully exploit the powerful learning ability of this network.

[0007] There are relatively few existing multi-label remote sensing image classification methods based on the ResNest network. Most similar methods only utilize the top-level features extracted by the network. Some methods add the features of different layers or add an attention mechanism after the network layer. These methods are too simple in utilizing the features extracted by the network and do not consider the mutual relationship between the features extracted from adjacent layers of the network, resulting in insufficient performance of the neural network in feature learning, a simple classification model obtained, and limited accuracy in actual classification. Summary of the Invention

[0008] The technical solution adopted by the present invention is to design an automatic classification network for multi-semantic label image data based on ResNest-50 to solve the above problems and improve the classification effect of multi-semantic label remote sensing images. In view of this, the present invention provides a method and device for classifying multi-label data of remote sensing images.

[0009] The technical solution of the present invention, a method for classifying multi-label data of remote sensing images, includes: Step S1, obtaining multi-semantic label remote sensing images, dividing them into a training data set and a test data set according to a proportion, and performing preprocessing; Step S2, using the configured ResNest-50 network to determine the upper, middle, and lower layer convolutional features corresponding to the current training data, where the ResNest-50 network includes pre-trained parameters imported on the ImageNet data set; Step S3, using the configured "top-down" feature fusion and focusing module to determine the fusion feature of the middle and lower layer convolutional features; Step S4, using the configured "bottom-up" feature guiding and enhancing module to determine the enhanced feature of the middle and upper layer fusion features; Step S5, using the configured multi-scale feature fusion module to fuse the lower layer fusion feature, the middle and upper layer enhanced features to obtain classification features; Step S6, inputting the classification features into the classification training and label prediction unit to determine the best label prediction model for multi-label remote sensing images.

[0010] In one embodiment, the step S1 includes: Step S101, randomly cropping the training data, where the cropping covers at least 80% of the training data picture; Step S102, randomly rotate the training data, where the rotation range is between -45 degrees and 45 degrees; Step S103, horizontally flip the current training data with a probability of 0.5; Step S104, crop from the center of the current training data to both sides; Step S105, convert the current training data into tensor data of a preset shape; Step S106, perform batch normalization on the tensor data channel by channel, changing the mean to 0 and the standard deviation to 1 to obtain preprocessed training data.

[0011] In one embodiment, the step S2 includes: Step S201, build a ResNest-50 network and import the pre-trained parameter "resnest50-528c19ca.pth" obtained from the ImageNet dataset; Step S202, input the preprocessed training data; Step S203, utilize the upper-layer features, middle-layer features, and lower-layer features extracted by the ResNest-50 network.

[0012] In one embodiment, the step S3 includes: Step S301, perform channel dimensionality reduction on the three features respectively using convolution, then perform feature normalization and network non-linear correction processing to obtain preprocessed upper, middle, and lower-layer features; Step S302, perform bilinear interpolation on the preprocessed upper and middle-layer features to obtain upper and middle-layer features with amplified spatial dimensions; Step S303, add the upper-layer features with amplified spatial dimensions and the preprocessed middle-layer features element by element at corresponding positions to obtain upper-middle-layer fusion features; Step S304, add an attention mechanism to the upper-middle-layer fusion features to focus on key information to obtain fused middle-layer features; Step S305, add the fused middle-layer features and the preprocessed lower-layer features element by element at corresponding positions to obtain middle-lower-layer fusion features; Step S306, add an attention mechanism to the middle-lower-layer fusion features to focus on key information to obtain fused lower-layer features.

[0013] In one embodiment, the step S4 includes: Step S401, perform bilinear interpolation on the fused middle-layer features to obtain middle-layer fusion features with amplified spatial dimensions; Step S402: Use the upper-layer features after size amplification to perform bilinear interpolation again to obtain the upper-layer features after further size amplification; Step S403: Calculate the guiding enhancement coefficient of the middle-layer features using the fused lower-layer features to obtain the middle-layer feature guiding enhancement coefficient; Step S404: Multiply the middle-layer feature guiding enhancement coefficient and the middle-layer fused features after size amplification element by element to obtain the middle-layer enhanced features; Step S405: Calculate the guiding enhancement coefficient of the upper-layer features using the middle-layer enhanced features to obtain the upper-layer feature guiding enhancement coefficient; Step S406: Multiply the upper-layer feature guiding enhancement coefficient and the upper-layer features after further size amplification element by element to obtain the upper-layer enhanced features.

[0014] In one embodiment, the step S5 includes: Step S501: Fuse the fused lower-layer features, middle-layer enhanced features, and upper-layer enhanced features using the Concat function in PyTorch along the channel dimension to obtain the classification features.

[0015] In one embodiment, the step S6 includes: Step S601: Use the global average pooling function to calculate the mean of all pixel values of each channel of the classification features to obtain the target pooling features; Step S602: Flatten the target pooling features to obtain a one-dimensional feature vector; Step S603: Input the one-dimensional feature vector into a fully connected layer to output the fully connected features.

[0016] Step S604: Input the features of the fully connected layer into the cross-entropy loss function to determine the distance loss between the true probability and the predicted probability; Step S605: Determine the semantic label corresponding to the minimum value of the distance loss as the best label prediction model for the image to be classified.

[0017] On the other hand, the present invention also provides a remote sensing image multi-label data classification device, including: An image preprocessing unit, configured to obtain multi-label remote sensing image training data and perform preprocessing; An "upper-middle-lower" feature acquisition unit, configured to use the configured ResNest-50 network to determine the convolutional features corresponding to the current training data; An "upward-downward" feature fusion and focusing unit, configured to use the configured three-scale convolutional feature fusion and focusing modules to determine the fusion features corresponding to the multi-scale convolutional features; The "bottom-up" feature-guided enhancement unit is configured to determine the enhanced feature corresponding to the fused feature by using a guided enhancement module that fuses three configured features. The multi-scale feature fusion unit is configured to fuse the fused feature and the enhanced feature to obtain a final classification feature. The remote sensing image classification unit is configured to determine the predicted class of the remote sensing image to be classified based on the final classification feature.

[0018] Adopting the above technical solutions, the present invention has at least the following advantages: 1) The present invention uses high, medium, and low three-layer multi-scale features for classification, and designs a "top-down" feature fusion focusing module, a "bottom-up" feature-guided enhancement module, and a multi-scale feature fusion module, avoiding the problem of feature redundancy, removing invalid feature information, and obtaining richer feature information. Therefore, it has a high classification accuracy.

[0019] 2) The method proposed by the present invention selects the ResNest network and improves and optimizes the model. There is no more complex backbone network, and only a 50-layer network is selected to extract features, so the number of model parameters is small. In addition, in the "top-down" feature fusion focusing and "bottom-up" feature-guided enhancement stages, most of the calculation processes involved are multiplication and addition operations, so it is more convenient. Description of the Drawings

[0020] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 It is a basic flowchart of a remote sensing image multi-label data classification method according to an embodiment of the present invention; Figure 2 It is a schematic diagram of the network architecture of a remote sensing image multi-label data classification method according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the "top-down" feature fusion focusing module according to an embodiment of the present invention; Figure 4 It is a schematic diagram of the "bottom-up" feature-guided enhancement module according to an embodiment of the present invention. Detailed Embodiments

[0021] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined purpose, the present invention will be described in detail below with reference to the accompanying drawings and preferred embodiments.

[0022] In the accompanying drawings, although exemplary embodiments of the present invention are shown, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be completely conveyed to those skilled in the art. The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0023] The first embodiment of the present invention is a method for classifying multi-label data of remote sensing images. As Figure 1 shown, it includes the following steps: Step S1, obtain a multi-semantic label remote sensing image, divide it into a training data set and a test data set according to a ratio, and perform preprocessing; Step S2, use the configured ResNest-50 network to determine the upper, middle, and lower layer convolutional features corresponding to the current training data. Among them, the ResNest-50 network includes imported pre-trained parameters on the ImageNet data set; Step S3, use the configured "top-down" feature fusion focusing module to determine the fusion feature of the middle and lower layer convolutional features; Step S4, use the configured "bottom-up" feature-guided enhancement module to determine the enhanced feature of the middle and upper layer fusion features; Step S5, use the configured multi-scale feature fusion module to fuse the lower layer fusion feature, the middle and upper layer enhanced features to obtain a classification feature; Step S6, input the classification feature into the classification training and label prediction unit to determine the optimal label prediction model for the multi-label remote sensing image.

[0024] Refer to Figure 2 below, and the method provided in this embodiment will be described in detail step by step.

[0025] Step S1, obtain a multi-semantic label remote sensing image, divide it into a training data set and a test data set according to a ratio, and perform preprocessing.

[0026] Data preprocessing is to reduce the influence of noise, increase the richness of data, and reduce the computational amount in the network learning process, so that the model has stronger generalization ability. Use the transforms function in PyTorch to perform the following processing on the training data: Step S101, randomly crop the training data to 256×256 pixels, and cover at least 80% of the picture during the cropping process; Step S102, randomly rotate the training data within the range of -45 degrees to 45 degrees; Step S103, horizontally flip the training data with a probability of 0.5; Step S104, cropping along both sides starting from the data center, the cropped image size is 224×224 pixels; Step S105, converting the format of the training data into a tensor format; Step S106, batch normalize the data channel by channel, change the mean to 0, and the standard deviation to 1; Step S2, using the configured ResNest-50 network to determine the multi-scale convolutional features corresponding to the current training data, wherein the ResNest-50 network includes pre-trained parameters imported on the ImageNet dataset.

[0027] In this embodiment, the ResNest-50 network is built, the pre-trained parameters "resnest50-528c19ca.pth" obtained on the ImageNet dataset are imported, and then the training data is input. The ResNest-50 network is used to perform convolution operations on the training data in sequence, and the output convolution features of the conv3_x layer, conv4_x layer, and conv5_x layer are extracted in sequence. , and .

[0028] Step S3, such as Figure 3 As shown, it is mainly used to calculate the feature fusion process from the upper layer (ie, L3) to the lower layer (ie, L1).

[0029] Step S301, using convolution to perform channel dimension reduction processing, normalization processing and network nonlinear correction processing on the upper layer, middle layer and lower layer features respectively, to obtain preprocessed upper layer, middle layer and lower layer features; For upper-level features , middle-level features , use the convolution kernel of 1×1, the step size of 1, the number of output feature channels of 512, and the two-dimensional convolution with 0 padding on both sides to perform "inner product operation" to obtain the features after channel dimensionality reduction.

[0030] In order to avoid network instability caused by excessive feature data after channel dimensionality reduction, batch normalization is performed on the features after channel dimensionality reduction to make them satisfy the distribution law with a mean of 0 and a variance of 1.

[0031] In order to increase the nonlinearity of the network and make the network fit better, a rectified linear unit is added to keep only the output greater than 0 and set the other inputs to 0.

[0032] At this point, the upper-layer features after preprocessing are obtained , the middle-level features after preprocessing And the lower layer features after preprocessing .

[0033] Step S302: Perform bilinear interpolation on the preprocessed upper-layer features to obtain the upper and middle-layer features with amplified spatial dimensions; Perform bilinear interpolation on the preprocessed upper-layer features by a factor of 2 to obtain features with amplified spatial dimensions .

[0034] Step S303: Add the upper-layer features with amplified spatial dimensions and the preprocessed middle-layer features element by element at corresponding positions to obtain upper-middle fusion features; Add the upper-layer features with amplified spatial dimensions and the preprocessed middle-layer features element by element at corresponding positions to enhance the features and obtain upper-middle fusion features . These features incorporate and important feature information.

[0035] Step S304: Incorporate a self-attention mechanism into the upper-middle fusion features to focus on key information and obtain fused middle-layer features; This part is mainly divided into three parts: calculating local self-attention features, calculating global self-attention features, and calculating self-attention focused features.

[0036] (1) Calculate local self-attention features First, perform "inner product operation" on the upper-middle fusion features using a 2D convolution with a kernel size of 1×1, a stride of 1, an output feature channel number of 128, and padding of 0 on both sides to obtain .

[0037] To avoid network instability caused by excessive data, perform batch normalization on to obtain . The element values of this feature follow a distribution pattern with a mean of 0 and a variance of 1.

[0038] To increase the network's non-linearity and enable better fitting, pass the convolutional feature through a rectified linear unit to obtain , where the element values are all non-negative.

[0039] Then continue to perform "inner product operation" using a 2D convolution with a kernel size of 1×1, a stride of 1, an output feature channel number of 512, and padding of 0 on both sides to obtain .

[0040] To avoid network instability caused by excessive data, perform batch normalization on to obtain the local self-attention feature of the upper-middle fusion feature . The element values of this feature satisfy the distribution law with a mean of 0 and a variance of 1.

[0041] (2)Calculate the global self-attention First, perform a two-dimensional average adaptive pooling operation on the upper-middle fusion feature to obtain the adaptive average pooling feature .

[0042] Perform an "inner product operation" on the feature using a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, an output feature channel number of 128, and padding of 0 on both sides to obtain .

[0043] To avoid network instability caused by excessive data, perform batch normalization on to obtain . The element values of this feature satisfy the distribution law with a mean of 0 and a variance of 1.

[0044] To increase the non-linearity of the network and enable the network to better fit, pass the feature through the rectified linear unit to obtain , and the element values of this feature are all not less than 0.

[0045] Then, continue to perform an "inner product operation" on the feature using a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, an output feature channel number of 512, and padding of 0 on both sides to obtain .

[0046] To avoid network instability caused by excessive data, perform batch normalization on to obtain the global self-attention feature of the upper-middle fusion feature .

[0047] (3)Calculate the self-attention focusing feature Add the global self-attention feature of the upper-middle fusion feature and the local self-attention feature of the upper-middle fusion feature element-wise and then input the result into the Sigmoid function in PyTorch to obtain the self-attention focusing feature coefficient .

[0048] Multiply the upper-middle fusion feature by the self-attention focusing feature coefficient Perform element-wise multiplication to obtain the focused features .

[0049] To avoid gradient explosion, smooth the focused features by performing convolution calculations in sequence (using a 3×3 convolution kernel, a stride of 1, 512 output feature channels, and padding of 1 on both sides), normalization processing, and adding a rectified linear unit to finally obtain the fused and focused middle-level features .

[0050] Step S305: Perform bilinear interpolation on the fused and focused middle-level features to obtain the fused and focused middle-level features with increased spatial dimensions; Perform 2-fold interpolation on the fused and focused middle-level features using the bilinear interpolation algorithm to obtain the fused and focused middle-level features with increased spatial dimensions .

[0051] Step S306: Add the fused and focused middle-level features with increased spatial dimensions and the preprocessed lower-level features element by element at corresponding positions to obtain the middle-lower fused features; The fused and focused middle-level features with increased spatial dimensions and the preprocessed lower-level features are added element by element at corresponding positions to achieve feature enhancement and obtain the middle-lower fused features . This feature fuses the important elements of the features , , .

[0052] Step S307: Add an attention mechanism to the middle-lower fused features to focus on key information and perform smoothing to obtain the fused and focused lower-level features.

[0053] This part is mainly divided into three parts: calculating local self-attention features, calculating global self-attention features, and calculating self-attention focused features.

[0054] (1) Calculate local self-attention features First, perform "inner product operation" on the middle-lower fused features using a 2D convolution with a 1×1 convolution kernel, a stride of 1, 128 output feature channels, and padding of 0 on both sides to obtain the feature .

[0055] To avoid network instability caused by excessive feature data, perform batch normalization on to obtain The element values of this feature satisfy the distribution law with a mean of 0 and a variance of 1.

[0056] To increase the non-linearity of the network and enable the network to better fit, the feature is obtained through the activation operation of the rectified linear unit The element values of this feature are all not less than 0.

[0057] Then, continue to perform an "inner product operation" using a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, an output feature channel number of 512, and padding of 0 on both sides to obtain .

[0058] To avoid the network instability caused by excessive data, perform batch normalization on to obtain the local self-attention feature of the middle-lower fusion feature The element values of this feature satisfy the distribution law with a mean of 0 and a variance of 1.

[0059] (2) Calculate the global self-attention Perform a two-dimensional average adaptive pooling operation on the middle-lower fusion feature first to obtain the adaptive average pooling feature .

[0060] Perform an "inner product operation" on the feature using a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, an output feature channel number of 128, and padding of 0 on both sides to obtain .

[0061] To avoid the network instability caused by excessive data, perform batch normalization on to obtain The element values of this feature satisfy the distribution law with a mean of 0 and a variance of 1.

[0062] To increase the non-linearity of the network and enable the network to better fit, the feature is obtained through the activation operation of the rectified linear unit The element values of this feature are all not less than 0.

[0063] Then, perform an "inner product operation" on the feature continuing to use a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, an output feature channel number of 512, and padding of 0 on both sides to obtain .

[0064] To avoid the network instability caused by excessive data, perform batch normalization on Perform batch normalization on the data to obtain the global self-attention feature of the middle-lower fusion feature 。

[0065] (3)Calculate the self-attention focused feature The global self-attention feature of the middle-lower fusion feature and the local self-attention feature of the middle-lower fusion feature are element-wise added and then input into the Sigmoid function to obtain the self-attention focused feature coefficient 。

[0066] The middle-lower fusion feature and the self-attention focused feature coefficient are element-wise multiplied to obtain the focused feature 。

[0067] To avoid gradient explosion, smooth processing is performed on the focused feature by performing convolution calculations in sequence (using a convolution kernel of 3×3, a stride of 1, an output feature channel number of 512, and padding of 1 on both sides), normalization processing, and adding a rectified linear unit, and finally obtaining the fused and focused lower-layer feature 。

[0068] Step S4, as Figure 4 shown, use the configured "bottom-up" feature-guided enhancement module to determine the enhanced features of the middle and upper layer fusion features; Step S401, calculate the guided enhancement coefficient of the fused and focused middle-layer feature using the fused and focused lower-layer feature to obtain the guided enhancement coefficient of the fused and focused middle-layer feature; This part is divided into three parts: global relationship feature calculation, local relationship feature calculation, and guided enhancement coefficient calculation: (1)Global relationship feature calculation Perform a two-dimensional average adaptive pooling operation on the fused and focused lower-layer feature to obtain the adaptive average pooling feature; perform an "inner product operation" using a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, an output feature channel number of 32, and padding of 0 on both sides; after activation by a rectified linear unit, perform an "inner product operation" using a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, an output feature channel number of 512, and padding of 0 on both sides to obtain the global relationship feature 。

[0069] (2)Local relationship feature calculation Perform a two-dimensional average adaptive pooling operation on the fused and focused lower-layer feature Perform a two-dimensional adaptive maximum pooling operation to obtain an adaptive maximum pooling feature; perform an "inner product operation" using a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, an output feature channel number of 32, and padding of 0 on both sides; after activation by a rectified linear unit, perform an "inner product operation" using a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, an output feature channel number of 512, and padding of 0 on both sides to obtain a local relationship feature 。

[0070] (3)Guided enhancement coefficient calculation Add the global relationship feature and the local relationship feature element-wise to obtain a relationship feature 。 Then input the relationship feature into the Sigmoid function to obtain the guided enhancement coefficient 。

[0071] Step S402: Multiply the guided enhancement coefficient for fusing and focusing the middle-level features with the middle-level fused and focused feature after size amplification channel-wise to obtain a middle-level enhanced feature; Multiply the guided enhancement coefficient with the middle-level fused and focused feature after spatial size amplification channel-wise to obtain a middle-level enhanced feature 。

[0072] Step S403: Perform bilinear interpolation on the upper-level feature after spatial size amplification to obtain an upper-level feature after further size amplification; Perform 2-fold interpolation on the upper-level feature after spatial size amplification using the bilinear interpolation algorithm to obtain an upper-level feature after further size amplification 。

[0073] Step S404: Calculate the guided enhancement coefficient for the upper-level feature after further size amplification using the middle-level enhanced feature to obtain the guided enhancement coefficient for the upper-level feature; This part is divided into three parts: global relationship feature calculation, local relationship feature calculation, and guided enhancement coefficient calculation: (1)Global relationship feature calculation Perform a two-dimensional average adaptive pooling operation on the middle-level enhanced feature to obtain an adaptive average pooling feature; perform an "inner product operation" using a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, an output feature channel number of 32, and padding of 0 on both sides; after activation by a rectified linear unit, perform an "inner product operation" using a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, an output feature channel number of 512, and padding of 0 on both sides to obtain a global relationship feature 。

[0074] (2) Local relationship feature calculation Perform two-dimensional adaptive max pooling operation on the middle-layer enhanced features to obtain the adaptive max pooling features; perform "inner product operation" using a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, an output feature channel number of 32, and zero padding on both sides; after activation by a rectified linear unit, perform "inner product operation" using a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, an output feature channel number of 512, and zero padding on both sides to obtain the local relationship features .

[0075] (3) Guided enhancement coefficient calculation Add the global relationship features and the local relationship features element-wise to obtain the relationship features . Then input the relationship features into the Sigmoid function to obtain the guided enhancement coefficient .

[0076] Step S405: Multiply the guided enhancement coefficient with the upper-layer features after size re-augmentation channel by channel to obtain the upper-layer enhanced features; Multiply the guided enhancement coefficient with the upper-layer features after spatial size re-augmentation channel by channel to obtain the upper-layer enhanced features .

[0077] Step S5: Use the configured multi-scale feature fusion module to fuse the lower-layer fusion focused features, middle-layer enhanced features, and upper-layer enhanced features to obtain the classification features; Fusion features and Use the Concat function in PyTorch to fuse the pooled relationship attention feature maps to generate the final classification features. The calculation formula is as follows: Finally, obtain the classification feature L 123 .

[0078] Step S6: Based on the classification features, determine the predicted category of the remote sensing image to be measured.

[0079] After obtaining the classification features , successively determine the distance between the true probability and the predicted probability through global average pooling, a fully connected network layer, and a cross-entropy loss function. The predicted probability corresponding to the smaller loss is the best prediction model for multi-label remote sensing image data.

[0080] Compared with the prior art, this embodiment has at least the following advantages: 1) The present invention classifies using high, medium, and low three-layer multi-scale features, and designs a "top-down" feature fusion focusing module, a "bottom-up" feature guiding enhancement module, and a multi-scale feature fusion module, avoiding the problem of feature redundancy, eliminating invalid feature information, and obtaining richer feature information. Therefore, it has a high classification accuracy.

[0081] 2) The method proposed by the present invention selects the ResNest network and improves and optimizes the model. There is no more complex backbone network, and only a 50-layer neural network is selected to extract features, so the model parameter quantity is small. In addition, in the "top-down" feature fusion focusing and "bottom-up" feature guiding enhancement stages, most of the calculation processes involved are multiplication and addition operations, so it is more convenient.

[0082] The second embodiment of the present invention corresponds to the first embodiment. This embodiment introduces a multi-label data classification device for remote sensing images, including the following components: An image preprocessing unit, configured to obtain multi-label remote sensing image data, divide it into a training data set and a test data set according to a ratio, and perform preprocessing; A "top-middle-bottom" feature acquisition unit, configured to use the configured ResNest-50 network to determine the convolutional features corresponding to the current training data; A "top-down" feature fusion focusing unit, configured to use the configured three-scale convolutional feature fusion focusing module to determine the fusion features corresponding to the multi-scale convolutional features; A "bottom-up" feature guiding enhancement unit, configured to use the configured three fusion feature guiding enhancement modules to determine the enhanced features corresponding to the fusion features; A multi-scale feature fusion unit, configured to fuse the fusion features and the enhanced features to obtain the final classification features; A remote sensing image classification unit, configured to determine the best label prediction model for the multi-label remote sensing image based on the final classification features.

[0083] Through the description of the specific implementation manners, it should be possible to understand more deeply and specifically the technical means and effects adopted by the present invention to achieve the predetermined purpose. However, the attached drawings are only for reference and illustration, and are not used to limit the present invention.

Claims

1. A remote sensing image multi-label data classification method, characterized in that: include: Step S1, obtaining a multi-semantic label remote sensing image, dividing it into a training data set and a test data set in proportion, and performing preprocessing; Step S2, using the configured ResNest-50 network to determine the upper, middle, and lower convolutional features corresponding to the current training data, wherein the ResNest-50 network includes pre-trained parameters imported on the ImageNet dataset; Step S3, using the configured "top-down" feature fusion focusing module to determine the fusion features of the middle and lower convolutional features; Step S4, using the configured "bottom-up" feature to guide the enhancement module, determine the enhancement features of the middle layer and upper layer fusion features; Step S5, using the configured multi-scale feature fusion module to fuse the lower layer fusion features, the middle layer and the upper layer enhancement features to obtain the classification features; Step S6, inputting the classification features into the classification training and label prediction unit to determine the label prediction model of the multi-label remote sensing image.

2. A remote sensing image multi-label data classification method as claimed in claim 1, characterized in that: The step S1 comprises: Step S101, randomly cropping the training data, wherein the cropping covers at least 80% of the training data image; Step S102, randomly rotating the training data, wherein the rotation range is between -45 degrees and 45 degrees; Step S103, horizontally flipping the current training data with a probability of 0.5; Step S104, cutting from the current training data center to both sides; Step S105, converting the current training data into tensor data of a preset shape; Step S106, batch normalization is performed on the tensor data channel by channel, the mean is changed to 0, and the standard deviation is changed to 1, so as to obtain preprocessed training data.

3. A remote sensing image multi-label data classification method as claimed in claim 2, characterized in that: The step S2 comprises: Step S201, build the ResNest-50 network and import the pre-trained parameters "resnest50-528c19ca.pth" on the ImageNet dataset; Step S202, input current training data; Step S203, using the upper-layer features, middle-layer features and lower-layer features extracted by the ResNest-50 network.

4. A remote sensing image multi-label data classification method and device as claimed in claim 3, characterized in that: The step S3 comprises: Step S301, using convolution to perform channel dimension reduction processing, normalization processing and network nonlinear correction processing on the upper layer, middle layer and lower layer features respectively, to obtain preprocessed upper layer, middle layer and lower layer features; Step S302, performing bilinear interpolation processing using the preprocessed upper-layer features to obtain the upper-layer features after size enlargement; Step S303, adding the upper-layer features after the size expansion and the middle-layer features after the preprocessing one by one according to the elements at the corresponding positions to obtain the upper-middle fusion features; Step S304, adding a self-attention mechanism to the upper-middle fusion features to focus on key information, and performing smoothing to obtain fused and focused middle-level features; Step S305, performing bilinear interpolation processing using the fused focused middle-level features to obtain the fused focused middle-level features after size enlargement; Step S306, adding the size-enlarged fused focused middle-layer features and the preprocessed lower-layer features one by one according to the elements at corresponding positions to obtain middle-lower fused features; Step S307, adding an attention mechanism to the middle-lower fusion features to focus on key information, and performing smoothing to obtain fused and focused lower-layer features.

5. A remote sensing image multi-label data classification method as claimed in claim 4, characterized in that: The step S4 comprises: Step S401, performing bilinear interpolation processing using the fused focused middle-level features to obtain the size-enlarged middle-level fused focused features; Step S402, using the upper-layer features after the size expansion to perform bilinear interpolation processing again to obtain the upper-layer features after the size expansion; Step S403, using the fused lower-layer features to calculate the guided enhancement coefficient of the middle-layer features to obtain the guided enhancement coefficient of the middle-layer features; Step S404, using the middle-level feature to guide the enhancement coefficient and multiply the middle-level fusion feature after the size expansion element by element to obtain the middle-level enhancement feature; Step S405, the middle-layer enhancement feature calculates the guided enhancement coefficient of the upper-layer feature to obtain the guided enhancement coefficient of the upper-layer feature; Step S406, using the upper-layer feature to guide the enhancement coefficient and the upper-layer feature after the size expansion to multiply the element by element, so as to obtain the upper-layer enhanced feature.

6. A remote sensing image multi-label data classification method as claimed in claim 5, characterized in that: The step S5 comprises: Step S501, the fused lower-layer features, middle-layer enhanced features, and upper-layer enhanced features are fused according to the channel dimension using the Concate function in PyTorch to obtain classification features.

7. A remote sensing image multi-label data classification method as claimed in claim 6, characterized in that: The step S6 comprises: Step S601, using a global average pooling function to average all pixel values ​​of each channel of the classification feature to obtain a target pooling feature; Step S602, flattening the target pooling feature to obtain a one-dimensional feature vector; Step S603, inputting the one-dimensional feature vector into a fully connected layer to output a fully connected feature; Step S604, inputting the features of the fully connected layer into a cross entropy loss function to determine the distance loss between the true probability and the predicted probability; Step S605: Determine the category corresponding to the minimum value of the distance loss as the best label prediction model for the image to be classified.

8. A remote sensing image multi-label data classification device, characterized in that: include: An image preprocessing unit is configured to obtain multi-label remote sensing image data, divide it into a training data set and a test data set in proportion, and perform preprocessing; The "upper, middle and lower" feature acquisition unit is configured to use the configured ResNest-50 network to determine the convolutional features corresponding to the current training data; A "top-down" feature fusion focusing unit is configured to determine a fusion feature corresponding to the multi-scale convolution feature by using the configured three-scale convolution feature fusion residual module; A "bottom-up" feature guided enhancement unit, configured to use the three configured guided enhancement modules of the fused features to determine the enhanced features corresponding to the fused features; A multi-scale feature fusion unit, configured to fuse the fused features and the enhanced features to obtain a final classification feature; The remote sensing image classification unit is configured to determine the best label prediction model for the multi-label remote sensing image based on the final classification features.