High-resolution remote sensing image scene classification method based on FACNNCN

By using the FACNNCN method to aggregate features of high-resolution remote sensing images and fuse capsule module information, the scene classification problem of complex remote sensing images that is difficult to distinguish in existing technologies is solved, and scene classification with higher accuracy and robustness is achieved.

CN120808001APending Publication Date: 2025-10-17SUZHOU ZHIPRICE YUNHUI INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510867353.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing high-resolution remote sensing image scene classification methods have difficulty effectively distinguishing different images of the same scene category and similar images of different scene categories when faced with large-scale, multi-category and complex morphology remote sensing images.

Method used

A high-resolution remote sensing image scene classification method based on FACNNCN is adopted. The features of multiple convolutional layers of the VGG-16 network are aggregated through the feature aggregation module, and the capsule module is combined for information fusion. The low-level and high-level capsule layers are used to encode the attributes of the ground objects and scenes. The coupling coefficient is optimized through the dynamic routing algorithm to achieve accurate classification of the scene.

Benefits of technology

It improves the network's ability to distinguish complex scenes, accurately expresses the spatial information of ground objects, and improves the accuracy and robustness of scene classification, especially showing excellent classification performance in large-scale and multi-category remote sensing image data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808001A_ABST
    Figure CN120808001A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing, in particular to a high-resolution remote sensing image scene classification method based on FACNNCN. According to the technical scheme, the method comprises the following steps: firstly, carrying out unified preprocessing on an input RGB high-resolution remote sensing image; the method comprises the steps that firstly, a VGG-16 network is utilized to extract multi-layer convolution features, convolution features with the distinguishing enhancing capacity are generated through cross-layer aggregation, then the aggregation features are input into a vector capsule network, modeling of ground feature local features and the spatial relation of the ground feature local features is achieved by constructing low-layer capsules and high-layer capsules, and coupling between the capsules is optimized through a dynamic routing algorithm; and finally, judging the scene category to which the image belongs through the module length of the output vector of the high-level capsule. Experimental results show that the method has excellent performance on two remote sensing scene classification data sets, and the defects that scene image feature extraction is insufficient and ground feature space features are not considered in a current high-resolution remote sensing image scene classification method based on the convolutional neural network are effectively overcome.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of remote sensing, in particular to a high-resolution remote sensing image scene classification method based on FACNNCN. BACKGROUND

[0002] With the rapid development of remote sensing technology, high-resolution remote sensing images have been widely applied in the fields of city planning, land use, environmental monitoring and military reconnaissance. In the face of the increasing amount of remote sensing data and the complex and diverse ground object scenes, the traditional pixel-based classification method is difficult to meet the demand for semantic understanding of remote sensing images, and the scene-level classification method of remote sensing images emerges as the times require and becomes an important research direction of current remote sensing intelligent interpretation.

[0003] The existing high-resolution remote sensing image scene classification methods can be roughly divided into two categories: artificial feature-based and deep learning-based. The former relies on color histograms, direction gradient histograms, local binary patterns and other manual features, but its expression ability is limited and is susceptible to noise interference. The latter uses deep neural networks, especially convolutional neural networks (CNN), which have achieved remarkable results in computer vision tasks such as image classification and are widely applied in remote sensing scene classification. For example, the classic CNN architectures such as VGGNet and ResNet have achieved certain improvement in classification accuracy in the field of remote sensing.

[0004] However, the existing CNN-based methods still have the following shortcomings in remote sensing scene classification:

[0005] When the scale of remote sensing image scene data is large, the number of categories is large, and the form is complex, CNN and its extended models still have difficulty in effectively distinguishing different images of the same scene category and similar images of different scene categories.

[0006] Therefore, we propose a high-resolution remote sensing image scene classification method based on FACNNCN to solve the existing problems. SUMMARY

[0007] The purpose of the present application is to propose a high-resolution remote sensing image scene classification method based on FACNNCN to solve the problems in the background art.

[0008] To achieve the above purpose, the present application provides the following technical scheme: a high-resolution remote sensing image scene classification method based on FACNNCN, comprising the following steps:

[0009] Input layer

[0010] Step one input preprocessing: obtaining a high-resolution remote sensing scene image with three color channels of RGB, and performing uniform size cropping and normalization processing to form a structured input sample set;

[0011] Feature aggregation module

[0012] The main calculation steps of the feature aggregation module are as follows:

[0013] Step two, convolution feature extraction: for the input high-resolution remote sensing scene image, VGG-16 network is used to perform average pooling on the last convolution layer conv3_3, conv4_3 and conv5_3 of the third, fourth and fifth convolution pooling groups respectively, to obtain new intermediate features with the same size, which are the supplements of the original features, and are denoted as X1 R H×W×D1 , X2 R H×W×D2 and X3 R H×W×D3 , and the obtained features are aggregated in turn to obtain a new convolution feature denoted as AF1:

[0014] AF1 = [X1; X2; X3] R H×W×(D1+D2+D3) (formula 1);

[0015] Step three, first stage feature aggregation: 1x1 convolution and ReLU operation are performed on the convolution feature AF1 of step two to realize information fusion between aggregated features across channels and increase the non-linear interaction between different channels. The number of convolution kernels is set to N, and at this time the dimension of AF1 changes from D1+D2+D3 to N

[0016] Step four, second stage feature aggregation: AF1 is aggregated with the feature X4 obtained by maximum pooling of the convolution layer conv5-3, denoted as X4 R H×W×D4 , L = N + D4, and the aggregated convolution feature AF2 output by the feature aggregation module is taken as the output of the feature aggregation module:

[0017] AF2 = [AF1; X4] R H×W×L (formula 2);

[0018] Capsule module

[0019] The capsule module mainly consists of a convolution layer, a low-level capsule layer and a high-level capsule layer. The main calculation steps of the capsule module are as follows:

[0020] Step five, low-level capsule layer and high-level capsule layer construction: the number of neurons S1 contained in each capsule of the low-level capsule layer is set, and a low-level capsule layer containing HxWxL / S1 capsules is constructed, where any capsule i corresponds to a feature vector with a length of S1 in the convolution feature AF2, i.e. the output vector of the capsule i. The low-level capsule layer is used to describe smaller ground objects in the scene image and encode the attributes of the ground objects to provide the probability that the ground object belongs to a certain scene type.

[0021] The value of dimension S2 of each capsule of the high-level capsule layer and the number T of capsules are set, and the value of T is the number of prediction categories of the scene image. A high-level capsule layer containing T capsules is constructed, wherein any one capsule j corresponds to an S2-dimensional feature vector. The high-level capsule layer is used to describe the entire scene, and the category to which the scene belongs is determined based on the encoded attributes.

[0022] For example, for a train station scene image, the low-level capsules are used to describe the platform, train, rail, building and other components of the scene, and encode the attributes of the entities, and the high-level capsules are used to describe the entire scene and encode the attributes of the scene. Through the information transmission between the low-level capsule layer and the high-level capsule layer, the capsule module can learn the relationship between the smaller ground objects in the scene image and the overall scene.

[0023] Step six input vector calculation: the low-level capsules are subjected to affine transformation by the weight matrix to calculate the prediction vector of the low-level capsules to the high-level capsules, i.e. the input vector of the high-level capsules, i.e. the connection between the capsule layers. The prediction vector of the low-level capsule i to the high-level capsule j is denoted as

[0024]

[0025] W ij is a weight matrix, which can be optimized by back propagation. All low-level capsules predict the high-level capsule j, therefore, the input vector of the high-level capsule j is is the weighted sum of all low-level capsule prediction vectors:

[0026]

[0027] wherein c ij is a coupling coefficient, which is optimized in the dynamic routing iteration process, and c ij is calculated by formula 5, and the sum of the coupling coefficients of capsule i and all capsules in the high-level is 1:

[0028]

[0029] b ij represents the logarithmic prior probability of the coupling of the low-level capsule i and the high-level capsule j, and the initial value is set to 0.

[0030] Step seven output vector acquisition: the coupling coefficients and the logarithmic prior probability are iteratively updated between the low-level capsules and the high-level capsules through the dynamic routing algorithm, and the high-level capsule input vector and the output vector are calculated, the output vector is processed by the nonlinear squashing function squash, so that its length does not exceed 1, and the length of the output vector of the high-level capsule represents the prediction probability of the scene image category, therefore, in order to ensure that the length of the vector does not exceed 1, the nonlinear squashing function squash is needed, so that the short vector is almost shrunk to 0, and the long vector is shrunk to slightly less than 1, the high-level capsule j obtains the output vector v j :

[0031]

[0032] Step eight parameter update: the logarithmic probability is updated according to the inner product of the output vector and the prediction vector, the coupling coefficient is iteratively optimized, the parameter update of the capsule module is realized, and the spatial information transmission between the ground object and the overall scene is completed.

[0033] In the capsule module, the update of the logarithmic probability b ij depends on the consistency of the vector and the vector , that is, the two vectors have a large inner product, therefore, the update formula of b ij is as follows:

[0034]

[0035] Formula 3-7 constitutes a complete dynamic routing process for calculating , after updating the logarithmic probability b ij , the coupling coefficient c ij is updated according to formula (5), after step 8, step 6 is turned to, and the iteration of the multiple dynamic routing processes is performed, the parameter update of the capsule module is realized, and the spatial information transmission between the ground object and the overall scene is completed.

[0036] Step nine loss function calculation: the loss function of each capsule k in the high-level capsule layer is represented as L k:

[0037] L k =T k ·max(0,m + -||v k ||) 2 +λ(1-T k )·max(0,||v k ||-m - ) 2 (Formula 8)

[0038] The length of the high-level capsule vector corresponding to each category is taken as the scene prediction probability, and the loss function Lk Model parameter training is performed.

[0039] T k = 1, the total loss in the training process is the sum of the losses of all capsules in the layer, m+, m-, and λ are hyperparameters that need to be set according to the input scene image.

[0040] Output layer

[0041] Step ten, category decision: compare the lengths of all high-level capsule output vectors, and the category corresponding to the longest vector is the predicted scene category C k ; the output layer determines the category C of the predicted scene image by the length of the vector output by the high-level capsule layer (formula 10). k

[0042]

[0043]

[0044] Compare the lengths of multiple output vectors, and the category corresponding to the longest vector is the predicted category C k of the scene image.

[0045] Step eleven, output result: output the predicted scene image category C k , and complete the classification task of high-resolution remote sensing images.

[0046] Preferably, the parameters of the VGG-16 model used to extract convolutional features are pre-trained on a large general image dataset before use to enhance its initial feature extraction capability for natural scenes, and then fine-tuned on a high-resolution remote sensing image scene dataset using stochastic gradient descent to adapt to the structure features of ground objects in remote sensing scenes and complex backgrounds, thereby improving the adaptability of the model to remote sensing data and the classification accuracy.

[0047] Preferably, in the first-stage feature aggregation step, convolution and nonlinear activation operations are applied to multiple intermediate convolutional feature maps to realize deep fusion of features between channels, the number of convolution kernels used is set in the medium to high dimension range according to the complexity of the feature map, and the rectified linear unit function is selected as the nonlinear activation function to enhance the nonlinear and discriminative ability of the aggregated feature representation.

[0048] ​Preferably, in the second stage feature enhancement process, the final aggregated feature is obtained by splicing the first stage aggregated feature and the fifth convolution group pooled main feature map, and further unifying the dimension and enhancing the semantic expression ability through a layer of channel compression convolution, so as to obtain the final aggregated feature with more rich context information, and enhance the recognition ability of complex scenes.

[0049] Preferably, the low-layer capsule layer is composed of a plurality of capsule units with fixed dimensions, each capsule unit being used to express a plurality of attribute information of a local ground object entity in the remote sensing image, including but not limited to position, direction, scale, texture, a plurality of local regions are divided on the aggregated feature map in a sliding window manner, and the features in each region are encoded, thereby forming a low-layer capsule set for supporting subsequent modeling of spatial relationships.

[0050] Preferably, the number of capsules in the high-layer capsule layer is equal to the total number of categories involved in scene classification, each capsule is used to represent the global semantic information of a specific scene category, and the dimension of each capsule is higher than that of the low-layer capsule, so as to carry more comprehensive attributes, and the mapping relationship is established between the low-layer capsule and the high-layer capsule through the mapping weight matrix, thereby realizing the abstract expression of multi-level spatial semantics.

[0051] Preferably, the dynamic routing process adopts an iterative optimization-based mechanism to establish coupling weights between the low-layer capsule and the high-layer capsule, the initial weights are set according to a preset logarithmic priori, and the weight distribution is adjusted according to the current output result in each iteration, so that the low-layer capsule with highly similar information contributes more weight to the specific high-layer capsule, thereby gradually optimizing the information flow and expression path, and the iteration process is usually performed for three to five rounds to ensure convergence effect and stability.

[0052] Preferably, the loss function of the capsule module is constructed based on a two-sided boundary strategy controlled by a category label, the length of the capsule output of the correct category is respectively constrained and the capsule output of the wrong category is suppressed, so as to realize stable convergence and improved discrimination ability of the network in the training process, and related hyperparameters are set according to the distribution characteristics of the input data to adapt to remote sensing images of different resolutions and complexity.

[0053] Compared with the prior art, the present application has the following advantages:

[0054] The present application obtains enhanced scene image classification features through feature aggregation, effectively improves the network's ability to distinguish complex scenes, and more accurately expresses the spatial information of ground objects such as position, angle, rotation, inclination and size based on the enhanced features, better learns the spatial relationship between ground objects and scenes, and further improves the accuracy of scene classification in complex backgrounds.

[0055] And its capsule module uses vector capsule to express and learn the spatial relationship between the whole and part of the scene, which can deeply mine the spatial features of the scene image, therefore, the FACNNCN fully utilizes the feature information contained in the scene image, and has better complex scene classification performance. BRIEF DESCRIPTION OF DRAWINGS

[0056] Figure 1 The network structure schematic diagram of the FACNNCN of the application;

[0057] Figure 2 The connection schematic diagram between the low-layer capsule and the high-layer capsule of the application;

[0058] Figure 3 The scene category schematic diagram of the UC Merced dataset of the application;

[0059] Figure 4 The scene category schematic diagram of the NWPU dataset of the application;

[0060] Figure 5 The classification accuracy and error comparison schematic diagram of different models under the 80% training ratio of the UC Merced dataset of the application;

[0061] Figure 6 The confusion matrix schematic diagram of the classification result of the FACNNCN method under the 80% training ratio of the UC Merced dataset of the application;

[0062] Figure 7 The classification accuracy and error comparison schematic diagram of different models under the 20% ratio of the NWPU dataset of the application;

[0063] Figure 8 The confusion matrix schematic diagram of the classification result of the FACNNCN method under the 10% training ratio of the NWPU dataset of the application;

[0064] Figure 9 The confusion matrix schematic diagram of the classification result of the FACNNCN method under the 20% training ratio of the NWPU dataset of the application. DETAILED DESCRIPTION

[0065] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the application.

[0066] Convolutional neural network (CNN)

[0067] Convolutional Neural Network (CNN) is a neural network model inspired by the mechanism of biological receptive field. CNN has achieved great success in various tasks in image processing field since its proposal, such as image classification, face recognition, image segmentation, etc. With the continuous development of CNN research, a series of excellent CNN models have emerged. AlexNet is the first modern deep CNN model, which is a breakthrough in the application of deep learning technology in image classification. After AlexNet, a number of deep CNN models have emerged, such as VGGNet, GoogLeNet, ResNet, DenseNet, etc. VGGNet is a new deep CNN model developed by the Computer Vision Group of Oxford University and Google DeepMind. The model uses smaller receptive fields and deep network structure to significantly improve the accuracy of image classification. The two commonly used structures of VGGNet are VGG-16 and VGG-19, of which VGG-16 is more widely used.

[0068] VGG-16 contains 5 convolutional pooling groups, each of which contains 2, 2, 3, 3, and 3 convolutional layers, respectively. Convolutional layer is the most core part of VGG-16. Convolutional layers obtained by different sizes of convolutional kernels can effectively extract multi-level and complex features of images. The size of the convolutional kernel is usually 3x3 or 5x5. At the end of each convolutional pooling group, there is a pooling layer. The pooling layer mainly compresses the feature map obtained by the convolutional layer to reduce the dimension of the feature. The commonly used pooling operations mainly include max pooling and average pooling. After the 5 convolutional pooling groups, VGG-16 contains 3 fully connected layers and 1 softmax output layer. The fully connected layer highly synthesizes the features obtained after convolution and pooling operations and inputs them into the classifier for classification. Generally, the output dimension of the last layer of the fully connected layer is the number of classes of the classification task. In order to increase the non-linear relationship between layers, the convolutional layer and the fully connected layer of VGG-16 use ReLU function to increase the non-linear expression ability of the network.

[0069] Vector Capsule Network CapsNet

[0070] Capsule Network is a new deep learning model proposed by Hinton et al. in 2011, which is a variant of CNN

[18] . In 2017, Hinton et al. proposed an improved model of Capsule Network, Vector Capsule Network and dynamic routing algorithm between capsules

[15] . The capsule in Vector Capsule Network is a vector composed of a group of neurons. The parameters of the vector represent different attributes of entities, such as position, size, direction, deformation, velocity, reflectivity, chroma, texture, etc. The length of the vector represents the probability of the existence of the entity.

[0071] The vector capsule network is composed of a convolutional layer, a low-level capsule layer, and a high-level capsule layer. The convolutional layer extracts image features and serves as an input to the low-level capsule layer. The low-level capsule layer vectorizes the convolutional features of the image to form multiple vector capsules, realizing the encoding of entity attribute information. The high-level capsule layer, also known as the class capsule layer, accepts the vector input from the low-level capsule layer, forms new vector capsules, and calculates the probability of the capsule belonging to a class through the length of the vector capsule, thereby completing the classification task. To realize information transmission and dynamic connection between different capsule layers, the vector capsule network uses a dynamic routing algorithm to automatically select more effective capsules and thereby improve the performance of the model. The dynamic routing algorithm uses an iterative method to connect capsules between different capsule layers. Since a capsule is a set of neurons representing different features, it can establish the positional relationship between different entities during the routing process. Therefore, the vector capsule network is more robust to changes in spatial information such as the position and angle of entities.

[0072] Fusion network CNNCN

[0073] Convolutional neural networks have strong representation learning capabilities and perform excellently in large-scale and complex tasks. With the rapid development of numerical computing devices, especially the powerful support of GPU computing clusters, the performance of convolutional neural networks has been more fully utilized. However, the pooling operation of convolutional neural networks loses a large amount of useful information and cannot identify the spatial structure between parts and the whole in images. The vector capsule network, as a variant of convolutional neural networks, is designed to address the above shortcomings of convolutional neural networks. However, the performance of the vector capsule network cannot surpass that of convolutional neural networks in most application tasks.

[0074] The CNNCN network, which fuses convolutional neural networks and vector capsule networks, fully utilizes the advantages of both networks. It can obtain deep convolutional features using convolutional neural networks, fully utilize the strong representation ability of convolutional neural networks, and obtain rich spatial information through the good spatial feature expression ability of vector capsule networks. The constructed network model has stronger feature expression ability and better robustness.

[0075] CNNCN has been successfully applied to high-resolution remote sensing image scene classification and has achieved good classification performance. However, with the continuous development of high-resolution remote sensing image scene classification applications, the complexity of scene image classification tasks is continuously increasing, and people's requirements for scene image classification accuracy are also continuously improving. Therefore, how to improve the CNNCN network to further improve the classification performance of CNNCN is worth further research.

[0076] The application proposes a new FACNNCN network on the basis of the CNNCN network and applies it to high-resolution remote sensing image scene classification. The FACNNCN adds new aggregated features to the traditional CNNCN network, further enhances the expression ability of features in scene classification through the aggregated features.

[0077] The FACNNCN mainly comprises an input layer, a feature aggregation module, a capsule module and an output layer. Figure 1 The network structure of the FACNNCN is described in detail. The input of the input layer is multiple high-resolution remote sensing scene images with class labels. The capsule module takes the convolutional features obtained by the feature aggregation module as the input of the vector capsule network, encodes the spatial information of the scene image features through vectorization of the convolutional features, and learns the spatial relationship between the local and the whole of the scene image.

[0078] The application is based on an experimental platform equipped with a 3.6GHz 8-core E5-1650 v4 CPU, 64GB memory and a Linux operating system, and uses GPU for acceleration.

[0079] The UC Merced Land-Use Dataset is a high-resolution remote sensing dataset containing 21 typical scene categories, which is the first publicly available high-resolution remote sensing image scene classification dataset. The 21 scene categories are: farm, airplane, baseball diamond, beach, buildings, chaparral, dense residential, forest, freeway, golf course, harbor, intersection, medium density residential, mobile home park, overpass, parking lot, river, runway, sparse residential, storage tank and tennis court.

[0080] NWPU-RESISC45 Dataset dataset has rich scene categories, is the largest high-resolution remote sensing image scene classification dataset at present, and has high intra-class diversity and inter-class similarity, the dataset contains 45 scene categories, each scene category contains 700 images with a size of 256*256, the spatial resolution is from about 30 to 0.2 meters, and the 45 scene categories are: airplane, airport, baseball field, basketball court, beach, bridge, jungle, church, circular farmland, cloud, commercial area, dense residential area, desert, forest, highway, golf course, track and field, port, industrial area, intersection, island, lake, grassland, medium-density residential area, mobile home park, mountain, overpass, palace, parking lot, railway, railway station, rectangular farmland, river, roundabout, runway, sea ice, ship, snow peak, sparse residential area, stadium, storage tank, tennis court, terrace, thermal power plant and wetland.

[0081] Embodiment one

[0082] As Figures 1-9 shown, the application provides a high-resolution remote sensing image scene classification method based on FACNNCN, which comprises:

[0083] Input layer

[0084] The step one input preprocessing: obtaining high-resolution remote sensing scene image with RGB three color channels, and carrying out uniform size cutting and normalization processing to form a structured input sample set;

[0085] The input of the method is a high-resolution remote sensing scene image with RGB three color channels, in order to facilitate model comparison, the scene image is randomly cropped, and the input size becomes 224*224*3.

[0086] Feature aggregation module

[0087] In the feature aggregation module, the size of the convolution kernel of the VGG-16 network is 3*3, each convolution pooling group contains a maximum pooling layer, and the size of the pooling operation window is 2*2 and the step is 2. The weight parameters of VGG-16 are initialized by the parameters pre-trained based on the ImageNet dataset, and the random gradient descent method is used to optimize the weight parameters. The initial value of the learning rate is 0.001, and the learning rate is divided by 10 every 30 epochs. The weight decay parameter and momentum in the training stage are 0.0005 and 0.9 respectively. The batch size of the training of the two datasets is 32 and 8 respectively.

[0088] The main calculation steps of the feature aggregation module are as follows:

[0089] The step two convolution feature extraction: for the input high-resolution remote sensing scene image, average pooling is performed on the last convolution layer conv3_3, conv4_3 and conv5_3 of the third, fourth and fifth convolution pooling groups of the VGG-16 network respectively to obtain new intermediate features with the same size, which are supplements of the original features, and are denoted as X1 R H×W×D1 , X2 R H ×W×D2 and X3 R H×W×D3 , and the obtained features are aggregated in turn to obtain a new convolution feature denoted as AF1:

[0090] AF1 = [X1; X2; X3] R H×W×(D1+D2+D3) (formula 1);

[0091] The step three first stage feature aggregation: 1x1 convolution and ReLU operation are performed on the convolution feature AF1 of step two to realize information fusion between aggregated features across channels and increase the nonlinear interaction between different channels, and the number of convolution kernels is set to N, at this time the dimension of AF1 is changed from D1+D2+D3 to N

[0092] The step four second stage feature aggregation: AF1 is aggregated with the feature X4 after the maximum pooling of the convolution layer conv5-3, denoted as X4 R H×W×D4 , L = N + D4, and the aggregated convolution feature AF2 output by the feature aggregation module is finally output as the output of the feature aggregation module:

[0093] AF2 = [AF1; X4] R H×W×L (formula 2);

[0094] Capsule module

[0095] The capsule module mainly consists of a convolution layer, a low-level capsule layer and a high-level capsule layer, and in the capsule module, the dimensions of the low-level capsule layer and the high-level capsule layer are 8 and 16 respectively; the values of the parameters m+, m- and λ in the loss function are set to 0.9, 0.1 and 0.5 respectively.

[0096] The main calculation steps of the capsule module are as follows:

[0097] The step five low-level capsule layer and high-level capsule layer construction: the value of the number S1 of neurons contained in each capsule of the low-level capsule layer is set, and a low-level capsule layer containing HxWxL / S1 capsules is constructed, wherein any capsule i corresponds to a feature vector with a length of S1 in the convolution feature AF2, that is, the output vector of the capsule i, the low-level capsule layer is used to describe smaller ground objects in the scene image and encode the attributes of the ground objects to provide the probability that the ground object belongs to a certain scene type;

[0098] The value of the dimension S2 of each capsule of the high-level capsule layer and the number T of capsules are set, and the value of T is the number of prediction categories of the scene image. A high-level capsule layer containing T capsules is constructed, wherein any one capsule j corresponds to an S2-dimensional feature vector. The high-level capsule layer is used to describe the entire scene, and the category to which the scene belongs is determined based on the encoded attributes.

[0099] For example, for a train station scene image, the low-level capsules are used to describe the platform, train, rail, building and other components of the scene, and encode the attributes of the entities, and the high-level capsules are used to describe the entire scene and encode the attributes of the scene. Through the information transmission between the low-level capsule layer and the high-level capsule layer, the capsule module can learn the relationship between the smaller ground objects and the overall scene in the scene image.

[0100] The step six input vector calculation: the low-level capsules are calculated by affine transformation through the weight matrix to obtain the prediction vector of the low-level capsules to the high-level capsules, that is, the input vector of the high-level capsules, that is, the connection between the capsule layers. The prediction vector of the low-level capsule i to the high-level capsule j is denoted as

[0101]

[0102] W ij is a weight matrix, which can be optimized by back propagation. All low-level capsules need to predict the high-level capsule j, therefore, the input vector of the high-level capsule j is is the weighted sum of all low-level capsule prediction vectors:

[0103]

[0104] wherein c ij is a coupling coefficient, which is optimized in the dynamic routing iteration process, and c ij is calculated by formula 5, and the sum of the coupling coefficients of the capsule i and all capsules in the high-level is 1:

[0105]

[0106] b ij represents the logarithmic prior probability of the coupling of the low-level capsule i and the high-level capsule j, and the initial value is set to 0.

[0107] The step seven outputs the vector acquisition: the coupling coefficients and the logarithmic prior probability are iteratively updated between the low-level capsules and the high-level capsules through the dynamic routing algorithm, and the high-level capsule input vector and the output vector are calculated, the output vector is processed by the nonlinear squashing function squash, so that its length is not more than 1, and the length of the output vector of the high-level capsule represents the prediction probability of the scene image category, therefore, in order to ensure that the length of the vector is not more than 1, the nonlinear squashing function squash is needed, so that the short vector is almost shrunk to 0, and the long vector is shrunk to slightly less than 1, and the high-level capsule j obtains the output vector v by squashing the input vector j :

[0108]

[0109] The step eight parameter update: the logarithmic probability is updated according to the inner product of the output vector and the prediction vector, the coupling coefficient is iteratively optimized, the parameter update of the capsule module is realized, and the spatial information transmission between the ground objects and the overall scene is completed.

[0110] In the capsule module, the update of the logarithmic probability b ij depends on the consistency of the vector and the vector , that is, the two vectors have a large inner product, therefore, the update formula of b ij is as follows:

[0111]

[0112] Formula 3-7 constitutes a complete dynamic routing process for calculating After the update of the logarithmic probability b ij , the coupling coefficient c ij is updated according to formula (5), after step 8, step 6 is turned to, and the iteration of the multiple dynamic routing processes is performed, the parameter update of the capsule module is realized, and the spatial information transmission between the ground objects and the overall scene is completed.

[0113] The step 9 loss function calculation: the loss function of each capsule k in the high-level capsule layer is represented as L k:

[0114] L k = T k ·max(0, m + -||v k ||) 2 +λ(1-T k )·max(0,||v k ||-m - ) 2 (Formula 8)

[0115] The length of the high-level capsule vector corresponding to each category is used as the scene prediction probability, and L k Model parameter training is performed.

[0116] When the corresponding category k exists, T k The value is 1, and during the training process, the total loss is the sum of the losses of all capsules in the layer, m+, m-, and λ are hyperparameters that need to be set according to the input scene image.

[0117] Output layer

[0118] Step ten category decision: compare the lengths of all high-level capsule output vectors, and the category corresponding to the longest vector is the predicted scene category C k ; the output layer judges the category C k of the predicted scene image according to the length of the vector (formula 9) output by the high-level capsule layer (formula 10).

[0119]

[0120] Compare the lengths of multiple output vectors, and the category corresponding to the longest vector is the predicted category C k of the scene image.

[0121] Step eleven output result: output the predicted scene image category C k , and complete the classification task of high-resolution remote sensing images.

[0122] By comparing the classification accuracy and error of the latest multiple CNN-based high-resolution scene image classification methods, the effectiveness of the method is verified. In order to realize the objectivity of evaluation, 10 experiments are repeated for each data set to reduce the influence of randomness on the experimental results.

[0123] UC Merced data set

[0124] The experiment randomly selects 80% of the images from each scene category of the UC Merced dataset as the training set, and the remaining 20% of the images as the test set. The proposed FACNNCN method is compared with 10 scene classification methods. Among the comparison methods, GoogLeNet, CaffeNet, VGG-VD-16, and Fine-tuned VGG-16 belong to the classic CNN-based methods; LGFBOVW, Fusion by addition, Two-Stream Fusion, MSCP, and FACNN belong to the feature aggregation-based methods; CNN-CapsNet and D-CapsNet belong to the methods combining CNN and vector capsule networks. The classification accuracy comparison results based on different scene classification methods are shown in the following figure:

[0125]

[0126] As shown in the above figure, the classification accuracy of the classic CNN-based scene classification method is relatively low, but compared with the four methods, the classification accuracy of Fine-tuned VGG-16 is about 2% higher than that of the other three models, which reflects the effectiveness of the fine-tuning method. FACNNCN also uses Fine-tuned VGG-16 for feature extraction in the feature fusion module, and the classification accuracy is about 2% higher than that of Fine-tuned VGG-16, which proves the effectiveness of the feature aggregation module and the capsule module in FACNNCN.

[0127] Compared with the four classic CNN-based scene classification methods, the classification accuracy of the five feature aggregation-based classification methods is relatively high. The CNN-CapsNet method combining CNN and vector capsule networks also achieves higher classification accuracy than the classic CNN-based classification method. This shows that feature aggregation and vector capsule networks are effective ways to improve the classification accuracy of scene images. Feature aggregation can obtain more discriminative convolutional features of scene images by aggregating multiple intermediate features, while vector capsule networks can obtain multiple spatial features of geographical entities of scene images through vectorized representation of capsules. The fusion of these features enriches the input information of the model classification, which is more conducive to distinguishing complex scene categories and thus improving the overall classification accuracy of scene images.

[0128] FACNNCN utilizes both aggregated features and multiple spatial features when performing remote sensing scene classification, and the input information of the model is more abundant, so the classification performance will be improved. The experimental results also verify the effectiveness of FACNNCN. Compared with the best feature aggregation-based classification method FACNN, the average precision, maximum precision and minimum precision of FACNNCN are improved by 0.44%, 0.45% and 0.43% respectively. Compared with the CNN-CapsNet method combining CNN and vector capsule network, the average precision, maximum precision and minimum precision of FACNNCN are improved by 0.44%, 0.47% and 0.41% respectively. In summary, on the UCMerced data, the average precision, maximum precision and minimum precision of the FACNNCN method proposed in this paper are 99.25%, 99.50% and 99.00% respectively, which are higher than all the comparison methods.

[0129] Figure 5 The intuitive display shows the comparison of classification accuracy error between different classification methods. It can be seen that the classification error of FACNNCN is smaller, which is obviously better than the 7 comparison methods. Compared with the remaining 3 methods Fine-tuned VGG-16, FACNN and CNN-CapsNet, the classification error is very close, which shows that the FACNNCN method not only can achieve higher classification accuracy, but also has good stability.

[0130] Figure 6 The confusion matrix shows the best classification results of FACNNCN on the UCMerced dataset. As can be seen from the figure, among the 21 scene categories, the classification accuracy of 19 scene categories is 100%, and the prediction accuracy of the remaining 2 scene categories, building and medium-density residential area, is 95%. The model predicts 5% of the scene images classified as buildings to be dense residential areas, and 5% of the scene images classified as medium-density residential areas to be dense residential areas. The reason for the prediction error is that the constituting entities of the three scene images of building, medium-density residential area and dense residential area are mainly houses, and the optical images and spatial structure features are very similar, which easily causes confusion of the prediction results.

[0131] NWPU dataset

[0132] The data size of the UC Merced dataset is relatively small, and the number of scene categories is also relatively small. In order to more objectively evaluate the performance of the method, comparative experiments are carried out on the NWPU dataset which has more scene categories and larger size. The experiments are divided into two groups: (1) randomly select 10% of the scene images of each scene category as the training set, and the rest of the images as the test set; (2) randomly select 20% of the scene images of each scene category as the training set, and the rest of the images as the test set. On the NWPU dataset, 10 scene classification methods are also selected for comparative experiments. The comparison results of the scene classification accuracy of different methods are shown in the following figure:

[0133]

[0134] As can be seen from the above figure, due to the significant increase in the size of the NWPU dataset and the number of scene categories, the classification accuracy of all methods is significantly lower than that of the UC Merced dataset. However, compared with the 10 comparison methods, the average accuracy, maximum accuracy and minimum accuracy of FACNNCN are still the highest under the training ratio of 10% and 20%, indicating that FACNNCN can still achieve better classification performance on large-scale datasets.

[0135] Figure 7 The difference between the classification errors of different methods under the training ratios of 10% and 20% is shown in the following figure. It can be seen that when the training ratio is 10%, the classification errors of each method are relatively close, and when the training ratio increases from 10% to 20%, the difference in classification error is smaller. Under the training ratio of 20%, the classification error of FACNNCN is lower than that of all comparison methods, indicating that the method in this paper can still maintain good stability on large-scale and multi-class remote sensing image scene data.

[0136] Figure 8 and Figure 9 The confusion matrices of the classification results of FACNNCN under the training ratios of 10% and 20% are given. Figure 8 The prediction accuracy of 29 scene categories in the first confusion matrix is more than 90%, Figure 9 The prediction accuracy of 34 scene categories in the second confusion matrix is more than 90%. In the two confusion matrices, the two scene categories with the worst prediction accuracy and the most confusion are palaces and churches, mainly because their architectural styles are very similar.

[0137] To verify the performance of the feature aggregation module and the capsule module in the FACNNCN method, ablation experiments were conducted on the UC Merced and NWPU datasets. For the UC Merced dataset, when the FACNNCN only retains the feature aggregation module or the capsule module, the average classification accuracy of the FACNNCN is 98.81%. When the feature aggregation module and the capsule module are retained simultaneously, the average classification accuracy of the FACNNCN increases to 99.25%. For the NWPU dataset, the same ablation experiments were conducted. Under the training proportions of 10% and 20%, when the feature aggregation module is retained, the average classification accuracy of the FACNNCN is 86.41% and 90.68% respectively. When the capsule module is retained, the average classification accuracy of the FACNNCN is 85.08% and 89.18% respectively. When the feature aggregation module and the capsule module are retained simultaneously, the average classification accuracy of the FACNNCN increases to 87.34% and 91.12% respectively.

[0138] It can be found from the ablation experiments that the feature aggregation module and the capsule module in the FACNNCN method proposed in this paper both exhibit good performance in the scene classification process. On the NWPU dataset with large-scale scene images, the feature aggregation module plays a better role than the vector capsule module, and the organic combination of the two can obtain a classification performance better than that of any module.

[0139] The above specific embodiments are only several preferred embodiments of the present application, and based on the technical solutions of the present application and the related inspiration of the above embodiments, those skilled in the art can make various alternative improvements and combinations on the above specific embodiments.

[0140] It is apparent for those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, but can be implemented in other concrete forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present application is defined by the appended claims rather than the above description, and all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application.

Claims

1. A high-resolution remote sensing image scene classification method based on FACNNCN, characterized in that: The following steps are involved: Step 1: Input preprocessing: obtain high-resolution remote sensing scene images with three RGB color channels, crop them to a uniform size, and normalize them to form a structured input sample set; Step 2: Convolutional feature extraction: For the input high-resolution remote sensing scene image, the VGG-16 network is used to perform average pooling from the last convolutional layers conv3_3, conv4_3, and conv5_3 of the 3rd, 4th, and 5th convolutional pooling groups respectively to obtain new intermediate features of the same size. These features are complementary to the original features. The obtained features are aggregated in turn to obtain new convolutional features recorded as AF1; Step 3: First stage feature aggregation: Perform 1*1 convolution and ReLU operations on the convolution feature AF1 in step 2 to achieve information fusion between aggregated features across channels and increase nonlinear interaction between different channels; Step 4: Second stage feature aggregation: Aggregate AF1 with the feature X4 after the maximum pooling of the convolution layer conv5-3. The aggregated convolution feature AF2 finally output by the feature aggregation module is used as the output of the feature aggregation module. Step 5: Construct the low-level capsule layer and the high-level capsule layer: Set the value of the number of neurons S1 contained in each capsule of the low-level capsule layer, and construct a low-level capsule layer containing H×W×L / S1 capsules, where any capsule i corresponds to a feature vector of length S1 in the convolution feature AF2, that is, the output vector of capsule i. The low-level capsule layer is used to describe smaller objects in the scene image and encode the attributes of the objects, providing the probability that the objects belong to a certain scene type; Set the dimension S2 of each capsule in the high-level capsule layer and the number of capsules T, where T is the number of predicted categories of the scene image. Construct a high-level capsule layer containing T capsules, where any capsule j corresponds to an S2-dimensional feature vector. The high-level capsule layer is used to describe the entire scene and determine the category of the scene based on the encoded attributes. Step 6 Input vector calculation: Perform affine transformation on the low-level capsule through the weight matrix to obtain the prediction vector of the low-level capsule for the high-level capsule, that is, the input vector of the high-level capsule; Step 7: Output vector acquisition: The coupling coefficient and the logarithmic prior probability are iteratively updated between the low-level capsule and the high-level capsule through the dynamic routing algorithm, and the input vector and output vector of the high-level capsule are calculated. The output vector is processed by the nonlinear squash function so that its length does not exceed 1. Step 8: Parameter update: Update the logarithmic probability based on the inner product of the output vector and the predicted vector, iteratively optimize the coupling coefficient, update the capsule module parameters, and complete the transfer of spatial information between the ground object and the overall scene; Step 9 Loss function calculation: The loss function of each capsule k in the high-level capsule layer is expressed as L k , the length of the high-level capsule vector corresponding to each category is the scene prediction probability, through L k Perform model parameter training; Step 10 Category decision: Compare the modulus lengths of all high-level capsule output vectors. The category corresponding to the one with the largest modulus length is the predicted scene category C. k ; Step 11 Output result: Output the predicted scene image category C k , complete the classification task of high-resolution remote sensing images.

2. The high-resolution remote sensing image scene classification method based on FACNNCN according to claim 1 is characterized by: The parameters of the VGG-16 model for extracting convolutional features are pre-trained on a large-scale general image dataset before use to enhance its preliminary feature extraction capability for natural scenes. Subsequently, the model is fine-tuned using the stochastic gradient descent method on a high-resolution remote sensing image scene dataset to adapt to the structural features of objects and complex backgrounds in remote sensing scenes, thereby improving the model's adaptability to remote sensing data and classification accuracy.

3. The high-resolution remote sensing image scene classification method based on FACNNCN according to claim 1 is characterized by: In the first-stage feature aggregation step, convolution and nonlinear activation operations are applied to multiple intermediate convolutional feature maps to achieve ×-deep fusion of inter-channel features. The number of convolution kernels used is set in the medium to high dimensional range according to the complexity of the feature map, and the nonlinear activation function uses the rectified linear unit function to enhance the nonlinear ability and discrimination ability of the aggregated feature expression.

4. The high-resolution remote sensing image scene classification method based on FACNNCN according to claim 1 is characterized by: During the second-stage feature enhancement process, the final aggregated features are obtained by splicing the first-stage aggregated features with the main feature map after pooling of the fifth convolution group, and further unifying the dimensions and enhancing the semantic expression ability through a layer of channel compression convolution, thereby obtaining the final aggregated features with richer contextual information and enhancing the recognition ability of complex scenes.

5. The high-resolution remote sensing image scene classification method based on FACNNCN according to claim 1 is characterized in that: The low-level capsule layer is composed of multiple capsule units with fixed dimensions. Each capsule unit is used to express multiple attribute information of a local feature entity in the remote sensing image, including but not limited to position, direction, scale, and texture. By dividing multiple local areas in a sliding window manner on the aggregated feature map and encoding the features in each area, a low-level capsule set is formed to support the subsequent modeling of spatial relationships.

6. The high-resolution remote sensing image scene classification method based on FACNNCN according to claim 1 is characterized by: The number of capsules in the high-level capsule layer is equal to the total number of categories involved in scene classification. Each capsule is used to represent the global semantic information of a specific scene category, and its dimension is higher than that of the low-level capsule, so as to carry richer comprehensive attributes. By establishing a mapping relationship with the mapping weight matrix between the low-level capsules, the abstract expression of multi-level spatial semantics is realized.

7. The high-resolution remote sensing image scene classification method based on FACNNCN according to claim 1 is characterized by: The dynamic routing process adopts an iterative optimization-based mechanism to establish coupling weights between low-level capsules and high-level capsules. The initial weights are set according to a preset logarithmic prior, and the weight distribution is adjusted according to the current output results in each iteration, so that low-level capsules with highly similar information contribute more weight to specific high-level capsules, thereby gradually optimizing the information flow and expression path. The iterative process is usually performed for three to five rounds to ensure convergence and stability.

8. The high-resolution remote sensing image scene classification method based on FACNNCN according to claim 1 is characterized by: The loss function of the capsule module is constructed based on the bilateral boundary strategy controlled by category labels. By respectively strengthening the output length of capsules of the correct category and suppressing the output of capsules of the wrong category, the network's stable convergence and improved discrimination ability are achieved during training. The relevant hyperparameters are set according to the distribution characteristics of the input data to adapt to remote sensing images of different resolutions and complexities.

Citation Information

Cited By

  • Dynamic sparse feature enhancement-based extremely few sample remote sensing scene classification method

    CN121962700A