Remote sensing image classification method fusing KAN architecture and convolutional neural network
By integrating the KAN architecture and convolutional neural network, a remote sensing image classification model KCN with a fused fully convolution mask autoencoder and a global response normalization layer is built, which solves the problems of large parameters, low training efficiency and overfitting of deep learning models in high-resolution remote sensing image processing, and realizes efficient and high-precision remote sensing image classification.
Patent Information
- Application Number
- CN202510020186.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-16
AI Technical Summary
When existing deep learning models process high-resolution remote sensing images, the number of parameters is huge, resulting in high processing costs and memory requirements, low training efficiency, and possible overfitting problems.
Using a method of fusion KAN architecture and convolutional neural network, a convolutional neural network is constructed by constructing a convolutional neural network that integrates a fully convolutional mask autoencoder and a global response normalization layer, and integrating the KAN architecture with it, a remote sensing image classification model KCN is constructed.
It realizes higher efficiency and higher accuracy remote sensing image classification, has self-supervised learning ability and perceived features of different scales, reducing the amount of parameters and calculations, and avoiding overfitting.
Smart Images

Figure CN120014440A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image classification, and in particular to a remote sensing image classification method integrating a KAN architecture and a convolutional neural network. Background Art
[0002] The rapid development of satellite technology has led to a rapid increase in high-resolution remote sensing image datasets, which have been widely analyzed by deep learning models and have been widely used in land cover classification, object detection, and environmental monitoring. However, there are still certain challenges in applying deep learning to high-resolution datasets, mainly the huge number of parameters required for deep neural networks, especially for high-resolution images, which leads to higher processing costs and memory requirements, making the training process inefficient and may not be feasible with standard hardware resources.
[0003] In addition, over-parameterization leads to longer training times and potential overfitting, where the model memorizes the training data rather than generalizing from it, limiting its effectiveness in practical applications. To address these issues, one way is to adopt more effective training methods, such as model pruning, quantization, and using more compact network designs, but each method has its trade-offs and disadvantages. Another way is to build a classification model with a new network structure to provide a more effective learning framework for remote sensing image classification, in order to achieve efficient and accurate remote sensing image classification accuracy. Summary of the invention
[0004] In view of the shortcomings of the prior art, an embodiment of the present invention aims to provide a remote sensing image classification method that integrates a KAN architecture and a convolutional neural network to solve the problems in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions: A remote sensing image classification method integrating KAN architecture and convolutional neural network includes the following steps: Step S1: Collect remote sensing image datasets from different countries and divide them into training set, validation set and test set; Step S2: construct a convolutional neural network integrating a fully convolutional masked autoencoder and a global response normalization layer; Step S3: Integrate the KAN architecture with the convolutional neural network constructed in step S2 to construct a remote sensing image classification model KCN that integrates the KAN architecture and the convolutional neural network; Step S4: Use the training set and validation set in step S1 to train and validate the KCN model constructed, and evaluate the classification accuracy of the KCN model on the test set.
[0006] As a further solution of the present invention, the remote sensing image dataset from different countries in step S1 includes 10 different categories, and each category contains 2000 to 3000 images.
[0007] As a further solution of the present invention, the image acquisition source is the Sentinel-2A satellite, the size of each image is 64×64 pixels, and the spatial resolution of each pixel is 10 meters.
[0008] As a further solution of the present invention, the 27,000 remote sensing images collected in step S1 are divided into three parts: a training set, a validation set and a test set, wherein the training set contains 18,900 remote sensing images, the validation set contains 4,050 remote sensing images, and the test set contains 4,050 remote sensing images.
[0009] As a further solution of the present invention, when constructing a convolutional neural network integrating a fully convolutional masked autoencoder and a global response normalization layer in step S2, the fully convolutional masked autoencoder is based on a self-supervised learning method of a convolutional neural network, which increases the model's perception ability for features of different scales and processes remote sensing image data with masks.
[0010] As a further solution of the present invention, when constructing a convolutional neural network integrating a fully convolutional masked autoencoder and a global response normalization layer in step S2, the feature map is normalized on each channel to enhance feature competition between channels. The global response normalization layer calculation mainly includes three steps: (1) Global feature aggregation of the input image: For a given input image, it can be represented as a feature map , where H is the feature map height, W is the feature map height, and C is the number of feature map channels. First, pass the global function The feature information of the i-th channel Aggregate to get the aggregate information of the i-th channel , the feature vector obtained by aggregating the information of all C channels , It can be expressed as formula 1: Formula 1: , Based on formula 1, the dimension can be The feature map is transformed into a dimension of The eigenvector of ; (2) Aggregation information for the i-th channel Normalization is performed, and the normalization process can be expressed as formula 2: Formula 2: , Represents the normalized score after normalizing the aggregate information of the i-th channel; (3) Finally, the calculated feature normalization score is used to calculate the feature information of the i-th channel After calibration and updating, the characteristic information of the final i-th channel is obtained , the process can be expressed as formula 3: Formula 3: , After integrating the global response normalization layer into the convolutional neural network, the integrated convolutional neural network can be pre-trained based on the fully convolutional masked autoencoder.
[0011] As a further solution of the present invention, when the KAN architecture is integrated with the convolutional neural network constructed in step S2, the KAN architecture used has a learnable activation function at the edge, and the activation function is parameterized by a B-spline function, wherein the B-spline function is a piecewise polynomial function defined by control points and nodes, and for each input feature , based on the B-spline function After parameterization, the aggregate value of each intermediate variable q is obtained, and the function Transform and finally get the sum of the transformed values , which can be expressed as formula 4: Formula 4: .
[0012] As a further solution of the present invention, when the KAN architecture is integrated with the convolutional neural network constructed in step S2, the specific integration method is to replace the MLP classifier of the convolutional neural network integrating the full convolutional masked autoencoder and the global response normalization layer with the KANLinear layer based on two KAN architectures, thereby constructing a remote sensing image classification model KCN integrating the KAN architecture, the full convolutional masked autoencoder and the convolutional neural network with the global response normalization layer; The KCN model replaces the MLP classifier of the convolutional neural network with the KANLinear layer, and then uses the spline function of the KANLinear layer to replace the linear weights of the MLP classifier of the convolutional neural network.
[0013] As a further solution of the present invention, the KCN model constructed in step S4 is trained and verified using the divided training set and verification set, and the classification accuracy of the KCN model is evaluated on the test set. The classification accuracy evaluation index is selected as Accuracy, and its calculation formula can be expressed as Formula 5: Formula 5: , In Formula 5, TP represents the number of samples that are actually positive examples and predicted as positive examples by the model; TN represents the number of samples that are actually negative examples and predicted as negative examples by the model; FP represents the number of samples that are actually negative examples but are incorrectly predicted as positive examples by the model; and FN represents the number of samples that are actually positive examples but are incorrectly predicted as negative examples by the model.
[0014] In summary, compared with the prior art, the embodiments of the present invention have the following beneficial effects: At the same time, it has the self-supervised learning ability of the fully convolutional masked autoencoder, the improvement of the potential feature collapse problem by the global response normalization layer, the ability to perceive features of different scales and the ability to adapt to input data more efficiently, and can complete remote sensing image classification tasks with higher efficiency and higher accuracy.
[0015] In order to more clearly illustrate the structural features and effects of the present invention, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flow chart of an embodiment of the invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0018] The specific implementation of the present invention is described in detail below in conjunction with specific embodiments.
[0019] In one embodiment, a remote sensing image classification method integrating KAN architecture and convolutional neural network is described in detail. Figure 1 , including the following steps: Step S1: Collect remote sensing image datasets from different countries and divide them into training set, validation set and test set; Step S2: construct a convolutional neural network integrating a fully convolutional masked autoencoder and a global response normalization layer; Step S3: Integrate the KAN architecture with the convolutional neural network constructed in step S2 to construct a remote sensing image classification model KCN that integrates the KAN architecture and the convolutional neural network; Step S4: Use the training set and validation set in step S1 to train and validate the KCN model constructed, and evaluate the classification accuracy of the KCN model on the test set.
[0020] For further information, see Figure 1In step S1, the remote sensing image dataset from different countries includes 10 different categories, each category contains 2000 to 3000 images.
[0021] For further information, see Figure 1 The image is obtained from the Sentinel-2A satellite. The size of each image is 64×64 pixels, and the spatial resolution of each pixel is 10 meters.
[0022] For further information, see Figure 1 The 27,000 remote sensing images collected in step S1 are divided into three parts: a training set, a validation set, and a test set, wherein the training set contains 18,900 remote sensing images, the validation set contains 4,050 remote sensing images, and the test set contains 4,050 remote sensing images.
[0023] For further information, see Figure 1 When constructing a convolutional neural network integrating a fully convolutional masked autoencoder and a global response normalization layer in step S2, the fully convolutional masked autoencoder is based on a self-supervised learning method of a convolutional neural network, which increases the model's perception ability for features of different scales and processes remote sensing image data with masks.
[0024] For further information, see Figure 1 In step S2, when constructing a convolutional neural network integrating a fully convolutional masked autoencoder and a global response normalization layer, the feature map is normalized on each channel to enhance the feature competition between channels. The calculation of the global response normalization layer mainly includes three steps: (1) Global feature aggregation of the input image: For a given input image, it can be represented as a feature map , where H is the feature map height, W is the feature map height, and C is the number of feature map channels. First, pass the global function The feature information of the i-th channel Aggregate to get the aggregate information of the i-th channel , the feature vector obtained by aggregating the information of all C channels , It can be expressed as formula 1: Formula 1: , Based on formula 1, the dimension can be The feature map is transformed into a dimension of The eigenvector of ; (2) Aggregation information for the i-th channel Normalization is performed, and the normalization process can be expressed as formula 2: Formula 2: , Represents the normalized score after normalizing the aggregate information of the i-th channel; (3) Finally, the calculated feature normalization score is used to calculate the feature information of the i-th channel After calibration and updating, the characteristic information of the final i-th channel is obtained , the process can be expressed as formula 3: Formula 3: , After integrating the global response normalization layer into the convolutional neural network, the integrated convolutional neural network can be pre-trained based on the fully convolutional masked autoencoder.
[0025] For further information, see Figure 1 When the KAN architecture is integrated with the convolutional neural network constructed in step S2, the KAN architecture used has a learnable activation function at the edge, and the activation function is parameterized by a B-spline function, wherein the B-spline function is a piecewise polynomial function defined by control points and nodes, and for each input feature , based on the B-spline function After parameterization, the aggregate value of each intermediate variable q is obtained, and the function Transform and finally get the sum of the transformed values , which can be expressed as formula 4: Formula 4: .
[0026] For further information, see Figure 1 When the KAN architecture is integrated with the convolutional neural network constructed in step S2, the specific integration method is to replace the MLP classifier of the convolutional neural network integrating the full convolutional masked autoencoder and the global response normalization layer with the KANLinear layer based on two KAN architectures, thereby constructing a remote sensing image classification model KCN integrating the KAN architecture, the full convolutional masked autoencoder and the convolutional neural network with the global response normalization layer; The KCN model replaces the MLP classifier of the convolutional neural network with the KANLinear layer, and then uses the spline function of the KANLinear layer to replace the linear weights of the MLP classifier of the convolutional neural network.
[0027] For further information, see Figure 1 In step S4, the divided training set and validation set are used to train and validate the constructed KCN model, and the classification accuracy of the KCN model is evaluated on the test set. The classification accuracy evaluation index is selected as Accuracy, and its calculation formula can be expressed as Formula 5: Formula 5: , In Formula 5, TP represents the number of samples that are actually positive examples and predicted as positive examples by the model; TN represents the number of samples that are actually negative examples and predicted as negative examples by the model; FP represents the number of samples that are actually negative examples but are incorrectly predicted as positive examples by the model; and FN represents the number of samples that are actually positive examples but are incorrectly predicted as negative examples by the model.
[0028] In this embodiment, the remote sensing image dataset from different countries in step S1 mainly includes 27,000 remote sensing images collected from different cities in 34 European countries, covering 10 different categories, including annual crops, forests, herbaceous vegetation, roads, industries, pastures, perennial crops, residential areas, rivers, and seas and lakes. Each category contains 2,000 to 3,000 images, and the image acquisition source is the Sentinel-2A satellite, covering complex land cover scenes with high diversity. The size of each image is 64×64 pixels, and the spatial resolution of each pixel is 10 meters;
[0029] The 27,000 remote sensing images collected in step S1 are divided into three parts: a training set, a validation set, and a test set, wherein the training set contains 18,900 remote sensing images, the validation set contains 4,050 remote sensing images, and the test set contains 4,050 remote sensing images; When constructing a convolutional neural network integrating a fully convolutional masked autoencoder and a global response normalization layer in step S2, the fully convolutional masked autoencoder is a self-supervised learning method based on a convolutional neural network. On the one hand, it uses a fully convolutional method to process remote sensing images instead of using a fully connected layer to generate masks and reconstruct images, which significantly reduces the number of parameters and calculations while maintaining spatial information; on the other hand, this method uses a multi-scale mask strategy instead of using a fixed-size mask, which effectively increases the model's perception of features of different scales, and is particularly suitable for processing remote sensing image data with masks; When constructing a convolutional neural network that integrates a fully convolutional masked autoencoder and a global response normalization layer in step S2, the global response normalization layer is a new convolutional neural network layer that can normalize feature maps on each channel, thereby enhancing feature competition between channels. On the one hand, the global response normalization layer only normalizes feature maps and does not require additional parameters; on the other hand, the global response normalization layer can process batches of any size without dynamically adjusting parameters, and the amount of calculation is small;
[0030] When integrating the KAN architecture with the convolutional neural network constructed in step S2, the KAN architecture adopted is mainly inspired by the Kolmogorov-Arnold representation theorem, and its edges have learnable activation functions that are parameterized by B-spline functions, where B-spline functions are piecewise polynomial functions defined by control points and nodes. For each input feature , based on the B-spline function After parameterization, the aggregate value of each intermediate variable q is obtained, and the function Transform and finally get the sum of the transformed values ,This process can help the KAN architecture to capture complex data patterns more ,flexibly and effectively;
[0031] When integrating the KAN architecture with the convolutional neural network constructed in step S2, the specific integration method is to use the KANLinear layer based on two KAN architectures to replace the MLP classifier of the convolutional neural network that integrates the full convolutional masked autoencoder and the global response normalization layer, thereby constructing a remote sensing image classification model KCN that integrates the KAN architecture, the full convolutional masked autoencoder, and the convolutional neural network with the global response normalization layer. After the KCN model replaces the MLP classifier with the KANLinear layer, the spline function of the KANLinear layer can be used to replace the linear weights of the MLP classifier, so that the model can adapt to the input data more effectively, and thus can complete the remote sensing image classification task with higher efficiency and higher accuracy;
[0032] Based on the KCN method proposed in the present invention, when its KANLinear layer is set to 32 nodes and the number of Epoch iterations is set to 3, its accuracy on the test set of the remote sensing image dataset in step S1 is 96%.
[0033] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A remote sensing image classification method integrating KAN architecture and convolutional neural network, characterized in that: The steps include: Step S1: Collect remote sensing image datasets from different countries and divide them into training set, validation set and test set; Step S2: construct a convolutional neural network integrating a fully convolutional masked autoencoder and a global response normalization layer; Step S3: Integrate the KAN architecture with the convolutional neural network constructed in step S2 to construct a remote sensing image classification model KCN that integrates the KAN architecture and the convolutional neural network; Step S4: Use the training set and validation set in step S1 to train and validate the KCN model constructed, and evaluate the classification accuracy of the KCN model on the test set.
2. The remote sensing image classification method integrating the KAN architecture and the convolutional neural network according to claim 1, characterized in that: The remote sensing image dataset from different countries in step S1 includes 10 different categories, each category contains 2000 to 3000 images.
3. The remote sensing image classification method integrating the KAN architecture and the convolutional neural network according to claim 2 is characterized in that: The images are acquired from the Sentinel-2A satellite. The size of each image is 64×64 pixels, and the spatial resolution of each pixel is 10 meters.
4. The remote sensing image classification method integrating the KAN architecture and the convolutional neural network according to claim 3 is characterized in that: The 27,000 remote sensing images collected in step S1 are divided into three parts: a training set, a validation set and a test set, wherein the training set contains 18,900 remote sensing images, the validation set contains 4,050 remote sensing images, and the test set contains 4,050 remote sensing images.
5. The remote sensing image classification method integrating KAN architecture and convolutional neural network according to claim 1, characterized in that: When constructing a convolutional neural network integrating a fully convolutional masked autoencoder and a global response normalization layer in step S2, the fully convolutional masked autoencoder is based on a self-supervised learning method of a convolutional neural network, which increases the model's perception ability for features of different scales and processes remote sensing image data with masks.
6. The remote sensing image classification method integrating KAN architecture and convolutional neural network according to claim 5, characterized in that: When constructing the convolutional neural network integrating the full convolutional masked autoencoder and the global response normalization layer in step S2, the feature map is normalized on each channel to enhance the feature competition between channels. The calculation of the global response normalization layer mainly includes three steps: (1) Global feature aggregation of the input image: For a given input image, it can be represented as a feature map , where H is the feature map height, W is the feature map height, and C is the number of feature map channels. First, pass the global function The feature information of the i-th channel Aggregate to get the aggregate information of the i-th channel , the feature vector obtained by aggregating the information of all C channels , It can be expressed as formula 1: Formula 1: , Based on formula 1, the dimension can be The feature map is transformed into a dimension of The eigenvector of ; (2) Aggregation information for the i-th channel Normalization is performed, and the normalization process can be expressed as formula 2: Formula 2: , Represents the normalized score after normalizing the aggregate information of the i-th channel; (3) Finally, the calculated feature normalization score is used to calculate the feature information of the i-th channel After calibration and updating, the characteristic information of the final i-th channel is obtained , the process can be expressed as formula 3: Formula 3: , After integrating the global response normalization layer into the convolutional neural network, the integrated convolutional neural network can be pre-trained based on the fully convolutional masked autoencoder.
7. The remote sensing image classification method integrating KAN architecture and convolutional neural network according to claim 6, characterized in that: When the KAN architecture is integrated with the convolutional neural network constructed in step S2, the KAN architecture used has a learnable activation function at the edge, and the activation function is parameterized by a B-spline function, wherein the B-spline function is a piecewise polynomial function defined by control points and nodes, and for each input feature , based on the B-spline function After parameterization, the aggregate value of each intermediate variable q is obtained, and the function Transform and finally get the sum of the transformed values , which can be expressed as formula 4: Formula 4: .
8. The remote sensing image classification method integrating KAN architecture and convolutional neural network according to claim 7, characterized in that: When the KAN architecture is integrated with the convolutional neural network constructed in step S2, the specific integration method is to replace the MLP classifier of the convolutional neural network integrating the full convolutional masked autoencoder and the global response normalization layer with the KANLinear layer based on the two KAN architectures, thereby constructing a remote sensing image classification model KCN integrating the KAN architecture, the full convolutional masked autoencoder and the convolutional neural network with the global response normalization layer; The KCN model replaces the MLP classifier of the convolutional neural network with the KANLinear layer, and then uses the spline function of the KANLinear layer to replace the linear weights of the MLP classifier of the convolutional neural network.
9. The remote sensing image classification method integrating KAN architecture and convolutional neural network according to claim 1, characterized in that: In step S4, the constructed KCN model is trained and verified using the divided training set and verification set, and the classification accuracy of the KCN model is evaluated on the test set. The classification accuracy evaluation index is selected as Accuracy, and its calculation formula can be expressed as Formula 5: Formula 5: , In Formula 5, TP represents the number of samples that are actually positive examples and predicted as positive examples by the model; TN represents the number of samples that are actually negative examples and predicted as negative examples by the model; FP represents the number of samples that are actually negative examples but are incorrectly predicted as positive examples by the model; and FN represents the number of samples that are actually positive examples but are incorrectly predicted as negative examples by the model.