Multi-spectral image classification method based on LO-MLPRNN
Through the LO-MLPRNN network framework, combined with ODC, LSK, GRU and MLP networks, the problem of insufficient information utilization in traditional algorithms in multispectral image classification is solved, and efficient and accurate multispectral remote sensing image classification is achieved.
Patent Information
- Application Number
- CN202510072589.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional deep learning algorithms cannot fully utilize the context information in multispectral remote sensing images. Recurrent neural networks are complex in the multispectral image classification and are difficult to process in parallel, resulting in inefficient classification.
The LO-MLPRNN network framework is adopted, combined with ODC network, LSK network, GRU network and MLP network, and efficient classification of multispectral images is achieved through parallel feature fusion and full connection layer mapping.
It improves the accuracy and stability of multispectral image classification, can efficiently process multispectral remote sensing images, and enhances feature extraction capabilities and model practicality.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a multi-spectral image classification method based on LO-MLPRNN. Background Art
[0002] Spectral data information can be used for ground object recognition and classification. The ground object recognition and classification based on remote sensing multi-spectral images rely on image classification. With the continuous development of space technology and sensor technology, it has become increasingly easy for researchers to obtain a large number of high-quality multi-spectral remote sensing images, which present the state of the ecological environment and the traces of human activities. Efficiently learning and extracting the information contained therein has become the core task of intelligent interpretation of remote sensing images. Semantic segmentation, as an effective coping strategy, has attracted much attention, and its main purpose is to determine the semantic categories of each pixel in the image. Remote sensing image semantic segmentation has been widely applied in many practical scenarios such as urban planning, disaster assessment, and agricultural production. In the field of multi-spectral image classification, traditional deep learning algorithms cannot fully utilize the context information in multi-spectral remote sensing images. Although the recurrent neural network (RNN) has achieved certain results, it also has deficiencies. Although RNN is good at processing time-series data, its calculation is complex and it is difficult to process in parallel, which to a certain extent restricts its application in multi-spectral image classification. To solve the above problems, the present invention proposes a network framework LO-MLPRNN mainly composed of a multi-layer perceptron and a recurrent neural network, which integrates a large-scale selective convolution network (LSK) and an omnidirectional dynamic convolution (ODC) for pixel classification of multi-spectral images. In the LO-MLPRNN network framework, the band information of multi-spectral images is processed by the parallel fusion of ODC and LSK modules. After the generated parallel features are fused, the features extracted by GRU are further mapped to a high-dimensional feature space through a fully connected layer, and then the non-linear characteristics are strengthened through an activation function, and finally fine classification of multi-spectral images is achieved through a classification layer.
[0003] In summary, the ground object recognition and classification technology based on multi-spectral data has higher accuracy and reliability, and at the same time provides more useful surface information, which makes it an important tool for modern vegetation management and environmental monitoring, and it is necessary to conduct further research. Summary of the Invention
[0004] The present invention aims to provide a multi-spectral image classification method based on LO-MLPRNN, which has the capabilities of efficient parallel processing, multi-level feature fusion, noise robustness, sample imbalance processing, etc., and can perform high-performance classification, and can better achieve high-performance pixel-level multi-spectral remote sensing image classification.
[0005] The technical solution of the present invention is as follows:
[0006] The described multi-spectral image classification method based on LO-MLPRNN includes the following steps:
[0007] A. Construct a deep neural network model, which includes: ODC network, LSK network, GRU network, and MLP network;
[0008] B. Train the deep neural network model to obtain a trained deep neural network model;
[0009] C. Input the original image into the ODC network and the LSK network for processing respectively. After the processing results of the ODC network and the LSK network are added and fused, input them into the GRU network for processing, and the obtained results are input into the MLP network for processing to obtain the final output result.
[0010] The process of training the deep neural network model in step B is as follows:
[0011] a. Preprocess the multi-spectral images from the GF-2 dataset and divide them into a training set and a validation set;
[0012] b. Use the data in the training set to train the LO-MLPRNN neural network model to obtain a trained and improved neural network model;
[0013] c. Use the data in the validation set to test the LO-MLPRNN neural network model to obtain a trained improved LO-MLPRNN neural network model.
[0014] In step a, the process of preprocessing the multi-spectral remote sensing data is as follows:
[0015] According to the high-resolution GF-2 dataset, select ROI label data in the remote sensing image, and randomly divide the image data into a training set and a validation set. The ratio of the number of image data in the training set to the validation set is 5 - 7:3 - 5.
[0016] In the ODC network, processing is performed through dynamic convolution to enhance the extraction of complex features of multi-spectral images. The dynamic convolution formula is as follows:
[0017] y = (α w1 ⊙α f1 ⊙α c1 ⊙α s1 ⊙W1 +... + α wn ⊙α fn ⊙α cn ⊙α sn ⊙W n ) * x(1)
[0018] In the formula, α wnDenotes the kernel intelligent multiplication operation along the kernel dimension of the convolutional kernel, and assigns an attention scalar, α, to the entire convolutional kernel fn Denotes the channel intelligent multiplication operation along the output channel dimension, and assigns different attention scalars, α, to the convolution cn Denotes the filter intelligent multiplication operation along the input channel dimension, and assigns different attention scalars, α, to each convolution sn Denotes the position intelligent multiplication operation along the spatial dimension, and assigns different attention scalars to the convolution parameters at the k×k spatial positions
[0019] The processing in the LSK network is as follows:
[0020] The input result is successively processed by the first large convolution and the second large convolution. After the results of the first large convolution and the second large convolution are concatenated by the channel concatenation module; the obtained concatenated result is divided into two paths. The first path is successively processed by average pooling, max pooling, the spatial feature description module, convolution, and the Sigmoid function to obtain the function processing result; the obtained function processing result is divided into two branches. The first branch is multiplied and fused with the result of the first large convolution to obtain the result of the first branch, and the second branch is multiplied and fused with the result of the second large convolution to obtain the result of the second branch; after the result of the first branch and the result of the second branch are added and fused, the result of the first path is obtained; the spatial feature description module is the Figure 3 SA module in , full name: spatial feature descriptors spatial feature description module;
[0021] The second path is processed by the multi-head attention module to obtain the result of the second path;
[0022] After the result of the first path and the result of the second path are added and fused, they are multiplied and fused with the input result to obtain the output result
[0023] In the MLP network described, the input layer is set with 64 - 256 nodes, and the hidden layer is set with 128 - 512 nodes. The number of hidden layer nodes is twice that of the input layer
[0024] The beneficial effects of the present invention are as follows:
[0025] In the method of the present invention, the ODC module learns complementary attention along the four dimensions of the convolutional kernel space, effectively captures the spectral and spatial feature correlations of the hyperspectral image, enhances the feature extraction ability, improves the pre-training performance in hyperspectral image classification, and enhances the practicability and reliability of the model
[0026] In the method of the present invention, the LSK network is improved. The improved LSK network can adaptively adjust the feature extraction strategy according to the uniqueness of different ground objects in terms of spectrum and space, greatly improving the accuracy and stability of classification. At the same time, through the multi-head attention mechanism, it can more precisely focus on the key information in different spectral bands and spatial positions, enabling the model to fully capture the complex correlations between the rich spectral and spatial features of the multi-spectral image. Performing parallel fusion processing on the improved LSK network and the ODC network can effectively mine the global correlation information of the image and the feature information between different channels.
[0027] The present invention combines the advantages of the MLP network and the RNN network. The powerful non-linear mapping of the MLP network is combined with the ability of the RNN network to process sequence context and long-term dependencies, enabling more comprehensive extraction of data features and better achieving high-performance pixel-level multi-spectral image classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is a schematic structural diagram of the network model Figure 1 ;
[0029] Figure 2 is a schematic structural diagram of the network model Figure 2 ;
[0030] Figure 3 is a schematic structural diagram of the network model Figure 3 ;
[0031] Figure 4 is a schematic structural diagram of the network model Figure 4 ;
[0032] Figure 5 is a schematic structural diagram of the deep neural network model in Embodiment 1;
[0033] Figure 6 is a schematic structural diagram of the ODC network in Embodiment 1; Figure 7 is a processing flow chart of the LSK network in Embodiment 1; Figure 8 is a schematic structural diagram of the GRU network in Embodiment 1; Figure 9 is a schematic structural diagram of the MLP network in Embodiment 1; Figure 10 is a comparison chart of classification test results of different algorithms for the GF-2 dataset in Embodiment 2. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The present invention will be specifically described below in conjunction with the drawings and embodiments.
[0035] Embodiment 1
[0036] The multi-spectral image classification method based on LO-MLPRNN in this embodiment includes the following steps:
[0037] A. Construct a deep neural network model as shown in Figure 1 , where the deep neural network model includes: 0DC network, LSK network, GRU network, and MLP network;
[0038] B. Train the deep neural network model to obtain a trained deep neural network model;
[0039] The process of training the deep neural network model is as follows:
[0040] a. Preprocess the multi-spectral images from the GF-2 dataset and divide them into a training set and a validation set;
[0041] b. Use the data in the training set to train the LO-MLPRNN neural network model to obtain a trained and improved neural network model;
[0042] c. Use the data in the validation set to test the LO-MLPRNN neural network model to obtain a trained and improved LO-MLPRNN neural network model.
[0043] In the above step a, the process of preprocessing the multi-spectral remote sensing data is as follows:
[0044] According to the high-resolution GF-2 dataset, select ROI label data in the remote sensing image, randomly divide the image data into a training set and a validation set, and the ratio of the number of image data in the training set to the validation set is 6:4.
[0045] C. Input the original images into the ODC network and the LSK network for processing respectively. After adding and fusing the processing results of the ODC network and the LSK network, input them into the GRU network for processing, and input the obtained results into the MLP network for processing to obtain the final output result.
[0046] The structure of the 0DC network is as shown in Figure 2 . In the ODC network, dynamic convolution is used for processing to enhance the extraction of complex features of multi-spectral images. The dynamic convolution formula is as follows:
[0047] y = (α w1 ⊙ α f1 ⊙ α c1 ⊙ α s1 ⊙ W1 +... + α wm ⊙ α fn ⊙ α cn ⊙ α sn ⊙ W n ) * x(1)
[0048] where α wn represents the kernel intelligent multiplication operation along the kernel dimension of the convolution kernel space, and assigns an attention scalar to the entire convolution kernel. α fn represents the channel intelligent multiplication operation along the output channel dimension, and assigns different attention scalars to the convolution. α cn represents the filter intelligent multiplication operation along the input channel dimension, and assigns different attention scalars to each convolution. α sn represents the position intelligent multiplication operation along the spatial dimension, and assigns different attention scalars to the convolution parameters at the k×k spatial positions.
[0049] As Figure 3 shown, the processing process in the LSK network is as follows:
[0050] The input result is successively processed by the first large convolution and the second large convolution. After the results of the first large convolution and the second large convolution are concatenated by the channel concatenation module; the obtained concatenated result is divided into two paths. The first path is successively processed by average pooling, max pooling, the spatial feature description module, convolution, and the Sigmoid function to obtain the function processing result; the obtained function processing result is divided into two branches. The first branch is multiplied and fused with the result of the first large convolution to obtain the result of the first branch, and the second branch is multiplied and fused with the result of the second large convolution to obtain the result of the second branch; after the result of the first branch and the result of the second branch are added and fused, the result of the first path is obtained;
[0051] The second path is processed by the multi-head attention module to obtain the result of the second path;
[0052] After the result of the first path and the result of the second path are added and fused, they are multiplied and fused with the input result to obtain the output result.
[0053] The GRU network is a prior art and is a variant of the RNN network, and its structure is as Figure 4 shown;
[0054] As Figure 5 shown, 128 nodes are set in the input layer and 256 nodes are set in the hidden layer of the MLP network.
[0055] Embodiment 2
[0056] The selected satellite remote sensing training study area in the remote sensing multispectral image database is Liucheng County, Liuzhou City, Guangxi Zhuang Autonomous Region. The Liucheng County area belongs to the subtropical monsoon region, with hot summers and cold winters, distinct seasons, abundant light energy and water volume. The forest coverage is relatively wide and is very suitable for sugarcane crop cultivation. Therefore, this experiment selects the Liucheng County area for the following experiments. GF-2 data with a resolution of 0.75m was downloaded respectively, and the data sets were taken in Liucheng County on October 15, 2022. The GF-2 satellite data with a resolution of 0.75m is L3C-level image data, and its remote sensing satellite products include 4 spectral bands (B, G, R, NIR). After orthorectification, geometric correction, atmospheric correction and other processes, the spatial resolution and spectral resolution are relatively high. The training set is used to train the network model for the task of ground object recognition and classification, and the test set is used to test whether the network model is accurate.
[0057] The method of Example 1 (i.e., the LO-MLPRNN group) was compared and tested with the prior art method. The test results are shown in Table 1 below, and Figure 6 are as follows:
[0058] Table 1 Comparison of classification results of different classification methods on the GF-2 data set. The OA of the proposed LO-MLPRNN in this paper reaches 99.11%, which is the best among all the compared classification methods.
[0059]
[0060] Figure 6 It is a comparison chart of classification test results of different algorithms for the GF-2 data set. The comparison algorithm experiments in the figure are KNN, VIT, VITMLP, LO-MLPRNN in turn; among them, LO-MLPRNN is the algorithm of Example 1 of the present invention. From Figure 6 it can be seen that the results of the LO-MLPRNN method are better both globally and locally.
Claims
1. A multi-spectral image classification method based on LO-MLPRNN, characterized in that, It includes the following steps: A. Construct a deep neural network model, which includes: ODC network, LSK network, GRU network, and MLP network; B. Train the deep neural network model to obtain a trained deep neural network model; C. Input the original image into the ODC network and the LSK network for processing respectively. After the processing results of the ODC network and the LSK network are added and fused, they are input into the GRU network for processing, and the obtained result is input into the MLP network for processing to obtain the final output result.
2. The multi-spectral image classification method based on LO-MLPRNN according to claim 1, wherein: The process of training the deep neural network model in step B is as follows: a. Preprocess the multi-spectral images in the GF-2 dataset and divide them into a training set and a validation set; b. Use the data in the training set to train the LO-MLPRNN neural network model to obtain a trained and improved neural network model; c. Use the data in the validation set to test the LO-MLPRNN neural network model to obtain a trained and improved LO-MLPRNN neural network model.
3. The multi-spectral image classification method based on LO-MLPRNN according to claim 2, wherein: In step a, the process of preprocessing the multi-spectral remote sensing data is as follows: According to the high-resolution GF-2 dataset, select ROI label data in the remote sensing image, randomly divide the image data into a training set and a validation set, and the ratio of the number of image data in the training set to the validation set is 5-7:3-5.
4. The multi-spectral image classification method based on LO-MLPRNN according to claim 1, wherein: In the ODC network, it is processed through dynamic convolution to enhance the extraction of complex features of multi-spectral images. The dynamic convolution formula is as follows: y = (α w1 ⊙ α f1 ⊙ α c1 ⊙ α s1 ⊙ W1 +... + α wn ⊙ α fn ⊙ α cn ⊙ α sn ⊙ W n ) * x (1) where α wn represents a kernel intelligent multiplication operation along the kernel dimension of the convolution kernel space, and assigns an attention scalar to the entire convolution kernel, α fn represents a channel intelligent multiplication operation along the output channel dimension, and assigns different attention scalars to the convolution, α cn represents a filter intelligent multiplication operation along the input channel dimension, and assigns different attention scalars to each convolution, α sn represents a position intelligent multiplication operation along the spatial dimension, and assigns different attention scalars to the convolution parameters at the k×k spatial positions.
5. The multi-spectral image classification method based on LO-MLPRNN according to claim 1, wherein: The processing process in the LSK network is as follows: The input result is successively processed by the first large convolution and the second large convolution. After the results of the first large convolution and the second large convolution are spliced by the channel splicing module; the obtained spliced result is divided into two paths. The first path is successively processed by average pooling, max pooling, spatial feature description module, convolution, and Sigmoid function to obtain a function processing result; the obtained function processing result is divided into two branches. The first branch is multiplied and fused with the result of the first large convolution to obtain the result of the first branch, and the second branch is multiplied and fused with the result of the second large convolution to obtain the result of the second branch; after the result of the first branch and the result of the second branch are added and fused, the result of the first path is obtained; The second path is processed by the multi-head attention module to obtain the result of the second path; After the result of the first path and the result of the second path are added and fused, they are multiplied and fused with the input result to obtain the output result.
6. The multi-spectral image classification method based on LO-MLPRNN according to claim 1, wherein: In the MLP network described above, the input layer is set with 64 - 256 nodes, the hidden layer is set with 128 - 512 nodes, and the number of hidden layer nodes is twice that of the input layer.