An airborne and spaceborne remote sensing scene classification method based on a high-precision full adder network
By adopting high-precision full-added network and generation learning methods on drones and satellite-based platforms, the deployment problem of convolutional neural networks in resource-constrained environments is solved, and efficient and accurate remote sensing scenario classification is achieved, which is suitable for resource-sensitive application environments.
Patent Information
- Application Number
- CN202211680294.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-26
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-12-26
AI Technical Summary
In application scenarios where power consumption is strictly limited, such as drone and satellite, the high resource overhead and performance losses of convolutional neural networks are difficult to effectively deploy on edge devices such as FPGAs, and the feature extraction capability of the full-added network is insufficient, resulting in poor accuracy in remote sensing image processing tasks.
A high-precision full-added network structure is adopted, combined with a generative learning and mixed knowledge distillation training strategy based on knowledge matching, and a false sample with a data distribution similar to the real sample is generated through the generator. The loss function optimization generator and the full-added network are used to realize the efficient transfer of knowledge from a convolutional neural network to a full-added network, and improve the performance of the full-added network.
While ensuring the resource efficiency of hardware deployment, it significantly improves the accuracy of remote sensing scene classification of the full-plus network, meets the online intelligent scene classification needs of multifunctional remote sensing images for airborne and satellite-based platforms, and is suitable for resource-sensitive application environments.
Smart Images

Figure CN116363409B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to an airborne and spaceborne remote sensing scene classification method based on a high-precision fully-added network. Background Art
[0002] In recent years, convolutional neural networks have demonstrated powerful feature extraction capabilities in the field of image processing. Convolutional neural networks are computationally intensive, so most current deep neural network algorithms are deployed on high-performance devices such as CPUs or GPUs. However, the high performance of CPUs and GPUs requires a large amount of power consumption support, so it is difficult to apply them in application scenarios with strict power consumption limitations such as unmanned aerial vehicles and spaceborne platforms. High-performance, low-power embedded hardware devices provide a new solution to solve this problem. Therefore, building a deep neural network implementation platform based on embedded devices has increasingly become a research hotspot in engineering applications. Designing a deep neural network hardware accelerator based on FPGA is the most widespread solution among them. However, this solution also has certain technical difficulties. The on-chip computing resources of FPGA are limited and it is difficult to meet the resource overhead required for the intensive multiplication operations in convolutional neural networks during deployment. The fully-added network replaces all standard convolutional kernels in the traditional convolutional neural network with addition kernels, converting most of the multiplication operations in the network into addition operations that are more hardware-friendly, reducing the consumption of hardware resources and being suitable for deployment on edge devices such as FPGA. However, the network performance of the fully-added network in image processing tasks has a huge loss compared with convolutional neural networks, making it technically difficult to obtain a fully-added network suitable for deployment on edge devices such as FPGA and with excellent performance. Summary of the Invention
[0003] In view of this, the present invention provides an airborne and spaceborne remote sensing scene classification method based on a high-precision adder network, which includes proposing a novel adder network suitable for hardware deployment and a hybrid knowledge distillation training strategy based on incremental learning. Compared with traditional convolutional neural networks, the proposed adder network abandons all multiplication operations in the network convolutional layer and is completely composed of adder kernels; therefore, this network has lower resource overhead in hardware deployment than existing convolutional neural networks. However, in specific applications, such as intelligent interpretation of remote sensing images, the adder network has a large gap in accuracy compared with convolutional neural networks. The hybrid knowledge distillation training strategy based on generative learning significantly improves the network performance of the adder network by expanding teaching samples and using different teaching teachers. A large number of experiments and analyses on publicly available remote sensing scene classification datasets show that the finally obtained high-precision adder network achieves performance comparable to that of convolutional neural networks on most datasets. Therefore, the method proposed by the present invention minimizes the computing resource overhead required for hardware deployment while ensuring network accuracy. An airborne and spaceborne remote sensing scene classification method based on a high-precision adder network includes the following steps:
[0004] Step 1: Design an adder network that does not contain convolutional kernels but is composed of adder kernels according to the classical convolutional neural network structure as the basic network structure;
[0005] Step 2: Use the picture data of airborne and spaceborne remote sensing scenes as real samples to train a convolutional neural network and an adder network with the same structure as the target adder network, and obtain a well-trained convolutional neural network denoted as f CNN , and a well-trained adder network denoted as f ANN ;
[0006] Step 3: Use the generator G θ to generate fake samples with a data distribution similar to that of real samples as generated samples
[0007] Step 4: Optimize the generator G G through the loss function L θ :
[0008]
[0009] where α and β are hyperparameters for balancing the three loss functions;
[0010] The loss function is as follows:
[0011]
[0012] where CE[·] represents the cross-entropy loss; Denote the output of the well-trained convolutional neural network for the generated samples ; Denote the label of the generated samples;
[0013] The loss function L BN is as follows:
[0014]
[0015] where N represents the number of BN layers in the convolutional neural network, and are the mean and variance of the distribution of the generated samples in the i-th BN layer of the well-trained convolutional neural network knowledge source model, μ i and σ i are the mean and variance of the distribution of the real samples in the well-trained convolutional neural network stored in the i-th BN layer respectively, and MSE(·) is the mean square error operator;
[0016] The loss function L D is expressed as:
[0017]
[0018] and represent the outputs of the convolutional neural network and the addition network for the generated samples respectively; MAE(·) is the mean absolute error operator and is expressed as:
[0019]
[0020] where B represents the batch size of the generated samples set in each training;
[0021] Step Five: Design a joint loss function for optimizing the student fully additive network. The joint loss function is as follows:
[0022]
[0023] where λ is a hyperparameter for balancing the loss function;
[0024] The loss function is expressed as:
[0025]
[0026] where, is the output of the fully additive network for the input real samples, and l represents the real sample category of the real samples;
[0027] The loss function is as follows:
[0028]
[0029] where T is a hyperparameter that determines the softness of the labels. The higher the value of T, the softer the probability distribution over the classes. KLDiv(·) is the introduced Kullback-Leibler (KL) divergence function; is the class probability generated by the fully additive network using real samples as training data via the softmax function with T; P CNN is the class probability generated by the convolutional neural network using real samples as training data via the softmax function with T;
[0030] Loss function is as follows:
[0031]
[0032] where, is the class probability generated by the fully additive network using generated samples as training data via the softmax function with T; is the class probability generated by the convolutional neural network using generated samples as training data via the softmax function with T;
[0033] Step 6. After the generator generates a batch of generated samples each time, perform Step 4 and Step 5 to respectively complete one optimization of the generator and the fully additive network using their respective loss functions. Among them, each time the generator completes an optimization, the optimized generator generates a new batch of generated samples and continues to execute Step 4 and Step 5; and so on. During the iteration of multiple training batches, both the generator and the fully additive network are continuously optimized and updated; finally, a fully additive network is constructed.
[0034] Step 7. Use the optimized fully additive network to classify the airborne and spaceborne scenarios.
[0035] The present invention has the following beneficial effects:
[0036] The present invention provides an airborne and spaceborne remote sensing scene classification method based on a high-precision fully additive network. Compared with the traditional convolutional neural network, the proposed fully additive network significantly reduces the resource overhead required for hardware deployment; in addition, for this novel network structure of the fully additive network, the proposed training method greatly alleviates its performance loss. Using the proposed training method, a fully additive network with high resource efficiency and excellent performance can be obtained, which is suitable for hardware deployment in resource-sensitive scenarios.
[0037] In an application environment with severely limited resources and power consumption, the present invention can meet the huge computational and storage requirements of deep neural networks, providing support for realizing online intelligent scene classification of multi-functional remote sensing images on airborne and small satellite platforms, and playing an important role in civil and military fields such as disaster warning and emergency response, environmental monitoring, and intelligence collection. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flowchart of a high-precision full adder network for airborne and spaceborne remote sensing scene classification tasks of the present invention;
[0039] Figure 2 It is a schematic structural diagram of a well-trained convolutional neural network and adder network according to an embodiment of the present invention;
[0040] Figure 3 It is a training schematic diagram of a hybrid knowledge distillation training method based on generative learning provided in an embodiment of the present invention;
[0041] Figure 4 It is a training schematic diagram of a generative learning method based on knowledge matching provided in an embodiment of the present invention;
[0042] Figure 5 It is an operation schematic diagram of the inference stage of the full adder network obtained in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application.
[0044] The present invention provides an airborne and spaceborne remote sensing scene classification method based on a high-precision fully additive network. With the development of deep learning, a large number of different types of deep convolutional neural networks have emerged. Convolutional neural networks are characterized by a large number of network parameters and intensive calculations. Therefore, most current convolutional neural network algorithms use high-performance devices such as CPUs and GPUs to complete training and inference. Although CPUs and GPUs can achieve high performance during the deployment of convolutional neural networks, the huge power consumption that follows limits their application in scenarios with strict power consumption constraints such as unmanned aerial vehicles and spaceborne applications. Edge devices have the characteristics of low power consumption and high energy efficiency, so they are increasingly used by researchers for the deployment of network algorithms in resource-sensitive scenarios. However, the intensive multiplication operations in convolutional neural networks consume a large amount of resources during deployment, which limits their deployment on edge devices. Table 1 shows the computational resource overhead required for deploying addition operations and multiplication operations with different precisions on an FPGA. It can be seen that deploying multiplication operations requires more resource overhead than addition operations. Therefore, the provided fully additive network fundamentally solves the problem of high resource overhead caused by intensive multiplication operations when deploying complex neural networks.
[0045] Table 1 Resource overhead for deploying different arithmetic units on an FPGA
[0046]
[0047] However, since the feature extraction ability of the addition kernel lags behind that of the convolutional kernel, the performance of the fully additive network is severely lost. The provided high-precision fully additive network for airborne and spaceborne remote sensing scene classification tasks combines a generative learning method based on knowledge matching and a hybrid knowledge distillation method to alleviate the performance degradation of the fully additive network. In the hybrid knowledge distillation method, a well-trained additive network with both addition kernels and convolutional kernels is introduced, and the generated samples obtained by the generative learning method based on knowledge matching are used to guide the knowledge in the high-performance convolutional neural network to be successfully transferred to the fully additive network using real training samples. The generative learning method based on knowledge matching can train a generator to generate fake samples with a data distribution similar to that of real samples by matching the knowledge of the feature distribution of real samples in the knowledge source, and use the generated samples and real samples as two different training databases for hybrid knowledge distillation. Using the proposed training method, a fully additive network with high resource efficiency and excellent performance can be obtained, which is suitable for hardware deployment in resource-sensitive scenarios.
[0048] The flowchart of an airborne and spaceborne remote sensing scene classification method based on a high-precision fully additive network provided by the present invention is as Figure 1 shown, and includes the following steps:
[0049] Step 1: Design a fully additive network in the network structure that does not include convolutional kernels but is composed of addition kernels.
[0050] Based on the classical convolutional neural network structure as the basic network structure, a fully additive network is designed in which the network structure does not contain convolutional kernels but is composed of additive kernels, thereby reducing the resource overhead required for hardware deployment.
[0051] Step 2: Train a convolutional neural network and an additive network that are consistent with the target fully additive network structure.
[0052] Specifically, use real training samples to train a convolutional neural network and an additive network that are consistent with the target fully additive network structure. The well-trained convolutional neural network and additive network as shown Figure 2 can be obtained, where the well-trained convolutional neural network is denoted as f CNN , and the well-trained additive network is denoted as f ANN .
[0053] As shown Figure 3 , the well-trained convolutional neural network and additive network play an important role in a high-precision fully additive network for airborne and spaceborne remote sensing scene classification tasks proposed, and are used as the knowledge source in the knowledge matching-based generative learning method in Step 4, and the teacher network in the hybrid knowledge distillation method in Step 5.
[0054] Step 3: Use the knowledge matching-based generative learning method to train a generator to generate fake samples with a data distribution similar to that of real samples.
[0055] Specifically, in order to provide a training database for knowledge transfer to the additive network in the subsequent hybrid knowledge distillation method, the generator G θ matches the knowledge about real samples in the knowledge source network to generate new valid samples
[0056]
[0057] where n represents random noise, represents a random label, represents the generated samples of preset categories, and M represents the number of classes of real samples.
[0058] Step 4: Optimize the generator G G through the loss function L θ :
[0059]
[0060] where α and β are hyperparameters for balancing the three loss functions, and use L BN and L D to design the joint loss function L GOptimize the generator G θ as shown in Figure 4 .
[0061] Since the convolutional neural network has a better ability to extract features from samples than the additive network, by introducing the loss function L BN , extract the classification boundary and distribution knowledge from a well-trained convolutional neural network. The loss function is designed to use the cross-entropy loss CE to minimize the prediction result of the generated samples in a well-trained convolutional neural network and the preset generation category . Because effective generated samples should be classified into the preset generation category through a well-trained convolutional neural network The loss function is as follows:
[0062]
[0063] where CE[·] represents the cross-entropy loss;
[0064] In addition, the BN layer in the network is used to stabilize the data distribution, learn the classification boundary, and thus allow the data to be divided into different classes. Therefore, the knowledge of the classification boundary and data distribution can be obtained by using the BN layer to guide the distribution of the generated samples. The loss function L BN is designed to optimize the generator to ensure that the means and variances of each BN layer obtained by the generated samples through the knowledge source match as much as possible the means and variances of each BN layer obtained by the real samples. The loss function L BN is as follows:
[0065]
[0066] where N represents the number of BN layers in the convolutional neural network, and are the mean and variance of the distribution of the generated samples in the i-th BN layer of the well-trained convolutional neural network knowledge source model, respectively, and μ i and σ i are the mean and variance of the distribution of the real samples in the well-trained convolutional neural network knowledge source model stored in the i-th BN layer, respectively. MSE(·) is the mean square error operator, which can be expressed as:
[0067]
[0068] Through the above two loss functions, the feature distribution of the generated samples on the knowledge source of the convolutional neural network can gradually approach the feature distribution of the real samples. The role of the generated samples in a high-precision fully additive network for airborne and spaceborne remote sensing scene classification tasks is to serve as a training database for the knowledge transfer of the teacher neural network in hybrid knowledge distillation. Since the additive network and the fully additive network have similar network structures because they both contain additive kernels in their network structures, some knowledge about real samples is extracted from the well-trained additive network model to ensure that the generated samples also conform to the prior knowledge of the additive network. The model difference is used to quantitatively measure the classification boundary difference between the generated samples in the well-trained convolutional neural network and the additive network. Minimizing this model difference can ensure that the convolutional neural network and the additive network can classify the generated samples into the same label. The loss function L D is used to measure the difference between the classification boundaries of the well-trained convolutional neural network and the additive network when the generated samples are used as inputs, and is expressed as:
[0069]
[0070] where the log function is used to slow down the training of the generator and make the training more stable, and MAE(·) is the mean absolute error operator, which can be expressed as:
[0071]
[0072] where B represents the batch size of the generated samples set in each training.
[0073] By minimizing the loss function L D , the convolutional neural network and the additive network classify the generated samples into the same label. In this way, the feature distributions of the generated samples on the knowledge sources of the convolutional neural network and the additive network can gradually approach the feature distribution of the real samples.
[0074] This part of the generated samples can augment the training samples and play an important role in the hybrid knowledge distillation method in Step Four.
[0075] Step Five: Use the hybrid knowledge distillation method to improve the performance of the target fully additive network, thereby obtaining a fully additive network with high resource efficiency and excellent performance.
[0076] Specifically, a fully additive network with excellent performance can correctly classify real samples. Therefore, the real training samples are used as the inputs of the fully additive network, and the loss function is designed to minimize the distance between the prediction result of the fully additive network and the real sample category. The loss function is as follows:
[0077]
[0078] Among them is the output of the full adder network for the input real sample, and l represents the real sample category of the real sample.
[0079] In step two, the well-trained convolutional neural network and adder network are used as the two teacher networks in the hybrid knowledge distillation method, and the full adder network is used as the student network to improve the network performance of the full adder network. Since the feature extraction ability of the full adder network is weak, the feature knowledge extracted by it through the real sample category of the real sample is limited. Since the well-trained convolutional neural network has better feature extraction ability and classification performance than the well-trained adder network, the real training samples are used as the training database for the knowledge transfer of the teacher convolutional neural network to guide the full adder network to improve its network performance. Design the loss function to optimize the full adder network to ensure that the full adder network mimics the classification boundary of the teacher convolutional neural network on the real training samples. The loss function is as follows:
[0080]
[0081] where T is a hyperparameter that determines the softness of the label. The higher the value of T, the softer the probability distribution over the classes; KLDiv(·) is the introduced Kullback-Leibler (KL) divergence function, which enables the full adder network to mimic the classification boundary of the teacher convolutional neural network. is the class probability generated by the softmax function with T for the full adder network using real samples as training data; P CNN is the class probability generated by the softmax function with T for the convolutional neural network using real samples as training data.
[0082] Since both the full adder network and the adder network contain adder kernels, the feature extraction methods of the two models are similar. Therefore, in the embodiment of the present invention, a well-trained adder network is introduced as a teacher to guide the weight distribution of the full adder network. In particular, since there is a certain gap in network performance between the well-trained adder network and the well-trained convolutional neural network, using the same real training samples as the training database for the teacher adder network will have an adverse impact on the knowledge transfer performance of the teacher convolutional neural network. Therefore, the loss function is designed to optimize the full adder network to ensure that the student full adder network can mimic the weight distribution of the teacher adder network on the generated samples, specifically as follows:
[0083]
[0084] where, similar to formula (9), the difference is that the training samples used are the generated samples generated in step 3.
[0085] In the hybrid knowledge distillation method, a well-trained addition network with both addition kernels and convolutional kernels is introduced. Using the generated samples obtained by the generative learning method based on knowledge matching, the knowledge in the high-performance convolutional neural network is guided to smoothly transfer to the fully-additive network using real training samples. Through the combination of the above three loss functions, the constructed joint loss function can be used to optimize the student fully-additive network, and the loss function is as follows:
[0086]
[0087] where λ is a hyperparameter that balances the loss function.
[0088] Step 6: After the generator generates a batch of generated samples each time, perform Step 4 and Step 5 to respectively complete one optimization of the generator and the fully-additive network using their respective loss functions; among them, every time the generator completes an optimization, the optimized generator generates a new batch of generated samples and continues to execute Step 4 and Step 5; and so on. During the iteration of multiple training batches, both the generator and the fully-additive network are continuously optimized and updated. Finally, a fully-additive network with high resource efficiency and excellent performance is constructed.
[0089] Step 7: Use the optimized fully-additive network to classify airborne and spaceborne scenarios.
[0090] Example:
[0091] Taking the actual application scenario of on-orbit remote sensing scene classification as an example, on multiple publicly available datasets, compare the classification accuracies of the fully-additive network and the convolutional neural network obtained by the proposed high-precision fully-additive network for airborne and spaceborne remote sensing scene classification tasks. As shown in Table 2, it can be seen that the obtained fully-additive network has better classification accuracy compared to the convolutional neural network. In addition, as can be seen from Table 3, the fully-additive network and the convolutional neural network with the same structure have the same number of operation times, but the multiplication operations in the convolutional neural network are replaced by the same number of addition operations in the fully-additive network, thereby reducing the resource overhead required when the network is deployed on hardware devices such as FPGAs.
[0092] Table 2 Comparison of the classification accuracies [accuracy ± variance (%)] of the fully-additive network and the convolutional neural network obtained by the proposed high-precision fully-additive network for airborne and spaceborne remote sensing scene classification tasks on the UCM, SIRI-WHU, and AID datasets
[0093]
[0094] Table 3 Number of various operations used in different models [accuracy ± variance (%)]
[0095]
[0096] Of course, the present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can certainly make various corresponding changes and modifications according to the present invention. However, these corresponding changes and modifications should all fall within the protection scope of the appended claims of the present invention.
Claims
1. An airborne and spaceborne remote sensing scene classification method based on a high-precision full adder network, characterized in that, It includes the following steps: Step 1: Based on the classical convolutional neural network structure as the basic network structure, design a fully additive network in the network structure that does not contain convolutional kernels but is composed of additive kernels; Step 2: Use the image data of airborne and spaceborne remote sensing scenes as real samples to train a convolutional neural network and an addition network with the same structure as the target full adder network, and obtain a well-trained convolutional neural network denoted as f CNN , and a well-trained addition network denoted as f ANN ; Step 3: Use the generator G θ Generate fake samples with a data distribution similar to that of the real samples as the generated samples Step 4. Optimize the generator G G through the loss function L θ : where α and β are hyperparameters for balancing the three loss functions; Loss function As follows: where CE[·] represents the cross-entropy loss; represents the output of a well-trained convolutional neural network for the generated samples and represents the label of the generated samples; Loss function L BN is as follows: where N represents the number of BN layers in the convolutional neural network, and are the mean and variance of the distribution of the i-th BN layer in the well-trained convolutional neural network knowledge source model for the generated samples, respectively, μ i and σ i are the mean and variance of the distribution of the real samples in the well-trained convolutional neural network stored in the i-th BN layer, respectively, and MSE(·) is the mean square error operator; Loss function L D is expressed as: and respectively represent the outputs of the convolutional neural network and the addition network for the generated samples ; MAE(·) is the mean absolute error operator, expressed as: where B represents the batch size of the generated samples set in each training; Step 5: Design a joint loss function for optimizing the student fully additive network. The joint loss function is as follows: where λ is a hyperparameter for balancing the loss function; Loss function It is expressed as: Among them, is the output of the full addition network for the input real sample, and l represents the real sample category of the real sample; Loss function As follows: where T is a hyperparameter that determines the softness of the labels. The higher the value of T, the softer the probability distribution over the classes; KLDiv(·) is the introduced Kullback-Leibler (KL) divergence function; is the class probability generated by the softmax function with T for the fully additive network using real samples as training data; P CNN is the class probability generated by the softmax function with T for the convolutional neural network using real samples as training data; Loss function Specifically as follows: wherein, is the class probability generated by the full addition network using the generated samples as training data through the softmax function with T; is the class probability generated by the convolutional neural network using the generated samples as training data through the softmax function with T; Step 6: After the generator generates a batch of generated samples each time, execute Step 4 and Step 5, and use their respective loss functions to complete one optimization of the generator and the fully additive network respectively; among them, after the generator completes one optimization, the optimized generator generates a new batch of generated samples, and continue to execute Step 4 and Step 5; and so on. During the iteration of multiple training batches, the generator and the fully additive network are continuously optimized and updated; finally, a fully additive network is constructed; Step 7: Use the optimized fully additive network to classify the scenarios of airborne and spaceborne drones.
Citation Information
Patent Citations
Convolutional neural network accelerator based on FPGA and acceleration method thereof
CN115374929A