A remote sensing image scene classification method based on high performance quantized full-add network

By quantizing and debiasing the addition core of the initial full-added network, the target quantitative full-added network is obtained, which solves the resource overhead problem of deploying convolutional neural networks and full-added networks on satellite-on-board and airborne platforms, and realizes efficient remote sensing image scene classification.

CN118429686BActive Publication Date: 2025-05-13BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410324714.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2025-05-13
Estimated Expiration
2044-03-21

AI Technical Summary

Technical Problem

On satellite-on-board and airborne platforms, it is difficult to implement the deployment of convolutional neural networks due to the computational complexity and large parameters; while the full-added network reduces the overhead of computing resources, the parameters are still large, resulting in memory challenges.

Method used

A remote sensing image scene classification method based on high-sex energy-generated full-added network is proposed. By obtaining the initial full-added network, quantizing the addition core, performing debiased processing, replacing the addition core in the initial full-added network, obtaining the quantitative full-added network to be trained, and obtaining the target quantitative full-added network through debiased training.

Benefits of technology

On the premise of ensuring network accuracy, the resource overhead required for hardware deployment is minimized. It is suitable for satellite-on-board and airborne platforms with severe resources, and realizes the online intelligent interpretation task of remote sensing images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118429686B_ABST
    Figure CN118429686B_ABST
Patent Text Reader

Abstract

The present application provides a remote sensing image scene classification method based on a high-performance quantized full-additive network, which belongs to the field of image processing technology. The method includes: obtaining an initial full-additive network based on an initial additive kernel; quantizing the initial additive kernel to obtain a quantized additive kernel; performing debiased quantization processing on the quantized additive kernel to obtain a target quantized additive kernel, replacing the initial additive kernel with the target quantized additive kernel in the initial full-additive network to obtain a quantized full-additive network to be trained; performing debiased quantization training on the quantized full-additive network based on a benchmark full-additive network model to obtain a trained target quantized full-additive network, the debiased quantization training uses remote sensing image samples as training samples and scene classification of remote sensing images as downstream tasks; based on the target quantized full-additive network, scene classification is performed on any remote sensing image to be identified. By adopting the present application, the resource overhead required for hardware deployment can be minimized while ensuring network accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a remote sensing image scene classification method based on a high-performance quantized full-add network. Background Art

[0002] With the development of aerospace technology, earth observation and intelligent interpretation play an important role in disaster emergency response, national defense and military, and neural networks have become the mainstream method of remote sensing image interpretation in ground stations.

[0003] Completing remote sensing image interpretation tasks such as scene classification directly on satellite and airborne platforms can reduce the pressure on satellite-to-ground (air-to-ground) transmission links and improve response speed. It is difficult to carry common high-performance devices such as CPU (Central Processing Unit) or GPU (Graphics Processing Unit) in application scenarios with strict power consumption such as satellite and airborne. High-performance, low-power embedded hardware devices such as FPGA (Field Programmable Gate Array) are common implementation platforms in these application scenarios. However, their resources are very limited, and convolutional neural networks with high computational complexity and large number of parameters are difficult to deploy on these platforms. The full-add network is a novel basic network that replaces all standard convolution kernels in traditional convolutional neural networks with addition kernels, reducing the computational resource overhead during deployment. However, the number of parameters of the full-add network is comparable to that of the convolutional neural network, and its deployment on satellite and airborne platforms still faces memory challenges.

[0004] Quantized full-additive networks can significantly reduce storage overhead and further reduce computational overhead compared to full-additive networks. Compared with convolutional neural networks and full-additive networks, they have the advantages of low computational overhead and low storage overhead. However, there are certain technical difficulties in quantizing the full-additive network and obtaining a quantized full-additive network with image processing performance, so as to be deployed on edge devices such as FPGA to complete the online intelligent interpretation of satellite and airborne remote sensing images. Summary of the invention

[0005] In order to solve the existing technical problems, the embodiment of the present application provides a remote sensing image scene classification method based on a high-performance quantized full-add network, which can minimize the resource overhead required for hardware deployment while ensuring network accuracy. The technical solution is as follows:

[0006] According to one aspect of the present application, a remote sensing image scene classification method based on a high-performance quantized full-add network is provided, the method comprising:

[0007] Acquire an initial full-addition network based on an initial addition kernel, wherein the initial full-addition network has a preset convolutional neural network structure;

[0008] quantizing the initial additive kernel to obtain a quantized additive kernel;

[0009] Performing a debiased quantization process on the quantized addition kernel to obtain a target quantized addition kernel, and replacing the initial addition kernel with the target quantized addition kernel in the initial full-addition network to obtain a quantized full-addition network to be trained;

[0010] Based on the benchmark full-additive network model, debiasing quantization training is performed on the quantized full-additive network to be trained to obtain a trained target quantized full-additive network, wherein the debiasing quantization training uses remote sensing image samples as training samples and uses scene classification of remote sensing images as a downstream task;

[0011] Based on the target quantized full-add network, scene classification is performed on any remote sensing image to be identified.

[0012] According to another aspect of the present application, a remote sensing image scene classification device based on a high-performance quantized full-add network is provided, the device comprising:

[0013] An acquisition module, used for acquiring an initial full-addition network based on an initial addition kernel, wherein the initial full-addition network has a preset convolutional neural network structure;

[0014] A quantization module, used for performing quantization processing on the initial additive kernel to obtain a quantized additive kernel;

[0015] A first debiasing quantization module is used to perform debiasing quantization processing on the quantized addition kernel to obtain a target quantized addition kernel, and replace the initial addition kernel with the target quantized addition kernel in the initial full-addition network to obtain a quantized full-addition network to be trained;

[0016] A second debiasing quantization module is used to perform debiasing quantization training on the quantized full-additive network to be trained based on a reference full-additive network model to obtain a trained target quantized full-additive network, wherein the debiasing quantization training uses remote sensing image samples as training samples and uses scene classification of remote sensing images as a downstream task;

[0017] The recognition module is used to perform scene classification on any remote sensing image to be recognized based on the target quantization full-add network.

[0018] According to another aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above-mentioned remote sensing image scene classification method based on a high-performance quantized full-add network.

[0019] This application can achieve the following beneficial effects:

[0020] The present application provides a remote sensing image scene classification method based on a high-performance quantized full-additive network. The quantized full-additive network obtained by using the proposed quantization scheme and quantization strategy has the characteristics of low computational overhead, low storage overhead, and low performance loss compared to convolutional neural networks and full-additive networks. It is suitable for hardware deployment in scenarios with severe resource constraints, so as to complete remote sensing intelligent interpretation tasks in harsh environments such as satellites and airborne.

[0021] In application environments where resources and power consumption are strictly limited, this application can meet the deployment requirements of complex deep neural networks, provide support for online intelligent interpretation of remote sensing images such as scene classification on satellite and airborne platforms, and play an important role in civil and military fields such as disaster warning and emergency response, environmental monitoring, and intelligence collection. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Further details, features and advantages of the present application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0023] Figure 1 A flow chart of a remote sensing image scene classification method based on a high-performance quantized full-add network provided according to an exemplary embodiment of the present application is shown;

[0024] Figure 2 A schematic diagram of an initial full-add network provided according to an exemplary embodiment of the present application is shown;

[0025] Figure 3 A flowchart of constructing a quantized full-addition network according to an exemplary embodiment of the present application is shown;

[0026] Figure 4 A parameter count histogram of the first layer addition kernel of the full-add network of the well-trained VGG11 base network provided according to an exemplary embodiment of the present application is shown;

[0027] Figure 5 A schematic diagram of weight quantization before debiasing quantization provided according to an exemplary embodiment of the present application is shown;

[0028] Figure 6 A schematic diagram of weight quantization after debiasing quantization according to an exemplary embodiment of the present application is shown;

[0029] Figure 7 A schematic diagram of debiasing quantization training provided according to an exemplary embodiment of the present application is shown;

[0030] Figure 8 A schematic diagram of a scene classification application provided according to an exemplary embodiment of the present application is shown;

[0031] Fig. 9 A schematic block diagram of a remote sensing image scene classification device based on a high-performance quantized full-add network provided according to an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION

[0032] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not intended to limit the scope of protection of the present application.

[0033] It should be understood that the various steps described in the method implementation of the present application can be performed in different orders and / or performed in parallel. In addition, the method implementation may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0034] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". Relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0035] It should be noted that the modifications of "one" and "plurality" mentioned in the present application are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0036] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0037] With the development of deep learning, a large number of different types of deep convolutional neural networks have been generated. Convolutional neural networks have the characteristics of large network parameters and intensive computation. Therefore, most convolutional neural network algorithms currently use high-performance devices such as CPUs and GPUs to complete training and inference. Although CPUs and GPUs can achieve high performance when deploying convolutional neural networks, the huge power consumption that comes with it limits their application in application scenarios with strict power constraints such as satellite and airborne. Edge devices have the characteristics of low power consumption and high energy efficiency, so they are increasingly used by researchers to deploy network algorithms in resource-sensitive scenarios. However, the dense multiplication operations in convolutional neural networks require a lot of resources when deployed, which limits their deployment on edge devices. The full-add network is a novel basic network that replaces all standard convolution kernels in traditional convolutional neural networks with addition kernels, reducing the computational resource overhead during deployment. However, the number of parameters of the full-add network is comparable to that of the convolutional neural network, and deployment on satellite and airborne platforms still faces memory challenges. Table 1 shows the computational resource overhead of LUT (Logic Unit Table) and FF (Flip-Flop) required for deploying addition and multiplication operations of different precisions on FPGA. It can be seen that deploying multiplication operations requires more resource overhead than adding operations, and as the data bit width decreases, the resource overhead of deploying addition operations can be reduced by 2 orders of magnitude. Therefore, the quantized full-addition network provided in this application fundamentally solves the problem of high resource overhead when deploying complex neural networks.

[0038]

[0039] However, since the feature extraction capability of the addition kernel lags behind that of the convolution kernel, the performance loss of the full-add network is serious. This application proposes a debiasing quantization method based on a shared scale factor designed specifically for a quantized full-add network to obtain a resource-efficient and high-performance quantized full-add network. The debiasing quantization method based on a shared scale factor combines a shared scale factor quantization scheme based on a power exponent and a multidimensional debiasing quantization strategy. The shared scale factor quantization scheme based on a power exponent converts an addition kernel with floating-point input activation, weights, and operations into a quantized addition kernel with hardware-friendly integer input activation, weights, and operations. During hardware deployment, a quantized full-add network composed of quantized adder filters has lower computational and memory overhead than a full-add network. The multidimensional debiasing quantization strategy combines a weight debiasing strategy and a feature debiasing strategy to avoid the performance degradation of the quantized full-add network caused by the offset of weights and features during the quantization process. Using the quantization method proposed in this application, a resource-efficient and high-performance quantized full-add network can be obtained, which is suitable for hardware deployment in resource-sensitive scenarios, so as to complete the online intelligent interpretation task of remote sensing images in harsh environments such as satellite and airborne.

[0040] The present application provides a remote sensing image scene classification method based on a high-performance quantized full-add network, which proposes a debiasing quantization method based on a shared scale factor designed specifically for a quantized full-add network. The debiasing quantization method based on a shared scale factor combines a shared scale factor quantization scheme based on a power exponent and a multidimensional debiasing quantization strategy. The shared scale factor quantization scheme based on a power exponent converts adder filters with floating-point input activations, weights, and operations into quantized adder filters with hardware-friendly integer input activations, weights, and operations. During hardware deployment, a quantized full-add network composed of quantized adder filters has lower computational and memory overhead than a full-add network. However, the accuracy of image interpretation tasks such as remote sensing scene classification of low-bit-width quantized full-add networks is significantly reduced. The multidimensional debiasing quantization strategy combines a weight debiasing strategy and a feature debiasing strategy to avoid the performance degradation of the quantized full-add network due to the offset of weights and features during the quantization process. Among them, the former avoids the performance degradation caused by the offset of quantization weights, and the latter improves the performance of the quantized full-add network in remote sensing image interpretation tasks such as scene classification by minimizing the offset of the output features of each layer. A large number of experiments and analyses on public remote sensing scene classification datasets show that the quantized full-add network obtained by using the proposed quantization scheme and quantization strategy has the characteristics of low computational overhead, low storage overhead, and low performance loss compared to convolutional neural networks and full-add networks, and is suitable for hardware deployment in resource-constrained scenarios, so as to complete remote sensing intelligent interpretation tasks in harsh environments such as satellite and airborne. Therefore, the method proposed in this application minimizes the resource overhead required for hardware deployment while ensuring network accuracy.

[0041] The following will refer to Figure 1 A flow chart of a remote sensing image scene classification method based on a high-performance quantized full-add network is shown, and the method is introduced.

[0042] Step 101, obtaining an initial full-add network based on an initial addition core.

[0043] Among them, the initial full-add network has a preset convolutional neural network structure.

[0044] In one possible implementation, the initial full-add network is as follows: Figure 2 As shown in the figure, according to the preset classic convolutional neural network structure, an initial full-additive network can be designed that does not contain convolution kernels but consists of multiple layers of addition kernels and BN (Batch Normalization) layers. Using the existing training method, an initial full-additive network with superior performance that can be used for airborne and satellite-borne remote sensing image scene classification can be obtained after pre-training.

[0045] Optionally, the first expression of the initial additive kernel is:

[0046]

[0047] in, represents the input activation tensor, represents the output feature tensor, represents the weight tensor of the initial additive kernel, h in and w in Represent the length and width of the input activation tensor, respectively. c in represents the number of channels of the input activation tensor, h out and w out Represent the length and width of the output feature tensor respectively, c out represents the number of channels of the output feature tensor, m , n Represent the length and width of the weight tensor respectively, k Represents the number of input channels, t Represents the number of output channels, x , y Represent the length and width of the output feature tensor respectively. d Represents the length and width of the initial additive kernel, that is, the length and width of the initial additive kernel are consistent.

[0048] Step 102, quantize the initial additive kernel to obtain a quantized additive kernel.

[0049] In one possible implementation, the initial additive kernel has input activations, weights, and operations in floating-point form, and a power-based shared scale factor quantization scheme can be used to convert the initial additive kernel into a quantized additive kernel having input activations, weights, and operations in hardware-friendly integer form.

[0050] Specifically, in order to convert the addition cores that occupy the vast majority of operations in the addition network from floating-point operations to hardware-friendly integer forms, it is necessary to convert the addition cores with floating-point input activations, weights, and operations to quantized addition cores with hardware-friendly integer forms of input activations, weights, and operations. Therefore, it is necessary to propose a quantization scheme suitable for the addition cores, such as Figure 3 The flowchart for constructing a quantitative full-additive network is shown.

[0051] The specific steps are as follows:

[0052] Converting the above initial addition kernel into integer form, the second expression is as follows:

[0053]

[0054] The activation scale factor will be entered Approximately in the form of a power of 2, the third expression is as follows:

[0055]

[0056] The weight scaling factor Approximately in the form of a power of 2, the fourth expression is as follows:

[0057]

[0058] in, max(·) and min(·) are functions for determining the largest and smallest elements in a given vector, respectively. N Indicates the bit width of the quantized value;

[0059] i It is expressed as the following fifth expression:

[0060]

[0061] j It is expressed as the following sixth expression:

[0062]

[0063] in, round(·) Indicates rounding to the nearest integer value;

[0064] According to the third expression above, the quantized input activation value I int It is expressed as the following seventh expression:

[0065]

[0066] According to the fourth expression above, the quantized weight value is F int It is expressed as the following eighth expression:

[0067]

[0068] in, clamp(·) The function is used to limit the quantized value to the target quantization range. ;

[0069] The initial addition kernel in integer form is further rewritten as the following ninth expression:

[0070]

[0071] To further transform the operation of the adder filter into a hardware-friendly form, the scale factor can be activated according to the input and weight scaling factor The relative size of the scale factor will be input to activate the scale factor and weight scaling factor The shared scale factor is defined as the following tenth expression:

[0072]

[0073] According to i and j The relationship between can be divided into three cases. The operation of the addition kernel is converted into a quantized addition kernel operation with an inverse quantization operation. The initial addition kernel of the ninth expression above is converted into a quantized addition kernel as follows:

[0074] like i < j , then it is expressed as the following eleventh expression:

[0075]

[0076] like i = j , then it is expressed as the following twelfth expression:

[0077]

[0078] like i > j , then it is expressed as the following thirteenth expression:

[0079] .

[0080] With this quantization scheme, all floating-point addition operations in the addition core are converted to integer addition operations. It is worth noting that no other operations are introduced except for hardware-friendly bit shift operations. The quantized addition core uses integer input activations and weights to perform integer addition operations, resulting in quantized output features. Then, the power-based shared scaling factor s A The quantized output features are inversely quantized to obtain the final output features of the additive kernel.

[0081] In order to reduce the accuracy degradation caused by the deviation of weights and features during the quantization process, this application proposes a multi-dimensional debiasing quantization strategy that combines feature debiasing and weight debiasing. The weight debiasing strategy mitigates the performance degradation of the quantized full-additive network by correcting the deviation in the quantization weight distribution. It redefines the weight scaling factor when the weight distribution is skewed and greatly exceeds the target quantization range to ensure that there is sufficient quantization range for densely distributed weights close to zero. The feature debiasing strategy improves the classification performance of the quantized full-additive network by minimizing the deviation between the output features of each layer. It aligns the output features of the intermediate and last layers in the quantized full-additive network with the output features of the corresponding layers in the benchmark full-additive network model, reduces the quantization error layer by layer, and improves the feature extraction capability of the quantized full-additive network.

[0082] The corresponding processing of the above-mentioned weight debiasing strategy will be introduced in step 103 below.

[0083] Step 103, de-biasing and quantizing the quantized additive kernel to obtain a target quantized additive kernel, and replacing the initial additive kernel with the target quantized additive kernel in the initial full-additive network to obtain a quantized full-additive network to be trained.

[0084] In one possible implementation, in a well-trained full-additive network in step 101 , the weights of the additive kernels generally follow a Laplace distribution, and many weights are concentrated near zero. Figure 4 The parameter count histogram of the first layer of the adder kernel of a well-trained fully-additive network - VGG11 model. It can be seen that the weight distribution of the adder kernel is skewed. The weight distribution of this filter ranges from -50 to 40, but most of the weights are distributed between -2 and 2, accounting for more than 80%. In contrast, there are very few weights less than -30 or greater than 20, accounting for less than 1%. Due to the scale factor of the weights s F According to the weight distribution F min and F max Therefore, in a low-bit-width quantized full-add network, the quantized weights may have serious deviations. Figure 5 The schematic diagram of weight quantization before debiasing quantization is shown. When the target quantization range is much smaller than the actual weight distribution range, many weights densely distributed near zero are compressed to integer zero after quantization. Therefore, in a well-trained full-precision full-add network model, most of the information is lost during the quantization process, resulting in a serious drop in accuracy. It is worth noting that the deviation of the weight distribution after quantization increases as the quantization bit width decreases. On this basis, the scale factor of the weight can be redefined according to the current weight distribution s F , in order to effectively alleviate the problem of serious degradation of the accuracy of the quantized full-add network caused by the weight distribution deviation during the quantization process.

[0085] Specifically, refer to Figure 6 The schematic diagram of weight quantization after de-biased quantization is shown. When the weight distribution is skewed and the weight distribution range is much larger than the target quantization range, the median F Me Ratio to boundary value F min and F max It can better reflect the actual weight distribution. Therefore, in the quantized additive kernel, if the weight scale factor If it is greater than 1, the weight scale factor The fourteenth expression is redefined as follows:

[0086]

[0087] j Re-expressed as the following fifteenth expression:

[0088]

[0089] in, and Represent the negative median and positive median in the weight distribution respectively.

[0090] like Figure 6 As shown, using the redefined weight scaling factor s F , the median of the weight distribution can be F Me The weight debiasing strategy ensures that the weights densely distributed around zero in the additive kernel have sufficient quantization range, effectively alleviating the performance degradation caused by the large amount of information loss during the quantization process. In particular, the proposed weight debiasing strategy does not introduce additional computational or memory overhead in the inference phase of the quantized full-additive network model.

[0091] In the quantization process of the full-additive network, not only will there be obvious deviations in the weight distribution, but also due to the accumulation of quantization errors, the output features of each layer will also have obvious deviations. Therefore, this application also introduces another feature debiasing strategy combined with the weight debiasing strategy to improve the classification performance of the quantized full-additive network by reducing the deviation between the output features of each layer of the quantized full-additive network and the benchmark full-additive network model.

[0092] The corresponding processing of the above-mentioned feature debiasing strategy will be introduced in step 104 below.

[0093] Step 104 , performing debiasing quantization training on the quantized full-add network to be trained based on the reference full-add network model to obtain a trained target quantized full-add network.

[0094] The debiased quantization training uses remote sensing image samples as training samples and the scene classification of remote sensing images as downstream tasks. The benchmark full-additive network model may refer to a full-additive network that has not been quantized, has superior performance, and can be used for airborne and spaceborne remote sensing image scene classification. In some possible implementations, the benchmark full-additive network model may be the same as the initial full-additive network.

[0095] In a possible implementation, during the debiased quantization training, the quantization error of the intermediate layer can be effectively reduced layer by layer by aligning the output features of each intermediate layer in the quantized full-additive network model with the output features of the corresponding layer in the reference full-additive network model. In this way, the accumulation of quantization errors can be reduced, which has a great impact on subsequent layers in the quantized full-additive network model. Furthermore, the feature extraction capability of the target full-additive network model can be made as close as possible to that of the reference full-additive network model by aligning the output features of the last layer of the target quantized full-additive network model with the output features of the last layer of the reference full-additive network model.

[0096] Specifically, refer to Figure 7 The debiasing quantization training schematic diagram is shown, and the processing of the above step 104 can be as follows:

[0097] The remote sensing image samples are processed by the benchmark full-add network model and the quantized full-add network to be trained respectively;

[0098] Calculate the first loss based on the difference between the first scene classification result output by the quantized full-add network to be trained and the scene classification truth value of the remote sensing image sample L Q ;

[0099] Calculate the second loss based on the difference between the first intermediate layer output features of the quantized full-add network to be trained and the second intermediate layer output features of the benchmark full-add network model L IF ;

[0100] Calculate the third loss based on the difference between the first scene classification result output by the quantized full-add network to be trained and the second scene classification result output by the baseline full-add network model L LF ;

[0101] The first loss L Q Second loss L IF and the third loss L LF , the joint loss function is determined by the following sixteenth expression :

[0102]

[0103] in, λ is a hyperparameter used to balance the loss function;

[0104] Based on the joint loss function , the quantized full-add network to be trained is updated, and after multiple iterations of training batches, the trained target quantized full-add network is obtained.

[0105] Loss Function L Q The purpose is to minimize the distance between the classification result of the quantized full-additive network and the true value. A quantized full-additive network model with excellent classification performance is used for training samples. S The classification accuracy should be high. Optionally, the loss function can be constructed using cross entropy loss L Q , then the first loss L Q The seventeenth expression is as follows:

[0106]

[0107] in, is the cross entropy loss function, is the function set of the quantized full-add network to be trained, S is the remote sensing image sample to be tested, It is the scene classification truth value of remote sensing image samples.

[0108] Loss Function L IF It is designed to measure the difference between the output features of the intermediate batch normalization (BN) layer in the quantized full-additive network and the baseline full-additive network model. The full-additive network model adds a BN layer after each adder layer to stabilize the data distribution and facilitate the learning of classification boundaries. Therefore, by minimizing L IF , can minimize the deviation between the output features of each intermediate layer due to quantization error, thereby reducing the adverse effect of accumulated quantization error on subsequent layers. Optionally, the second loss L IF The eighteenth expression is as follows:

[0109]

[0110] in, M represents the number of batch normalization (BN) layers in the quantized full-add network, and They are the first and second BN layers in the quantized full-add network and the benchmark full-add network models. b Layer output features; MSE(·)is the mean square error operator, and its nineteenth expression is as follows:

[0111]

[0112] in, B Indicates the batch size during the training phase.

[0113] Formulate loss function L LF To determine the difference between the class probabilities of the last layer in the quantized fully additive network model and the baseline fully additive network model. Due to the information loss in the quantization process, the feature extraction ability in the quantized fully additive network model is worse than that of the full precision fully additive network model. Minimize L LF To optimize the quantized full-additive network model, ensure that the classification boundary of the last layer in the quantized full-additive network model is similar to the classification boundary of the benchmark full-additive network model, thereby improving the feature extraction ability of the quantized full-additive network model. Optionally, the Kullback-Leibler (KL) divergence can be introduced to construct the loss function L LF , then the third loss L LF The twentieth expression is as follows:

[0114]

[0115] in, C Indicates the number of categories for scene classification, It is the class probability generated by the quantized full-add network using the normalized exponential (softmax) function; is the class probability generated by the baseline full-additive network model via a normalized exponential function; means KL Divergence;

[0116] The k The elements are determined by the following twenty-first expression:

[0117]

[0118] The k The elements are determined by the following twenty-second expression:

[0119]

[0120] in, represents the processing function of the quantized full-add network, represents the processing function of the benchmark full-additive network model, T Represents the temperature coefficient in the class probability generation process.

[0121] Step 105: Based on the target quantization full-add network, scene classification is performed on any remote sensing image to be identified.

[0122] In a possible implementation, after the iteration of the above-mentioned multiple training batches, the quantized full-add network is continuously optimized and updated. Finally, a resource-efficient and high-performance target quantized full-add network is constructed, and then the target quantized full-add network can be deployed on hardware devices such as FPGA. Whenever any satellite-borne or airborne remote sensing image needs to be interpreted, the remote sensing image can be processed by the target quantized full-add network to identify the scene classification corresponding to the remote sensing image, such as Figure 8 As shown, scene classification may include airports, beaches, bridges, train stations, ports, viaducts, etc.

[0123] Taking the actual application scenario of on-orbit remote sensing scene classification as an example, the classification accuracy of the quantized full-addition network proposed in this application is compared with that of the unquantized full-addition network and convolutional neural network on multiple public data sets, as shown in Table 2. It can be seen that until quantized to 6-bit, the quantized full-addition network can still achieve performance indicators that exceed those of the convolutional neural network and the full-addition network on multiple data sets. In addition, it can be seen from Table 3 that the quantized full-addition network, the unquantized full-addition network and the convolutional neural network of the same structure have the same number of operations, but the multiplication operation in the convolutional neural network is replaced by the same number of addition operations in the full-addition network and the quantized full-addition network, thereby reducing the computing resource overhead required when the network is deployed on hardware devices such as FPGAs; on this basis, the quantized full-addition network can significantly reduce the storage resource overhead required when deployed on hardware devices such as FPGAs.

[0124]

[0125]

[0126] The embodiments of the present application can achieve the following beneficial effects:

[0127] The present application provides a remote sensing image scene classification method based on a high-performance quantized full-additive network. The quantized full-additive network obtained by using the proposed quantization scheme and quantization strategy has the characteristics of low computational overhead, low storage overhead, and low performance loss compared to convolutional neural networks and full-additive networks. It is suitable for hardware deployment in scenarios with severe resource constraints, so as to complete remote sensing intelligent interpretation tasks in harsh environments such as satellites and airborne.

[0128] In application environments where resources and power consumption are strictly limited, this application can meet the deployment requirements of complex deep neural networks, provide support for online intelligent interpretation of remote sensing images such as scene classification on satellite and airborne platforms, and play an important role in civil and military fields such as disaster warning and emergency response, environmental monitoring, and intelligence collection.

[0129] The present application embodiment provides a remote sensing image scene classification device based on a high-performance quantized full-add network, which is used to implement the above remote sensing image scene classification method. Fig. 9 The schematic block diagram of the remote sensing image scene classification device shown in FIG. 900 , the remote sensing image scene classification device 900 includes: an acquisition module 901 , a quantization module 902 , a first debiasing quantization module 903 , a second debiasing quantization module 904 , and an identification module 905 .

[0130] An acquisition module 901 is used to acquire an initial full-addition network based on an initial addition kernel, wherein the initial full-addition network has a preset convolutional neural network structure;

[0131] A quantization module 902, configured to perform quantization processing on the initial additive kernel to obtain a quantized additive kernel;

[0132] A first debiasing quantization module 903 is used to perform debiasing quantization processing on the quantized addition kernel to obtain a target quantized addition kernel, and replace the initial addition kernel with the target quantized addition kernel in the initial full-addition network to obtain a quantized full-addition network to be trained;

[0133] A second debiasing quantization module 904 is used to perform debiasing quantization training on the quantized full-additive network to be trained based on the reference full-additive network model to obtain a trained target quantized full-additive network, wherein the debiasing quantization training uses remote sensing image samples as training samples and uses scene classification of remote sensing images as a downstream task;

[0134] The recognition module 905 is used to perform scene classification on any remote sensing image to be recognized based on the target quantization full-add network.

[0135] Optionally, the first expression of the initial additive kernel is:

[0136]

[0137] in, represents the input activation tensor, represents the output feature tensor, represents the weight tensor of the initial additive kernel, h in and w in Represent the length and width of the input activation tensor, respectively. cin represents the number of channels of the input activation tensor, h out and w out Represent the length and width of the output feature tensor respectively, c out represents the number of channels of the output feature tensor, d represents the length and width of the initial additive kernel, m , n Represent the length and width of the weight tensor respectively, k Represents the number of input channels, t Represents the number of output channels, x , y Represent the length and width of the output feature tensor respectively.

[0138] Optionally, the quantization module 902 is used to:

[0139] The initial addition kernel is converted into integer form, and the second expression is obtained as follows:

[0140]

[0141] The activation scale factor will be entered Approximately in the form of a power of 2, the third expression is as follows:

[0142]

[0143] The weight scaling factor Approximately in the form of a power of 2, the fourth expression is as follows:

[0144]

[0145] in, max(·) and min(·) are functions for determining the largest and smallest elements in a given vector, respectively. N Indicates the bit width of the quantized value;

[0146] i It is expressed as the following fifth expression:

[0147]

[0148] j It is expressed as the following sixth expression:

[0149]

[0150] in, round(·) Indicates rounding to the nearest integer value;

[0151] According to the third expression, the quantized input activation value I int It is expressed as the following seventh expression:

[0152]

[0153] According to the fourth expression, the quantized weight value F int It is expressed as the following eighth expression:

[0154]

[0155] in, clamp(·) The function is used to limit the quantized value to the target quantization range. ;

[0156] The initial addition kernel in integer form is further rewritten as the following ninth expression:

[0157]

[0158] The input activation scale factor and the weight scaling factor The shared scale factor is defined as the following tenth expression:

[0159]

[0160] According to i and j The relationship between , converting the initial addition kernel of the ninth expression into a quantized addition kernel is as follows:

[0161] like i < j , then it is expressed as the following eleventh expression:

[0162]

[0163] like i = j , then it is expressed as the following twelfth expression:

[0164]

[0165] like i > j , then it is expressed as the following thirteenth expression:

[0166] .

[0167] Optionally, the first de-biasing and quantization module 903 is configured to:

[0168] In the quantized additive kernel, if the weight scale factor is greater than 1, the weight scale factor The fourteenth expression is redefined as follows:

[0169]

[0170] j Re-expressed as the following fifteenth expression:

[0171]

[0172] in, and Represent the negative median and positive median in the weight distribution respectively.

[0173] Optionally, the second de-biasing and quantization module 904 is configured to:

[0174] Processing the remote sensing image samples respectively through the benchmark full-addition network model and the quantized full-addition network to be trained;

[0175] Based on the difference between the first scene classification result output by the quantized full-add network to be trained and the scene classification true value of the remote sensing image sample, a first loss is calculated. L Q ;

[0176] Based on the difference between the first intermediate layer output feature of the quantized full-add network to be trained and the second intermediate layer output feature of the benchmark full-add network model, a second loss is calculated. L IF ;

[0177] Based on the difference between the first scene classification result output by the quantized full-add network to be trained and the second scene classification result output by the benchmark full-add network model, a third loss is calculated. L LF ;

[0178] The first loss L Q The second loss L IF and the third loss L LF , the joint loss function is determined by the following sixteenth expression :

[0179]

[0180] in, λ is a hyperparameter used to balance the loss function;

[0181] Based on the joint loss function , the quantized full-add network to be trained is updated, and after multiple training batches of iterations, a trained target quantized full-add network is obtained.

[0182] The embodiments of the present application can achieve the following beneficial effects:

[0183] The present application provides a remote sensing image scene classification method based on a high-performance quantized full-additive network. The quantized full-additive network obtained by using the proposed quantization scheme and quantization strategy has the characteristics of low computational overhead, low storage overhead, and low performance loss compared to convolutional neural networks and full-additive networks. It is suitable for hardware deployment in scenarios with severe resource constraints, so as to complete remote sensing intelligent interpretation tasks in harsh environments such as satellites and airborne.

[0184] In application environments where resources and power consumption are strictly limited, this application can meet the deployment requirements of complex deep neural networks, provide support for online intelligent interpretation of remote sensing images such as scene classification on satellite and airborne platforms, and play an important role in civil and military fields such as disaster warning and emergency response, environmental monitoring, and intelligence collection.

[0185] Of course, the present application may have many other embodiments. Without departing from the spirit and essence of the present application, technicians familiar with the field can certainly make various corresponding changes and modifications based on the present application, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present application.

Claims

1. A remote sensing image scene classification method based on a high-performance quantized full-add network, characterized in that: The method includes: Obtain an initial full adder network based on an initial adder kernel, where the initial full adder network has a preset convolutional neural network structure; wherein, the first expression of the initial adder kernel is: in, represents the input activation tensor, represents the output feature tensor, represents the weight tensor of the initial additive kernel, h in and w in Respectively represent the length and width of the input activation tensor, c in represents the number of channels of the input activation tensor, h out and w out Respectively represent the length and width of the output feature tensor, c out represents the number of channels of the output feature tensor, d represents the length and width of the initial additive kernel, m and n represent the length and width of the weight tensor, k represents the number of input channels, t represents the number of output channels, and x and y represent the length and width of the output feature tensor, respectively; Convert the initial adder kernel into an integer form to obtain the following second expression: The input activation scale factor S I Approximately in the form of a power of 2, the third expression is as follows: The weight scaling factor S F Approximately in the form of a power of 2, the fourth expression is as follows: where max(·) and min(·) are functions for determining the maximum and minimum elements in a given vector respectively, and N represents the number of bits of the quantized value; i is expressed as the following fifth expression: j is expressed as the following sixth expression: where round(·) represents rounding to the nearest integer value; According to the third expression, the quantized input activation value I int It is expressed as the following seventh expression: I int =clamp(2 -i I,-2 N-1 +1,2 N-1 -1) According to the fourth expression, the quantized weight value F int It is expressed as the following eighth expression: F int =clamp(2 -j F,-2 N-1 +1,2 N-1 -1) The clamp(·) function is used to limit the quantized value to the target quantization range [-2 N-1 +1,2 N-1 -1]; Rewrite the initial adder kernel in integer form into the following ninth expression: The input activation scale factor S I and the weight scaling factor S F The shared scale factor is defined as the following tenth expression: s A =min(2 i ,2 j ) Then, according to the relationship between i and j, convert the initial adder kernel of the ninth expression into a quantized adder kernel as follows: If i < j, it is expressed as the following eleventh expression: If i = j, it is expressed as the following twelfth expression: If i > j, it is expressed as the following thirteenth expression: Perform debiasing quantization processing on the quantized adder kernel to obtain a target quantized adder kernel, and replace the initial adder kernel with the target quantized adder kernel in the initial full adder network to obtain a quantized full adder network to be trained; wherein, the performing debiasing quantization processing on the quantized adder kernel to obtain a target quantized adder kernel includes: In the quantized additive kernel, if the weight scaling factor S F is greater than 1, then the weight scale factor S F The fourteenth expression is redefined as follows: j is re-expressed as the following fifteenth expression: Among them, F′ Me and F″ Me Represent the negative median and positive median in the weight distribution respectively; Perform debiasing quantization training on the quantized full adder network to be trained based on a reference full adder network model to obtain a trained target quantized full adder network, where the debiasing quantization training uses remote sensing image samples as training samples and scene classification of remote sensing images as a downstream task; Based on the target quantized full adder network, perform scene classification on any remote sensing image to be recognized.

2. The method according to claim 1, characterized in that The performing debiasing quantization training on the quantized full adder network to be trained based on a reference full adder network model to obtain a trained target quantized full adder network includes: Process the remote sensing image samples through the reference full adder network model and the quantized full adder network to be trained respectively; Based on the difference between the first scene classification result output by the quantized full-add network to be trained and the scene classification true value of the remote sensing image sample, a first loss L is calculated. Q ; Based on the difference between the first intermediate layer output feature of the quantized full-add network to be trained and the second intermediate layer output feature of the benchmark full-add network model, the second loss L is calculated. IF ; Based on the difference between the first scene classification result output by the quantized full-add network to be trained and the second scene classification result output by the benchmark full-add network model, a third loss L is calculated. LF ; The first loss L Q The second loss L IF and the third loss L LF , the joint loss function L is determined by the following sixteenth expression: FD : L FD =L Q +λ(L IF +L LF ) where λ is a hyperparameter for balancing the loss function; Based on the joint loss function L FD , the quantized full-add network to be trained is updated, and after multiple training batches of iterations, a trained target quantized full-add network is obtained.

3. A remote sensing image scene classification device based on a high-performance quantized full-add network, characterized in that: The apparatus includes: An acquisition module, configured to acquire an initial full adder network based on an initial adder kernel, where the initial full adder network has a preset convolutional neural network structure; wherein, the first expression of the initial adder kernel is: in, represents the input activation tensor, represents the output feature tensor, represents the weight tensor of the initial additive kernel, h in and w in Respectively represent the length and width of the input activation tensor, c in represents the number of channels of the input activation tensor, h out and w out Respectively represent the length and width of the output feature tensor, c out represents the number of channels of the output feature tensor, d represents the length and width of the initial additive kernel, m and n represent the length and width of the weight tensor, k represents the number of input channels, t represents the number of output channels, and x and y represent the length and width of the output feature tensor, respectively; A quantization module, configured to convert the initial adder kernel into an integer form to obtain the following second expression: The input activation scale factor S I Approximately in the form of a power of 2, the third expression is as follows: The weight scaling factor S F Approximately in the form of a power of 2, the fourth expression is as follows: where max(·) and min(·) are functions for determining the maximum and minimum elements in a given vector respectively, and N represents the number of bits of the quantized value; i is expressed as the following fifth expression: j is expressed as the following sixth expression: where round(·) represents rounding to the nearest integer value; According to the third expression, the quantized input activation value I int It is expressed as the following seventh expression: I int =clamp(2 -i L,-2 N-1 +1,2 N-1 -1) According to the fourth expression, the quantized weight value F int It is expressed as the following eighth expression: F int =clamp(2 -j F,-2 N-1 +1,2 N-1 -1) The clamp(·) function is used to limit the quantized value to the target quantization range [-2 N-1 +1,2 N-1 -1]; Rewrite the initial adder kernel in integer form into the following ninth expression: The input activation scale factor S I and the weight scaling factor S F The shared scale factor is defined as the following tenth expression: s A =min(2 i ,2 j ) Then, according to the relationship between i and j, convert the initial adder kernel of the ninth expression into a quantized adder kernel as follows: If i < j, it is expressed as the following eleventh expression: If i = j, it is expressed as the following twelfth expression: If i > j, it is expressed as the following thirteenth expression: A first debiasing quantization module is used to perform debiasing quantization processing on the quantized addition kernel to obtain a target quantized addition kernel, and replace the initial addition kernel with the target quantized addition kernel in the initial full addition network to obtain a quantized full addition network to be trained; wherein, when the quantized addition kernel is debiased quantized to obtain the target quantized addition kernel, the first debiasing quantization module is used to: In the quantized additive kernel, if the weight scaling factor S F is greater than 1, then the weight scale factor S F The fourteenth expression is redefined as follows: j is re-expressed as the following fifteenth expression: Among them, F′ Me and F″ Me Represent the negative median and positive median in the weight distribution respectively; A second debiasing quantization module is used to perform debiasing quantization training on the quantized full-additive network to be trained based on a reference full-additive network model to obtain a trained target quantized full-additive network, wherein the debiasing quantization training uses remote sensing image samples as training samples and uses scene classification of remote sensing images as a downstream task; The recognition module is used to perform scene classification on any remote sensing image to be recognized based on the target quantization full-add network.

4. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to claim 1 or 2.

Citation Information

Patent Citations

  • Scene classification method and system based on Lie-Fisher remote sensing image

    CN111026897A

  • Airborne and satellite-borne remote sensing scene classification method based on high-precision full-plus network

    CN116363409A