Neural Network Quantization Method, Device, Equipment and Storage Medium
By analyzing the frequency domain and spatial distribution differences of tensors before and after quantization, and optimizing the quantization strategy, the problem of poor performance of the quantization strategy in the existing technology is solved, and the quantization effect with low resource consumption and little performance impact is achieved.
Patent Information
- Application Number
- CN202210102781.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-01-27
AI Technical Summary
When performing quantization processing of neural networks, it is difficult to minimize performance impact while reducing computing and storage resource consumption. The existing measurement methods are too single, resulting in poor performance of the quantization strategy.
By analyzing the frequency domain distribution differences of tensors before and after quantization, selecting the target quantization strategy from a variety of quantization strategies, optimizing the quantization strategy with frequency domain and spatial distribution differences, quantization perception training is used to adjust the quantization parameters to ensure the frequency domain distribution differences and performance losses are minimized.
It realizes the minimization of performance impact while reducing the consumption of neural network computing and storage resources, and improves the effectiveness and accuracy of quantitative strategies.
Smart Images

Figure CN114492792B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a neural network quantization method, apparatus, device, and storage medium. Background Art
[0002] Due to the complex structure of neural networks, which contain a large number of network parameters, a large amount of storage and computing resources are consumed during the calculation process. In order to deploy neural networks to ordinary terminal devices and expand their application scenarios, the tensors in neural networks can be quantized, converting the original high-precision tensors in neural networks into low-precision tensors. When quantizing the tensors in neural networks, it is necessary to optimize the quantization strategy so that the neural network quantized using the optimized quantization strategy consumes as little computing and storage resources as possible, while maintaining its performance without being greatly affected. Therefore, it is necessary to provide a solution for optimizing the quantization strategy. Summary of the Invention
[0003] The present disclosure provides a neural network quantization method, apparatus, device, and storage medium.
[0004] According to a first aspect of an embodiment of the present disclosure, a neural network quantization method is provided. The method includes:
[0005] Obtain a tensor set composed of tensors of a neural network;
[0006] Quantize the tensors in the tensor set using a plurality of quantization strategies respectively;
[0007] Based on the frequency domain distribution of the tensors before quantization and the frequency domain distribution of the tensors after quantization, determine the target quantization strategy for the tensor set from the plurality of quantization strategies, and use the result of quantizing the tensors in the tensor set using the target quantization strategy as the quantization result of the tensor set.
[0008] In some embodiments, the determining the target quantization strategy for the tensor set from the plurality of quantization strategies based on the frequency domain distribution of the tensors before quantization and the frequency domain distribution of the tensors after quantization includes:
[0009] For each quantization strategy in the plurality of quantization strategies, determine the frequency domain distribution difference of the tensor set based on the difference between the frequency domain distribution of the tensors before quantization and the frequency domain distribution of the tensors after quantization;
[0010] Determine the target quantization strategy for the tensor set from the plurality of quantization strategies based on the frequency domain distribution difference.
[0011] In some embodiments, there are multiple sets of tensors, and each set of tensors is composed of tensors of one layer of the neural network or adjacent multiple layers of the neural network;
[0012] Determining the target quantization strategy for the set of tensors from the multiple quantization strategies based on the frequency domain distribution difference includes:
[0013] Determining the target quantization strategy for each set of tensors based on the frequency domain distribution difference of each set of tensors; or
[0014] Determining the target quantization strategy for each set of tensors among the multiple sets of tensors based on the cumulative result of the frequency domain distribution differences of the multiple sets of tensors.
[0015] In some embodiments, determining the target quantization strategy for each set of tensors based on the frequency domain distribution difference of each set of tensors includes:
[0016] For each set of tensors, using the quantization strategy adopted when the frequency domain distribution difference of each set of tensors is minimized as the target quantization strategy for each set of tensors.
[0017] In some embodiments, determining the target quantization strategy for each set of tensors among the multiple sets of tensors based on the cumulative result of the frequency domain distribution differences of the multiple sets of tensors includes:
[0018] Determining a first loss based on the cumulative result of the frequency domain distribution differences of the multiple sets of tensors;
[0019] Determining a second loss based on the difference between the prediction result of the sample data output by the neural network and the label of the sample data;
[0020] Obtaining a first target loss based on the first loss and the second loss, and adjusting the quantization parameters in the quantization strategy until the first target loss meets a first preset condition;
[0021] Determining the quantization parameters when the first target loss meets the first preset condition as the quantization parameters of the target quantization strategy.
[0022] In some embodiments, determining the target quantization strategy for the set of tensors from the multiple quantization strategies based on the frequency domain distribution of the tensors before quantization and the frequency domain distribution of the tensors after quantization includes:
[0023] Determining the target quantization strategy for the set of tensors from the multiple quantization strategies based on the frequency domain distribution of the tensors before quantization and the frequency domain distribution of the tensors after quantization, and the spatial distribution of the tensors before quantization and the frequency domain distribution of the tensors after quantization.
[0024] In some implementations, determining the target quantization strategy for the tensor set from the multiple quantization strategies based on the frequency-domain distribution of the tensor before quantization processing and the frequency-domain distribution of the tensor after quantization processing, as well as the spatial distribution of the tensor before quantization processing and the frequency-domain distribution of the tensor after quantization processing, includes:
[0025] For each quantization strategy among the multiple quantization strategies, determine the frequency-domain distribution difference of the tensor set based on the difference between the frequency-domain distribution of the tensor before quantization processing and the frequency-domain distribution of the tensor after quantization processing; and determine the spatial distribution difference of the tensor set based on the difference between the spatial distribution of the tensor before quantization processing and the spatial distribution of the tensor after quantization processing;
[0026] Based on the frequency-domain distribution difference and the spatial distribution difference, determine the target quantization strategy for the tensor set from the multiple quantization strategies.
[0027] In some implementations, there are multiple tensor sets, and each tensor set is composed of tensors of one layer of the neural network or tensors of adjacent multiple layers of the neural network;
[0028] The determining the target quantization strategy for the tensor set from the multiple quantization strategies based on the frequency-domain distribution difference and the spatial distribution difference includes:
[0029] Determine the target quantization strategy for each tensor set based on the respective frequency-domain distribution difference and spatial distribution difference of each tensor set; or
[0030] Determine the target quantization strategy for each tensor set among the multiple tensor sets based on the cumulative result of the frequency-domain distribution differences of the multiple tensor sets and the cumulative result of the spatial distribution differences.
[0031] In some implementations, the determining the target quantization strategy for each tensor set based on the respective frequency-domain distribution difference and spatial distribution difference of each tensor set includes:
[0032] For each tensor set, determine the weighted average value of the frequency-domain distribution difference and the spatial distribution difference of the each tensor set;
[0033] Take the quantization strategy adopted when the weighted average value is the smallest as the target quantization strategy for each tensor set; or
[0034] For each tensor set, take the quantization strategy adopted when the frequency-domain distribution difference is less than a first threshold and the spatial distribution difference is less than a second threshold as the target quantization strategy for each tensor set.
[0035] In some embodiments, determining the target quantization strategy for each tensor set in the multiple tensor sets based on the cumulative result of the frequency-domain distribution differences and the cumulative result of the spatial distribution differences of the multiple tensor sets includes:
[0036] Based on the cumulative result of the frequency-domain distribution differences and the cumulative result of the spatial distribution differences of the multiple tensor sets, determine a third loss;
[0037] Based on the difference between the prediction result of the sample data output by the neural network and the label of the sample data, determine a fourth loss;
[0038] Obtain a second target loss based on the third loss and the fourth loss, and adjust the quantization parameters in the quantization strategy based on the second target loss until the second target loss meets a second preset condition;
[0039] Take the quantization parameters when the second target loss meets the second preset condition as the quantization parameters of the target quantization strategy.
[0040] In some embodiments, the frequency-domain distribution of each tensor before quantization is determined in the following manner:
[0041] Extract the frequency-domain features of the tensor before quantization to obtain the frequency-domain distribution of the tensor before quantization; and / or
[0042] The frequency-domain distribution of each tensor after quantization is determined in the following manner:
[0043] Extract the frequency-domain features of the tensor after quantization to obtain the frequency-domain distribution of the tensor after quantization.
[0044] In some embodiments, the dimension of each tensor is greater than or equal to 2. Extracting the frequency-domain features of the tensor before quantization includes:
[0045] Convert the tensor before quantization into a one-dimensional tensor;
[0046] Extract the frequency-domain features of the obtained one-dimensional tensor to obtain the frequency-domain distribution of the tensor before quantization; and / or
[0047] Extracting the frequency-domain features of the tensor after quantization includes:
[0048] Convert the tensor after quantization into a one-dimensional tensor;
[0049] Extract the frequency-domain features of the obtained one-dimensional tensor to obtain the frequency-domain distribution of the tensor after quantization.
[0050] In some embodiments, the multiple quantization strategies include:
[0051] Multiple quantization strategies with different quantization types; and / or
[0052] Multiple quantization strategies with the same quantization type but different quantization parameters.
[0053] In some embodiments, the quantization type includes: uniform quantization and / or non-uniform quantization; the quantization parameters include: quantization step size and / or zero value point.
[0054] According to a second aspect of the embodiments of the present disclosure, there is provided a neural network quantization device, the device includes:
[0055] An acquisition module, configured to acquire a set of tensors composed of tensors of a neural network;
[0056] A quantization processing module, configured to perform quantization processing on the tensors in the set of tensors by using multiple quantization strategies respectively;
[0057] A quantization strategy determination module, configured to determine a target quantization strategy for the set of tensors from the multiple quantization strategies based on the frequency domain distribution of the tensors before quantization processing and the frequency domain distribution of the tensors after quantization processing, and use the result of quantizing the tensors in the set of tensors by using the target quantization strategy as the quantization result of the set of tensors. According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, the electronic device includes a processor, a memory, and computer instructions stored in the memory and executable by the processor. When the processor executes the computer instructions, the method mentioned in the first aspect above can be implemented.
[0058] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are executed, the method mentioned in the first aspect above is implemented.
[0059] In the embodiments of the present disclosure, for a set of tensors composed of tensors in a neural network, when determining the quantization strategy of the set of tensors, a target quantization strategy can be selected from multiple quantization strategies based on the frequency domain distribution of each tensor in the set of tensors before and after quantization, so as to quantize the tensors in the neural network by using the target quantization strategy. By using the frequency domain distribution of the tensors before and after quantization to measure the quality of the quantization strategy, the quantization strategy can be measured more carefully from the frequency domain dimension, so that the neural network after quantization occupies less resources and its performance is not affected too much.
[0060] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. Description of the Drawings
[0061] The accompanying drawings here are incorporated into the specification and form a part of this specification. These drawings show embodiments consistent with the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure.
[0062] Figure 1 It is a schematic diagram of quantifying a neural network according to an embodiment of the present disclosure.
[0063] Figure 2 It is a flowchart of a method for quantifying a neural network according to an embodiment of the present disclosure.
[0064] Figure 3 It is a schematic diagram of quantifying a neural network based on the difference in the frequency-domain distribution of tensors before and after quantization and the difference between the output result and the true result of the neural network according to an embodiment of the present disclosure.
[0065] Figure 4 It is a schematic diagram of the logical structure of a neural network quantization device according to an embodiment of the present disclosure.
[0066] Figure 5 It is a schematic diagram of the logical structure of a device according to an embodiment of the present disclosure. Detailed Embodiments
[0067] Here, exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0068] The terms used in the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit the present disclosure. The singular forms "a", "the", and "said" used in the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items. Additionally, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality.
[0069] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0070] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of this disclosure and make the above-mentioned objects, features, and advantages of the embodiments of this disclosure more obvious and understandable, the technical solutions in the embodiments of this disclosure will be further described in detail below with reference to the accompanying drawings.
[0071] Neural networks are widely used in various fields. In order to enable neural networks to have higher precision and accuracy, when designing neural networks, their network structures are becoming increasingly complex, and thus the number of network parameters included is also extremely large, which results in neural networks consuming a large amount of storage and computing resources during the calculation process, posing a huge challenge to the devices on which neural networks are deployed.
[0072] In order to enable neural networks to be deployed on ordinary terminal devices so that neural networks can be applied in more scenarios, the tensors in neural networks can be quantized. Quantization processing is to convert the original high-precision tensors in neural networks into low-precision tensors. For example, Figure 1 As shown, it is a schematic diagram of quantizing the tensors of a neural network. The original float32 (32-bit floating-point number) tensors in the neural network (such as the input data and weights in the figure) can be converted into int 8 (8-bit integers), and then corresponding calculations and storage can be performed. By using fewer bits to represent the tensors in the neural network, the storage space occupied by these tensors can be reduced, and the calculation speed can also be accelerated.
[0073] In order to enable the quantized neural network to reduce the consumption of computing and storage resources while maintaining its performance without being greatly affected, the quantization strategy can be optimized. For example, the quantization type can be adjusted, or the quantization parameters can be changed to obtain the optimal quantization strategy. Usually, after quantizing the tensors in the tensor set, the original data distribution will change. Therefore, measuring the data distribution difference before and after quantization can guide the quality of quantization to a certain extent. When optimizing the quantization strategy, an easy-to-think way is to measure the quality of the quantization strategy based on the data difference or spatial distribution difference of the tensors before and after quantization. This measurement method is too single, and the quantization effect of the determined quantization strategy also needs to be improved.
[0074] Based on this, embodiments of the present disclosure provide a quantization method for a neural network. For a tensor set composed of tensors in a neural network, when determining the quantization strategy for the tensor set, the target quantization strategy can be selected from multiple quantization strategies based on the frequency domain distributions of the tensors in the tensor set before and after quantization, so as to use the target quantization strategy to quantize the tensors in the neural network. By using the frequency domain distributions of the tensors before and after quantization to measure the quality of the quantization strategy, the quantization strategy can be measured more meticulously from the frequency domain dimension, so that the quantized neural network occupies less resources and its performance is not greatly affected.
[0075] The neural network provided by the embodiments of the present disclosure can be executed by various devices. For example, cloud servers, laptop computers, various user terminals, etc. For example, the neural network can be used for image classification, speech recognition, etc. Therefore, the device executing this method can be a mobile phone or a computer on which the neural network is deployed. After quantizing the neural network, the occupation of computing resources and storage resources of the terminal by the neural network can be reduced, and then the terminal can use the deployed neural network for various applications such as image recognition and speech recognition.
[0076] In the embodiments of the present disclosure, the tensor can be a network parameter of the neural network. For example, it can be a weight, a bias, a convolution kernel, etc. It can also be the input data in the neural network. For example, the input image data, speech data, etc. It can also be the intermediate operation result of the neural network. For example, the output data of each layer. For example, it can be an activation, an extracted feature map, etc. The embodiments of the present disclosure do not make any restrictions.
[0077] The quantization strategy in the embodiments of the present disclosure can be some quantization schemes for quantizing tensors. For example, the quantization strategy can include which quantization type to use to quantize the tensor. For example, the quantization type can be uniform quantization or non-uniform quantization. And under each quantization type, which quantization parameters to use to quantize the tensor. For example, when performing uniform quantization, how to set the quantization step size and the zero point value, etc.
[0078] Specifically, as Figure 2 shown, the neural network quantization method provided by the embodiments of the present disclosure can include the following steps:
[0079] S202, obtain a tensor set composed of tensors of the neural network;
[0080] In step S202, when quantifying the tensors in the neural network, it can be in units of a tensor set composed of multiple tensors, and each tensor in the tensor set is quantified. For example, for each tensor set, a quantization strategy is determined to uniformly quantify each tensor in the tensor set using this quantization strategy. Therefore, a tensor set composed of the tensors of the neural network can be obtained, where the tensor set can be one or more. For example, the tensors of each channel in the neural network can be used as a tensor set, or the tensors of each layer or multiple layers in the neural network can be used as a tensor set, or all the tensors in the neural network can be used as a tensor set, which can be flexibly set according to actual needs.
[0081] S204. Quantize the tensors in the tensor set using multiple quantization strategies respectively;
[0082] In step S204, after obtaining the tensor set, the tensors in the tensor set can be quantized using multiple quantization strategies respectively. Among them, the multiple quantization strategies can be preset or generated in real time during the quantization process. For example, multiple quantization strategies can be preset, and these quantization strategies can include different quantization types and multiple groups of quantization parameters under each quantization type, and then the tensors are quantized using these preset quantization strategies respectively. Or, the appropriate quantization type can be automatically determined based on the distribution of the tensors in the tensor set, and then the initial quantization parameters are determined based on the minimum and maximum values of the tensors in the tensor set, and then the quantization parameters are continuously adjusted according to a certain step size to obtain multiple quantization strategies. Or, a fixed quantization type can be preset, and then the quantization parameters are automatically changed to obtain multiple quantization strategies.
[0083] S206. Based on the frequency domain distribution of the tensors before quantization and the frequency domain distribution of the tensors after quantization, determine the target quantization strategy for the tensor set from the multiple quantization strategies, and use the result of quantizing the tensors in the tensor set using the target quantization strategy as the quantization result of the tensor set.
[0084] In step S206, after quantizing the tensors in the tensor set using multiple quantization strategies respectively, the frequency-domain distribution of each tensor before quantization processing can be determined, as well as the frequency-domain distribution of each tensor after quantization processing using each quantization strategy. Then, based on the frequency-domain distribution of each tensor before and after quantization processing, the target quantization strategy can be determined from multiple quantization strategies. Then, the result after quantizing the tensors in the tensor set using the target quantization strategy can be used as the final quantization result of the tensor set, so as to deploy the quantized neural network to a terminal device or a chip for application, so as to reduce the computing resources, storage resources, and data transmission resources required by the neural network during the execution of processing tasks, reduce the requirements for the operating environment required for the deployment of the neural network, and improve the processing efficiency of the terminal device or the chip.
[0085] Among them, the frequency-domain distribution refers to the distribution of the characteristics of the tensor in the frequency domain. Taking a tensor as an input image in the neural network as an example (the corresponding tensor set can be composed of all or part of multiple image feature tensors and / or neural network-related parameters when quantizing the image-related tensors in the neural network), generally, an image includes a high-frequency part (the part where the color changes violently in the image, such as the contour and edge of the image) and a low-frequency part (the part where the color changes slowly in the image, such as the background in the image). The frequency-domain distribution of this image characterizes the intensity of signals with different frequencies (such as high-frequency signals and low-frequency signals) in the image. Similarly, for other tensors, such as weights, feature maps, biases, etc., which are usually matrices and are similar to image data, their frequency-domain distributions also represent similar meanings.
[0086] Generally, we hope that the differences in the frequency-domain distributions of the tensors in the tensor set before and after quantization processing are as small as possible. Therefore, after obtaining the frequency-domain distributions of the tensors before and after quantization processing, the similarity of the frequency-domain distributions before and after quantization processing can be statistically analyzed, or the deviation degree of the frequency-domain distributions before and after quantization processing can be statistically analyzed, etc. Then, based on the statistical results, the optimal quantization strategy can be selected from multiple quantization strategies as the target quantization strategy.
[0087] In some embodiments, the multiple quantization strategies can be multiple quantization strategies with different quantization types. For example, the quantization types can include uniform quantization and non-uniform quantization. Uniform quantization can be further divided into different types such as symmetric quantization and non-symmetric quantization. In some implementations, the multiple quantization strategies can be multiple quantization strategies with the same quantization type but different quantization parameters. For example, the quantization parameters can be one or more of the zero point value and the quantization step size, and the categories of the quantization parameters vary with different quantization types.
[0088] In some embodiments, when determining the frequency-domain distribution of each tensor in the pre-quantization tensor set, the frequency-domain characteristics of the tensor before quantization can be extracted to obtain the frequency-domain distribution of the tensor before quantization. Similarly, for each tensor after quantization processing, its frequency-domain characteristics can also be extracted to obtain the frequency-domain distribution of each tensor after quantization. Among them, when extracting the frequency-domain characteristics of the tensors before and after quantization, various frequency-domain feature extraction algorithms can be used. For example, the Fourier transform can be used to extract frequency-domain characteristics, or the discrete cosine transform (DCT) can also be used to extract frequency-domain characteristics. It is not difficult to understand that all methods that can extract the frequency-domain characteristics of data are applicable, and the embodiments of the present disclosure do not make any limitations.
[0089] In some embodiments, the dimension of the tensors in the tensor set may be greater than or equal to 2. For multi-dimensional tensors, when performing frequency-domain feature extraction on the tensors before or after quantization, in order to improve the processing efficiency, the tensor can be first dimension-reduced. For example, the tensor can be converted into a one-dimensional tensor, and then the frequency-domain distribution of the tensor can be obtained by performing feature extraction on the one-dimensional tensor.
[0090] For example, taking a convolutional neural network as an example, the quantization of the convolutional neural network is mainly performed on the fully connected layer and the convolutional layer. Among them, the parameter dimension of the fully connected layer is two-dimensional and can be expressed as using W fc (H f ,W f ). The parameters of the convolutional layer are four-dimensional and can be expressed as W conv (Co,Ci,H c ,W c ), where H f and W f respectively represent the height and width of the weight matrix of the fully connected layer, and Co, Ci, H c ,W c respectively correspond to the output channel, input channel, convolutional kernel height, and convolutional kernel width. In addition, high-precision feature maps consume a large amount of computing resources in model calculations, so the quantization of feature maps is also necessary. Feature maps are usually four-dimensional tensor signals and can be expressed as A(N,C a ,H a ,W a ), where N represents the input batch size, C a corresponds to the number of channels of the feature map, and H a ,W a respectively represent the height and width of the feature map. Since the dimensions of the parameters and the feature maps are both multi-dimensional, they can be first converted into one-dimensional data, and then the Fourier transform or the discrete cosine transform (DCT) can be used to extract the frequency-domain characteristics of the one-dimensional data, so as to obtain the corresponding frequency-domain distribution.
[0091] Of course, in some embodiments, for a multi-dimensional tensor, in order to obtain a more accurate frequency domain distribution, feature extraction can also be directly performed on the multi-dimensional tensor, and specific selection can be flexibly made according to actual requirements. That is, if a more accurate frequency domain distribution is desired, direct extraction is performed; if the processing efficiency is to be improved, it is converted into a one-dimensional tensor and then extracted.
[0092] In some embodiments, when determining the target quantization strategy of a tensor set from multiple quantization strategies based on the frequency domain distributions of tensors before and after quantization processing, for each quantization strategy, the difference between the frequency domain distribution of each tensor in the tensor set after quantization processing using this quantization strategy and the frequency domain distribution of each tensor before quantization processing can be statistically calculated to obtain the frequency domain distribution difference of the tensor set. Among them, this frequency domain distribution difference can describe the overall deviation situation of the frequency domain distributions of all tensors in the tensor set before and after quantization. For example, the standard deviation of the deviation between the frequency domain distribution of each tensor in the tensor set after quantization and the frequency domain distribution before quantization can be statistically calculated, and this standard deviation can be used as the frequency domain distribution difference of the tensor set, or the KL divergence algorithm can also be used to statistically calculate the deviation between the frequency domain distribution of each tensor in the tensor set after quantization and the frequency domain distribution before quantization as the frequency domain distribution difference of the tensor set. It is not difficult to understand that various methods for statistically calculating the distribution difference of data in a dataset can be used to determine the frequency domain distribution difference of the tensor set in the embodiments of the present disclosure, and the embodiments of the present disclosure do not make any restrictions. After determining the frequency domain distribution difference of the tensor set, the target quantization strategy of the tensor set can be determined from multiple quantization strategies using this frequency domain distribution difference as a measurement criterion.
[0093] In some embodiments, multiple tensor sets can be obtained, and each tensor set is composed of tensors of one layer of a neural network or adjacent multiple layers of a neural network. For example, the tensors (such as weights, biases, inputs, and outputs) in each layer of a neural network can be used to form a tensor set, or the tensors of adjacent multiple layers of a neural network can be used to form a tensor set. For example, the tensors of every adjacent three layers of a neural network can be used to form a tensor set.
[0094] In some embodiments, after obtaining multiple tensor sets, the target quantization strategy of each tensor set can be determined based on the frequency domain distribution difference of each tensor set. For example, the quantization strategy of each tensor set can be optimized separately. When determining the target quantization strategy of this tensor set, only the frequency domain distribution difference of this tensor set needs to be considered. For example, in some embodiments, for each tensor set, the quantization strategy adopted when the frequency domain distribution difference of this tensor set is the smallest can be used as the target quantization strategy of this tensor set. Of course, when determining the target quantization strategy, the embodiments of the present disclosure are not limited to the quantization strategy adopted when the frequency domain distribution difference is the smallest. For example, some other factors can also be combined for joint screening. For example, the quantization strategy adopted when the frequency domain distribution difference is less than a certain value and other factors also meet the preset conditions can be used as the target quantization strategy.
[0095] In some embodiments, the target quantization strategy for each tensor set may also be determined based on the cumulative result of the frequency domain distribution differences of the multiple tensor sets. When determining the quantization strategy for each tensor set, the frequency domain distribution differences of other tensor sets may be jointly considered to optimize the quantization strategies of all tensor sets in the neural network as a whole. For example, when the cumulative result of the frequency domain distribution differences of all tensor sets is minimized, the quantization strategy adopted by each tensor set may be used as the target quantization strategy for that tensor set. Herein, the cumulative result may be the sum of the frequency domain distribution differences of all tensor sets, or may also be the product result. The embodiments of the present disclosure do not make any limitations thereto.
[0096] When quantizing a neural network, it is generally divided into two methods: post-training quantization (PTQ) and quantization-aware training (QAT). Among them, post-training quantization means that during the training process, a raw model is first trained based on high-precision calculations. In the inference stage, the high-precision raw model can be quantized to low precision to reduce the computing and storage space occupied by the model, facilitating the deployment of the model to various terminal devices. For example, assume that the precision of the model is 32-bit floating point numbers. That is, when training the model, the input data (such as voice, image, etc.), the parameters of the model, and the intermediate data during the training process are all represented by 32-bit floating point numbers, and then corresponding calculations are performed to train the model to obtain a trained high-precision raw model (at this time, the parameters of the model are still 32-bit floating point numbers). Then, the trained raw model with float32 precision can be quantized to an int8 (as Figure 1 shown) model. For example, the model parameters of the raw model are 32-bit floating point numbers, and then a certain quantization strategy can be used to quantize the model parameters into 8-bit integers, reducing the computing and storage space occupied by the model. Of course, this quantization process can be performed before using the model for inference (only quantizing the model parameters at this time), or can also be performed during the process of using the model for inference (at this time, the model input, model parameters, and intermediate data can be quantized simultaneously). The embodiments of the present disclosure do not make any limitations thereto.
[0097] Quantization-aware training is to introduce the fake-quantization operation while training the model, that is, to quantize the tensors of the model during the training process. For example, during the forward propagation stage of the training process, the input data, weights, and activation values can be quantized, and during the backpropagation stage, the gradients can also be quantized, and then the original parameters are updated. Of course, in order to make the trained model have higher accuracy, the model can be first trained with high-precision data to make the model have a better effect, and then the quantization-aware training method can be used to train the model. When performing quantization-aware training, in some scenarios, the network parameters of the neural network can be kept unchanged, and only the quantization parameters are optimized, that is, the original values of the network parameters remain unchanged, only the quantization parameters are optimized, and the neural network is trained by continuously adjusting the quantization parameters to obtain a better quantization strategy. In some scenarios, the network parameters and quantization parameters can also be optimized simultaneously, that is, the parameter values of the network parameters and the quantization parameters can be adjusted simultaneously to train the neural network to obtain a better quantization strategy. However, this method has a large amount of calculation and a large training difficulty. Therefore, in some embodiments, when determining the target quantization strategy of each tensor set by combining the cumulative results of the frequency domain distribution differences of multiple tensor sets, the quantization-aware training method can also be used, that is, the first loss can be determined based on the cumulative results of the frequency domain distributions of each tensor set, and then the second loss can be determined based on the difference between the prediction result of the sample data output by the neural network and the sample data label. Combining the first loss and the second loss can obtain a total target loss, hereinafter referred to as the first target loss. Then, the first target loss can be used as the optimization target, and the quantization parameters in the quantization strategies of each tensor set are continuously adjusted until the first target loss meets the first preset condition. Then, the quantization strategy composed of the quantization parameters adopted by each tensor set when the target loss meets the first preset condition can be determined as the target quantization strategy of the tensor set. Among them, the first preset condition can be that the target loss converges, or the target loss is less than the preset threshold, which can be specifically set according to actual needs.
[0098] For example, such as Figure 3As shown, assuming that the neural network is a network for classifying images, images carrying class labels can be used as the input of the neural network, and the prediction results of the class to which the image belongs can be output through the neural network. In this process, the input data, network parameters, and output of the neural network can be quantized before corresponding calculations are performed to obtain the final prediction result of the neural network. Taking the tensors of each layer of the neural network (including the input data, weights, biases, and activations of each layer) as a tensor set as an example, for each tensor set of the neural network, a quantization type can be preset. For example, the same quantization type can be uniformly adopted for all tensor sets of the neural network, or different quantization types can be set for each tensor set according to actual needs. Then, for each tensor set, an initial quantization parameter can be determined respectively according to the data distribution in each tensor set, and the tensors of the tensor set can be quantized using the initial quantization parameter, and then corresponding calculations can be performed to obtain the output of each layer. At the same time, the frequency domain distribution difference of the tensor set under the initial quantization parameter can be determined, and then the frequency domain distribution differences of each tensor set can be accumulated to obtain the total frequency domain distribution difference of the tensors in the entire neural network before and after quantization as the first loss. At the same time, the difference between the predicted result and the true result of the class of the image output by the neural network under the initial quantization parameter can be combined to obtain the second loss, and then the first target loss can be determined based on the first loss and the second loss. For example, the two can be added together, or multiplied by a certain weight and then added together to obtain the first target loss. Then, based on the first target loss, backpropagation can be performed. For example, the quantization parameters of each tensor set can be continuously adjusted according to the preset gradient to obtain the updated quantization strategy of each tensor set. Then, the tensors of each layer of the neural network can be quantized using the updated quantization strategy, and then the above process can be repeated to obtain the first target loss under the updated quantization strategy until the first target loss converges. Then, the quantization strategy composed of the quantization parameters adopted by each tensor set at this time is used as the target quantization strategy of each tensor set. By determining the total loss based on the cumulative result of the frequency domain distribution differences of the tensors of each layer before and after quantization, and the difference between the predicted result and the true result output by the neural network, and then optimizing the quantization parameters using the total loss, the finally quantized neural network can be obtained. At the same time, since the frequency domain distribution of the tensors in the neural network before and after quantization and the accuracy of the prediction result of the neural network are comprehensively considered during the process of optimizing the quantization strategy, an appropriate target quantization strategy can be obtained as much as possible, and the neural network quantized using the target quantization strategy can occupy as little storage and computing resources as possible, and have as little impact on the accuracy of the neural network as possible.
[0099] If only the frequency-domain distributions of the tensors in the tensor set before and after quantization processing are used to measure the pros and cons of the quantization strategy, it may still not be comprehensive enough. Therefore, in some embodiments, in order to make the determined quantization strategy more accurate, the frequency-domain distribution and the spatial distribution of the tensors before and after quantization processing can also be combined to optimize the quantization strategy. For example, the frequency-domain distribution of each tensor in the tensor set before quantization processing can be combined with the frequency-domain distribution of each tensor after quantization processing, and the spatial distribution of each tensor before quantization processing can be combined with the spatial distribution of each tensor after quantization processing to determine the target quantization strategy of the tensor set from multiple quantization strategies. Among them, the spatial distribution refers to the numerical distribution of the tensors before and after quantization, that is, the distribution of the numerical values of the tensors in the data space. The frequency-domain distribution is the distribution of the signal intensities of the tensors before and after quantization at different frequencies. By combining the frequency-domain distribution and the spatial distribution to screen the quantization strategy, the quantization strategy can be measured from multiple dimensions to obtain a more accurate quantization result. When screening the target quantization strategy based on the frequency-domain distribution and the spatial distribution, the overall principle is that the difference in the frequency-domain distribution before and after quantization is as small as possible, and the difference in the spatial distribution is also as small as possible.
[0100] In some embodiments, for each quantization strategy, the frequency-domain distribution difference of the tensor set can be determined based on the difference between the frequency-domain distribution of each tensor in the tensor set before quantization processing and the frequency-domain distribution of each tensor after quantization processing. Similarly, the spatial distribution difference of the tensor set can be determined based on the difference between the spatial distribution of the tensors in each tensor set before quantization and the spatial distribution of the tensors in the tensor set after quantization. Among them, the spatial distribution difference can describe the overall deviation of the spatial distributions of all tensors in the tensor set before and after quantization. The spatial distribution difference can be determined by the KL divergence algorithm, or the root mean square error of the tensors before and after quantization can also be calculated as the spatial distribution difference, and it can be specifically selected based on the actual situation. Then, based on the frequency-domain distribution difference and the spatial distribution difference, the target quantization strategy of the tensor set can be determined from multiple quantization strategies.
[0101] In some embodiments, there can be multiple such tensor sets, and each tensor set is composed of the tensors of one layer of the neural network or adjacent multiple layers of the network; when determining the quantization strategies of these multiple tensor sets, the quantization strategy of each tensor set can be optimized separately, that is, the target quantization strategy of each tensor set can be determined based on the frequency-domain distribution difference and the spatial distribution difference of each tensor set respectively.
[0102] For example, in some embodiments, for each tensor set, the weighted average value of the frequency-domain distribution difference and the spatial distribution difference of the tensor set can be determined, where the weights of the frequency-domain distribution difference and the spatial distribution difference can be flexibly set based on the actual scenario. Then, the quantization strategy adopted when the weighted average value is the smallest is used as the target quantization strategy of each tensor set.
[0103] In some embodiments, a threshold may also be set in advance for each of the frequency-domain distribution difference and the spatial distribution difference, and then the quantization strategy adopted when both meet the corresponding conditions may be used as the target quantization strategy for this tensor set. For example, for each tensor set, the quantization strategy adopted when the frequency-domain distribution difference is less than a first threshold and the spatial distribution difference is less than a second threshold may be used as the target quantization strategy for this tensor set. By determining the target quantization strategy by comprehensively considering the differences in both the frequency-domain distribution and the spatial distribution, a more accurate and appropriate quantization strategy can be obtained.
[0104] In some embodiments, the quantization strategies of multiple tensor sets may also be optimized as a whole by combining the frequency-domain distribution differences and the spatial distribution differences of the multiple tensor sets. For example, the target quantization strategy for each tensor set may be determined by combining the cumulative result of the frequency-domain distribution differences of the multiple tensor sets and the cumulative result of the spatial distribution differences of the multiple tensor sets. Among them, the cumulative result of the frequency-domain distribution differences may be obtained by adding or multiplying the frequency-domain distributions of each tensor set, and the cumulative result of the spatial distribution differences may be obtained by adding or multiplying the spatial distributions of each tensor set.
[0105] In some embodiments, a third loss may be determined based on the cumulative result of the frequency-domain distribution differences of multiple tensor sets and the cumulative result of the spatial distribution differences. For example, the sum of the cumulative result of the frequency-domain distribution differences and the cumulative result of the spatial distribution differences may be used as the third loss, or the larger one of the cumulative result of the frequency-domain distribution differences or the cumulative result of the spatial distribution differences may be used as the third loss, or the weighted average of the cumulative result of the frequency-domain distribution differences or the cumulative result of the spatial distribution differences may be used as the third loss. Then, a fourth loss may be determined based on the difference between the prediction result of the sample data output by the neural network and the sample data label. By combining the third loss and the fourth loss, a total target loss, hereinafter referred to as the second target loss, can be obtained. Then, with the second target loss as the optimization target, the quantization parameters in the quantization strategies of each tensor set may be continuously adjusted until the second target loss meets a second preset condition. For example, the second preset condition may be that the second target loss converges or is less than a preset threshold. Then, the quantization strategy composed of the quantization parameters adopted by each tensor set when the second target loss meets the second preset condition may be determined as the target quantization strategy for this tensor set.
[0106] By combining the frequency-domain distribution and the spatial distribution to optimize the quantization parameters of each layer of the neural network as a whole, a more appropriate target quantization strategy can be obtained, so that the neural network quantized using the target quantization strategy occupies as little storage and computing resources as possible, and the performance is not greatly affected.
[0107] In some embodiments, the neural network may be a network for various prediction processes such as image classification and recognition. When deploying the neural network to a terminal device, quantization processing can be performed on the neural network. For example, some image data with labels can be input into the neural network, and then different quantization strategies can be used to perform quantization processing on the input image data, the parameters of each layer of the neural network, and the output of each layer. Then, the quantization strategy can be adjusted by combining the frequency-domain distribution differences before and after quantization and the differences between the prediction results of the images output by the neural network and the labels to obtain the optimal quantization strategy. When performing various processes such as image classification and recognition using the neural network subsequently, various applications such as image classification and recognition can be performed based on the model parameters obtained by quantizing according to the optimal quantization strategy, thereby reducing the occupancy of computing and storage resources of the neural network on the terminal device and improving the processing efficiency.
[0108] Among them, it is not difficult to understand that the solutions described in the above embodiments can be combined in the case of no conflict, and they are not listed one by one in the embodiments of the present disclosure.
[0109] Correspondingly, the embodiments of the present disclosure also provide a neural network quantization device, as Figure 4 shown. The device 40 includes:
[0110] An acquisition module 41, configured to acquire a tensor set composed of tensors of a neural network;
[0111] A quantization processing module 42, configured to perform quantization processing on the tensors in the tensor set by using a plurality of quantization strategies respectively;
[0112] A quantization strategy determination module 43, configured to determine a target quantization strategy for the tensor set from the plurality of quantization strategies based on the frequency-domain distribution of the tensors before quantization processing and the frequency-domain distribution of the tensors after quantization processing, and use the result of performing quantization processing on the tensors in the tensor set by using the target quantization strategy as the quantization result of the tensor set.
[0113] Among them, the specific steps for the above device to execute the neural network quantization method can refer to the description in the above method embodiments and will not be elaborated here.
[0114] Furthermore, the embodiments of the present disclosure also provide a device, as Figure 5 shown. The device includes a processor, a memory, and computer instructions stored in the memory and executable by the processor. When the processor executes the computer instructions, the method described in any one of the above embodiments is implemented.
[0115] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in any one of the foregoing embodiments is implemented.
[0116] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0117] From the description of the above embodiments, those skilled in the art can clearly understand that the embodiments of this specification can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions of the embodiments of this specification, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of this specification.
[0118] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, laptop computer, cellular phone, camera phone, smart phone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or a combination of any several of these devices.
[0119] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. When implementing the solutions of the embodiments of this specification, the functions of the modules can be realized in one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solutions of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0120] The above is only the specific implementation manners of the embodiments of this specification. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the embodiments of this specification, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of this specification.
Claims
1. A neural network quantization method, characterized in that, The neural network is an image classification network, and the method includes: Obtaining a tensor set composed of tensors of the neural network, where the tensor set at least includes an image feature tensor of an image to be classified and network parameters of the neural network; Performing quantization processing on the tensors in the tensor set respectively by using a plurality of quantization strategies; Based on the frequency domain distribution of the tensors before quantization processing and the frequency domain distribution of the tensors after quantization processing, determining a target quantization strategy for the tensor set from the plurality of quantization strategies, and using the result of quantizing the tensors in the tensor set by using the target quantization strategy as the quantization result of the tensor set, so as to obtain a quantized neural network, and classifying the image to be classified by using the quantized neural network; where the frequency domain distribution of the image feature tensor represents the intensities of high-frequency signals and low-frequency signals in the image to be classified.
2. The method according to claim 1, characterized in that The determining the target quantization strategy for the tensor set from the plurality of quantization strategies based on the frequency domain distribution of the tensors before quantization processing and the frequency domain distribution of the tensors after quantization processing includes: For each quantization strategy in the plurality of quantization strategies, determining the frequency domain distribution difference of the tensor set based on the difference between the frequency domain distribution of the tensors before quantization processing and the frequency domain distribution of the tensors after quantization processing; Determining the target quantization strategy for the tensor set from the plurality of quantization strategies based on the frequency domain distribution difference.
3. The method according to claim 2, wherein There are a plurality of tensor sets, and each tensor set is composed of tensors of one layer of the network or adjacent multiple layers of the network in the neural network; The determining the target quantization strategy for the tensor set from the plurality of quantization strategies based on the frequency domain distribution difference includes: Determining the target quantization strategy for each tensor set based on the respective frequency domain distribution differences of each tensor set; Or Determining the target quantization strategy for each tensor set in the plurality of tensor sets based on the cumulative result of the frequency domain distribution differences of the plurality of tensor sets.
4. The method according to claim 3, wherein The determining the target quantization strategy for each tensor set based on the respective frequency domain distribution differences of each tensor set includes: For each tensor set, using the quantization strategy adopted when the frequency domain distribution difference of each tensor set is the smallest as the target quantization strategy for each tensor set.
5. The method according to claim 3, characterized in that, The determining the target quantization strategy for each tensor set in the plurality of tensor sets based on the cumulative result of the frequency domain distribution differences of the plurality of tensor sets includes: Determining a first loss based on the cumulative result of the frequency domain distribution differences of the plurality of tensor sets; Determining a second loss based on the difference between the prediction result of the sample data output by the neural network and the label of the sample data; Obtaining a first target loss based on the first loss and the second loss, and adjusting the quantization parameters in the quantization strategy until the first target loss meets a first preset condition; Determining the quantization parameters when the first target loss meets the first preset condition as the quantization parameters of the target quantization strategy.
6. The method according to any one of claims 1 to 5, characterized in that, The determining the target quantization strategy for the tensor set from the plurality of quantization strategies based on the frequency domain distribution of the tensors before quantization processing and the frequency domain distribution of the tensors after quantization processing includes: Determine the target quantization strategy for the tensor set from the multiple quantization strategies based on the frequency-domain distribution of the tensor before quantization processing and the frequency-domain distribution of the tensor after quantization processing, as well as the spatial distribution of the tensor before quantization processing and the frequency-domain distribution of the tensor after quantization processing.
7. The method according to claim 6, wherein The determining the target quantization strategy for the tensor set from the multiple quantization strategies based on the frequency-domain distribution of the tensor before quantization processing and the frequency-domain distribution of the tensor after quantization processing, as well as the spatial distribution of the tensor before quantization processing and the frequency-domain distribution of the tensor after quantization processing, includes: For each quantization strategy among the multiple quantization strategies, determine the frequency-domain distribution difference of the tensor set based on the difference between the frequency-domain distribution of the tensor before quantization processing and the frequency-domain distribution of the tensor after quantization processing; and determine the spatial distribution difference of the tensor set based on the difference between the spatial distribution of the tensor before quantization processing and the spatial distribution of the tensor after quantization processing; Determine the target quantization strategy for the tensor set from the multiple quantization strategies based on the frequency-domain distribution difference and the spatial distribution difference.
8. The method according to claim 7, wherein There are multiple tensor sets, and each tensor set is composed of tensors of one layer of the neural network or adjacent multiple layers of the neural network; The determining the target quantization strategy for the tensor set from the multiple quantization strategies based on the frequency-domain distribution difference and the spatial distribution difference, includes: Determine the target quantization strategy for each tensor set based on the respective frequency-domain distribution difference and spatial distribution difference of each tensor set; or Determine the target quantization strategy for each tensor set among the multiple tensor sets based on the cumulative result of the frequency-domain distribution differences of the multiple tensor sets and the cumulative result of the spatial distribution differences.
9. The method according to claim 8, wherein The determining the target quantization strategy for each tensor set based on the respective frequency-domain distribution difference and spatial distribution difference of each tensor set, includes: For each tensor set, determine the weighted average value of the frequency-domain distribution difference and the spatial distribution difference of the each tensor set; Take the quantization strategy when the weighted average value is the smallest as the target quantization strategy for each tensor set; or For each tensor set, take the quantization strategy when the frequency-domain distribution difference is less than the first threshold and the spatial distribution difference is less than the second threshold as the target quantization strategy for each tensor set.
10. The method according to claim 8, wherein The determining the target quantization strategy for each tensor set among the multiple tensor sets based on the cumulative result of the frequency-domain distribution differences of the multiple tensor sets and the cumulative result of the spatial distribution differences, includes: Determine the third loss based on the cumulative result of the frequency-domain distribution differences of the multiple tensor sets and the cumulative result of the spatial distribution differences; Determine the fourth loss based on the difference between the prediction result of the sample data output by the neural network and the label of the sample data; Obtain the second target loss based on the third loss and the fourth loss, and adjust the quantization parameters in the quantization strategy based on the second target loss until the second target loss meets the second preset condition; Use the quantization parameter when the second target loss meets the second preset condition as the quantization parameter of the target quantization strategy.
11. The method according to any one of claims 1-5, characterized in that, The frequency domain distribution of each tensor before quantization is determined based on the following method: Extract the frequency domain features of the tensor before quantization to obtain the frequency domain distribution of the tensor before quantization; and / or The frequency domain distribution of each tensor after quantization is determined based on the following method: Extract the frequency domain features of the tensor after quantization to obtain the frequency domain distribution of the tensor after quantization.
12. The method according to claim 11, wherein The dimension of each tensor is greater than or equal to 2. The extraction of the frequency domain features of the tensor before quantization includes: Convert the tensor before quantization into a one-dimensional tensor; Extract the frequency domain features of the obtained one-dimensional tensor to obtain the frequency domain distribution of the tensor before quantization; and / or The extraction of the frequency domain features of the tensor after quantization includes: Convert the tensor after quantization into a one-dimensional tensor; Extract the frequency domain features of the obtained one-dimensional tensor to obtain the frequency domain distribution of the tensor after quantization.
13. The method according to any one of claims 1-5, characterized in that, The multiple quantization strategies include: Multiple quantization strategies with different quantization types; and / or Multiple quantization strategies with the same quantization type and different quantization parameters.
14. The method according to claim 13, wherein The quantization type includes: uniform quantization and / or non-uniform quantization; the quantization parameters include: quantization step size and / or zero value point.
15. A neural network quantization device, characterized in that, The neural network is an image classification network, and the device includes: An acquisition module, configured to acquire a tensor set composed of tensors of the neural network, where the tensor set at least includes an image feature tensor of the image to be classified and network parameters of the neural network; A quantization processing module, configured to perform quantization processing on the tensors in the tensor set by using multiple quantization strategies respectively; A quantization strategy determination module, configured to determine the target quantization strategy of the tensor set from the multiple quantization strategies based on the frequency domain distribution of the tensor before quantization processing and the frequency domain distribution of the tensor after quantization processing, and use the result of performing quantization processing on the tensors in the tensor set by using the target quantization strategy as the quantization result of the tensor set, so as to obtain a quantized neural network, and classify the image to be classified by using the quantized neural network; where the frequency domain distribution of the image feature tensor represents the intensities of high-frequency signals and low-frequency signals in the image to be classified.
16. An electronic device, characterized in that, The electronic device includes a processor, a memory, and computer instructions stored in the memory and executable by the processor. When the processor executes the computer instructions, the method described in any one of claims 1-14 is implemented.
17. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and when the computer instructions are processed, the method described in any one of claims 1-14 is implemented.
Citation Information
Patent Citations
Electroencephalogram based behavior decision prediction system
CN106175757A
Method and system for realizing end-to-end fixed-point fast Fourier transform quantization by neural network
CN113626756A