Dynamic Quantization for Energy-Efficient Deep Learning
By using a dynamic quantization method, the quantization level of each layer of a deep convolutional neural network is adjusted based on the input content, which solves the problems of resource waste and accuracy degradation in traditional methods and achieves efficient inference on systems with limited resources.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2021-09-29
- Publication Date
- 2026-07-31
AI Technical Summary
Traditional deep convolutional neural networks are difficult to deploy on systems with limited resources. Conventional quantization methods cannot dynamically adapt to the input content, resulting in a waste of computing and storage resources and a decrease in accuracy.
A dynamic quantization method is adopted, which dynamically selects the quantization level of each layer based on the input content. By adjusting the bit width of activation and weights, resource utilization and inference accuracy are optimized. Dynamic bit depth selection is achieved using a bit width selector layer and regularization learning.
It effectively reduces the use of computing and storage resources, improves the inference accuracy and efficiency of the model on systems with limited resources, and dynamically adapts to different input conditions.
Smart Images

Figure CN116210009B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit of U.S. Patent Application No. 17 / 488,261, filed September 28, 2021, entitled "Dynamic Quantization for Energy-Efficient Deep Learning," and U.S. Provisional Patent Application No. 63 / 084,902, filed September 29, 2020, entitled "Dynamic Quantization for Energy-Efficient Deep Learning," the disclosures of which are expressly incorporated herein by reference in their entirety. Background Technology
[0003] field
[0004] The various aspects of this disclosure generally relate to dynamic quantization for energy-efficient deep learning neural networks.
[0005] background
[0006] Convolutional neural networks (such as deep convolutional neural networks (DCNNs)) can utilize significant computational and storage resources. This can make it difficult to deploy traditional neural networks on systems with limited resources, such as cloud systems, embedded systems, or federated learning systems. Some traditional neural networks are pruned and / or quantized to reduce processor load and / or memory usage. Improvements in quantization methods are desired to further reduce computational and storage resources.
[0007] Overview
[0008] In one aspect of this disclosure, a method performed by a deep neural network (DNN) is disclosed. The method includes receiving, during an inference phase, a layer input at a layer of the DNN, comprising content associated with a DNN input received at the DNN. The method further includes quantizing one or more parameters of a plurality of parameters associated with the layer based on the content of the layer input. The method further includes performing a task corresponding to the DNN input, the task being performed using the one or more quantized parameters.
[0009] Another aspect of this disclosure relates to an apparatus including means for receiving, during an inference phase, a layer of a DNN including content associated with a DNN input received at that DNN. The apparatus further includes means for quantizing one or more of a plurality of parameters associated with the layer based on the content of the layer input. The apparatus further includes means for performing a task corresponding to the DNN input, the task being performed using the one or more quantized parameters.
[0010] In another aspect of this disclosure, a non-transient computer-readable medium having non-transient program code recorded thereon is disclosed. This program code is used for a deep neural network (DNN). The program code is executed by a processor and includes program code for receiving, during an inference phase, layer inputs at a layer of the DNN, comprising content associated with DNN inputs received at the DNN. The program code also includes program code for quantizing one or more of a plurality of parameters associated with the layer based on the content of the layer inputs. The program code further includes program code for performing a task corresponding to the DNN inputs, the task being performed using the one or more quantized parameters.
[0011] Another aspect of this disclosure relates to an apparatus. The apparatus has a memory, one or more processors coupled to the memory, and instructions stored in the memory. These instructions, when executed by the processor, are operable to cause the apparatus to receive, during an inference phase, a layer input at a layer of a DNN, comprising content associated with a DNN input received at that DNN. These instructions also cause the apparatus to quantize one or more parameters of a plurality of parameters associated with that layer based on the content of the layer input. Additionally, these instructions cause the apparatus to perform a task corresponding to the DNN input, the task being performed using the one or more quantized parameters.
[0012] The aspects generally include, as substantially described herein with reference to the accompanying drawings and description, methods, apparatus (equipment), systems, computer program products, non-transient computer-readable media, user equipment, base stations, wireless communication equipment, and processing systems.
[0013] The foregoing has broadly outlined the features and technical advantages of the examples according to this disclosure in an effort to facilitate a better understanding of the following detailed description. Additional features and advantages will be described thereafter. The disclosed concepts and specific examples can be readily used as the basis for modifications or the design of other structures for implementing the same purposes as this disclosure. Such equivalent constructions do not depart from the scope of the appended claims. The characteristics of the disclosed concepts, in both their organization and manner of operation, and their associated advantages, will be better understood by considering the following description in conjunction with the accompanying drawings. Each drawing is provided for illustrative and descriptive purposes and not for defining limitations on the claims. Brief description of the attached diagram
[0015] The features, nature, and advantages of this disclosure will become more apparent when understood in conjunction with the accompanying drawings, in which the same reference numerals are always used to indicate the subject.
[0016] Figure 1 An example implementation of designing a neural network using a system-on-a-chip (SoC) (including a general-purpose processor) according to certain aspects of this disclosure is described.
[0017] Figure 2A , 2B 2C are illustrations explaining various aspects of the neural network according to this disclosure.
[0018] Figure 2D This is a diagram illustrating an exemplary deep convolutional network (DCN) according to various aspects of this disclosure.
[0019] Figure 3 This is a block diagram illustrating an exemplary deep convolutional network (DCN) according to various aspects of this disclosure.
[0020] Figure 4A This is a block diagram illustrating examples of neural network models based on various aspects of this disclosure.
[0021] Figure 4B This is a block diagram illustrating examples of bit-width selectors according to various aspects of this disclosure.
[0022] Figure 5 This is a block diagram illustrating an example of an input-based bit-width selector for selecting the quantization level of one or more channels of a layer, according to various aspects of this disclosure.
[0023] Figure 6 It is a flowchart illustrating an example process performed, for example, by a trained deep neural network, according to various aspects of this disclosure.
[0024] Detailed description
[0025] The detailed description that follows, taken in conjunction with the accompanying drawings, is intended as a description of various configurations and is not intended to represent only the configurations in which the described concepts can be practiced. This detailed description includes specific details to provide a thorough understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts can be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring such concepts.
[0026] Based on this teaching, those skilled in the art will appreciate that the scope of this disclosure is intended to cover any aspect of this disclosure, whether implemented independently of or in combination with any other aspect of this disclosure. For example, any number of the aspects described may be used to implement an apparatus or method of practice. Furthermore, the scope of this disclosure is intended to cover such apparatus or methods practiced using other structures, functionalities, or structures and functionalities that complement or differ from the aspects of the described disclosure. It should be understood that any aspect of this disclosure may be implemented by one or more elements of the claims.
[0027] The word “exemplary” is used to mean “serving as an example, instance, or explanation.” Any aspect described as “exemplary” need not be construed as superior to or better than other aspects.
[0028] While specific aspects are described, numerous variations and substitutions of these aspects fall within the scope of this disclosure. Although some benefits and advantages of preferred aspects are mentioned, the scope of this disclosure is not intended to be limited to a particular benefit, use, or objective. Rather, the aspects of this disclosure are intended to be broadly applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated as examples in the accompanying drawings and the following description of preferred aspects. The detailed description and accompanying drawings are merely illustrative and not limiting of this disclosure, the scope of which is defined by the appended claims and their equivalents.
[0029] Deep convolutional neural networks (DCNNs) can be specified for a variety of tasks, such as, for example, computer vision, speech recognition, and natural language processing. Traditional DCNNs can include a large number of weights and parameters and can be computationally intensive. Therefore, deploying DCNNs on embedded devices with limited resources, such as limited computing power and / or limited memory capacity, can be challenging. DCNNs can also be referred to as deep neural networks (DNNs) or deep convolutional networks (DCNs). A DCNN can include a total of three or more layers, one of which is a hidden layer.
[0030] Due to the computationally intensive nature of DNNs, it may be desirable to dynamically reduce model size and computational cost at runtime. Conventional systems reduce model size and computational complexity by applying separable filters, pruning weights, and / or reducing bit width. Additionally, conventional systems can reduce model size and computational complexity during training. These conventional systems do not dynamically reduce model size and computational cost during testing (e.g., at runtime) based on the content of the input (such as features).
[0031] In some examples, the bit width can be reduced to decrease model size and computational cost. Bit width reduction involves a quantization process that maps continuous real values to discrete integers. However, reducing the bit width may increase quantization error, thereby reducing the accuracy of the DNN. Thus, there is a trade-off between model accuracy and model efficiency. To optimize this trade-off, conventional systems train (e.g., customize) models based on resource budgets to optimize accuracy. That is, different models can be trained for different resource budgets. Training different models for different resource budgets may hinder dynamic bit width adjustment.
[0032] Some conventional systems train a single model that is flexible and scalable. For example, the number of channels can be adjusted by changing the width multiplier in each layer. As another example, depth, width, and kernel size can be dynamically adjusted. However, these conventional systems do not dynamically adjust the weight bit width and / or activation bit width.
[0033] Various aspects of this disclosure relate to a dynamic quantization method that dynamically selects the quantization level for each layer of a trained model to optimize a trade-off between reducing the amount of resources used (e.g., processor, battery, and / or memory resources) and inference accuracy. In such respects, the dynamic quantization method can be performed during testing (e.g., runtime) of the trained model.
[0034] Figure 1 An example implementation of a system-on-a-chip (SOC) 100 according to certain aspects of this disclosure is described, which may include a central processing unit (CPU) 102 or a multi-core CPU configured for dynamic bit-width quantization. Variables (e.g., neural signals and synaptic weights), system parameters associated with computing devices (e.g., weighted neural networks), latency, frequency slot information, and task information may be stored in memory blocks associated with a neural processing unit (NPU) 108, a memory block associated with the CPU 102, a memory block associated with a graphics processing unit (GPU) 104, a memory block associated with a digital signal processor (DSP) 106, a memory block 118, or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from the program memory associated with the CPU 102 or may be loaded from memory block 118.
[0035] SOC 100 may also include additional processing blocks tailored to specific functions, such as GPU 104, DSP 106, connectivity block 110 (which may include fifth-generation (5G) connectivity, fourth-generation LTE (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and multimedia processor 112, for example, capable of detecting and recognizing gestures. In one implementation, the NPU is implemented within the CPU, DSP, and / or GPU. SOC 100 may also include sensor processor 114, image signal processor (ISP) 116, and / or navigation module 120 (which may include a global positioning system).
[0036] The SOC 100 may be based on the ARM instruction set. In one aspect of this disclosure, the instructions loaded into the general-purpose processor 102 may include code for: identifying the content of an input received at the DNN; quantizing one or more parameters of a layer of the DNN based on the content of the input; and performing a task corresponding to the input. This task may be performed using the quantized parameters.
[0037] Deep learning architectures perform object recognition tasks by learning to represent inputs at progressively higher levels of abstraction in each layer, thereby constructing useful feature representations of the input data. In this way, deep learning addresses a major bottleneck in traditional machine learning. Before deep learning, machine learning methods for object recognition problems often relied heavily on human-engineered features, perhaps combined with shallow classifiers. Shallow classifiers could be two-class linear classifiers, where a weighted sum of feature vector components is compared to a threshold to predict which class the input belongs to. Human-engineered features could be templates or kernels customized for a specific problem domain by engineers with domain expertise. In contrast, deep learning architectures can learn to represent features similar to those that human engineers might design, but this learning is achieved through training. Furthermore, deep networks can learn to represent and recognize novel types of features that humans might not have considered.
[0038] Deep learning architectures can learn hierarchical levels of features. For example, if visual data is presented to the first layer, it can learn to recognize relatively simple features (such as edges) in the input stream. In another example, if auditory data is presented to the first layer, it can learn to recognize spectral power at specific frequencies. A second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as recognizing simple shapes in visual data or sound combinations in auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to recognize common visual objects or spoken phrases.
[0039] Deep learning architectures can perform particularly well when applied to problems with a naturally hierarchical structure. For example, the classification of motor vehicles can benefit from first learning to identify wheels, windshields, and other features. These features can then be combined in different ways at higher levels to identify cars, trucks, and airplanes.
[0040] Neural networks can be designed with various connectivity patterns. In feedforward networks, information is passed from lower layers to higher layers, where each neuron in a given layer communicates to neurons in higher layers. As mentioned above, hierarchical representations can be constructed in successive layers of a feedforward network. Neural networks can also have backflow or feedback (also known as top-down) connections. In a backflow connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Backflow architectures can help identify patterns across more than one block of input data sequentially delivered to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be beneficial when the recognition of higher-level concepts can help discern specific lower-level features of the input.
[0041] The connections between layers of a neural network can be fully connected or locally connected. Figure 2A An example of a fully connected neural network 202 is explained. In a fully connected neural network 202, a neuron in the first layer can pass its output to each neuron in the second layer, so that each neuron in the second layer receives input from each neuron in the first layer. Figure 2B An example of a locally connected neural network 204 has been explained. In the locally connected neural network 204, neurons in the first layer can connect to a finite number of neurons in the second layer. More generally, the locally connected layers of the locally connected neural network 204 can be configured such that each neuron in a layer will have the same or similar connectivity pattern, but its connection strength can have different values (e.g., 210, 212, 214, and 216). The connectivity patterns of locally connected networks may produce spatially dissimilar receptive fields in higher layers because higher-layer neurons in a given region can receive inputs that are tuned to a restricted portion of the total input to the network through training.
[0042] An example of a locally connected neural network is a convolutional neural network. Figure 2C An example of a convolutional neural network 206 has been explained. The convolutional neural network 206 can be configured such that the connection strength associated with the input for each neuron in the second layer is shared (e.g., 208). Convolutional neural networks may be well-suited for problems where the spatial location of the input is meaningful.
[0043] One type of convolutional neural network is the deep convolutional network (DCN). Figure 2D A detailed example of a DCN 200 designed to recognize visual features from an image 226 input from an image capture device 230 (such as an in-vehicle camera) is explained. The DCN 200 of this example can be trained to identify traffic signs and the numbers provided on them. Of course, the DCN 200 can be trained for other tasks, such as identifying lane markings or traffic lights.
[0044] The DCN 200 can be trained using supervised learning. During training, an image (such as image 226 of a speed limit sign) can be presented to the DCN 200, and a "forward pass" can then be computed to produce output 222. The DCN 200 may include feature extraction segments and classification segments. Upon receiving image 226, a convolutional layer 232 may apply a convolutional kernel (not shown) to image 226 to generate a first set of feature maps 218. As an example, the convolutional kernel of the convolutional layer 232 may be a 5x5 kernel that generates a 28x28 feature map. In this example, since four different feature maps are generated in the first set of feature maps 218, four different convolutional kernels are applied to image 226 at the convolutional layer 232. The convolutional kernel may also be referred to as a filter or convolutional filter.
[0045] The first set of feature maps 218 can be subsampled by a max-pooling layer (not shown) to generate a second set of feature maps 220. The max-pooling layer reduces the size of the first set of feature maps 218. That is, the size of the second set of feature maps 220 (e.g., 14x14) is smaller than the size of the first set of feature maps 218 (e.g., 28x28). The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 220 can be further convolved via one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).
[0046] exist Figure 2D In the example, the second set of feature maps 220 is convolved to generate a first feature vector 224. Furthermore, the first feature vector 224 is further convolved to generate a second feature vector 228. Each feature of the second feature vector 228 may include a number corresponding to a possible feature of the image 226 (such as "sign", "60", and "100"). A softmax function (not shown) converts the numbers in the second feature vector 228 into probabilities. Thus, the output 222 of the DCN 200 is the probability that the image 226 includes one or more features.
[0047] In this example, the probabilities of “sign” and “60” in output 222 are higher than the probabilities of other features in output 222 (such as “30”, “40”, “50”, “70”, “80”, “90”, and “100”). Before training, output 222 generated by DCN 200 is likely incorrect. Therefore, the error between output 222 and the target output can be calculated. The target output is the ground truth of image 226 (e.g., “sign” and “60”). The weights of DCN 200 can then be adjusted so that output 222 of DCN 200 is more closely aligned with the target output.
[0048] To adjust the weights, the learning algorithm can compute gradient vectors for the weights. This gradient indicates how much the error will increase or decrease as the weights are adjusted. At the top layers, this gradient directly corresponds to the values of the weights connecting the activated neurons in the penultimate layer to the neurons in the output layer. In lower layers, the gradient depends on the values of the weights and the error gradients computed from the higher layers. The weights can then be adjusted to reduce the error. This method of adjusting weights is called "backpropagation" because it involves a "back pass" in neural networks.
[0049] In practice, the error gradient of the weights may be calculated on a small number of examples, thus approximating the true error gradient. This approximation method is called stochastic gradient descent. Stochastic gradient descent can be repeated until the error rate achievable by the entire system stops decreasing or until the error rate reaches the target level. After learning, a new image (e.g., the speed limit sign in image 226) can be presented to the DCN and forward propagation through the network produces output 222, which can be considered an inference or prediction of the DCN.
[0050] Deep Belief Networks (DBNs) are probabilistic models that include multiple layers of hidden nodes. DBNs can be used to extract hierarchical representations of training datasets. DBNs can be obtained by stacking multiple layers of Restricted Boltzmann Machines (RBMs). RBMs are a class of artificial neural networks that can learn probability distributions on an input set. Because RBMs can learn probability distributions without information about which class each input should be classified into, they are often used in unsupervised learning. Using a hybrid unsupervised and supervised paradigm, the bottom RBM of a DBN can be trained unsupervised and used as a feature extractor, while the top RBM can be trained supervised (on the joint distribution of inputs from previous layers and the target class) and used as a classifier.
[0051] Deep convolutional networks (DCNs) are networks of convolutional networks configured with additional pooling and normalization layers. DCNs have achieved state-of-the-art performance on many tasks. DCNs can be trained using supervised learning, where both the input and output targets are known for many paradigms and are used to modify the network's weights using gradient descent.
[0052] DCNs can be feedforward networks. Furthermore, as mentioned above, the connections from neurons in the first layer of a DCN to the neuron group in the next higher layer are shared across neurons in the first layer. The feedforward and shared connections of a DCN can be used for fast processing. The computational burden of a DCN can be much smaller than, for example, a similarly sized neural network that includes backflow or feedback connections.
[0053] The processing of each layer in a convolutional network can be considered as a spatially invariant template or a fundamental projection. If the input is first decomposed into multiple channels, such as the red, green, and blue channels of a color image, then a convolutional network trained on that input can be considered three-dimensional, having two spatial dimensions along the axis of the image and a third dimension capturing color information. The output of the convolutional connections can be considered as forming a feature map in subsequent layers, where each element receives input from a range of neurons in the previous layer (e.g., feature map 220) and from each of those multiple channels. The values in the feature map can be further processed non-linearly (such as correction, max(0,x)). Values from neighboring neurons can be further pooled (which corresponds to downsampling) and provide additional local invariance and dimensionality reduction. Normalization, corresponding to whitening, can also be applied through lateral inhibition between neurons in the feature map.
[0054] The performance of deep learning architectures can improve as more labeled data points become available or as computational power increases. Modern deep neural networks are routinely trained with thousands of times more computational resources than were available to a typical researcher just fifteen years ago. New architectures and training paradigms can further boost the performance of deep learning. Corrected linear units reduce the training problem known as vanishing gradients. New training techniques reduce overfitting and thus enable larger models to achieve better generalization. Encapsulation techniques can abstract the data within a given receptive field and further improve overall performance.
[0055] Figure 3 This is a block diagram illustrating a Deep Convolutional Network 350. A Deep Convolutional Network 350 can include multiple layers of different types based on connectivity and weight sharing. For example... Figure 3As shown, the deep convolutional network 350 includes convolutional blocks 354A and 354B. Each of the convolutional blocks 354A and 354B may be configured with a convolutional layer (CONV) 356, a normalization layer (LNorm) 358, and a max pooling layer (MAX POOL) 360.
[0056] Convolutional layer 356 may include one or more convolutional filters that can be applied to the input data to generate feature maps. Although only two convolutional blocks 354A and 354B are shown, this disclosure is not limited thereto, and any number of convolutional blocks 354A and 354B may be included in the deep convolutional network 350 according to design preferences. Normalization layer 358 may normalize the output of the convolutional filters. For example, normalization layer 358 may provide whitening or lateral suppression. Max pooling layer 360 may provide spatial downsampling aggregation to achieve local invariance and dimensionality reduction.
[0057] For example, the parallel filter set of the deep convolutional network can be loaded onto the CPU 102 or GPU 104 of the SOC 100 to achieve high performance and low power consumption. In an alternative embodiment, the parallel filter set can be loaded onto the DSP 106 or ISP 116 of the SOC 100. Additionally, the deep convolutional network 350 can access other processing blocks that may exist on the SOC 100, such as the sensor processor 114 and navigation module 120, respectively dedicated to sensors and navigation.
[0058] The deep convolutional network 350 may also include one or more fully connected layers 362 (FC1 and FC2). The deep convolutional network 350 may further include logistic regression (LR) layers 364. Weights (not shown) to be updated are located between each layer 356, 358, 360, 362, and 364 of the deep convolutional network 350. The output of each layer (e.g., 356, 358, 360, 362, and 364) can be used as input to a subsequent layer in the deep convolutional network 350 (e.g., 356, 358, 360, 362, and 364) to learn a hierarchical feature representation from the input data 352 (e.g., image, audio, video, sensor data, and / or other input data) supplied from the first convolutional block 354A. The output of the deep convolutional network 350 is a classification score 366 for the input data 352. The classification score 366 may be a set of probabilities, where each probability is the probability that the input data includes features from a feature set.
[0059] As discussed, deep neural networks (DNNs) can be computationally intensive. Thus, DNNs can increase the use of system resources, such as processor load and / or memory usage. Quantization can reduce the amount of computation (such as binary computation) by reducing the weights and / or parameter bit widths. As a result, quantization can reduce the amount of system resources used by DNNs, thereby improving system performance.
[0060] Conventional quantization methods are static. For example, conventional quantization methods are not suitable for inputs to a DNN. Additionally or alternatively, conventional quantization methods are not suitable for the energy level of the device executing the DNN. Aspects of this disclosure relate to a dynamic quantization method that adaptively selects the quantization level based on the content of the input, such as features. As discussed, in some implementations, bits of one or more layers of a neural network (e.g., a deep neural network (DNN)) can be quantized during the inference phase (e.g., the testing phase).
[0061] This disclosure proposes a dynamic quantization method that adaptively changes the quantization precision (e.g., bit depth) based on the input or the content of the input (such as features extracted at each layer of a DNN). In one configuration, the bit depth is learned based on L0 regularization learning that penalizes large bit depths (see Equation 6).
[0062] Quantization can be layer-by-layer quantization or layer-by-layer and channel-by-channel quantization. In one configuration, a module (e.g., a computational / circuit system) examines the input (e.g., the content of the input) of the neural network layers. The output of this module determines the quantization amount of the activations and / or weights of the neural network layers. In some examples, the output of each layer can be considered a feature. In some configurations, a separate module can be specified for each layer to determine the quantization amount for that particular layer.
[0063] Therefore, the quantization of activations and weights can be varied based on the input. That is, the quantization of activations and / or weights used by layers(s) of a neural network can be varied by different inputs to that neural network. Different examples of inputs can allow for different quantization bit depths for activations and / or weights, producing statistically significant increments in bit-level computational complexity. Examples of different inputs include inputs used for classifying different categories (such as classifying cats and dogs) or different inputs used for image processing (such as image restoration), where the input can be a low-light image and a normal-light image. In some examples, inputs can be distinguished based on the background (such as the background of an image). In such examples, some inputs can be classified as simple background inputs, while others can be classified as complex background inputs. Different bit depths can be chosen for simple background inputs and complex background inputs. In some examples, operators using higher bit depths may take longer than those using lower bit depths. For example, 8-bit multiplication and accumulation operations can be composed of 4-bit multiplication and accumulation operations, so that an 8-bit version of a convolution can take longer than a convolution using 4 bits or less.
[0064] In one configuration, the gating mechanism can be specified to decrease / discard bits starting from the least significant bit of the activation bits before the activation bits are fed into the neural network processing layers(e.g., activations that can be input into convolutional layers). Another gating mechanism can be specified to decrease / discard bits starting from the least significant bit before the remaining bits are used by the neural network processing layers(e.g., weights / biases used by convolutional layers).
[0065] In one configuration, to leverage per-input bit-level computational complexity, the inference hardware can be modified to support dynamic configuration of bit depth at the layer and / or channel granularity. The inference hardware can utilize dynamic bit depth selection to improve power consumption and / or latency. As an example, the inference hardware can use a lower bit depth to reduce power consumption and / or latency.
[0066] In some respects, bit-width selector layers can be associated with each layer of a neural network model (e.g., a DNN). The bit-width selector can be trained to minimize the classification loss. At the same time, the width of each layer is reduced. Figure 4A This is a diagram illustrating an example of a neural network model 400 according to various aspects of this disclosure. The neural network model 400 can be an example of a DNN. Figure 4AIn the example, the neural network model 400 can dynamically adjust the bit width of each layer based on the input (denoted as In) (such as the content of the input). In some examples, the content can be features of the input, where these features can be extracted at one or more layers 402A, 402B, 402C of the neural network model 400. Layers 402A, 402B, 402C of the neural network model 400 are provided for illustrative purposes. The neural network model 400 is not limited to three layers 402A, 402B, 402C, as... Figure 4A As shown, additional or fewer layers can be used in neural network model 400. Figure 4A In the example, each layer 402A, 402B, 402C can be an example of one of the convolutional layers 232, 356, normalization layer 358, max pooling layer 360, fully connected layer 362, or logistic regression layer 364 described with reference to Figures 2 and 3.
[0067] exist Figure 4A In the example, each layer 402A, 402B, 402C can be associated with bit-width selectors 404A, 404B, 404C. Each bit-width selector 404A, 404B, 404C can be data-dependent and can be a layer of the neural network model 400. Additionally, each bit-width selector 404A, 404B, 404C can also implement a bit-width selection function (σ) for selecting the quantization depth of the corresponding layer 402A, 402B, 402C. i For example, such as Figure 4A As shown, the first layer 402A and the first bit width selector 404A can receive input (In). Based on the input (e.g., the content of the input), the first bit width selector 404A can select the bit width (φ1) for the first layer 402A. In the current example, the first layer 402A can process the input and generate the output received at the second layer 402B and the second bit width selector 404B. In this example, the second bit width selector 404B selects the bit width (φ2) for the second layer 402B based on the output of the first layer 402A. This process continues for the third layer 402C, such that the third bit width selector 404C selects the bit width (φ3) for the third layer 402C based on the output of the second layer 402B.
[0068] According to various aspects of this disclosure, the bit width φ of each layer can be selected based on training. i The regularization term λ can be specified during training and can be omitted during inference. Figure 4A In the example, the output (λφ) corresponds to the downward arrow from each bit-width selector 404A, 404B, 404C. iThe regularization term λ can be generated during training and can be omitted during inference. The regularization term λ can be set by the user and can vary depending on the device or task. Bit width φ i It can also be referred to as depth. Figure 4A In the example, the upward arrows from each bit-width selector 404A, 404B, and 404C correspond to the bit width φ of the corresponding layer. i As discussed, the bit width selection function (σ) implemented at each bit width selector 404A, 404B, 404C i The selectors can be trained to penalize large bit widths and / or one or more associated complexity metrics when optimizing performance. During training and inference, the output of each bit-width selector 404A, 404B, 404C can be similar to the reference... Figure 4B Or the output described in 5.
[0069] Figure 4B This is a block diagram illustrating an example of a data-related bit-width selector 450 based on various aspects of this disclosure. Figure 4B In the example, bit-width selector 450 can be Figure 4A Examples of each of the bit-width selectors 404A, 404B, and 404C described herein. Figure 4B As shown, the bit-width selector 450 includes a global pooling layer 452, a first fully connected (FC) layer 454 that implements the rectified linear unit (ReLU) activation function, and a second fully connected layer 456 that implements the sigmoid function.
[0070] exist Figure 4B In the example, the input to the bit-width selector 450 can be received at the global pooling layer 452 and collapsed into a 1x1xC vector, where the parameter C represents the number of channels of the layer corresponding to the bit-width selector 450. The first fully connected layer 454 receives the 1x1xC vector to generate a vector with dimension... The output of, where The number corresponds to the number of strobes 460A, 460B, 460C, and 460D, and the parameter r is a reduction factor used as an integer divisor of C. In one configuration, r is a hyperparameter. The output of the first fully connected layer 454 is received at the second fully connected layer 456. Several strobes (g) are specified for the bit-width selector 450. i The output of the second fully connected layer 456 is received at points 460A, 460B, 460C, and 460D. Figure 4B In the example, in the first strobe 460A (g 3 ), second strobe 460B (g 2 ), third strobe 460C (g 1 ) and the fourth strobe 460D (g 0 The output of the second fully connected layer 456 is received at ().
[0071] exist Figure 4B In the example, the global pooling layer 452 (e.g., a global average pooling layer) can combine information from the input activations / features and generate per-channel statistics (such as the average). A global descriptor for the input features can be generated based on the per-channel statistics. The statistics generated at the global pooling layer 452 can be combined using a first fully connected layer 454 with ReLU or other non-linear activations. A second fully connected layer 456 can implement a sigmoid activation providing a range of 0 to 1. Bit extension layer ( Figure 4B (Not shown in the diagram) The outputs of the bit-width selector gates 460A, 460B, 460C, and 460D are generated by multiplying the sigmoid output of the second fully connected layer 456 with the learned weights in the bit-width dimension. For example, the bit-expansion layer may have eight weights for determining the bit-width of an 8-bit network.
[0072] In one configuration, the lower significant bit is selected only if the higher significant bit is selected. In the example above, the maximum bit width of the layer associated with the data-dependent bit-width selector 450 can be based on several strobes (g) specified for the data-dependent bit-width selector 450. i 460A, 460B, 460C, 460D. In this example, if the higher significant bit is selected (such as with the first strobe 460A (g...)... 3 If the bit associated with it is used, then the lower significant bit can be selected, such as the bit associated with the second strobe 460B (g). 2 ), third strobe 460C (g 1 ), Fourth strobe 460D (g 0 The associated bits. In one configuration, strobes 460A, 460B, 460C, and 460D are linear or power-based (based on 2). Figure 4B In the example, the architecture of the bit-width selector 450 is exemplary. The neural network model disclosed herein is not limited to... Figure 4B The diagram shows a global pooling layer 452, a first fully connected layer 454, and a second fully connected layer 456. Other layer configurations are envisioned.
[0073] During training, each layer of the neural network model (such as the reference layer) Figure 4A The described layers 402A, 402B, and 402C) caused losses. This penalty applies during the large bit width inference period. Loss It can be sparse regularization L0. Each layer can also incur a performance loss L for a specific task (such as prediction or regression accuracy). pThis disclosure is not limited to the network configurations described. Other network configurations are envisioned. In some examples, the regularization loss L0 can penalize bit-level operations (BOPs). A BOP can be the product of the selected bit width and the layer complexity. The loss function can be defined as:
[0074]
[0075] Where parameters Represents the loss, parameter x i It is the input, the parameter It is the activation bit width, and the parameter θ represents the parameters of the model (such as neural network model 400 in Figure 4). It is the weighted bit width, parameter y i It is a label, parameter ||z i || bit It is a bitwise regularization of L0, and the parameter λ represents the weighting factor for regularization. In Equation 1, N represents the total number of training examples. The first term of Equation 1 To predict f() and label y i The model training loss between them. The second term (λ‖z) i || bit ) is the bit width count / penalty for each training sample.
[0076] The regularization loss L0 is discrete, and therefore, non-differentiable. The bit width z can be represented as discrete samples from a Bernoulli distribution. In one configuration, z ~ π = Bernoulli(φ), where the parameter φ represents the number of bits selected for the layer such that the total loss (e.g., the regularization loss L0 and the performance loss) is such that... ) can be written as:
[0077]
[0078] Equation 2 is not differentiable and can be discretized. In some implementations, the original problem can be recovered by restricting the number of bits per layer to 0 and 1 (e.g., φ∈{0,1}). By selecting the quantization bits, the original problem can be recovered to solve for dynamic quantization. The original problem can be rewritten to make it solvable. The rewritten regularization term can be optimized using mini-batch gradient descent. Additionally, the rewritten regularization term can be evaluated analytically. Equation 3 is a rewritten regularization term. Additionally, Equation 3 can be differentiable:
[0079]
[0080] As mentioned above, the bit depth z is non-differentiable. The bit depth z can be relaxed by applying the sigmoid function. A binary random variable can be sampled as follows:
[0081] L = logu - log(1 - u). (4)
[0082] Based on Equation 4, if (logφ+L)>0, then the bit depth z can be 1, and if (logφ+L)<0, then the bit depth z can be 0. Discontinuous functions can be replaced by sigmoid:
[0083]
[0084] The total loss (e.g., performance loss and regularization loss) can be differentiable with respect to the number of bits φ, thus enabling optimization based on stochastic gradients. In Equation 5, the number of bits φ represents the total number of bits in the neural network. The parameter φ represents the number of bits during inference (e.g., bit width). During training, the parameter φ represents the average / expected bit width chosen based on the probability distribution.
[0085] As stated in Equation 1, the second term (λ‖z) i ||0) is the bit width count / penalty per training sample. In some examples, the penalty can be a bit-by-bit L0 regularization loss, which is based on bit-level operations determined by the adjusted bit width and a complexity metric (e.g., a complexity metric associated with a layer), the number of bits allocated to the adjusted bit width, or one of the complexity metrics. The complexity metric can be based on one or more of the adjusted bit width, the network's binary operations (BOPs), and memory footprint (e.g., intermediate activations, computational power, or other complexity metrics). Computational power can be a combination of BOPs of the hardware associated with the neural network, memory bandwidth, and power scaling. The regularization loss can be defined as:
[0086]
[0087] Where the parameter q k This represents the complexity measure at bit level k, where K represents the maximum bit width. As an example, if the maximum bit width K equals 8, then bit level k can range from 1 to 8. Each bit level k results in a customized complexity (qk). k For example, using all eight bits (q8) may be more expensive than using four bits (q4). In Equation 6, the metric can be one or more of bit width, BOP, computational complexity, memory footprint, computational power, or another metric at the selected bit level.
[0088] For reference Figure 4A As described, in some respects, the quantization level for each layer can be selected based on the input. Additionally or alternatively, the quantization level for each channel can be selected based on the input. Figure 5 This is a block diagram illustrating an example of a bit-width selector 500 that selects the quantization level for one or more channels of a layer based on input, according to various aspects of this disclosure. Figure 5In the example, bit-width selector 500 can be Figure 4A Examples of each of the bit-width selectors 404A, 404B, and 404C described herein. Figure 5 As shown, the bit width selector 500 includes a global pooling layer 552, a first fully connected (FC) layer 554 that implements the rectified linear unit (ReLU) activation function, a second fully connected layer 556 that implements the sigmoid function, and a bit extension layer 558. Figure 5 Global pooling layer 552, first fully connected layer 554, second fully connected layer 556 execution and reference Figure 4B The global pooling layer 452, the first fully connected layer 454, and the second fully connected layer 456 operate identically. For simplicity, we will start from... Figure 5 The description omits the operations of the global pooling layer 552, the first fully connected layer 554, and the second fully connected layer 556. The bit extension layer 558 can be based on the bit width and the number of output channels C. out The product is used to expand the bits. The bit expansion layer 558 can generate for these multiple output channels (C out Each channel in the array has several gating options. For example, such as... Figure 5 As shown in the example, the bit extension layer 558 generates (C out *Bit width) Output (shown as channel 1 to channel C) out The output of the first group of strobes (560A, 560B, 560C, 560D) can be associated with the first channel (channel 1), and the output of the second group of strobes (562A, 562B, 562C, 562D) can be associated with the last channel (C). out This is related to the fact that, in such examples of channel-by-channel bit selection networks, bit widths can be generated for each output layer of the neural network.
[0089] Figure 6 This is a flowchart illustrating an example process 600 performed, for example, by a deep neural network (DNN), according to various aspects of this disclosure. Example process 600 is an example of dynamically quantizing layer parameters based on input to the deep neural network.
[0090] like Figure 6 As shown in box 602, during the inference phase, the DNN receives layer inputs at its layers, including content associated with the DNN inputs received at that DNN. This layer may be a reference layer. Figure 4A An example of one of the described layers 402A, 402B, and 402C. In box 604, the DNN quantizes one or more of a plurality of parameters associated with that layer based on the content of the layer's input. In some examples, these plurality of parameters may be determined by a bit-width selector (such as a reference). Figure 4BThe described bit-width selector (450) is quantized. These multiple parameters may include a set of weights and a set of activations. In some examples, quantizing the one or more parameters involves generating an adjusted bit-width by adjusting the size of the original bit-width associated with the one or more parameters. The adjusted bit-width can be generated by discarding bits of the original bit-width from least significant bit to most significant bit until the size of the original bit-width equals the size of the adjusted bit-width determined based on the content of the input. In some such examples, the DNN can be trained to determine the size used to adjust the original bit-width based on a total loss, which is a function of performance loss and regularization loss. The performance loss may be determined by cross-entropy loss or mean squared error or other supervised loss functions. Additionally, the regularization loss may be a bitwise L0 regularization loss, which penalizes either the adjusted bit-width and a complexity metric associated with bit-level operations or the number of bits allocated to the adjusted bit-width. The complexity metric may include one or more of the number of binary operations of the DNN, the memory footprint of the DNN, or the computational power of the DNN.
[0091] In some other examples, quantizing the one or more parameters includes quantizing one or both of the corresponding weight sets or corresponding activation sets of one or more output channels associated with the layer. The corresponding weight sets and / or corresponding activation sets of one or more output channels associated with the layer can be determined by a bit-width selector (such as a reference). Figure 5 The bit-width selector 500 described is used for quantization. In some examples, the first quantization value of the first parameter among the plurality of parameters is different from the second quantization value of the second parameter among the plurality of parameters.
[0092] like Figure 6 As shown in box 606, the DNN performs a task corresponding to its input, which is performed using one or more quantized parameters. In some examples, this layer is one of many layers in the DNN. In such examples, one or more parameters of the corresponding parameters of each of the multiple layers can be quantized based on the content of the input. In some such examples, the quantization amount is different for each of the multiple layers.
[0093] Examples of implementations are described in the following numbered clauses.
[0094] 1. A method performed by a deep neural network (DNN), comprising: receiving, during an inference phase, a layer input at a layer of the DNN including content associated with a DNN input received at the DNN; quantizing one or more parameters of a plurality of parameters associated with the layer based on the content of the layer input; and performing a task corresponding to the DNN input, the task being performed using the one or more quantized parameters.
[0095] 2. The method as described in Clause 1, wherein the plurality of parameters includes a set of weights and a set of activations.
[0096] 3. The method of any of Clause 2, wherein quantizing the one or more parameters includes quantizing one or both of the corresponding weight set or the corresponding activation set of one or more output channels associated with the layer.
[0097] 4. The method of any of Clauses 1-2, wherein the first quantization of the first parameter among the plurality of parameters is different from the second quantization of the second parameter among the plurality of parameters.
[0098] 5. The method of any of Clauses 1-4, wherein quantizing the one or more parameters includes generating an adjusted bit width by adjusting the size of the original bit width associated with the one or more parameters.
[0099] 6. The method of Clause 5, wherein generating the adjusted bit width includes discarding bits of the original bit width from least significant bit to most significant bit until the size of the original bit width is equal to the size of the adjusted bit width determined based on the content of the input.
[0100] 7. The method of Item 6 further includes: training the DNN to determine the size for adjusting the original bit width based on a total loss, which is a function of performance loss and regularization loss.
[0101] 8. The method of Clause 7, wherein the performance loss is determined as either cross-entropy loss or mean squared error.
[0102] 9. The method of Clause 8, wherein the regularization loss is a bitwise L0 regularization loss that penalizes one of the following: the adjusted bit width and the complexity measure associated with the bit-level operation; or the number of bits allocated to the adjusted bit width.
[0103] 10. The method of Clause 9, wherein the complexity metric includes one or more of the number of binary operations of the DNN, the memory footprint of the DNN, or the computational power of the DNN.
[0104] 11. The method of Clause 9 further comprises: recompiling the bitwise L0 regularization loss into a Bernoulli distribution; relaxing the recompiling bitwise L0 regularization loss based on the sigmoid function; and minimizing the performance loss and regularization loss based on the number of bits selected for the adjusted bit width.
[0105] 12. The method of any one of clauses 1-11, wherein the layer is one of a plurality of layers of the DNN, and the method further includes quantizing one or more parameters of a corresponding plurality of parameters of each of the plurality of layers based on the content of the input.
[0106] 13. The method as described in Clause 12, wherein the quantization quantity is different for each of the plurality of layers.
[0107] The various operations of the methods described above can be performed by any suitable means capable of performing the corresponding functions. These means may include various hardware and / or software components and / or modules, including but not limited to circuits, application-specific integrated circuits (ASICs), or processors. Generally, in the cases where operations are illustrated in the accompanying drawings, those operations may have corresponding paired means with similar numbers plus functional components.
[0108] As used, the term "determine" encompasses a wide variety of actions. For example, "determine" can include calculation, computation, processing, derivation, research, searching (e.g., looking in a table, database, or other data structure), ascertaining, and the like. Additionally, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), and similar actions. Furthermore, "determine" can include parsing, selecting, choosing, establishing, and similar actions.
[0109] As used, the phrase “at least one of” refers to any combination of these items, including a single member. As an example, “at least one of a, b, or c” is intended to cover: a, b, c, ab, ac, bc, and abc.
[0110] The various illustrative logic blocks, modules, and circuits described in this disclosure can be implemented or executed using a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the described functions. The general-purpose processor may be a microprocessor, but in alternatives, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.
[0111] The steps of the methods or algorithms described in this disclosure can be implemented directly in hardware, in a software module executed by a processor, or in a combination of both. The software module can reside in any form of storage medium known in the art. Some examples of usable storage media include random access memory (RAM), read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs, and so on. The software module may include a single instruction or many instructions, and may be distributed across several different code segments, across different programs, and across multiple storage media. The storage medium may be coupled to the processor so that the processor can read and write information from / to the storage medium. In an alternative, the storage medium may be integrated into the processor.
[0112] The disclosed methods include one or more steps or actions for achieving the described methods. These method steps and / or actions may be interchanged without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims.
[0113] The described functionality can be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system within the device. The processing system can be implemented using a bus architecture. Depending on the specific application and overall design constraints of the processing system, the bus may include any number of interconnect buses and bridges. The bus can link together various circuits, including processors, machine-readable media, and bus interfaces. The bus interface can be used, in particular, to connect network adapters and the like to the processing system via the bus. The network adapter can be used to implement signal processing functions. In some respects, user interfaces (e.g., keypads, displays, mice, joysticks, etc.) may also be connected to the bus. The bus can also link various other circuits, such as timing sources, peripherals, regulators, power management circuits, and similar circuits, which are well known in the art and will not be described further.
[0114] A processor is responsible for managing the bus and general processing, including executing software stored on a machine-readable medium. A processor may be implemented using one or more general-purpose and / or special-purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuit systems capable of executing software. Software should be interpreted broadly as instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. As examples, a machine-readable medium may include random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, disks, optical disks, hard drives, or any other suitable storage medium, or any combination thereof. The machine-readable medium may be implemented in a computer program product. This computer program product may include packaging materials.
[0115] In hardware implementations, machine-readable media can be a separate part of the processing system from the processor. However, as those skilled in the art will readily appreciate, machine-readable media or any part thereof can be external to the processing system. As examples, machine-readable media may include transmission lines, carrier waves modulated by data, and / or computer components separate from the device, all accessible to the processor via a bus interface. Alternatively or additionally, machine-readable media or any part thereof may be integrated into the processor, such as caches and / or general-purpose register files. While the various components discussed may be described as having a specific location, such as local components, they can also be configured in various ways, such as certain components being configured as part of a distributed computing system.
[0116] The processing system can be configured as a general-purpose processing system having one or more microprocessors providing processor functionality, and external memory providing at least a portion of machine-readable medium, all linked to other supporting circuitry via an external bus architecture. Alternatively, the processing system may include one or more neuromorphic processors for implementing the described neuron and nervous system models. As another alternative, the processing system can be implemented using an application-specific integrated circuit (ASIC) with a processor, bus interface, user interface, supporting circuitry, and at least a portion of machine-readable medium integrated on a single chip, or using one or more field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic, discrete hardware components, or any other suitable circuitry, or any combination of circuitry capable of performing the various functionalities described throughout this disclosure. Depending on the specific application and the overall design constraints imposed on the network or system, those skilled in the art will recognize how best to implement the functionality described with respect to the processing system.
[0117] Machine-readable media may include several software modules. These software modules include instructions that, when executed by a processor, cause the processing system to perform various functions. These software modules may include transfer modules and receive modules. Each software module may reside in a single storage device or be distributed across multiple storage devices. As an example, when a trigger event occurs, a software module may be loaded from a hard drive into RAM. During the execution of a software module, the processor may load some instructions into a cache to improve access speed. One or more cache lines may subsequently be loaded into a general-purpose register file for processor execution. In the context of the functionality of the software modules described below, it will be understood that such functionality is implemented by the processor when the processor executes the instructions from the software module. Furthermore, it should be understood that aspects of this disclosure result in improvements to the functionality of a processor, computer, machine, or other system implementing such aspects.
[0118] If implemented in software, the functions can be stored or transmitted as one or more instructions or codes on or through a computer-readable medium. Computer-readable media includes both computer storage media and communication media, encompassing any medium that facilitates the transfer of a computer program from one location to another. Storage media can be any available medium accessible to a computer. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and is accessible to a computer. Additionally, any connection is also legitimately referred to as computer-readable media. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared (IR), radio, and microwave), then that coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) is included in the definition of medium. The disks and discs used include CDs, laser discs, optical discs, DVDs, floppy disks, and Blu-ray discs. Disks, where disks often magnetically reproduce data, and discs optically reproduce data using lasers. Therefore, in some aspects, computer-readable media may include non-transient computer-readable media (e.g., tangible media). Additionally, in other aspects, computer-readable media may include transient computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.
[0119] Therefore, some aspects may include a computer program product for performing the given operations. For example, such a computer program product may include a computer-readable medium on which instructions are stored (and / or encoded) that can be executed by one or more processors to perform the described operations. In some aspects, the computer program product may include packaging material.
[0120] Furthermore, it should be understood that modules and / or other suitable means for performing the described methods and techniques may be downloaded and / or otherwise obtained by the user terminal and / or base station where applicable. For example, such devices can be coupled to a server to facilitate the transfer of means for performing the described methods. Alternatively, the various methods described can be provided via a storage device (e.g., RAM, ROM, physical storage media such as CDs or floppy disks, etc.) so that the device can obtain the various methods once the storage device is coupled to or provided to the user terminal and / or base station. In addition, any other suitable techniques suitable for providing the described methods and techniques to the device may be utilized.
[0121] It will be understood that the claims are not limited to the precise configurations and components described above. Various modifications, substitutions, and variations can be made to the layout, operation, and details of the methods and apparatus described above without departing from the scope of the claims.
Claims
1. A method performed by a deep neural network (DNN) for computer vision, auditory feature recognition, or natural language processing, comprising: During the inference phase, layer inputs are received at layers of the DNN, including content associated with DNN inputs received at the DNN, each output channel of the layer having an original bit width before receiving the layer inputs, the DNN inputs being visual data, auditory data, or spoken phrases; Based on the content of the layer input, quantify one or more parameters of a corresponding plurality of parameters associated with each of the one or more output channels of the layer; The original bit width of each of the one or more output channels of the layer is reduced by quantizing the one or more parameters to generate a corresponding adjusted bit width for each of the one or more output channels, the original bit width being associated with the one or more parameters, the number of bits processed by one or more processors executing the DNN corresponding to the size of the adjusted bit width, such that one or more of the computing power of the one or more processors or the memory bandwidth of the one or more processors is reduced by reducing the corresponding original bit width of each of the one or more output channels; as well as Perform computer vision, auditory feature recognition, or natural language processing tasks corresponding to the DNN input, the tasks being performed via the one or more processors based on the adjusted bit width.
2. The method of claim 1, wherein the corresponding plurality of parameters includes a weight set and an activation set.
3. The method of claim 2, wherein quantizing the one or more parameters comprises quantizing one or both of a corresponding set of weights or a corresponding set of activations for each of the one or more output channels associated with the layer.
4. The method of claim 1, wherein the first quantization value of the first parameter among the corresponding plurality of parameters is different from the second quantization value of the second parameter among the corresponding plurality of parameters.
5. The method of claim 1, wherein generating the adjusted bit width includes discarding bits of the original bit width from least significant bit to most significant bit until the size of the original bit width is equal to the size of the adjusted bit width determined based on the content of the layer input.
6. The method of claim 5, further comprising: The DNN is trained to determine the amount used to adjust the size of the original bit width based on a total loss, which is a function of performance loss and regularization loss.
7. The method of claim 6, wherein the performance loss is determined as cross-entropy loss or mean square error.
8. The method of claim 7, wherein the regularization loss is a bitwise L0 regularization loss that penalizes one of the following: The adjusted bit width and the complexity measure associated with bit-level operations; or The number of bits allocated to the adjusted bit width.
9. The method of claim 8, wherein the complexity metric includes one or more of the number of binary operations of the DNN, the memory footprint of the DNN, or the computational power of the DNN.
10. The method of claim 8, further comprising: The bitwise L0 regularization loss is recompiled into a Bernoulli distribution; The re-compiled bitwise L0 regularization loss is relaxed based on the sigmoid function; as well as The performance loss and the regularization loss are minimized based on the number of bits allocated to the adjusted bit width.
11. The method of claim 1, wherein the layer is one of a plurality of layers of the DNN, and The method further includes quantifying one or more parameters of a plurality of parameters for each of the plurality of layers based on the content of the input of the respective layer.
12. The method of claim 11, wherein the quantization amount is different for each of the plurality of layers.
13. An apparatus for use by a deep neural network (DNN) for computer vision, auditory feature recognition, or natural language processing, comprising: processor; Memory coupled to the processor; as well as Instructions, which are stored in the memory and, when executed by the processor, are operable to cause the device to: During the inference phase, layer inputs are received at layers of the DNN, including content associated with DNN inputs received at the DNN, each output channel of the layer having an original bit width before receiving the layer inputs, the DNN inputs being visual data, auditory data, or spoken phrases; Based on the content of the layer input, quantize one or more parameters of a corresponding plurality of parameters associated with each of the one or more output channels of the layer; The original bit width of each of the one or more output channels of the layer is reduced by quantizing the one or more parameters to generate a corresponding adjusted bit width for each of the one or more output channels, the original bit width being associated with the one or more parameters, the number of bits processed by one or more processors executing the DNN corresponding to the size of the adjusted bit width, such that one or more of the computing power of the one or more processors or the memory bandwidth of the one or more processors is reduced by reducing the corresponding original bit width of each of the one or more output channels; as well as Perform computer vision, auditory feature recognition, or natural language processing tasks corresponding to the DNN input, the tasks being performed via the one or more processors based on the adjusted bit width.
14. The apparatus of claim 13, wherein the respective plurality of parameters includes a weight set and an activation set.
15. The apparatus of claim 14, wherein the instructions further cause the apparatus to quantize the one or more parameters by quantizing one or both of a corresponding set of weights or a corresponding set of activations for each of the one or more output channels associated with the layer.
16. The apparatus of claim 13, wherein the first quantization value of the first parameter among the corresponding plurality of parameters is different from the second quantization value of the second parameter among the corresponding plurality of parameters.
17. The apparatus of claim 13, wherein the instructions further cause the apparatus to generate the adjusted bit width by discarding bits of the original bit width from least significant bit to most significant bit until the size of the original bit width is equal to the size of the adjusted bit width determined based on the content of the layer input.
18. The apparatus of claim 17, wherein the instructions further cause the apparatus to determine, during the training phase, an amount for adjusting the size of the original bit width based on a total loss, the total loss being a function of performance loss and regularization loss.
19. The apparatus of claim 18, wherein the regularization loss is a bitwise L0 regularization loss that penalizes one of the following: The adjusted bit width and the complexity measure associated with bit-level operations; or The number of bits allocated to the adjusted bit width.
20. The apparatus of claim 19, wherein the complexity metric includes one or more of the number of binary operations of the DNN, the memory footprint of the DNN, or the computational power of the DNN.
21. The DNN of claim 19, wherein the instructions further cause the DNN to: The bitwise L0 regularization loss is recompiled into a Bernoulli distribution; The loss is relaxed based on the sigmoid function, using a re-compiled bitwise L0 regularization loss; and The performance loss and the regularization loss are minimized based on the number of bits allocated to the adjusted bit width.
22. The DNN of claim 13, wherein the layer is one of a plurality of layers of the DNN, and The instructions further enable the DNN to quantize one or more parameters from a plurality of parameters for each of the plurality of layers based on the content of the input to the corresponding layer.
23. The DNN of claim 22, wherein the quantization amount is different for each of the plurality of layers.
24. A non-transient computer-readable medium having program code recorded thereon for execution by a deep neural network (DNN) for computer vision, auditory feature recognition, or natural language processing, said program code being executed by a processor and comprising: Program code for receiving layer inputs at layers of the DNN during the inference phase, including content associated with DNN inputs received at the DNN, each output channel of the layer having an original bit width before receiving the layer inputs, the DNN inputs being visual data, auditory data, or spoken phrases; Program code for quantizing one or more parameters of a plurality of parameters associated with each of one or more output channels of the layer based on the content of the layer input; Program code for reducing the original bit width of each of the one or more output channels of the layer according to quantizing the one or more parameters to generate a corresponding adjusted bit width for each of the one or more output channels, the original bit width being associated with the one or more parameters, the number of bits processed by one or more processors executing the DNN corresponding to the size of the adjusted bit width, such that one or more of the computing power of the one or more processors or the memory bandwidth of the one or more processors is reduced by reducing the corresponding original bit width of each of the one or more output channels; as well as Program code for performing computer vision, auditory feature recognition, or natural language processing tasks corresponding to the DNN input, the tasks being performed via the one or more processors according to the adjusted bit width.
25. The non-transient computer-readable medium of claim 24, wherein the respective plurality of parameters includes a weight set and an activation set.
26. The non-transient computer-readable medium of claim 25, wherein the program code for quantizing the one or more parameters includes program code for quantizing one or both of a corresponding set of weights or a corresponding set of activations for each of the one or more output channels associated with the layer.
27. An apparatus for use by a deep neural network (DNN) for computer vision, auditory feature recognition, or natural language processing, comprising: A means for receiving, during the inference phase, a layer of the DNN including content associated with DNN input received at the DNN, each output channel of the layer having an original bit width before receiving the layer input, the DNN input being visual data, auditory data, or spoken phrases; A means for quantifying one or more parameters of a plurality of parameters associated with each of one or more output channels of the layer based on the content of the layer input; A means for reducing the original bit width of each of the one or more output channels of the layer according to quantizing the one or more parameters to generate a corresponding adjusted bit width for each of the one or more output channels, the original bit width being associated with the one or more parameters, the number of bits processed by one or more processors executing the DNN corresponding to the size of the adjusted bit width, such that one or more of the computing power of the one or more processors or the memory bandwidth of the one or more processors is reduced by reducing the corresponding original bit width of each of the one or more output channels; as well as A means for performing a computer vision, auditory feature recognition, or natural language processing task corresponding to the DNN input, the task being performed via the one or more processors according to the adjusted bit width.
28. The apparatus of claim 27, wherein the means for quantizing the one or more parameters includes means for quantizing one or both of a corresponding set of weights or a corresponding set of activations for each of the one or more output channels associated with the layer.
29. An apparatus for use by a deep neural network (DNN) for computer vision, auditory feature recognition, or natural language processing, comprising: One or more processors; One or more memories coupled to the one or more processors; as well as Instructions stored in the one or more memories, which, when executed by the one or more processors, are operable to cause the device to perform an inference phase: At the layer of the DNN, a layer input including content associated with the DNN input received at the DNN is received, each output channel of the layer having an original bit width before receiving the layer input, the DNN input being visual data, auditory data, or spoken phrases; Based on the content of the layer input, quantify one or more parameters of a corresponding plurality of parameters associated with each of the one or more output channels of the layer; The original bit width of each of the one or more output channels of the layer is reduced by quantizing the one or more parameters to generate a corresponding adjusted bit width for each of the one or more output channels, the original bit width being associated with the one or more parameters, the number of bits processed by the one or more processors executing the DNN corresponding to the size of the adjusted bit width, such that one or more of the computing power of the one or more processors or the memory bandwidth of the one or more processors is reduced by reducing the corresponding original bit width of each of the one or more output channels; and Perform a computer vision, auditory feature recognition, or natural language processing task corresponding to the DNN input, the task being performed via the one or more processors according to the adjusted bit width, wherein the instructions, when executed by the one or more processors, are operable to cause the bit width selector of the device to generate channel-by-channel statistics about the layer input to perform the reduction of the corresponding original bit width of each of the one or more output channels.
30. The apparatus of claim 29, wherein the respective plurality of parameters includes a weight set and an activation set.
31. The apparatus of claim 30, wherein the instructions further cause the apparatus to quantize the one or more parameters by quantizing one or both of a corresponding weight set or a corresponding activation set for each of the one or more output channels associated with the layer.