Simulated low-bitwidth quantization using bit-shifted neural network parameters

JP7866071B2Active Publication Date: 2026-05-26QUALCOMM INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
QUALCOMM INC
Filing Date
2023-01-31
Publication Date
2026-05-26

Smart Images

  • Figure 0007866071000008
    Figure 0007866071000008
  • Figure 0007866071000009
    Figure 0007866071000009
  • Figure 0007866071000010
    Figure 0007866071000010
Patent Text Reader

Abstract

The processor-implemented method includes bit-shifting a binary representation of a neural network parameter, the neural network parameter having bits b that are less than the number of hardware bits B supported by hardware that processes the neural network parameter. The bit-shifting adds 2 bits to the neural network parameter. B-b The method also multiplies the quantization scale by 2 to get the updated quantization scale. B-b The method further includes quantizing the bit-shifted binary representation with the updated quantization scale to obtain a value of the neural network parameter.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Cross-reference of related applications)

[0001] This application claims priority to U.S. Patent Application No. 18 / 103,428, filed on 30 January 2023, also titled "SIMULATED LOW BIT-WIDTH QUANTIZATION USING BIT SHIFTED NEURAL NETWORK PARAMETERS," and claims the benefit of U.S. Provisional Patent Application No. 63 / 323,450, filed on 24 March 2022, the disclosures of which are expressly incorporated herein by reference in their entirety.

[0002]

[0002] The aspects of the present disclosure generally relate to reducing power consumption by neural networks, and more specifically to bit-shifting neural network parameters to simulate low bit-width quantization. [Background technology]

[0003]

[0003] An artificial neural network may comprise an interconnected group of artificial neurons (e.g., neuron models). An artificial neural network may be a computing device or represented as a method to be performed by a computing device. A convolutional neural network (CNN) is one type of feedforward artificial neural network. A convolutional neural network may comprise a collection of neurons, each having a receptive field and tiling the input space together. Convolutional neural networks, such as deep convolutional neural networks (DCNs), have numerous applications. Specifically, these neural network architectures are used in a variety of technologies, including image recognition, speech recognition, acoustic scene classification, keyword spotting, autonomous driving, and other classification tasks.

[0004]

[0004] Neural networks have achieved impressive breakthroughs in various fields, but they consume a considerable amount of power. In recent years, neural network quantization (for example, running neural networks on dedicated low-bit-width integer hardware) has been used to reduce the power consumption of neural networks, as lower bit widths can result in proportionally reduced power consumption. However, designing and manufacturing dedicated low-bit-width hardware is difficult and expensive. [Overview of the Initiative]

[0005]

[0005] Each of the present disclosures is described in an independent claim. Some aspects of the present disclosures are described in dependent claims.

[0006]

[0006] In aspects of the present disclosure, a method implemented by a processor includes the processor bit-shifting the binary representation of the neural network parameters. The neural network parameters have fewer bits b than the number of hardware bits B supported by the hardware processing the neural network parameters. The bit shift is performed on the neural network parameters. B-b The method implemented by the processor is also to obtain an updated quantization scale, so the processor multiplies the quantization scale by 2. B-b This includes division by a factor. The method implemented by the processor further includes the processor quantizing the bit-shifted binary representation using an updated quantization scale in order to obtain the values ​​of the neural network parameters.

[0007]

[0007] Other aspects of the present disclosure relate to a device. The device comprises a memory and one or more processors coupled to the memory. The processor(s) are configured to bit-shift the binary representation of neural network parameters. The neural network parameters have fewer bits b than the number of hardware bits B supported by the hardware processing the neural network parameters. The bit shift is performed on the neural network parameters. B-b Effectively multiplies. To obtain the updated quantization scale, the processor(s) also multiplies the quantization scale by 2. B-b It is configured to divide by . The processor(s) are further configured to quantize the bit-shifted binary representation using the updated quantization scale in order to obtain the values ​​of the neural network parameters.

[0008]

[0008] Other aspects of the present disclosure relate to an apparatus. The apparatus includes means for bit-shifting a binary representation of a neural network parameter. The neural network parameter has fewer bits b than the number of hardware bits B supported by the hardware processing the neural network parameter. The bit shift is performed on the neural network parameter. B-b It effectively multiplies. The device also doubles the quantization scale to obtain an updated quantization scale. B-b The apparatus includes means for division by a factor. The apparatus further includes means for quantizing a bit-shifted binary representation using an updated quantization scale to obtain the values ​​of neural network parameters.

[0009]

[0009] In other aspects of the present disclosure, a non-temporary computer-readable medium recording program code is disclosed. The program code is executed by a processor and includes program code for bit-shifting the binary representation of neural network parameters. The neural network parameters had fewer bits b than the number of hardware bits B supported by the hardware processing the neural network parameters. The bit shift gives the neural network parameters 2 B-b The program code also multiplies the quantization scale by 2 to obtain the updated quantization scale. B-b The program code includes program code for division by a factor of 1. The program code further includes program code for quantizing the bit-shifted binary representation using the updated quantization scale to obtain the values ​​of the neural network parameters.

[0010]

[0010] Additional features and advantages of the present disclosure are described below. It should be understood by those skilled in the art that the present disclosure may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. It should also be recognized by those skilled in the art that such equivalent constructions do not depart from the teachings of the present disclosure as set forth in the appended claims. Novel features that are considered to be characteristics of the present disclosure, both as to its organization and method of operation, will be better understood from the following description when considered in connection with the accompanying figures, which are provided for purposes of illustration and description only and are not intended to limit the scope of the present disclosure. However, it should be clearly understood that each of the figures is provided for illustrative and explanatory purposes only and does not define the scope of the present disclosure.

Brief Description of the Drawings

[0011]

[0011] The features, nature, and advantages of the present disclosure will become more apparent from the detailed description set forth below when read in conjunction with the drawings in which like reference numerals identify correspondingly throughout. [Figure 1]

[0012] An exemplary implementation of a neural network using a system-on-a-chip (SOC) including a general-purpose processor according to some aspects of the present disclosure is shown. [Figure 2A]

[0013] A diagram showing a neural network according to an aspect of the present disclosure. [Figure 2B] A diagram showing a neural network according to an aspect of the present disclosure. [Figure 2C] A diagram showing a neural network according to an aspect of the present disclosure. [Figure 2D]

[0014] A diagram showing an exemplary deep convolutional network (DCN) according to an aspect of the present disclosure. [Figure 3]

[0015] A block diagram showing an exemplary deep convolutional network (DCN) according to an aspect of the present disclosure. [Figure 4]

[0016] Block diagram showing an exemplary software architecture that allows for the modularization of artificial intelligence (AI) functions according to the aspects of this disclosure. [Figure 5]

[0017] This is a process flow diagram illustrating a method for reducing the power consumption of a neural network by bit shifting in a processor, according to the aspects of this disclosure. [Modes for carrying out the invention]

[0012]

[0018] The modes for carrying out the invention described below with respect to the attached drawings illustrate various configurations and do not represent only configurations in which the described concepts can be implemented. The “modes for carrying out the invention” include specific details intended to provide a complete understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts can be implemented without these specific details. In some cases, well-known structures and components are shown in block diagram form to avoid obscuring such concepts.

[0013]

[0019] Those skilled in the art will understand, based on the teachings, that the scope of this disclosure is intended to encompass any aspect of the disclosure, whether implemented independently of or in combination with any other aspect of the disclosure. For example, an apparatus may be implemented or a method may be practiced using any number of the described aspects. In addition, the scope of this disclosure shall encompass any such apparatus or method practiced using other structures, functions, or structures and functions in addition to or other than the various aspects of the disclosure described. It will be understood that any aspect of the disclosed disclosure may be embodied by one or more elements of the claims.

[0014]

[0020] The word "exemplary" is used to mean "serving as an example, case, or illustration." Any aspect described as "exemplary" should not necessarily be interpreted as being preferable or more advantageous than other aspects.

[0015]

[0021] While specific embodiments are described, many variations and substitutions of these embodiments fall within the scope of this disclosure. While some advantages and benefits of preferred embodiments are stated, the scope of this disclosure is not limited to any particular advantage, use, or purpose. Rather, the embodiments of this disclosure are broadly applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated as examples in the figures and the following description of preferred embodiments. The detailed description and drawings are illustrative and not limiting, and the scope of this disclosure is defined by the appended claims and their equivalents.

[0016]

[0022] Some conventional matrix multiplication hardware accelerators divide large matrix multiplications into chunks of weights and activations. These chunks can be computed in a dedicated multiply-accumulate (MAC) array, which can reduce power consumption as a result of data transfer via enhanced locality. If a bit in an individual MAC array multiplier unit is zero for two or more consecutive cycles (referred to as "consecutive zero bits"), the multiplier may consume less power.

[0017]

[0023] Aspects of this disclosure introduce low-bitwidth quantization using the most significant bits (MSBs). In these aspects, the neural network parameter bits are shifted such that the least significant bits (LSBs) are always 0. These aspects can be applied to any type of neural network.

[0018]

[0024] According to an aspect of the present disclosure, in the case of b bits on B-bit hardware where b < B, all values are bit-shifted by only B - b bits. In other words, all values can be multiplied by 2 (B-b) . This multiplication can be canceled by dividing by the corresponding quantization scale by 2 (B-b) . This approach can ensure that the least significant bits of B - b are always 0 for both positive and negative values.

[0019]

[0025] Aspects of the present disclosure may be applicable to per-channel quantization or to any per-block quantization technique. Further, aspects of the present disclosure can enable different simulated bit widths in different channels / blocks. In this case, each block may be shifted by a different number of bits, and each corresponding quantization scale is divided by a different correction factor 2 (b’-B) , where b’ represents the per-channel / per-block bit width.

[0020]

[0026] Thus, aspects of the present disclosure can be advantageously employed on existing hardware without explicit support for (simulated) low-bitwidth quantization.

[0021]

[0027] Figure 1 shows an exemplary implementation of a system-on-a-chip (SOC) 100, which may include a central processing unit (CPU) 102 or a multicore CPU equipped with a low-precision multiplier for evaluating a low-bit-width quantized neural network. Variables (e.g., neural signals and synaptic weights), system parameters associated with computing devices (e.g., neural networks with weights), delays, frequency bin information, and task information may be stored in a memory block associated with the neural processing unit (NPU) 108, a memory block associated with the CPU 102, a memory block associated with the graphics processing unit (GPU) 104, a memory block associated with the digital signal processor (DSP) 106, or in memory block 118, or distributed across multiple blocks. Instructions executed in the CPU 102 may be loaded from program memory associated with the CPU 102 or from memory block 118.

[0022]

[0028] The SOC100 may also include a connectivity block 110 which may include a GPU 104, a DSP 106, fifth-generation (5G) connectivity, fourth-generation long-term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, and additional processing blocks adapted to specific functions, such as a multimedia processor 112 which may detect and recognize gestures. In one implementation, the NPU 108 is implemented in the CPU 102, DSP 106, and / or GPU 104. The SOC100 may also include a navigation module 120 which may include a sensor processor 114, image signal processors (ISPs) 116, and / or a global positioning system.

[0023]

[0029] The SOC100 may be based on the ARM instruction set. In aspects of this disclosure, instructions loaded into the general-purpose processor 102 may include code for bit-shifting the binary representation of neural network parameters. The general-purpose processor 102 may also quantize the scale to obtain an updated quantization scale. B-b The general-purpose processor 102 may also include code for division by a certain value.

[0024]

[0030] A deep learning architecture may perform object recognition tasks by learning to represent the input at successively higher levels of abstraction within each layer, thereby constructing a useful feature representation of the input data. In this way, deep learning addresses a major bottleneck in traditional machine learning. Prior to the advent of deep learning, machine learning methods for object recognition problems may have relied heavily on human-designed features, sometimes in combination with shallow classifiers. A shallow classifier may be, for example, a two-class linear classifier that can compare a weighted sum of feature vector components to a threshold to predict which class the input belongs to. Human-designed features may be templates or kernels adapted to a specific problem domain by engineers with domain expertise. In contrast, a deep learning architecture may learn, through training, to represent features that are similar to those that human engineers can design. Furthermore, deep networks may learn to represent and recognize novel types of features that humans may not have considered.

[0025]

[0031] A deep learning architecture may learn a hierarchy of features. If visual data is presented, for example, the first layer may learn to recognize relatively simple features such as edges in the input stream. In another example, if auditory data is presented, the first layer may learn to recognize spectral power at a specific frequency. A second layer, taking the output of the first layer as input, may learn to recognize combinations of features such as simple shapes in the case of visual data, or combinations of sounds in the case of auditory data. For example, a higher layer may learn to represent complex shapes in visual data or words in auditory data. A higher layer may learn to recognize common visual objects or spoken phrases.

[0026]

[0032] Deep learning architectures can sometimes work particularly well when applied to problems with natural hierarchical structures. For example, classifying electric vehicles may benefit from initially learning to recognize wheels, windshields, and other features. These features may then be combined in different ways in higher layers to recognize cars, trucks, and airplanes.

[0027]

[0033] Neural networks may be designed using various connectivity patterns. In a feedforward network, each neuron in a given layer communicates with neurons in higher layers to pass information from lower to higher layers. As mentioned above, a hierarchical representation may be constructed within the consecutive layers of a feedforward network. Neural networks may also have recursive or feedback (also called top-down) connectivity. In recursive connectivity, the output from a neuron in a given layer may be transmitted to another neuron in the same layer. Recursive architectures can be useful when recognizing patterns across two or more chunks of input data delivered to the neural network in a sequence. Connectivity from a neuron in a given layer to a neuron in a lower layer is called feedback (or top-down) connectivity. Networks with a large number of feedback connectivitys can be useful when the recognition of a higher-level concept can help discriminate certain lower-level features of the input.

[0028]

[0034] The connections between layers of a neural network may be fully connected or locally connected. Figure 2A shows an example of a fully connected neural network 202. In the fully connected neural network 202, neurons in the first layer may transmit their output to any neuron in the second layer, so that each neuron in the second layer receives input from any neuron in the first layer. Figure 2B shows an example of a locally connected neural network 204. In the locally connected neural network 204, neurons in the first layer may be connected to a limited number of neurons in the second layer. More generally, the locally connected layers of the locally connected neural network 204 may be configured so that each neuron in a given layer has the same or similar connectivity pattern, but with different connection strengths (e.g., 210, 212, 214, and 216). Since higher-layer neurons within a given region can receive inputs that, through training, are tuned to the characteristics of a limited portion of the total inputs to the network, the connectivity patterns of local connections can give rise to spatially distinct receptive fields within the higher layers.

[0029]

[0035] An example of a locally connected neural network is a convolutional neural network. Figure 2C shows an example of a convolutional neural network 206. The convolutional neural network 206 may be configured such that the connection strengths associated with the input for each neuron in the second layer are shared (e.g., 208). Convolutional neural networks may be suitable for problems where the spatial location of the input is meaningful.

[0030]

[0036] One type of convolutional neural network is the deep convolutional network (DCN). Figure 2D shows a detailed example of DCN200, designed to recognize visual features from an image 226 input from an image capture device 230, such as an in-vehicle camera. DCN200 in this example may be trained to identify traffic signs and the numbers written on them. Of course, DCN200 may also be trained for other tasks, such as identifying lane markings or traffic signals.

[0031]

[0037] DCN200 may be trained using supervised learning. During training, DCN200 may be presented with images, such as image 226 of a speed limit sign, and then a forward pass may be computed to produce output 222. DCN200 may include a feature extraction section and a classification section. Upon receiving image 226, the convolutional layer 232 may apply a convolutional kernel (not shown) to image 226 to produce a first set 218 of feature maps. For example, the convolutional kernel for convolutional layer 232 may be a 5x5 kernel that produces a 28x28 feature map. In this example, four different feature maps are produced in the first set 218 of feature maps, so four different convolutional kernels were applied to image 226 in convolutional layer 232. Convolutional kernels are also sometimes called filters or convolutional filters.

[0032]

[0038] The first set of feature maps 218 may be subsampled by a max pooling layer (not shown) to generate a second set of feature maps 220. The max pooling layer reduces the size of the first set of feature maps 218; that is, the size of the second set of feature maps 220, such as 14×14, is smaller than the size of the first set of feature maps 218, such as 28×28. The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 220 may be further convolved through one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).

[0033]

[0039] In the example in Figure 2D, a second set of feature maps 220 is convolved to generate a first feature vector 224. Further convolved, the first feature vector 224 generates a second feature vector 228. Each feature in the second feature vector 228 may contain a number corresponding to a possible feature of image 226, such as "label", "60", and "100". A softmax function (not shown) may convert the numbers in the second feature vector 228 into probabilities. Thus, the output 222 of DCN200 is the probability that image 226 contains one or more features.

[0034]

[0040] In this example, the probabilities in output 222 for "label" and "60" are higher than the probabilities for other outputs 222 such as "30", "40", "50", "70", "80", "90", and "100". Before training, the outputs 222 generated by DCN200 may be inaccurate. Therefore, an error may be calculated between output 222 and the target output. The target output is the ground truth of image 226 (e.g., "label" and "60"). The weights of DCN200 may then be adjusted so that the output 222 of DCN200 is more closely matched to the target output.

[0035]

[0041] To adjust the weights, the learning algorithm may calculate gradient vectors for the weights. The gradient can indicate the amount by which the error will increase or decrease when the weights are adjusted. In the top layer, the gradient can directly correspond to the weight values ​​connecting the activated neurons in the second-to-last layer to the neurons in the output layer. In lower layers, the gradient can depend on the weight values ​​and the error gradient calculated in the upper layers. The weights can then be adjusted to reduce the error. This method of adjusting weights is sometimes called "backpropagation" because it involves a "backward path" through the neural network.

[0036]

[0042] In practice, the error gradient of the weights may be calculated over a small number of examples so that the calculated gradient approximates the true error gradient. This approximation method is sometimes called stochastic gradient descent. Stochastic gradient descent can be repeated until the achievable error rate for the entire system stops decreasing or until the error rate reaches a target level. After training, DCN200 may be presented with new images, and a forward pass through the DCN200 network may yield an output 222 that can be considered DCN200's inference or prediction.

[0037]

[0043] Deep belief networks (DBNs) are probabilistic models with multiple layers of hidden nodes. DBNs may be used to extract a hierarchical representation of a training dataset. DBNs may also be obtained by stacking layers of restricted Boltzmann machines (RBMs). RBMs are a type of artificial neural network that can learn a probability distribution over a set of inputs. Because RBMs can learn a probability distribution without any prior information about the class to which each input should be categorized, RBMs are frequently used in unsupervised learning. Using a hybrid paradigm of unsupervised and supervised learning, the lower RBM of a DBN can be trained unsupervised and can function as a feature extractor, while the upper RBM can be trained supervised (on a joint distribution of inputs from previous layers and the target class) and can function as a classifier.

[0038]

[0044] Deep convolutional networks (DCNs) are networks of convolutional networks constructed with additional pooling and normalization layers. DCNs achieve state-of-the-art performance for a wide range of tasks. DCNs can be trained using supervised learning, where both the input and output targets are known for a large number of samples, and the network weights are modified using gradient descent.

[0039]

[0045] A DCN may also be a feedforward network. In addition, as described above, connections from neurons in the first layer of a DCN to groups of neurons in the next higher layer are shared across the neurons in the first layer. The feedforward and shared connections of a DCN can be leveraged for high-speed processing. The computational burden of a DCN may be significantly less than that of a similarly sized neural network, for example, one that includes recursive or feedback connections.

[0040]

[0046] The processing of each layer of a convolutional network may be considered as a spatially invariant template or basis projection. If the input is initially decomposed into multiple channels, such as the red, green, and blue channels of a color image, the convolutional network trained on that input may be considered three-dimensional, having two spatial dimensions along the image axes and a third dimension that captures color information. The output of the convolutional connections may be considered as forming a feature map in subsequent layers, where each element of the feature map (e.g., 220) receives input from a range of neurons in the previous layer (e.g., feature map 218) and from each of the multiple channels. The values ​​in the feature map may be further processed using rectification, max(0,x), or other nonlinearities. Values ​​from adjacent neurons may further undergo pooling corresponding to downsampling, which can result in additional local invariance and dimensionality reduction. Normalization corresponding to whitening may also be applied through lateral inhibition between neurons in the feature map.

[0041]

[0047] The performance of deep learning architectures can improve as more labeled data points become available or as computational power increases. Modern deep neural networks are routinely trained using computational resources thousands of times greater than what was available to the typical researcher just 15 years ago. New architectures and training paradigms can further enhance the performance of deep learning. Rectified linear units may mitigate the training problem known as vanishing gradients. New training techniques may reduce overfitting, thus allowing larger models to achieve better generalization. Encapsulation techniques may extract data within a given receptive field, further improving overall performance.

[0042]

[0048] Figure 3 is a block diagram of a deep convolutional network 350. The deep convolutional network 350 may include several different types of layers based on connectivity and weight sharing. As shown in Figure 3, the deep convolutional network 350 includes convolutional blocks 354A and 354B. Each of the convolutional blocks 354A and 354B may consist of a convolutional layer (CONV) 356, a normalization layer (LNorm) 358, and a maximum pooling layer (MAX POOL) 360. Although only two of the convolutional blocks 354A and 354B are shown, this disclosure is not limited in that way, and instead, any number of convolutional blocks 354A and 354B may be included in the deep convolutional network 350 according to design preferences.

[0043]

[0049] The convolutional layer 356 may include one or more convolutional filters that can be applied to the input data to generate a feature map. The normalization layer 358 may normalize the output of the convolutional filters. For example, the normalization layer 358 may result in whitening or lateral suppression. The maximal pooling layer 360 may result in downsampling aggregation across space for local invariance and dimensionality reduction.

[0044]

[0050] For example, the parallel filter bank of the deep convolutional network may be loaded on the CPU 102 or GPU 104 of the SOC 100 (e.g., Figure 1) to achieve high performance and low power consumption. In an alternative embodiment, the parallel filter bank may be loaded on the DSP 106 or ISP 116 of the SOC 100. In addition, the deep convolutional network 350 may access other processing blocks that may reside on the SOC 100, such as a sensor processor 114 and a navigation module 120 dedicated to sensors and navigation, respectively.

[0045]

[0051] The deep convolutional network 350 may also include one or more fully connected layers 362 (FC1 and FC2). The deep convolutional network 350 may further include a logistic regression (LR) layer 364. Between each layer 356, 358, 360, 362, 364 of the deep convolutional network 350 are weights (not shown) that will be updated. The output of each layer (e.g., 356, 358, 360, 362, 364) may serve as input to one of the subsequent layers (e.g., 356, 358, 360, 362, 364) in the deep convolutional network 350 to learn a hierarchical feature representation from the initially supplied input data 352 (e.g., image, audio, video, sensor data, and / or other input data) of the convolutional block 354A. The output of the deep convolutional network 350 is a classification score 366 for the input data 352. The classification score 366 may be a set of probabilities, where each probability is the probability that the input data contains a feature from the set of features.

[0046]

[0052] Figure 4 is a block diagram showing an exemplary software architecture 400 that can modularize artificial intelligence (AI) functionality. Using this architecture 400, applications can be designed that enable various processing blocks of the SOC 420 (which may be similar to the SoC 100 in Figure 1) (e.g., CPU 422, DSP 424, GPU 426, and / or NPU 428) to support low-bit-width quantization simulation for an AI application 402 according to the embodiments of this disclosure. The architecture 400 may be included in a computing device such as a smartphone, for example.

[0047]

[0053] The AI ​​application 402 can be configured to invoke functions defined in user space 404, for example, which can provide scene detection and recognition indicating the location where a computing device including architecture 400 is currently operating. The AI ​​application 402 may configure the microphone and camera differently depending on whether the scene to be recognized is an office, an auditorium, a restaurant, or an outdoor setting such as a lake. The AI ​​application 402 may make requests for compiled program code associated with libraries defined in the AI ​​Functional Application Programming Interface (API) 406. These requests may ultimately rely on the output of a deep neural network configured to provide inference responses based on video and positioning data, for example.

[0048]

[0054] The runtime engine 408, which may be compiled code of a runtime framework, may be further accessible to the AI ​​application 402. The AI ​​application 402 may request the runtime engine 408 to perform inference, for example, at a specific time interval or triggered by an event detected by the user interface of the AI ​​application 402. The runtime engine 408 may then signal an operating system in the operating system (OS) space 410, such as a kernel 412 running on the SOC 420, when it is ready to provide an inference response. In some embodiments, the kernel 412 may be a Linux kernel. The operating system may then perform a continuous relaxation of quantization on the CPU 422, DSP 424, GPU 426, NPU 428, or any combination thereof. The CPU 422 may be accessed directly by the operating system, while the other processing blocks may be accessed through drivers, such as drivers 414, 416, or 418 for the DSP 424, GPU 426, or NPU 428, respectively. In an exemplary example, the deep neural network may be configured to operate on a combination of processing blocks such as CPU422, DSP424, and GPU426, or it may operate on NPU428.

[0049]

[0055] The AI ​​application 402 can be configured to invoke functions defined in user space 404, for example, to provide scene detection and recognition indicating the location where a computing device including architecture 400 is currently operating. The AI ​​application 402 may configure the microphone and camera differently depending on whether the scene to be recognized is an office, auditorium, restaurant, or an outdoor setting such as a lake. The AI ​​application 402 can make requests to compiled program code associated with a library defined in the SceneDetect application programming interface (API) 406 to provide an estimate of the current scene. This request may ultimately rely on the output of a differential neural network configured to provide a scene estimate based on video and positioning data, for example.

[0050]

[0056] The runtime engine 408, which may be compiled code for a runtime framework, may be further accessible to the AI ​​application 402. The AI ​​application 402 may request the runtime engine 408 to perform scene estimation, for example, at a specific time interval or triggered by an event detected by the application's user interface. The runtime engine 408 may then signal an operating system 410, such as a Linux kernel 412 running on the SOC 420, when it has performed scene estimation. The operating system 410 may then perform computations on the CPU 422, DSP 424, GPU 426, NPU 428, or any combination thereof. The CPU 422 may be accessed directly by the operating system, while the other processing blocks may be accessed through drivers, such as drivers 414-418 for the DSP 424, GPU 426, or NPU 428. In an exemplary embodiment, the differential neural network may be configured to run on a combination of processing blocks such as the CPU 422 and GPU 426, or on the NPU 428.

[0051]

[0057] Generally speaking, neural networks can consume a considerable amount of power. High power consumption in computing cores can lead to shorter battery life. As a result, computing cores may be power-thinned, or in other words, their clock speed may be reduced to lower power consumption.

[0052]

[0058] A neural signal processor (NSP) multiplier consumes less power when the received input is zero for two consecutive cycles. This effect operates at the bit level; that is, power consumption is reduced when the same bit is zero in the same multiplier for two consecutive cycles. Using a lower bit width on hardware with a higher bit width increases the number of bits that have zero values. For example, when using a 4-bit value on 8-bit hardware, for a positive value, the four most significant bits are always zero. A 4-bit integer (INT4) has an unsigned range: [0, 15] in decimal, which is [00000000, 00001111] in binary for an 8-bit integer (INT8).

[0053]

[0059] However, this solution does not work for signed negative values ​​because negative integers are encoded using two's complement coding. Continuing with the example of a 4-bit value on 8-bit hardware in decimal, the signed INT4 range is [-8, 7], which in INT8 binary is [11111000, 00000111]. It is desirable that some bits are always 0. However, this is not possible with existing solutions when the range of possible values ​​includes negative numbers.

[0054]

[0060] To utilize efficient low-bit integer hardware, networks can be quantized to an appropriate bit width. As an example, 8-bit quantization of a floating-point tensor X can be given by the following equation:

[0055]

number

[0056] In the formula, (int_min, int_max) is (0, 255) for an unsigned tensor and (-128, 127) for a signed tensor, s xis the quantization scale for a tensor X, set heuristically or by gradient descent on some target loss function. The matrix-matrix product WX can be approximated as follows:

[0057]

number

[0058] In the formula, W represents the weight, s w This is the quantization scale of the weights W. Only integer multiplication can be used to multiply the product W. int X int Since this calculation is used to compute [the result], it can be performed on efficient low-precision hardware. Matrix-matrix multiplication is used, but this is merely an example for illustrative purposes. Alternatively, certain extensions such as convolution and channel-by-channel quantization can also be used.

[0059]

[0061] To improve quantized performance, quantization-aware training (QAT) can be employed. In some conventional methods, (non-differentiable) quantization operations are used in the forward pass but ignored in the reverse pass. However, these conventional quantization procedures introduce noise into the weights and activation tensors. Since some neural network layers are more sensitive to noise than others, some conventional methods utilize mixed-precision quantization (MPQ). Mixed-precision quantization uses tensor-specific bit widths instead of fixed bit widths for each tensor in the network. However, the space resulting from an MPQ configuration increases exponentially with the number of layers in the network, eliminating exhaustive search.

[0060]

[0062] The computation of a neural network layer can be viewed as a matrix-matrix multiplication Y = WX between input activations X and weights W. The product WX involves multiplying individual elements in W by elements in X, and then adding (accumulating) the resulting scalar products. A neural network accelerator can be implemented in hardware as a multiplicative-accumulation (MAC) array, where a subset of weights is multiplied in parallel with a subset of activations in each cycle and then accumulated. For example, with 16 weights W as input... 1:4,1:4 and four activated X 1:4,k A MAC array having this will have 16 multipliers and 4 accumulators. In each cycle, the scalar product W nm X mk is a multiplier M nm It can be computed in parallel for (n,m∈[1,4]). Y 1:4,k The (partial) result for this can be stored in the accumulator.

[0061]

[0063] Multiplier M mn If one of the inputs to is 0 for two or more consecutive cycles, multiplier M mn In some cases, no power is consumed. This power saving occurs at the bit level. That is, if the same input bit in the multiplier is 0 for two consecutive cycles, the activity in the gate toggling the bit between 0 and 1 can be reduced, and in some embodiments, avoided. Thus, the power consumption per such input bit that remains 0 for consecutive cycles can be significantly reduced, and in some embodiments, avoided.

[0062]

[0064] Therefore, in order to reduce power consumption, aspects of this disclosure aim to increase the number of consecutive zeros. According to aspects of this disclosure, low-bit-width (e.g., 4-bit) quantization can be simulated in hardware with higher bit widths (e.g., 8-bit or 16-bit). Since the activation tensor in a rectifier linear unit (ReLU) network can be sparser than its corresponding weight, significant gains can be achieved by increasing the number of consecutive zeros in the weight. However, this is only an example and not limiting. Aspects of this disclosure can be applied to other tensors and other types of neural networks.

[0063]

[0065] Low-bit-width weights can be simulated on higher-bit-width hardware by limiting the range of (integer) values ​​that the weights can take. For example, simulating 4-bit integer quantization on 8-bit hardware can be achieved by limiting the range of integer weight values ​​to [-8, 7]. However, due to sign extension resulting from two's complement coding of negative numbers, negative values ​​can be represented by 1 in the most significant bit (MSBs). This property can be undesirable, as one objective is to increase, and in some aspects maximize, the number of consecutive 0 bits.

[0064]

[0066] To avoid sign extension in two's complement coding, the low-bit integer weights can be bit-shifted by an appropriate amount instead. Bit-shifting an integer by b bits is 2 b This can be considered equivalent to multiplying by 2. Therefore, when simulating signed b-bit quantization on B-bit hardware, each value is multiplied by 2. B-b By multiplying by this, each value can be effectively bit-shifted by Bb bits. To offset the effects of the bit shift, the quantization scale s is:

[0065]

number

[0066] It can be adjusted accordingly.

[0067]

[0067] In one embodiment, a signed 4-bit weight can be simulated on 8-bit hardware. A signed weight tensor W having values ​​in the range [-8, 7] int and associated scales w This can be specified. The binary representation of the range [-8,7] can be given by [11111000,00000111]. In this example, consecutive zero bits can only occur if the consecutive values ​​are either both positive or both negative. To avoid this scenario,

[0068]

number

[0069] It may also be defined as, and thereby, W int Each value within may be effectively shifted by 4 bits. Doing so can result in a decimal representation ranging from [-128, 112] or a binary representation of [10000000, 01110000]. Since the resulting value is a multiple of 16, the least significant 4 bits here become 0, regardless of the represented value. In some embodiments, the effect of multiplication by 16 is the quantization scale.

[0070]

number

[0071] This can be canceled by using . Therefore, the resulting integer matrix product remains mathematically equivalent, as given by the following:

[0072]

number

[0073]

[0068] Figure 5 is a process flow diagram showing a processor-based method 500 for reducing the power consumption of a neural network by bit shifting, according to an aspect of the present disclosure. In some aspects, the processor-based method 500 can be performed by a processor such as a CPU 102 or an NPU 108. As shown in Figure 5, in block 502, the processor can bit shift the binary representation of the neural network parameters. The neural network parameters have fewer bits b than the number of hardware bits B supported by the hardware processing the neural network parameters. The bit shift is performed on the neural network parameters. B-b Multiply effectively.

[0074]

[0069] In block 504, in order to obtain the updated quantization scale, the processor sets the quantization scale to 2 B-b Divide by . As explained, in order to offset the effects of bit shifts, the quantization scale s is

[0075]

number

[0076] It can be adjusted accordingly.

[0077]

[0070] In block 506, in order to obtain the values ​​of the neural network parameters, the processor quantizes the bit-shifted binary representation using the updated quantization scale.

[0078] Exemplary aspects

[0071] Embodiment 1: A processor bit-shifts the binary representation of a neural network parameter having fewer bits b than the number of hardware bits B supported by the hardware processing the neural network parameter, wherein the neural network parameter has 2 B-b The processor multiplies the quantization scale by 2 to effectively multiply, perform a bit shift, and obtain an updated quantization scale. B-b A method performed by a processor, which includes dividing by and quantizing the bit-shifted binary representation using an updated quantization scale by the processor in order to obtain the values ​​of the neural network parameters.

[0079]

[0072] Embodiment 2: A method to be implemented using the processor according to Embodiment 1, wherein the neural network parameters have 7 bits or less and the hardware supports 8-bit values.

[0080]

[0073] Embodiment 3: A method to be implemented with the processor according to Embodiment 1 or 2, wherein the neural network parameters have 4 bits and the hardware supports 8-bit values.

[0081]

[0074] Embodiment 4: A method of implementing the method using the processor according to any one of Embodiments 1 to 3, wherein the binary representation of the neural network parameters is bit-shifted such that the difference Bb of the least significant bits has a value of 0.

[0082]

[0075] Embodiment 5: A method of performing the neural network parameter using the processor according to any one of Embodiments 1 to 4, wherein the neural network parameter includes a neural network weight tensor or a neural network activation tensor.

[0083]

[0076] Embodiment 6: A device comprising a memory and at least one processor coupled to the memory, wherein the at least one processor is a bit shift of a binary representation of a neural network parameter, the binary representation of the neural network parameter having fewer bits b than the number of hardware bits B supported by the hardware processing the neural network parameter, and the neural network parameter has 2 B-b To effectively multiply, perform a bit shift and obtain the updated quantization scale, set the quantization scale to 2. B-b A device configured to divide by and quantize the bit-shifted binary representation using an updated quantization scale to obtain the values ​​of the neural network parameters.

[0084]

[0077] Embodiment 7: The apparatus according to Embodiment 6, wherein the neural network parameters have 7 bits or less, and the hardware supports 8-bit values.

[0085]

[0078] Embodiment 8: The apparatus according to Embodiment 6 or 7, wherein the neural network parameters have 4 bits and the hardware supports 8-bit values.

[0086]

[0079] Embodiment 9: The apparatus according to any one of Embodiments 6 to 8, wherein at least one processor is further configured to bit-shift the binary representation of neural network parameters such that the difference Bb of the least significant bits has a value of 0.

[0087]

[0080] Embodiment 10: The apparatus according to any one of Embodiments 6 to 9, wherein the neural network parameters include a neural network weight tensor or a neural network activation tensor.

[0088]

[0081] Embodiment 11: A means for bit-shifting the binary representation of a neural network parameter, wherein the neural network parameter has fewer bits b than the number of hardware bits B supported by the hardware processing the neural network parameter, and the neural network parameter has 2 B-b A means of effectively multiplying, bit shifting, and obtaining an updated quantization scale, the quantization scale is set to 2 B-b An apparatus comprising: a means for division by; and a means for quantizing a bit-shifted binary representation using an updated quantization scale in order to obtain the values ​​of neural network parameters.

[0089]

[0082] Embodiment 12: The apparatus according to Embodiment 11, wherein the neural network parameters have 7 bits or less, and the hardware supports 8-bit values.

[0090]

[0083] Embodiment 13: The apparatus according to Embodiment 11 or 12, wherein the neural network parameters have 4 bits and the hardware supports 8-bit values.

[0091]

[0084] Embodiment 14: The apparatus according to any one of Embodiments 11 to 13, further comprising means for bit-shifting the binary representation of neural network parameters such that the difference Bb of the least significant bits has a value of 0.

[0092]

[0085] Embodiment 15: The apparatus according to any one of Embodiments 11 to 14, wherein the neural network parameters include a neural network weight tensor or a neural network activation tensor.

[0093]

[0086] Embodiment 16: A non-temporary computer-readable medium recording program code, wherein the program code is executed by a processor and is a bit shift of a binary representation of a neural network parameter having fewer bits b than the number of hardware bits B supported by the hardware that processes the neural network parameter, and the neural network parameter has 2 B-b Program code for bit shifting, which effectively multiplies the quantization scale, and to obtain the updated quantization scale, the quantization scale is set to 2. B-b A non-temporary computer-readable medium including program code for division by and program code for quantizing a bit-shifted binary representation using an updated quantization scale to obtain the values ​​of neural network parameters.

[0094]

[0087] Embodiment 17: The non-temporary computer-readable medium according to Embodiment 16, wherein the neural network parameters have 7 bits or less and the hardware supports 8-bit values.

[0095]

[0088] Embodiment 18: A non-temporary computer-readable medium according to Embodiment 16 or 17, wherein the neural network parameters have 4 bits and the hardware supports 8-bit values.

[0096]

[0089] Embodiment 19: A non-temporary computer-readable medium according to any one of Embodiments 16 to 18, further comprising program code for bit-shifting a binary representation of neural network parameters such that the difference Bb of the least significant bits has a value of 0.

[0097]

[0090] Embodiment 20: A non-temporary computer-readable medium according to any one of Embodiments 16 to 19, wherein the neural network parameters include a neural network weight tensor or a neural network activation tensor.

[0098]

[0091] In one embodiment, the bit shifting means, the division means, and / or quantizing means may be a CPU 102, a program memory associated with the CPU 102, a dedicated memory block 118, a fully connected layer 362, an NPU 428, and / or a routing connection processing unit 216 configured to perform the listed functions. In another configuration, the means described above may be any module or any device configured to perform the functions listed by the means described above.

[0099]

[0092] The various operations of the above-described method can be performed by any preferred means capable of performing the corresponding function. The means may include, but are not limited to, various hardware and / or software components and / or modules, including circuits, application-specific integrated circuits (ASICs), or processors. Generally, where there are operations shown in the figures, those operations may have corresponding relative means-plus-function components that are similarly numbered.

[0100]

[0093] When used, the term “determining” encompasses a wide variety of actions. For example, “determining” may include calculating, calculating, processing, deriving, investigating, looking up (e.g., looking up in a table, database, or other data structure), and confirming. In addition, “determining” may include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), and so on. Furthermore, “determining” may include resolving, selecting, choosing, and establishing.

[0101]

[0094] When used, the phrase “at least one of” an enumeration of items refers to any combination of those items that includes a single member. For example, “at least one of a, b, or c” would include a, b, c, ab, ac, bc, and abc.

[0102]

[0095] Various exemplary logic blocks, modules, and circuits described in this disclosure may be implemented or run using general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices (PLDs), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described. The general-purpose processor may be a microprocessor, but alternatively, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0103]

[0096] The steps of the methods or algorithms described in this disclosure may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in any form of storage medium known in the art. Some examples of storage mediums that may be used include random access memory (RAM), read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs, and the like. The software module may consist of a single instruction or a number of instructions and may be distributed across several different code segments, between different programs, and across multiple storage mediums. The storage medium may be coupled to the processor so that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integrated with the processor.

[0104]

[0097] The disclosed methods include one or more steps or actions for achieving the described method. The steps and / or actions of those methods can be replaced with one another without departing from the claims. In other words, unless a particular order of steps or actions is specified, the order of any particular steps and / or actions, and / or the use of those steps and / or actions, can be modified without departing from the claims.

[0105]

[0098] The functions described may be implemented in hardware, software, firmware, or any combination thereof. When implemented in hardware, the exemplary hardware configuration may include a processing system within the device. The processing system may be implemented using a bus architecture. The bus may include any number of interconnection buses and bridges, depending on the specific application of the processing system and the overall design constraints. The bus may link various circuits to each other, including processors, machine-readable media, and bus interfaces. The bus interface may, among other things, be used to connect a network adapter to the processing system via the bus. The network adapter may be used to implement signal processing functions. In some embodiments, a user interface (e.g., a keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also link various other circuits, such as timing sources, peripherals, voltage regulators, and power management circuits, but these circuits are well known in the art and are therefore not described further.

[0106]

[0099] The processor may be responsible for managing the bus and general processing, including executing software stored on a machine-readable medium. The processor may be implemented using one or more general-purpose processors and / or dedicated processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuits capable of executing software. Software is broadly interpreted to mean instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or by any other name. Machine-readable medium may include, for example, random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, magnetic disks, optical disks, hard drives, or any other suitable storage medium, or any combination thereof. The machine-readable medium may be embodied in a computer program product. The computer program product may include packaging materials.

[0107]

[0100] In hardware implementations, the machine-readable medium may be part of a processing system separate from the processor. However, as will be readily apparent to those skilled in the art, the machine-readable medium or any part thereof may be outside the processing system. For example, the machine-readable medium may include transmission lines, data-modulated carriers, and / or computer products separate from the device, all of which may be accessed by the processor through a bus interface. Alternatively, or in addition, the machine-readable medium or any part thereof may be integrated into the processor, such as caches and / or general-purpose register files. The various components discussed may be described as having a specific location, such as local components, but these components may also be configured in various ways, such as several components configured as part of a distributed computing system.

[0108]

[0101] The processing system may be configured as a general-purpose processing system having one or more microprocessors providing processor functions, all of which are linked to each other with other support circuits through an external bus architecture, and external memory providing at least a portion of a machine-readable medium. Alternatively, the processing system may comprise one or more neuromorphological processors for implementing the described neuron model and neural system model. Another alternative is that the processing system may be implemented using an application-specific integrated circuit (ASIC) having a processor, a bus interface, a user interface, support circuits, and at least a portion of a machine-readable medium integrated on a single chip, or using one or more field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gate logic, individual hardware components, or any other suitable circuits, or any combination of circuits capable of performing the various functions described throughout this disclosure. Those skilled in the art will recognize how to best implement the functions described for the processing system depending on the specific application and the overall design constraints imposed on the entire system.

[0109]

[0102] The machine-readable medium may comprise several software modules. The software modules, when executed by the processor, contain instructions that cause the processing system to perform various functions. The software modules may also include transmit modules and receive modules. Each software module may reside in a single storage device or be distributed across multiple storage devices. For example, a software module may be loaded from a hard drive into RAM when a trigger event occurs. While a software module is executing, the processor may load some of the instructions into a cache to increase access speed. One or more cache lines may then be loaded into a general-purpose register file for execution by the processor. When the functions of a software module are referred to below, it will be understood that such functions are implemented by the processor when it executes instructions from that software module. Furthermore, it should be understood that aspects of this disclosure result in improvements to the functionality of processors, computers, machines, or other systems that implement such aspects.

[0110]

[0103] When implemented in software, the functions may be stored on or transmitted via computer-readable media as one or more instructions or codes. Computer-readable media include both computer storage media and communication media, including any media that facilitates the transfer of computer programs from one place to another. The storage media may be any available media that can be accessed by a computer. Such computer-readable media may include, but are not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other media that can be accessed by a computer and that can be used to carry or store desired program code in the form of instructions or data structures. In addition, any connection is appropriately referred to as computer-readable media. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared (IR), radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. When used, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray® disc, where a disk typically reproduces data magnetically, and a disc optically reproduces data using a laser. Thus, in some embodiments, computer-readable medium may include non-temporary computer-readable medium (e.g., tangible medium). In addition, in other embodiments, computer-readable medium may include temporary computer-readable medium (e.g., signals). The combinations of the above are also considered to fall within the scope of computer-readable media.

[0111]

[0104] Accordingly, some embodiments may include a computer program product for performing the operations described. For example, such a computer program product may include a computer-readable medium storing (and / or encoding) instructions, the instructions being executable by one or more processors to perform the operations described. In some embodiments, the computer program product may include packaging material.

[0112]

[0105] Furthermore, it should be understood that modules and / or other suitable means for performing the described methods and techniques may be downloaded and / or otherwise obtained by user terminals and / or base stations, where applicable. For example, such devices may be coupled to a server to facilitate the transfer of means for performing the described methods. Alternatively, the various methods described may be provided via storage means so that user terminals and / or base stations can obtain the various methods by coupling or providing storage means (e.g., physical storage media such as RAM, ROM, compact disks (CDs) or floppy disks) to the device. Furthermore, any other suitable techniques for providing the described methods and techniques to the device may be utilized.

[0113]

[0106] It should be understood that the claims are not limited to the exact configurations and components illustrated above. Various modifications, changes, and variations may be made to the configuration, operation, and details of the methods and apparatus described above without departing from the claims. The invention described in the original claims of this application is listed below. [C1] The processor performs a bit shift on the binary representation of a neural network parameter having fewer bits b than the number of hardware bits B supported by the hardware processing the neural network parameter, thereby 2 B-b Effectively multiplying, bit shifting, To obtain an updated quantization scale, the processor sets the quantization scale to 2 B-b Dividing by and In order to obtain the values ​​of the neural network parameters, the processor quantizes the bit-shifted binary representation using the updated quantization scale, A method that includes implementing it on the processor. [C2] A method of implementation using the processor described in C1, wherein the neural network parameters have 7 bits or less, and the hardware supports 8-bit values. [C3] A method of implementing the process described in C1 using a processor, wherein the neural network parameters have 4 bits and the hardware supports 8-bit values. [C4] A method of performing the process on the processor described in C1, wherein the binary representation of the neural network parameters is bit-shifted such that the difference Bb of the least significant bits is 0. [C5] A method performed by the processor described in C1, wherein the neural network parameters include a neural network weight tensor or a neural network activation tensor. [C6] A device, Memory and The system comprises at least one processor coupled to the memory, wherein the at least one processor is A bit shift of the binary representation of a neural network parameter, wherein the number of bits b is less than the number of hardware bits B supported by the hardware processing the neural network parameter, and the number of bits b is less than the number of hardware bits B supported by the hardware processing the neural network parameter. B-b Perform a bit shift to effectively multiply, To obtain the updated quantization scale, set the quantization scale to 2 B-b Divide by, To obtain the values ​​of the neural network parameters, the bit-shifted binary representation is quantized using the updated quantization scale. It is structured in such a way. Device. [C7] The apparatus according to C6, wherein the neural network parameters have 7 bits or less, and the hardware supports 8-bit values. [C8] The apparatus according to C6, wherein the neural network parameters have 4 bits and the hardware supports 8-bit values. [C9] The apparatus according to C6, wherein at least one processor is further configured to bit-shift the binary representation of the neural network parameters such that the difference Bb of the least significant bits is a value of 0. [C10] The apparatus according to C6, wherein the neural network parameters include a neural network weight tensor or a neural network activation tensor. [C11] Means for bit-shifting the binary representation of a neural network parameter, wherein the neural network parameter has fewer bits b than the number of hardware bits B supported by the hardware processing the neural network parameter, wherein the neural network parameter has 2 B-b Means for effectively multiplying and bit-shifting, To obtain the updated quantization scale, set the quantization scale to 2 B-b The means of division by, A means for quantizing the bit-shifted binary representation using the updated quantization scale in order to obtain the values ​​of the neural network parameters, A device equipped with the following features. [C12] The apparatus according to C11, wherein the neural network parameters have 7 bits or less, and the hardware supports 8-bit values. [C13] The apparatus according to C11, wherein the neural network parameters have 4 bits and the hardware supports 8-bit values. [C14] The apparatus according to C11, further comprising means for bit-shifting the binary representation of the neural network parameters such that the difference Bb of the least significant bits has a value of 0. [C15] The apparatus according to C11, wherein the neural network parameters include a neural network weight tensor or a neural network activation tensor. [C16] A non-temporary computer-readable medium on which program code is recorded, wherein the program code is executed by a processor, A bit shift of the binary representation of a neural network parameter, wherein the number of bits b is less than the number of hardware bits B supported by the hardware processing the neural network parameter, and the number of bits b is less than the number of hardware bits B supported by the hardware processing the neural network parameter. B-b Program code for effectively multiplying and bit shifting, To obtain the updated quantization scale, set the quantization scale to 2 B-b Program code for division by, A program code for quantizing the bit-shifted binary representation using the updated quantization scale in order to obtain the values ​​of the neural network parameters, including, Non-temporary computer-readable media. [C17] The non-temporary computer-readable medium according to C16, wherein the neural network parameters have 7 bits or less, and the hardware supports 8-bit values. [C18] The non-temporary computer-readable medium according to C16, wherein the neural network parameters have 4 bits and the hardware supports 8-bit values. [C19] A non-temporary computer-readable medium according to C16, further comprising program code for bit-shifting the binary representation of the neural network parameters such that the difference Bb of the least significant bits has a value of 0. [C20] The non-temporary computer-readable medium according to C16, wherein the neural network parameters include a neural network weight tensor or a neural network activation tensor.

Claims

1. The processor performs a bit shift on the binary representation of a neural network parameter, wherein the binary representation of the neural network parameter has fewer bits b than the number of hardware bits B supported by the hardware processing the neural network parameter, thereby 2 B-b Effectively multiplying, bit shifting, To obtain an updated quantization scale, the processor modifies the quantization scale by 2 B-b Dividing by and In order to obtain the values ​​of the neural network parameters, the processor quantizes the bit-shifted binary representation using the updated quantization scale, A method that includes implementing it on the processor.

2. The method of implementing the process using the processor according to claim 1, wherein the neural network parameters have 7 bits or less, and the hardware supports 8-bit values.

3. The method of implementing the process according to claim 1, wherein the neural network parameters have 4 bits and the hardware supports 8-bit values.

4. The method of implementing the process using the processor according to claim 1, wherein the binary representation of the neural network parameters is bit-shifted such that the difference B-b of the least significant bits is 0.

5. The method of implementation using the processor according to claim 1, wherein the neural network parameters include a neural network weight tensor or a neural network activation tensor.

6. It is a device, Memory and The system comprises at least one processor coupled to the memory, wherein the at least one processor A bit shift of the binary representation of a neural network parameter, wherein the neural network parameter has fewer bits b than the number of hardware bits B supported by the hardware processing the neural network parameter, and the neural network parameter has two bits b. B-b Perform a bit shift to effectively multiply, To obtain the updated quantization scale, set the quantization scale to 2 B-b Divide by, To obtain the values ​​of the neural network parameters, the bit-shifted binary representation is quantized using the updated quantization scale. It is structured in such a way. Device.

7. The apparatus according to claim 6, wherein the neural network parameters have 7 bits or less, and the hardware supports 8-bit values.

8. The apparatus according to claim 6, wherein the neural network parameters have 4 bits and the hardware supports 8-bit values.

9. The apparatus according to claim 6, wherein the at least one processor is further configured to bit-shift the binary representation of the neural network parameters such that the difference B-b of the least significant bits is zero.

10. The apparatus according to claim 6, wherein the neural network parameters include a neural network weight tensor or a neural network activation tensor.

11. A non-temporary computer-readable medium on which program code is recorded, wherein the program code is When executed by the processor, the processor provides a bit shift of the binary representation of a neural network parameter, wherein the neural network parameter has fewer bits b than the number of hardware bits B supported by the hardware processing the neural network parameter, and the neural network parameter has two bits B-b Program code that performs bit shifts to effectively multiply, When executed by the processor, the processor is given two quantization scales in order to obtain the updated quantization scale. B-b Program code that performs division by, When executed by the processor, the program code causes the processor to quantize the bit-shifted binary representation using the updated quantization scale in order to obtain the values ​​of the neural network parameters, including, Non-temporary computer-readable media.

12. The non-temporary computer-readable medium according to claim 11, wherein the neural network parameters have 7 bits or less, and the hardware supports 8-bit values.

13. The non-temporary computer-readable medium according to claim 11, wherein the neural network parameters have 4 bits and the hardware supports 8-bit values.

14. The non-temporary computer-readable medium according to claim 11, further comprising program code that, when executed by a processor, causes the processor to bit-shift the binary representation of the neural network parameters such that the difference B-b of the least significant bits has a value of 0.

15. The non-temporary computer-readable medium according to claim 11, wherein the neural network parameters include a neural network weight tensor or a neural network activation tensor.