Method and apparatus for quantizing a floating point neural network to obtain a fixed point neural network

By quantizing the floating-point neural network and selecting the moment of the input distribution to determine the quantizer parameters, the problem of high computational complexity of fixed-point neural networks is solved, and an effective trade-off between performance and complexity is achieved.

CN112949839BActive Publication Date: 2025-05-09QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110413148.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2015-10-22
Filing Date
2016-04-14
Publication Date
2025-05-09
Estimated Expiration
2036-04-14

AI Technical Summary

Technical Problem

When implementing fixed-point neural networks, it is difficult to effectively reduce the computing complexity and improve the performance-complexity trade-off.

Method used

By quantizing the floating-point neural network, selecting the moment of the input distribution to determine the quantizer parameters, thereby obtaining a fixed-point neural network and reducing the computational complexity.

Benefits of technology

On the basis of maintaining performance, it significantly reduces the computing complexity of fixed-point neural networks and improves the efficiency of hardware and software implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112949839B_ABST
    Figure CN112949839B_ABST
Patent Text Reader

Abstract

Methods and apparatus for quantizing a floating point neural network to obtain a fixed point neural network are disclosed. A method for quantizing a floating point machine learning network to obtain a fixed point machine learning network using a quantizer may include: selecting at least one moment of an input distribution of the floating point machine learning network. The method may also include: determining quantizer parameters for quantizing values ​​of the floating point machine learning network to obtain corresponding values ​​of the fixed point machine learning network based at least in part on the at least one selected moment of the input distribution of the floating point machine learning network.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a Chinese invention patent application with an application date of April 14, 2016, application number 201680026295.7 (international application number PCT / US2016 / 027589), and invention name “Method and device for quantizing floating-point neural networks to obtain fixed-point neural networks”.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 159,079, filed on May 8, 2015, entitled “FIXED POINT NEURAL NETWORK BASED ON FLOATING POINT NEURAL NETWORK QUANTIZATION,” the disclosure of which is expressly incorporated herein by reference in its entirety. background Technical Field

[0004] Certain aspects of the present disclosure relate generally to machine learning and, more particularly, to quantizing a floating point neural network to obtain a fixed point neural network. Background Art

[0006] An artificial neural network, which may include a group of interconnected artificial neurons (e.g., a neuron model), is a computing device or represents a method to be performed by the computing device. Individual nodes in an artificial neural network can mimic biological neurons by obtaining input data and performing simple operations on the data. The results of performing simple operations on the input data are selectively passed to other neurons. The output of each node is called its "activation". Weight values ​​are associated with each vector and node in the network, and these values ​​constrain how the input data is related to the output data. The weight values ​​associated with individual nodes are also called biases. These weight values ​​are determined by the iterative flow of training data through the network (e.g., weight values ​​are established during the training phase in which the network learns how to identify specific categories through typical input data features of these specific categories).

[0007] A convolutional neural network is a feed-forward artificial neural network. A convolutional neural network may include a collection of neurons, each of which has a receptive field and collectively forms an input space. Convolutional neural networks (CNNs) have many applications. In particular, CNNs have been widely used in the fields of pattern recognition and classification.

[0008] Deep learning architectures (such as deep belief networks and deep convolutional networks) are layered neural network architectures where the outputs of the first layer of neurons become the inputs of the second layer of neurons, the outputs of the second layer of neurons become the inputs of the third layer of neurons, and so on. Deep neural networks can be trained to recognize hierarchies of features and therefore they are increasingly used in object recognition applications. Similar to convolutional neural networks, computations in these deep learning architectures can be distributed across a population of processing nodes, which can be configured into one or more computation chains. These multi-layer architectures can be trained one layer at a time and can be fine-tuned using backpropagation.

[0009] Other models can also be used for object recognition. For example, support vector machine (SVM) is a learning tool that can be applied to classification. Support vector machine includes a separating hyperplane (e.g., decision boundary) that classifies data. The hyperplane is defined by supervised learning. The desired hyperplane increases the margin of training data. In other words, the hyperplane should have the maximum minimum distance to the training examples.

[0010] Although these solutions have achieved excellent results on several classification benchmarks, their computational complexity can be extremely high. In addition, training these models can be challenging. Summary of the invention

[0011] A method of quantizing a floating point machine learning network using a quantizer to obtain a fixed point machine learning network may include: selecting at least one moment of an input distribution of the floating point machine learning network. The method may also include: determining a quantizer parameter for quantizing a value of the floating point machine learning network based at least in part on the at least one selected moment of the input distribution of the floating point machine learning network to obtain a corresponding value of the fixed point machine learning network.

[0012] An apparatus for quantizing a floating point machine learning network using a quantizer to obtain a fixed point machine learning network may include: means for selecting at least one moment of an input distribution of the floating point machine learning network. The apparatus may also include: means for determining, based at least in part on the at least one selected moment of the input distribution of the floating point machine learning network, quantizer parameters for quantizing a value of the floating point machine learning network to obtain a corresponding value of the fixed point machine learning network.

[0013] An apparatus for quantizing a floating point machine learning network using a quantizer to obtain a fixed point machine learning network may include: a memory unit and at least one processor coupled to the memory unit. The at least one processor may be configured to: select at least one moment of an input distribution of the floating point machine learning network. The at least one processor may be further configured to: determine a quantizer parameter for quantizing a value of the floating point machine learning network based at least in part on the at least one selected moment of the input distribution of the floating point machine learning network to obtain a corresponding value of the fixed point machine learning network.

[0014] A non-transitory computer-readable medium having recorded thereon program code for, when executed by a processor, using a quantizer to quantize a floating point machine learning network to obtain a fixed point machine learning network, the program code may include: program code for selecting at least one moment of an input distribution of the floating point machine learning network. The non-transitory computer-readable medium may further include: program code for determining quantizer parameters for quantizing values ​​of the floating point machine learning network based at least in part on the at least one selected moment of the input distribution of the floating point machine learning network to obtain corresponding values ​​of the fixed point machine learning network.

[0015] Additional features and advantages of the present disclosure will be described below. Those skilled in the art will appreciate that the present disclosure can be easily used as a basis for modifying or designing other structures for implementing the same purpose as the present disclosure. Those skilled in the art will also recognize that such equivalent constructions do not depart from the teachings of the present disclosure as set forth in the appended claims. The novel features that are considered to be characteristic of the present disclosure, both in terms of its organization and method of operation, together with further objects and advantages, will be better understood when the following description is considered in conjunction with the accompanying drawings. However, it is to be clearly understood that each of the drawings is provided for illustration and description purposes only and is not intended to be a definition of limitations of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The features, nature and advantages of the present disclosure will become more apparent when the detailed description set forth below is read in conjunction with the accompanying drawings, in which like reference numerals designate correspondingly throughout.

[0017] Figure 1 An example implementation of designing a neural network using a system on a chip (SOC) including a general-purpose processor in accordance with certain aspects of the present disclosure is illustrated.

[0018] Figure 2 Example implementations of systems according to aspects of the present disclosure are illustrated.

[0019] Figure 3A is a diagram illustrating a neural network according to aspects of the present disclosure.

[0020] Figure 3B is a block diagram illustrating an exemplary deep convolutional network (DCN) according to aspects of the present disclosure.

[0021] Figure 4 An exemplary probability distribution function showing a distribution for defining a range of a fixed-point representation is illustrated.

[0022] Figure 5A and 5B Illustrate the distribution of activation values ​​and weights in different layers of an exemplary deep convolutional network.

[0023] Fig. 6A Illustrated input distribution for an exemplary deep convolutional network.

[0024] Figure 6B Illustrated are modified input distributions for an exemplary deep convolutional network in accordance with aspects of the present disclosure.

[0025] Fig. 7A and 7B Transforming a first machine learning network into a second machine learning network by incorporating a mean of a distribution of activation values ​​into a network bias of the second machine learning network according to aspects of the present disclosure is illustrated.

[0026] Figure 8 A method of quantizing a floating-point machine learning network using a quantizer to obtain a fixed-point machine learning network according to aspects of the present disclosure is illustrated.

[0027] Fig. 9 A method of determining a step size of a second machine learning network according to a first machine learning network to reduce the computational complexity of the second machine learning network according to aspects of the present disclosure is explained. DETAILED DESCRIPTION

[0028] The detailed description set forth below in conjunction with the accompanying drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein may be practiced. This detailed description includes specific details in order to provide a thorough understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts may be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring such concepts.

[0029] Based on this teaching, it will be appreciated by those skilled in the art that the scope of the present disclosure is intended to cover any aspect of the present disclosure, whether it is implemented independently or in combination with any other aspect of the present disclosure. For example, any number of aspects described can be used to implement a device or practice method. In addition, the scope of the present disclosure is intended to cover such devices or methods that are practiced using other structures, functionality, or structures and functionality that are supplemented or different from the various aspects of the present disclosure described. It should be understood that any aspect of the present disclosure disclosed can be implemented by one or more elements of the claims.

[0030] The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.

[0031] Although specific aspects are described herein, numerous variations and permutations of these aspects fall within the scope of the present disclosure. Although some benefits and advantages of preferred aspects are mentioned, the scope of the present disclosure is not intended to be limited to specific benefits, uses or goals. On the contrary, various aspects of the present disclosure are intended to be broadly applicable to different technologies, system configurations, networks and protocols, some of which are illustrated as examples in the accompanying drawings and the following description of preferred aspects. The detailed description and drawings merely illustrate the present disclosure and do not limit the present disclosure, and the scope of the present disclosure is defined by the attached claims and their equivalent technical solutions.

[0032] Quantization is the process of mapping a set of input values ​​to a smaller set of values. For example, an input value may be rounded to a given unit of precision. Specifically, in one example, the conversion of a floating point number to a fixed point number may be a quantization process.

[0033] In some artificial neural networks (ANNs), such as deep convolutional networks (DCNs), quantization may be applied to the activations of normalization layers; weights, biases, and activations of fully connected layers; and / or weights, biases, and activations of convolutional layers. In addition, for DCNs, quantization may not be applied to pooling layers if max pooling is specified; and / or not be applied to neuron layers if rectified linear units (ReLUs) are specified.

[0034] Various aspects of the present disclosure are directed to improving the quantization of weights, biases, and / or activations in an ANN. That is, various aspects of the present disclosure are directed to quantizing weights, biases, and / or activation values ​​in an ANN to improve the tradeoff between performance and complexity when implementing the ANN using fixed-point numbers.

[0035] Figure 1 An example implementation of the aforementioned reduction in computational complexity by quantizing a floating-point neural network to obtain a fixed-point neural network using a system on a chip (SOC) 100, which may include a general purpose processor (CPU) or a multi-core general purpose processor (CPU) 102, in accordance with certain aspects of the present disclosure, is illustrated. Variables (e.g., neural signals and synaptic weights), system parameters associated with a computing device (e.g., a neural network with weights), delays, frequency bin information, and task information may be stored in a memory block associated with a neural processing unit (NPU) 108, in a memory block associated with the CPU 102, in a memory block associated with a graphics processing unit (GPU) 104, in a memory block associated with a digital signal processor (DSP) 106, in a dedicated memory block 118, or may be distributed across multiple blocks. Instructions executed at the general purpose processor 102 may be loaded from a program memory associated with the CPU 102 or may be loaded from a dedicated memory block 118.

[0036] The SOC 100 may also include additional processing blocks tailored for specific functions, such as a GPU 104, a DSP 106, a connectivity block 110 (which may include fourth generation long term evolution (4G LTE) connectivity, unlicensed Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 112 that may, for example, detect and recognize gestures. In one implementation, the NPU is implemented in the CPU, DSP, and / or GPU. The SOC 100 may also include a sensor 114, an image signal processor (ISP), and / or a navigation 120 (which may include a global positioning system).

[0037] SOC 100 may be based on the ARM instruction set. In one aspect of the present disclosure, the instructions loaded into the general purpose processor 102 may include code for quantizing a floating point neural network to obtain a fixed point neural network. The instructions loaded into the general purpose processor 102 may also include code for providing a fixed point representation when quantizing weights, biases, and activation values ​​in the network.

[0038] Figure 2 An example implementation of a system 200 according to certain aspects of the present disclosure is illustrated. Figure 2 As illustrated in , the system 200 may have a plurality of local processing units 202 that may perform various operations of the methods described herein. Each local processing unit 202 may include a local state memory 204 and a local parameter memory 206 that may store parameters of a neural network. In addition, the local processing unit 202 may have a local (neuron) model program (LMP) memory 208 for storing a local model program, a local learning program (LLP) memory 210 for storing a local learning program, and a local connection memory 212. In addition, as Figure 2 As illustrated in , each local processing unit 202 may interface with a configuration processor unit 214 for providing configuration for each local memory of the local processing unit, and with a routing connection processor unit 216 for providing routing between each local processing unit 202 .

[0039] Deep learning architectures can perform object recognition tasks by learning to represent the input at successively higher levels of abstraction in each layer, thereby building useful feature representations of the input data. In this way, deep learning addresses a major bottleneck of traditional machine learning. Before the advent of deep learning, machine learning approaches to object recognition problems might rely heavily on human-engineered features, perhaps combined with shallow classifiers. A shallow classifier could be a two-class linear classifier, for example, where a weighted sum of the components of a feature vector is compared to a threshold to predict which class the input belongs to. A human-engineered feature could be a template or kernel that is customized for a specific problem domain by an engineer with domain expertise. In contrast, a deep learning architecture can learn to represent features similar to what a human engineer might design, but it does so through training. In addition, deep networks can learn to represent and recognize new types of features that humans might not have considered.

[0040] A deep learning architecture can learn a hierarchy of features. For example, if a first layer is presented with visual data, the first layer can learn to recognize relatively simple features (such as edges) in the input stream. In another example, if a first layer is presented with auditory data, the first layer can learn to recognize spectral power in specific frequencies. A second layer that takes the output of the first layer as input can learn to recognize combinations of features, such as simple shapes for visual data or combinations of sounds for auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to recognize common visual objects or spoken phrases.

[0041] Deep learning architectures can perform particularly well when applied to problems that have a natural hierarchical structure. For example, the classification of motor vehicles can benefit from first learning to recognize wheels, windshields, and other features. These features can be combined in different ways at higher levels to recognize cars, trucks, and airplanes.

[0042] Neural networks can be designed to have various connectivity patterns. In a feedforward network, information is passed from lower layers to higher layers, with each neuron in a given layer communicating to neurons in a higher layer. As described above, hierarchical representations can be constructed in successive layers of a feedforward network. Neural networks can also have reflow or feedback (also known as top-down) connections. In a reflow connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. The reflow architecture can help identify patterns that span more than one block of input data delivered sequentially to the neural network. The connection from a neuron in a given layer to a neuron in a lower layer is called a feedback (or top-down) connection. A network with many feedback connections may be helpful when the recognition of high-level concepts can assist in discerning specific low-level features of the input.

[0043] Reference Figure 3A , the connections between the layers of the neural network can be fully connected (302) or locally connected (304). In the fully connected network 302, a neuron in the first layer can communicate its output to each neuron in the second layer, so that each neuron in the second layer will receive input from each neuron in the first layer. Alternatively, in the locally connected network 304, the neurons in the first layer can be connected to a limited number of neurons in the second layer. The convolutional network 306 can be locally connected and further configured so that the connection strengths associated with the input to each neuron in the second layer are shared (e.g., 308). More generally, the locally connected layers of the network can be configured so that each neuron in a layer will have the same or similar connectivity pattern, but its connection strengths may have different values ​​(e.g., 310, 312, 314, and 316). The connectivity patterns of local connections may produce spatially distinct receptive fields in higher layers because higher layer neurons in a given area may receive inputs that are tuned through training to be a restricted portion of the total input to the network.

[0044] Locally connected neural networks may be well suited for problems where the spatial location of the input is meaningful. For example, a network 300 designed to recognize visual features from an onboard camera may develop high-level neurons with different properties depending on whether they are associated with the lower portion of the image or the upper portion of the image. For example, neurons associated with the lower portion of the image may learn to recognize lane markings, while neurons associated with the upper portion of the image may learn to recognize traffic lights, traffic signs, etc.

[0045] The DCN can be trained using supervised learning. During training, the DCN can be presented with an image (such as a cropped image of a speed limit sign), and then a "forward pass" can be calculated to produce output 322. Output 322 can be a vector of values ​​corresponding to features (such as "sign", "60", and "100"). The network designer may want the DCN to output high scores for some of the neurons in the output feature vector, such as those neurons corresponding to "sign" and "60" shown in output 322 of the trained network 300. Prior to training, the output produced by the DCN is likely to be incorrect, and the error between the actual output and the target output can be calculated. The weights of the DCN can then be adjusted to align the DCN's output scores more closely with the target.

[0046] To adjust the weights, the learning algorithm may calculate a gradient vector for the weights. The gradient may indicate the amount by which the error would increase or decrease if the weights were adjusted slightly. At the top layer, the gradient may correspond directly to the value of the weights connecting the activated neurons in the penultimate layer to the neurons in the output layer. In lower layers, the gradient may depend on the value of the weights and the calculated error gradient for the higher layers. The weights may then be adjusted to reduce the error. This way of adjusting weights may be referred to as "backward propagation" because it involves a "backward pass" through a neural network.

[0047] In practice, the error gradient of the weights may be computed on a small number of examples so that the computed gradient approximates the true error gradient. This approximation may be called stochastic gradient descent. Stochastic gradient descent may be repeated until the achievable error rate of the entire system has stopped decreasing or until the error rate has reached a target level.

[0048] After learning, the DCN may be presented with new images and a forward pass through the network may produce output 322, which may be considered an inference or prediction of the DCN.

[0049] Deep belief network (DBN) is a probabilistic model including multiple layers of hidden nodes. DBN can be used to extract the hierarchical representation of training data set. DBN can be obtained by stacking multiple layers of restricted Boltzmann machine (RBM). RBM is a kind of artificial neural network that can learn probability distribution on input set. Since RBM can learn probability distribution when there is no information about which class each input should be classified into, RBM is often used for unsupervised learning. Using mixed unsupervised and supervised paradigm, the bottom RBM of DBN can be trained in an unsupervised manner and can be used as a feature extractor, while the top RBM can be trained in a supervised manner (on the joint distribution of input and target class from previous layer) and can be used as a classifier.

[0050] A deep convolutional network (DCN) is a network of convolutional networks configured with additional pooling and normalization layers. DCN has achieved state-of-the-art performance on many tasks. DCN can be trained using supervised learning, where both the input and output targets are known for many examples and are used to modify the weights of the network by using gradient descent.

[0051] The DCN can be a feed-forward network. In addition, as described above, the connections from the neurons in the first layer of the DCN to the neuron groups in the next higher layer are shared across the neurons in the first layer. The feed-forward and shared connections of the DCN can be utilized for fast processing. The computational burden of the DCN can be much smaller than, for example, a neural network of similar size that includes recurrent or feedback connections.

[0052] The processing of each layer of the convolutional network can be considered as a spatially invariant template or basis projection. If the input is first decomposed into multiple channels, such as the red, green and blue channels of a color image, the convolutional network trained on the input can be considered to be three-dimensional, with two spatial dimensions along the axis of the image and a third dimension that captures color information. The output of the convolutional connection can be considered to form a feature map in subsequent layers 318 and 320, each element in the feature map (e.g., 320) receives input from a certain range of neurons in the previous layer (e.g., 318) and from each of the multiple channels. The values ​​in the feature map can be further processed with nonlinearity (such as correction) max (0, x). The values ​​from adjacent neurons can be further pooled (which corresponds to downsampling) and can provide additional local invariance and dimensionality reduction. Normalization can also be applied by lateral inhibition between neurons in the feature map, which corresponds to whitening.

[0053] The performance of deep learning architectures can improve as more labeled data points become available or as computing power increases. Modern deep neural networks are routinely trained with thousands of times more computing resources than were available to a typical researcher just fifteen years ago. New architectures and training paradigms can further boost the performance of deep learning. Rectified linear units can reduce the training problem known as vanishing gradients. New training techniques can reduce over-fitting and thus enable larger models to achieve better generalization. Encapsulation techniques can abstract the data within a given receptive field and further improve overall performance.

[0054] Figure 3B is a block diagram illustrating a deep convolutional network 350. The deep convolutional network 350 may include multiple layers of different types based on connectivity and weight sharing. Figure 3B As shown, the deep convolutional network 350 includes a plurality of convolutional blocks (e.g., C1 and C2). Each convolutional block may be configured with a convolutional layer (CONV), a normalization layer (LNorm), and a pooling layer. The convolutional layer may include one or more convolutional filters, which may be applied to input data to generate a feature map. Although only two convolutional blocks are shown, the present disclosure is not limited thereto, but, depending on design preferences, any number of convolutional blocks may be included in the deep convolutional network 350. The normalization layer may be used to normalize the output of the convolutional filter. For example, the normalization layer may provide whitening or lateral suppression. The pooling layer may provide downsampling aggregation in space to achieve local invariance and dimensionality reduction.

[0055] For example, the parallel filter banks of the deep convolutional network may be optionally loaded onto the CPU 102 or GPU 104 of the SOC 100 based on the ARM instruction set to achieve high performance and low power consumption. In an alternative embodiment, the parallel filter banks may be loaded onto the DSP 106 or ISP 116 of the SOC 100. In addition, the DCN may access other processing blocks that may exist on the SOC, such as processing blocks dedicated to sensors 114 and navigation 120.

[0056] The deep convolutional network 350 may also include one or more fully connected layers (e.g., FC1 and FC2). The deep convolutional network 350 may further include a logistic regression (LR) layer. Between each layer of the deep convolutional network 350 are weights (not shown) to be updated. The output of each layer can be used as an input to subsequent layers in the deep convolutional network 350 to learn a hierarchical feature representation from the input data (e.g., image, audio, video, sensor data, and / or other input data) provided at the first convolutional block C1.

[0057] In one configuration, a machine learning model (such as a neural model) is configured to quantize a floating point neural network to obtain a fixed point neural network. The model includes a reducing device and / or a balancing device. In one aspect, the reducing device and / or the balancing device can be a general purpose processor 102, a program memory associated with the general purpose processor 102, a memory block 118, a local processing unit 202, and / or a routing connection processing unit 216 configured to perform the functions recited. In another configuration, the aforementioned means can be any module or any device configured to perform the functions recited by the aforementioned means.

[0058] According to certain aspects of the present disclosure, each local processing unit 202 may be configured to determine parameters of the model based on one or more desired functional characteristics of the model, and to develop the one or more functional characteristics toward the desired functional characteristics as the determined parameters are further adapted, tuned, and updated.

[0059] Fixed-point neural network based on floating-point neural network quantization

[0060] Floating point representations of weights, biases, and / or activation values ​​in an artificial neural network (ANN) may increase the complexity of the hardware and / or software implementation of the network. In some cases, a fixed point representation of a network (such as a deep convolutional network (DCN) or an artificial neural network) may provide an improved performance-complexity tradeoff. Various aspects of the present disclosure relate to quantizing weights, biases, and activation values ​​in an artificial neural network to improve the tradeoff between performance and complexity when the network is implemented using fixed point numbers.

[0061] Fixed-point numbers can be specified to use less complex software and / or hardware design at the expense of reduced accuracy, because floating-point numbers have a larger dynamic range than fixed-point numbers. Converting floating-point numbers to fixed-point numbers by a quantization process can reduce the complexity of hardware and / or software implementations. Floating-point numbers can take a single-precision binary format that includes a sign bit, an 8-bit exponent, and 23 fractional components.

[0062] Various aspects of the present disclosure relate to using a Q number format to represent fixed-point numbers. However, other formats are contemplated. The Q number format is represented as Qm.n, where m is the number of bits in the integer portion and n is the number of bits in the fractional portion. In one configuration, m does not include a sign bit. Each Qm.n format may use an m+n+1 bit signed integer container with n fractional bits. In one configuration, the range is [-(2 m ),2 m -2 -n )] and the resolution is 2 -n For example, Q14.1 format numbers can use sixteen bits. In this example, the range is [-2 14 ,2 14 -2 -1 ] (for example, [-16384.0, +16383.5]) and the resolution is 2 -1 (For example, 0.5).

[0063] In one configuration, an extension of the Q number format is specified to support instances where the resolution is greater than one or the maximum range is less than one. In some cases, the fractional bits of a negative number may be specified for a resolution greater than one. Additionally, the integer bits of a negative number may be specified for a maximum range less than one.

[0064] A deep convolutional network (DCN) is an example of an artificial neural network to which quantization may be applied according to aspects of the present disclosure. Quantization may be applied to the activations of normalization layers; weights, biases, and activations of fully connected layers; and / or weights, biases, and activations of convolutional layers. However, quantization need not be applied to pooling layers when max pooling is specified, and / or need not be applied to neuron layers when rectified linear units (ReLUs) are specified. Aspects of the present disclosure relate to improving quantization of weights, biases, and / or activations in artificial neural networks by applying various optimizations.

[0065] According to aspects of the present disclosure, the quantization efficiency in an artificial neural network can be obtained by Figure 4 The quantization of the probability distribution function 400 shown in FIG. 4 can be better understood by viewing the quantization. For example, the input to the quantizer can be uniformly distributed in [X min ,X max ], where X min (i.e. X 最小 , Figure 4Indicated as -X max ) and X max (i.e. X 最大 ) defines the range of fixed-point representation. When the input of the quantizer is uniformly distributed in [X min ,X max ], the quantization noise is:

[0066]

[0067] The signal power is:

[0068]

[0069] And the signal to quantization noise ratio (SQNR) assuming that M is an integer number of bits is:

[0070]

[0071] Figure 5A and 5B Illustrate the distribution of activation values ​​and weights in different layers of an exemplary deep convolutional network. Figure 5A The activation values ​​of convolutional layers one to five (conv1, ..., conv5) and fully connected layers one and two (fc1, fc2) are shown. Figure 5B Weights 550 for convolutional layers one to five (conv1, ..., conv5) and fully connected layers one and two (fc1, fc2) are shown. Applying quantization to weights, biases, and activation values ​​in an artificial neural network includes determining the step size. For example, the step size of a symmetric uniform quantizer for Gaussian, Laplace, and gamma (γ) distributions can be calculated using a deterministic function of the standard deviation of the input distribution assuming that these distributions have zero mean and unit variance. Accordingly, aspects of the present disclosure relate to modifications to the calculation of weights and / or activation values ​​so that each distribution has a zero mean (e.g., a mean of approximately zero). In one configuration, both weights and activation values ​​are assumed to have a Gaussian distribution, however, other distributions are also contemplated.

[0072] In various aspects of the present disclosure, the modification of the weight and / or activation value calculations so that each distribution has a zero mean is performed by removing the mean. In one configuration, quantization can be performed after removing the mean value (μ) of the input distribution. If the distribution has a large mean value, the removal of the mean value may have a large impact.

[0073] Fig. 6A An input distribution 600 for an exemplary deep convolutional network is illustrated. In this example, the input distribution 600 includes a variance (σ) and a mean value (μ). Aspects of the present disclosure relate to specifying a zero mean (μ=0) for the distribution of weights, biases, and activation values, e.g., Figure 6B as shown in .

[0074] Figure 6B A modified input distribution 650 for an exemplary deep convolutional network is illustrated. In this configuration, the mean value (μ) is added to the standard (std) deviation (variance (σ)) to reduce the computational overhead in determining the range of the encoding, such that:

[0075] Input stdσ′=|μ|+σ (4)

[0076] However, Figure 6B The modified input distribution 650 of quantizes the input data quantized by ...

[0077] Fig. 7A and 7B A method of converting a first machine learning network into a second machine learning network by incorporating the mean value of the activation value distribution of the first machine learning network into the network bias of the second machine learning network is explained. Fig. 7A An input distribution 700 of activation values ​​for an exemplary deep convolutional network is illustrated that also has a mean (μ) and a variance (σ). In various aspects of the present disclosure, a modification of the input distribution 700 for activation value calculations is performed so that the activation values ​​have a zero mean. In this aspect of the present disclosure, the modification of the weight and / or activation value calculations is performed by incorporating the mean value (μ) into the bias of the artificial neural network, for example, as Figure 7B as shown in .

[0078] Figure 7B A modified input distribution 750 of activation values ​​of an exemplary deep convolutional network with variance (σ) and mean (μ) is illustrated. In this configuration, the mean (μ) is absorbed into the bias value of the deep convolutional network model. Absorbing the mean (μ) into the bias value can reduce the computational burden of mean removal. Absorbing the mean (μ) into the bias of the deep convolutional network model substantially removes the mean activation without additional computational burden. The calculation of the modified bias value can be performed as follows.

[0079] In some cases, when the ANN has multiple layers, the activation value of the i-th neuron in layer l+1 can be calculated as follows:

[0080]

[0081] Where (l) represents the lth layer, N represents the number of additions, and w i,jrepresents the weight from neuron j to neuron i in layer l, and b i Indicates bias.

[0082] activation can be expressed as the mean component μ (l) and the zero mean part The sum of , then:

[0083]

[0084]

[0085] In one configuration, the bias values ​​are modified to specify a zero mean throughout the network for the activation distribution. Additionally, the nonlinear activation function is modified when the mean value is incorporated into the network bias. In this configuration, the original bias values Replaced with the modified value

[0086]

[0087] The resulting network has zero-mean activations:

[0088]

[0089] For some layers, such as at the output of an ANN, one can specify non-zero mean output activations such that:

[0090]

[0091] in,

[0092] Applying quantization to weights, biases, and activation values ​​in an artificial neural network can include a fixed-point converter. In some cases, for an ANN model, and is known. In addition, can be measured from floating-point simulations. Thus, in one configuration, the floating-point to fixed-point model converter can be implemented by measuring the average activation μ at each layer (l) and calculate the new value as follows to calculate the new value and assign it to

[0093]

[0094] In some cases, it is assumed that the weights in the ANN have a substantially zero mean for each observation. Thus, it is assumed that the foregoing examples involve activation values. Still, the foregoing examples may also be applied to weight values. Application to weight values ​​may include shifting activation values ​​to create a zero mean distribution for each layer of a fixed-point machine learning network. When activations throughout the network are shifted to create a zero mean distribution for each layer, if a subsequent nonlinear function is applied to the activation, the coordinates of the function are shifted by the same amount so that the output is a shift of the original output without bias modification. That is, the quantization process includes shifting the coordinates of any nonlinear function applied to the shifted activation value by an amount corresponding to the shifted activation value.

[0095] Applying quantization to weights, biases, and activation values ​​in an artificial neural network includes determining the step size. For fixed-point representation, the step size can be limited to a power of 2. Determining the step size as a power of 2 can correspond to determining the number of fractional digits in a fixed-point number representation. The formula for determining the step size can be specified as follows:

[0096]

[0097] Where μ and σ are the mean and standard deviation of the inputs used to calculate the effective ∑ value σ'. Next, the effective step size is calculated based on the effective ∑ value σ' as follows:

[0098] S 浮点 =σ′×C 缩放 (w)×α (15)

[0099] Where S 浮点 is the calculated step size in floating point format, C 缩放(w) is the scaling constant for the bit width w, and α is the adjustment factor for the step size. Finally, the nearest power of 2 is determined for the step size as follows:

[0100] n=-S 浮点 [log2S 浮点 ] (16) where n is the number of fractional bits that can be specified to represent the quantizer input, and 2 -n Can be specified as a step size. In addition to (ceiling operation), other rounding functions can be used to obtain integer n, including round(·) and (Floor operation).

[0101] The scaling function and adjustment factor function C from equation (15) 缩放(w) It can be specified as a function of the bit width (W) according to Table I.

[0102] Table I: Uniform quantizer for Gaussian distribution

[0103]

[0104] The additional adjustment factor α is a value that can be adjusted to improve classification performance. For example, in certain scenarios, α can be specified as a value different from 1, such as: (1) the input distribution is not Gaussian (e.g., potentially long tail); or (2) the fixed-point representation calculated for DCN is inconsistent with the representation calculated based on signal-to-quantization noise ratio (SQNR) considerations. In an exemplary DCN model involving scene detection, α different from 1 (such as α=1.5) improves performance.

[0105] In addition, the step size adjustment factor α can be specified differently throughout the model. For example, α can be specified individually for the weights and activations of each layer. In addition, the weights and biases can have very different dynamic ranges. For example, the weights and biases can be specified to have different Q number representations and different bit widths. In addition, the bit widths of weights and biases in the same layer can be the same. In one configuration, for a given layer, the weights have a format of Q 3.18 and the biases have a format of Q 6.9.

[0106] In one configuration, after the floating point model is quantized to a fixed point model, the fixed point network is fine-tuned via additional training to further improve network performance. Fine-tuning may include training via back propagation. In addition, the step size and Q number representation identified in the floating point to fixed point conversion can be continued to the fine-tuned network. In this example, no additional step size optimization is specified. The quantizer for the ANN can maintain the fidelity of the network while reducing resource usage. For the exemplary network, there may be almost no difference in accuracy between a 32-bit floating point network implementation and a 16-bit fixed point network implementation. The fixed-point implementation of the ANN based on the disclosed quantizer design can reduce model size, processing time, memory bandwidth, and power consumption.

[0107] Figure 8 A method 800 of quantizing a floating point machine learning network using a quantizer to obtain a fixed point machine learning network in accordance with aspects of the present disclosure is illustrated. At block 802, at least one moment of an input distribution of a floating point machine learning network is selected. The at least one moment of the input distribution of the floating point machine learning network may include a mean, variance, or other similar moment of the input distribution. At block 804, quantizer parameters for quantizing values ​​of the floating point machine learning network are determined based on the selected moment of the input distribution of the floating point machine learning network. At block 806, it is determined whether additional moments are available for the input distribution. If so, blocks 802 and 804 may be repeated for each moment of the input distribution, including, for example, a mean, variance, or other similar moment of the input distribution.

[0108] In various aspects of the present disclosure, determination of quantizer parameters for quantizing values ​​of a floating point machine learning network is performed to obtain corresponding values ​​of a fixed point machine learning network. The quantizer parameters for quantizing values ​​of a floating point machine learning network include a maximum dynamic encoding range, a quantizer step size, a bit width, a signed / unsigned indicator value, or other similar quantization parameters of a fixed point machine learning network. Additionally, the corresponding values ​​of a fixed point machine learning network may include, but are not limited to, biases, weights, and / or activation values.

[0109] Fig. 9 A method 900 for determining a step size of a second machine learning network (e.g., a fixed-point neural network) based on a first machine learning network (e.g., a floating-point neural network) to reduce the computational complexity of the second machine learning network according to various aspects of the present disclosure is illustrated. Applying quantization to weights, biases, and activation values ​​in an artificial neural network includes determining the step size. At box 902, an effective ∑ is calculated. For example, as shown in equation (14), the absolute value of the mean μ of the input distribution is added to the standard deviation (σ) of the input distribution to calculate an effective ∑ value (σ'). At box 904, for example, according to equation (15), an effective step size (S ) is calculated based on the effective ∑ value (σ'). 浮点 ). At block 906, the nearest power of 2 is determined for the step size, for example according to equation (16).

[0110] The various operations of the methods described above may be performed by any suitable device capable of performing the corresponding functions. These devices may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. In general, where there are operations illustrated in the accompanying drawings, those operations may have corresponding paired device-plus-function components with similar numbers.

[0111] As used herein, the term "determining" encompasses a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, searching (e.g., searching in a table, database, or other data structure), ascertaining, and the like. Additionally, "determining" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and the like. Furthermore, "determining" may include resolving, selecting, choosing, establishing, and the like.

[0112] As used herein, a phrase referring to "at least one" of a list of items refers to any combination of those items, including single members. As an example, "at least one of a, b, or c" is intended to cover: a, b, c, ab, ac, bc, and abc.

[0113] The various illustrative logical blocks, modules, and circuits described in conjunction with the present disclosure may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array signal (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0114] The steps of the method or algorithm described in conjunction with the present disclosure may be implemented directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in any form of storage medium known in the art. Some examples of usable storage media include random access memory (RAM), read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs, etc. The software module may include a single instruction, or many instructions, and may be distributed over several different code segments, distributed between different programs, and distributed across multiple storage media. A storage medium may be coupled to a processor so that the processor can read and write information from / to the storage medium. In an alternative, a storage medium may be integrated into a processor.

[0115] The methods disclosed herein include one or more steps or actions for achieving the described methods. These method steps and / or actions may be interchangeable with each other without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be changed without departing from the scope of the claims.

[0116] The functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system in a device. The processing system may be implemented with a bus architecture. Depending on the specific application and overall design constraints of the processing system, the bus may include any number of interconnected buses and bridges. The bus may link together various circuits including a processor, a machine-readable medium, and a bus interface. The bus interface may be used to connect a network adapter, etc., to the processing system via the bus, in particular. The network adapter may be used to implement signal processing functions. For some aspects, a user interface (e.g., a keypad, a display, a mouse, a joystick, etc.) may also be connected to the bus. The bus may also link various other circuits (such as timing sources, peripherals, voltage regulators, power management circuits, etc.), which are well known in the art and will not be described in detail.

[0117] The processor may be responsible for managing the bus and general processing, including executing software stored on a machine-readable medium. The processor may be implemented with one or more general and / or special processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuit systems that can execute software. Software should be broadly interpreted as meaning instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or other. As an example, a machine-readable medium may include a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a disk, an optical disk, a hard drive, or any other suitable storage medium, or any combination thereof. The machine-readable medium may be implemented in a computer program product. The computer program product may include packaging material.

[0118] In hardware implementation, the machine-readable medium can be a part separated from the processor in the processing system. However, as those skilled in the art will readily appreciate, the machine-readable medium or any part thereof can be outside the processing system. As an example, the machine-readable medium may include a transmission line, a carrier wave modulated by data, and / or a computer product separated from the device, all of which can be accessed by the processor through a bus interface. Alternatively or in addition, the machine-readable medium or any part thereof can be integrated into the processor, such as a cache and / or a general register file may be this case. Although the various components discussed can be described as having specific locations, such as local components, they can also be configured in various ways, such as some components are configured as a part of a distributed computing system.

[0119] The processing system can be configured as a general processing system having one or more microprocessors providing processor functionality and external memory providing at least a portion of the machine-readable medium, all linked together with other supporting circuitry via an external bus architecture. Alternatively, the processing system can include one or more neuromorphic processors for implementing the neuron model and neural system model described herein. As another alternative, the processing system can be implemented with an application specific integrated circuit (ASIC) with a processor, bus interface, user interface, supporting circuitry, and at least a portion of the machine-readable medium integrated in a single chip, or with one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic, discrete hardware components, or any other suitable circuitry, or any combination of circuits capable of performing the various functionalities described throughout this disclosure. Depending on the specific application and the overall design constraints imposed on the overall system, those skilled in the art will recognize how to best implement the functionality described with respect to the processing system.

[0120] The machine-readable medium may include several software modules. These software modules include instructions that cause the processing system to perform various functions when executed by the processor. These software modules may include a transmission module and a receiving module. Each software module may reside in a single storage device or be distributed across multiple storage devices. As an example, when a triggering event occurs, the software module can be loaded into the RAM from the hard drive. During the execution of the software module, the processor can load some instructions into the cache to increase the access speed. Subsequently, one or more cache lines can be loaded into the general register file for execution by the processor. When the functionality of the software module is described below, it will be understood that such functionality is implemented by the processor when the processor executes the instructions from the software module. In addition, it should be appreciated that various aspects of the present disclosure produce improvements in the functionality of processors, computers, machines, or other systems that implement such aspects.

[0121] If implemented in software, each function can be stored on or transmitted by one or more instructions or codes on a computer-readable medium. Computer-readable media include both computer storage media and communication media, and these media include any media that facilitate the transfer of a computer program from one place to another. Storage media can be any available media that can be accessed by a computer. As an example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage devices, or any other media that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer. In addition, any connection is also properly referred to as a computer-readable medium. For example, if the software is transmitted from a website, a server, or other remote sources using a coaxial cable, a fiber optic cable, a twisted pair, a digital subscriber line (DSL), or a wireless technology (such as infrared (IR), radio, and microwaves), the coaxial cable, the fiber optic cable, the twisted pair, the DSL, or the wireless technology (such as infrared, radio, and microwaves) are included in the definition of the medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Disks, where disks often reproduce data magnetically, and discs reproduce data optically with lasers. Thus, in some aspects, computer-readable media may include non-transitory computer-readable media (e.g., tangible media). Additionally, for other aspects, computer-readable media may include transient computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.

[0122] Thus, some aspects may include a computer program product for performing the operations presented herein. For example, such a computer program product may include a computer-readable medium having stored (and / or encoded) thereon instructions that can be executed by one or more processors to perform the operations described herein. For some aspects, the computer program product may include packaging materials.

[0123] In addition, it should be appreciated that the modules and / or other appropriate means for performing the methods and techniques described herein can be downloaded and / or otherwise obtained by the user terminal and / or base station where applicable. For example, such a device can be coupled to a server to facilitate the transfer of the means for performing the methods described herein. Alternatively, the various methods described herein can be provided via a storage device (e.g., RAM, ROM, a physical storage medium such as a compact disc (CD) or a floppy disk, etc.) so that once the storage device is coupled to or provided to the user terminal and / or base station, the device can obtain the various methods. In addition, any other suitable technology suitable for providing the methods and techniques described herein to the device can be utilized.

[0124] It will be understood that the claims are not limited to the precise configuration and components illustrated above. Various changes, substitutions and variations may be made in the arrangement, operation and details of the methods and apparatus described above without departing from the scope of the claims.

Claims

1. A method for machine learning a neural network, the method comprising: determining an average of input activation values ​​for a layer of the neural network; shifting the input activation values ​​of the layer by a shift value based on the mean to achieve a zero mean distribution for the layer of the neural network; shifting the coordinates of a nonlinear function applied to the shifted input activation values ​​by an amount corresponding to the shift value; as well as A shifted nonlinear function is applied to the shifted input activation values ​​to generate an output of the layer of the neural network.

2. The method of claim 1, further comprising: determining a variance of the input activation values ​​for the layer of the neural network; as well as A quantizer parameter for quantizing at least one of the shifted activation values, weight values ​​of the layer, and bias values ​​of the layer of the neural network is determined based on at least one of the variance and the mean.

3. The method of claim 2, wherein the quantizer parameters include one or more of a maximum dynamic coding range, a quantizer step size, a bit width, and a signed / unsigned indicator value.

4. The method of claim 1, further comprising: determining a second average of input activation values ​​for different layers of the neural network; shifting the input activation values ​​of the different layers by a second shift value based on the second mean to achieve a zero mean distribution for the different layers of the neural network; shifting the coordinates of a second nonlinear function applied to the shifted input activation values ​​of the different layers by an amount corresponding to the second shift value; as well as A shifted second non-linear function is applied to the shifted input activation values ​​to generate outputs of the different layers of the neural network.

5. A method for machine learning a neural network, the method comprising: determining an average of input activation values ​​for a layer of the neural network; shifting the input activation values ​​of the layer based on the mean to achieve a zero-mean distribution for the layer of the neural network; shifting corresponding bias values ​​of the layers based on the determined average; as well as The shifted input activation values ​​are processed using the shifted corresponding bias values ​​to generate an output of the layer of the neural network.

6. The method of claim 5, further comprising: determining a variance of the input activation values ​​for the layer of the neural network; as well as A quantizer parameter for quantizing at least one of the shifted activation values, the weight values ​​of the layer, and the bias value of the layer of the neural network is determined based on at least one of the variance and the mean.

7. The method of claim 6, wherein the quantizer parameters include one or more of a maximum dynamic coding range, a quantizer step size, a bit width, and a signed / unsigned indicator value.

8. The method of claim 6, further comprising quantizing the weight values ​​and bias values ​​of at least the layer using the determined quantizer parameters, wherein the quantized weight values ​​and bias values ​​of the layer have different bit widths.

9. The method of claim 6, further comprising quantizing the weight values ​​and bias values ​​of at least the layer using the determined quantizer parameters, wherein the quantized weight values ​​and bias values ​​of the layer have different Q-number representations.

10. The method of claim 5, further comprising: determining a second average of input activation values ​​for different layers of the neural network; shifting the input activation values ​​of the different layers by a second shift value based on the second mean to achieve a zero mean distribution for the different layers of the neural network; shifting corresponding bias values ​​of the different layers based on the determined second average; as well as The shifted input activation values ​​of the different layers are processed to generate outputs of the different layers of the neural network.

11. An apparatus comprising a processor and a memory, the processor being adapted to: Determine the average of input activation values ​​for a layer of a machine learning neural network; shifting the input activation values ​​of the layer by a shift value based on the mean to achieve a zero mean distribution for the layer of the neural network; shifting the coordinates of a nonlinear function applied to the shifted input activation values ​​by an amount corresponding to the shift value; as well as A shifted nonlinear function is applied to the shifted input activation values ​​to generate an output of the layer of the neural network.

12. The apparatus of claim 11, wherein the processor is further adapted to: determining a variance of the input activation values ​​for the layer of the neural network; and A quantizer parameter for quantizing at least one of the shifted activation values, weight values ​​of the layer, and bias values ​​of the layer of the neural network is determined based on at least one of the variance and the mean.

13. The apparatus of claim 12, wherein the quantizer parameters include one or more of a maximum dynamic coding range, a quantizer step size, a bit width, and a signed / unsigned indicator value.

14. The apparatus of claim 11, wherein the processor is further adapted to: determining a second average of input activation values ​​for different layers of the neural network; shifting the input activation values ​​of the different layers by a second shift value based on the second mean to achieve a zero mean distribution for the different layers of the neural network; shifting the coordinates of a second nonlinear function applied to the shifted input activation values ​​of the different layers by an amount corresponding to the second shift value; as well as A shifted second non-linear function is applied to the shifted input activation values ​​to generate outputs of the different layers of the neural network.

15. An apparatus comprising a processor and a memory, the processor being adapted to: Determine the average of input activation values ​​for a layer of a machine learning neural network; shifting the input activation values ​​of the layer based on the mean to achieve a zero-mean distribution for the layer of the neural network; shifting corresponding bias values ​​of the layers based on the determined average; as well as The shifted input activation values ​​are processed using the shifted corresponding bias values ​​to generate an output of the layer of the neural network.

16. The apparatus of claim 15, wherein the processor is further adapted to: determining a variance of the input activation values ​​for the layer of the neural network; and A quantizer parameter for quantizing at least one of the shifted activation values, the weight values ​​of the layer, and the bias value of the layer of the neural network is determined based on at least one of the variance and the mean.

17. The apparatus of claim 16, wherein the quantizer parameters include one or more of a maximum dynamic coding range, a quantizer step size, a bit width, and a signed / unsigned indicator value.

18. The apparatus of claim 16, the processor being further adapted to quantize the weight values ​​and bias values ​​of at least the layer using the determined quantizer parameters, wherein the quantized weight values ​​and bias values ​​of the layer have different bit widths.

19. The apparatus of claim 16, the processor being further adapted to quantize the weight values ​​and bias values ​​of at least the layer using the determined quantizer parameters, wherein the quantized weight values ​​and bias values ​​of the layer have different Q-number representations.

20. The apparatus of claim 15, wherein the processor is further adapted to: determining a second average of input activation values ​​for different layers of the neural network; shifting the input activation values ​​of the different layers by a second shift value based on the second mean to achieve a zero mean distribution for the different layers of the neural network; shifting corresponding bias values ​​of the different layers based on the determined second average; as well as The shifted input activation values ​​of the different layers are processed to generate outputs of the different layers of the neural network.

21. A non-transitory computer readable medium having program code recorded thereon, the program code being adapted to cause a computer to execute the method of any one of claims 1-10.