Optimized quantization for resolution-reduced neural networks
Patent Information
- Application Number
- CN202110022078.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-10
- Filing Date
- 2021-01-08
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2041-01-08
Smart Images

Figure CN113112013B_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to neural networks that use fixed-point value computation. Background Technology
[0002] In recent years, deep learning methods have enabled most of the breakthroughs in computer vision and speech processing / recognition based on machine learning. The task of classifying input data using these deep learning-based classifiers has been extensively studied and applied to many different applications. Depending on the application, the neural networks required for classification can be enormous, containing tens of millions of variables. Such large networks require significant computational and data storage resources, resulting in a high energy / power footprint. Due to these high resource requirements, many deep learning tasks are primarily performed in the cloud (most computation is carried out on GPUs or specialized hardware such as neural network accelerators). Due to computational and power constraints, deep learning networks cannot be deployed in resource-constrained environments in many cases. A recent trend is expanding applications from imagers and phones to other types of sensors (e.g., inertial sensors). Due to battery life limitations, these sensors can become part of wearable devices without permanent cloud connectivity—the so-called edge computing. Therefore, novel concepts for local classification on edge devices are needed. Summary of the Invention
[0003] A method for converting floating-point weighting factors of a neural network to fixed-point weighting factors includes: selecting a predetermined number of candidate scaling factors, which are multiples of a predetermined base. The method includes: evaluating each candidate scaling factor in a cost function. The method includes: selecting a scaling factor as the one among the candidate scaling factors that results in the minimum value of the cost function. The method includes: generating fixed-point weighting factors by scaling the floating-point weighting factors using the scaling factors. The method includes: operating the neural network using the fixed-point weighting factors.
[0004] The predetermined cardinality can be two. The method may further include: providing fixed-point weighting factors to the inference phase in response to the completion of the training phase of the neural network. The predetermined number of candidate scaling factors may include a larger number of candidates having values exceeding the average of the absolute values of the floating-point weighting factors. The predetermined number of candidate scaling factors may include only one candidate whose value is less than the average of the absolute values of the floating-point weighting factors. The cost function may be the mean squared error between the floating-point weighting factor and the product of the candidate scaling factor and the corresponding fixed-point weighting factor. The method may further include: updating the scaling factors during the training phase of the neural network after a predetermined number of training intervals.
[0005] The machine learning system includes a controller programmed to convert the floating-point weighting factors of a neural network to fixed-point weighting factors using a scaling factor, which is a predetermined base. b The scaling factor is a multiple of the scaling factor and minimizes the cost function, which is the mean square error between the floating-point weighting factor and the product of the candidate scaling factor and the corresponding fixed-point weighting factor, and the scaling factor is changed after a predetermined number of iterations during the training phase.
[0006] The controller can be further programmed to implement the neural network using fixed-point operations. Candidate scaling factors can include those with exponential values. L and L -1 The first and second candidate values make the average of the absolute values of the floating-point weighting factors in b L and b L-1 Between. The controller can be further programmed to evaluate the cost function using candidate scaling factors, which are derived from a predetermined base. b L-1 arrive b L+4 The controller can be further programmed to evaluate the cost function for a first number of candidate scaling factors and a second number of candidate scaling factors, wherein the first number of candidate scaling factors is greater than the average of the absolute values of the floating-point weighted factors, and the second number of candidate scaling factors is less than the average, with the first number being greater than the second number. The controller can be further programmed to provide fixed-point weighted factors to the inference phase configured to implement the neural network after the training phase is complete. Predetermined cardinality. b It can be two. The controller can be further programmed to define scaling factors for layers that include more than one node.
[0007] One method includes: selecting a predetermined number of candidate scaling factors, which are multiples of two, and evaluating a cost function for each candidate scaling factor, the cost function being the mean squared error between a predetermined set of floating-point weighting factors of the neural network and the product of the evaluated candidate scaling factor and a fixed-point weighting factor defined by the evaluated candidate scaling factor. The method further includes: selecting a scaling factor as one of the candidate scaling factors that results in a minimum cost function, and generating a set of fixed-point weighting factors by scaling each floating-point weighting factor by the scaling factor. The method also includes: implementing the neural network using the set of fixed-point weighting factors.
[0008] Candidate scaling factors can include those with exponential values. L and L-1 The first and second candidate values, such that the average of the absolute values of the floating-point weighting factors of the predetermined set is within 2L and 2 L-1 Between. Candidate scaling factors can include two from. 2 L-1 arrive 2 L +4 The candidate scaling factor can include a larger number of candidate scaling factors greater than the average of the absolute values of the floating-point weighted factors compared to the average of the candidate scaling factors. A predetermined set can correspond to the nodes of a neural network. Attached Figure Description
[0009] Figure 1 An example depicting a single node of a neural network.
[0010] Figure 2 An example of a single node in a neural network using fixed-point resolution is depicted.
[0011] Figure 3 A graph showing the accuracy associated with different weighting factor transformation strategies is presented.
[0012] Figure 4 A possible block diagram of a machine learning system is depicted.
[0013] Figure 5 A possible flowchart for selecting a scaling factor to convert a weighted factor into a fixed-point representation is depicted. Detailed Implementation
[0014] Embodiments of this disclosure are described herein. However, it is to be understood that the disclosed embodiments are merely examples, and other embodiments may take various forms and alternatives. These figures are not necessarily drawn to scale; some features may be enlarged or minimized to show detail of particular components. Therefore, the specific structural and functional details disclosed herein should not be construed as limiting, but merely as a representative basis for teaching those skilled in the art to employ the invention in various ways. As will be understood by those skilled in the art, various features illustrated and described with reference to any one of the figures may be combined with features illustrated in one or more other figures to produce embodiments not explicitly illustrated or described. The combinations of illustrated features provide representative embodiments for typical applications. However, for a particular application or implementation, various combinations and modifications of features consistent with the teachings of this disclosure may be desired.
[0015] Machine learning systems are being incorporated into a wide variety of modern systems. Machine learning systems are attractive because they can be trained or adapted to different situations. For example, by applying a training dataset to a machine learning algorithm, the system can adjust its internal weighting factors to achieve the desired result. The training dataset can include the set of inputs to the machine learning algorithm and the corresponding set of expected outputs. During training, the system can monitor the error between the expected output and the actual output generated by the machine learning algorithm to adjust or calibrate the weighting factors within the algorithm. Training can be repeated until the error falls below a predetermined level.
[0016] A machine learning system can be part of a larger application or system. For example, a machine learning system can be incorporated into a robotics application. In other examples, a machine learning system can be part of a vision system. For instance, a machine learning system can be trained to identify specific objects in a field of view from video images. A machine learning system can be further configured to provide control signals for controlling devices, such as robotic arms.
[0017] Neural networks can be used as part of machine learning systems. A neural network can consist of different interconnected stages between input and output. A neural network may include sequence encoder layers, prototype layers, and fully connected layers. Furthermore, a neural network may include one or more convolutional layers for feature extraction. Neural networks can be implemented on embedded controllers with limited memory and processing resources. Consequently, the amount of computation that can be performed for each time interval may be limited. As the size and number of neural networks increase, computational resources may become strained. That is, there may not be enough processing time to complete the computation within the desired time interval. Therefore, methods to reduce the computational load may be helpful.
[0018] A major limiting factor for practical deployment in resource-constrained environments is the need to maintain the precision of the weights. The precision of neural network weights can directly impact network performance and is therefore typically maintained as floating-point values. Consequently, mathematical operations related to these weights are also performed with floating-point precision. Floating-point variables may require more memory space than fixed-point variables. Furthermore, operations on floating-point variables generally require more processor cycles than operations on fixed-point variables. Note that the above discussion regarding weights can also be applied to the inputs and outputs of nodes / neurons. The strategies disclosed in this paper are applicable to any system with a set of coefficients or factors.
[0019] The observations above may suggest some paths to improve the performance of neural networks. The first improvement could be storing variables / data in fixed-point representations with reduced precision. The second improvement could be performing mathematical operations using fixed-point arithmetic. Thus, both data and variables can be maintained as fixed-point values with fewer bits (1 to 8), thereby reducing storage and computational complexity.
[0020] Typically, neural networks and all their input variables are represented in floating-point form (based on the processor requiring 16 / 32 / 64 bits, and another variable is needed to store the decimal place). Therefore, the corresponding mathematical operations are also performed in floating-point. This incurs considerable storage and processing overhead. For example, evaluating the output of a single node in a fully connected layer of a neural network can be done as follows: Figure 1 It is represented as shown. Figure 1 An example of a single node 100 in a fully connected layer is depicted. Node 100 can be configured to sum multiple weighted input values to generate an intermediate output value 108. Each input value 102 can be multiplied by a corresponding weighting factor 104. The product of each input value 102 and weighting factor 104 can be fed into a summing element 106 and summed together. In a typical implementation, each of the input value 102, weighting factor 104, and intermediate output value 108 can be implemented as a floating-point value. Note that the corresponding neural network layer can consist of many nodes among these nodes. The intermediate output value 108 can be passed through an activation function to generate the final output, and non-linear behavior can be introduced. For example, a rectified linear unit (RELU) can be used as the activation function to generate the final output. Figure 1 The configurations depicted can represent other elements of a neural network. For example, Figure 1 The diagram can represent any structure with a weighting factor that is applied to the input and fed into the summing element.
[0021] It may be beneficial to consider the case where the input 102 already has fixed-point precision. Note that when the input 102 is expressed in floating-point form, the strategy disclosed herein can be applied to convert the value to a fixed-point representation. Then, it can be expressed in two stages with fixed-point precision (using...). k (bit) represents the quantization process of weights. The first stage could be finding the factor. a Scale the data to a reduced range (e.g., [-1 to 1]), as shown below: The second stage could involve breaking down this scope into... n There are intervals, among which n =2 k -1, and the quantization / fixed-point scaling factor can be expressed as: .
[0022] A second stage can be implemented to make all intervals have the same size (linear quantization) or have different sizes (non-linear quantization). Since the goal is to perform mathematical operations after quantization, linear quantization is more common. k The value can represent the number of bits used to express the weight value as an integer or fixed-point value. Fixed-point weighting factors can be expressed by scaling factors. α The output is derived by scaling each floating-point weighting factor. In this example, scaling is performed by dividing the floating-point weighting factor by the scaling factor.
[0023] In this example, linear quantization can be used. To efficiently represent values using fixed-point variables, a scaling factor can be appropriately chosen. a The scaling factor can be adjusted. a The quantized weights are chosen such that they are as close as possible to the original floating-point weights in a mean-square sense. The following cost function can be minimized: in, w Represents the original weights. w q The weights representing the quantization, and a It is a scaling factor. The cost function can be described as a floating-point weighted factor ( w The product of the candidate scaling factor and the corresponding fixed-point weighting factor ( a*w q The mean square error between ( ).
[0024] Therefore, in the weights of the quantization layer (to k During the period (bit), one of the main operations affecting the quantization loss in equation (3) involves scaling the weights. Several different methods can be used to solve this optimization problem.
[0025] After training the neural network, several methods are applied to the weighting factor. These strategies can maintain the overhead of floating-point representation during the training process. The first method is the binary neural network approach. This method focuses on 1-bit neural networks (…). k =1). This method takes into account a =1 is an oversimplified approach. Therefore, the resulting weights are quantized based solely on the sign of the weights: .
[0026] The second method can be XOR (Exclusive NOR). This method again focuses primarily on 1-bit networks, but the same strategy also applies to larger numbers of bits. Under the assumption of Gaussian distributed weights, this method solves the optimization problem in equation (3) for the 1-bit case. The solution in closed form is derived as follows: a = E(|w|) It is the average of the absolute values of all weights in the layer.
[0027] The third method is the Ternary Weight Network (TWN) method. This method uses three levels for quantization, { -a, 0, a The scaling factor can be selected as in the XNOR method. a = E(|w|) .
[0028] The fourth method could be statistics-aware weight binning (SAWB). This method may be suitable for more than one ( k >1). Instead of using factors that rely solely on first-order statistics, second-order statistics are also applied. This method uses a heuristic approach, which derives the weights by assuming they can be derived from a fixed set of probability distributions. The scaling factor in this method can be given by: Among them, it was found through experiments c 1 and c 2 And for a given k The values are fixed. The above method is applied after training. That is, the training process is completed to learn the entire set of weights before the scaling factor is determined and the weights are quantized. Accordingly, the benefits of fixed-point representation are not realized during the training process.
[0029] exist Figure 2 The diagram depicts the operation of a single node 200 used to evaluate the output of a neural network layer. Node 200 may include a fixed-point input 202. The fixed-point input 202 is a value expressed as an integer or fixed-point representation. Node 200 further includes multiple fixed-point weighting factors 204 (…). W q1 ,…W qN The fixed-point weighting factor 204 can be derived as described above. Node 200 includes a summation block or function 206 configured to sum the product of the corresponding fixed-point input 202 and the fixed-point weighting factor 204. The output of the summation block 206 can be a fixed-point output 208. The output of the summation block can be expressed as: .
[0030] The discrete weighting factor 204 can be defined as: The fixed-point weighting factor 204 can be derived from the quantization function Q applied to the original floating-point weighting factor. The quantization function can be as described above.
[0031] The fixed-point output 208 can be multiplied by a scaling factor of 210 (shown as...). a ), to generate node output 212 as follows: The node output 2^12 can be represented as a floating-point value. The fixed-point output 2^08 multiplied by the scaling factor... a 210, to make the node output 212 reach the actual scale. Note that additional activation functions may be applied.
[0032] Quantization strategies can be configured to reduce the number of floating-point multiplications and divisions by choosing a scaling factor that is a multiple of two. By forcing the scaling factor to a multiple of two, multiplications and divisions involving the scaling factor can be performed via shift operations in the microprocessor. Shift operations are generally faster and / or more efficient than floating-point operations in the microprocessor. Solving the optimization problem of equation (3) can provide a more robust solution than using heuristic functions developed based on sample data or designed for a specific distribution. The specific problem then becomes determining what multiple of two should be chosen as the scaling factor.
[0033] A batch can be defined as a set of training samples or a training set applied to a neural network. During training, the weighting factor can be updated after each batch has been processed. A batch may result in multiple iterations through the neural network. Batch updates can be updates to the weighting factor after processing a predetermined number of training samples. Updating the scaling factor after each batch of training samples may unnecessarily utilize computational resources. For example, due to the discrete nature of the scaling factor, it is unlikely to change as rapidly as the weighting factor during training. Accordingly, the scaling factor can be changed after processing a predetermined number of batches during the training phase. The predetermined number of batches can depend on the set of training data and the application. The weighting factor can be updated more frequently than the scaling factor. For example, the scaling factor can be updated once for a predetermined number of updates to the weighting factor (e.g., 2-100).
[0034] The first step is to identify the scaling factor that will improve the weight quantization process. a To accomplish this, the optimization problem in equation (3) can be solved to select the scaling factor. aHowever, the optimization problem in equation (3) may not be solvable to the global minimum because it is non-convex and non-smooth. This may happen, for example, when the space of scaling factors is large and the weights are also optimized. An alternative is to brute-force search the scaling factors after each batch. Therefore, iterating through the scaling factors... a An infinite value is not a viable option. This is because existing methods rely on developing heuristics to estimate the scaling factor. a The main reason.
[0035] The method disclosed in this paper involves selecting candidate scaling factors from a finite set during training, following a predetermined number of batch updates. S Solve the problem in equation (3). The candidate scaling factors for a finite set can be restricted to multiples of two: Since the set is defined as having only a finite number of values (e.g., 10–20), the cost in Equation (3) can be evaluated for each member of the set. The weights can be quantized by choosing a scaling factor that minimizes the cost. The scaling factor may be the same for all layers, or it may be different for each layer, or it may be different for each kernel or neuron within a layer. While using a single scaling factor for each layer may be more common, using a different factor for each kernel or neuron is also acceptable. This strategy can be applied to any configuration in which the set of coefficients is to be discretized. For example, a kernel can describe the set of coefficients for the filters used in a convolutional layer. The disclosed strategy can be applied to the coefficients to generate a set of fixed-point coefficients.
[0036] Due to scaling factor a It can be chosen as a multiple of 2, so the scaling operation in equation (1) can be performed using a shift operation. Figure 2 The inverse scaling / rescaling operation shown is performed instead of using the floating-point multiplier in existing schemes. Therefore, Figure 2 The set of operations represented in the code does not use any floating-point multiplication. Depending on the size of the neural network, this will result in significant cost savings, and there could be tens of thousands to millions of such operations across all layers of the network. This can be beneficial during both the training and inference phases when implemented in real-time systems.
[0037] Since the scaling factor chosen through the proposed method may be a multiple of two, updating the weights in each iteration / batch during the training phase will not affect the optimal scaling factor set. Therefore, it is not necessary to evaluate equation (3) relative to S at each batch / iteration. Thus, the scaling factor can be chosen. αFurthermore, the cost function is minimized every predetermined number of batches (e.g., 10-100). Therefore, even though the complexity of evaluating candidate scaling factors relative to a single update might be larger than existing methods (averaging over dozens of iterations), this complexity is significantly lower, representing another advantage of the disclosed method. Moreover, it should be remembered that the scaling factor only needs to be iteratively updated during the training phase. During the inference phase, the scaling factor learned during training is used for different layers and is not updated.
[0038] As described, the cost function can be minimized by evaluating the cost function for each candidate scaling factor. Then, the scaling factor is selected as the candidate scaling factor that minimizes the cost function. This reduces the number of scaling factors that can be identified. a Another aspect of the complexity is reducing the number of sets. S The scaling factor is determined by the number of candidate scaling factors. It can be proven that for many common distributions (e.g., linear and Gaussian distributions), the scaling factor... a The following conditions must be met: a>M=mean(abs(W)) Therefore, sets S The selection can be made such that there are more candidate elements greater than the average and fewer candidate elements less than the average. The set of candidate scaling factors can be defined as follows: ,in The set defined above was used in experiments and found to be effective. However, it is understood that the described method is not limited to this scope. The number of candidate scaling factors can include a larger number of candidates with values exceeding the average of the absolute values of the set of floating-point weighted factors. The number of candidate scaling factors can include only one candidate that is less than the average of the absolute values of the set of floating-point weighted factors.
[0039] Although the scaling factor is reasonably defined as a multiple of two, it can be used as any other number (e.g., p∈{real integers} Scaling factors that are multiples of a predetermined base other than two are also within the scope of this disclosure. The scaling factor can be a multiple of a predetermined base other than two. For example, if the embedded system uses tri-state signals, the following set can be used to optimize the scaling factor. a : ,in This can be further generalized to any predetermined cardinality. b As: ,in .
[0040] Since the optimization problem in Equation (3) is solved for a reduced set of candidate scaling factors without making any assumptions about the probability distribution of the weights, the method remains reliable even if the base weights are derived from some random distribution. Therefore, this approach has a much wider range of applicability compared to existing heuristic or distribution-specific methods.
[0041] Figure 3 An example plot of the bimodal distribution is shown, illustrating how the root mean square error (RMSE) varies with the distance between the two modes for three different methods. Both modes are Gaussian, with a standard deviation of 0.02. The first curve 302 depicts the performance of the XNOR algorithm used by XNOR. The second curve 304 depicts the performance of the second strategy. The third curve 306 depicts the performance of the method presented in this paper. One can see that for most bimodal modes, the proposed method has a significantly lower error than the other strategies.
[0042] The method mentioned above learns a scaling factor for each layer of the neural network (or for the entire network or a part of the network) during training. a (which is a multiple of two (or any other base), and is the best scaling factor in the set of multiples of two.
[0043] The proposed scaling factor a This reduces computational complexity because it remains invariant across several batches of updates during the training phase and can be updated less frequently. Furthermore, since the proposed scaling factor can be a multiple of two, its implementation for quantization is hardware-friendly.
[0044] Scaling factor obtained from the method above a It is robust to changes in the probability density of the weights (e.g., from...) Figure 3 It is seen and designed to be applicable, regardless of the underlying distribution.
[0045] Figure 4 A block diagram of a possible machine learning system 400 is depicted. The machine learning system 400 can be implemented in one or more controllers. The controller may include a processor configured to execute instructions. The controller may further include volatile and non-volatile memory for storing programs and data. In a configuration with multiple controllers, the controllers may include circuitry and software for communicating with each other via a communication channel (e.g., Ethernet or others).
[0046] Machine learning system 400 may include neural network 402. Neural network 402 may include multiple layers and may consist of the aforementioned multiple nodes and / or kernels. Machine learning system 400 may include trainer 404. Trainer 404 may perform operations for training neural network 402. Machine learning system 400 may include training database 406, which includes input sets and corresponding outputs or labels of neural network 402. Training database 406 may include the expected output for each input set of neural network 402. Trainer 404 may coordinate the application of inputs and outputs of training database 406. For example, trainer 404 may cause neural network 402 to process a batch 418 of input data and output an updated weighting factor 424 after processing the batch. Neural network 402 may receive input data 418 from training database 406. Trainer 404 may receive expected output data 420 from training database 406. Neural network 402 may process input data 418 according to a neural network strategy to generate output data 416. For example, neural network 402 may include multiple nodes and / or kernels or some combination thereof. The trainer 404 can receive output data 416 from the neural network 402 for comparison with expected output data 420. The neural network can operate on the input data 418 using a set of fixed-point weighting factors 423.
[0047] The trainer 404 can monitor the performance of the neural network 402. When the output data 416 is not closely related to the expected output data 420, the trainer 404 can generate a weighting factor adjustment 424. Various known strategies can be used to adjust the weighting factors. The trainer 404 can iterate through the training data until the output data 416 is within a predetermined range of the expected output data 420. The training phase can be completed when the error is less than a predetermined threshold. The weighting factor adjustment 424 can be input to the shadow weight function 405. The shadow weight function 405 can maintain the set of weighting factors of the neural network 402 in full precision. The output of the shadow weight function 405 can be a set of full-precision weighting factors 422. For example, the full-precision weighting factors 422 can be represented as floating-point variables.
[0048] The machine learning system 400 may further include a quantizer 410 configured to convert a full-precision weighting factor 422 learned during the training phase into a fixed-point weighting factor 423 used in the neural network 402. For example, the quantizer 410 may apply equations (1) and (2) to generate the fixed-point weighting factor 423. The fixed-point weighting factor 423 may be provided to the neural network 402 during the training phase. The quantizer 410 may use a scaling factor determined as described above. During the training phase, the quantizer 410 may continuously convert the full-precision weighting factor 422 into the fixed-point weighting factor 423 at each iteration. The fixed-point weighting factor 423 may be generated using a currently provided scaling factor 425.
[0049] The machine learning system 400 may further include a scaling factor determination function 411, which is configured to generate a scaling factor for the quantizer 410. a 425. For example, the scaling factor determination function 411 may be implemented periodically during the training phase. For example, the scaling factor determination function 411 may be implemented once every tenth batch or iteration of the training phase. More generally, the scaling factor determination function 411 may be implemented after a predetermined number of batches or iterations during the training phase. The scaling factor determination function 411 may include selecting a predetermined set of candidate scaling factors for evaluation in the cost function. The scaling factor determination function 411 may further include evaluating each candidate scaling factor in the cost function to determine the scaling factor 425 that minimizes the cost function. The scaling factor determination function 411 may use the scaling factor... a The scaling factor 425 is output to quantizer 410. Quantizer 410 can output a fixed-point weighting factor 423 for use by neural network 402. After the training phase is complete, the fixed-point weighting factor 423 can be provided to the inference phase 408. Using the fixed-point weighting factor 423 during the training phase can improve the performance of the training operation. While some floating-point operations may still be used during the training phase to maintain the full-precision weighting factor 422, fewer floating-point operations are used in neural network 402. Another advantage is that the same fixed-point weighting factor 423 is used in both the training phase (in neural network 402) and the inference phase 408. During the training phase, as the full-precision weighting factor 422 is updated, the fixed-point weighting factor 423 can change with each iteration or batch. After a predetermined number of batches or iterations, the scaling factor 425 can be updated to change the scaling operation of quantizer 410.
[0050] The inference phase 408 can implement the neural network algorithm as part of the real-time system. In this way, the neural network can be implemented in an optimal manner for real-time operation. The inference phase 408 can receive actual input 412 and generate actual output 414 based on the operation of the neural network. The actual input 412 can come from sensor input. The inference phase 408 can be incorporated into the real-time system to process the input data set. For example, the inference phase 408 can be incorporated into a machine vision system and configured to recognize specific objects in image frames. During the inference phase, the weighting factors of the neural network can be maintained at the fixed-point values learned during the training phase.
[0051] Since the inference phase 408 can be a real-time system, it may be desirable to use fixed-point or integer operations at runtime. Accordingly, the inference phase 408 can be configured to operate using fixed-point or integer operations to improve computational throughput. The inference phase 408 may include logic for processing layers and nodes. Figure 5 A possible flowchart 500 is depicted for a set of operations used to convert floating-point weighted factors of a neural network into a fixed-point representation. At operation 502, a set of weighted factors can be generated. This set of weighted factors can be generated during the training phase and can correspond to one or more elements of the neural network (e.g., nodes, neurons, kernels, layers). During this phase, the weighted factors can be represented as floating-point values. This set of weighted factors can correspond to a set of nodes or an entire layer of the neural network. This set of weighted factors can be the output of the training operations and can be updated after one or more iterations of the training phase.
[0052] At operation 504, a set of candidate scaling factors can be selected. A predetermined number of candidate scaling factors can be selected. This candidate scaling factor can be a multiple of a predetermined base (e.g., two). The candidate scaling factor can include a larger number of candidates whose values exceed the average of the absolute values of the floating-point weighted factors. The average can have a set of weighted factors associated with a node or layer. Candidate scaling factors can include only one candidate whose absolute value is less than the average of the absolute values of the floating-point weighted factors. Candidate scaling factors can include those with exponents... L and L -1 The first and second candidates make the average of the absolute values of the set of floating-point weighted factors in b L and b L-1 Between. Candidate scaling factors can include a predetermined base from... b L-1 arrive b L +4 Those multiples. Minimizing the number of candidate scaling factors to reduce the execution time used to evaluate the cost function may be useful.
[0053] At operation 506, candidate scaling factors can be evaluated in the cost function. The cost function can be the mean squared error between the set of floating-point weighted factors and the product of the candidate scaling factor being evaluated and the fixed-point weighted factor defined by the candidate scaling factor being evaluated. The cost function can be expressed as equation (3) above.
[0054] At operation 508, candidate scaling factors that minimize the cost function can be selected. Each candidate scaling factor can be evaluated within the cost function to determine the numerical values used for comparison. For example, the cost function can be evaluated for each candidate scaling factor to generate numerical values. These values can then be compared to determine the minimum value.
[0055] At operation 510, a fixed-point weighting factor can be generated using the selected scaling factor. For example, a fixed-point weighting factor can be generated using equations (1) and (2) above. During the training phase, operations 502 through 510 can be repeated periodically. After the training phase is complete, the final set of quantized weighting factors can be used for the inference phase.
[0056] At operation 512, the fixed-point weighting factor and scaling factor can be provided to the inference phase. The fixed-point weighting factor and scaling factor can be passed via a communication channel in a configuration where the inference phase is implemented in a separate controller. At operation 514, the fixed-point weighting factor and scaling factor can be used during the operation of the inference engine.
[0057] The system and method disclosed in this paper present an improved approach to quantizing weighting factors for neural networks. This method generates scaling factors that minimize the cost function. Furthermore, it reduces the number of candidate scaling factors to be evaluated in the cost function and decreases the computational overhead for transforming the weighting factors.
[0058] The processes, methods, or algorithms disclosed herein may be deliverable to, or implemented by, a processing device, controller, or computer, which may include any existing programmable electronic control unit or dedicated electronic control unit. Similarly, processes, methods, or algorithms may be stored in many forms as data and instructions executable by a controller or computer, including but not limited to information permanently stored on a non-writable storage medium such as a ROM device, and information reproducibly stored on a writable storage medium such as a floppy disk, magnetic tape, CD, RAM device, and other magnetic and optical media. These processes, methods, or algorithms may also be implemented in a software executable object. Alternatively, these processes, methods, or algorithms may be embodied, in whole or in part, using suitable hardware components such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), state machines, controllers, or other hardware components or devices, or combinations of hardware, software, and firmware components.
[0059] While exemplary embodiments have been described above, they are not intended to describe all possible forms covered by the claims. The terms used in this specification are descriptive and not limiting, and it is understood that various changes may be made without departing from the spirit and scope of this disclosure. As previously stated, features of various embodiments may be combined to form other embodiments of the invention, which may not be explicitly described or illustrated. While various embodiments may be described as offering advantages over other embodiments or prior art implementations with respect to one or more desired characteristics, or being preferred over other embodiments or prior art implementations, those skilled in the art will recognize that one or more features or characteristics may be compromised to achieve desired overall system properties, depending on the specific application and implementation. These properties may include, but are not limited to, cost, strength, durability, lifecycle cost, marketability, appearance, packaging, size, availability, weight, manufacturability, ease of assembly, etc. Accordingly, embodiments described with respect to one or more characteristics as less desirable compared to other embodiments or prior art implementations are not outside the scope of this disclosure and may be desirable for a particular application.
Claims
1. A method for implementing a neural network for real-time operation, the method comprising: Train the neural network, converting its floating-point weighting factors to fixed-point weighting factors during the training process; Image frames are received as actual input by a trained neural network; as well as The image frame is manipulated by a trained neural network using a fixed-point weighting factor; The conversion of the floating-point weighting factor of the neural network to the fixed-point weighting factor includes: Select a predetermined number of candidate scaling factors, which are multiples of a predetermined base; Evaluate each candidate scaling factor in the cost function; The scaling factor is selected as one of the candidate scaling factors that results in the minimum value of the cost function; Fixed-point weighting factors are generated by scaling floating-point weighting factors using scaling factors.
2. The method according to claim 1, wherein, The predetermined base number is two.
3. The method according to claim 1, wherein, A predetermined number of candidate scaling factors includes a larger number of candidates whose values exceed the average of the absolute values of the floating-point weighted factors.
4. The method according to claim 1, wherein, A predetermined number of candidate scaling factors include only one candidate that is less than the average of the absolute values of the associated floating-point weighted factors.
5. The method according to claim 1, wherein, The cost function is the mean square error between the product of the floating-point weighting factor and the candidate scaling factor and the corresponding fixed-point weighting factor.
6. The method of claim 1, further comprising: The scaling factor is updated during the training phase of the neural network after a predetermined number of training intervals.
7. A machine learning system, comprising: A neural network that is trained using input data from a training database; The controller is programmed to convert the floating-point weighting factors of the neural network to fixed-point weighting factors using a scaling factor that is a predetermined base. b The scaling factor is a multiple of the scaling factor and minimizes the cost function, which is the mean square error between the floating-point weighting factor and the product of the candidate scaling factor and the corresponding fixed-point weighting factor, and the scaling factor is changed after a predetermined number of iterations during the training phase; wherein the controller is further programmed to implement the neural network using fixed-point operations, wherein an image frame is received as the actual input and the image frame is operated on using the fixed-point weighting factor.
8. The machine learning system according to claim 7, wherein, Candidate scaling factors include: those with exponential values respectively. L and L - 1 The first and second candidate values make the average of the absolute values of the floating-point weighting factors in b L and b L-1 between.
9. The machine learning system according to claim 8, wherein, The controller is further programmed to evaluate the cost function using candidate scaling factors, which are derived from a predetermined base. b L-1 arrive b L+4 Multiples of.
10. The machine learning system according to claim 7, wherein, The controller is further programmed to evaluate the cost function for a first number of candidate scaling factors and a second number of candidate scaling factors, wherein the first number of candidate scaling factors is greater than the average of the absolute values of the floating-point weighted factors, and the second number of candidate scaling factors is less than the average value, with the first number being greater than the second number.
11. The machine learning system according to claim 7, wherein, The controller is further programmed to provide fixed-point weighting factors to the inference phase, which is configured to implement the neural network, after the training phase is completed.
12. The machine learning system according to claim 7, wherein, Predetermined base b It is two.
13. The machine learning system according to claim 7, wherein, The controller is further programmed to define scaling factors for layers that include more than one node.
14. A method for implementing a neural network for real-time operation, comprising: Train the neural network, during which the floating-point weighting factors of the neural network are converted into a set of fixed-point weighting factors; Image frames are received as actual input by a trained neural network; as well as The trained neural network operates on the image frame using this set of fixed-point weighting factors; The conversion of the floating-point weighting factors of the neural network into a set of fixed-point weighting factors includes: Select a predetermined number of candidate scaling factors, which are multiples of two; For each candidate scaling factor, a cost function is evaluated, which is the mean squared error between the product of a predetermined set of floating-point weighting factors of the neural network and the product of the candidate scaling factor being evaluated and a fixed-point weighting factor defined by the candidate scaling factor being evaluated. Choose the scaling factor as one of the candidate scaling factors that leads to the minimum of the cost function; A set of fixed-point weighting factors is generated by scaling each floating-point weighting factor according to the scaling factor.
15. The method according to claim 14, wherein, Candidate scaling factors include: those with exponential values respectively. L and L-1 The first and second candidate values, such that the average of the absolute values of the floating-point weighting factors of the predetermined set is within 2 L and 2 L-1 between 。 16. The method according to claim 15, wherein, Candidate scaling factors include two from 2 L-1 arrive 2 L+4 Multiples of.
17. The method of claim 14, wherein, Candidate scaling factors include a larger number of candidate scaling factors that are greater than the average of the absolute values of the floating-point weighted factors compared to the average of the candidate scaling factors that are less than the average of the absolute values of the floating-point weighted factors.
18. The method according to claim 16, wherein, The predefined set corresponds to the nodes of the neural network.
Citation Information
Patent Citations
Implementing neural networks in fixed point arithmetic computing systems
CN108427991A
Method and apparatus for learning low-precision neural network
CN109754063A