Method, device and related product for processing data
By adopting multiple pairs of asymmetric truncation threshold quantization methods in the machine learning model and selecting the optimal truncation threshold to reduce the quantization accuracy loss, the problem of large accuracy loss in the quantization process in the existing technology is solved, and the processing efficiency and accuracy are improved.
Patent Information
- Application Number
- CN201910804618.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-08-28
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2039-11-25
AI Technical Summary
When existing technologies perform large-scale data processing, there is a large loss of precision in the quantization process of machine learning models. Traditional methods such as KL divergence cannot effectively reduce the quantization precision loss, resulting in low processing efficiency.
A multi-pair asymmetric truncation threshold quantization method is adopted. By selecting truncation thresholds with different absolute values for the upper and lower limits, and selecting the optimal truncation threshold based on the mean absolute value difference of the data before and after quantization, the loss of quantization accuracy can be reduced.
This achieves smaller precision loss during the quantization process, improves the efficiency and accuracy of data processing, and reduces the consumption of computing resources.
Smart Images

Figure CN112446496B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to the field of computer technology, and more particularly to methods, devices, and related products for processing data. Background Art
[0002] With the continuous development of artificial intelligence (AI) technology, its application areas are becoming increasingly broad, with successful applications in areas such as image recognition, speech recognition, and natural language processing. However, as the complexity and accuracy of AI algorithms increase, machine learning models are becoming larger and larger, requiring an ever-increasing amount of data to be processed. Processing large amounts of data requires significant computational and time overhead, resulting in low processing efficiency. Summary of the Invention
[0003] In view of this, embodiments of the present disclosure provide a method, apparatus, and related products for processing data.
[0004] In a first aspect of the present disclosure, a method for processing data is provided. The method includes: obtaining a set of data to be quantified for a machine learning model; determining multiple sets of quantized data by quantizing the set of data to be quantified using multiple pairs of truncation thresholds, wherein each pair of truncation thresholds in the multiple pairs of truncation thresholds includes a truncation upper limit and a truncation lower limit, and the truncation upper limit and the truncation lower limit in at least one pair of truncation thresholds in the multiple pairs of truncation thresholds have different absolute values; and selecting a pair of truncation thresholds from the multiple pairs of truncation thresholds based on a difference between a mean of the absolute values of each set of quantized data in the multiple sets of quantized data and a mean of the absolute values of the set of data to be quantified, for quantizing the set of data to be quantized.
[0005] In a second aspect of the present disclosure, a device for processing data is provided. The device includes: a data acquisition unit for obtaining a set of data to be quantified for a machine learning model; a quantized data determination unit for determining multiple sets of quantized data by quantizing the set of data to be quantized using multiple pairs of truncation thresholds, wherein each pair of truncation thresholds in the multiple pairs of truncation thresholds includes a truncation upper limit and a truncation lower limit, and the truncation upper limit and the truncation lower limit in at least one pair of truncation thresholds in the multiple pairs of truncation thresholds have different absolute values; and a truncation threshold selection unit for selecting a pair of truncation thresholds from the multiple pairs of truncation thresholds based on a difference between the mean of the absolute values of each set of quantized data in the multiple sets of quantized data and the mean of the absolute values of the set of data to be quantized, for quantizing the set of data to be quantized.
[0006] In a third aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed, the method according to each embodiment of the present disclosure is implemented.
[0007] In a fourth aspect of the present disclosure, an artificial intelligence chip is provided, which includes a device for processing data according to various embodiments of the present disclosure.
[0008] In a fifth aspect of the present disclosure, an electronic device is provided, which includes the artificial intelligence chip according to various embodiments of the present disclosure.
[0009] In a sixth aspect of the present disclosure, a board is provided, comprising: a memory device, an interface device, a control device, and an artificial intelligence chip according to various embodiments of the present disclosure. The artificial intelligence chip is connected to the memory device, the control device, and the interface device; the memory device is used to store data; the interface device is used to enable data transmission between the artificial intelligence chip and external devices; and the control device is used to monitor the status of the artificial intelligence chip.
[0010] By deducing the technical features in the claims, the beneficial effects of the technical problems in the background technology can be achieved. Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.
[0012] Figure 1 A schematic diagram showing a processing system of a method for processing data according to an embodiment of the present disclosure;
[0013] Figure 2 A schematic diagram illustrating an example architecture of a neural network according to an embodiment of the present disclosure;
[0014] Figure 3 A schematic diagram illustrating a process for quantizing data according to an embodiment of the present disclosure is shown;
[0015] Figure 4A A schematic diagram for symmetrically quantizing data according to an embodiment of the present disclosure is shown;
[0016] Figure 4B FIG2 shows a schematic diagram for symmetrically quantizing data based on a truncation threshold according to an embodiment of the present disclosure;
[0017] Figure 4C FIG2 shows a schematic diagram for asymmetrically quantizing data according to an embodiment of the present disclosure;
[0018] Figure 4D FIG2 shows a schematic diagram for asymmetrically quantizing data based on a truncation threshold according to an embodiment of the present disclosure;
[0019] Figure 5 A flowchart of a method for processing data according to an embodiment of the present disclosure is shown;
[0020] Figure 6 A flowchart of a method for searching a truncation threshold for asymmetric quantization according to an embodiment of the present disclosure is shown;
[0021] Figure 7A A schematic diagram illustrating a coarse-grained search for a truncation threshold for asymmetric quantization according to an embodiment of the present disclosure is shown;
[0022] Figure 7B A schematic diagram illustrating a fine-grained search for a truncation threshold for asymmetric quantization according to an embodiment of the present disclosure is shown;
[0023] Figure 8 A flowchart of a method for iteratively searching for an optimal truncation threshold according to an embodiment of the present disclosure is shown;
[0024] Figure 9 A block diagram showing an apparatus for processing data according to an embodiment of the present disclosure; and
[0025] Figure 10 A structural block diagram of a board according to an embodiment of the present disclosure is shown. Specific embodiments
[0026] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the embodiments described are part of the embodiments of the present disclosure, not all of them. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0027] It should be understood that the terms "first," "second," "third," and "fourth," etc. in the claims, specification, and drawings of the present disclosure are used to distinguish different objects rather than to describe a specific order. The terms "include" and "comprising" used in the specification and claims of the present disclosure indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.
[0028] It should also be understood that the terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the present disclosure. As used in this disclosure and the claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should further be understood that the term "and / or" as used in this disclosure and the claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.
[0029] As used in this specification and claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0030] Generally speaking, when quantizing data, choosing a wide range of values results in lower precision after quantization. However, choosing a narrow range results in excessive data truncation, leading to information loss in the data on both sides. The range of values refers to the range between the lower and upper truncation limits used to quantize the data. Therefore, it is necessary to find a suitable pair of truncation thresholds to minimize or minimize data quantization loss. Traditionally, the optimal truncation threshold is determined using the KL divergence (Kullback–Leibler divergence). The KL divergence measures the correlation between the data before and after quantization. KL divergence is also known as relative entropy, information divergence, or information gain. The KL divergence measures the difference between two probability distributions, P and Q. Assuming that P is the distribution of 32-bit floating-point numbers before quantization and Q is the distribution of 8-bit integers after quantization, the smaller the KL divergence between P and Q, the closer the distributions before and after quantization, and the more effective the quantization. However, the inventors of the present application have found that the quantization effect achieved by the truncation threshold obtained by the traditional KL method is not good, and usually causes a large loss of accuracy.
[0031] To this end, an embodiment of the present disclosure proposes a new scheme for determining a truncation threshold for asymmetric quantization, which can achieve a smaller quantization accuracy loss than traditional techniques (such as the KL method). According to an embodiment of the present disclosure, after obtaining a set of data to be quantized for a machine learning model, a set of data to be quantized is quantized using multiple pairs of truncation thresholds to determine multiple sets of quantized data, wherein each pair of truncation thresholds in the multiple pairs of truncation thresholds includes a truncation upper limit and a truncation lower limit, and the truncation upper limit and the truncation lower limit in at least one pair of truncation thresholds in the multiple pairs of truncation thresholds have different absolute values, i.e., an asymmetric pair of truncation thresholds. Then, the difference between the mean of the absolute values of each set of quantized data and the mean of the absolute values of a set of data to be quantized is used as an evaluation index to select a suitable pair of truncation thresholds from the multiple pairs of truncation thresholds. In this way, a pair of truncation thresholds that is more suitable for quantization can be found. In addition, compared with symmetric quantization, asymmetric quantization can further reduce the accuracy loss of quantization.
[0032] The following references Figures 1 to 10 It should be understood that these exemplary embodiments are provided only to enable those skilled in the art to better understand and implement the embodiments of the present disclosure, and are not intended to limit the scope of the present disclosure in any way.
[0033] Figure 1 FIG. 1 is a schematic diagram of a processing system 100 for a method for processing data according to an embodiment of the present disclosure. Figure 1 As shown, processing system 100 includes multiple processors 101-1, 101-2, and 101-3 (collectively referred to as processors 101) and memory 102. Processors 101 are used to execute instruction sequences, and memory 102 is used to store data, which may include random access memory (RAM) and register files. Multiple processors 101 in processing system 100 can share some memory space, such as some RAM storage space and register files, or have their own memory space.
[0034] It should be understood that the various methods according to the embodiments of the present disclosure can be applied to any processor of a processing system 100 (e.g., an artificial intelligence chip) including multiple processors (multi-core). The processor can be a general-purpose processor, such as a CPU (Central Processing Unit), or an artificial intelligence processor (IPU) for performing artificial intelligence operations. Artificial intelligence operations may include machine learning operations, brain-like operations, etc. Among them, machine learning operations include neural network operations, k-means operations, support vector machine operations, etc. The artificial intelligence processor may, for example, include one or a combination of a GPU (Graphics Processing Unit), an NPU (Neural-Network Processing Unit), a DSP (Digital Signal Processing Unit), and a Field-Programmable Gate Array (FPGA) chip. The present disclosure does not limit the specific type of processor. In addition, the types of multiple processors in the processing system 100 may be the same or different, and the present disclosure does not limit this.
[0035] In one possible implementation, the processor mentioned in this disclosure may include multiple processing units, each of which can independently execute various assigned tasks, such as convolution tasks, pooling tasks, or fully connected tasks. This disclosure does not limit the processing units or the tasks they execute.
[0036] Figure 2A schematic diagram illustrates an example architecture of a neural network 200 according to an embodiment of the present disclosure. A neural network (NN) is a mathematical model that mimics the structure and function of biological neural networks. A neural network performs computations by connecting a large number of neurons. Therefore, a neural network is a computational model composed of a large number of interconnected nodes (or "neurons"). Each node represents a specific output function, called an activation function. Each connection between two neurons represents a weighted value of the signal passing through that connection, called a weight, which acts as the memory of the neural network. The output of a neural network varies depending on the connections between neurons, as well as the weights and activation functions. In a neural network, a neuron is the basic unit of the network. It receives a certain number of inputs and a bias, and when a signal (value) arrives, it is multiplied by a weight. A connection connects a neuron to another neuron in another layer or the same layer, and the connection is accompanied by an associated weight. Additionally, a bias is an additional input to a neuron that is always 1 and has its own connection weight. This ensures that a neuron will activate even if all inputs are empty (all 0s).
[0037] In practice, if a nonlinear function isn't applied to the neurons in a neural network, the neural network is simply a linear function and is no more powerful than a single neuron. If the output of a neural network is set between 0 and 1, for example, in the example of cat-dog identification, outputs close to 0 can be considered cats, and outputs close to 1 can be considered dogs. To achieve this goal, activation functions, such as the sigmoid activation function, are introduced into the neural network. All you need to know about this activation function is that its return value is a number between 0 and 1. Therefore, the activation function is used to introduce nonlinearity into the neural network, narrowing the results of the neural network's calculations to a smaller range. In reality, how the activation function is expressed is not important; what is important is that a nonlinear function is parameterized through weights, and the nonlinear function can be modified by changing these weights.
[0038] like Figure 2 As shown in FIG, it is a structural diagram of the neural network 200. Figure 2 The neural network shown in FIG. 1 includes three layers, namely, an input layer 210, a hidden layer 220, and an output layer 230. Figure 2 The hidden layer 220 shown has three layers. Of course, the hidden layer 220 can also include more or fewer layers. The neurons in the input layer 210 are called input neurons. As the first layer in the neural network, the input layer receives input signals (values) and passes them to the next layer. It does not perform any operations on the input signals (values) and has no associated weights and biases. Figure 2In the neural network shown, it can receive 4 input signals (values).
[0039] Hidden layer 220 is a neuron (node) used to apply different transformations to the input data. A hidden layer is a collection of neurons arranged vertically (Representation). Figure 2 The neural network shown includes three hidden layers. The first hidden layer has four neurons (nodes), the second layer has six neurons, and the third layer has three neurons. Finally, the hidden layers pass values to the output layer 230. Figure 2 The neural network 200 shown in FIG. 2 has three hidden layers in which each neuron is fully connected. Each neuron in each of the three hidden layers is connected to each neuron in the next layer. It should be noted that not every hidden layer of a neural network is fully connected.
[0040] The neurons in the output layer 230 are called output neurons. The output layer receives the output from the last hidden layer. Through the output layer 230, the desired value and the desired range can be determined. Figure 2 In the neural network shown, the output layer has 3 neurons, that is, there are 3 output signals (values).
[0041] In practical applications, the role of neural networks is to provide a large amount of sample data (including input and output) for training in advance. After the training is completed, the neural network is used to obtain an accurate output for the input of the future real environment.
[0042] Before discussing neural network training, we need to define the loss function. A loss function is a function that indicates how well a neural network performs on a particular task. The most straightforward approach is to pass each sample through the neural network during training, obtaining a number. The difference between this number and the desired actual value is then squared. This yields the distance between the predicted value and the true value. The goal of training a neural network is to minimize this distance, or the value of the loss function.
[0043] When training a neural network, the weights are randomly initialized. Obviously, a randomly initialized neural network will not produce good results. During training, we might start with a poorly trained neural network and then train it to produce a highly accurate one. At the same time, we also want the loss function to be extremely small by the end of training.
[0044] The training process of a neural network is divided into two stages. The first stage is the forward processing of the signal, from the input layer 210 through the hidden layer 220, and finally to the output layer 230. The second stage is the back propagation of the gradient, from the output layer 230 to the hidden layer 220, and finally to the input layer 210. The weights and biases of each layer in the neural network are adjusted in turn according to the gradient.
[0045] During the forward process, input values are fed into the neural network's input layer 210, and outputs, known as predicted values, are obtained from the neural network's output layer 230. When the input values are provided to the neural network's input layer 210, no operations are performed. Within the hidden layers, the second hidden layer receives the predicted intermediate values from the first hidden layer, performs computations and activations, and then passes the resulting predicted intermediate values to the next hidden layer. Subsequent layers perform the same operations, ultimately obtaining output values at the neural network's output layer 230.
[0046] After forward processing, an output value, called a predicted value, is obtained. To calculate the error, a loss function is used to compare the predicted value with the actual output value, resulting in the corresponding error value. Backpropagation uses the chain rule from differential calculus. In this chain rule, the derivatives of the error value with respect to the weights of the last layer of the neural network are first calculated. These derivatives are called gradients. These gradients are then used to calculate the gradients of the penultimate layer of the neural network. This process is repeated until the gradients for each weight in the neural network are obtained. Finally, the corresponding gradients are subtracted from the weights, thereby updating the weights to reduce the error value.
[0047] Furthermore, for neural networks, fine-tuning involves loading a trained neural network. The fine-tuning process is identical to the training process and consists of two phases: the first is forward signal processing, and the second is backpropagation of gradients to update the weights of the trained neural network. The difference between training and fine-tuning is that training randomly processes an initialized neural network, training it from scratch, while fine-tuning does not.
[0048] During the training or fine-tuning of a neural network, each time the network undergoes forward signal processing and the corresponding backpropagation of errors, the weights in the network are updated using the gradient. This process is called an iteration. To achieve a neural network with the desired accuracy, the training process requires a very large sample dataset. In this case, it is impossible to input the sample dataset into the computer all at once. Therefore, to solve this problem, the sample dataset is divided into multiple blocks and passed to the computer. After each block is forward processed, the neural network weights are updated. When a complete sample dataset passes through the neural network once and returns a corresponding weight update, this process is called an epoch. In practice, passing the complete dataset through the neural network once is not sufficient; it needs to be passed through the same network multiple times, requiring multiple epochs, to ultimately achieve a neural network with the desired accuracy.
[0049] When training or fine-tuning a neural network, speed and accuracy are generally desired. Neural network data is represented in high-precision formats, such as floating-point numbers. Therefore, during training or fine-tuning, all data involved is in high-precision formats. The trained neural network is then quantized. For example, consider the weights of the entire neural network, where the quantized weights are all 8-bit fixed-point numbers. Since a neural network often has millions of connections, almost all the space is occupied by the weights of these connections. Furthermore, these weights are all different floating-point numbers. The weights of each layer tend to be normally distributed within a certain interval, such as (-3.0, 3.0). The maximum and minimum values of the weights for each layer in the neural network are stored, and each floating-point value is represented as an 8-bit fixed-point number. Within the range of the maximum and minimum values, the space is linearly divided into 256 quantization intervals, each represented by an 8-bit fixed-point number. For example, in the interval (-3.0, 3.0), byte 0 represents -3.0, and byte 255 represents 3.0. Similarly, byte 128 represents 0.
[0050] For data represented in high-precision formats, such as floating-point numbers, computer architecture shows that based on the arithmetic representation rules of floating-point and fixed-point numbers, for fixed-point and floating-point operations of the same length, floating-point calculations are more complex and require more logic devices to construct the floating-point unit. Therefore, floating-point units are physically larger than fixed-point units. Furthermore, floating-point units require more processing resources, resulting in a power consumption difference of orders of magnitude between fixed-point and floating-point operations. In short, floating-point units occupy many times more chip area and consume significantly more power than fixed-point units.
[0051] Figure 3 FIG. 3 is a schematic diagram illustrating a process 300 for quantifying data according to an embodiment of the present disclosure. Figure 3 The input data 310 may be a floating point number to be quantized, such as a 32-bit floating point number. If the input data 310 is directly input into the neural network model 340 for processing, it will consume more computing resources and the processing speed will be slow. Therefore, at block 320, the input data may be quantized to obtain quantized data 330 (e.g., an 8-bit integer). If the quantized data 330 is input into the neural network model 340 for processing, since 8-bit integers are faster to calculate, the neural network model 340 will complete the processing of the input data more quickly and generate the corresponding output result 350.
[0052] The quantization process from the input data 310 to be quantized to the quantized data 330 will cause some precision loss to a certain extent, and the degree of precision loss will directly affect the accuracy of the output result 350. Therefore, during the quantization process of the input data 330, it is necessary to ensure that the precision loss of the quantization process is minimized or as small as possible.
[0053] Figure 4A 4 shows a diagram 400 for symmetrically quantizing data according to an embodiment of the present disclosure. Figure 4A As shown in Figure 1, this is the simplest symmetric quantization method. It directly selects the absolute maximum value of all the data to be quantized, namely |max|, and then quantizes within the range of -|max| to |max| to generate the quantized data. However, this method does not perform any truncation, resulting in low accuracy in the quantized data. Symmetric quantization also results in some data waste. For example, there are no data points around the quantized maximum value of 127.
[0054] Figure 4B 450 is shown for symmetrically quantizing data based on a truncation threshold according to an embodiment of the present disclosure. Figure 4A Direct quantization method in Figure 4B Select a cutoff threshold T in , and the data outside the range of -|T| to |T| will be set to -|T| or |T|. For example, Figure 4B In the example, the three values to be quantized in circle 460 are outside the truncation range and are therefore treated as -|T| for quantization, resulting in data point 470. In this way, by using the truncation threshold to narrow the range of values for the data to be quantized, the accuracy of the quantized data can be improved.
[0055] Figure 4C 480 is shown for asymmetrically quantizing data according to an embodiment of the present disclosure. Figure 4CAs shown in Figure 1, this is an asymmetric quantization method that directly selects the maximum value (max) and the minimum value (min) of all the data to be quantized, and then quantizes within the range from min to max to generate the quantized data. However, this method does not perform any truncation, resulting in lower accuracy in the quantized data.
[0056] Figure 4D Graph 490 is shown for asymmetrically quantizing data based on a truncation threshold according to an embodiment of the present disclosure. Figure 4C Direct quantization method in Figure 4D Select a cutoff upper limit T and a cutoff lower limit min, and then the data outside the range of min to T will be set to min or T. For example, in Figure 4D In the example, the two values to be quantized in circle 492 are outside the truncation range and are therefore treated as value T for quantization, resulting in data point 495. By using asymmetric upper and lower truncation limits to narrow the range of values for the quantized data, the accuracy of the quantized data can be improved. However, determining a pair of asymmetric truncation thresholds that minimizes loss in quantization accuracy remains a pressing technical challenge.
[0057] Figure 5 FIG. 5 is a flow chart of a method 500 for processing data according to an embodiment of the present disclosure. It should be understood that the method 500 may be implemented by referring to FIG. Figure 1 The described one or more processors 101 are executed.
[0058] In block 502, a set of data to be quantified for a machine learning model is obtained. Figure 3 , the input data to be quantized 310 can be obtained and the input data can be quantized, thereby speeding up the processing speed of the neural network model 340. In addition, some parameters of the neural network model itself (such as weights, etc.) can also be quantized. By quantizing the network parameters, the size of the neural network model can be reduced. In some embodiments, the data to be quantized can be a 32-bit floating point number. Alternatively, the data to be quantized can also be a floating point number of other bits, or other data types.
[0059] In block 504, multiple groups of quantized data are determined by using multiple pairs of truncation thresholds to quantize a group of data to be quantized, respectively. Each pair of truncation thresholds in the multiple pairs of truncation thresholds includes a truncation upper limit and a truncation lower limit, and the truncation upper limit and the truncation lower limit in at least one pair of the multiple pairs of truncation thresholds have different absolute values. In other words, the multiple pairs of truncation thresholds include at least one pair of asymmetric truncation thresholds. In an asymmetric quantization scheme, each pair of truncation thresholds includes a truncation upper limit and a truncation lower limit, and each pair of the truncation upper limit and the truncation lower limit are typically asymmetric, i.e., their absolute values are different. However, in some cases, one or more of the multiple pairs of truncation thresholds determined may be symmetric, but at least one pair of truncation thresholds may be asymmetric. In some embodiments, the truncation lower limit may not be the minimum value in the data to be quantized, but may be other values.
[0060] According to embodiments of the present disclosure, multiple pairs of truncation thresholds can be selected to quantize the data to be quantized. In some embodiments, truncation thresholds can be selected at fixed intervals. For example, based on the range between the maximum and minimum values in the data to be quantized, an upper truncation threshold can be selected at predetermined intervals, while the lower truncation threshold can always be the minimum value of the data to be quantized. In some embodiments, truncation thresholds can be selected at only a few specific locations, for example, only a few predetermined percentages of the maximum value can be selected as the upper truncation thresholds.
[0061] In some embodiments, one or more corresponding quantization parameters may be calculated based on each pair of truncation thresholds, and then the calculated quantization parameters may be used to quantize the data to be quantized. Alternatively, the data to be quantized may be directly quantized using various formulas or models based on a pair of truncation thresholds, without having to individually calculate the values of each quantization parameter.
[0062] In block 506, a pair of truncation thresholds is selected from multiple pairs of truncation thresholds based on the difference between the mean of the absolute values of each group of quantized data in the multiple groups of quantized data and the mean of the absolute values of the group of data to be quantized, for quantizing the group of data to be quantized. The inventors of this application have discovered through research and extensive experiments that the difference in the mean of the absolute values of the data before and after quantization can reflect the loss of precision before and after quantization, wherein the smaller the difference in the mean of the absolute values, the smaller the loss of precision in the quantization operation. Therefore, the embodiments of the present disclosure use the difference in the mean of the absolute values of the data before and after quantization as an indicator for selecting the optimal truncation threshold, which can achieve a smaller loss of precision than the traditional KL method.
[0063] In some embodiments, the difference between the mean of the absolute values of the quantized data and the mean of the absolute values of the data to be quantized may be the difference between the means of the two absolute values. Alternatively, the difference between the mean of the absolute values of the quantized data and the mean of the absolute values of the data to be quantized may be the difference between the means of the two absolute values divided by the mean of the absolute values of the data to be quantized, and then the absolute value of the resultant difference is taken.
[0064] In some embodiments, after selecting an optimal pair of truncation thresholds, the selected pair of truncation thresholds can be used to quantize a set of data to be quantized to obtain quantized data, including: truncating the values in the set of data to be quantized that are greater than the selected truncation upper limit to the truncation upper limit, and truncating the values in the set of data to be quantized that are less than the selected truncation lower limit to the truncation lower limit; and then inputting the obtained quantized data into a neural network model for processing.
[0065] Figure 6 A flowchart of a method 600 for searching for a truncation threshold for asymmetric quantization according to an embodiment of the present disclosure is shown. The method 600 determines an optimal pair of asymmetric truncation thresholds for quantization of data based on data to be quantized.
[0066] In block 602, the mean Data_mean of the absolute values of the data to be quantized, as well as the maximum Data_max and minimum Data_min in the data to be quantized, are determined. The mean of the absolute values is the sum of the absolute values of all the data in the data to be quantized divided by the number of elements. Furthermore, the minimum mean difference needs to be initialized, for example, by initially setting the maximum value in floating-point numbers, and the search order i of the cyclic search is initialized (for example, to 0). In some embodiments, the search order i can also be initialized to half of the total number of searches, that is, starting the search from the middle, which can improve search efficiency. According to embodiments of the present disclosure, one or more rounds of threshold search processes can be set, and each round of threshold search can have the same or different total number of searches. In some embodiments, the total number of searches per round can be set between 10 and 32. Generally speaking, the greater the total number of searches, the longer the search time spent and the more accurate the truncation threshold found. However, when the total number of searches reaches a certain limit, the search effect may no longer be substantially improved.
[0067] Next, the first round of coarse-grained truncation threshold search process begins. For example, Figure 7A FIG700 shows an example of a coarse-grained search for a truncation threshold for asymmetric quantization according to an embodiment of the present disclosure. Figure 7A As shown, 10 candidate cutoff thresholds can be determined in the data to be quantified (by Figure 7A The dotted line in the figure) is used to sequentially use these 10 pairs of truncation thresholds ( Figure 7AOnly 10 upper truncation limits are shown; the lower truncation limit can always be the minimum value of the data to be quantized) to perform the quantization process, and the optimal pair of truncation thresholds is determined based on the difference in the mean absolute values of the data before and after quantization. The inventors of this application have discovered that in a neural network model, the input data is usually concentrated in small values and dispersed in large values. Therefore, directly setting the lower truncation limit to the minimum value of the data to be quantized does not cause much loss of accuracy, while also avoiding the complex process of selecting the lower truncation limit.
[0068] At block 604, a determination is made as to whether search order i is less than the predetermined total number of searches, search_grid. Specifically, a determination is made as to whether all calculations for the truncation thresholds have been completed as each pair of truncation thresholds is sequentially selected for quantization. If search order i is less than the total number of searches, a pair of truncation thresholds is determined at block 606 based on the current search order i. The upper truncation limit for this pair of truncation thresholds is, for example, Data_max - i * (Data_max - Data_min) / search_grid, while the lower truncation limit is simply the minimum value in the data to be quantized. Alternatively, the upper truncation limit for search order i can be selected as Data_max * (i + 1) / search_grid.
[0069] In box 608, the pair of truncation thresholds are used to quantize the data to be quantized to obtain the corresponding quantized data Quant_data_i. Then, in box 610, the difference between the mean of the absolute values of the quantized data Quant_data_mean_i and the mean of the absolute values of the data to be quantized Data_mean is calculated as Distance_i = abs(Quant_data_mean_i - Data_mean) / Data_mean.
[0070] In box 612, determine whether the calculated difference Distance_i is less than the current minimum difference. If so, then in box 614, set the calculated difference Distance_i to the current minimum difference, and record the truncation threshold when the difference is the smallest, and then increment the search order i in box 616. If the judgment in box 612 is no, directly increment the search order i in box 616 (i.e., i++), that is, continue to determine the difference when the next pair of truncation thresholds are met. Next, continue to loop steps 604 to 616 until the value of the search order i reaches the total number of searches, then in box 618, exit the first round of truncation threshold search process. Figure 7AAs shown, after the first round of searching, it is determined that the difference corresponding to the truncation upper limit at the dotted line 770 is the smallest. Thus, the truncation threshold search process is as follows: quantizing the data to be quantized using multiple pairs of truncation thresholds, determining the set of quantized data with the smallest mean difference in absolute value from the data to be quantized among the multiple sets of quantized data, and then selecting a pair of truncation thresholds corresponding to this set of quantized data from the multiple pairs of truncation thresholds.
[0071] Optionally, a second round of fine-grained truncation threshold search can be performed. The second round of search process can also refer to method 600, except that the second round of search is performed within a certain range around the first round optimal truncation upper limit 770 (for example, between the truncation upper limit before and the truncation upper limit after the selected truncation upper limit 770), further refining the search results of the first round. For example, in the second round of search, the interval between the truncation upper limits can be ((Data_max-Data_min)*2) / (search_grid1*search_grid2), where search_grid1 represents the total number of searches in the first round, and search_grid2 represents the total number of searches in the second round. Figure 7B A diagram 750 showing a fine-grained search for a truncation threshold for asymmetric quantization according to an embodiment of the present disclosure is shown. Figure 7B After the second round of search, the optimal upper limit of fine-grained truncation is determined to be 772, and the lower limit of truncation can be selected as the minimum value of the data to be quantized, 778. Through two rounds of search, a more accurate and precise truncation threshold can be obtained, further reducing the precision loss caused by data quantization.
[0072] Figure 8 FIG. 8 is a flow chart of a method 800 for iteratively searching for an optimal truncation threshold according to an embodiment of the present disclosure. In block 802, three pairs of truncation thresholds are determined. For example, the data to be quantized F x The maximum value Data_max and the minimum value Data_min of all data in Z max =Data_max, and Z min =Data_min, then the three pairs of truncation thresholds may be (Data_min, Data_max / 2), (Data_min, Data_max*3 / 4) and (Data_min, Data_max). In block 804, the three pairs of truncation thresholds are used to quantize the data to be quantized F. x , and obtain the quantified data Then calculate the data F x 、 The mean of the corresponding absolute values Then according to the formula Select the minimum difference diff_min. In box 806, it is necessary to determine whether the minimum difference diff_min is less than a predetermined threshold value set in advance. If not, then in box 808, based on the selected pair of truncation thresholds (the value corresponding to the minimum difference diff_min is set to the new maximum value), the three pairs of truncation thresholds are re-determined, and the above process is repeated until the minimum difference diff_min is less than the predetermined threshold value, then in box 810, the iterative process of the truncation threshold is exited. In some embodiments, in addition to the iterative stop condition that the minimum difference diff_min is less than the predetermined threshold value, other iterative stop conditions can also be set, such as the maximum number of iterations, reaching a predetermined minimum interval, etc. In addition, although Figure 8 Method 800 shows an iterative selection of the best pair of truncation thresholds. Alternatively, the iteration may be omitted and performed only once. Then, the pair of truncation thresholds corresponding to the minimum difference diff_min is directly used as the final truncation thresholds, thereby determining the quantization parameters and completing the quantization of the data.
[0073] In some embodiments, the quantization parameters when quantizing data using each pair of truncation thresholds may be determined by the following equations (1)-(4).
[0074]
[0075]
[0076]
[0077]
[0078] Where n represents the number of binary bits after quantization, o, S, and f represent quantization parameters, and ceil represents rounding up.
[0079] According to an embodiment of the present disclosure, by max Select Data_max / 2, Data_max*3 / 4 and Data_max respectively, and the quantization parameters o1, S1, f1, o2, S2, f2, o3, S3 and f3 can be obtained, thus obtaining the quantized data Correspondingly, after a pair of truncation thresholds are selected, o, S, and f corresponding to the pair of truncation thresholds are directly taken as quantization parameters of the data to be quantized.
[0080] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present disclosure.
[0081] It should be further noted that, although the various steps in the flowchart are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps may be performed in other orders. Moreover, at least a portion of the steps in the flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but may be performed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but may be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0082] Figure 9 FIG. 9 is a block diagram of an apparatus 900 for processing data according to an embodiment of the present disclosure. Figure 9 As shown, the device 900 includes a data acquisition unit 910 for quantization, a quantized data determination unit 920, and a truncation threshold selection unit 930. The data acquisition unit 910 for quantization is used to obtain a set of data for quantization for a machine learning model. The quantized data determination unit 920 is used to determine multiple sets of quantized data by quantizing a set of data for quantization using multiple pairs of truncation thresholds, wherein each pair of truncation thresholds in the multiple pairs of truncation thresholds includes a truncation upper limit and a truncation lower limit, and the truncation upper limit and the truncation lower limit in at least one pair of truncation thresholds in the multiple pairs of truncation thresholds have different absolute values. The truncation threshold selection unit 930 is used to select a pair of truncation thresholds from the multiple pairs of truncation thresholds based on the difference between the mean of the absolute values of each set of quantized data in the multiple sets of quantized data and the mean of the absolute values of the set of data for quantization.
[0083] In addition, the to-be-quantized data acquiring unit 910 , the quantized data determining unit 920 , and the truncation threshold selecting unit 930 in the apparatus 900 may also be configured to perform steps and / or actions according to various embodiments of the present disclosure.
[0084] It should be understood that the above-described device embodiments are merely illustrative, and the devices of the present disclosure may also be implemented in other ways. For example, the division of units / modules described in the above-described embodiments is merely a logical functional division, and actual implementations may employ other division methods. For example, multiple units, modules, or components may be combined or integrated into another system, or some features may be omitted or not implemented.
[0085] In addition, unless otherwise specified, the functional units / modules in the various embodiments of the present disclosure may be integrated into a single unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The aforementioned integrated units / modules may be implemented in the form of hardware or software program modules.
[0086] If the integrated unit / module is implemented in hardware, the hardware may be a digital circuit, an analog circuit, or the like. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, and the like. Unless otherwise specified, the artificial intelligence processor may be any appropriate hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC. Unless otherwise specified, the storage unit may be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), and the like.
[0087] If the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a memory, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned memory includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0088] In one embodiment, a computer-readable storage medium is disclosed, on which a computer program is stored. When the program is executed, the method according to each embodiment of the present disclosure is implemented.
[0089] In one embodiment, an artificial intelligence chip is also disclosed, which includes the above-mentioned device for processing data.
[0090] In one embodiment, a board is also disclosed, which includes a storage device, an interface device, a control device and the above-mentioned artificial intelligence chip; wherein the artificial intelligence chip is connected to the storage device, the control device and the interface device respectively; the storage device is used to store data; the interface device is used to realize data transmission between the artificial intelligence chip and an external device; the control device is used to monitor the status of the artificial intelligence chip.
[0091] Figure 10 The structural block diagram of the board 1000 according to the embodiment of the present disclosure is shown. Figure 10 In addition to the above-mentioned chips 1030-1 and 1030-2 (collectively referred to as chip 1030), the above-mentioned board 1000 may also include other supporting components, including but not limited to: a storage device 1010, an interface device 1040 and a control device 1020. The interface device 1040 can be connected to an external device 1060. The storage device 1010 is connected to the artificial intelligence chip 1030 via a bus 1050, which is used to store data. The storage device 1010 may include multiple groups of storage units 1010-1 and 1010-2. Each group of storage units is connected to the artificial intelligence chip via a bus 1050. It can be understood that each group of storage units can be DDR SDRAM (English: Double Data Rate SDRAM, double data rate synchronous dynamic random access memory).
[0092] DDR can double the speed of SDRAM without increasing the clock frequency. DDR allows data to be read out on the rising and falling edges of the clock pulse. The speed of DDR is twice that of standard SDRAM. In one embodiment, the storage device may include 4 groups of storage units. Each group of storage units may include multiple DDR4 particles (chips). In one embodiment, the artificial intelligence chip may include 4 72-bit DDR4 controllers, and 64 bits of the above 72-bit DDR4 controllers are used for data transmission and 8 bits are used for ECC verification. It can be understood that when DDR4-3200 particles are used in each group of storage units, the theoretical bandwidth of data transmission can reach 25600MB / s.
[0093] In one embodiment, each group of the memory cells includes a plurality of double data rate synchronous dynamic random access memories (DDRs) connected in parallel. DDRs can transmit data twice within one clock cycle. A controller for controlling the DDRs is provided in the chip to control data transmission and data storage in each of the memory cells.
[0094] The interface device is electrically connected to the artificial intelligence chip. The interface device is used to realize data transmission between the artificial intelligence chip and an external device (such as a server or a computer). For example, in one embodiment, the interface device can be a standard PCIE interface. For example, the data to be processed is transmitted to the chip by the server through the standard PCIE interface to realize data transfer. Preferably, when the PCIE 3.0 X 16 interface is used for transmission, the theoretical bandwidth can reach 16000MB / s. In another embodiment, the interface device can also be other interfaces. The present disclosure does not limit the specific form of the above-mentioned other interfaces. The interface unit can realize the switching function. In addition, the calculation results of the artificial intelligence chip are still transmitted back to the external device (such as a server) by the interface device.
[0095] The control device is electrically connected to the artificial intelligence chip. The control device is used to monitor the status of the artificial intelligence chip. Specifically, the artificial intelligence chip and the control device can be electrically connected via an SPI interface. The control device may include a single-chip microcomputer (MCU). For example, the artificial intelligence chip may include multiple processing chips, multiple processing cores or multiple processing circuits, which can drive multiple loads. Therefore, the artificial intelligence chip can be in different working states such as multi-load and light load. The control device can realize the regulation of the working states of multiple processing chips, multiple processing and / or multiple processing circuits in the artificial intelligence chip.
[0096] In one possible implementation, an electronic device is disclosed that includes the aforementioned artificial intelligence chip. The electronic device includes a data processing device, a robot, a computer, a printer, a scanner, a tablet computer, a smart terminal, a mobile phone, a driving recorder, a navigation system, a sensor, a camera, a server, a cloud server, a camera, a video camera, a projector, a watch, a headset, a mobile storage device, a wearable device, a vehicle, a household appliance, and / or a medical device.
[0097] The transportation vehicles include airplanes, ships and / or vehicles; the household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, and range hoods; and the medical equipment include magnetic resonance imaging (MRI), ultrasound machines and / or electrocardiographs.
[0098] In the above embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined in any way. To keep the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0099] The foregoing content can be better understood in accordance with the following terms:
[0100] A1. A method for processing data, comprising:
[0101] Obtain a set of quantified data for machine learning models;
[0102] Determining multiple sets of quantized data by respectively quantizing the set of data to be quantized using multiple pairs of truncation thresholds, wherein each pair of truncation thresholds in the multiple pairs of truncation thresholds includes an upper truncation limit and a lower truncation limit, and the upper truncation limit and the lower truncation limit in at least one pair of truncation thresholds in the multiple pairs of truncation thresholds have different absolute values; and
[0103] Based on the difference between the mean of the absolute values of each group of quantized data in the multiple groups of quantized data and the mean of the absolute values of the group of data to be quantized, a pair of truncation thresholds is selected from the multiple pairs of truncation thresholds to quantize the group of data to be quantized.
[0104] A2. The method of clause A1, wherein determining the plurality of sets of quantized data comprises:
[0105] Determine the maximum value and the minimum value of all data in the set of data to be quantified; and
[0106] Based on the maximum value and the minimum value, the pairs of cutoff thresholds are determined.
[0107] A3. The method according to clause A2, wherein determining the plurality of sets of quantized data further comprises:
[0108] determining a first truncation upper limit based on the maximum value, a predetermined total number of searches, and a current search order;
[0109] determining a first set of quantized data by quantizing the set of data to be quantized using a first pair of truncation thresholds, the first pair of truncation thresholds comprising the first truncation upper limit and a first truncation lower limit that is the same as the minimum value; and
[0110] A first difference between a mean of the absolute values of the first set of quantized data and a mean of the absolute values of the set of data to be quantized is determined.
[0111] A4. The method according to clause A3, wherein determining the plurality of sets of quantized data further comprises:
[0112] Incrementing the current search order;
[0113] determining a second truncation upper limit based on the maximum value, the predetermined total number of searches, and the current search order;
[0114] determining a second set of quantized data by quantizing the set of data to be quantized using a second pair of truncation thresholds, the second pair of truncation thresholds comprising the second upper truncation limit and a second lower truncation limit that is the same as the minimum value; and
[0115] A second difference between a mean of the absolute values of the second set of quantized data and a mean of the absolute values of the set of data to be quantized is determined.
[0116] A5. The method according to any one of clauses A1-A4, wherein selecting a pair of truncation thresholds from the plurality of pairs of truncation thresholds comprises:
[0117] Determining a set of quantized data among the multiple sets of quantized data that has the smallest difference in mean absolute value from the set of data to be quantized; and
[0118] A pair of truncation thresholds corresponding to the set of quantized data is selected from the plurality of pairs of truncation thresholds.
[0119] A6. The method according to clause A5, further comprising:
[0120] determining a cutoff search range associated with the selected pair of cutoff thresholds;
[0121] determining new pairs of cutoff thresholds within the cutoff search range;
[0122] Determine multiple new groups of quantized data by quantizing the group of data to be quantized using the new multiple pairs of cutoff thresholds respectively; and
[0123] A new pair of truncation thresholds is selected from the new multiple pairs of truncation thresholds based on a difference between a mean of the absolute values of each group of quantized data in the new multiple groups of quantized data and a mean of the absolute values of the group of data to be quantized.
[0124] A7. The method of clause A1, wherein determining the plurality of sets of quantized data comprises:
[0125] Determine the maximum and minimum values of all data in the set of data to be quantified;
[0126] determining, based on the maximum value and the minimum value, three pairs of cutoff thresholds, a first pair of the three pairs of cutoff thresholds including half of the maximum value and the minimum value, a second pair of the three pairs of cutoff thresholds including three-quarters of the maximum value and the minimum value, and a third pair of the three pairs of cutoff thresholds including the maximum value and the minimum value; and
[0127] The set of data to be quantized is quantized respectively by using three pairs of cutoff thresholds to determine three sets of quantized data.
[0128] A8. The method of clause A7, wherein selecting a pair of truncation thresholds from the plurality of pairs of truncation thresholds comprises:
[0129] Iterate the following actions until the stopping condition is met:
[0130] Selecting a pair of truncation thresholds from the three pairs of truncation thresholds;
[0131] determining whether a difference corresponding to the selected pair of cutoff thresholds is less than a predetermined threshold;
[0132] In response to the difference being less than a predetermined threshold, stopping iteratively performing the action; and
[0133] In response to the difference being greater than a predetermined threshold, three pairs of cutoff thresholds are re-determined based on the selected pair of cutoff thresholds.
[0134] A9. The method according to any one of clauses A1-A8, wherein the set of data to be quantized is a set of floating-point numbers in a neural network model, and the method further comprises:
[0135] quantizing the set of data to be quantized using the selected pair of truncation thresholds to obtain quantized data, wherein quantizing the set of data to be quantized comprises: setting values in the set of data to be quantized that are greater than the selected truncation upper limit as the truncation upper limit, and setting values in the set of data to be quantized that are less than the selected truncation lower limit as the truncation lower limit; and
[0136] The obtained quantized data is input into the neural network model for processing.
[0137] A10. A device for processing data, comprising:
[0138] A data acquisition unit for quantification, configured to acquire a set of data for quantification used in a machine learning model;
[0139] a quantized data determining unit, configured to determine multiple groups of quantized data by respectively quantizing the group of to-be-quantized data using multiple pairs of truncation thresholds, wherein each pair of truncation thresholds in the multiple pairs of truncation thresholds includes a truncation upper limit and a truncation lower limit, and the truncation upper limit and the truncation lower limit in at least one pair of the multiple pairs of truncation thresholds have different absolute values; and
[0140] A truncation threshold selection unit is used to select a pair of truncation thresholds from the multiple pairs of truncation thresholds based on the difference between the mean of the absolute values of each group of quantized data in the multiple groups of quantized data and the mean of the absolute values of the group of data to be quantized, so as to quantize the group of data to be quantized.
[0141] A11. The apparatus according to clause A10, wherein the quantized data determination unit comprises:
[0142] a maximum value and a minimum value determining unit, configured to determine the maximum value and the minimum value of all data in the set of data to be quantized; and
[0143] A plurality of pairs of truncation threshold value determining unit is configured to determine the plurality of pairs of truncation threshold values based on the maximum value and the minimum value.
[0144] A12. The apparatus according to clause A11, wherein the quantized data determination unit further comprises:
[0145] a first truncation upper limit determining unit, configured to determine a first truncation upper limit based on the maximum value, a predetermined total number of searches, and a current search order;
[0146] a first-group-of-quantized-data determining unit, configured to determine a first group of quantized data by quantizing the group of to-be-quantized data using a first pair of truncation thresholds, the first pair of truncation thresholds comprising the first truncation upper limit and a first truncation lower limit that is the same as the minimum value; and
[0147] The first difference determining unit is configured to determine a first difference between a mean of the absolute values of the first group of quantized data and a mean of the absolute values of the group of data to be quantized.
[0148] A13. The apparatus according to clause A12, wherein the quantized data determination unit further comprises:
[0149] an incrementing unit, configured to increment the current search order;
[0150] a second truncation upper limit determining unit, configured to determine a second truncation upper limit based on the maximum value, the predetermined total number of searches, and the current search order;
[0151] a second group of quantized data determining unit, configured to determine a second group of quantized data by quantizing the group of to-be-quantized data using a second pair of truncation thresholds, the second pair of truncation thresholds comprising the second truncation upper limit and a second truncation lower limit that is the same as the minimum value; and
[0152] The second difference determining unit is configured to determine a second difference between a mean of the absolute values of the second group of quantized data and a mean of the absolute values of the group of data to be quantized.
[0153] A14. The apparatus according to any one of clauses A10-A13, wherein the truncation threshold selection unit comprises:
[0154] a minimum difference determining unit, configured to determine a set of quantized data having the smallest difference in mean absolute value from the set of data to be quantized among the multiple sets of quantized data; and
[0155] The second truncation threshold selection unit is configured to select a pair of truncation thresholds corresponding to the set of quantized data from the multiple pairs of truncation thresholds.
[0156] A15. The apparatus of clause A14, further comprising:
[0157] a truncated search range determining unit, configured to determine a truncated search range associated with the selected pair of truncation thresholds;
[0158] a new multiple pairs of truncation threshold value determination unit, configured to determine multiple new pairs of truncation threshold values within the truncation search range;
[0159] a second quantized data determining unit, configured to determine multiple new groups of quantized data by respectively quantizing the group of data to be quantized using the multiple new pairs of cutoff thresholds; and
[0160] a third truncation threshold selection unit, configured to select a new pair of truncation thresholds from the new multiple pairs of truncation thresholds based on a difference between a mean of the absolute values of each group of quantized data in the new multiple groups of quantized data and a mean of the absolute values of the group of data to be quantized.
[0161] A16. The apparatus of clause A10, wherein the quantized data determining unit comprises:
[0162] A maximum value and minimum value determining unit, configured to determine the maximum value and the minimum value of all data in the set of data to be quantized;
[0163] a three-pair truncation threshold determination unit, configured to determine three pairs of truncation thresholds based on the maximum value and the minimum value, wherein a first pair of the three pairs of truncation thresholds includes half of the maximum value and the minimum value, a second pair of the three pairs of truncation thresholds includes three-quarters of the maximum value and the minimum value, and a third pair of the three pairs of truncation thresholds includes the maximum value and the minimum value; and
[0164] The three-group quantized data determination unit is configured to determine three groups of quantized data by respectively quantizing the group of data to be quantized using three pairs of truncation thresholds.
[0165] A17. The apparatus of clause A16, wherein the truncation threshold selection unit comprises:
[0166] The iteration unit is used to iteratively perform the following actions until the stopping condition is met:
[0167] Selecting a pair of truncation thresholds from the three pairs of truncation thresholds;
[0168] determining whether a difference corresponding to the selected pair of cutoff thresholds is less than a predetermined threshold;
[0169] In response to the difference being less than a predetermined threshold, stopping iteratively performing the action; and
[0170] In response to the difference being greater than a predetermined threshold, three pairs of cutoff thresholds are re-determined based on the selected pair of cutoff thresholds.
[0171] A18. The apparatus of any one of clauses A10-A17, wherein the set of data to be quantized is a set of floating-point numbers in a neural network model, and further comprising:
[0172] a data quantization unit, configured to quantize the set of data to be quantized using the selected pair of truncation thresholds to obtain quantized data, wherein quantizing the set of data to be quantized comprises: setting values in the set of data to be quantized that are greater than the selected truncation upper limit as the truncation upper limit, and setting values in the set of data to be quantized that are less than the selected truncation lower limit as the truncation lower limit; and
[0173] A data input unit is used to input the obtained quantized data into the neural network model for processing.
[0174] A19. A computer-readable storage medium, characterized in that a computer program is stored thereon, and when the program is executed, it implements the method according to any one of clauses A1 to A9.
[0175] A20. An artificial intelligence chip, characterized in that the chip includes a device for processing data according to any one of clauses A10-A18.
[0176] A21. An electronic device, characterized in that the electronic device includes an artificial intelligence chip according to clause A20.
[0177] A22. A board comprising: a memory device, an interface device, a control device, and an artificial intelligence chip according to clause A20;
[0178] Wherein, the artificial intelligence chip is connected to the storage device, the control device and the interface device;
[0179] The storage device is used to store data;
[0180] The interface device is used to realize data transmission between the artificial intelligence chip and external equipment; and
[0181] The control device is used to monitor the status of the artificial intelligence chip.
[0182] A23. The board according to clause A22, wherein:
[0183] The memory device includes: multiple groups of memory units, each group of memory units is connected to the artificial intelligence chip via a bus, and the memory units are DDR SDRAM;
[0184] The chip includes: a DDR controller for controlling data transmission and data storage of each of the storage units;
[0185] The interface device is a standard PCIE interface.
[0186] The embodiments of the present disclosure are described in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method and core ideas of the present disclosure. At the same time, changes or modifications made by those skilled in the art based on the ideas of the present disclosure, on the specific implementation methods and application scope of the present disclosure, all fall within the scope of protection of the present disclosure. In summary, the contents of this specification should not be understood as limiting the present disclosure.
Claims
1. A method for processing data, characterized in that include: Obtaining a set of to-be-quantified data for a machine learning model, wherein the input data of the machine learning model includes at least one of the following field data: image recognition, speech recognition, and natural language processing; Determining multiple sets of quantized data by respectively quantizing the set of data to be quantized using multiple pairs of truncation thresholds, wherein each pair of truncation thresholds in the multiple pairs of truncation thresholds includes an upper truncation limit and a lower truncation limit, and the upper truncation limit and the lower truncation limit in at least one pair of truncation thresholds in the multiple pairs of truncation thresholds have different absolute values; and Based on the difference between the mean of the absolute values of each group of quantized data in the multiple groups of quantized data and the mean of the absolute values of the group of data to be quantized, a pair of truncation thresholds is selected from the multiple pairs of truncation thresholds to quantize the group of data to be quantized.
2. The method according to claim 1, characterized in that Determine the multiple sets of quantified data including: Determine the maximum value and the minimum value of all data in the set of data to be quantified; and Based on the maximum value and the minimum value, the pairs of cutoff thresholds are determined.
3. The method according to claim 2, characterized in that Determine the multiple sets of quantified data also include: determining a first truncation upper limit based on the maximum value, a predetermined total number of searches, and a current search order; determining a first set of quantized data by quantizing the set of data to be quantized using a first pair of truncation thresholds, the first pair of truncation thresholds comprising the first truncation upper limit and a first truncation lower limit that is the same as the minimum value; and A first difference between a mean of the absolute values of the first set of quantized data and a mean of the absolute values of the set of data to be quantized is determined.
4. The method according to claim 3, characterized in that Determine the multiple sets of quantified data also include: Incrementing the current search order; determining a second truncation upper limit based on the maximum value, the predetermined total number of searches, and the current search order; determining a second set of quantized data by quantizing the set of data to be quantized using a second pair of truncation thresholds, the second pair of truncation thresholds comprising the second upper truncation limit and a second lower truncation limit that is the same as the minimum value; and A second difference between a mean of the absolute values of the second set of quantized data and a mean of the absolute values of the set of data to be quantized is determined.
5. The method according to any one of claims 1 to 4, characterized in that Selecting a pair of truncation thresholds from the plurality of pairs of truncation thresholds comprises: Determining a set of quantized data among the multiple sets of quantized data that has the smallest difference in mean absolute value from the set of data to be quantized; and A pair of truncation thresholds corresponding to the set of quantized data is selected from the plurality of pairs of truncation thresholds.
6. The method according to claim 5, characterized in that Also includes: determining a cutoff search range associated with the selected pair of cutoff thresholds; determining new pairs of cutoff thresholds within the cutoff search range; Determine multiple new groups of quantized data by quantizing the group of data to be quantized respectively using the new multiple pairs of cutoff thresholds; as well as A new pair of truncation thresholds is selected from the new multiple pairs of truncation thresholds based on a difference between a mean of the absolute values of each group of quantized data in the new multiple groups of quantized data and a mean of the absolute values of the group of data to be quantized.
7. The method according to claim 1, characterized in that Determine the multiple sets of quantified data including: Determine the maximum and minimum values of all data in the set of data to be quantified; determining, based on the maximum value and the minimum value, three pairs of cutoff thresholds, a first pair of the three pairs of cutoff thresholds including half of the maximum value and the minimum value, a second pair of the three pairs of cutoff thresholds including three-quarters of the maximum value and the minimum value, and a third pair of the three pairs of cutoff thresholds including the maximum value and the minimum value; and The set of data to be quantized is quantized respectively by using three pairs of cutoff thresholds to determine three sets of quantized data.
8. The method according to claim 7, characterized in that Selecting a pair of truncation thresholds from the plurality of pairs of truncation thresholds comprises: Iterate the following actions until the stopping condition is met: Selecting a pair of truncation thresholds from the three pairs of truncation thresholds; determining whether a difference corresponding to the selected pair of cutoff thresholds is less than a predetermined threshold; In response to the difference being less than a predetermined threshold, stopping iteratively performing the action; and In response to the difference being greater than a predetermined threshold, three pairs of cutoff thresholds are re-determined based on the selected pair of cutoff thresholds.
9. The method according to claim 1, characterized in that The set of data to be quantized is a set of floating-point numbers in a neural network model, and the method further includes: quantizing the set of data to be quantized using the selected pair of truncation thresholds to obtain quantized data, wherein quantizing the set of data to be quantized comprises: setting values in the set of data to be quantized that are greater than the selected truncation upper limit as the truncation upper limit, and setting values in the set of data to be quantized that are less than the selected truncation lower limit as the truncation lower limit; and The obtained quantized data is input into the neural network model for processing.
10. A device for processing data, characterized in that include: a data acquisition unit for quantification, configured to acquire a set of data for quantification used in a machine learning model, wherein the input data of the machine learning model includes at least one of the following field data: image recognition, speech recognition, and natural language processing; a quantized data determining unit, configured to determine multiple groups of quantized data by respectively quantizing the group of data to be quantized using multiple pairs of truncation thresholds, wherein each pair of truncation thresholds in the multiple pairs of truncation thresholds includes a truncation upper limit and a truncation lower limit, and the truncation upper limit and the truncation lower limit in at least one pair of truncation thresholds in the multiple pairs of truncation thresholds have different absolute values; as well as A truncation threshold selection unit is used to select a pair of truncation thresholds from the multiple pairs of truncation thresholds based on the difference between the mean of the absolute values of each group of quantized data in the multiple groups of quantized data and the mean of the absolute values of the group of data to be quantized, so as to quantize the group of data to be quantized.
11. The device according to claim 10, characterized in that The quantized data determining unit includes: a maximum value and a minimum value determining unit, configured to determine the maximum value and the minimum value of all data in the set of data to be quantized; and A plurality of pairs of truncation threshold value determining unit is configured to determine the plurality of pairs of truncation threshold values based on the maximum value and the minimum value.
12. The device according to claim 11, characterized in that The quantized data determination unit further includes: a first truncation upper limit determining unit, configured to determine a first truncation upper limit based on the maximum value, a predetermined total number of searches, and a current search order; a first-group-of-quantized-data determining unit, configured to determine a first group of quantized data by quantizing the group of to-be-quantized data using a first pair of truncation thresholds, the first pair of truncation thresholds comprising the first truncation upper limit and a first truncation lower limit that is the same as the minimum value; and The first difference determining unit is configured to determine a first difference between a mean of the absolute values of the first group of quantized data and a mean of the absolute values of the group of data to be quantized.
13. The device according to claim 12, characterized in that The quantized data determination unit further includes: an incrementing unit, configured to increment the current search order; a second truncation upper limit determining unit, configured to determine a second truncation upper limit based on the maximum value, the predetermined total number of searches, and the current search order; a second group of quantized data determining unit, configured to determine a second group of quantized data by quantizing the group of to-be-quantized data using a second pair of truncation thresholds, the second pair of truncation thresholds comprising the second truncation upper limit and a second truncation lower limit that is the same as the minimum value; and The second difference determining unit is configured to determine a second difference between a mean of the absolute values of the second group of quantized data and a mean of the absolute values of the group of data to be quantized.
14. The device according to any one of claims 10 to 13, characterized in that The truncation threshold selection unit includes: a minimum difference determining unit, configured to determine a set of quantized data having the smallest difference in mean absolute value from the set of data to be quantized among the multiple sets of quantized data; and The second truncation threshold selection unit is configured to select a pair of truncation thresholds corresponding to the set of quantized data from the multiple pairs of truncation thresholds.
15. The device according to claim 14, characterized in that Also includes: a truncated search range determining unit, configured to determine a truncated search range associated with the selected pair of truncation thresholds; a new multiple pairs of truncation threshold value determination unit, configured to determine multiple new pairs of truncation threshold values within the truncation search range; a second quantized data determining unit, configured to determine multiple new groups of quantized data by respectively quantizing the group of data to be quantized using the multiple new pairs of cutoff thresholds; as well as a third truncation threshold selection unit, configured to select a new pair of truncation thresholds from the new multiple pairs of truncation thresholds based on a difference between a mean of the absolute values of each group of quantized data in the new multiple groups of quantized data and a mean of the absolute values of the group of data to be quantized.
16. The device according to claim 10, characterized in that The quantized data determining unit includes: A maximum value and minimum value determining unit, configured to determine the maximum value and the minimum value of all data in the set of data to be quantized; a three-pair truncation threshold determination unit, configured to determine three pairs of truncation thresholds based on the maximum value and the minimum value, wherein a first pair of the three pairs of truncation thresholds includes half of the maximum value and the minimum value, a second pair of the three pairs of truncation thresholds includes three-quarters of the maximum value and the minimum value, and a third pair of the three pairs of truncation thresholds includes the maximum value and the minimum value; and The three-group quantized data determination unit is configured to determine three groups of quantized data by respectively quantizing the group of data to be quantized using three pairs of truncation thresholds.
17. The device according to claim 16, characterized in that The truncation threshold selection unit includes: The iteration unit is used to iteratively perform the following actions until the stopping condition is met: Selecting a pair of truncation thresholds from the three pairs of truncation thresholds; determining whether a difference corresponding to the selected pair of cutoff thresholds is less than a predetermined threshold; In response to the difference being less than a predetermined threshold, stopping iteratively performing the action; and In response to the difference being greater than a predetermined threshold, three pairs of cutoff thresholds are re-determined based on the selected pair of cutoff thresholds.
18. The device according to claim 10, characterized in that The set of data to be quantized is a set of floating-point numbers in a neural network model, and the device further comprises: a data quantization unit, configured to quantize the set of data to be quantized using the selected pair of truncation thresholds to obtain quantized data, wherein quantizing the set of data to be quantized comprises: setting values in the set of data to be quantized that are greater than the selected truncation upper limit as the truncation upper limit, and setting values in the set of data to be quantized that are less than the selected truncation lower limit as the truncation lower limit; and A data input unit is used to input the obtained quantized data into the neural network model for processing.
19. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the program is executed, the method according to any one of claims 1 to 9 is implemented.
20. An artificial intelligence chip, characterized in that: The chip comprises the device for processing data according to any one of claims 10-18.
21. An electronic device, characterized in that: The electronic device includes the artificial intelligence chip according to claim 20.
22. A board, characterized in that: The board includes: a storage device, an interface device, a control device, and an artificial intelligence chip according to claim 20; Wherein, the artificial intelligence chip is connected to the storage device, the control device and the interface device; The storage device is used to store data; The interface device is used to realize data transmission between the artificial intelligence chip and external equipment; and The control device is used to monitor the status of the artificial intelligence chip.
23. The board according to claim 22, wherein: The memory device includes: multiple groups of memory units, each group of memory units is connected to the artificial intelligence chip via a bus, and the memory units are DDR SDRAM; The chip includes: a DDR controller for controlling data transmission and data storage of each of the storage units; The interface device is a standard PCIE interface.
Citation Information
Patent Citations
Ultra-high-speed static gesture recognition method based on depth model optimization
CN110096968A
Computing device and method
CN110163353A