Method, apparatus, and related product for processing data

By using multiple point locations to quantize the data to be quantized in the machine learning model and selecting the optimal point location, the problem of high computational resource consumption and time overhead in large-scale data processing is solved, and more efficient data processing is achieved.

CN112446460BActive Publication Date: 2025-12-16SHANGHAI CAMBRICON INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910804627.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-08-28
Publication Date
2025-12-16
Estimated Expiration
2040-01-10

AI Technical Summary

Technical Problem

As the complexity and accuracy of artificial intelligence algorithms increase, machine learning models become larger, leading to increased data processing volume, higher computational and time costs, and lower processing efficiency.

Method used

By quantizing the data to be quantized at multiple points, multiple sets of quantized data are determined. The optimal point position is selected for quantization based on the difference between the quantized data and the data to be quantized, thereby reducing the loss of accuracy.

Benefits of technology

It improves the accuracy of data quantification, reduces the consumption of computing resources and processing time, and enhances processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112446460B_ABST
    Figure CN112446460B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method and apparatus for processing data and related products. Embodiments of the present disclosure relate to a board card comprising a memory device, an interface device, a control device, and an artificial intelligence chip; wherein the artificial intelligence chip is connected with the memory device, the control device, and the interface device respectively; the memory device is configured to store data; the interface device is configured to realize data transmission between the artificial intelligence chip and an external device; and the control device is configured to monitor a state of the artificial intelligence chip. The board card can be used to perform artificial intelligence operation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure generally relate to the field of computer technology, and more particularly to a method, an apparatus and a related product for processing data. BACKGROUND

[0002] With the continuous development of artificial intelligence technology, its application fields are becoming more and more extensive, and it has been well applied in the fields of image recognition, speech recognition, natural language processing, etc. However, with the increase of complexity and accuracy of artificial intelligence algorithms, machine learning models are becoming larger and larger, and the amount of data to be processed is also increasing. When a large amount of data processing is performed, a large amount of computation and time is required, and the processing efficiency is low. SUMMARY

[0003] In view of this, embodiments of the present disclosure provide a method, an apparatus and a related product for processing data.

[0004] In a first aspect of the present disclosure, a method for processing data is provided. The method comprises: obtaining a set of to-be-quantized data for a machine learning model; determining a plurality of sets of quantized data by quantizing the set of to-be-quantized data using a plurality of point positions respectively, each of the plurality of point positions specifying a position of a decimal point in the plurality of sets of quantized data; and selecting a point position from the plurality of point positions for quantizing the set of to-be-quantized data based on a difference between each of the plurality of sets of quantized data and the set of to-be-quantized data.

[0005] In a second aspect of the present disclosure, an apparatus for processing data is provided. The apparatus comprises: an obtaining unit configured to obtain a set of to-be-quantized data for a machine learning model; a determining unit configured to determine a plurality of sets of quantized data by quantizing the set of to-be-quantized data using a plurality of point positions respectively, each of the plurality of point positions specifying a position of a decimal point in the plurality of sets of quantized data; and a selecting unit configured to select a point position from the plurality of point positions for quantizing the set of to-be-quantized data based on a difference between each of the plurality of sets of quantized data and the set of to-be-quantized data.

[0006] In a third aspect of the present disclosure, a computer-readable storage medium is provided, which stores a computer program, the program being executed to implement the method according to various embodiments of the present disclosure.

[0007] In a fourth aspect of the present disclosure, an artificial intelligence chip is provided, which comprises the apparatus for processing data according to various embodiments of the present disclosure.

[0008] In a fifth aspect of the present disclosure, an electronic device is provided, which comprises the artificial intelligence chip according to various embodiments of the present disclosure.

[0009] In a sixth aspect of the present disclosure, a board card is provided, comprising: a memory device, an interface device and a control device, and an artificial intelligence chip according to various embodiments of the present disclosure. The artificial intelligence chip is connected to the memory device, the control device and the interface device; the memory device is configured to store data; the interface device is configured to realize data transmission between the artificial intelligence chip and an external device; and the control device is configured to monitor the state of the artificial intelligence chip.

[0010] The technical features in the claims can deduce the beneficial effects of solving the technical problems in the background art. Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS

[0011] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and serve to explain the principles of the present disclosure.

[0012] Figure 1 A schematic diagram of a processing system for processing data according to embodiments of the present disclosure is shown;

[0013] Figure 2 A schematic diagram of an example architecture of a neural network according to embodiments of the present disclosure is shown;

[0014] Figure 3 A schematic diagram of a process for quantizing data according to embodiments of the present disclosure is shown;

[0015] Figure 4 A schematic diagram of a quantization process according to embodiments of the present disclosure is shown;

[0016] Figure 5 A schematic diagram of a process for processing data according to embodiments of the present disclosure is shown;

[0017] Figure 6 A flowchart of a method for processing data according to embodiments of the present disclosure is shown;

[0018] Figure 7 A schematic diagram of different quantization schemes based on different point positions according to embodiments of the present disclosure is shown;

[0019] Figure 8 A flowchart of a method for processing data according to embodiments of the present disclosure is shown;

[0020] Figure 9 A block diagram of an apparatus for processing data according to embodiments of the present disclosure is shown; and

[0021] Figure 10 A structural block diagram of a board card according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0022] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by a person skilled in the art without creative labor fall within the scope of protection of the present disclosure.

[0023] It should be understood that the terms “first”, “second”, “third”, and “fourth” and the like in the claims, specification, and drawings of the present disclosure are used to distinguish different objects, and are not used to describe a particular order. The terms “include” and “contain” used in the specification and claims of the present disclosure indicate the presence of described features, whole, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, whole, steps, operations, elements, components, and / or sets thereof.

[0024] It should also be understood that the terms used in the specification of the present disclosure herein are only for the purpose of describing specific embodiments, and are not intended to limit the present disclosure. As used in the specification and claims of the present disclosure, the singular forms “a”, “an” and “the” are intended to include plural forms, unless the context clearly indicates otherwise. It should be further understood that the term “and / or” used in the specification and claims of the present disclosure refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.

[0025] As used in the specification and claims, the term “if” can be interpreted as “when” or “upon” or “in response to a determination” or “in response to detecting” depending on the context. Similarly, the phrase “if determined” or “if detected [described condition or event]” can be interpreted to mean “upon determining” or “in response to determining” or “upon detecting [described condition or event]” or “in response to detecting [described condition or event]” depending on the context.

[0026] Generally, when quantizing data, scaling needs to be performed on the data to be quantized. For example, in a case where it has been determined how many bits of binary are used to represent the quantized data, a point position can be used to describe the position of the decimal point. At this time, the decimal point can divide the quantized data into an integer part and a decimal part. Therefore, a suitable point position needs to be found to quantize the data, so that the loss of data quantization is minimal or small.

[0027] Traditionally, techniques have been proposed to determine the point position based on the value range of a set of data to be quantized. However, since the data to be quantized can not always be uniformly distributed, the point position determined based on the value range can not always enable accurate quantization to be performed, but can result in a large precision loss for some data to be quantized in the set of quantized data.

[0028] To this end, embodiments of the present disclosure propose a new scheme to determine the point position used in the quantization process. The scheme can achieve a smaller quantization precision loss than traditional techniques. According to embodiments of the present disclosure, after obtaining a set of data to be quantized for a machine learning model, a plurality of sets of quantized data are determined by quantizing the set of data to be quantized using a plurality of point positions, each of the plurality of point positions specifying the position of the decimal point in the plurality of sets of quantized data. Then, based on the difference between each set of quantized data in the plurality of sets of quantized data and the set of data to be quantized as an evaluation assignment, a point position is selected from the plurality of point positions for quantizing the set of data to be quantized. In this way, a more suitable point position can be found.

[0029] The basic principles and several example implementations of the present disclosure are described below with reference to Figures 1 to 10 It should be understood that these example embodiments are given only so that those skilled in the art can better understand and implement embodiments of the present disclosure, and do not limit the scope of the present disclosure in any way.

[0030] Figure 1 A schematic diagram of a processing system 100 for processing data according to embodiments of the present disclosure is shown. As Figure 1 shown, the processing system 100 includes a plurality of processors 101-1, 101-2, 101-3 (collectively referred to as processors 101) for executing sequences of instructions and a memory 102 for storing data, which can include random access memory (RAM) and a register file. The plurality of processors 101 in the processing system 100 can both share part of the storage space, such as part of the RAM storage space and the register file, and also have their own storage space at the same time.

[0031] It should be understood that various methods according to embodiments of the present disclosure can be applied to any one of the plurality of processors (multi-core) of the processing system 100 (for example, an artificial intelligence chip). The processor can be a general-purpose processor, such as a CPU (Central Processing Unit), or an artificial intelligence processor (IPU) for performing artificial intelligence operations. The artificial intelligence operations can include machine learning operations, brain-like operations, and the like. The machine learning operations can include neural network operations, k-means operations, support vector machine operations, and the like. The artificial intelligence processor can include one or a combination of, for example, a GPU (Graphics Processing Unit), a NPU (Neural-Network Processing Unit), a DSP (Digital Signal Process), and an FPGA (Field-Programmable Gate Array) chip. The present disclosure does not limit the specific type of processor. In addition, the types of the plurality of processors in the processing system 100 can be the same or different, and the present disclosure does not limit this.

[0032] In one possible implementation, the processor mentioned in the present disclosure can include a plurality of processing units, each of which can independently run various tasks assigned to it, such as convolution operation tasks, pooling tasks, or fully connected tasks, and the like. The present disclosure does not limit the processing units and the tasks run by the processing units.

[0033] Figure 2A schematic diagram of an example architecture of a neural network 200 according to an embodiment of the present disclosure is shown. A neural network (NN) is a mathematical model that mimics the structure and function of biological neural networks. Neural networks perform computations through a large number of interconnected neurons. Therefore, a neural network is a computational model consisting of a large number of interconnected nodes (or "neurons"). Each node represents a specific output function called an activation function. Each connection between two neurons represents a weighted value of the signal passing through that connection, called a weight, which is equivalent to the memory of the neural network. The output of the neural network varies depending on the connections between neurons and the weights and activation functions. In a neural network, a neuron is the basic unit. It receives a certain number of inputs and a bias, which is multiplied by a weight when a signal (value) arrives. A connection connects a neuron to another neuron in another layer or the same layer, and the connection is accompanied by an associated weight. Additionally, the bias is an extra input to the neuron; it is always 1 and has its own connection weight. This ensures that the neuron is activated even if all inputs are empty (all 0).

[0034] In applications, if a non-linear function isn't applied to the neurons in a neural network, the network is merely a linear function and no more powerful than a single neuron. If we want the output of a neural network to be between 0 and 1—for example, in cat / dog identification—outputs close to 0 can be interpreted as cats, and outputs close to 1 as dogs. To achieve this, activation functions, such as the sigmoid activation function, are introduced into the neural network. Regarding this activation function, all we need to know is that its return value is a number between 0 and 1. Therefore, activation functions introduce non-linearity into the neural network, narrowing the results of the network's computation to a smaller range. In reality, how the activation function is expressed is not important; what matters is parameterizing a non-linear function with weights, which can be changed to alter the non-linear function.

[0035] like Figure 2 The diagram shown is a schematic representation of the structure of neural network 200. Figure 2 The neural network shown includes three layers: an input layer 210, a hidden layer 220, and an output layer 230. Figure 2 The hidden layer 220 shown has three layers; however, it can also include more or fewer layers. The neurons in the input layer 210 are called input neurons. As the first layer in the neural network, the input layer requires input signals (values) and passes them to the next layer. It does not perform any operations on the input signals (values) and has no associated weights or biases. Figure 2In the illustrated neural network, 4 input signals (values) can be received.

[0036] The hidden layer 220 is used to apply different transformations to the input data by neurons (nodes). One hidden layer is a collection of neurons (Representation) arranged vertically. In the illustrated neural network, the hidden layer 220 has 4 neurons (nodes). Figure 2 In the illustrated neural network, there are 3 hidden layers. The first hidden layer has 4 neurons (nodes), the second layer has 6 neurons, and the third layer has 3 neurons. Finally, the hidden layers pass the values to the output layer. Figure 2 The illustrated neural network 200 has full connections between each neuron in the 3 hidden layers, i.e., each neuron in the 3 hidden layers is connected to every neuron in the next layer. It should be noted that not every hidden layer of a neural network is fully connected.

[0037] The neurons of the output layer 230 are called output neurons. The output layer receives the output from the last hidden layer. Through the output layer 230, the desired value and the desired range can be determined. In the illustrated neural network, the output layer has 3 neurons, i.e., there are 3 output signals (values). Figure 2

[0038] In practical applications, the role of a neural network is to train a large amount of sample data (containing input and output) in advance, and after training, use the neural network to obtain an accurate output for the input of the real environment in the future.

[0039] Before starting to discuss the training of a neural network, a loss function needs to be defined. The loss function is a function that indicates how well a neural network performs on a certain task. The most direct way is to, during the training process, pass each sample data through the neural network to obtain a number, and then subtract the actual number that is expected to be obtained from the number to square it. This calculated distance between the predicted value and the actual value is the distance between the predicted value and the actual value, and the training of the neural network is to reduce the value of the loss function.

[0040] When starting to train a neural network, the weights are randomly initialized. Obviously, the initialized neural network will not provide a good result. During the training process, it is assumed that a very poor neural network is started, and through training, a network with high accuracy can be obtained. At the same time, it is also desired that the function value of the loss function becomes very small at the end of the training.

[0041] ​The training process of the neural network is divided into two stages. The first stage is the forward processing of the signal, from the input layer 210 through the hidden layer 220, and finally to the output layer 230. The second stage is the backward propagation of the gradient, from the output layer 230 to the hidden layer 220, and finally to the input layer 210, according to the gradient to adjust the weight and bias of each layer in the neural network in turn.

[0042] In the process of forward processing, the input value is input to the input layer 210 of the neural network, and the output called the predicted value is obtained from the output layer 230 of the neural network. When the input value is provided to the input layer 210 of the neural network, it does not perform any operation. In the hidden layer, the second hidden layer obtains the predicted intermediate result value from the first hidden layer and performs the calculation operation and the activation operation, and then passes the obtained predicted intermediate result value to the next hidden layer. The same operation is performed in the subsequent layers, and finally the output value is obtained at the output layer 230 of the neural network.

[0043] After the forward processing, an output value called the predicted value is obtained. In order to calculate the error, a loss function is used to compare the predicted value with the actual output value to obtain the corresponding error value. Backpropagation uses the chain rule of differential calculus, in which the derivative of the error value corresponding to the last layer weight of the neural network is first calculated. These derivatives are called gradients, and then these gradients are used to calculate the gradients of the second-to-last layer in the neural network. This process is repeated until the gradients of each weight in the neural network are obtained. Finally, the corresponding gradients are subtracted from the weights, thereby updating the weights once to reduce the error value.

[0044] In addition, for the neural network, fine-tuning is to load the trained neural network, and the fine-tuning process is the same as the training process, which is divided into two stages. The first stage is the forward processing of the signal, and the second stage is the backward propagation of the gradient to update the weight of the trained neural network. The difference between training and fine-tuning is that training is to randomly process the initialized neural network, and the neural network is trained from scratch, while fine-tuning is not to train from scratch.

[0045] In the process of training or fine-tuning of the neural network, the weight value in the neural network is updated by gradient once per forward processing of the signal and corresponding backward propagation of the error, which is called an iteration. In order to obtain a neural network with expected accuracy, a very large sample data set is required in the training process. In this case, it is impossible to input the sample data set into the computer at one time. Therefore, in order to solve this problem, the sample data set needs to be divided into multiple blocks, each block is transmitted to the computer, and the weight value of the neural network is updated once after each block of data set is forward processed. When a complete sample data set is forward processed once in the neural network and the corresponding weight value is returned once, this process is called an epoch. In practice, it is not enough to transmit a complete data set in the neural network, and the complete data set needs to be transmitted multiple times in the same neural network, that is, multiple epochs are required to finally obtain a neural network with expected accuracy.

[0046] In the process of training or fine-tuning of the neural network, it is generally desired to be faster and more accurate. The data of the neural network is represented by high-precision data format, such as floating-point numbers, so in the training or fine-tuning process, the data involved are all high-precision data format, and then the trained neural network is quantized. Taking the weight value of the entire neural network as the quantization object and the quantized weight value as an 8-bit fixed-point number as an example, since there are usually millions of connections in a neural network, almost all the space is occupied by the weight values of the neuron connections. Moreover, these weight values are different floating-point numbers. The weight values of each layer tend to be normally distributed in a certain interval, for example, (-3.0, 3.0). The maximum and minimum values of the weight values of each layer in the neural network are saved, and each floating-point value is represented by an 8-bit fixed-point number. Among them, 256 quantization intervals are linearly divided in the maximum and minimum value range, and each quantization interval is represented by an 8-bit fixed-point number. For example: in the interval (-3.0, 3.0), byte 0 represents -3.0, and byte 255 represents 3.0. In this way, byte 128 represents 0.

[0047] For data represented by high-precision data format, taking floating-point numbers as an example, according to the computer architecture, the operation representation rules based on floating-point numbers and fixed-point numbers, for fixed-point operations and floating-point operations of the same length, the floating-point operation mode is more complex, and more logic devices are required to form a floating-point operation unit. In this way, the volume of the floating-point operation unit is larger than that of the fixed-point operation unit. Moreover, the floating-point operation unit consumes more resources for processing, so that the power consumption difference between fixed-point operation and floating-point operation is usually of orders of magnitude. In short, the chip area and power consumption occupied by the floating-point operation unit are much larger than those of the fixed-point operation unit.

[0048] Figure 3 A schematic diagram of a process 300 for quantizing data according to an embodiment of the disclosure is shown. Referring to Figure 3 , the input data 310 is an unquantized floating point number, for example, a 32-bit floating point number. If the input data 310 is directly input into a neural network model 340 for processing, it would consume more computing resources and the processing speed would be slower. Therefore, at block 320, the input data can be quantized to obtain quantized data 330 (e.g., 8-bit integer). If the quantized data 330 is input into the neural network model 340 for processing, since the 8-bit integer calculation is faster, the neural network model 340 would complete the processing of the input data faster and generate a corresponding output result 350.

[0049] During the quantization process from the unquantized input data 310 to the quantized data 330, some precision loss would be caused to some extent, and the precision loss would directly affect the accuracy of the output result 350. Therefore, during the quantization process of the input data 330, it is necessary to ensure that the precision loss of the quantization process is minimal or as small as possible.

[0050] In the following, the quantization process will be described in detail with reference to Figure 4 . Figure 4 A schematic diagram 400 of the quantization process according to an embodiment of the disclosure is shown. As Figure 4 A simple quantization process is shown, which can map each of a set of to-be-quantized data to a set of quantized data. At this time, the range of the set of to-be-quantized data is -|max| to |max|, and the range of the generated set of quantized data is -(2 n-1 -1) to +(2 n-1 -1). Here, n represents the predefined data width 410, i.e., how many bits are used to represent the quantized data. Continuing the example above, when 8 bits are used to represent the quantized data, if the first bit represents the sign bit, the range of the quantized data can be -127 to +127.

[0051] It will be understood that, in order to more accurately represent the quantized data, the range of the quantized data can also be -(2 Figure 4The n-bit data structure is shown to represent quantized data. As shown, n bits can be employed to represent the quantized data, where the left-most bit can represent a sign bit 430 indicating whether the data is positive or negative. A point 420 can be set. The point 420 in this case represents the boundary between an integer portion 432 and a fractional portion 434 of the quantized data. To the left of the point is a positive power of two, and to the right of the point is a negative power of two. In the context of the present disclosure, the position of the point can be represented by a point position. It will be appreciated that, given a predetermined data width 410, by adjusting the point position (represented as an integer) to move the position of the point 420, the range and precision represented by the n-bit data structure will change.

[0052] For example, assume that the point 420 is located after the right-most bit, then the sign bit 430 comprises one bit, the integer portion 432 comprises n-1 bits, and the fractional portion 434 comprises zero bits. Thus, the range represented by the n-bit data structure will be from -(2 n-1 -1) to +(2 n-1 -1), and the precision is an integer. For another example, assume that the point 420 is located before the right-most bit, then the sign bit 430 comprises one bit, the integer portion 432 comprises n-2 bits, and the fractional portion 434 comprises one bit. Thus, the range represented by the n-bit data structure will be from -(2 n-2 -1) to +(2 n-2 -1), and the precision is a decimal fraction "0.5". In this case, the point position needs to be determined so that the range and precision represented by the n-bit data structure better match the range and precision of the data to be quantized.

[0053] According to an embodiment of the present disclosure, a method for processing data is provided. Referring first to Figure 5 Embodiments of the present disclosure are generally described. Figure 5 A schematic diagram 500 of a process for processing data according to an embodiment of the present disclosure is shown. According to an embodiment of the present disclosure, a plurality of quantization processes can be performed based on a plurality of point positions 520. For example, for data to be quantized 510, a respective quantization process can be performed based on each of the plurality of point positions 520, respectively, to obtain a plurality of sets of quantized data 530. Then, each of the plurality of sets of quantized data 530 can be compared with the data to be quantized 510, respectively, to determine a difference therebetween. By selecting a point position 550 corresponding to a smallest difference from the plurality of differences 540 obtained, the point position 550 that is most suitable for the data to be quantized 510 can be determined. With embodiments of the present disclosure, quantized data can be represented with higher precision.

[0054] In the following, reference will be made to Figure 6 Further details regarding data processing are described in detail.Figure 6 A flowchart of a method 600 for processing data according to an embodiment of the present disclosure is shown. As shown, at block 610, a set of data to be quantized for a machine learning model is obtained. For example, with reference to the above description of the neural network model 340, the set of data to be quantized can be the input data 310. Figure 6 Figure 3 The set of data to be quantized obtained here can be the input data 310, and the input data 310 is quantized to speed up the processing of the neural network model 340. In addition, some parameters (such as weights, etc.) of the neural network model itself can also be quantized, and by quantizing the parameters of the neural network, the size of the neural network model can be reduced. In some embodiments, each data to be quantized in the set of data to be quantized can be a 32-bit floating point number. Alternatively, the data to be quantized can also be a floating point number of other bits, or other data types.

[0055] At block 620, a plurality of sets of quantized data can be determined by quantizing the set of data to be quantized using a plurality of point positions respectively. Here, each point position in the plurality of point positions specifies the position of the decimal point in the plurality of sets of quantized data. According to an embodiment of the present disclosure, each point position in the plurality of point positions is represented by an integer. One point position can be determined first, and then expansion can be performed for the point position to obtain more point positions.

[0056] According to an embodiment of the present disclosure, one point position in the plurality of point positions can be obtained based on a range associated with the set of data to be quantized. In the following, the point position will be represented by an integer S for convenience of description, and the value of the integer S represents the number of bits included in the integer part 432. For example, S = 3 represents that the integer part 432 includes 3 bits. Assuming that the original data to be quantized is represented as F x The data to be quantized represented by an n-bit data structure is I x Then, there will be the following formula 1.

[0057] F x ≈ I x × 2 s Formula 1

[0058] At this time, the quantized data can be represented as the following formula 2:

[0059]

[0060] In formula 2, round represents the floor operation. Thus, the point position S here can be represented as the following formula 3:

[0061]

[0062] ​In Formula 3, p represents the maximum absolute value in a set of data to be quantified. Alternatively and / or additionally, p may represent a range determined in other ways. In Formula 3, ceil represents the round-up budget. A point position (e.g., S0) among multiple point positions can be determined based on Formula 3 above. According to embodiments of this disclosure, other point positions among multiple point positions can be determined based on integers adjacent to the obtained point position S0. Here, "adjacent" integers refer to integers having values ​​adjacent to integer S0. According to embodiments of this disclosure, an increment operation can be performed on the integer representing the point position to determine a point position among the other point positions. According to embodiments of this disclosure, a decrement operation can also be performed on the integer representing the point position to determine a point position among the other point positions. For example, assuming the value of S0 is 3, an increment operation can obtain another adjacent integer 3+1=4, and a decrement operation can obtain another adjacent integer 3-1=2.

[0063] By utilizing embodiments of this disclosure, and considering multiple point locations near a given point location, the quantization effects of quantization processes based on multiple point locations can be compared, and the most suitable point location for a set of data to be quantized can be selected from among the multiple point locations. Compared to a technical solution that determines a point location solely based on Formula 3, embodiments of this disclosure can improve the accuracy of the quantization process.

[0064] See below. Figure 7 Further details of embodiments according to this disclosure are described in detail. Figure 7 A schematic diagram 700 illustrates different quantization schemes based on different point locations according to embodiments of the present disclosure. For example... Figure 7 As shown, in the first quantization scheme, the decimal point is located at... Figure 7 The first position 710 is shown. This first point position 712 can be determined according to Formula 3 described above. Then, by performing a decrement operation on the first point position 712, the second point position 722 can be determined, at which point the decimal point is moved to the left to the second position 720.

[0065] Will understand, although Figure 7 This illustration only shows a single decrement operation performed on the first point position 712 to determine a point position. According to embodiments of this disclosure, increment and / or decrement operations can also be performed on the first point position 712 to determine more point positions. According to embodiments of this disclosure, an even greater number of increment and decrement operations can be performed to determine more point positions. For example, different point positions can be determined separately: S1 = S0 + 1, S2 = S0 - 1, S3 = S0 + 2, S4 = S0 - 2, and so on.

[0066] In a case where a plurality of point positions S0, S1, S2, S3, S4, etc. have been determined, a difference between each of the plurality of point positions S0, S1, S2, S3, S4, etc. and the data F to be quantized can be determined based on Formula 2 described above. x The quantization operation is performed. Specifically, in Formula 2, F x represents the data to be quantized, and the corresponding quantized data F , etc.

[0067] It will be understood that the F x represents only one data to be quantized in a group of data to be quantized, and there can be a plurality of (e.g., m) data to be quantized in the group of quantized data. At this time, each data to be quantized can be processed based on the process described above to obtain the corresponding quantized data. At this time, based on each point position, a corresponding group of quantized data (m) can be obtained.

[0068] At block 530, one point position can be selected from the plurality of point positions based on a difference between each of the plurality of groups of quantized data and the group of data to be quantized for quantizing the group of data to be quantized. The inventors of the present application have found that the difference between the data before and after quantization can reflect the loss of accuracy before and after quantization, and the smaller the difference, the smaller the loss of accuracy of the quantization operation, according to research and a large number of experiments. Therefore, the embodiments of the present disclosure use the difference between the data before and after quantization as an index for selecting the best point position, and can achieve a smaller loss of accuracy than the conventional scheme.

[0069] Continuing the example above, the difference can be determined based on a comparison of the quantized data and the data F x to be quantized. According to embodiments of the present disclosure, the difference can be determined based on a plurality of ways. For example, Formula 4 or Formula 5 shown below can be used to determine the difference between the data before and after quantization for each data.

[0070]

[0071]

[0072] In Formulas 4 and 5, Diff represents the difference for one data to be quantized, F x represents the data to be quantized, represents the operation of taking absolute value. For example, for each point position, the absolute value of the difference between the data before and after the quantization operation can be determined for each of the set of data to be quantized. For a set of m data to be quantized, m differences can be obtained. Then, the difference for the point position can be determined based on the obtained m differences.

[0073] For example, for the point position S0, m differences between the data before and after the quantization operation performed using the point position S0 can be determined. Then, for example, the m differences are summed (alternatively and / or additionally, other operations can be employed), the difference Diff0 for the point position S0 can be obtained. Similarly, the differences Diff1, Diff2, Diff3, Diff4, etc. for the other point positions S1, S2, S3, S4, etc. can also be obtained respectively.

[0074] According to embodiments of the present disclosure, the minimum difference can be selected from the plurality of differences, and the point position corresponding to the minimum difference can be selected from the plurality of point positions for performing the quantization operation. For example, assuming that the difference Diff1 determined based on the point position S1 is the minimum difference, the point position S1 can be selected for the subsequent quantization processing.

[0075] According to embodiments of the present disclosure, the plurality of differences for the plurality of point positions can also be determined based on the mean value for simplification. For example, the mean value F mean of the set of data to be quantized (for example, it can be referred to as the original mean value) can be calculated. The mean value herein can be determined based on the average of each of the set of data to be quantized, for example. Similarly, the mean value F of the set of quantized data (for example, it can be referred to as the quantized mean value) can be calculated. Further, one of the plurality of differences can be determined based on the quantized mean value and the original mean value. Specifically, the difference for each of the plurality of point positions can be determined based on the following equation 6 or equation 7.

[0076]

[0077]

[0078] In the equation 6 and the equation 7, F mean represents the mean value of the set of data to be quantized, F meanThe mean value representing a set of quantized data is shown. Specifically, the difference Diff0 for the point position S0 can be obtained based on the above formula 6 or formula 7. Similarly, the differences Diff1, Diff2, Diff3, Diff4, etc. for other point positions S1, S2, S3, S4, etc. can also be obtained respectively. Then, one of the point positions S0, S1, S2, S3, S4 corresponding to the minimum difference can be selected for performing the quantization operation. In this way of using the mean value, it is not necessary to determine the difference before and after quantization for each of the data to be quantized, so that the data processing efficiency can be improved and the speed of determining the point position can be accelerated.

[0079] The above has described a plurality of formulas that can be involved during processing. In the following, the formulas will be referred to as Figure 8 The specific flow of data processing is described. Figure 8 A flowchart of a method 800 for processing data according to an embodiment of the present disclosure is shown. At block 810, a first point position (e.g. S0) can be obtained based on a range associated with a set of data to be quantized. Here, the point position S0 can be obtained based on formula 3. At block 820, an increment / decrement operation can be performed for the first point position to obtain a second point position (e.g. S1 = S0 + 1).

[0080] At block 830, a first set of quantized data and a second set of quantized data can be determined based on the first point position S0 and the second point position S1 respectively. Specifically, for each of the data to be quantized, the corresponding quantized data can be obtained based on formula 2. At block 840, a first difference Diff0 and a second difference Diff1 between the first set of quantized data and the second set of quantized data and the set of data to be quantized can be determined respectively. For example, the first difference Diff0 and the second difference Diff1 can be determined based on any one of formulas 4 to 7. At block 850, the first difference Diff0 and the second difference Diff1 can be compared, and if the first difference is less than the second difference, the method 800 proceeds to block 852 to select the first point position. If the first difference is greater than (or equal to) the second difference, the method 800 proceeds to block 854 to select the second point position. As shown by the dashed block 860, the selected point position can be used for performing the quantization processing for the data to be quantized.

[0081] It will be appreciated that the quantization process can be performed on the initial set of to-be-quantized data at block 860. In the case that the distribution of subsequent to-be-quantized data is similar to that of the initial set of to-be-quantized data, the quantization process can also be performed on other sets of subsequent to-be-quantized data. In the following, the specific application environment of the neural network model will be described. According to an embodiment of the present disclosure, the set of to-be-quantized data can include a set of floating-point numbers in the neural network model. The quantization operation can be performed using the selected point position, so as to convert the floating-point numbers with higher complexity to lower complexity. According to an embodiment of the present disclosure, the set of to-be-quantized data is quantized using the selected point position to obtain a set of quantized data. Specifically, the set of to-be-quantized data is mapped to the set of quantized data based on the selected point position, and the position of the decimal point in the set of quantized data is determined by the selected point position. Assuming that the selected point position is 4, 4 bits can be used to represent the integer part of the quantized data in the quantization process. Then, the obtained set of quantized data can be input to the neural network model for processing.

[0082] According to an embodiment of the present disclosure, the selected point position can also be used to quantize other to-be-quantized data subsequently. Specifically, another set of to-be-quantized data including a set of floating-point numbers in the neural network model can be obtained. The other set of to-be-quantized data can be quantized using the selected point position to obtain another set of quantized data, and the obtained another set of quantized data can be input to the neural network model for processing.

[0083] It should be noted that for the foregoing method embodiments, in order to simply describe, they are all expressed as a combination of a series of actions, but those skilled in the art should know that the present disclosure is not limited by the order of the described actions, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present disclosure.

[0084] It should be further noted that although the steps in the flowchart are displayed in sequence according to the arrows, these steps are not necessarily executed in sequence according to the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed in rotation or alternation with other steps or sub-steps or stages of other steps.

[0085] Figure 9 A block diagram of an apparatus 900 for processing data according to an embodiment of the present disclosure is shown. As shown, the apparatus 900 includes an obtaining unit 910, a determining unit 920, and a selecting unit 930. The obtaining unit 910 is configured to obtain a set of to-be-quantized data for a machine learning model. The determining unit 920 is configured to determine a plurality of sets of quantized data by respectively quantizing the set of to-be-quantized data using a plurality of point positions, each of the plurality of point positions specifying a position of a decimal point in the plurality of sets of quantized data. The selecting unit 930 is configured to select, from the plurality of point positions, a point position for quantizing the set of to-be-quantized data based on a difference between each of the plurality of sets of quantized data and the set of to-be-quantized data. Figure 9

[0086] In addition, the obtaining unit 910, the determining unit 920, and the selecting unit 930 in the apparatus 900 can also be configured to perform the steps and / or actions according to various embodiments of the present disclosure.

[0087] It should be understood that the above-described apparatus embodiments are merely illustrative, and the apparatus of the present disclosure can also be implemented in other manners. For example, the division of units / modules in the above-described embodiments is merely a logical function division, and the actual implementation can be in another division manner. For example, a plurality of units / modules or components can be combined, or can be integrated into another system, or some features can be omitted or not executed.

[0088] In addition, each functional unit / module in each embodiment of the present disclosure can be integrated in one unit / module, or can be physically present separately, or two or more units / modules can be integrated together. The above-mentioned integrated unit / module can be realized in the form of hardware or in the form of a software program module.

[0089] ​If the integrated units / modules are implemented in the form of hardware, the hardware can be a digital circuit, an analog circuit, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the artificial intelligence processor can be any appropriate hardware processor, such as a CPU, a GPU, an FPGA, a DSP, an ASIC, etc. Unless otherwise specified, the storage unit can be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.

[0090] If the integrated units / modules are implemented in the form of software program modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the present disclosure, essentially or the part that contributes to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods in the embodiments of the present disclosure. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0091] In one embodiment, a computer-readable storage medium is disclosed, which stores a computer program. The program is executed to implement the method according to the embodiments of the present disclosure

[0092] In one embodiment, an artificial intelligence chip is also disclosed, which includes the above-mentioned data processing device.

[0093] In one embodiment, a board card is also disclosed, which comprises a memory device, an interface device, a control device and the artificial intelligence chip as described above; wherein the artificial intelligence chip is connected with the memory device, the control device and the interface device respectively; the memory device is used for storing data; the interface device is used for realizing data transmission between the artificial intelligence chip and an external device; and the control device is used for monitoring the state of the artificial intelligence chip.

[0094] Figure 10 A structural block diagram of the board card 1000 according to an embodiment of the present disclosure is shown, referring to Figure 10 The board card 1000 as described above can comprise other supporting components in addition to the chips 1030-1 and 1030-2 (collectively referred to as chips 1030), which include but are not limited to: a memory device 1010, an interface device 1040 and a control device 1020. The interface device 1040 can be connected with an external device 1060. The memory device 1010 is connected with the artificial intelligence chip 1030 through a bus 1050, and is used for storing data. The memory device 1010 can comprise a plurality of groups of storage units 1010-1 and 1010-2. Each group of storage units is connected with the artificial intelligence chip through the bus 1050. It can be understood that each group of storage units can be a DDR SDRAM (English: Double Data Rate SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory).

[0095] DDR can double the speed of SDRAM without increasing the clock frequency. DDR allows data to be read on the rising and falling edges of the clock pulse. The speed of DDR is twice that of standard SDRAM. In one embodiment, the storage device can comprise 4 groups of storage units. Each group of storage units can comprise a plurality of DDR4 particles (chips). In one embodiment, the artificial intelligence chip can internally comprise 4 72-bit DDR4 controllers, of which 64 bits are used for data transmission and 8 bits are used for ECC verification. It can be understood that when DDR4-3200 particles are used in each group of storage units, the theoretical bandwidth of data transmission can reach 25600 MB / s.

[0096] In one embodiment, each group of storage units comprises a plurality of double rate synchronous dynamic random access memories arranged in parallel. DDR can transmit data twice in one clock cycle. A controller for controlling DDR is arranged in the chip, which is used for controlling the data transmission and data storage of each storage unit.

[0097] The interface device is electrically connected to the artificial intelligence chip. The interface device is used to realize data transmission between the artificial intelligence chip and an external device (for example, a server or a computer). For example, in an embodiment, the interface device can be a standard PCIE interface. For example, the data to be processed is transmitted from the server to the chip through the standard PCIE interface, so as to realize data transfer. Preferably, when the PCIE 3.0X 16 interface is used for transmission, the theoretical bandwidth can reach 16000 MB / s. In another embodiment, the interface device can also be other interfaces, and the disclosure does not limit the specific forms of the above-mentioned other interfaces. The interface unit can realize the switching function. In addition, the calculation result of the artificial intelligence chip is still transmitted back to the external device (for example, the server) by the interface device.

[0098] The control device is electrically connected to the artificial intelligence chip. The control device is used to monitor the state of the artificial intelligence chip. Specifically, the artificial intelligence chip and the control device can be electrically connected through an SPI interface. The control device can include a single-chip microcomputer (MCU). For example, the artificial intelligence chip can include multiple processing chips, multiple processing cores or multiple processing circuits, and can drive multiple loads. Therefore, the artificial intelligence chip can be in different working states such as multiple loads and light loads. Through the control device, the working states of the multiple processing chips, the multiple processing and / or the multiple processing circuits in the artificial intelligence chip can be regulated.

[0099] In a possible implementation, an electronic device including the artificial intelligence chip is disclosed. The electronic device includes a data processing device, a robot, a computer, a printer, a scanner, a tablet computer, a smart terminal, a mobile phone, a vehicle record device, a navigation device, a sensor, a camera, a server, a cloud server, a camera, a video camera, a projector, a watch, a headset, a mobile storage, a wearable device, a vehicle, a household appliance, and / or a medical device.

[0100] The vehicle includes an airplane, a ship and / or a vehicle; the household appliance includes a television, an air conditioner, a microwave oven, a refrigerator, an electric rice cooker, a humidifier, a washing machine, an electric lamp, a gas stove, an exhaust hood; the medical device includes a nuclear magnetic resonance instrument, a B-ultrasound instrument and / or an electrocardiograph.

[0101] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments. The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the disclosure.

[0102] The foregoing can be better understood in accordance with the following clauses:

[0103] A1. A method for processing data, comprising: obtaining a set of data to be quantized for a machine learning model;

[0104] determining a plurality of sets of quantized data by quantizing the set of data to be quantized using a plurality of point positions, each of the plurality of point positions specifying a position of a decimal point in the plurality of sets of quantized data; and

[0105] selecting, from the plurality of point positions, a point position for quantizing the set of data to be quantized based on a difference between each of the plurality of sets of quantized data and the set of data to be quantized.

[0106] A2. The method of clause A1, wherein each of the plurality of point positions is represented by an integer, the method further comprising:

[0107] obtaining a point position from the plurality of point positions based on a range associated with the set of data to be quantized; and

[0108] determining other point positions from the plurality of point positions based on integers adjacent to the obtained point position.

[0109] A3. The method of clause A2, wherein determining the other point positions from the plurality of point positions comprises at least one of:

[0110] incrementing the integer representing the point position to determine one of the other point positions; and

[0111] decrementing the integer representing the point position to determine one of the other point positions.

[0112] A4. The method of any one of clauses A1-A3, wherein selecting the point position from the plurality of point positions comprises:

[0113] determining a plurality of differences between the plurality of sets of quantized data and the set of data to be quantized, respectively;

[0114] selecting a smallest difference from the plurality of differences; and

[0115] selecting, from the plurality of point positions, a point position corresponding to the smallest difference.

[0116] A5. The method of clause A4, wherein determining the plurality of differences between the plurality of sets of quantized data and the set of data to be quantized, respectively, comprises, for a given set of quantized data from the plurality of sets of quantized data,

[0117] determining a set of relative differences between the given set of quantized data and the set of data to be quantized, respectively; and

[0118] determining one of the plurality of differences based on the set of relative differences.

[0119] A6. The method of clause A4, wherein determining the plurality of differences between the plurality of sets of quantized data and the set of data to be quantized respectively comprises, for a given set of quantized data of the plurality of sets of quantized data,

[0120] determining a quantization mean of the given set of quantized data and a raw mean of the set of data to be quantized respectively; and

[0121] determining one of the plurality of differences based on the quantization mean and the raw mean.

[0122] A7. The method of any one of clauses A1-A6, wherein the set of data to be quantized comprises a set of floating point numbers in a neural network model, and the method further comprises:

[0123] quantizing the set of data to be quantized using the selected point position to obtain a set of quantized data, wherein quantizing the set of data to be quantized comprises mapping the set of data to be quantized to the set of quantized data based on the selected point position, the position of the decimal point in the set of quantized data being determined by the selected point position; and

[0124] inputting the obtained set of quantized data to the neural network model for processing.

[0125] A8. The method of any one of clauses A1-A6, further comprising:

[0126] obtaining another set of data to be quantized comprising a set of floating point numbers in the neural network model;

[0127] quantizing the other set of data to be quantized using the selected point position to obtain another set of quantized data, wherein quantizing the other set of data to be quantized comprises mapping the other set of data to be quantized to the other set of quantized data based on the selected point position, the position of the decimal point in the other set of quantized data being determined by the selected point position; and

[0128] inputting the obtained other set of quantized data to the neural network model for processing.

[0129] A9. An apparatus for processing data, comprising:

[0130] an obtaining unit configured to obtain a set of data to be quantized for a machine learning model;

[0131] determining a plurality of sets of quantized data by quantizing a set of data to be quantized using a plurality of point positions, each of the plurality of point positions specifying a position of a decimal point in the plurality of sets of quantized data; and

[0132] selecting a point position from the plurality of point positions for quantizing the set of data to be quantized based on a difference between each of the plurality of sets of quantized data and the set of data to be quantized.

[0133] A10. The apparatus of clause A9, wherein each of the plurality of point positions is represented by an integer, the apparatus further comprising:

[0134] obtaining a point position from the plurality of point positions based on a range associated with the set of data to be quantized; and

[0135] determining other point positions from the plurality of point positions based on integers adjacent to the obtained point position.

[0136] A11. The apparatus of clause A10, wherein the point position determining module comprises:

[0137] incrementing the integer representing the point position to determine one of the other point positions; and

[0138] decrementing the integer representing the point position to determine one of the other point positions.

[0139] A12. The apparatus of any one of clauses A9-A11, wherein the selecting module comprises:

[0140] determining a plurality of differences between the plurality of sets of quantized data and the set of data to be quantized, respectively;

[0141] selecting a minimum difference from the plurality of differences; and

[0142] selecting a point position from the plurality of point positions corresponding to the minimum difference.

[0143] A13. The apparatus of clause A12, wherein the difference determining module comprises:

[0144] determining a plurality of relative differences between a given set of quantized data from the plurality of sets of quantized data and the set of data to be quantized,

[0145] determining a plurality of relative differences between a given set of quantized data from the plurality of sets of quantized data and the set of data to be quantized,

[0146] determine one of the plurality of differences based on the set of relative differences.

[0147] A14. The apparatus according to clause A12, characterized in that the difference determining unit comprises:

[0148] a mean determining unit configured to determine, for a given set of quantized data among the plurality of sets of quantized data, a quantized mean of the given set of quantized data and an original mean of the set of data to be quantized, respectively; and

[0149] a mean difference determining unit configured to determine one of the plurality of differences based on the quantized mean and the original mean.

[0150] A15. The apparatus according to any one of clauses A9-A14, characterized in that the set of data to be quantized comprises a set of floating point numbers in a neural network model, and the apparatus further comprises:

[0151] a quantizing unit configured to quantize the set of data to be quantized using the selected point position to obtain a set of quantized data, wherein quantizing the set of data to be quantized comprises: mapping the set of data to be quantized to the set of quantized data based on the selected point position, a position of a decimal point in the set of quantized data being determined by the selected point position; and

[0152] an input unit configured to input the obtained set of quantized data to the neural network model for processing.

[0153] A16. The apparatus according to any one of clauses A9-A14, characterized in that it further comprises:

[0154] a data obtaining unit configured to obtain another set of data to be quantized comprising a set of floating point numbers in the neural network model;

[0155] a quantizing unit configured to quantize the other set of data to be quantized using the selected point position to obtain another set of quantized data, wherein quantizing the other set of data to be quantized comprises: mapping the other set of data to be quantized to the other set of quantized data based on the selected point position, a position of a decimal point in the other set of quantized data being determined by the selected point position; and

[0156] an input unit configured to input the obtained other set of quantized data to the neural network model for processing.

[0157] A17. A computer readable storage medium, characterized in that it has stored thereon a computer program which, when executed, implements the method according to any one of clauses A1-A8.

[0158] A18. An artificial intelligence chip, characterized in that the chip comprises the apparatus for processing data according to any one of clauses A9-A16.

[0159] A19. An electronic device, comprising the artificial intelligence chip according to clause A18.

[0160] A20. A board card, comprising a memory device, an interface device, a control device, and the artificial intelligence chip according to clause A18.

[0161] The artificial intelligence chip is connected with the memory device, the control device, and the interface device.

[0162] The memory device is configured to store data.

[0163] The interface device is configured to realize data transmission between the artificial intelligence chip and an external device.

[0164] The control device is configured to monitor a state of the artificial intelligence chip.

[0165] A21. The board card according to clause A20, wherein

[0166] The memory device comprises a plurality of groups of memory units, each group of memory units is connected with the artificial intelligence chip through a bus, and the memory unit is a DDR SDRAM.

[0167] The chip comprises a DDR controller configured to control data transmission and data storage of each memory unit.

[0168] The interface device is a standard PCIE interface.

[0169] The above describes the embodiments of the present disclosure in detail. The principles and implementation manners of the present disclosure are described by applying specific examples. The above description of the embodiments is only used to help understand the method and core idea of the present disclosure. Meanwhile, the changes or deformations made by the person skilled in the art according to the idea of the present disclosure, based on the specific implementation manners and application scope of the present disclosure, all belong to the protection scope of the present disclosure. In summary, the content of the specification should not be understood as a limitation of the present disclosure.

Claims

1. A method for processing data, characterized in that, include: A set of data to be quantized is obtained for a machine learning model, the set of data to be quantized including the input data of the machine learning model, the input data including one or more of the following: image data, voice data or text data, the machine learning model is implemented using a fixed-point arithmetic unit; By quantizing the set of data to be quantized using multiple point positions, multiple sets of quantized data are determined, and each point position in the multiple sets of quantized data specifies the position of the decimal point in the multiple sets of quantized data. as well as Based on the difference between each set of quantized data and the set of data to be quantized, a point position is selected from the plurality of point positions to be used for quantizing the set of data to be quantized.

2. The method according to claim 1, characterized in that, Each of the plurality of point positions is represented by an integer, and the method further includes: Based on the range associated with the set of data to be quantized, obtain the position of one of the plurality of point positions; and Based on the integers that are close to the obtained point position, determine the other point positions among the plurality of point positions.

3. The method according to claim 2, characterized in that, Determining the other point locations among the plurality of point locations includes at least one of the following: Incrementing the integer representing the point position to determine one of the other point positions; and Decrement the integer representing the point position to determine one of the other point positions.

4. The method according to any one of claims 1-3, characterized in that, Selecting a point location from the plurality of point locations includes: Each of the multiple sets of quantized data and the set of data to be quantized is determined as a different difference. Select the smallest difference from the plurality of differences; and Select one point location from the plurality of point locations that corresponds to the smallest difference.

5. The method according to claim 4, characterized in that, Determining multiple differences between the multiple sets of quantized data and the set of data to be quantized includes: for a given set of quantized data from the multiple sets of quantized data, Determine a set of relative differences between the given set of quantized data and the set of data to be quantized; and One of the plurality of differences is determined based on the set of relative differences.

6. The method according to claim 4, characterized in that, Determining multiple differences between the multiple sets of quantized data and the set of data to be quantized includes: for a given set of quantized data from the multiple sets of quantized data, Determine the quantized mean of the given set of quantized data and the original mean of the set of data to be quantized, respectively; and One of the plurality of differences is determined based on the quantized mean and the original mean.

7. The method according to any one of claims 1-3, characterized in that, The set of data to be quantized includes a set of floating-point numbers in the neural network model, and the method further includes: The selected point locations are used to quantize the set of data to be quantized to obtain a set of quantized data, wherein quantizing the set of data to be quantized includes: mapping the set of data to be quantized to the set of quantized data based on the selected point locations, wherein the position of the decimal point in the set of quantized data is determined by the selected point locations; and The obtained set of quantized data is input into the neural network model for processing.

8. The method according to any one of claims 1-3, characterized in that, Further includes: Obtain another set of data to be quantized, including a set of floating-point numbers from a neural network model; The selected point position is used to quantize the other set of data to be quantized to obtain another set of quantized data, wherein quantizing the other set of data to be quantized includes: mapping the other set of data to be quantized to the other set of quantized data based on the selected point position, wherein the position of the decimal point in the other set of quantized data is determined by the selected point position; as well as The obtained other set of quantized data is input into the neural network model for processing.

9. An apparatus for processing data, characterized in that, include: An acquisition unit is used to acquire a set of data to be quantized for a machine learning model. The set of data to be quantized includes the input data of the machine learning model. The input data includes one or more of the following: image data, voice data, or text data. The machine learning model is implemented using a fixed-point arithmetic unit. A determining unit is used to determine multiple sets of quantized data by quantizing the set of data to be quantized using multiple point positions, wherein each point position in the multiple point positions specifies the position of the decimal point in the multiple sets of quantized data. as well as The selection unit is used to select a point position from the plurality of point positions based on the difference between each group of quantized data and the group of data to be quantized, so as to quantize the group of data to be quantized.

10. The apparatus according to claim 9, characterized in that, Each of the plurality of point positions is represented by an integer, and the device further includes: A point location acquisition unit is configured to acquire the location of one of the plurality of point locations based on a range associated with the set of data to be quantized; and A point position determination unit is used to determine other point positions among the plurality of point positions based on integers that are adjacent to the obtained point positions.

11. The apparatus according to claim 10, characterized in that, The point location determination unit includes: An incrementing unit is used to increment the integer representing the point position to determine one of the other point positions; and A decrementing unit is used to decrement the integer representing the point position to determine one of the other point positions.

12. The apparatus according to any one of claims 9-11, characterized in that, The selection module includes: A difference determination unit is used to determine multiple differences between the multiple sets of quantized data and the set of data to be quantized; A difference selection unit is configured to select the smallest difference from the plurality of differences; and A point location selection unit is used to select a point location from the plurality of point locations that corresponds to the minimum difference.

13. The apparatus according to claim 12, characterized in that, The difference determination unit includes: The relative difference determination unit is used to determine a given set of quantized data from the plurality of quantized data sets. The overall difference determination unit is used to determine a set of relative differences between the given set of quantized data and the set of data to be quantized; and One of the plurality of differences is determined based on the set of relative differences.

14. The apparatus according to claim 12, characterized in that, The difference determination unit includes: The mean determination unit is configured to, for a given set of quantized data from the plurality of quantized data sets, determine the quantized mean of the given set of quantized data sets and the original mean of the set of data to be quantized; and The mean difference determination unit is used to determine one of the multiple differences based on the quantized mean and the original mean.

15. The apparatus according to any one of claims 9-11, characterized in that, The set of data to be quantized includes a set of floating-point numbers in a neural network model, and the device further includes: A quantization unit is configured to quantize a set of data to be quantized using selected point positions to obtain a set of quantized data, wherein quantizing the set of data to be quantized includes: mapping the set of data to be quantized to the set of quantized data based on the selected point positions, wherein the position of the decimal point in the set of quantized data is determined by the selected point positions; and An input unit is used to input the obtained set of quantized data into the neural network model for processing.

16. The apparatus according to any one of claims 9-11, characterized in that, Further includes: The data acquisition unit is used to acquire another set of data to be quantized, including a set of floating-point numbers from the neural network model. A quantization unit is used to quantize the other set of data to be quantized using the selected point position to obtain another set of quantized data, wherein quantizing the other set of data to be quantized includes: mapping the other set of data to be quantized to the other set of quantized data based on the selected point position, wherein the position of the decimal point in the other set of quantized data is determined by the selected point position; as well as An input unit is used to input the obtained other set of quantized data into the neural network model for processing.

17. A computer-readable storage medium, characterized in that... It contains a computer program that, when executed, implements the method according to any one of claims 1-8.

18. An artificial intelligence chip, characterized in that, The chip includes a means for processing data according to any one of claims 9-16.

19. An electronic device, characterized in that, The electronic device includes the artificial intelligence chip according to claim 18.

20. A circuit board, characterized in that, The board includes: storage devices, interface devices, and control devices, as well as the artificial intelligence chip according to claim 18; The artificial intelligence chip is connected to the storage device, the control device, and the interface device. The storage device is used to store data; The interface device is used to realize data transmission between the artificial intelligence chip and external devices; and The controller is used to monitor the state of the artificial intelligence chip.

21. The circuit board according to claim 20, characterized in that, The storage device includes: multiple sets of storage units, each set of storage units being connected to the artificial intelligence chip via a bus, and the storage units being DDR SDRAM; The chip includes: a DDR controller for controlling data transmission and data storage of each memory cell; The interface device is a standard PCIe interface.

Citation Information

Patent Citations

  • Quantification realization method and related product

    CN109993296A

  • Computing device and method

    CN110163353A