Data processing apparatus and method, electronic device, and storage medium
The data processing device addresses the challenge of high computing power and accuracy loss in large-scale models by using quantization and inverse quantization to convert floating-point data to fixed-point data, improving performance and accuracy in matrix operations.
Patent Information
- Application Number
- JP2025072088
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-03
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Existing technologies face challenges in achieving both performance and accuracy in matrix operations during the execution of large-scale models due to the high computing power and storage requirements of floating-point calculations, and existing quantization methods lead to accuracy loss or hardware complexity.
A data processing device with a quantization unit that converts floating-point data to fixed-point data using extreme values, allowing for inverse quantization to maintain accuracy while reducing bandwidth and computing power demands, and a first calculation unit that performs operations on these fixed-point data to achieve precise results.
This approach maintains calculation accuracy while optimizing bandwidth and computing power, enhancing the performance of artificial intelligence chips by utilizing fixed-point numbers for matrix operations.
Smart Images

Figure 2025108728000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to technical fields such as chips, mixed-precision computing, and distributed computing platforms. More specifically, the present disclosure provides a data processing device, method, electronic device, and storage medium.
Background Art
[0002] With the development of artificial intelligence technology, the application scenarios of large-scale models are increasing. It is difficult to achieve both performance and accuracy in matrix operations performed during the execution of large-scale models.
Summary of the Invention
[0003] The present disclosure provides a data processing device, method, electronic device, and storage medium.
[0004] According to an aspect of the present disclosure, there is provided a data processing device including a quantization unit and a first calculation unit. The quantization unit quantizes the floating-point data to be processed into first fixed-point data based on a first extreme value corresponding to the floating-point data to be processed, and quantizes second floating-point data into second fixed-point data based on a second extreme value corresponding to the floating-point data to be processed, where the second floating-point data is obtained based on the floating-point data to be processed and first floating-point data, and the first floating-point data is obtained by inverse quantizing the first fixed-point data. The first calculation unit is configured to obtain a first calculation result based on the fixed-point data to be processed and the first fixed-point data, obtain a second calculation result based on the fixed-point data to be processed and the second fixed-point data, and obtain a target calculation result based on the first calculation result and the second calculation result.
[0005] According to another aspect of the present disclosure, there is provided an electronic device including the device according to the present disclosure.
[0006] According to another aspect of the present disclosure, there is provided a data processing method, including: quantizing floating-point data to be processed into first fixed-point data by a quantization unit based on a first extreme value corresponding to the floating-point data to be processed; quantizing second floating-point data into second fixed-point data by the quantization unit based on a second extreme value corresponding to the floating-point data to be processed, where the second floating-point data is obtained based on the floating-point data to be processed and first floating-point data, and the first floating-point data is obtained by inverse quantization of the first fixed-point data; obtaining a first calculation result by a first calculation unit based on the fixed-point data to be processed and the first fixed-point data; obtaining a second calculation result by the first calculation unit based on the fixed-point data to be processed and the second fixed-point data; and obtaining a target calculation result by the first calculation unit based on the first calculation result and the second calculation result.
[0007] According to another aspect of the present disclosure, there is provided an electronic device, including at least one processor and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to the present disclosure.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to execute the method according to the present disclosure.
[0009] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, which, when executed by a processor, implements the method according to the present disclosure.
[0010] It should be understood that the content described in this part is not intended to indicate the key points or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will be easily understood from the following description.
Brief Description of the Drawings
[0011] The drawings are for a better understanding of the present invention and do not limit the present disclosure.
Figure 1A
Figure 1B
Figure 1C
Figure 1D
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0012] Exemplary embodiments of the present disclosure will be described below with reference to the drawings. Here, various details of the embodiments of the present disclosure are included for easier understanding, and they should be considered exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and configurations are omitted in the following description for clarity and brevity.
[0013] General Matrix to Matrix Multiplication (GEMM) is one of the basic operations related to neural networks and is also an important basic operation required for the execution of large-scale models. General Matrix to Matrix Multiplication is a calculation process that multiplies two matrices to obtain one output matrix. The matrices as inputs for General Matrix to Matrix Multiplication may be a weight matrix and a data matrix respectively, and after performing the matrix multiplication, an output result matrix is obtained. The general calculation flow of General Matrix to Matrix Multiplication is shown in FIG. 1A.
[0014] FIG. 1A is a schematic diagram of General Matrix to Matrix Multiplication according to an embodiment of the present disclosure.
[0015] In some embodiments, the related functional components for performing general matrix multiplication may include a Direct Memory Access (DMA) unit 111, a Multiply Accumulate (MAC) unit 120, and an Element Wise (EW) unit 130. The Direct Memory Access unit 111 can load or store data. When the Direct Memory Access unit 111 executes a load instruction, the Direct Memory Access unit 111 can load data from, for example, the global storage unit 112 to the local storage unit 113. When the Direct Memory Access unit 111 executes a store instruction, the Direct Memory Access unit 111 can store data from, for example, the local storage unit 113 to the global storage unit 112.
[0016] In the embodiments of the present disclosure, with the development of artificial intelligence technology, the number of parameters of deep neural networks has increased rapidly. If floating-point storage and calculation are continuously adopted in general matrix multiplication, the requirements for the computing power and storage capacity of the chip may become too high. Therefore, in order to achieve higher computing power and smaller storage space, quantizing floating-point data into fixed-point data has become one of the mainstream means to accelerate model inference. The quantization unit can perform data type conversion. For example, the quantization unit may perform operations such as cast, quant, and dequant according to the data type.
[0017] In some embodiments, to ensure arithmetic precision, the data matrix of the general matrix multiplication may include a plurality of floating-point values. To reduce the occupation of memory resources, the weight matrix of the general matrix multiplication may include a plurality of fixed-point values. For example, the floating-point value may be a 16-bit floating-point value (FP16). The fixed-point value may be an 8-bit fixed-point value (int8). As shown in FIG. 1A, the input matrix A10 may be the above data matrix. The input matrix B11 may be the above weight matrix. By performing general matrix multiplication by the sum-of-products operation unit 120, the output result matrix C10 can be obtained. The per-element processing unit 130 can perform processing (such as bias addition or scaling, etc.) on the output result matrix C, and then the direct memory access unit 111 can write the processed output result matrix C10 to the global storage unit 114.
[0018] In some embodiments, the general matrix multiplication performed by the sum-of-products operation unit 120 can support inputs of the same data type. When performing general matrix multiplication on two matrices of different data types, the data types of the two matrices can be converted to the same type at the load stage. In this case, the data matrix may be quantized from floating-point numbers to fixed-point numbers, or the weight matrix may be dequantized from fixed-point numbers to floating-point numbers. Next, general matrix multiplication can be performed using the sum-of-products operation unit 120. The following will be described with reference to FIGS. 1B to 1C.
[0019] FIG. 1B is a schematic diagram of general matrix multiplication according to another embodiment of the present disclosure.
[0020] The input matrix A10 may include one or more 16-bit floating-point values. The input matrix B11 may include one or more 8-bit fixed-point values. In the loading stage, the 8-bit fixed-point values of the input matrix B can be dequantized into 16-bit floating-point values, and the dequantized input matrix dB11 can be obtained. The multiply-accumulate unit 120 can calculate based on the input matrix A10 and the dequantized input matrix dB11 to obtain the output result matrix C10. The per-element processing unit 130 can process the output result matrix C10. The direct memory access unit 111 can store the processed output result matrix C10 in the global storage unit 112. However, for the input matrix B11, what is read from the global storage unit 112 is an 8-bit fixed-point number, and what is written to the local storage unit 113 is a 16-bit floating-point number. The amount of written data is about twice the amount of read data, and the amount of written data may reach the bandwidth bottleneck, thereby slowing down the loading process of the input matrix B11 and affecting the chip performance.
[0021] FIG. 1C is a schematic diagram of general matrix multiplication according to another embodiment of the present disclosure.
[0022] In some embodiments, it is different from the embodiment shown in FIG. 1B in that in the loading stage, the 16-bit floating-point values of the input matrix A10 can be quantized into 8-bit fixed-point values. The multiply-accumulate unit 120 can calculate based on the quantized input matrix qA10 and the input matrix B11. The per-element processing unit 130 can process the output result matrix C10. Next, the direct memory access unit 111 can store the processed output result matrix C in the global storage unit 114. In this case, it is difficult for the amount of written data to reach the bandwidth bottleneck, and the risk of performance problems is small. However, when quantizing 16-bit floating-point values into 8-bit fixed-point values in the loading stage, the data accuracy is lost, and further the calculation accuracy of the chip is reduced.
[0023] The above has described several forms of general matrix multiplication between input matrices with different precisions based on quantization and inverse quantization respectively. In some embodiments, the multiply-accumulate unit can be adjusted, which will be described below with reference to FIG. 1D.
[0024] FIG. 1D is a schematic diagram of general matrix multiplication according to another embodiment of the present disclosure.
[0025] In some embodiments, the multiply-accumulate unit 120 may be modified so that the multiply-accumulate unit 120 can support operations on input matrices with different precisions. For example, the modified multiply-accumulate unit can directly multiply the input matrix A10 and the input matrix B11. In the load stage, the direct memory access unit 111 can directly read the input matrix A10 and can also directly read the input matrix B11. Next, the multiply-accumulate unit 120 can perform general matrix multiplication on the input matrix A10 and the input matrix B11 to obtain the output result matrix C10. Performing hardware modification on the multiply-accumulate unit will increase the hardware complexity, increase the modification cost, and lengthen the iteration period.
[0026] In order to improve the precision and performance of the artificial intelligence chip, the present disclosure provides a data processing device, which will be described below.
[0027] FIG. 2 is a schematic block diagram of a data processing device according to an embodiment of the present disclosure.
[0028] As shown in FIG. 2, the data processing device 200 may include a quantization unit 210 and a first calculation unit 220.
[0029] The quantization unit 210 may be configured to quantize the floating-point data to be processed into first fixed-point data based on a first extreme value corresponding to the floating-point data to be processed. For example, the floating-point data to be processed may be the input matrix A10 described above. The input matrix A10 may include one or more 16-bit floating-point values. The first extreme value may be a predetermined value corresponding to the floating-point data to be processed, or may be determined based on the extreme value within the floating-point data to be processed. The first fixed-point data may include one or more 8-bit fixed-point values.
[0030] The quantization unit 210 may be configured to quantize second floating-point data into second fixed-point data based on a second extreme value corresponding to the floating-point data to be processed. For example, the second extreme value may be a predetermined value corresponding to the floating-point data to be processed, or may be determined based on the extreme value within the floating-point data to be processed. The second fixed-point data may include one or more 8-bit fixed-point values. Also, for example, the extreme value may be the maximum value, the minimum value, or the maximum value after taking the absolute value.
[0031] In an embodiment of the present disclosure, the second floating-point data is obtained based on the floating-point data to be processed and the first floating-point data, and the first floating-point data is obtained by inverse quantizing the first fixed-point data.
[0032] For example, when converting floating-point data to fixed-point data, the precision of the data decreases. Therefore, when quantizing the floating-point data to be processed into first fixed-point data, the precision of the first fixed-point data is smaller than that of the floating-point data to be processed. After inverse quantizing the first fixed-point data into floating-point data, there may be a difference between the obtained first floating-point data and the floating-point data to be processed. In this case, the first floating-point data can be used as the upper floating-point component of the floating-point data to be processed. Based on the floating-point data to be processed and the first floating-point data, second floating-point data can be obtained. The second floating-point data corresponds to the lower floating-point component of the floating-point data to be processed. The second floating-point data can be quantized to obtain second fixed-point data.
[0033] The first calculation unit 220 may be configured to obtain a first calculation result based on the fixed-point data to be processed and the first fixed-point data. A second calculation result is obtained based on the fixed-point data to be processed and the second fixed-point data.
[0034] For example, the fixed-point data to be processed may be the above input matrix B11.
[0035] For example, the first calculation unit 220 may be the above product-sum operation unit. Matrix operations can be performed based on the fixed-point data to be processed and the first fixed-point data to obtain a first calculation result. Matrix operations can be performed based on the fixed-point data to be processed and the second fixed-point data to obtain a second calculation result. The matrix operations can include various matrix operations such as matrix multiplication, matrix addition, and matrix subtraction.
[0036] Also, for example, the first calculation unit 212 may be configured to obtain a target calculation result based on the first calculation result and the second calculation result. For example, the target calculation result can be obtained by adding, subtracting, or multiplying the first calculation result and the second calculation result.
[0037] According to an embodiment of the present disclosure, using a quantization and inverse quantization method supported by a chip, one floating-point input data is quantized into a plurality of fixed-point data. Then, by performing a correlation operation on another fixed-point input data with each of the plurality of fixed-point data obtained by quantization, it is possible to reduce the accuracy loss that occurs when directly converting floating-point data to fixed-point data. At the same time, by maintaining the use of fixed-point numbers for calculations when performing matrix calculations, the bandwidth and computing power of the artificial intelligence chip can be fully utilized, and the performance of the chip and the calculation accuracy can be maintained.
[0038] As described above, the data processing apparatus of the present disclosure has been described. Hereinafter, the quantization unit of the present disclosure will be described with reference to FIG. 3.
[0039] FIG. 3 is a schematic diagram of data accuracy conversion according to an embodiment of the present disclosure.
[0040] In some embodiments, the quantization unit can implement conversions between floating-point values of different precisions and between floating-point values and fixed-point values. As shown in FIG. 3, the quantization unit can convert a 32-bit floating-point value (FP32) to a 16-bit floating-point value (cast), and can also convert a 16-bit floating-point value to a 32-bit floating-point value.
[0041] In an embodiment of the present disclosure, quantization is performed based on the floating-point value to be quantized, the target maximum value corresponding to the fixed-point value obtained by quantization, and the minimum and maximum values corresponding to the floating-point number to be quantized, and a fixed-point value can be obtained. For example, a floating-point value can be quantized into a fixed-point value according to the following formula.
[0042]
Equation
[0043] The quantized value may be a fixed-point value obtained by quantization, and the value_f may be a floating-point value to be quantized. The int_max may be a target maximum value. The fixed-point value obtained by quantization can correspond to a certain target accuracy. The target maximum value may be the maximum value indicated by the target accuracy. The max_value_general may be the extreme value corresponding to the floating-point value to be quantized, and the extreme value may be a predetermined value, or may be determined based on the extreme value among a plurality of floating-point values to be quantized. As shown in FIG. 3, the quantization unit can quantize a 32-bit floating-point value into a 16-bit fixed-point value (int16) or an 8-bit fixed-point value (int8). The quantization unit can also quantize a 16-bit floating-point value into a 16-bit fixed-point value or an 8-bit fixed-point value. Based on the above formula 1, a 16-bit floating-point value can be quantized into an 8-bit fixed-point value by the following formula.
[0044] [Number]
[0045] The quant_int8 may be an 8-bit fixed-point value obtained by quantization, and the max_value may be the extreme value corresponding to a 16-bit floating-point number. The int_max may be 127, which is the target maximum value corresponding to the 8-bit fixed-point value.
[0046] In the embodiments of the present disclosure, inverse quantization can be performed based on the fixed-point value obtained by quantization, the target maximum value corresponding to the fixed-point value, and the extreme value corresponding to the floating-point value used during quantization to obtain the floating-point value. For example, a fixed-point value can be quantized into a floating-point value by the following formula.
[0047] [Number]
[0048] The dequant_value is a floating-point value obtained by inverse quantization, and the quant_value is a fixed-point value obtained by quantization. As shown in FIG. 3, the quantization unit can inverse-quantize a 16-bit fixed-point value into a 32-bit floating-point value or a 16-bit floating-point value. The quantization unit can also inverse-quantize an 8-bit fixed-point value into a 32-bit floating-point value or a 16-bit floating-point value.
[0049] As described above, the forms of quantization and inverse quantization of the present disclosure have been described. Hereinafter, the apparatus of the present disclosure will be further described.
[0050] FIG. 4A is a schematic block diagram of a data processing apparatus according to another embodiment of the present disclosure.
[0051] As shown in FIG. 4A, the data processing apparatus may include a quantization unit 410, a first calculation unit 420, a second calculation unit 430, a global storage unit 412, and a local storage unit 413. The above descriptions of the quantization unit 210 and the first calculation unit 210 are similarly applicable to the quantization unit 410 and the first calculation unit 420, and it should be understood that the present disclosure will not repeat the description here. The above descriptions of the processing unit 130, the global storage unit 112, and the local storage unit 113 for each element are similarly applicable to the second calculation unit 430, the global storage unit 412, and the local storage unit 413, and the present disclosure will not repeat the description here.
[0052] In some embodiments, the quantization unit may be integrated with the data transfer unit. For example, the data transfer unit 410 may be a direct memory access unit in the chip. According to an embodiment of the present disclosure, by integrating the quantization unit and the data transfer unit, quantization can be performed when transferring data, and the computing efficiency of the chip can be improved.
[0053] In some embodiments, the data transfer unit may be configured to transfer the floating-point data to be processed from the global storage unit to the local storage unit, and quantize the floating-point data to be processed into first fixed-point data by the quantization unit. As shown in FIG. 4A, the data transfer unit 411 can read the floating-point data F40 to be processed and the fixed-point data I41 to be processed from the global storage unit 412 respectively, and store the floating-point data F40 to be processed and the fixed-point data I41 to be processed in the local storage unit 413. According to the embodiments of the present disclosure, by transferring data to the local storage unit, the efficiency of the quantization unit in processing floating-point data can be improved, and the memory resources of the chip can be fully utilized.
[0054] In the embodiments of the present disclosure, the floating-point data F40 to be processed may include at least one floating-point value to be processed. The quantization unit 410 may correspond to the target accuracy. The maximum value represented by the target accuracy is the target maximum value. The target accuracy may be the accuracy of the fixed-point value after quantization supported by the quantization unit 410. For example, when the target accuracy is the accuracy of an 8-bit fixed-point value, the target maximum value is 127, which is the maximum value that can be represented by an 8-bit fixed-point value.
[0055] In an embodiment of the present disclosure, as described above, the quantization unit can quantize a 16-bit floating-point value into an 8-bit fixed-point value, but this will cause a loss of data accuracy. Quantizing the 16-bit floating-point value into a fixed-point value with higher precision can reduce the accuracy loss. Further, by using the quantization unit to quantize the floating-point value into a plurality of fixed-point values, it is possible to realize higher-precision quantization using a combination of the plurality of fixed-point values. For example, the sign bits of two 8-bit fixed-point values are the same. Two 8-bit fixed-point values can share one sign bit. Therefore, two 8-bit fixed-point values can be combined into a 15-bit fixed-point value (INT15). Also, for example, based on the above formula 1, a 16-bit floating-point value can be quantized into a 15-bit fixed-point value by the following formula.
[0056] [Number]
[0057] quant_int15 may be a 15-bit fixed-point value obtained by quantization, 16383 may be the maximum value represented by a 15-bit fixed-point value, and max_value may be the extreme value corresponding to the floating-point data u to be processed.
[0058] It should be understood that the above quantization unit can quantize a 16-bit floating-point value into an 8-bit fixed-point value, but it is difficult to directly quantize a 16-bit floating-point value into a 15-bit floating-point value. In order to fully utilize the quantization ability of the quantization unit and reduce the modification of the hardware, a 16-bit floating-point value can be quantized into two 8-bit fixed-point values. As shown in FIG. 4A, based on the quantization unit 410 and the second calculation unit 430, the floating-point data F40 to be processed can be quantized into the first fixed-point data I401 and the second fixed-point data I402. The following will be described with reference to FIG. 4B.
[0059] FIG. 4B is a schematic diagram of a quantization unit and a second calculation unit according to an embodiment of the present disclosure.
[0060] In some embodiments, the quantization unit may be configured to quantize the floating-point data to be processed into first fixed-point data based on a first extreme value corresponding to the floating-point data to be processed.
[0061] In an embodiment of the present disclosure, quantization is performed based on the floating-point data to be processed, a target maximum value, and a first extreme value, and first fixed-point data is obtained.
[0062] In an embodiment of the present disclosure, the first fixed-point data may include at least one first fixed-point value.
[0063] In an embodiment of the present disclosure, the first extreme value can be obtained by multiplying the extreme value to be processed, the target maximum value, and the effective parameter value to obtain a first multiplication result, and then dividing the first multiplication result by a predetermined extreme value to obtain the first extreme value. The extreme value to be processed may be a predetermined value or the extreme value among at least one floating-point value to be processed. The effective parameter value can be determined based on the number of effective bits corresponding to the target accuracy. For example, the target maximum value may be 127. The target accuracy may be the accuracy of an 8-bit fixed-point number. The 8-bit fixed-point number may include a sign bit and seven effective bits. Accordingly, the effective parameter value may be 128. The predetermined extreme value may be the maximum value that can be represented with a predetermined accuracy. When the predetermined accuracy is the accuracy of a 15-bit fixed-point value, the predetermined extreme value may be 16383. The first fixed-point value may be an 8-bit fixed-point value or the upper fixed-point component of the above 15-bit fixed-point value. When this upper fixed-point component changes by 1, the 15-bit fixed-point value can be caused to change by 2 7 times (128 times). Thereby, the effective parameter value can be associated with the number of effective bits of the target accuracy.
[0064] In an embodiment of the present disclosure, the quantization unit is further configured to multiply the floating-point value to be processed by a target maximum value to obtain a first multiplication result to be processed, and divide the first multiplication result to be processed by a first extreme value, so as to perform operations of quantizing the floating-point value to be processed into a first fixed-point value, thereby quantizing the floating-point data to be processed into first fixed-point data based on a first extreme value corresponding to the floating-point data to be processed. For example, the floating-point value to be processed can be quantized into a first fixed-point value according to the following formula.
[0065]
Number
[0066]
Number
[0067] In some embodiments, the quantization unit may be configured to inverse-quantize the first fixed-point data to obtain first floating-point data. The first floating-point data may include at least one first floating-point value. As shown in FIG. 4B, using the above formula 3, based on the first extreme value or the extreme value to be processed, the quantization unit 410 can inverse-quantize the first fixed-point value i401 into the first floating-point value df400.
[0068] In some embodiments, the second calculation unit may be configured to subtract the first floating-point data from the floating-point data to be processed to obtain the second floating-point data. The second floating-point data may include at least one second floating-point value. As shown in FIG. 4B, the second calculation unit 430 can subtract the first floating-point value df400 from the floating-point value f400 to be processed to obtain the second floating-point value f400'.
[0069] In some embodiments, the second calculation unit may be configured to provide a second floating-point value to the quantization unit. As shown in FIG. 4B, the second calculation unit 420 can provide the second floating-point value to the quantization unit 410.
[0070] In some embodiments, the quantization unit may further be configured to quantize the second floating-point data into second fixed-point data based on a second extreme value corresponding to the floating-point data to be processed.
[0071] In an embodiment of the present disclosure, the second fixed-point data may include at least one second fixed-point value.
[0072] In an embodiment of the present disclosure, the second extreme value can be obtained by multiplying the extreme value to be processed by the target maximum value to obtain a second multiplication result, and dividing the second multiplication result by a predetermined extreme value to obtain the second extreme value.
[0073] In an embodiment of the present disclosure, the quantization unit may further be configured to multiply the second floating-point value by the target maximum value to obtain a second multiplication result for the object to be processed. By dividing the second multiplication result for the object to be processed by the second extreme value, the floating-point value to be processed is quantized into the second fixed-point value. For example, the second floating-point value can be quantized into the second fixed-point value by the following formula.
[0074]
Equation
[0075]
Equation
[0076] As shown in FIG. 4A, after quantization of one or more floating-point values to be processed in the floating-point data F40 to be processed is completed, the first fixed-point data I401 and the second fixed-point data I402 can be obtained. The first fixed-point data I401 and the second fixed-point data I402 may be provided to the first calculation unit 420. The first calculation unit 420 can perform a correlation operation based on the fixed-point data I41 to be processed, the first fixed-point data I401, and the second fixed-point data I402, which will be described below with reference to FIG. 4C.
[0077] FIG. 4C is a schematic diagram of a first calculation unit according to an embodiment of the present disclosure.
[0078] In some embodiments, the first calculation unit may be configured to obtain a first calculation result based on the fixed-point data to be processed and the first fixed-point data. As shown in FIG. 4C, the first calculation unit 420 can multiply the fixed-point data I41 to be processed by the first fixed-point data I401 to obtain a first calculation result.
[0079] In some embodiments, the first calculation unit may be configured to obtain a second calculation result based on the fixed-point data to be processed and the second fixed-point data. As shown in FIG. 4C, the first calculation unit 410 can multiply the fixed-point data I41 to be processed by the second fixed-point data I402 to obtain a second calculation result.
[0080] In some embodiments, the first calculation unit may be configured to obtain a target calculation result based on the first calculation result and the second calculation result. As shown in FIG. 4C, the first calculation unit 410 can add the first calculation result and the second calculation result to obtain a target calculation result C40.
[0081] The above description has taken the case where the floating-point value to be processed is a 16-bit floating-point value as an example to explain the present disclosure. The present disclosure is not limited thereto, and the floating-point value to be processed may be a 32-bit floating-point value, and the corresponding first fixed-point value and second fixed-point value may be 16-bit fixed-point values.
[0082] The above describes the device of the present disclosure. Hereinafter, an electronic device including the device will be described.
[0083] FIG. 5 is a schematic block diagram of an electronic device according to an embodiment of the present disclosure.
[0084] As shown in FIG. 5, the device 50 may include a data processing device 500. The data processing device 500 may be the above-described data processing device 200.
[0085] The above describes the device and the equipment of the present disclosure. Hereinafter, the method of the present disclosure will be described.
[0086] FIG. 6 is a flowchart of a data processing method according to an embodiment of the present disclosure.
[0087] As shown in FIG. 6, the data processing method 600 of the present disclosure may include operations S610 to S650.
[0088] In operation S610, based on the first extreme value corresponding to the floating-point data to be processed, the quantization unit quantizes the floating-point data to be processed into first fixed-point data.
[0089] In operation S620, based on the second extreme value corresponding to the floating-point data to be processed, the quantization unit quantizes the second floating-point data into second fixed-point data. The second floating-point data is obtained based on the floating-point data to be processed and the first floating-point data, and the first floating-point data is obtained by inverse quantizing the first fixed-point data.
[0090] In operation S630, based on the fixed-point data to be processed and the first fixed-point data, the first calculation unit obtains a first calculation result.
[0091] In operation S640, based on the fixed-point data to be processed and the second fixed-point data, the first calculation unit obtains a second calculation result.
[0092] In operation S650, based on the first calculation result and the second calculation result, the first calculation unit obtains a target calculation result.
[0093] For example, method 600 may be executed by data processing apparatus 200.
[0094] In an embodiment of the present disclosure, the method further includes subtracting first floating-point data from the floating-point data to be processed by a second calculation unit to obtain second floating-point data. The second floating-point data is supplied to a quantization unit by the second calculation unit.
[0095] In an embodiment of the present disclosure, the floating-point data to be processed includes at least one floating-point value to be processed, the quantization unit corresponds to a target accuracy, and the maximum value that can be represented with the target accuracy is the target maximum value.
[0096] In an embodiment of the present disclosure, the method further includes multiplying the extreme value to be processed, the target maximum value, and the effective parameter value to obtain a first multiplication result. The extreme value to be processed is a predetermined value or the extreme value among at least one floating-point value to be processed, and the effective parameter value is determined based on the number of effective bits corresponding to the target accuracy. The first multiplication result is divided by a predetermined extreme value to obtain a first extreme value.
[0097] In an embodiment of the present disclosure, the first fixed-point data includes at least one first fixed-point value. Quantizing the floating-point data to be processed into the first fixed-point data by a quantization unit based on the first extreme value corresponding to the floating-point data to be processed includes multiplying the floating-point value to be processed by a target maximum value to obtain a first multiplication result to be processed, and dividing the first multiplication result to be processed by the first extreme value to quantize the floating-point value to be processed into the first fixed-point value.
[0098] In an embodiment of the present disclosure, the method further includes multiplying a target extreme value and a target maximum value to obtain a second multiplication result. The target extreme value is a predetermined value or the extreme value among at least one floating-point value to be processed. Dividing the second multiplication result by a predetermined extreme value to obtain a second extreme value.
[0099] In an embodiment of the present disclosure, the first floating-point data includes at least one first floating-point value, the second floating-point data includes at least one second floating-point value, and the second fixed-point data includes at least one second fixed-point value. Quantizing the second floating-point data into the second fixed-point data by a quantization unit based on the second extreme value corresponding to the floating-point data to be processed includes multiplying the second floating-point value by a target maximum value to obtain a second multiplication result to be processed. Dividing the second multiplication result to be processed by the second extreme value to quantize the floating-point value to be processed into the second fixed-point value.
[0100] In an embodiment of the present disclosure, the quantization unit is integrated with a data transfer unit.
[0101] In an embodiment of the present disclosure, the method further includes transferring, by a data transfer unit, the floating-point data to be processed from a global storage unit to a local storage unit for the quantization unit to quantize the floating-point data to be processed into the first fixed-point data.
[0102] In the technical solution of the present disclosure, any processing such as obtaining, storing, using, processing, transmitting, providing, and disclosing relevant user personal information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0103] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a computer-readable storage medium, and a computer program product.
[0104] FIG. 7 shows a schematic block diagram of an exemplary electronic device 700 for implementing an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, for example, a laptop computer, a desktop computer, a workstation, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may further represent various forms of mobile devices, for example, a personal digital assistant, a mobile phone, a smart phone, a wearable device, and other similar computing devices. The components shown in this specification, their connections and relationships, and their functions are merely exemplary and do not limit the implementation of the present disclosure described and / or claimed in this specification.
[0105] As shown in FIG. 7, the device 700 includes a computing unit 701, and the computing unit 701 may execute various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 may further store various programs and data required for the operation of the device 700. The computing unit 701, the ROM 702, and the RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0106] The multiple components in device 700 are connected to the I / O interface 705, and include an input unit 706 such as a keyboard and a mouse, an output unit 707 such as various types of displays and speakers, a storage unit 708 such as a magnetic disk and an optical disk, and a communication unit 709 such as a network card, a modem, and a wireless communication transceiver. The communication unit 709 enables the device 700 to exchange information and data with other devices via a computer network such as the Internet and / or various electrical networks.
[0107] The computing unit 701 may be various general-purpose and / or dedicated processing modules having processing and computing capabilities. Some examples of the computing unit 701 include a central processing unit (CPU), a GPU (Graphics Processing Unit), various dedicated artificial intelligence (AI) computing chips, a computing unit running various machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc., but are not limited thereto. The computing unit 701 executes each of the methods and processes described above, such as a data processing method. For example, in some embodiments, the data processing method may be implemented as a computer software program tangibly included in a machine-readable medium such as the storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed into the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the data processing method described above may be executed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute the data processing method in any other suitable manner (e.g., via firmware).
[0108] The various embodiments of the systems and techniques described above in this specification may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may be implemented in one or more computer programs, which may be executed and / or interpreted in a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, and which receives data and instructions from, and transmits data and instructions to, a memory system, at least one input device, and at least one output device.
[0109] The program code for implementing the methods of the present disclosure may be created in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a dedicated computer, or other programmable data processing device, such that, when the program code is executed by the processor or controller, the functions and operations defined in the flowchart and / or block diagram are implemented. The program code may be executed entirely on the device, partially on the device, partially on the device as an independent software package, and partially on a remote device or entirely on a remote device or server.
[0110] In the context of the present disclosure, a machine-readable medium may be a tangible medium that includes or stores a program for use in or in combination with an instruction execution system, apparatus, or electronic device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or electronic device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium include electrical connections by one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.
[0111] To provide for interaction with a user, the computer may implement the systems and techniques described herein, and the computer may include a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or a trackball), and the user may provide input to the computer via the keyboard and the pointing device. Other kinds of devices may further provide for interaction with the user, for example, feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input received from the user may be in any form (including voice input, speech input, or tactile input).
[0112] The systems and techniques described herein may be implemented in a computing system that includes background components (e.g., a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such background components, middleware components, or front-end components. The components of the system can be connected to each other by digital data communication in any form or medium (e.g., a communication network). Exemplary communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0113] The computer system may include clients and servers. The clients and servers are generally remote from each other and typically interact via a communication network. The relationship between the client and the server is generated by a computer program running on the corresponding computer and having a client-server relationship.
[0114] It should be understood that various forms of the flows shown above may be used, and the operations may be sorted, added, or deleted again. For example, each operation described in this disclosure may be executed in parallel, sequentially, or in a different order, and this specification is not limited herein as long as the desired results of the technical solutions disclosed in this disclosure can be achieved.
[0115] The above specific embodiments do not limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and alternatives can be made according to design requirements and other factors. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure should all be included within the protection scope of the present disclosure.
Claims
1. A data processing apparatus, comprising a quantization unit and a first calculation unit, wherein the quantization unit quantizes the floating-point data to be processed into first fixed-point data based on a first extreme value corresponding to the floating-point data to be processed, and quantizes second floating-point data into second fixed-point data based on a second extreme value corresponding to the floating-point data to be processed, wherein the second floating-point data is obtained based on the floating-point data to be processed and first floating-point data, and the first floating-point data is obtained by inverse quantization of the first fixed-point data, and the first calculation unit acquires a first calculation result based on the fixed-point data to be processed and the first fixed-point data, acquires a second calculation result based on the fixed-point data to be processed and the second fixed-point data, and acquires a target calculation result based on the first calculation result and the second calculation result.
2. further comprising a second calculation unit, wherein the second calculation unit subtracts the first floating-point data from the floating-point data to be processed to obtain the second floating-point data, and supplies the second floating-point data to the quantization unit. The apparatus according to claim 1.
3. The floating-point data to be processed includes at least one floating-point value to be processed, the quantization unit corresponds to a target accuracy, and the maximum value that can be represented by the target accuracy is a target maximum value. The apparatus according to claim 1.
4. The first extreme value is obtained by multiplying a processing target extreme value, the target maximum value, and an effective parameter value to obtain a first multiplication result, where the processing target extreme value is a predetermined value or the extreme value among at least one floating-point value to be processed, and the effective parameter value is determined based on the number of effective bits corresponding to the target accuracy, and the first extreme value is obtained by dividing the first multiplication result by a predetermined extreme value. The apparatus according to claim 3.
5. The first fixed-point data includes at least one first fixed-point value, and the quantization unit further Multiply the floating-point value to be processed by the target maximum value to obtain a first multiplication result to be processed. Divide the first multiplication result to be processed by the first extreme value to quantize the floating-point value to be processed into the first fixed-point value. It is configured to perform the operation of The device according to claim 4.
6. The second extreme value is Multiply the extreme value to be processed by the target maximum value to obtain a second multiplication result. The extreme value to be processed is the extreme value among a predetermined value or at least one floating-point value to be processed. The second extreme value is determined by an operation of dividing the second multiplication result by a predetermined extreme value to obtain the second extreme value. The device according to claim 3.
7. The first floating-point data includes at least one first floating-point value, the second floating-point data includes at least one second floating-point value, and the second fixed-point data includes at least one second fixed-point value. The quantization unit is further configured to quantize the second floating-point data into the second fixed-point data based on the second extreme value corresponding to the floating-point data to be processed. Multiply the second floating-point value by the target maximum value to obtain a second multiplication result to be processed. Divide the second multiplication result to be processed by the second extreme value to quantize the second floating-point value into the second fixed-point value. It is configured to perform the operation of The device according to claim 6.
8. The quantization unit is integrated with the data transfer unit. The device according to claim 1.
9. The data transfer unit is configured to transfer the floating-point data to be processed from the global storage unit to the local storage unit so that the quantization unit quantizes the floating-point data to be processed into the first fixed-point data. The device according to claim 8.
10. An electronic device including the data processing device according to any one of claims 1 to 9.
11. A data processing method, comprising: Quantizing, by a quantization unit, the floating-point data to be processed into first fixed-point data based on a first extreme value corresponding to the floating-point data to be processed. Quantizing the second floating-point data into second fixed-point data by the quantization unit based on a second extreme value corresponding to the floating-point data to be processed, wherein the second floating-point data is obtained based on the floating-point data to be processed and first floating-point data, and the first floating-point data is obtained by inverse quantizing the first fixed-point data, Obtaining a first calculation result by a first calculation unit based on the fixed-point data to be processed and the first fixed-point data, Obtaining a second calculation result by the first calculation unit based on the fixed-point data to be processed and the second fixed-point data, Obtaining a target calculation result by the first calculation unit based on the first calculation result and the second calculation result, and A data processing method.
12. Subtracting the first floating-point data from the floating-point data to be processed by a second calculation unit to obtain the second floating-point data, and Further including supplying the second floating-point data to the quantization unit by the second calculation unit, The method according to claim 11.
13. The floating-point data to be processed includes at least one floating-point value to be processed, the quantization unit corresponds to a target accuracy, and the maximum value that can be represented by the target accuracy is a target maximum value, The method according to claim 11.
14. The first extreme value is Obtaining a first multiplication result by multiplying an extreme value of the data to be processed, the target maximum value, and a valid parameter value, wherein the extreme value of the data to be processed is an extreme value among a predetermined value or at least one floating-point value to be processed, and the valid parameter value is determined based on the number of valid bits corresponding to the target accuracy, The first extreme value is obtained by dividing the first multiplication result by a predetermined extreme value, The method according to claim 13.
15. The first fixed-point data includes at least one first fixed-point value, Quantizing the floating-point data to be processed into first fixed-point data by a quantization unit based on a first extreme value corresponding to the floating-point data to be processed is Obtaining a first multiplication result of the floating-point value to be processed by multiplying the target maximum value, Dividing the first multiplication result of the processing target by the first extreme value to quantize the floating-point value of the processing target into the first fixed-point value, and The method according to claim 14.
16. The second extreme value is Multiplying the extreme value of the processing target by the target maximum value to obtain a second multiplication result, where the extreme value of the processing target is the extreme value among a predetermined value or at least one floating-point value of the processing target, The second extreme value is determined by an operation of dividing the second multiplication result by a predetermined extreme value to obtain the second extreme value. The method according to claim 13.
17. The first floating-point data includes at least one first floating-point value, the second floating-point data includes at least one second floating-point value, and the second fixed-point data includes at least one second fixed-point value, Based on the second extreme value corresponding to the floating-point data of the processing target, quantizing the second floating-point data into the second fixed-point data by the quantization unit is Multiplying the second floating-point value by the target maximum value to obtain a second multiplication result of the processing target, and Dividing the second multiplication result of the processing target by the second extreme value to quantize the floating-point value of the processing target into the second fixed-point value, and The method according to claim 16.
18. The quantization unit is integrated with the data transfer unit, The method according to claim 11.
19. In order to quantize the floating-point data of the processing target into the first fixed-point data by the quantization unit, the data transfer unit transfers the floating-point data of the processing target from the global storage unit to the local storage unit, The method according to claim 18.
20. At least one processor, and A memory communicatively connected to the at least one processor, including The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 11 to 19. An electronic device.
21. A non-transitory computer-readable storage medium storing computer instructions, where The computer instructions are used to cause the computer to execute the method according to any one of claims 11 to 19. Non-transitory computer-readable storage medium. **Claim 22** A computer program product including a computer program which, when executed by a processor, implements the method according to any one of claims 11 to 19.
Citation Information
Patent Citations
Arithmetic method for floating point display data
JP1992140827A
Arithmetic unit and operation method
JP2006302180A
Neural network processing unit, neural network processing method and device thereof
JP2022116266A