Data Processing Method, Apparatus, Electronic Device, Medium and Chip
By converting the floating-point data format to a fixed-point data format, the problem of insufficient computing power of traditional CPUs is solved, more efficient matrix computing is achieved, and the calculation accuracy and efficiency of artificial intelligence models are improved.
Patent Information
- Application Number
- CN202210945376.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-08
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-08-08
AI Technical Summary
In the prior art, in artificial intelligence model computing, traditional CPU computing power is insufficient, and the calculation complexity of floating-point data formats is high, resulting in low computing efficiency and serious waste of hardware resources.
Using a fixed-point data format with more mantissa digits, floating-point data is mapped to fixed-point calculations that do not contain exponents, and element distribution is converted by determining the maximum value in the matrix to achieve fixed-point calculations.
It improves calculation accuracy and efficiency, saves hardware resources, improves the training and inference accuracy of artificial intelligence models, and reduces the difficulty of computing.
Smart Images

Figure CN115310035B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and more particularly to the field of artificial intelligence chip technology. Specifically, it relates to a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] Artificial intelligence is a discipline that studies how to make a computer simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0003] Artificial intelligence models have a large number of computationally intensive operators, mainly including matrix multiplication, convolution, pooling, activation, etc. These calculations are very time-consuming, and the computing power of traditional CPUs is difficult to meet the requirements in terms of performance. Therefore, heterogeneous computing has become the mainstream, and various artificial intelligence processors including GPUs, FPGAs, and ASICs have been widely applied to artificial intelligence model calculations. At the same time, the selection of data types also plays a very important role in the accuracy, performance, etc. of artificial intelligence calculations.
[0004] The methods described in this section are not necessarily methods that have been previously envisioned or adopted. Unless otherwise specified, no method described in this section should be considered to be prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention
[0005] The present disclosure provides a method, apparatus, electronic device, computer-readable storage medium, computer program product, and chip.
[0006] According to one aspect of the present disclosure, there is provided a data processing method, including: obtaining a first matrix and a second matrix, wherein each element in the first matrix and the second matrix is stored in a first storage unit in a first data format, and wherein the first data format is a floating-point type and has 1 sign bit and m mantissa bits, where m is an integer greater than 1; determining a first element with the largest absolute value in the first matrix, and denoting the largest absolute value of the first element as a first maximum value Max1; based on the first maximum value Max1, mapping each element in the first matrix to [0, 2 nInterval to convert the element into a corresponding converted element, where the converted element is stored in a second storage unit in a second data format, and where the second data format has 1 sign bit, 1 exponent bit, and n mantissa bits, n is an integer greater than m, and where the second data format and the first data format have the same number of bits; determine the second element with the largest absolute value in the second matrix, and denote the largest absolute value of the second element as the second maximum value Max2; based on the second maximum value Max2, map each element in the second matrix to [0, 2 n interval to convert the element into a corresponding converted element; and calculate the product of the first matrix and the second matrix as a third matrix based on the converted elements corresponding to each element in the first matrix and the second matrix respectively.
[0007] According to another aspect of the present disclosure, there is provided a data processing device, including: an acquisition module configured to acquire a first matrix and a second matrix, where each element in the first matrix and the second matrix is stored in a first storage unit in a first data format, and where the first data format is a floating-point type and has 1 sign bit and m mantissa bits, m is an integer greater than 1; a first determination module configured to determine the first element with the largest absolute value in the first matrix, and denote the largest absolute value of the first element as the first maximum value Max1; a first mapping module configured to map each element in the first matrix to [0, 2 n interval to convert the element into a corresponding converted element, where the converted element is stored in a second storage unit in a second data format, and where the second data format has 1 sign bit, 1 exponent bit, and n mantissa bits, n is an integer greater than m, and where the second data format and the first data format have the same number of bits; a second determination module configured to determine the second element with the largest absolute value in the second matrix, and denote the largest absolute value of the second element as the second maximum value Max2; a second mapping module configured to map each element in the second matrix to [0, 2 n interval to convert the element into a corresponding converted element; and a calculation module configured to calculate the product of the first matrix and the second matrix as a third matrix based on the converted elements corresponding to each element in the first matrix and the conversion matrix respectively.
[0008] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above method.
[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the above method.
[0010] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the above method.
[0011] According to another aspect of the present disclosure, there is provided an electronic circuit, including: a circuit configured to execute the above method.
[0012] According to one or more embodiments of the present disclosure, there is provided a data processing method, which adopts a data format with more mantissa digits than traditional floating-point data, so as to improve the calculation accuracy and further improve the accuracy of artificial intelligence model training and inference. At the same time, by mapping the original floating-point data to all mantissa digits, the floating-point calculation is converted into a fixed-point calculation without an exponent, thereby reducing the calculation difficulty, improving the calculation efficiency, and saving the hardware resources for calculation.
[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings exemplarily show embodiments and constitute a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The shown embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0015] Figure 1 A flowchart of the data processing method according to the embodiments of the present disclosure is shown;
[0016] Figure 2 A flowchart of the method for converting elements in a matrix into conversion elements according to the embodiments of the present disclosure is shown.
[0017] Figure 3 A block diagram of the structure of the data processing device according to the embodiments of the present disclosure is shown; and
[0018] Figure 4 The block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. Detailed implementation manners
[0019] The exemplary embodiments of the present disclosure will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0020] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, timing relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.
[0021] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.
[0022] In the related art, during the calculation process of an artificial intelligence model, the standard IEEE float type is mainly used. With the continuous development of technology, some new calculation types have emerged to replace the standard float, such as new half-precision calculation types like bfloat16, fp16, and int16. Fp16 and bfloat16 include a sign bit, an exponent bit, and a mantissa bit. Because of the exponent bit, the range of values that can be represented is relatively large, but due to the relatively large number of exponent bits, the calculation is relatively complex. Therefore, the calculation efficiency is relatively low compared to int16. Since int16 is fixed-point calculation, the calculation efficiency is higher than that of fp16 / bfloat16, but because there is no exponent bit, the range of numbers that can be represented is smaller than that of fp16 / bfloat16. In the training process of many artificial intelligence models, the situation of non-convergence will occur, so the range of use is relatively limited.
[0023] To solve the above problems, the present disclosure provides a data processing method, which adopts a data format with more mantissa bits than traditional floating-point data, thereby being able to improve the calculation accuracy and further improve the accuracy of artificial intelligence model training and inference. At the same time, by mapping the original floating-point data to all mantissa bits, the floating-point calculation is converted into a fixed-point calculation without an exponent, thereby reducing the calculation difficulty, improving the calculation efficiency, and saving the hardware resources for calculation.
[0024] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0025] Figure 1 The flowchart of the data processing method according to an embodiment of the present disclosure is shown. As Figure 1 shown, the data processing method 100 includes: Step S101, obtaining a first matrix and a second matrix, wherein each element in the first matrix and the second matrix is stored in a first storage unit in a first data format, and wherein the first data format is a floating-point type and has 1 sign bit and m mantissa bits, and m is an integer greater than 1; Step S102, determining a first element with the largest absolute value in the first matrix, and denoting the largest absolute value of the first element as a first maximum value Max1; Step S103, based on the first maximum value Max1, mapping each element in the first matrix to the interval [0, 2 n , so as to convert the element into a corresponding converted element, wherein the converted element is stored in a second storage unit in a second data format, and wherein the second data format has 1 sign bit, 1 exponent bit and n mantissa bits, n is an integer greater than m, and wherein the second data format and the first data format have the same number of bits.
[0026] Thus, through Step S103, each element stored in the first data format in the first matrix is respectively converted into a converted element stored in the second data format, so as to represent each element in the first matrix with more mantissa bits.
[0027] Step S104, determining a second element with the largest absolute value in the second matrix, and denoting the largest absolute value of the second element as a second maximum value Max2; Step S105, based on the second maximum value Max2, mapping each element in the second matrix to the interval [0, 2 n , so as to convert the element into a corresponding converted element.
[0028] Thus, through Step S105, each element stored in the first data format in the second matrix is respectively converted into a converted element stored in the second data format, so as to represent each element in the second matrix with more mantissa bits.
[0029] Step S106: Calculate the product of the first matrix and the second matrix as a third matrix based on the conversion elements corresponding to each element in the first matrix and the second matrix respectively.
[0030] By determining the maximum value of the absolute values in the matrix to determine the data distribution in the matrix, and mapping the value of each element to the trailing digits of the corresponding conversion element, the floating-point type matrix calculation is converted into a fixed-point type matrix operation, which can effectively reduce the difficulty of matrix calculation and save the hardware resources for calculation. Thus, by converting the floating-point data in the first data format into the second data format with more trailing digits, the calculation accuracy can be improved, and the accuracy of training and inference of the artificial intelligence model using this method for matrix operation can be improved. At the same time, by mapping the original floating-point data to all the trailing digits of the conversion element stored in the second data format, the floating-point matrix calculation is converted into a fixed-point calculation without an exponent. Compared with the exponent calculation, the fixed-point calculation is less difficult, so that the calculation efficiency can be improved and the hardware resources for calculation can be saved.
[0031] According to some embodiments, both the first data format and the second data format have 16 bits. Thus, by providing the second data format with a total of 16 bits and 14 exponent bits therein, the product operation of the floating-point data in the first data format such as fp16, bfloat16, etc. is converted into a fixed-point calculation on 14 trailing digits, so that it is simpler and more efficient in hardware implementation, and the energy efficiency can be close to that of the integer type int16 with the same number of bits. Compared with the calculation types such as fp16 and bfloat16, the use of hardware resources can be reduced, thereby improving the peak performance of the artificial intelligence chip.
[0032] For ease of description, the following will be described by taking the second data format as an example that includes 16 bits and the number of trailing digits n is 14. However, it can be understood that the present disclosure is not limited to the conversion of 16-bit data, and can also be used for the conversion of 32-bit single-precision floating-point numbers or 64-bit double-precision floating-point numbers to convert the floating-point operation including exponent operation into a fixed-point operation without an exponent, thereby improving the operation efficiency and saving the hardware calculation resources.
[0033] Figure 2 The flowchart of the method for converting the elements in the matrix into conversion elements according to the embodiments of the present disclosure is shown. As Figure 2As shown, step S103 includes: step S201, determining a first splitting point based on the first maximum value to split the interval [0, Max1] into two sub-intervals; and step S202, for each element in the first matrix, determining the exponent bit of the corresponding transformed element based on the sub-interval where the element is located, and mapping the element to [0, 2 n , to determine the mantissa bit of the corresponding transformed element of the element.
[0034] It can be understood that the distribution range of each element in the matrix, i.e., [0, Max1], can be determined by determining the maximum value of the absolute value of the elements in the first matrix. By determining the splitting point to further segment this distribution range, mapping calculations can be performed separately according to the sub-intervals into which each element falls, making the distribution range of the data more carefully divided, thereby obtaining higher calculation accuracy.
[0035] According to some embodiments, step S201 includes: based on the first maximum value Max1, determining the first splitting point as to split the interval [0, Max1] into two sub-intervals and
[0036] Taking n as 14 as an example, that is, when the second data format is 16-bit data with 14 mantissa bits, each element in the first matrix needs to be mapped to the interval [0, 2 14 to represent the absolute value of each element using 14 mantissa bits. The distribution interval [0, Max1] of the elements in the first matrix can be divided into 2 14 parts, and using the position of the first part as the splitting point to divide the distribution interval [0, Max1] into two sub-intervals and so that the data represented in the second data format has at least precision. Compared with the first data format, the second data format has more mantissa bits, so it has higher precision and can meet the precision requirements of most artificial intelligence models.
[0037] According to some embodiments, step S202 includes: for each element in the first matrix, determining the absolute value a of the element; in response to the element being in the sub-interval inside, determining that the exponent bit of the corresponding transformed element of the element is 0, and mapping the element to [0, 2 n , to determine that the mantissa bit of the corresponding transformed element of the element is or in response to the element being in the sub-interval inside, determine that the exponent bit of the conversion element corresponding to this element is 1, and map this element to [0, 2 n , to determine that the least significant bit of the conversion of this element to the corresponding conversion element is
[0038] Furthermore, by determining the sub-interval where the element is located, the distribution range of the element can be determined more precisely, thereby obtaining higher calculation accuracy. When it is determined that the element falls into the interval with a smaller value distribution, this interval is divided into 2 14 parts, so that the accuracy within the interval can be improved to while the elements with larger values fall into the sub-interval with a larger range and still have the accuracy.
[0039] It can be understood that the position of the set segmentation point can determine the maximum value corresponding to the elements in each sub-interval, and thus determine the accuracy of the elements. Another embodiment of obtaining different accuracies by setting different segmentation points will be given below.
[0040] According to some embodiments, step S201 includes: based on the first maximum value Max1, determining the first segmentation point as to divide the interval [0, Max1] into two sub-intervals and where i is a positive integer less than n.
[0041] According to some embodiments, step S202 includes: for each element in the first matrix, determining the absolute value a of this element; in response to the absolute value a being within the sub-interval inside, determine that the exponent bit of the conversion element corresponding to this element is 0, and map this element to [0, 2 n , to determine that the least significant bit of the conversion of this element to the corresponding conversion element is or in response to the absolute value a being within the sub-interval inside, determine that the exponent bit of the conversion element corresponding to this element is 1, and map this element to [0, 2 n , to determine that the least significant bit of the conversion of this element to the corresponding conversion element is
[0042] Taking n as 14 as an example, that is, when the second data format is a 16-bit data with 14 least significant bits, each element in the first matrix needs to be mapped to the interval of [0, 2 14 , to represent the absolute value of each element with 14 least significant bits. When the first segmentation point is determined to be greater than When compared with the case where the first segmentation point is determined as the maximum value corresponding to the elements in the interval is accurately increased from Max1 to Furthermore, the precision of the elements in this interval can be improved, that is, the precision of the elements in the interval is improved from to In addition, different i values can be set to determine which interval of data precision to improve. For an artificial intelligence model with a certain data distribution, the i value can be set according to the data distribution of the artificial intelligence model, so as to improve the precision of the data in the target interval.
[0043] It can be understood that the conversion process of each element in the second matrix is the same as the conversion process of the elements in the first matrix above, that is, the data in the first data format is converted into the data in the second data format through a mapping method, which will not be elaborated here.
[0044] According to some embodiments, the data processing method 100 further includes: for each element in the first matrix and each element in the second matrix, determining a recovery factor corresponding to the element, and the recovery factor satisfies the following condition: the conversion element corresponding to the element multiplied by the recovery factor corresponding to the element is equal to the absolute value of the element.
[0045] It can be understood that the data processing method 100 is used to convert the traditional floating-point data type into a new data type with more mantissa bits, so as to convert the matrix calculation of the floating-point type into a fixed-point calculation, thereby reducing the complexity of the calculation and achieving the purpose of saving hardware resources. After converting to fixed-point for matrix calculation, the calculation result still needs to be converted into the original data type to make this conversion process unknown to the user perspective and improve the user experience.
[0046] The process of converting the matrix calculation result into the original data type requires the above-mentioned recovery factor to achieve. Specifically, when converting an element in the first data format into a conversion element in the second data format, the element is mapped to [0, 2 n by multiplying the absolute value of the element by a factor, then the matrix calculation result can be converted into the original data type by multiplying the calculation result by the reciprocal of this factor, and the reciprocal of this factor is the recovery factor. The recovery factor satisfies the following condition: the conversion element corresponding to the element multiplied by the recovery factor corresponding to the element is equal to the absolute value of the element, so as to convert the matrix calculation result into the original data type.
[0047] According to some embodiments, the absolute value of each element in the third matrix is equal to the least significant digit of the conversion element of the third element corresponding to this element in the first matrix multiplied by the least significant digit of the conversion element of the fourth element corresponding to this element in the second matrix multiplied by the recovery factor corresponding to the third element multiplied by the recovery factor corresponding to the fourth element, and the sign bit of this element in the third matrix is the exclusive OR value of the sign bit of the third element and the sign bit of the fourth element.
[0048] Thus, the matrix calculation of floating-point numbers is converted into fixed-point calculation with more least significant digits, reducing the calculation complexity and improving the calculation efficiency. Due to the increase in the least significant digits, the precision of the data is also improved. At the same time, through the above process, automatic conversion of data is achieved, and the calculation result can be automatically converted back to the original data type. Users can't feel the specific data type used in the calculation and the data conversion process during use, and can obtain calculation results and calculation processes with higher precision and higher efficiency.
[0049] According to another aspect of the present disclosure, a data processing device is provided. As Figure 3 shown, the data processing device 300 includes: an acquisition module 301 configured to acquire a first matrix and a second matrix, wherein each element in the first matrix and the second matrix is stored in a first storage unit in a first data format, and wherein the first data format is a floating-point type and has 1 sign bit and m least significant digits, where m is an integer greater than 1; a first determination module 302 configured to determine a first element with the largest absolute value in the first matrix and denote the largest absolute value of the first element as a first maximum value Max1; a first mapping module 303 configured to map each element in the first matrix to the interval [0, 2 n based on the first maximum value Max1 to convert this element into a corresponding conversion element, wherein the conversion element is stored in a second storage unit in a second data format, and wherein the second data format has 1 sign bit, 1 exponent bit, and n least significant digits, where n is an integer greater than m, and wherein the second data format and the first data format have the same number of bits; a second determination module 304 configured to determine a second element with the largest absolute value in the second matrix and denote the largest absolute value of the second element as a second maximum value Max2; a second mapping module 305 configured to map each element in the second matrix to the interval [0, 2 n based on the second maximum value Max2 to convert this element into a corresponding conversion element; and a calculation module 306 configured to calculate the product of the first matrix and the second matrix as a third matrix based on the first conversion matrix and the second conversion matrix.
[0050] Accordingly, each element stored in the first matrix in the first data format is respectively converted into a converted element stored in the second data format through the first mapping module 303, so as to represent each element in the first matrix with more trailing bits. Each element stored in the second matrix in the first data format is respectively converted into a converted element stored in the second data format through the second mapping module 306, so as to represent each element in the second matrix with more trailing bits.
[0051] The maximum value of the absolute value in the matrix is determined by the first determination module 302 and the second determination module 304 to determine the data distribution in the matrix, so as to map the value of each element to the trailing bits of the corresponding converted element, converting the floating-point matrix calculation into a fixed-point matrix operation, which can effectively reduce the difficulty of matrix calculation and save the hardware resources for calculation. Accordingly, by converting the floating-point data in the first data format into the second data format with more trailing bits, the data processing device 300 can improve the calculation accuracy and can improve the accuracy of training and inference of the artificial intelligence model using this method for matrix operations.
[0052] At the same time, by mapping the original floating-point data to all the trailing bits of the converted element stored in the second data format, the first mapping module 303 and the second mapping module 305 realize the conversion of the floating-point matrix calculation into a fixed-point calculation without an exponent. Compared with the exponent calculation, the fixed-point calculation is less difficult, so that the calculation efficiency can be improved and the hardware resources for calculation can be saved.
[0053] According to some embodiments, both the first data format and the second data format have 16 bits. Accordingly, by providing the second data format with a total of 16 bits and 14 exponent bits therein, the multiplication operation of the floating-point data in the first data format such as fp16, bfloat16, etc. is converted into a fixed-point calculation on 14 trailing bits, so that it is simpler and more efficient in hardware implementation, and the energy efficiency can be close to that of the integer type int16 with the same number of bits. Compared with calculation types such as fp16 and bfloat16, the use of hardware resources can be reduced, thereby improving the peak performance of the artificial intelligence chip.
[0054] It can be understood that the data processing device 300 provided in the present disclosure is not limited to the conversion of 16-bit data, and can also be used for the conversion of 32-bit single-precision floating-point numbers or 64-bit double-precision floating-point numbers, so as to convert the floating-point operation including exponent operation into a fixed-point operation without an exponent, thereby improving the operation efficiency and saving the hardware calculation resources.
[0055] According to some embodiments, the first mapping module 303 includes: a first determining unit configured to determine a first split point based on the first maximum value to divide the interval [0, Max1] into two sub-intervals; and a second determining unit configured to, for each element in the first matrix, determine the exponent bit of the corresponding transformed element based on the sub-interval in which the element is located, and map the element to [0, 2 n , to determine the mantissa bit of the corresponding transformed element of the element.
[0056] It can be understood that the first determining unit can determine the distribution range of each element in the first matrix, i.e., [0, Max1], by determining the maximum value of the absolute values of the elements in the first matrix. By further segmenting this distribution range by determining the split point and performing mapping calculations separately according to the sub-interval into which each element falls, the distribution range of the data is divided more carefully, thereby obtaining higher calculation accuracy.
[0057] According to some embodiments, the first determining unit is further configured to: based on the first maximum value Max1, determine the first split point as to divide the interval [0, Max1] into two sub-intervals and
[0058] Taking n as 14 as an example, that is, when the second data format is 16-bit data with 14 mantissa bits, the first determining unit needs to map each element in the first matrix to the interval [0, 2 14 to represent the absolute value of each element using 14 mantissa bits. The first determining unit can divide the distribution interval [0, Max1] of the elements in the first matrix into 2 14 parts, and use the position of the first part as the split point to divide the distribution interval [0, Max1] into two sub-intervals and so that the data represented in the second data format has at least precision. Compared with the first data format, the second data format has more mantissa bits and thus higher precision, which can meet the precision requirements of most artificial intelligence models.
[0059] According to some embodiments, the second determining unit includes: a first determining subunit configured to determine the absolute value a of each element in the first matrix; a second determining subunit configured to, in response to the element being in the sub-interval inside, determine that the exponent bit of the corresponding transformed element of the element is 0, and map the element to [0, 2 n , to determine that the mantissa bit of the corresponding transformed element of the element is Or a third determination subunit, configured to determine that the exponent bit of the conversion element corresponding to the element is 1 in response to the element being located in the sub-interval and map the element to [0, 2 n to determine that the least significant bit of the conversion of the element to the corresponding conversion element is
[0060] Furthermore, by determining the sub-interval where the element is located through the second determination subunit and the third determination subunit, the distribution range of the element can be determined more precisely, thereby obtaining higher calculation accuracy. When the second determination subunit determines that the element falls into the interval with a smaller value distribution, this interval is divided into 2 14 parts, so that the accuracy within the interval can be increased to while the elements with larger values fall into the sub-interval with a larger range and still have the accuracy.
[0061] According to some embodiments, the first determination unit is further configured to: based on the first maximum value Max1, determine the first split point as to divide the interval [0, Max1] into two sub-intervals and where i is a positive integer less than n.
[0062] According to some embodiments, the second determination unit includes: a fourth determination subunit, configured to determine the absolute value a of each element in the first matrix; a fifth determination subunit, configured to determine that the exponent bit of the conversion element corresponding to the element is 0 in response to the absolute value a being located in the sub-interval and map the element to [0, 2 n to determine that the least significant bit of the conversion of the element corresponding to the element is or a sixth determination subunit, configured to determine that the exponent bit of the conversion element corresponding to the element is 1 in response to the absolute value a being located in the sub-interval and map the element to [0, 2 n to determine that the least significant bit of the conversion of the element to the corresponding conversion element is
[0063] Taking n as 14 as an example, that is, when the second data format is 16-bit data with 14 least significant bits, each element in the first matrix needs to be mapped to the interval [0, 2 14 to represent the absolute value of each element using 14 least significant bits. When the first split point is determined to be greater than of When, compared with the case where the first segmentation point is determined as the maximum value of the elements located in the interval is accurately increased from Max1 to Furthermore, the precision of the elements within this interval can be improved, that is, the precision of the elements located in the interval is increased from to In addition, by setting different values of i, the precision of the data in which interval can be determined to be improved. For an artificial intelligence model with a certain data distribution, the value of i can be set according to the data distribution of the artificial intelligence model, so as to improve the precision of the data in the target interval.
[0064] It can be understood that the conversion process of each element in the second matrix by the data processing device 300 is the same as the conversion process of the elements in the first matrix described above, that is, the data in the first data format is converted into the data in the second data format by means of mapping, which will not be elaborated here.
[0065] According to some embodiments, the data processing device 300 further includes: a third determination module configured to determine, for each element in the first matrix and each element in the second matrix, a recovery factor corresponding to the element, and the recovery factor satisfies the following condition: the converted element corresponding to the element multiplied by the recovery factor corresponding to the element is equal to the absolute value of the element.
[0066] It can be understood that the data processing device 300 is used to convert the traditional floating-point data type into a new data type with more mantissa bits, so as to convert the matrix calculation of the floating-point type into a fixed-point calculation, thereby reducing the complexity of the calculation and achieving the purpose of saving hardware resources. After converting to fixed-point for matrix calculation, the calculation result still needs to be converted into the original data type so that this conversion process is unknown to the user to improve the user experience.
[0067] The process of converting the matrix calculation result by the data processing device 300 into the original data type requires the recovery factor determined by the above-mentioned fifth determination module to achieve. Specifically, when converting an element in the first data format into a converted element in the second data format, the absolute value of the element is multiplied by a factor by the first mapping module 303 to map the element to [0, 2 n , then the matrix calculation result can be converted into the original data type by multiplying the calculation result by the reciprocal of this factor, and the reciprocal of this factor is the recovery factor. The recovery factor satisfies the following condition: the converted element corresponding to the element multiplied by the recovery factor corresponding to the element is equal to the absolute value of the element, so as to convert the matrix calculation result into the original data type.
[0068] According to some embodiments, the absolute value of each element in the third matrix is equal to the tail digit of the conversion element of the third element corresponding to the element in the first matrix multiplied by the tail digit of the conversion element of the fourth element corresponding to the element in the second matrix multiplied by the recovery factor corresponding to the third element multiplied by the recovery factor corresponding to the fourth element, and the sign bit of the element in the third matrix is the exclusive OR value of the sign bit of the third element and the sign bit of the fourth element.
[0069] Thereby, the data processing device 300 converts the matrix calculation of floating-point numbers into fixed-point calculation with more tail digits, reduces the complexity of calculation, and improves the calculation efficiency. Due to the increase in the number of tail digits, the precision of the data is also improved. At the same time, through the above process, the data processing device 300 realizes the automatic conversion of data, and the calculation result can be automatically converted back to the original data type. In use, the user cannot feel the specific data type used in the calculation and the data conversion process, and can obtain higher-precision and higher-efficiency calculation results and calculation processes.
[0070] According to another aspect of the present disclosure, there is also provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data processing method.
[0071] According to another aspect of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the data processing method.
[0072] According to another aspect of the present disclosure, there is also provided a computer program product, including a computer program, wherein the computer program implements the data processing method when executed by a processor.
[0073] According to another aspect of the present disclosure, there is also provided an electronic circuit, including: a circuit configured to execute the data processing method, and the electronic circuit can be implemented as a chip.
[0074] As Figure 4As shown, the electronic device 400 includes a computing unit 401 which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0075] A plurality of components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, an output unit 407, a storage unit 408, and a communication unit 409. The input unit 406 can be any type of device capable of inputting information into the electronic device 400. The input unit 406 can receive input digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device, and can include, but is not limited to, a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 407 can be any type of device capable of presenting information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 408 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth TM device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0076] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the data processing method. For example, in some embodiments, the data processing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the data processing method described above can be executed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the data processing method by any other suitable means (e.g., by means of firmware).
[0077] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0078] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0079] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0080] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0081] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0082] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0083] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0084] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after the present disclosure.
Claims
1. A data processing method, comprising: Obtaining a first matrix and a second matrix, wherein each element in the first matrix and the second matrix is stored in a first storage unit in a first data format, and wherein the first data format is a floating-point type and has 1 sign bit and m mantissa bits, where m is an integer greater than 1; Determining a first element with the largest absolute value in the first matrix, and denoting the largest absolute value of the first element as a first maximum value Max1; Based on the first maximum value Max1, map each element in the first matrix to the interval [0, 2 n to convert the element into a corresponding converted element, where the converted element is stored in a second storage unit in a second data format, and where the second data format has 1 sign bit, 1 exponent bit, and n mantissa bits, n being an integer greater than m, and where the second data format and the first data format have the same number of bits; Determining a second element with the largest absolute value in the second matrix, and denoting the largest absolute value of the second element as a second maximum value Max2; Based on the second maximum value Max2, map each element in the second matrix to the interval [0, 2 n to convert the element into a corresponding converted element; and Calculating the product of the first matrix and the second matrix as a third matrix based on the conversion elements corresponding to each element in the first matrix and the second matrix.
2. The method according to claim 1, wherein Based on the first maximum value Max1, mapping each element in the first matrix to the interval [0, 2 n , to convert the element into a corresponding converted element, includes: Determining a first splitting point based on the first maximum value to split the interval [0, Max1] into two sub-intervals; and For each element in the first matrix, based on the sub-interval where the element is located, determine the exponent bit of the conversion element corresponding to the element, and map the element to [0, 2 n , to determine the mantissa bit of the conversion element corresponding to the element.
3. The method according to claim 2, wherein The determining a splitting point based on the first maximum value Max1 to split the interval [0, Max1] into two sub-intervals includes: Based on the first maximum value Max1, determine the first splitting point as to divide the interval [0, Max1] into two sub-intervals and 4. The method according to claim 3, wherein For each element in the first matrix, based on the sub-interval in which the element is located, determine the exponent bit of the conversion element corresponding to the element, and map the element to [0, 2 n , to determine the mantissa bit of the conversion element corresponding to the element, including: For each element in the first matrix, Determining the absolute value a of the element; In response to the element being located within the sub - interval determine that the exponent bit of the conversion element corresponding to the element is 0, and map the element to [0, 2 n to determine that the mantissa bit of the conversion element corresponding to the element is or In response to the element being located in the sub - interval determine that the exponent bit of the conversion element corresponding to the element is 1, and map the element to [0, 2 n to determine that the least - significant bit of the element converted to the corresponding conversion element is 5. The method according to claim 2, wherein The determining a splitting point based on the first maximum value Max1 to split the interval [0, Max1] into two sub-intervals includes: Based on the first maximum value Max1, determine the first splitting point as to divide the interval [0, Max1] into two sub-intervals and where i is a positive integer less than n.
6. The method according to claim 5, wherein For each element in the first matrix, based on the sub-interval in which the element is located, determine the exponent bit of the conversion element corresponding to the element, and map the element to [0, 2 n , to determine the mantissa bit of the conversion element corresponding to the element, including: For each element in the first matrix, Determining the absolute value a of the element; In response to the absolute value a being within the sub-interval , determine that the exponent bit of the conversion element corresponding to this element is 0, and map this element to [0, 2 n , to determine that the mantissa bit of the conversion element corresponding to this element is or In response to the absolute value a being within the sub-interval , determine that the exponent bit of the conversion element corresponding to this element is 1, and map this element to [0, 2 n to determine that the tail bit of the conversion of this element to the corresponding conversion element is 7. The method according to any one of claims 1-6, further comprising: For each element in the first matrix and each element in the second matrix, determining a recovery factor corresponding to the element, and the recovery factor satisfies the following condition: The conversion element corresponding to the element multiplied by the recovery factor corresponding to the element is equal to the absolute value of the element.
8. The method according to claim 7, wherein The absolute value of each element in the third matrix is equal to the mantissa bits of the conversion element of the third element corresponding to the element in the first matrix multiplied by the mantissa bits of the conversion element of the fourth element corresponding to the element in the second matrix multiplied by the recovery factor corresponding to the third element multiplied by the recovery factor corresponding to the fourth element, and the sign bit of the element in the third matrix is the exclusive OR value of the sign bit of the third element and the sign bit of the fourth element.
9. A data processing apparatus, comprising: An obtaining module configured to obtain a first matrix and a second matrix, wherein each element in the first matrix and the second matrix is stored in a first storage unit in a first data format, and wherein the first data format is a floating-point type and has 1 sign bit and m mantissa bits, where m is an integer greater than 1; A first determining module configured to determine a first element with the largest absolute value in the first matrix, and denoting the largest absolute value of the first element as a first maximum value Max1; The first mapping module is configured to map each element in the first matrix to the interval [0, 2 n based on the first maximum value Max1, so as to convert the element into a corresponding converted element, wherein the converted element is stored in the second storage unit in a second data format, and wherein the second data format has 1 sign bit, 1 exponent bit and n mantissa bits, n is an integer greater than m, and wherein the second data format and the first data format have the same number of bits; A second determining module configured to determine a second element with the largest absolute value in the second matrix, and denoting the largest absolute value of the second element as a second maximum value Max2; A second mapping module, configured to map each element in the second matrix to the interval [0, 2 n based on the second maximum value Max2, so as to convert the element into a corresponding converted element; and A calculation module, configured to calculate a product of the first matrix and the second matrix as a third matrix based on conversion elements corresponding to each element in the first matrix and the second matrix respectively.
10. The apparatus according to claim 9, wherein The first mapping module includes: A first determination unit, configured to determine a first segmentation point based on the first maximum value to divide the interval [0, Max1] into two sub-intervals; and A second determination unit, configured to, for each element in the first matrix, determine an exponent bit of a conversion element corresponding to the element based on a sub-interval where the element is located, and map the element to [0, 2 n , to determine a mantissa bit of the conversion element corresponding to the element.
11. The apparatus according to claim 10, wherein, The first determination unit is further configured to: Based on the first maximum value Max1, determine the first splitting point as to divide the interval [0, Max1] into two sub-intervals and 12. The apparatus according to claim 11, wherein, The second determination unit includes: A first determination subunit, configured to determine an absolute value a of each element in the first matrix; A second determination subunit, configured to determine that an exponent bit of a conversion element corresponding to the element is 0 in response to the element being located within a sub-interval and map the element to [0, 2 n to determine that a tail bit of the conversion element corresponding to the element is or The third determination subunit is configured to determine that the exponent bit of the conversion element corresponding to the element is 1 in response to the element being located in the sub-interval and map the element to [0, 2 n to determine that the least significant bit of the element converted into the corresponding conversion element is 13. The device according to claim 10, wherein, The first determination unit is further configured to: Based on the first maximum value Max1, determine the first segmentation point as to divide the interval [0, Max1] into two sub-intervals and where i is a positive integer less than n.
14. The device according to claim 13, wherein, The second determination unit includes: A fourth determination subunit, configured to determine an absolute value a of each element in the first matrix; A fifth determination subunit, configured to determine that the exponent bit of the conversion element corresponding to the element is 0 in response to the absolute value a being within the sub-interval and map the element to [0, 2 n to determine that the mantissa bit of the conversion element corresponding to the element is or The sixth determination subunit is configured to, in response to the absolute value a being within the sub-interval , determine that the exponent bit of the conversion element corresponding to this element is 1, and map this element to [0, 2 n , to determine that the mantissa bit of this element converted to the corresponding conversion element is 15. The apparatus according to any one of claims 9-14, further comprising: A third determination module, configured to determine a recovery factor corresponding to each element in the first matrix and each element in the second matrix, where the recovery factor satisfies the following formula: The conversion element corresponding to the element multiplied by the recovery factor corresponding to the element is equal to the absolute value of the element.
16. The apparatus according to claim 15, wherein, The absolute value of each element in the third matrix is equal to the tail digit of the conversion element of the third element corresponding to the element in the first matrix multiplied by the tail digit of the conversion element of the fourth element corresponding to the element in the second matrix multiplied by the recovery factor corresponding to the third element multiplied by the recovery factor corresponding to the fourth element, and the sign bit of the element in the third matrix is the exclusive OR value of the sign bit of the third element and the sign bit of the fourth element.
17. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; Wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.
19. A computer program product comprising a computer program, wherein, The computer program implements the method according to any one of claims 1-8 when executed by a processor.
20. An electronic circuit, comprising: A circuit configured to execute the method according to any one of claims 1-8.
Citation Information
Patent Citations
Automatic relationship mode conversion method and device, and storage medium
CN108776673A
Floating point operation circuit implementation method for real number matrix inversion
CN110162742A