Hybrid precision tensor calculation module, system-in-package and tensor calculation method
Through the dynamic bias and data representation configuration of the mixed precision tensor calculation module, the problems of calculation accuracy and efficiency are solved, efficient floating-point calculation is realized, and computing scenarios that support a variety of software requirements are required.
Patent Information
- Application Number
- CN202510503778.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art cannot improve large-scale data computing and storage efficiency while ensuring computing accuracy. Especially in deep learning models, low-precision floating-point number calculation results in a small data bit width and cannot effectively represent data accuracy.
The hybrid precision tensor calculation module is adopted, including operand information processing unit, dynamic configuration information processing unit and data calculation unit. Through the configuration of dynamic bias and data representation, different forms of floating-point calculation are supported to improve calculation performance and accuracy.
It realizes the calculation accuracy while improving computing performance, supports any form of floating-point calculation, solves the contradiction between hardware design dependence on software requirements, and is suitable for computing scenarios with different software requirements.
Smart Images

Figure CN120295602A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence computing technology, and particularly to a mixed-precision tensor calculation module, a system-in-package, and a tensor calculation method. Background Art
[0002] With the development of deep learning and artificial intelligence technologies, in application scenarios such as large language model data processing, the scale of data processed by computing devices and the complexity of data calculations have increased exponentially, and the computing resources, hardware performance, and energy consumption efficiency of computing devices have all been severely tested.
[0003] Due to the relatively small data bit width, low-precision floating-point numbers can effectively improve computing performance and reduce storage costs. For deep learning models that require large-scale parallel computing, computing with low-precision floating-point numbers can significantly optimize model performance. However, since normalization techniques are often used in the process of deep learning operations, normalization techniques scale floating-point values to very small values, and low-precision floating-point numbers have a relatively low representable data precision due to their small data bit width.
[0004] In the prior art, no computing hardware has been proposed that can both ensure the computing precision of floating-point numbers and improve the computing and storage efficiency of large-scale data. Summary of the Invention
[0005] The present invention provides a mixed-precision tensor calculation module, a system-in-package, and a tensor calculation method, which can solve the problem of low computing performance of existing computing hardware in application scenarios with a large computing scale such as large language models, and can achieve fast computing of floating-point numbers while ensuring computing precision.
[0006] According to one aspect of the present invention, a mixed-precision tensor calculation module is provided. The mixed-precision tensor calculation module is configured in a system-in-package and includes an operand information processing unit, a dynamic configuration information processing unit, and a data calculation unit; wherein,
[0007] The operand information processing unit is used to receive and store the operation type and operand information sent by the system-in-package;
[0008] The dynamic configuration information processing unit is used to determine the data representation form and dynamic bias of each operand according to the operand information in the operand information processing unit;
[0009] The data calculation unit is used to perform a calculation operation according to the value of the operand, the operation type, the data representation form of each operand, and the dynamic bias.
[0010] According to another aspect of the present invention, there is provided a system-in-package, including a mixed-precision tensor calculation module as described in any embodiment of the present invention and at least one special calculation module;
[0011] The calculation type of the special calculation module is different from that of the mixed-precision tensor calculation module;
[0012] The special calculation module is configured to, after responding to the call of the mixed-precision tensor calculation module, perform a calculation operation according to the operation type and operand information sent by the mixed-precision tensor calculation module.
[0013] According to another aspect of the present invention, there is provided a method for mixed-precision tensor calculation, which is executed by an artificial intelligence chip. The artificial intelligence chip includes a system-in-package as described in any embodiment of the present invention, including:
[0014] Generate a floating-point calculation instruction. During the compilation of the floating-point instruction, verify the dynamic bias of each operand, and send the floating-point instruction to the system-in-package;
[0015] Through the system-in-package, according to the floating-point calculation instruction, parse to obtain the operation type and operand information, and send the operation type and operand information to the mixed-precision tensor calculation module;
[0016] Through the mixed-precision tensor calculation module, perform a mixed-precision tensor calculation operation according to the operation type and operand information.
[0017] In the technical solution of the embodiment of the present invention, by parsing the operand information in the operand processing unit through the dynamic configuration information processing unit, the pre-configured dynamic bias and data representation form can be obtained in the operand information. The dynamic bias can make the exponent of the low-precision floating-point number have a smaller representation range, thereby improving the accuracy of numerical calculation. The low-precision data operation can improve the data operation performance. By configuring the dynamic bias while improving the calculation performance, the calculation accuracy can be guaranteed at the same time. And because both the data representation form and the dynamic bias can be configured, therefore, there is no limitation on the floating-point number form supported by the mixed-precision tensor calculation module for calculation, and it can support floating-point calculations in any form. In the prior art, the hardware design generally depends on software requirements, and software depends on hardware performance when constructing processing logic, and there is a certain contradiction between software and hardware. However, the mixed-precision tensor calculation module proposed by the present invention is not limited to providing calculation functions for floating-point numbers in a specified format, and can provide computing power support for different software requirements, with wide applications.
[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Brief Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 FIG. 1 is a schematic structural diagram of a mixed-precision tensor calculation module according to Embodiment 1 of the present invention;
[0021] Figure 2 FIG. 2 is a schematic diagram of the numerical representation of a floating-point number according to an embodiment of the present invention;
[0022] Figure 3 FIG. 3 is a schematic structural diagram of another mixed-precision tensor calculation module according to an embodiment of the present invention;
[0023] Figure 4 FIG. 4 is a schematic structural diagram of another mixed-precision tensor calculation module according to Embodiment 2 of the present invention;
[0024] Figure 5 FIG. 5 is a schematic structural diagram of a system-in-package according to Embodiment 3 of the present invention;
[0025] Figure 6 FIG. 6 is a flowchart of a mixed-precision tensor calculation method according to Embodiment 4 of the present invention. Detailed Description of the Embodiments
[0026] To enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0027] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0028] Embodiment 1
[0029] Figure 1 FIG. is a schematic structural diagram of a mixed-precision tensor calculation module provided in Embodiment 1 of the present invention. As Figure 1 shown, the mixed-precision tensor calculation module 100 includes an operand information processing unit 110, a dynamic configuration information processing unit 120, and a data calculation unit 130.
[0030] Among them, the mixed-precision tensor calculation module is configured in a System In a Package (SIP). The system-level package can highly integrate multiple integrated circuits and passive components in a single package body, and is closely connected through an internal wiring network to form a complete-function and collaborative system unit.
[0031] The operand information processing unit 110 is used to receive and store the operation type and operand information sent by the system-level package.
[0032] The dynamic configuration information processing unit 120 is used to determine the data representation form and dynamic bias of each operand according to the operand information in the operand information processing unit 110.
[0033] The data calculation unit 130 is used to perform calculation operations according to the values of the operands, the operation type, the data representation forms of the operands, and the dynamic bias.
[0034] Optionally, the system-level package can receive floating-point calculation instructions from an application or software. The system-level package can first parse the floating-point calculation instructions, obtain the operation type and operand information, and send the operation type and operand information to the mixed-precision tensor calculation module.
[0035] Optionally, the operation type may include addition operation, multiplication operation, exponential calculation, logarithmic operation, etc., and no further examples are given here.
[0036] Among them, the operand information may include left-value information, right-value information, and result information; the left-value information, right-value information, and result information respectively include the mathematical type, data representation form, dynamic offset, tensor pointer base address, and data length of the target operand to which the information belongs.
[0037] Optionally, the data representation form is used to represent the exponent bits and mantissa bits of a floating-point number.
[0038] Optionally, the mathematical type may indicate whether the target operand is a floating-point number.
[0039] Optionally, the dynamic configuration information processing unit 120 may further include a data port, which can be used to read the value of the operand according to the tensor pointer base address and the data length, and save the value of the operand in the cache of the mixed-precision tensor calculation module 100.
[0040] It can be understood that in a computer, a floating-point number is represented by a sign bit, an exponent, and a mantissa. Figure 2 For a schematic diagram of the numerical representation of a floating-point number. As Figure 2 shown, S represents the sign bit of the floating-point number, occupying 1 bit, E represents the exponent bit of the floating-point number, M represents the mantissa bit of the floating-point number. Taking an 8-bit floating-point number FP8 as an example, it can be represented by E4M3 or E5M2. E4M3 or E5M2 is the data representation form of this floating-point number. E4M3 can mean that the exponent bits are 4 and the mantissa bits are 3. Similarly, E5M2 can mean that the exponent bits are 5 and the mantissa bits are 2. This also applies to floating-point numbers with other bit widths such as FP16 and FP32. This is only an exemplary illustration here.
[0041] Optionally, in the floating-point number representation, the offset is a constant value used to adjust the representation range of the exponent part. In the traditional floating-point number representation method, there is usually a fixed offset value in the exponent part. The offset value is generally determined based on the data representation form of the floating-point number and the fixed value 2 (n-1) -1, where n is the number of exponent bits. For example, taking FP8 E4M3 as an example, in the traditional floating-point number representation method, its offset value is 2 (4-1) -1 = 7, and the general representation range of the exponent is [-7, 7], which is mapped to 0-14 in the 4-bit space. If the traditional offset value calculation method is not adopted, for example, the offset value is set to 12, at this time the actual exponent range is [-12, 2]. At this time, compared with the traditional representation method, FP8 can represent floating-point numbers with very small absolute values. It can be seen that by adjusting the offset, the exponent range of the floating-point number can be adjusted, and thus the representation precision of the floating-point number can be adjusted.
[0042] Optionally, the dynamic bias can be the bias value set for the pointer with respect to the floating-point number. For different floating-point number calculation instructions, or different operands in the same floating-point number calculation instruction, the dynamic bias can be set as needed.
[0043] Optionally, in the operand information processing unit 110, different storage subunits are used to store the left-value information, right-value information, and result information respectively. The left value, right value, and result are all operands in the floating-point number calculation. The dynamic configuration information processing unit 120 can be used to parse the operand information in the operand information processing unit 110 to respectively determine the data representation forms and dynamic biases of the left value, right value, and result.
[0044] Figure 3 It is a structural schematic diagram of another optional mixed-precision tensor calculation module. As Figure 3 shown, dbe_config can represent the operand information processing unit 110 in the embodiments of the present invention. Among them, lv_config, rv_config1, and re_config1 can all represent storage subunits and are respectively used to store the left-value information, right-value information, and result information; Figure 3 the db_sem_r_unit in can be used as the dynamic configuration information processing unit 120 in the embodiments of the present invention; Shiftert Unit, Multiple Unit, and add Unit can jointly be used as the data calculation unit in the embodiments of the present invention.
[0045] Optionally, the data calculation unit 130 can read the operand left value and operand right value from the cache of the mixed-precision tensor calculation module 100, determine the target calculation subunit for performing the calculation operation according to the operation type, and perform the calculation operation through the target calculation subunit according to the operand left value, operand right value, the data representation forms of each operand, and the dynamic bias.
[0046] It can be understood that since the data representation form and the dynamic bias can be configured, for a string of binary numbers, the numerical value of the binary numbers can be changed by changing the configuration. For example, for the 8-bit binary number 01001010, if its data representation form is E3M4, then 100 is used as the exponent and 1010 is used as the mantissa. In the traditional floating-point representation method, its bias is 3, and the general representation range of the E3M4 exponent is [-3, 3]. If its bias is set to 5, the actual exponent representation range is [-5, 1]. At this time, a floating-point number with a smaller absolute value can be represented. If its data representation form is E2M5, then 10 is used as the exponent and 01010 is used as the mantissa. In the traditional floating-point representation method, its bias is 1, and the general representation range of the E3M4 exponent is [-1, 1]. If its bias is set to 2, the actual exponent representation range is [-2, 0]. It can be seen that for the same 8-bit floating-point number, by configuring different data representation forms and dynamic biases, the data content that can be represented is different. At the same time, since the data representation form is configurable, that is, the number of bits of the exponent bit and the mantissa bit can be configured, floating-point calculations with different numbers of bits can be supported, not limited to the FP8, FP4, and FP16 examples above. Or when the software needs to calculate floating-point numbers in formats such as FP6 and FP10, support can still be provided. The proposed mixed-precision tensor calculation module in this application can also support the operation of floating-point numbers with these different data forms and dynamic biases. Those skilled in the art have not considered this currently. Designing such a hardware module that can support the operation of floating-point numbers in different data formats and at the same time ensure the data operation efficiency. In the prior art, hardware design generally relies on software requirements, and software depends on hardware performance when constructing processing logic. There is a certain contradiction between software and hardware. However, the proposed mixed-precision tensor calculation module in this invention is not limited to providing calculation functions for floating-point numbers in a specified format, and can provide computing power support for different software requirements, with wide applications.
[0047] In the technical solution of the embodiment of the present invention, by parsing the operand information in the operand processing unit through the dynamic configuration information processing unit, the pre-configured dynamic bias and data representation form can be obtained in the operand information. The dynamic bias can make the exponent of the low-precision floating-point number have a smaller representation range, thereby improving the accuracy of numerical calculation. The low-precision data operation can improve the data operation performance. By configuring the dynamic bias, the calculation accuracy can be ensured while improving the calculation performance. And since both the data representation form and the dynamic bias are configurable, therefore, the floating-point number forms supported by the mixed-precision tensor calculation module are not restricted, and it can support floating-point calculations in any form. In the prior art, hardware design generally depends on software requirements, and software depends on hardware performance when constructing processing logic. There is a certain contradiction between software and hardware. However, the proposed mixed-precision tensor calculation module in this invention is not limited to providing calculation functions for floating-point numbers in a specified format, and can provide computing power support for different software requirements, with wide applications.
[0048] Embodiment 2
[0049] Figure 4 This is a schematic structural diagram of another mixed - precision tensor calculation module provided in Embodiment 2 of the present invention. On the basis of the above - mentioned embodiment, this embodiment specifically describes the mixed - precision tensor calculation module. As Figure 4 shown, the operand information processing unit includes a left - value storage subunit 111, a right - value storage subunit 112, and a result storage subunit 113; the data calculation unit 130 includes a shifter subunit 131, an addition calculation subunit 132, and a multiplication calculation subunit 133; the mixed - precision tensor calculation module further includes an output processing unit 140.
[0050] Optionally, the data calculation unit 130 may include at least one calculation subunit. The calculation subunit includes an addition calculation subunit 132 and a multiplication calculation subunit 133, and may also include other types of calculation subunits, which are not limited herein.
[0051] As Figure 3 shown, Shiftert Unit may represent the shifter subunit 131, Multiple Unit may represent the multiplication calculation subunit 133, add Unit may represent the addition calculation subunit 132, and db_sem_w_unit may serve as the output processing unit 140 in the embodiments of the present invention.
[0052] Among them, the dynamic configuration information processing unit 120 is further configured to:
[0053] Verify the dynamic biases of each operand respectively according to the data representation forms and dynamic biases of each operand;
[0054] When it is determined according to the verification result that the dynamic biases of each operand meet the configuration requirements, call the data calculation unit and send the data representation forms and dynamic biases of each operand to the data calculation unit.
[0055] Optionally, the dynamic configuration information processing unit 120 may be specifically configured to: judge whether the dynamic biases of the left - value, right - value, and result have configuration values, and, according to the representation forms and dynamic biases of each operand, judge whether there is an out - of - range exponent, and when neither exists, determine that the configuration requirements are met.
[0056] In an optional example, if the right - value is a floating - point number of FP8 E4M3 but no dynamic bias is configured, at this time, the standard bias 7 of E4M3 can be configured as the dynamic bias of the right - value.
[0057] Among them, the target calculation subunit is used to obtain a first floating-point number and a second floating-point number, and perform a floating-point calculation operation according to the first floating-point number, the second floating-point number, the data representation forms of the operands, and the dynamic bias.
[0058] Optionally, the target calculation subunit may refer to the addition calculation subunit 132 or the multiplication calculation subunit 133. When the operation type is an addition operation, the target calculation subunit is the addition calculation subunit 132. When the operation type is a multiplication operation, the target calculation subunit is the multiplication calculation subunit 133.
[0059] Among them, the shift subunit 131 is used to perform a shift operation in response to a call from the target calculation subunit during the calculation process of the target calculation subunit.
[0060] Optionally, the first floating-point number is the left value of the operand, and the second floating-point number is the right value of the operand. The target calculation subunit can read the first floating-point number and the second floating-point number from the cache of the mixed-precision tensor calculation module. Taking an 8-bit floating-point number as an example, the stored first floating-point number can be 00110101, and the second floating-point number can be 01001010. Furthermore, according to the data representation form, the specific exponent and mantissa can be determined.
[0061] Optionally, the addition calculation subunit 132 can specifically be used for:
[0062] Determine the actual exponent of the left value and the actual exponent of the right value according to the data representation forms of the first floating-point number, the second floating-point number, the left value, and the right value, and the dynamic bias, and determine the aligned actual exponent according to the actual exponent of the left value and the actual exponent of the right value;
[0063] After determining to perform a restore exponent operation on the target floating-point number, call the shift subunit to perform alignment processing on the mantissa of the target floating-point number;
[0064] Perform a mantissa addition operation according to the mantissas of the aligned floating-point numbers;
[0065] When it is determined according to the data representation form of the result that the mantissa addition result overflows or is less than a preset value, call the shift subunit to perform normalization processing on the mantissa addition result, and update the actual exponent according to the aligned actual exponent and the normalization result;
[0066] Use the normalized mantissa addition result, the updated actual exponent, and the sign bit as the calculation result and send it to the output processing unit.
[0067] In a specific example, the first floating-point number representing the left value is 00110101, the data representation form of the left value is E2M5, and the offset is set to 1. At this time, the sign bit is 0, the exponent bit is 01, and the mantissa bit is 10101. According to the 1.M representation method, the implicit mantissa is , the actual exponent is the decimal value of the exponent bit minus the offset, that is, restoring the exponent based on dynamic biasing. The actual exponent is 1 - 1 = 0. After converting the first floating-point number to decimal, it is ; the second floating-point number representing the right value is 01001010, the data representation form of the right value is E3M4, and the offset is set to 3. At this time, the sign bit is 0, the exponent bit is 100, and the mantissa bit is 1010. According to the 1.M representation method, the implicit mantissa is , the actual exponent is the decimal value of the exponent bit minus the offset, that is, the actual exponent is 4 - 3 = 1. After converting the second floating-point number to decimal, it is , in this example, it is assumed that the data representation form of the result is E2M5 and the offset is set to 1.
[0068] Furthermore, since the actual exponents of the two numbers are different, it is necessary to adjust the first floating-point number with the smaller exponent to be the same as the actual exponent of the second floating-point number, that is, adjust the actual exponent 0 of the first floating-point number to 1, and the aligned actual exponent is 1; and since it is necessary to shift the first floating-point number to the right to align the exponents, at this time, it is necessary to call the shift sub-unit for mantissa alignment.
[0069] Optionally, the shift sub-unit 131, during the process of being called to perform bit alignment, is used to shift the mantissa to the right according to the number of bits moved for exponent alignment. Continuing the previous example, shift the mantissa 1.10101 of the first floating-point number to the right, that is , at this time, the mantissa representation of the first floating-point number is , and the mantissa of the second floating-point number can be correspondingly extended to .
[0070] Furthermore, after the mantissas are aligned, add the mantissas of the first floating-point number and the second floating-point number. The mantissa of the first floating-point number is , and its corresponding decimal is 0.828125. The mantissa of the second floating-point number is , and its corresponding decimal is 1.625. Adding them together gives 0.828125 + 1.625 = 2.453125. Since the added mantissa 2.453125 ≥ 2, for decimal numbers greater than 2 in binary, at least 2 binary digits are required for representation, and the 1.M representation method only supports 1 binary digit before the decimal point. At this time, it is determined that the mantissa overflows. Therefore, it is necessary to call the shift sub-unit 131 to perform normalization processing on the result of the mantissa addition.
[0071] Optionally, during the process of being called to perform the normalization process on the result of the mantissa addition, the shifter unit 131 is used to shift the result of the mantissa addition to the right. Continuing with the previous example, the result of the mantissa addition needs to be shifted to the right by one bit, that is, divide the decimal result of the mantissa addition by 2, 2.453125 ÷ 2 = 1.2265625, and the corresponding binary form is: .
[0072] Furthermore, since the mantissa is normalized by shifting it to the right by one bit, at this time, 1 needs to be added to the aligned actual exponent, and the updated actual exponent is 2. Therefore, the calculation result is sign bit 0 (positive), actual exponent 2, and actual mantissa .
[0073] Among them, the output processing unit 140 is used for:
[0074] After obtaining the calculation result sent by the data calculation unit, when it is determined according to the result information that the calculation result meets the normalization processing condition, perform the normalization processing operation;
[0075] Determine the return value according to the calculation result after the normalization processing operation and the result information, and return the return value to the system-level package.
[0076] Continuing with the previous example, since the data representation form of the result is E2M5, at this time, the number of bits of the mantissa in the calculation result is different from the number of bits of the mantissa in the result representation form, and it meets the normalization processing condition. Then, the mantissa needs to retain 5 bits after the decimal point. At this time, the output processing unit 140 performs the normalization processing operation and takes the mantissa .
[0077] Furthermore, the output processing unit 140 combines and generates the return value according to the normalized mantissa and the remaining result information. Continuing with the previous example, according to the data representation form of the result being E2M5 and the offset being set to 1, at this time, the exponent in the return value is the actual exponent plus the offset, that is, the exponent in the return value is 2 + 1 = 3, and the binary representation is 11. The normalized mantissa takes 00111 after the decimal point, and the combined output value is 0 11 00111.
[0078] Optionally, the multiplication calculation subunit 133 can be specifically used for:
[0079] Determine the sign bit of the output value according to the first floating-point number and the second floating-point number;
[0080] Determine the actual exponent and mantissa of the left value, the actual exponent and mantissa of the right value according to the data representation forms of the first floating-point number, the second floating-point number, the left value, the right value, and the dynamic bias;
[0081] Determine the exponent calculation result according to the actual exponent of the left value and the actual exponent of the right value, and perform a mantissa multiplication operation according to the mantissa of the left value and the mantissa of the right value to obtain the mantissa multiplication result;
[0082] When it is determined according to the data representation form of the result that the mantissa multiplication result overflows or is less than the preset value, call the shift sub-unit to perform normalization processing on the mantissa multiplication result, and determine the actual exponent of the result according to the exponent calculation result and the normalization result;
[0083] Take the mantissa multiplication result after normalization processing, the actual exponent of the result, and the sign bit of the output value as the calculation result together, and send it to the output processing unit.
[0084] Still taking the above-mentioned first floating-point number and second floating-point number as an example, the first floating-point number is 00110101, the data representation form of the left value is E2M5, the offset is set to 1, at this time, the sign bit is 0, the exponent bit is 01, the mantissa bit is 10101, and the implicit mantissa is , the actual exponent of the left value is 0; the second floating-point number is 01001010, the data representation form of the right value is E3M4, the offset is set to 3, at this time, the sign bit is 0, the exponent bit is 100, the mantissa bit is 1010, and the implicit mantissa is , the actual exponent of the right value is 1, and the data representation form of the result is E2M5, and the offset is set to 1.
[0085] Furthermore, perform an exclusive OR calculation on the sign bits, 0⊕0 = 0, and the sign bit of the output value is 0; add the actual exponent of the left value and the actual exponent of the right value to determine the exponent calculation result 0 + 1 = 1, and the mantissa multiplication result is , at this time, it is necessary to call the shift sub-unit to shift the mantissa two bits to the right to obtain the normalized mantissa , since the mantissa is shifted two bits to the right, the actual exponent of the result is the exponent calculation result plus 2, that is, the actual exponent of the result is 1 + 2 = 3. Therefore, the calculation result is the sign bit 0, the actual exponent of the result 3, and the mantissa multiplication result after normalization processing .
[0086] Continuing the previous example, since the data representation form of the result is E2M5, at this time, the mantissa bits in the calculation result are different from the mantissa bits in the representation form of the result, and the normalization processing condition is satisfied. Then, the mantissa should retain 5 bits after the decimal point. At this time, the output processing unit 140 performs the normalization processing operation and takes the mantissa .
[0087] Further, the output processing unit 140 combines and generates a return value based on the normalized mantissa and the remaining result information. Continuing with the previous example, assuming the data representation form of the result is E2M5 and the offset is set to 1, the exponent in the return value is the actual exponent plus the offset. That is, the exponent in the return value is 3 + 1 = 4, and its binary representation is 100. However, the exponent bit of the result only has two digits and will overflow. At this time, the output processing unit is called to take the maximum representable value 11, and the normalized mantissa takes 01001 after the decimal point. The combined output value is 0 11 01001.
[0088] In the technical solution of the embodiment of the present invention, by parsing the operand information in the operand processing unit through the dynamic configuration information processing unit, the dynamically configured bias and data representation form pre-configured in the operand information can be obtained. The dynamically configured bias can make the exponent of low-precision floating-point numbers have a smaller representation range, thereby improving the accuracy of numerical calculations. Low-precision data operations can improve data operation performance. By configuring the dynamic bias, the calculation accuracy can be ensured while improving the calculation performance. Moreover, since both the data representation form and the dynamic bias are configurable, the floating-point number forms supported by the mixed-precision tensor calculation module are not restricted, and it can support floating-point calculations in any form. In the prior art, hardware design generally depends on software requirements, and software depends on hardware performance when constructing processing logic, resulting in a certain contradiction between software and hardware. However, the mixed-precision tensor calculation module proposed in the present invention is not limited to providing calculation functions for floating-point numbers in a specified format and can provide computing power support for different software requirements, with wide applications.
[0089] Embodiment III
[0090] Figure 5 is a schematic structural diagram of a system-in-package provided for Embodiment III of the present invention. As Figure 5 shown, the system-in-package 200 includes a mixed-precision tensor calculation module 100 and at least one special calculation module 300.
[0091] Among them, the calculation type of the special calculation module 300 is different from that of the mixed-precision tensor calculation module 100.
[0092] Among them, the special calculation module 300 is configured to perform a calculation operation according to the operation type and operand information sent by the mixed-precision tensor calculation module 100 after responding to the call of the mixed-precision tensor calculation module 100.
[0093] Optionally, when the operation type is different from the calculation type of the mixed-precision tensor calculation module 100, the operand information will be forwarded to the special calculation module 300, and the special calculation module 300 will perform a special calculation operation according to the operand information.
[0094] Among them, the system - level package 200 is used for:
[0095] Obtain a floating - point calculation instruction, and according to the floating - point calculation instruction, parse to obtain the operation type and operand information.
[0096] Among them, the system - level package 200 is specifically used for:
[0097] Parse the floating - point calculation instruction, and according to the parsing result, judge whether there is a dynamic bias configuration with local tensor differentiation in the floating - point calculation instruction;
[0098] If so, construct the dynamic bias of the first operand according to the dynamic bias configuration with local tensor differentiation, and construct the dynamic bias of the second operand according to the global dynamic bias configuration;
[0099] If not, construct the dynamic bias of each operand according to the global dynamic bias configuration.
[0100] Optionally, judging whether there is a dynamic bias configuration with local tensor differentiation in the floating - point calculation instruction can refer to judging whether the left - hand value, right - hand value, and result in the floating - point calculation instruction are all configured with dynamic biases. The first operand can refer to the operand corresponding to the dynamic bias configuration with local tensor differentiation. For example, if the dynamic bias configuration with local tensor differentiation includes the dynamic bias configuration of the left - hand value, then the first operand is the left - hand value. The second operand refers to the operand without the dynamic bias configuration with local tensor differentiation. When the second operand is not configured with a dynamic bias, then according to the global dynamic bias configuration and the data representation form of the second operand, determine the dynamic bias of the second operand. The global dynamic bias stores the standard bias values of each data representation form. For example, for FP8 E3M4, the corresponding stored standard bias value is 3, and for FP8E5M2, the corresponding stored standard bias value is 15. According to the data representation form of the second operand, the corresponding standard bias value can be directly determined in the global dynamic bias, and the dynamic bias of the second operand is configured according to the standard bias value.
[0101] Among them, the system - level package 200 is also used for:
[0102] According to the operation type, determine the current calculation type, and according to the current calculation type, judge whether the mixed - precision tensor calculation module can perform a calculation operation;
[0103] If so, send the operation type and operand information to the mixed - precision tensor calculation module; if not, according to the current calculation type, call the target special calculation module that meets the calculation requirements to perform the calculation operation.
[0104] Optionally, the current calculation type may refer to the calculation type specified in the floating-point calculation instruction. If the current calculation type can be executed by any calculation subunit in the data calculation unit 130, it is determined that the mixed-precision tensor calculation module can perform the calculation operation. When the data calculation unit 130 includes an addition calculation subunit 132 and a multiplication calculation subunit 133, when the calculation type belongs to addition calculation or multiplication calculation, it is determined that the mixed-precision tensor calculation module can perform the calculation operation. When the calculation type belongs to addition calculation, the operation type and operand information are sent to the addition calculation subunit 132. When the calculation type belongs to multiplication calculation, the operation type and operand information are sent to the multiplication calculation subunit 133.
[0105] It can be understood that subtraction calculations can be merged into the addition calculation subunit 132 for execution because subtraction calculations can be converted into adding a negative value. Similarly, division calculations can be converted into multiplying by a reciprocal, so division calculations can be merged into the multiplication calculation subunit 133 for execution.
[0106] Optionally, the special calculation module 300 may refer to a module with special calculation functions. For example, it performs trigonometric function operations, logarithmic operations, remainder operations, etc. The target special calculation module can be configured in the same system-level package 200 as the mixed-precision tensor calculation module 100. When the system-level package 200 determines that the mixed-precision tensor calculation module 100 cannot perform the calculation operation, among the target special calculation modules that can execute the current calculation type, the target special calculation module is a module that can perform floating-point calculations corresponding to the current calculation type.
[0107] The technical solution of the embodiments of the present invention can provide a system-level package that supports any form of floating-point operation by setting a mixed-precision tensor calculation module and at least one special calculation module in the system-level package, and is not limited to the specified calculation type, but supports multiple types of operations. By configuring the dynamic bias of the operands through the system-level package, the data integrity and calculation accuracy during the floating-point calculation process can be ensured.
[0108] Embodiment 4
[0109] Figure 6 It is a flowchart of a mixed-precision tensor calculation method provided by Embodiment 4 of the present invention. This embodiment is applicable to the situation of implementing floating-point calculations with different precisions. This method can be executed by a computer or a processor configured with a system-level package as described in the embodiments of the present invention. As Figure 6 shown, the method includes:
[0110] S410. Generate a floating-point calculation instruction. During the compilation of the floating-point instruction, verify the dynamic bias of each operand, and send the floating-point instruction to the system-level package.
[0111] Optionally, the dynamic bias of each operand in the floating-point instruction can be configured by the user during the process of generating the floating-point calculation instruction. During the instruction compilation process, software verification of the dynamic bias of each operand is performed first, which can effectively prevent exponent out-of-bounds.
[0112] S420. Through system-level packaging, based on the floating-point calculation instruction, parse to obtain the operation type and operand information, and send the operation type and operand information to the mixed-precision tensor calculation module.
[0113] Optionally, the operand information includes left-value information, right-value information, and result information;
[0114] The left-value information, right-value information, and result information respectively include the mathematical type, data representation form, dynamic bias, tensor pointer base address, and data length of the target operand to which the information belongs;
[0115] The data representation form is used to represent the exponent bits and mantissa bits of the floating-point number.
[0116] Optionally, through system-level packaging, based on the floating-point calculation instruction, parse to obtain the operation type and operand information, and send the operation type and operand information to the mixed-precision tensor calculation module, which may include:
[0117] Parse the floating-point calculation instruction, and judge whether there is a dynamically biased configuration with local tensor differentiation in the floating-point calculation instruction according to the parsing result;
[0118] If so, construct the dynamic bias of the first operand according to the dynamically biased configuration with local tensor differentiation, and construct the dynamic bias of the second operand according to the global dynamic bias configuration;
[0119] If not, construct the dynamic bias of each operand according to the global dynamic bias configuration.
[0120] Optionally, when the operation type is addition or multiplication, send the operation type and operand information to the mixed-precision tensor calculation module.
[0121] S430. Through the mixed-precision tensor calculation module, perform the calculation operation of the mixed-precision tensor according to the operation type and operand information.
[0122] Optionally, through the mixed-precision tensor calculation module, performing the calculation operation of the mixed-precision tensor according to the operation type and operand information may include:
[0123] Through the operand information processing unit in the mixed-precision tensor calculation module, receive and store the operation type and operand information sent by the system-level packaging;
[0124] Through the dynamic configuration information processing unit in the mixed-precision tensor calculation module, determine the data representation form and dynamic bias of each operand according to the operand information in the operand information processing unit.
[0125] Through the data calculation unit in the mixed-precision tensor calculation module, perform a calculation operation according to the value of the operand, the operation type, the data representation form of each operand, and the dynamic bias.
[0126] The technical solution of the embodiment of the present invention generates a floating-point calculation instruction. During the compilation of the floating-point instruction, the dynamic bias of each operand is verified, and the floating-point instruction is sent to the system-in-package. Through the system-in-package, the operation type and operand information are parsed according to the floating-point calculation instruction, and the operation type and operand information are sent to the mixed-precision tensor calculation module. Through the mixed-precision tensor calculation module, according to the operation type and operand information, the way of performing the mixed-precision tensor calculation operation can support floating-point operations of different precisions and forms, can provide computing power support for different software requirements, and has a wide range of applications.
[0127] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0128] The above specific implementation manners do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A mixed-precision tensor calculation module, characterized in that, The mixed-precision tensor calculation module is configured within a system-in-package (SiP), and includes an operand information processing unit, a dynamic configuration information processing unit, and a data calculation unit. Among them, The operand information processing unit is used to receive and store the operation type and operand information sent by the system-in-package. The dynamic configuration information processing unit is used to determine the data representation form and dynamic bias of each operand according to the operand information in the operand information processing unit. The data calculation unit is used to perform a calculation operation according to the values of the operands, the operation type, the data representation forms of the operands, and the dynamic bias.
2. The computing module according to claim 1, wherein The operand information includes left value information, right value information, and result information. The left value information, right value information, and result information respectively include the mathematical type, data representation form, dynamic bias, tensor pointer base address, and data length of the target operand to which the information belongs. The data representation form is used to represent the exponent bits and mantissa bits of a floating-point number.
3. The calculation module according to claim 1, characterized in that, The dynamic configuration information processing unit is further used for: Checking the dynamic bias of each operand respectively according to the data representation form and dynamic bias of each operand. When it is determined according to the check result that the dynamic bias of each operand meets the configuration requirements, the data calculation unit is called, and the data representation form and dynamic bias of each operand are sent to the data calculation unit.
4. The calculation module according to claim 2, wherein The data calculation unit includes a shift subunit and at least one calculation subunit; the calculation subunit includes an addition calculation subunit and a multiplication calculation subunit. Among them, The target calculation subunit is used to obtain a first floating-point number and a second floating-point number, and perform a floating-point calculation operation according to the first floating-point number, the second floating-point number, the data representation forms of the operands, and the dynamic bias. The shift subunit is used to perform a shift operation in response to the call of the target calculation subunit during the calculation process of the target calculation subunit.
5. The computing module according to claim 2, wherein It further includes an output processing unit, which is used for: After obtaining the calculation result sent by the data calculation unit, when it is determined according to the result information that the calculation result meets the normalization processing condition, performing a normalization processing operation. Determining a return value according to the calculation result after the normalization processing operation and the result information, and returning the return value to the system-in-package.
6. A system-in-package, characterized in that, It includes the mixed-precision tensor calculation module according to any one of claims 1-5 and at least one special calculation module. The calculation type of the special calculation module is different from that of the mixed-precision tensor calculation module. The special calculation module is used to perform a calculation operation according to the operation type and operand information sent by the mixed-precision tensor calculation module after responding to the call of the mixed-precision tensor calculation module.
7. The system-level package according to claim 6, characterized in that, The system-in-package is used for: Obtaining a floating-point calculation instruction, and parsing the operation type and operand information according to the floating-point calculation instruction.
8. The system-level package according to claim 7, wherein Specifically, the system-in-package is used for: Parsing the floating-point calculation instruction, and judging whether there is a dynamic bias configuration with local tensor differentiation in the floating-point calculation instruction according to the parsing result. If so, construct the dynamic bias of the first operand according to the dynamic bias configuration differentiated by local tensors, and construct the dynamic bias of the second operand according to the global dynamic bias configuration; If not, construct the dynamic bias of each operand according to the global dynamic bias configuration.
9. The system-level package according to claim 7, wherein The system-level package is also used for: Determine the current calculation type according to the operation type, and judge whether the mixed-precision tensor calculation module can perform a calculation operation according to the current calculation type; If so, send the operation type and operand information to the mixed-precision tensor calculation module; if not, call the target special calculation module that meets the calculation requirements to perform the calculation operation according to the current calculation type.
10. A mixed-precision tensor calculation method, characterized in that, It includes: Generate a floating-point calculation instruction. During the compilation of the floating-point instruction, check the dynamic bias of each operand, and send the floating-point instruction to the system-level package; Through the system-level package, parse the operation type and operand information according to the floating-point calculation instruction, and send the operation type and operand information to the mixed-precision tensor calculation module; Through the mixed-precision tensor calculation module, perform the mixed-precision tensor calculation operation according to the operation type and operand information.