Operating method and arithmetic unit based on Point data format
By preprocessing and encoding the sign bit, exponent extension bit, mantissa bit and exponent bit of the Posit data format, the parallel execution of mantissa multiplication and exponent extension bit is achieved, which solves the problems of high computing energy consumption and large latency in AI computing chips and improves computing efficiency and adaptability.
Patent Information
- Application Number
- CN202511120228.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-12
AI Technical Summary
The computational circuits based on the Posit data format in AI computing chips have problems of high computing energy consumption and large computing delay, which makes it difficult to meet the low power consumption and real-time requirements of edge devices.
An operation method based on the Posit data format is adopted. By preprocessing the sign bit, exponent extension bit, mantissa bit and exponent bit, the bit width allocation priority rule is defined, and the data is encoded in the one's complement form. The parallel execution of mantissa multiplication and exponent extension bit is realized, the shift logic operation is reduced, and the number is directly converted into a fixed-point number for addition operation.
It reduces the latency and energy consumption of computing circuits, improves computing efficiency, and adapts to the low power consumption and real-time requirements of edge devices.
Smart Images

Figure CN120631440A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an operation method and an operator based on the Posit data format. Background Art
[0002] Artificial intelligence technologies, represented by deep neural networks, have been applied on a large scale in image recognition, natural language processing, autonomous driving and other fields. Their core relies on the efficient numerical computing capabilities of the computational circuits in AI computing chips. However, the inherent computational intensity and memory intensity of deep learning models lead to two key problems for the computational circuits of AI computing chips: first, the computing energy consumption is too high, making it difficult to meet the low power consumption requirements of edge devices; second, the computing delay is large, making it impossible to adapt to application scenarios with high real-time requirements.
[0003] In AI computing, the Posit data format, as a new data format, can achieve a dynamic range close to that of traditional floating-point numbers at a smaller bit width by dynamically allocating the bit width of the exponent and mantissa. Its conical data distribution is more compatible with the distribution characteristics of deep learning model parameters and data, and has significant potential in deep learning computing. However, the dynamic bit width characteristics of the Posit data format require real-time parsing of the boundaries of the exponent and mantissa, requiring additional logic modules, such as bit width detection circuits, to complete the extraction and verification of components, increasing the hardware circuit area and redundant operations, which directly leads to increased energy consumption. On the other hand, the mantissa multiplication operation must be started after the complete decoding of the exponent and mantissa is completed, resulting in increased calculation delay, which seriously restricts the practical application of the operation circuits in the Posit data format AI computing chip. At this stage, there is a need for an operation method and operator based on the Posit data format. Summary of the Invention
[0004] In order to solve the problems of high computing energy consumption and computing delay of the computing circuit in the AI computing chip, the present invention provides an operation method and an operator based on the Posit data format.
[0005] In a first aspect, the present invention provides an operation method based on the Posit data format, which adopts the following technical solution: A calculation method based on the Posit data format, comprising: Obtain the sign bit, exponent extension bit, mantissa bit, and exponent bit of the Posit data format, and preprocess each position of the Posit data format to obtain SR-Posit data; Define the bit width allocation priority rule based on the pre-processed sign bit, exponent extension bit, mantissa bit and exponent bit, and encode the SR-Posit data using the one's complement form; Use two encoded SR-Posit data as input data, perform decoding operation on the input data, and multiply the two decoded mantissa bits to obtain the mantissa product and mask sequence; determining the number of leading zeros of the mantissa product according to the mask sequence, and performing a uniform logical shift on the mantissa product based on the number of exponent extension bits of the product and the number of leading zeros of the mantissa product; The shifted mantissa is rounded, the mantissa bits are concatenated with the equivalent exponent bits, and then combined with the sign bit to obtain the product result of the SR-Posit data; The mantissa product is converted into a fixed-point format according to the sign bit, the exponent extension bit, and the exponent bit, and an addition operation is performed based on the fixed-point number to output the accumulated result.
[0006] Furthermore, the preprocessing of each position of the Posit data format includes setting the sign bit width of the Posit data format to 1-bit, the exponent extension bit is composed of a continuous digital sequence, and the end of the sequence is marked with a binary bit with a width of 1-bit and opposite to the sequence value, the mantissa bit is coded in a normalized manner, and the highest bit 1 is hidden, the exponent bit is coded with a fixed value and is encoded as an unsigned integer, and is sorted in the order of sign bit, exponent extension bit, mantissa bit and exponent bit.
[0007] Furthermore, the definition of the bit width allocation priority rule includes removing the sign bit from the total bit width to obtain the total available bit width, and first allocating the total available bit width to the exponent extension bit. The bit width requirement of the exponent extension bit includes the bit width occupied by the sequence itself and the end flag at the end of the sequence. The end flag bit is 1 bit and is opposite to the value of the continuous sequence. When the length of the continuous sequence is equal to the total available bit width, the bit width of the exponent extension bit is the length of the continuous sequence, and the end flag bit does not need to be allocated. When there is still remaining bit width after the total available bit width is allocated to the exponent extension bit, the remaining bit width is first allocated to the exponent bit, and then the bit width remaining after the allocation to the exponent bit is allocated to the mantissa bit.
[0008] Furthermore, the decoding operation on the input data includes receiving two encoded SR-Posit data IN1 and IN2, removing the sign bits of IN1 and IN2, calculating the number of leading zeros r1 and r2 of the remaining data through a leading zero counter, and determining the exponent extension bits Rg1 and Rg2 according to r1 and r2 respectively, extracting the fixed bit width bits of the lowest exponent bit after the end bits of Rg1 and Rg2 as exponent bits E1 and E2, removing the sign bit, exponent extension bit and exponent bit of IN1 and IN2 respectively, and using the remaining bits as mantissa bits F1 and F2, using the 1 of the end bits of Rg1 and Rg2 as a hidden bit, and splicing them with F1 and F2 respectively to obtain a complete valid mantissa; Among them, E1 represents the exponent bit of input data IN1, E2 represents the exponent bit of IN2, F1 represents the mantissa bit of IN1, F2 represents the mantissa bit of IN2, r1 and r2 represent the number of leading zeros of IN1 and IN2 respectively, and Rg1 and Rg2 represent the exponent extension bits of IN1 and IN2 respectively.
[0009] Furthermore, the two mantissa bits obtained by decoding are multiplied, including performing a multiplication operation on the complete valid mantissa to obtain a mantissa product, generating a correction vector according to the bit width occupied by the exponent extension bits in IN1 and IN2, performing a logical OR operation on the bits of the mantissa product with a lower total bit width and Pcorr to correct the unmantissa product, and finally generating a mask sequence according to the exponent extension length in IN1 and IN2.
[0010] Furthermore, determining the number of leading zeros of the mantissa product based on the mask sequence includes performing a bitwise logical AND operation on the mantissa product and the mask sequence to generate an intermediate result vector, performing a reduction logical OR operation on all bits of the intermediate result vector to obtain an overflow flag, and calculating the number of leading zeros of the mantissa product based on the overflow flag.
[0011] Furthermore, the mantissa product is subjected to a unified logical shift, including respectively concatenating the exponent extension bits and the exponent bits of IN1 and IN2 to obtain equivalent exponent bits, calculating the number of exponent extension bits of the product based on the equivalent exponent bits, adding the number of exponent extension bits of the product and the number of leading zeros to obtain a displacement, performing a logical left shift operation on the mantissa product according to the shift amount, moving the effective numerical part to the high bit and the leading zero sequence to the low bit, and using the leading zero sequence after the shift as the basis for the exponent extension bits of the final product, and outputting the normalized mantissa product, the number of leading zeros and the logical shift amount.
[0012] Furthermore, the product result of the SR-Posit data is obtained, including calculating the exponent extension bits of the product and the exponent bits of the product based on the equivalent exponent bits of IN1 and IN2, and sequentially splicing the exponent extension bits of the product, the rounded significant mantissa, and the exponent bits of the product in order from the most significant bit to the least significant bit to obtain a bit sequence without a sign bit, performing a splicing operation on the XOR result of the sign bits of IN1 and IN2 and the bit sequence to form a complete bit sequence and performing final encoding adjustment, and outputting the product result.
[0013] Furthermore, the addition operation based on fixed-point numbers includes converting the product into a fixed-point format of an exact accumulator based on the mantissa product, sign bit, exponent extension bit and exponent bit, placing the mantissa product to the right end of the lowest bit of the exact accumulator, sign-extending the mantissa product according to the XOR result of the sign bits of IN1 and IN2, performing a logical left shift on the mantissa product after sign extension according to the displacement amount to obtain a product in a fixed-point format, adding the product in the fixed-point format to the fixed-point number input externally, and outputting the final accumulated result.
[0014] In a second aspect, an operator based on the Posit data format includes: The data acquisition module is configured to: acquire a sign bit, an exponent extension bit, a mantissa bit, and an exponent bit of a Posit data format, and pre-process each position of the Posit data format to obtain SR-Posit data; An encoding module is configured to: define a bit width allocation priority rule based on the pre-processed sign bit, exponent extension bit, mantissa bit, and exponent bit, and encode the SR-Posit data using a one's complement form; The mantissa multiplication module is configured to: use two encoded SR-Posit data as input data, perform a decoding operation on the input data, and multiply the two decoded mantissa bits to obtain a mantissa product and a mask sequence; a normalization module configured to: determine a number of leading zeros of a mantissa product according to a mask sequence, and perform a uniform logical shift on the mantissa product based on the number of exponent extension bits of the product and the number of leading zeros of the mantissa product; The product encoding module is configured to: perform a rounding operation on the shifted mantissa, concatenate the mantissa bits with the equivalent exponent bits, and combine them with the sign bit to obtain the product result of the SR-Posit data; The accumulation module is configured to: convert the mantissa product into a fixed-point number format according to the sign bit, the exponent extension bit and the exponent bit, perform addition operation based on the fixed-point number, and output the accumulation result.
[0015] In summary, the present invention has the following beneficial technical effects: 1. The present invention places the exponent at the least significant bit, so that the exponent extension bit is directly connected to the mantissa bit. The mantissa bit can be extracted after completing a simple logical inversion, thereby achieving parallel execution of mantissa multiplication and exponent extension bit decoding, and solving the delay problem in the traditional Posit format where the mantissa multiplication must wait for complete decoding before starting.
[0016] 2. The mantissa product of the present invention performs a unified logical shift only once in the encoding stage, replacing the redundant operations of shifting in the decoding and encoding stages in the traditional method, reducing the number of operation layers of the shift logic and lowering the critical path delay.
[0017] 3. The present invention directly performs addition operations after converting the mantissa product into fixed-point numbers, eliminating the need for exponent alignment operations required for floating-point number accumulation, thereby further shortening the delay of the accumulation path.
[0018] 4. The present invention adopts one's complement encoding, and input and output processing only needs to complete sign-related operations through logical inversion, avoiding the adder required for the addition operation in the two's complement solution process, reducing circuit area and power consumption. The rounding operation is realized by a combinational circuit to implement simple rounding judgment bit logic, without the need for complex timing control, reducing logic resource consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a schematic diagram of the multiplication operation logic in an operation method based on the Posit data format in Example 1 of the present invention.
[0020] Figure 2 This is a schematic diagram of the data format in a calculation method based on the Posit data format in Example 1 of the present invention.
[0021] Figure 3 This is a schematic diagram of the accumulator format in an operation method based on the Posit data format in Example 1 of the present invention.
[0022] Figure 4 This is a schematic diagram of multiplication and accumulation logic in a calculation method based on the Posit data format in Example 1 of the present invention.
[0023] Figure 5 This is a schematic diagram of a logical shift operation in an operation method based on the Posit data format according to Example 1 of the present invention. DETAILED DESCRIPTION
[0024] The present invention will be further described in detail below with reference to the accompanying drawings.
[0025] Example 1 Reference Figure 1 , a calculation method based on the Posit data format of this embodiment includes: Obtain the sign bit, exponent extension bit, mantissa bit, and exponent bit of the Posit data format, and preprocess each position of the Posit data format to obtain SR-Posit data; Define the bit width allocation priority rule based on the pre-processed sign bit, exponent extension bit, mantissa bit and exponent bit, and encode the SR-Posit data using the one's complement form; Use two encoded SR-Posit data as input data, perform decoding operation on the input data, and multiply the two decoded mantissa bits to obtain the mantissa product and mask sequence; determining the number of leading zeros of the mantissa product according to the mask sequence, and performing a uniform logical shift on the mantissa product based on the number of exponent extension bits of the product and the number of leading zeros of the mantissa product; The shifted mantissa is rounded, the mantissa bits are concatenated with the equivalent exponent bits, and then combined with the sign bit to obtain the product result of the SR-Posit data; The mantissa product is converted into a fixed-point format according to the sign bit, the exponent extension bit, and the exponent bit, and an addition operation is performed based on the fixed-point number to output the accumulated result.
[0026] Specifically, a calculation method based on the Posit data format includes the following steps: like Figure 2 As shown in the figure, first, the sign bit, exponent extension bit, mantissa bit and the basic bit segment that can be used as the exponent bit are extracted from the original Posit data, wherein the sign bit is taken from the highest bit of the original Posit data, the exponent extension bit corresponds to the dynamic part of the original Posit used to represent the exponent range, the mantissa bit corresponds to the part of the original Posit that represents the numerical precision, and the exponent bit is separated from the exponent extension bit to form a fixed bit width part. Then, the extracted components are preprocessed: the sign bit is set to 1-bit width, 0 represents a positive number, 1 represents a negative number, and is placed in the most significant bit of the SR-Posit format.
[0027] The exponent extension bit consists of a continuous sequence of 0s or 1s, with a sequence length of r. To clearly define the boundary between the exponent extension bit and the subsequent bit segment, a 1-bit end flag is added to the end of the sequence, where 1-bit represents a binary bit. The end flag is the opposite of the sequence value. For example, the end flag of a sequence of 0s is 1. If the sequence length r already occupies the total available bit width, where the total available bit width is the total bit width nb minus the sign bit nb-1 bits, then there is no need to add an end flag. At the same time, the Rg value corresponding to the 0 sequence is defined as -r, and the Rg value corresponding to the 1 sequence is defined as r-1. The exponent extension bit is placed after the sign bit, and the mantissa bits are normalized. The most significant bit 1 of all significant mantissa bits is implicitly stored to save bit width and improve precision. Its bit width is the remaining bit width after subtracting the exponent extension bit (including the end flag) and the exponent bit from the total available bit width, and is placed after the exponent extension bit.
[0028] The exponent bit uses a fixed bit width (denoted as es, such as 2-bit) and is encoded as an unsigned integer to represent a small range of exponent adjustments. It is placed after the mantissa bit and serves as the least significant bit of the SR-Posit data. Subsequently, the preprocessed bit segments are concatenated from the most significant bit to the least significant bit in the order of the sign bit (S) to the exponent extension bit (Rg) to the mantissa bit (F) to the exponent bit (E). Through the above preprocessing, the sign bit and exponent bit width are fixed, the boundary of the exponent extension bit is clear, and the mantissa bit precision is improved. The overall sorting makes the exponent extension bit directly adjacent to the mantissa bit, providing a basis for the parallel execution of mantissa multiplication and decoding in subsequent operations. The generated SR-Posit data can directly enter the subsequent encoding step.
[0029] The total bit width, nb, includes the bit widths of all components. Since the sign bit is fixed to 1 bit during preprocessing, the total available bit width is the total bit width nb minus the sign bit width, that is, the total available bit width = nb - 1. This total available bit width is used to allocate the exponent extension bits, mantissa bits, and exponent bits. Next, we define the priority rule for bit width allocation: the total available bit width is allocated first to the exponent extension bits. The required width of the exponent extension bits is determined by the length of the continuous sequence itself and the end flag bit at the end of the sequence. The length of the continuous sequence is r, where r ≥ 1. The end flag bit is a 1-bit binary bit opposite to the sequence value. If the sequence does not occupy the total available bit width, the total bit width of the exponent extension bits is r + 1, that is, the sequence length plus the end flag bit. If the length of the continuous sequence r is equal to the total available bit width (that is, r = nb - 1), no additional end flag bit is required. In this case, the width of the exponent extension bits is directly r (that is, nb - 1), and the total available bit width is fully allocated.
[0030] When the total available bit width is still remaining after being allocated to the exponent extension bit, that is, the remaining bit width = the total available bit width - the exponent extension bit width > 0, the remaining bit width is allocated in the order of exponent bits taking precedence over mantissa bits: first allocated to the exponent bit, the bit width of the exponent bit is the fixed value es (such as 2-bit) defined in the preprocessing stage, if the remaining bit width is greater than or equal to es, the exponent bit is allocated es bit width; if the remaining bit width is less than es, the exponent bit is only allocated all the remaining bit widths, at this time the actual bit width of the exponent bit is less than the preset es, where nb is the total bit width and es is the fixed bit width of the exponent bit, but its encoding rule still follows the unsigned integer encoding, and the exponent bit is allocated After completion, if there is still remaining bit width, that is, remaining bit width = total available bit width - exponent extension bit width - exponent bit width > 0, then all remaining bit widths are allocated to the mantissa bits as the effective bit width of the mantissa bits, wherein the implicit highest bit 1 of the mantissa bits does not occupy this bit width. After completing the bit width allocation, the SR-Posit data is encoded in the form of one's complement. Specifically, for positive numbers (sign bit S=0), the exponent extension bit, mantissa bit, and exponent bit are encoded according to the actual value; for negative numbers, that is, the sign bit S=1, except for the sign bit, the exponent extension bit, mantissa bit, and exponent bit are all logically inverted, that is, the one's complement encoding feature, without the need for an additional add 1 operation. For positive numbers, that is, the sign bit S=0, the original values of the exponent extension bit, mantissa bit, and exponent bit are kept unchanged during encoding. That is, the continuous sequence of exponent extension bits (0 sequence or 1 sequence) and its end flag bit are directly encoded according to the actual binary value, the significant bits of the mantissa are encoded according to the normalized original value, and the exponent bit is encoded according to the actual value of the unsigned integer. For example, if the exponent extension bit of a positive number is 001, representing the 0 sequence "00" and the end flag "1", the mantissa bit is 101, and the exponent bit is 01, and the sign bit S=0, then its encoding result is "0 001 101 01" spliced in the order of sign bit, exponent extension bit, mantissa bit, and exponent bit.
[0031] For negative numbers (sign bit S=1), only the sign bit is retained as 1 during encoding, and the exponent extension bit, mantissa bit, and exponent bit are logically inverted for all bit segments except the sign bit. There is no need to perform an additional addition operation after inversion as in two's complement encoding. For example, if the positive number encoding (excluding the sign bit) corresponding to a negative number is "001 101 01", that is, the exponent extension bit is 001, the mantissa bit is 101, and the exponent bit is 01, then when encoding it, the sign bit is first set to 1, and then the remaining bit segments are inverted bit by bit to obtain the exponent extension bit "110", the mantissa bit "010", and the exponent bit "10". The final encoding result is 1 110 010 10.
[0032] The core advantage of this one's-complement encoding feature is that the conversion between positive and negative values requires only a logical inversion circuit, without the need for additional adder support, significantly simplifying the hardware encoding logic. At the same time, during decoding, restoring negative numbers only requires a single logical inversion of the bit segment outside the sign bit, avoiding the complex operation of subtracting 1 and then inverting in two's-complement encoding. At the same time, the encoding process must follow the encoding rules for special values: when all bits are 0, it represents the value 0; when the sign bit is 1 and all other bits are 0, it represents an undefined value, similar to NaN or NaR. Through the definition of the above-mentioned bit width allocation priority rules and one's-complement encoding, the SR-Posit format achieves flexible adaptation of the exponent and mantissa bit widths while ensuring the dynamic representation range. In addition, one's-complement encoding simplifies the conversion logic for positive and negative values.
[0033] like Figure 1 As shown, two encoded format data inputs 1 and 2, namely IN1 and IN2, are received. First, the sign bits S1 and S2 of the two are extracted. The highest bit is 0 for a positive number and 1 for a negative number. Based on the sign bit, it is determined whether preprocessing is required. If the sign bit is 1, all bits except the sign bit are logically inverted. Based on the one's complement encoding characteristics, the exponential extension bits are uniformly converted to the standard format of "000...01". If the sign bit is 0, the original data except the sign bit is directly retained. After completing the sign bit extraction and preprocessing, decoding is performed on IN1 and IN2 respectively. Operation: After removing the sign bit, the total bit width of the remaining data is nb-1, where nb is the total bit width of the SR-Posit format data. This part of the data includes the exponent extension bit, the mantissa bit, and the exponent bit, and the exponent extension bit is located on the leftmost side, that is, in the direction of the most significant bit. When the nb-1 bits of data are input into the leading zero counter LZC, the core function of the leading zero counter is to scan the input data from the most significant bit to the least significant bit in sequence, and count the number of consecutive 0s starting from the most significant bit. This number is the number of leading zeros. Specifically for the decoding of IN1: after removing its sign bit, we get nb-1 bits of data D1, and input D1 into the leading zero counter. The counter starts to detect bit by bit from the most significant bit of D1. If the first bit (most significant bit) is 0, the count is increased by 1; continue to detect the next bit, if it is still 0, the count continues to increase by 1 until the first non-zero bit is detected. This bit is the end flag of the exponential extension bit. Because the end of the 0 sequence of the exponential extension bit is 1 as the end flag, the counter stops counting at this time. The current count result is the number of leading zeros r1 of IN1. Similarly, after removing the sign bit of IN2, we get nb-1 bits of data D2, and input D2 into the leading zero counter. According to the same scanning rule, the number of consecutive 0s starting from the most significant bit is counted to get IN2 The number of leading zeros r2, where r1 represents the number of leading zeros of IN1 and r2 represents the number of leading zeros of IN2. This number is the length of the sequence of exponential extension bits Rg. The value of Rg is determined according to the number of leading zeros. If it is a 0 sequence, Rg1=-r1 and Rg2=-r2.
[0034] Then extract the exponent bit, at the end bit of the exponent extension bit Rg, that is, the first 1 after the leading zero sequence, then truncate the lowest es bit as the exponent bits E1 and E2, where E1 represents the exponent bit of IN1, E2 represents the exponent bit of IN2, and es is the predefined fixed width of the exponent bit. If the sequence length r1 or r2 of Rg has filled nb-1 bits, that is, there are no remaining bits allocated to the exponent bit, then the corresponding exponent bit E1 or E2 is set to 0, and then extract the mantissa bits: after removing the sign bit, the exponent extension bit Rg and the exponent bit, the remaining bits are the mantissa bits F1 and F2, where F1 represents the mantissa bit of IN1, and F2 represents the mantissa bit of IN2. The 1 at the end bit of Rg is used as the hidden bit of the mantissa, and is spliced with F1 and F2 respectively to obtain the complete valid mantissa, that is, 1+ F1 and 1+F2, while the decoding operation is being executed, the mantissa multiplication operation is started: first, the two complete valid mantissas after splicing are multiplied to obtain the initial mantissa product Pf, and then the correction vector Pcorr is generated according to the bit width occupied by the exponent extension bits of IN1 and IN2. If the Rg bit width of IN1 (r1+1) is greater than nb-3 and the Rg bit width of IN2 is less than or equal to nb-3, then Pcorr=the mantissa value of IN2; if the Rg bit width of IN2 is greater than nb-3 and the Rg bit width of IN1 is less than or equal to nb-3, then Pcorr=the mantissa value of IN1; if the Rg bit widths of both are greater than nb-3, then Pcorr=1, where Pcorr is represented as a correction vector, and the correction vector Pcorr is a compensation value generated according to the Rg bit widths of IN1 and IN2.
[0035] The lower-order nb bits of the mantissa product Pf are logically ORed with Pcorr to correct Pf and address the issue of missing valid mantissa information caused by the excessive bit width of Rg. Specifically, the initial result obtained by multiplying the complete valid mantissas of IN1 and IN2 (1+F1 and 1+F2) has a bit width greater than the total bit width nb of the SR-Posit format. The lower-order nb bits refer to the consecutive nb bits of the mantissa product Pf, starting from the least significant bit (rightmost). Finally, a mask sequence Mk is generated. The bit width of Mk is the same as the corrected mantissa product Pf. The r1+r2 bits from left to right are set to 1, and the remaining bits are all 0. This is used to subsequently determine the number of leading zeros in the mantissa product. Through this operation, decoding and mantissa multiplication are performed in parallel, avoiding the delay of waiting for decoding to complete before executing multiplication in the traditional Posit format. At the same time, the corrected mantissa product and mask sequence provide accurate data for subsequent normalization and can be directly used to determine the number of leading zeros and logical shift operations.
[0036] The number of leading zeros of the mantissa product is determined based on a mask sequence. The mask sequence Mk is generated in the preceding step based on the number of leading zeros r1 and r2 of the exponent extension bits of the two input data IN1 and IN2, where r1 and r2 are the exponent extension bits of IN1 and IN2, respectively. Their bit widths are consistent with the mantissa product Pf, and only the (r1+r2)th bit from left to right is 1, and the remaining bits are all 0.The mantissa product Pf is bitwise logically ANDed with the mask sequence Mk to obtain the intermediate result vector Pfm, that is, each bit of Pfm is the logical AND result of the corresponding bit of Pf and the corresponding bit of Mk. Since only the (r1+r2)th bit of Mk is 1, only this position in Pfm may be 1, and the remaining bits are all 0. Specifically, since the mask sequence Mk is generated based on the number of leading zeros r1 of IN1 and the number of leading zeros r2 of IN2, its bit width is exactly the same as the bit width of the mantissa product Pf, and the binary structure of Mk is unique. Only the (r1+r2)th bit from left to right (from the most significant bit to the least significant bit) is set to 1, and all other bits are 0. This structure determines that when Mk and the mantissa product Pf are bitwise logically ANDed, the value of each bit of the intermediate result vector Pfm is determined only by the corresponding bit of Pf and Mk. The logical AND result of the corresponding bits is determined. When both corresponding bits are 1, the result is 1. In other cases (either one is 0 or both are 0), the result is 0. Based on this, all bits except the r1+r2 bit in Mk are 0. Regardless of whether the bit value of the corresponding position of Pf is 0 or 1, the logical AND result of the two must be 0. For the (r1+r2)th bit in Mk that is the only 1, the logical AND result of it and the corresponding position of Pf (r1+r2th bit) depends entirely on Pf. The value of this bit performs a reduction logic OR operation on all bits of the intermediate result vector Pfm, that is, performs a logic OR operation on each bit of Pfm in turn to obtain the overflow flag Pfovf. If Pfovf is 1, it means that the (r1+r2)th bit of Pf is 1. At this time, the number of leading zeros LZcnt of the mantissa product Pf is equal to the sum of r1 and r2, that is, LZcnt=r1+r2; if Pfovf is 0, it means that the (r1+r2)th bit of Pf is 0. At this time, the number of leading zeros LZcnt is the sum of r1 and r2 plus 1, that is, LZcnt=r1+r2+1. After the number of leading zeros is determined, a unified logical shift operation is performed: first determine The number of exponent extension bits of the product is determined by concatenating the exponent extension bit Rg1 of IN1 with the exponent bit E1, and the exponent extension bit Rg2 of IN2 with the exponent bit E2, to obtain the equivalent exponent bits Eeff1 and Eeff2 of the two. The concatenation method is to combine them in bit order, retaining the numerical representation meaning, and numerically add Eeff1 and Eeff2 to obtain the total equivalent exponent bits Eeff_total. Then, according to the bit width allocation rule of the SR-Posit format (exponent extension bits are allocated first), the exponent extension bits Rg_prod of the product are separated from Eeff_total, and the number of its bits is the number of exponent extension bits of the product.
[0037] The logical shift amount shamt is then calculated using the formula shamt = number of exponent extension bits of the product + number of leading zeros LZcnt. The number of exponent extension bits of the product refers to the number of bits occupied by the product exponent extension bits (Rg_prod) separated according to the bit width allocation rule after the equivalent exponent bits of the two input data IN1 and IN2 are added. This number is determined by the large-scale adjustment requirements of the exponent and reflects the core distribution of the product result over the dynamic range of the exponent. The number of leading zeros LZcnt refers to the number of consecutive zeros that appear starting from the most significant bit after the mantissa product is corrected. This is determined by the logical AND operation and the reduced logical OR operation of the mask sequence and the mantissa product, and reflects the starting position of the significant numeric part of the mantissa product. The core purpose of the logical shift is to move the significant numeric part of the mantissa product to the high bit so that the mantissa meets the normalization requirements, and at the same time, move the leading zero sequence to the low bit as the basis for the product exponent extension bits.
[0038] A logical left shift operation is performed on the mantissa product Pf according to the shift amount, and the effective numerical part (non-leading zero) in Pf is moved to the high position, and the leading zero sequence is moved to the low position. After the shift, the original leading zero sequence will serve as the basis for the exponent extension bit Rg_prod of the final product, and the effective mantissa part will be retained for the subsequent rounding step. Finally, the normalized mantissa product (Pf after the logical left shift), the number of leading zeros LZcnt and the logical shift amount shamt are output. These data will be directly used for subsequent mantissa rounding, exponent and mantissa splicing and other operations.
[0039] A rounding operation is performed on the effective mantissa after the logical shift in the previous step. The shifted mantissa includes core effective bits and redundant bits, where the core effective bits are the number of bits to be retained, which is determined by the total bit width and bit width allocation rules of the SR-Posit format. The redundant bits are the parts that exceed the core effective bits. The highest bit of the redundant bits is taken as the rounding judgment bit. The rounding judgment bit is detected by the combinational circuit. If the bit is 1, the core effective bit is added by 1, that is, rounded up. If it is 0, the core effective bit remains unchanged, that is, rounded down. Then all redundant bits are truncated to retain the rounded core effective mantissa. For the combinational circuit, its core function is to generate the rounded core effective bit in real time based on the highest bit of the redundant bits, that is, the rounding judgment bit. The combinational circuit contains three key modules: input interface module, judgment logic module and add 1 logic module. Among them, the input interface module receives two sets of data: one is the core valid bit, which is determined by the total bit width nb and the bit width allocation rule; the other is the highest bit of the redundant bit, that is, the rounding judgment bit, which is R. The judgment logic module is the core of the circuit, and its interior consists of a NOT gate and an AND gate. When the rounding judgment bit R=1, the judgment logic module outputs a high-level signal (logic 1) to trigger the plus 1 logic module to work; when R=0, the judgment logic module outputs a low-level signal (logic 0), and the plus 1 logic module does not work.
[0040] The add-1 logic module is essentially an m-bit binary adder, whose input is connected to the core valid bit and the output signal of the judgment logic module: when a high-level trigger signal (R=1) is received, the adder performs a binary addition operation on the core valid bit and 1. For example, when the core valid bit is 1011, it becomes 1100 after adding 1; if a carry is generated during the addition process, such as when the core valid bit is 1111, it becomes 10000 after adding 1, then only the lower m bits are retained as the result, that is, 0000, and the carry is truncated. Since the bit width of the core valid bit is fixed, the excess part does not need to be retained. When the trigger signal R=0 is not received, the add-1 logic module directly outputs the original core valid bit.
[0041] The output of the circuit is connected to a truncation module. Regardless of whether an addition operation is performed, all redundant bits are truncated to retain only the m bits of the core significant bits, and the rounded core significant mantissa is finally output. For example, if the core significant bits are 101 (3 bits) and the highest redundant bit R=1, the combinational circuit outputs 110 through the addition logic; if R=0, it directly outputs 101.
[0042] Next, the mantissa bits and the equivalent exponent bits are concatenated. First, the exponent extension bits and exponent bits of the product are determined: based on the equivalent exponent bits Eeff1 of IN1 (the concatenation of Rg1 and E1) and the equivalent exponent bits Eeff2 of IN2 (the concatenation of Rg2 and E2) in the previous step, the two are added together to obtain the total equivalent exponent bits Eeff_total. Then, according to the bit width allocation priority rule of the SR-Posit format (the exponent extension bits are satisfied first, and the remaining bits are allocated to the exponent bits), the exponent extension bits Rg_prod of the product and the exponent bits E_prod of the product are separated from Eeff_total; Rg_prod, the rounded significant mantissa, and E_prod are concatenated in order from the most significant bit to the least significant bit to form a bit sequence without a sign bit.
[0043] Then the sign bit is combined and the final encoding is adjusted: the sign bit Spd of the product is the exclusive OR result of the sign bit S1 of IN1 and the sign bit S2 of IN2, that is, Spd=S1⊕S2, where ⊕ represents the exclusive OR sign, S1 and S2 represent the sign bit of IN1 and the sign bit of IN2 respectively, and Spd is placed in the most significant bit of the aforementioned bit sequence without the sign bit to form a complete bit sequence; according to the sign of Rg_prod, it is determined whether logical inversion is required. If Rg_prod is a negative number, that is, the sequence starts with 1, then logical inversion is performed on all bits except Spd, and the result after encoding adjustment is the defined product data, which includes the sign bit, the exponent extension bit Rg_prod, the rounded mantissa bit and the exponent bit E_prod, which can be directly used for storage or subsequent calculations.
[0044] like Figure 3 、 Figure 5 As shown, the mantissa product is finally converted to a fixed-point format, with the fixed-point format of the precise accumulator Qr as the target. The total bit width of the precise accumulator is 16nb, where nb is the total bit width of the SR-Posit format. This format includes a sign bit, a 31-bit extension bit, an 8nb-16-bit integer part, and an 8nb-16-bit fractional part, which is used to avoid precision loss during the accumulation process. During the conversion, the mantissa product Pf obtained in the previous step is first placed at the right end of the lowest bit of the precise accumulator Qr, that is, the rightmost low-order area, and then sign-extended according to the XOR result of the sign bits of IN1 and IN2, that is, the product sign bit Spd. If Spd=0 (positive number), add 0 to the high bit of Pf; if Spd=1, add 1 to the high bit of Pf until the total bit width after extension reaches 16nb, which is consistent with the bit width of the precise accumulator Qr. Then, according to the shift amount determined in the previous step, shamt=the number of exponent extension bits of the product + the number of leading zeros LZcnt, where shamt represents the shift amount. Perform a logical left shift on the sign-extended Pf to adjust the value to a position that matches the integer / fractional bit segment of the fixed-point number. Only the high 16nb bits of the result after the left shift are retained, and the redundant low-order bits are truncated to obtain the product Qr_prod in the fixed-point format.
[0045] like Figure 4 As shown, after the conversion is completed, an addition operation is performed to receive the third fixed-point number Qr_acc from the external input, that is, the current value of the accumulator, which is in the precise accumulator Qr format with a 16nb bit width, including a sign bit, an extension bit, an integer part, and a fractional part. The 16nb-bit-wide adder is used to perform an addition operation on Qr_prod and Qr_acc, where Qr_prod represents the product of the fixed-point format, and Qr_acc represents the value currently stored in the accumulator. The adder directly adds the corresponding bits of the two fixed-point numbers without the need for additional exponent alignment, because the fixed-point format has achieved numerical alignment through shifting, and the bit width covers all bit segments, including the 31-bit extension bit, to avoid precision loss caused by accumulation overflow.
[0046] After the addition operation is completed, the final accumulated result is output. The result is retained as a fixed-point number in the precise accumulator Qr format (16nb bit width) and can be directly used for the next multiplication and accumulation operation, that is, as a new Qr_acc, or converted to SR-Posit format for output as required.
[0047] In the arithmetic circuits of AI computing chips, the multiplication and accumulation and multiplication results of this embodiment are used in core calculations such as convolution operations and matrix multiplication in deep learning, which require multiplication and accumulation operations on large amounts of data. For example, the element-wise multiplication and accumulation of the convolution kernel and the input feature map, and the multiplication and accumulation of the weights of the fully connected layer and the input vector, can be directly used in scenarios requiring high-precision single-step multiplication, such as temporary storage of intermediate results. The multiplication and accumulation process (conversion of the mantissa product to fixed-point and fixed-point addition) is adapted to the needs of continuous multiplication and accumulation. The aforementioned fixed-point adder and shift logic can be integrated into the arithmetic circuits, and a hardware pipeline design can be used to implement pipeline execution of multiplication, conversion, and accumulation, reducing data exchange latency. Furthermore, the extended bit design of the precision accumulator adapts to the precision requirements of multiple accumulations of small and medium values in AI computing. The 16nb-wide fixed-point format can be directly connected to the adder via hardware registers, supporting parallel computing and significantly improving the computational efficiency and adaptability of the AI computing chip.
[0048] Example 2 The difference between this embodiment and embodiment 1 is that this embodiment provides an operator based on the Posit data format, including: The data acquisition module is configured to: acquire a sign bit, an exponent extension bit, a mantissa bit, and an exponent bit of a Posit data format, and pre-process each position of the Posit data format to obtain SR-Posit data; An encoding module is configured to: define a bit width allocation priority rule based on the pre-processed sign bit, exponent extension bit, mantissa bit, and exponent bit, and encode the SR-Posit data using a one's complement form; The mantissa multiplication module is configured to: use two encoded SR-Posit data as input data, perform a decoding operation on the input data, and multiply the two decoded mantissa bits to obtain a mantissa product and a mask sequence; a normalization module configured to: determine a number of leading zeros of a mantissa product according to a mask sequence, and perform a uniform logical shift on the mantissa product based on the number of exponent extension bits of the product and the number of leading zeros of the mantissa product; The product encoding module is configured to: perform a rounding operation on the shifted mantissa, concatenate the mantissa bits with the equivalent exponent bits, and combine them with the sign bit to obtain the product result of the SR-Posit data; The accumulation module is configured to: convert the mantissa product into a fixed-point number format according to the sign bit, the exponent extension bit and the exponent bit, perform addition operation based on the fixed-point number, and output the accumulation result.
[0049] The above are all preferred embodiments of the present invention, and are not intended to limit the scope of protection of the present invention. Therefore, any equivalent changes made based on the structure, shape, and principle of the present invention should be included in the scope of protection of the present invention.
Claims
1. A calculation method based on Posit data format, characterized in that: include: Obtain the sign bit, exponent extension bit, mantissa bit, and exponent bit of the Posit data format, and preprocess each position of the Posit data format to obtain SR-Posit data; Define the bit width allocation priority rule based on the pre-processed sign bit, exponent extension bit, mantissa bit and exponent bit, and encode the SR-Posit data using the one's complement form; Use two encoded SR-Posit data as input data, perform decoding operation on the input data, and multiply the two decoded mantissa bits to obtain the mantissa product and mask sequence; determining the number of leading zeros of the mantissa product according to the mask sequence, and performing a uniform logical shift on the mantissa product based on the number of exponent extension bits of the product and the number of leading zeros of the mantissa product; The shifted mantissa is rounded, the mantissa bits are concatenated with the equivalent exponent bits, and then combined with the sign bit to obtain the product result of the SR-Posit data; The mantissa product is converted into a fixed-point format according to the sign bit, the exponent extension bit, and the exponent bit, and an addition operation is performed based on the fixed-point number to output the accumulated result.
2. The calculation method based on the Posit data format according to claim 1, characterized in that: The preprocessing of each position of the Posit data format includes setting the sign bit width of the Posit data format to 1-bit, the exponent extension bit is composed of a continuous digital sequence, and the end of the sequence is marked with a binary bit with a width of 1-bit and opposite to the sequence value, the mantissa bit is coded in a normalized manner, and the highest bit 1 is hidden, the exponent bit is coded with a fixed value width and is encoded as an unsigned integer, and is sorted in the order of sign bit, exponent extension bit, mantissa bit and exponent bit.
3. The calculation method based on the Posit data format according to claim 1, characterized in that: The bit width allocation priority rule is defined, including removing the sign bit from the total bit width to obtain the total available bit width, first allocating the total available bit width to the exponent extension bits, the bit width requirement of the exponent extension bits includes the bit width occupied by the sequence itself and the end flag at the end of the sequence, the end flag of the exponent extension bits is 1 bit and has a value opposite to the sequence itself, when the length of the sequence itself is equal to the total available bit width, the bit width of the exponent extension bits is the length of the sequence itself, and no end flag bit needs to be allocated, and when there is still remaining bit width after the total available bit width is allocated to the exponent extension bits, the remaining bit width is first allocated to the exponent bits, and then the bit width remaining after the allocation to the exponent bits is allocated to the mantissa bits.
4. The calculation method based on the Posit data format according to claim 1, characterized in that: The decoding operation on the input data includes receiving two encoded SR-Posit data IN1 and IN2, removing the sign bits of IN1 and IN2, calculating the number of leading zeros r1 and r2 of the remaining data through a leading zero counter, and determining the exponent extension bits Rg1 and Rg2 according to r1 and r2 respectively, extracting the fixed bit width bits of the lowest exponent bit after the end bits of Rg1 and Rg2 as exponent bits E1 and E2, removing the sign bit, exponent extension bit and exponent bit of IN1 and IN2 respectively, and using the remaining bits as mantissa bits F1 and F2, using the 1 of the end bits of Rg1 and Rg2 as a hidden bit, and splicing them with F1 and F2 respectively to obtain a complete valid mantissa; Among them, E1 represents the exponent bit of input data IN1, E2 represents the exponent bit of IN2, F1 represents the mantissa bit of IN1, F2 represents the mantissa bit of IN2, r1 and r2 represent the number of leading zeros of IN1 and IN2 respectively, and Rg1 and Rg2 represent the exponent extension bits of IN1 and IN2 respectively.
5. The calculation method based on the Posit data format according to claim 1, characterized in that: The two mantissa bits obtained by decoding are multiplied, including performing a multiplication operation on the complete valid mantissas of IN1 and IN2 to obtain a mantissa product, generating a correction vector according to the bit width occupied by the exponent extension bits in IN1 and IN2, performing a logical OR operation on the bits of the mantissa product with a lower total bit width and the correction vector, correcting the unmantissa product, and finally generating a mask sequence according to the exponent extension length in IN1 and IN2.
6. The calculation method based on the Posit data format according to claim 1, characterized in that: The method of determining the number of leading zeros of the mantissa product according to the mask sequence includes performing a bitwise logical AND operation on the mantissa product and the mask sequence to generate an intermediate result vector, performing a reduction logical OR operation on all bits of the intermediate result vector to obtain an overflow flag, and calculating the number of leading zeros of the mantissa product according to the overflow flag.
7. The calculation method based on the Posit data format according to claim 1, characterized in that: The unified logical shift of the mantissa product includes respectively splicing the exponent extension bits and the exponent bits of IN1 and IN2 to obtain equivalent exponent bits, calculating the number of exponent extension bits of the product according to the equivalent exponent bits, adding the number of exponent extension bits of the product and the number of leading zeros to obtain a displacement, performing a logical left shift operation on the mantissa product according to the shift amount, moving the effective value part to the high bit and the leading zero sequence to the low bit, using the shifted leading zero sequence as the basis of the exponent extension bits of the final product, and outputting the normalized mantissa product, the number of leading zeros and the logical shift amount.
8. The calculation method based on the Posit data format according to claim 1, characterized in that: The product result of the SR-Posit data is obtained, including calculating the exponent extension bits of the product and the exponent bits of the product based on the equivalent exponent bits of IN1 and IN2, sequentially splicing the exponent extension bits of the product, the rounded significant mantissa, and the exponent bits of the product in order from the most significant bit to the least significant bit to obtain a bit sequence excluding a sign bit, performing a splicing operation on the XOR result of the sign bits of IN1 and IN2 and the bit sequence to form a complete bit sequence, performing final coding adjustment, and outputting the product result.
9. The calculation method based on the Posit data format according to claim 1, characterized in that: The addition operation based on fixed-point numbers includes converting the product into a fixed-point format of an accurate accumulator based on the mantissa product, the sign bit, the exponent extension bit and the exponent bit, and placing the product at the right end of the lowest bit of the accurate accumulator, sign-extending the mantissa product according to the exclusive OR result of the sign bits of IN1 and IN2, performing a logical left shift on the mantissa product after the sign extension according to the displacement to obtain a product in a fixed-point format, adding the product in the fixed-point format to an externally input fixed-point number, and outputting a final accumulated result.
10. An operator based on the Posit data format, executing the method according to claim 1, characterized in that: include: The data acquisition module is configured to: acquire a sign bit, an exponent extension bit, a mantissa bit, and an exponent bit of a Posit data format, and pre-process each position of the Posit data format to obtain SR-Posit data; An encoding module is configured to: define a bit width allocation priority rule based on the pre-processed sign bit, exponent extension bit, mantissa bit, and exponent bit, and encode the SR-Posit data using a one's complement form; The mantissa multiplication module is configured to: use two encoded SR-Posit data as input data, perform a decoding operation on the input data, and multiply the two decoded mantissa bits to obtain a mantissa product and a mask sequence; a normalization module configured to: determine a number of leading zeros of a mantissa product according to a mask sequence, and perform a uniform logical shift on the mantissa product based on the number of exponent extension bits of the product and the number of leading zeros of the mantissa product; The product encoding module is configured to: perform a rounding operation on the shifted mantissa, concatenate the mantissa bits with the equivalent exponent bits, and combine them with the sign bit to obtain the product result of the SR-Posit data; The accumulation module is configured to: convert the mantissa product into a fixed-point number format according to the sign bit, the exponent extension bit and the exponent bit, perform addition operation based on the fixed-point number, and output the accumulation result.
Citation Information
Patent Citations
Posit floating-point number processor
CN111538473A
Floating-point number processing method and device, computer equipment and processor
CN116974517A
Multiplying and adding operation hardware with mixed precision
CN119645346A
Multiplication device
CN119806473A
Extended floating-point range processors, methods, systems, and instructions
EP4478176A1
Cited By
Approximate multiplication method and arithmetic unit based on Point data
CN120821450A
Point matrix multiplication device for edge artificial intelligence chip
CN121433610A