Large number addition and / or subtraction with carry propagation
By designing a processing unit containing multiple adders and carry-generating circuits, the problem of low efficiency of large-number operations in the prior art is solved, efficient carry information processing and propagation is realized, and the performance of large-number operations is significantly improved.
Patent Information
- Application Number
- CN202380069492.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-29
- Filing Date
- 2023-09-28
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has inefficient problems when dealing with large-number operations, especially large-number addition and subtraction, especially due to the lack of an effective vector carry propagation mechanism, resulting in insufficient utilization of data paths and failure to realize performance potential.
A processing unit is designed, including multiple adders and carry bit generation circuits, and efficient carry information processing and propagation is achieved through circuit components such as multiple AND logic gates, 1-bit full adders and XOR logic gates.
This technology can efficiently process carry information on a single instruction multi-data (SIMD) data path, significantly improving the performance of large-number operations, especially when handling long carry situations, it can effectively utilize the data path to achieve higher execution efficiency.
Smart Images

Figure CN120153346A_ABST
Abstract
Description
Background Art
[0001] Arbitrary precision arithmetic (referred to herein as large number arithmetic) is an important computational primitive in cryptographic applications such as the RSA encryption algorithm (RSA). An important part of these workloads is large number addition and subtraction on large integers (e.g., 4096 bits). In addition to addition, large number addition operations are primitives used in other large number operations such as multiplication. Therefore, recently competing instruction set architectures (ISAs) have defined extensions to accelerate such operations and workloads (e.g., SVE2 of ARM, RVV of RISC-V).
[0002] ARM SVE2 provides a solution for large number arithmetic, where ARM has vector add top / bottom instructions with carry. Although these instructions can work in some cases, due to the lack of vector carry propagation, they cannot effectively handle cases where the carry propagation requires the entire width of the running data path (or more than one vector channel) (e.g., across all 512 bits). These instructions require additional software support to link them together to handle such cases. In addition, half of the vector channels remain unused, resulting in poor data path utilization and thus unfulfilled performance potential. RISC-V RVV also provides a solution for large number arithmetic. Similar to the ARM SVE2 solution, RISC-RVV has vector add instructions with carry, which require special handling for long carry cases. AVX-512F additionally provides a solution for large number arithmetic. AVX-512F instructions are used to perform vector large number addition and subtraction while handling full-width vector carry propagation. AVX-512F uses an iterative process, which means that the final carry bit for calculating the exact sum must be calculated one by one based on the results of the previous step.
[0003] Current large number workloads typically use libraries such as the GNU Multiple Precision (GMP) library based on scalar instructions, which limits their execution. Due to the large latency and area required to support carry propagation across the entire 512-bit data path, traditional methods / circuits are impractical. Another problem is that in order to support even larger integers (e.g., 1024 bits, 4096 bits), it will be necessary to support two register outputs per operation (i.e., one for the sum and one for the carry output to be fed into the calculation of the next valid 512 bits), but existing data paths and ISAs support only a single destination / output per instruction. Summary of the Invention
[0004] In the examples described herein, techniques are provided for solving inefficiencies associated with the use of scalar instructions and existing support for a single destination / output per instruction, providing wide addition (e.g., on a single instruction multiple data (SIMD) data path), and efficiently and perfectly handling carry information. In one exemplary embodiment, a processing unit includes a plurality of adders that add a first X-bit binary partial value and a second X-bit binary partial value of a first Y-bit binary value and a second Y-bit binary value to generate a first carry bit, where Y is a multiple of X. The processing unit further includes a plurality of carry bit generation circuits respectively coupled to the plurality of adders to receive the first carry bit and generate a second carry bit based on the first carry bit, where the second carry bit is used to add the first X-bit binary partial value and the second X-bit binary partial value of the first Y-bit binary value and the second Y-bit binary value, respectively.
[0005] In some embodiments, the plurality of carry bit generation circuits are configured to respectively receive the sums of the first X-bit binary partial value and the second X-bit binary partial value of the first Y-bit binary value and the second Y-bit binary value, and are configured to generate the second carry bit based on the sums of the first X-bit binary partial value and the second X-bit binary partial value of the first Y-bit binary value and the second Y-bit binary value.
[0006] In some embodiments, the first Y-bit binary value and the second Y-bit binary value are stored by at least one of a vector register, a memory operand, a double data rate (DDR) memory, a low power DDR (LPDDR), and Gen-Z. In some embodiments, a new instruction set architecture (ISA) instruction is added to an existing ISA instruction set to add the first X-bit binary partial value and the second X-bit binary partial value of the first Y-bit binary value and the second Y-bit binary value.
[0007] In some embodiments, the plurality of carry bit generation circuits include: a plurality of AND logic gates respectively; a plurality of 1-bit full adders respectively; and a plurality of XOR logic gates respectively. The plurality of AND logic gates are configured to: respectively receive the sums of the first X-bit binary partial value and the second X-bit binary partial value of the first Y-bit binary value and the second Y-bit binary value from the plurality of adders; and respectively output binary values to the plurality of 1-bit full adders and the plurality of XOR logic gates. The plurality of 1-bit full adders are configured to: receive the binary values output from the plurality of AND logic gates, the first carry bit from the plurality of adders, and binary values from adjacent 1-bit full adders; and output binary values. The plurality of XOR logic gates are configured to: receive the binary values output by the plurality of AND logic gates and the binary values output by the plurality of 1-bit full adders; and output the second carry bit to a vector register.
[0008] In some specific embodiments, the plurality of carry bit generation circuits include: a plurality of AND logic gates; and a plurality of 1-bit vector carry propagation (VCP) logic circuits. The plurality of AND logic gates are configured to: receive the sum of the first X-bit binary parts and the second X-bit binary parts of the first Y-bit binary value and the second Y-bit binary value from the plurality of adders; and output the binary values to the plurality of 1-bit vector carry propagation (VCP) logic circuits respectively. The plurality of 1-bit VCP logic circuits are configured to: receive the binary values from the plurality of AND logic gates, the first carry bit from the plurality of adders, and the binary values output by the adjacent 1-bit VCP logic circuits among the plurality of 1-bit VCP logic circuits; output other binary values to other adjacent 1-bit VCP logic circuits among the plurality of 1-bit VCP logic circuits; and output the second carry bit to the vector register.
[0009] In some specific embodiments, the plurality of AND logic gates are a plurality of first AND logic gates, and the 1-bit VCP logic circuits include: a plurality of OR logic gates respectively; and a plurality of second AND logic gates respectively. The plurality of OR logic gates are configured to respectively receive the binary values output from the plurality of first AND logic gates and the first carry bit from the plurality of adders. The plurality of second AND logic gates are configured to: receive the first carry bit from the plurality of adders and the binary values output by the plurality of OR logic gates; and output the second carry bit to the vector register respectively.
[0010] In some specific embodiments, the vector register is a first vector register, and the processing unit further includes: a plurality of bitwise inversion / NOT logic gates, the plurality of bitwise inversion / NOT logic gates are coupled to a second vector register, and are respectively configured to receive the second X-bit binary part, perform bitwise inversion on the second X-bit binary part of the second Y-bit binary value, and output the bitwise inverted X-bit binary part of the second X-bit binary part to the plurality of carry bit generation circuits.
[0011] In some specific embodiments, the plurality of carry bit generation circuits further include: a plurality of multiplexers, the plurality of multiplexers are coupled to the second vector register, and are configured to respectively receive the bitwise inverted X-bit binary part of the second Y-bit binary value from the plurality of bitwise inversion / NOT logic gates, receive the second X-bit binary part, receive binary values from adjacent multiplexers, and output the binary values to other adjacent multiplexers.
[0012] In another specific implementation, a system includes the processing unit according to claim 1. The system includes another processing unit, and the other processing unit includes the plurality of adders and the plurality of carry bit generation circuits.
[0013] In some specific implementations, the processing unit further includes: vector registers that respectively store the sums of the first X-bit binary parts and the second X-bit binary parts of the first Y-bit binary value and the second Y-bit binary value. In some specific implementations, the plurality of adders are a first plurality of adders, and the processing unit includes a second plurality of adders that respectively add the first Y-bit binary value and the second Y-bit binary value from the first vector register and the second vector register. In some specific implementations, the plurality of adders add the first Y-bit binary value and the second Y-bit binary value from the first vector register and the second vector register respectively.
[0014] In another specific implementation, a method includes adding the first X-bit binary part value and the second X-bit binary part value of a first Y-bit binary value and a second Y-bit binary value by a plurality of adders to generate a first carry bit, where Y is a multiple of X. The method further includes generating a second carry bit based on the first carry bit by a plurality of carry bit generation circuits; and adding the first X-bit binary part and the second X-bit binary part of the first Y-bit binary value and the second Y-bit binary value respectively using the second carry bit.
[0015] In some specific implementations, the method further includes: respectively receiving, by the plurality of carry bit generation circuits, the sums of the first X-bit binary parts and the second X-bit binary parts of the first Y-bit binary value and the second Y-bit binary value; and generating, by the plurality of carry bit generation circuits, the second carry bit based on the sums of the first X-bit binary parts and the second X-bit binary parts of the first Y-bit binary value and the second Y-bit binary value.
[0016] In some specific implementations, the method further includes storing the first Y-bit binary value and the second Y-bit binary value by at least one of a vector register, a memory operand, a double data rate (DDR) memory, a low power DDR (LPDDR), and Gen-Z.
[0017] In some specific embodiments, each carry bit generation circuit in the carry bit generation circuit includes a plurality of AND logic gates; a plurality of 1-bit full adders; a plurality of XOR logic gates, and the method further includes: receiving, by the plurality of AND logic gates respectively from the plurality of adders, the sum of the first X-bit binary parts and the second X-bit binary parts of the first Y-bit binary value and the second Y-bit binary value; outputting, by the plurality of AND logic gates, binary values to the plurality of 1-bit full adders and the plurality of XOR logic gates respectively; receiving, by the plurality of 1-bit full adders, the binary values output from the plurality of AND logic gates, the first carry bit from the plurality of adders, and the binary values from adjacent 1-bit full adders; outputting, by the plurality of 1-bit full adders, binary values respectively; receiving, by the plurality of XOR logic gates respectively, the binary values output from the plurality of AND logic gates and the binary values output from the plurality of 1-bit full adders; and outputting, by the plurality of XOR logic gates, the second carry bit to the vector register.
[0018] In some specific embodiments, each carry bit generation circuit in the carry bit generation circuit includes a plurality of AND logic gates and a plurality of 1-bit vector carry propagation (VCP) logic circuits, and the method further includes: receiving, by the plurality of AND logic gates respectively from the plurality of adders, the sum of the first X-bit binary parts and the second X-bit binary parts of the first Y-bit binary value and the second Y-bit binary value; outputting, by the plurality of AND logic gates, binary values to the plurality of 1-bit vector carry propagation (VCP) logic circuits respectively; receiving, by the plurality of 1-bit VCP logic circuits respectively, the binary values from the plurality of AND logic gates, the first carry bit from the plurality of adders, and the binary values output from adjacent 1-bit VCP logic circuits in the plurality of 1-bit VCP logic circuits; and outputting, by the plurality of 1-bit VCP logic circuits, other binary values to other adjacent 1-bit VCP logic circuits in the plurality of 1-bit VCP logic circuits, and outputting the second carry bit to the vector register.
[0019] In some specific embodiments, the method further includes storing, by the vector register respectively, the sum of the first X-bit binary parts and the second X-bit binary parts of the first Y-bit binary value and the second Y-bit binary value.
[0020] In another specific implementation, a processing unit includes a first vector register that stores a first Y-bit binary value including a plurality of first X-bit binary parts, where Y is a multiple of X. The processing unit further includes a second vector register that stores a second Y-bit binary value including a plurality of second X-bit binary parts. The processing unit further includes: a plurality of adders that add the first X-bit binary part values and the second X-bit binary part values of the first Y-bit binary value and the second Y-bit binary value to generate a first carry bit; and a plurality of carry bit generation circuits respectively coupled to the plurality of adders to receive the first carry bit and generate a second carry bit based on the first carry bit. The plurality of adders respectively add the first X-bit binary part and the second X-bit binary part of the first Y-bit binary value and the second Y-bit binary value using the second carry bit. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The present disclosure can be better understood by reference to the accompanying drawings, and its numerous features and advantages will be apparent to those skilled in the art. The same reference numerals are used in different drawings to denote similar or identical items.
[0022] Figure 1A and Figure 1B illustrates an exemplary processing unit for generating / propagating carry bits for integer large number addition (e.g., 512b) using 1-bit full adders (FAs) and XOR gates per vector channel.
[0023] Figure 2A and Figure 2B shows another exemplary processing unit for generating / propagating carry bits for integer large number addition (e.g., 512b) using vector carry propagation (VCP) circuits per vector channel.
[0024] Figure 3 illustrates according to some embodiments Figure 1A and Figure 1B an exemplary configuration of the 1-bit VCP logic circuit shown.
[0025] Figure 4 shows according to some embodiments from Figure 1A and Figure 1B as well as Figure 2A and Figure 2B exemplary hardware components of the processing unit configured to perform large number 512b addition.
[0026] Figure 5A and Figure 5BIllustrates an exemplary processing unit that can simultaneously perform carry generation / propagation for 2x256b large number addition according to some embodiments.
[0027] Figure 6A and Figure 6B Shows an exemplary processing unit that generates / propagates carry bits for integer large number subtraction (e.g., 512b) according to some embodiments.
[0028] Figure 7A and Figure 7B Illustrates an exemplary processing unit that generates / propagates carry bits for chaining multiple 512b subtraction operations according to some embodiments.
[0029] Figure 8A and Figure 8B Shows an exemplary processing unit that generates / propagates carry bits for both addition and subtraction of large numbers according to some embodiments.
[0030] Figure 9A and Figure 9B Illustrates an exemplary processing unit that can perform carry generation / propagation for adding and subtracting 2x256b large numbers according to some embodiments.
[0031] Figure 10A and Figure 10B Shows according to some embodiments Figure 8A and Figure 8B The illustrated processing unit configured to perform large number 512b addition / subtraction.
[0032] Figure 11A and Figure 11B Illustrates another configuration of an exemplary processing unit according to some embodiments, which adds two large number binary values based on Figure 2A and Figure 2B The configuration of the shown processing unit without first storing the carry in a vector register.
[0033] Figure 12 Shows an exemplary critical path for chained 512b addition according to some embodiments.
[0034] Figure 13 Illustrates an exemplary method of generating a carry bit and adding a first binary value to a second binary value according to some embodiments. Detailed Description
[0035] FIG. 1 to Figure 13Techniques are illustrated for solving inefficiencies associated with the use of scalar instructions and existing support for a single destination / output per instruction, providing wide addition (e.g., on a single instruction multiple data (SIMD) data path), and efficiently and perfectly handling carry information. Most data path components in modern CPUs have AVX-512 support (as well as other vector ISAs; although other vector ISAs do not support AVX-512 by definition), and are tested by comparison with MAX_WORD, which is a computation used in some large number addition algorithms. MAX_WORD is a binary number containing all 1s. The size of the number is defined by the data type (in this use case, the size of the number is 64 bits). In some embodiments, the processing unit augments the existing data path with a 1-bit full adder (FA) per vector channel, along with XOR and OR gates. This total additional logic is small compared to the overall size of the floating point (FP) SIMD data path / execution unit. The latency impact of this augmentation is also small, because the routing of the carry input most significant bit (MSB) in the first vector register can overlap with the latency of the vector 64-bit adder, and the sum of each adder is tested for a 64-bit sum relative to the MAX_WORD value (or alternatively three levels of four-input sum gates). The 1-bit FAs are coupled in a ripple-carry fashion, but alternative implementations may optimize the area / performance cost.
[0036] The processing unit also introduces a pair of new instructions to accelerate such workloads. The processing unit provides a single instruction dependency chain for chaining together multiple operations for larger bit widths (e.g., 2048 bits or more), and implements a modest modification to the existing data path (e.g., an existing AVX512 FP data path). The processing unit even accelerates 512b large number addition on existing AVX-512F-based solutions. In some embodiments, when adding 100,000 large numbers residing in registers, the processing unit provides approximately a 5x speedup over a scalar 512b addition implementation. These operations can be performed with the processing unit in fewer instructions, while avoiding the use of scalar registers that cause performance penalties due to moving data between vector and scalar registers.
[0037] Although the processing unit performs large number addition and subtraction (add / subtract) in hardware using AVX-512 registers, in some embodiments, the processing unit uses the same method to generate / propagate carry / borrow bits, and then uses those carry / borrow bits to calculate the correct large number sum / difference for other vector ISAs (such as ARM SVE / SVE2, ARM NEON, RISC-V RVV). In some embodiments, the processing unit implements a 512-bit data path, but in other embodiments, the processing unit implements a narrower data path (e.g., 256b) or a wider data path (e.g., 1024b). For integer large number addition / subtraction, the processing unit utilizes hardware to 1) perform vector carry / borrow generation and propagation and 2) use the propagated carry / borrow bits to complete the final arithmetic operation.
[0038] Additionally, vector large number multiplication relies on reduced radix calculations with filled zeros to overflow carry bits and then processes those carry bits using scalar add-with-carry instructions because the carry bits cannot be propagated using existing instructions. With these instructions and hardware support of the processing unit, the processing unit performs large number multiplication faster because it 1) does not need to perform two radix conversions, 2) does not require filling, and can utilize all bits in the vector arithmetic logic unit (ALU), and 3) does not particularly need to process overflow carry bits because there is no overflow. Thus, the processing unit also significantly accelerates large number multiplication. In some embodiments, the processing unit uses large number addition, multiplication, and subtraction operations to perform large number division and various other math operators more efficiently than conventional solutions.
[0039] If traditional vector addition (e.g., vpaddd) is used, each channel is calculated independently of the results / carries of other channels. The result of this vector operation (i.e., the "vector sum") can be significantly different from the "real sum". In some embodiments, the processing unit resolves this defect by performing carry propagation across all vector channels to provide an accurate "real sum" in a way that is not possible with traditional vector addition.
[0040] Figure 1A and Figure 1B(collectively referred to as FIG. 1) illustrates an exemplary processing unit 100 that uses 1-bit FAs and XOR gates per vector channel to generate / propagate carry bits for integer large number addition (e.g., 512b). In some other embodiments, for reasons related to performance, energy, and area, other types of multi-bit adders (e.g., carry look-ahead adders) may be used. The processing unit 100 cascades carry bits from one adder in one vector channel to another vector channel to generate carry bits for adding two large numerical values together. The processing unit 100 includes a plurality of vector channels 101-1 to 101-8, a first vector register 110, a second vector register 120, and a third vector register 130. In some embodiments, at least one of the vector registers 110, 120, 130 is a memory operand. In some embodiments, at least one of the first vector register 110, the second vector register 120, and the third vector register 130 is alternatively a conventional memory (e.g., double data rate (DDR) memory, low power DDR (LPDDR), etc.) and / or a "fabric-attached memory" (e.g., Compute Express Link (CXL) memory, Gen-Z). Each of the first vector register 110 ("zmm1"), the second vector register 120 ("zmm2"), and the third vector register 130 ("zmm3") includes a plurality of vector register portions, where the number of vector register portions is equal to the number of the plurality of vector channels 101-1 to 101-8.
[0041] The first vector register 110 includes vector register portions 110-1 to 110-8, the second vector register 120 includes vector register portions 120-1 to 120-8, and the third vector register 130 includes vector register portions 130-1 to 130-8 respectively disposed within the vector channels 101-1 to 101-8. The first vector register 110, the second vector register 120, and the third vector register 130 are each configured to store a Y-bit binary value (not shown), where each of the vector register portions 110-1 to 110-8, 120-1 to 120-8, 130-1 to 130-8 is configured to store an X-bit binary portion of the Y-bit binary value (not shown). In this example, the second vector register 120 stores the large numerical value A, and the third vector register 130 stores the large numerical value B, where Y is five hundred and twelve (512), X is sixty-four (64), so Y is eight (8) times X, but other multiples are also possible.
[0042] Vector channels 101-1 to 101-8 each further include data paths respectively including a plurality of adders 140-1 to 140-8. Adders 140-1 to 140-8 are configured to add the first X-bit binary partial values and the second X-bit binary partial values of the first Y-bit binary value and the second Y-bit binary value. Adders 140-1 to 140-8 also respectively generate carry or carry bits 151-1 to 151-8. In addition to adders 140-1 to 140-8, vector channels 101-1 to 101-8 each further include a plurality of carry bit generation circuits 150-1 to 150-8 respectively coupled to the plurality of adders 140-1 to 140-8 to respectively generate carry bits 154-1 to 154-8. Starting from adder 140-8, adder 140-8 generates a carry bit cascaded to vector channel 101-7, adder 140-7 generates a carry bit cascaded to vector channel 101-6, adder 140-6 generates a carry bit cascaded to vector channel 101-5, and so on. The carry bits 154-1 to 154-8 generated by the carry bit generation circuits 150-1 to 150-8 are ultimately used to add the first Y-bit binary value and the second Y-bit binary value, where any carry bits generated before the carry bits 154-1 to 154-8 are intermediate carry bits (i.e., intermediate carry bits subsequently used to formulate the carry bits 154-1 to 154-8).
[0043] Vector channels 101-1 to 101-8 each further include a plurality of carry bit generation circuits 150-1 to 150-8 respectively coupled to adders 140-1 to 140-8. The carry bit generation circuits 150-1 to 150-8 receive carry bits from adjacent adders among adders 140-1 to 140-8 and respectively generate carry bits 154-1 to 154-8 based on the carry bits 151-1 to 151-8. For example, carry bit generation circuit 150-1 receives carry bit 151-2 from adder 140-2, carry bit generation circuit 150-2 receives carry bit 151-3 from adder 140-3, carry bit generation circuit 150-3 receives carry bit 151-4 from adder 140-4, and so on. The carry bit generation circuits 150-1 to 150-8 are also configured to receive the sums of the first X-bit binary partial and the second X-bit binary partial of the first Y-bit binary value and the second Y-bit binary value respectively stored in the second vector register 120 and the third vector register 130. The carry bit generation circuits 150-1 to 150-8 respectively generate carry bits 154-1 to 154-8 based on the sums of the first X-bit binary partial and the second X-bit binary partial of the first Y-bit binary value and the second Y-bit binary value respectively stored in the second vector register 120 and the third vector register 130.
[0044] The carry bits 154-1 to 154-8 are respectively stored in the vector register sections 110-1 to 110-8 of the first vector register section 110. Then, adders 140-1 to 140-8 use the generated carry bits 154-1 to 154-8 stored by the vector register sections 110-1 to 110-8 to add the first Y-bit binary values and the second Y-bit binary values stored by the second vector register 120 and the third vector register 130 respectively. Details thereof will be explained in detail with respect to Figure 4 its details.
[0045] The carry bit generation circuits 150-1 to 150-8 adopt various configurations in various specific embodiments, and FIG. 1 shows an exemplary configuration of the carry bit generation circuits 150-1 to 150-8. In the configuration shown in FIG. 1, the carry bit generation circuits 150-1 to 150-8 respectively include a plurality of AND logic gates 152-1 to 152-8 (for example, 64-bit AND gates), a plurality of 1-bit FAs 153-1 to 153-8, and a plurality of XOR logic gates 154-1 to 154-8. The AND logic gates 152-1 to 152-8 are configured to respectively receive the sums of the first X-bit binary parts and the second X-bit binary parts of the first Y-bit binary values and the second Y-bit binary values from the adders 140-1 to 140-8. The AND logic gates 152-1 to 152-8 are also configured to respectively output the binary values to the 1-bit FAs 153-1 to 153-8 and the XOR logic gates 154-1 to 154-8. The XOR logic gates 154-1 to 154-8 are configured to respectively receive the binary values output by the AND logic gates 152-1 to 152-8 and the binary values output by the 1-bit FAs 153-1 to 153-8. The XOR logic gates 154-1 to 154-8 perform logical processing on these received binary values, and are also configured to respectively output the carry bits 154-1 to 154-8 to the first vector register 110, specifically to the vector register sections 110-1 to 110-8.
[0046] Multiple 1-bit FAs 153-1 to 153-7 are configured to receive carry bits 151-2 to 151-8 from adders 140-2 to 140-8, and are also configured to receive carry bits 155-2 to 155-8 from adjacent 1-bit FAs among 1-bit FAs 153-2 to 152-7, respectively. In contrast, 1-bit FA 152-8 is configured to receive a "0" binary value and the MSB 111 of the first vector register 110. For example, 1-bit FA 152-1 receives carry bit 155-2 from 1-bit FA 152-2, 1-bit FA 152-2 receives carry bit 155-3 from 1-bit FA 152-3, 1-bit FA 152-3 receives carry bit 155-4 from 1-bit FA 152-4, and so on. 1-bit FA 152-8 receives binary "0" instead of the carry bit from adder 140, like the carry bits received by other 1-bit FAs 152-1 to 152-7.
[0047] As can be seen in FIG. 1, the carry bit generation circuit 150-1 differs from the carry bit generation circuits 150-2 to 150-8 in that the carry bit generation circuit 150-1 includes additional logic circuitry. The carry bit generation circuit 150-1 also includes an OR logic gate 142 (e.g., a 64-bit logical OR gate), which receives the carry bit 151-1 from adder 140-1 and the binary value from 1-bit FA 153-1, and outputs the carry bit 143 to the vector register section 110-1 of the first vector register 110.
[0048] Figure 2A and Figure 2B (collectively referred to as FIG. 2) shows another exemplary processing unit 200 that uses a vector carry propagation (VCP) circuit per vector channel to generate / propagate carry bits for integer large number addition (e.g., 512b). The processing unit 200 includes a data path that includes a plurality of vector channels 201-1 to 201-8, and the plurality of vector channels include another configuration for carry bit generation circuits 250-1 to 250-8 that include VCP circuits. The processing unit 200 provides another configuration for generating and propagating carry bits for 512b integer large number addition in hardware. Then the processing unit 400 ( Figure 4 ) calculates the result of value A + value B + carry (as discussed above for processing unit 100) and includes some of the same features shown in FIG. 1. For the sake of brevity, these same features will not be discussed.
[0049] The carry bit generation circuits 250-1 to 250-8 respectively include a plurality of AND logic gates 252-1 to 252-8 (e.g., 64-bit AND gates) and a plurality of 1-bit VCP logic circuits 253-1 to 253-8. The plurality of AND logic gates 252-1 to 252-8 are configured to respectively receive the sum of the first X-bit binary part and the second X-bit binary part of the first Y-bit binary value and the second Y-bit binary value from the plurality of adders 140-1 to 140-8. The plurality of AND logic gates 252-1 to 252-8 are also configured to respectively output the binary values to the plurality of 1-bit VCP logic circuits 253-1 to 253-8.
[0050] The plurality of 1-bit VCP logic circuits 253-1 to 253-8 are configured to respectively receive the binary values output from the plurality of AND logic gates 252-1 to 252-8, the carry bits from the plurality of adders 140-1 to 140-8, and the binary values output from the adjacent 1-bit VCP logic circuits among the plurality of 1-bit VCP logic circuits 253-1 to 253-8. For example, the carry bit generation circuit 250-1 receives the carry bit 251-2 from the adder 140-2, the carry bit generation circuit 250-2 receives the carry bit 151-3 from the adder 140-3, the carry bit generation circuit 250-3 receives the carry bit 151-4 from the adder 140-4, and so on.
[0051] The plurality of 1-bit VCP logic circuits 253-1 to 253-7 are also configured to output other binary values to the other adjacent 1-bit VCP logic circuits among the plurality of 1-bit VCP logic circuits 253-2 to 253-8. The plurality of 1-bit VCP logic circuits 253-1 to 253-8 also respectively output the carry bits 257-1 to 252-8 to the first vector register 110 (specifically, the vector register parts 110-1 to 110-8), and respectively output the carry bits 255-2 to 255-8 to the adjacent 1-bit VCP logic circuits among the 1-bit VCP logic circuits 253-1 to 253-7. The carry bits 257-1 to 251-8 are respectively placed in the least significant bits 113-1 to 213-8 of the vector channels 201-1 to 201-8. The 1-bit VCP logic circuit 253-8 receives a binary "0" instead of the carry bit from the adder 140, as received by the other 1-bit VCP logic circuits 253-1 to 253-7, and also receives the MSB 111.
[0052] For example, the 1-bit VCP logic circuit 253-2 outputs a carry 255-2 to the 1-bit VCP logic circuit 253-1, the 1-bit VCP logic circuit 253-3 outputs a carry 255-2 to the 1-bit VCP logic circuit 253-2, the 1-bit VCP logic circuit 253-4 outputs a carry 255-3 to the 1-bit VCP logic circuit 253-2, and so on.
[0053] As can be seen in FIG. 2, the carry bit generation circuit 250-1 differs from the carry bit generation circuits 250-2 to 250-8 in that the carry bit generation circuit 250-1 includes additional logic circuitry. The carry bit generation circuit 250-1 also includes an OR logic gate 242 (e.g., a 64-bit logical OR gate) that receives a carry bit 251-1 from the adder 140-1 and outputs a carry bit 243 to the vector register portion 110-1 of the first vector register 110. In some embodiments, the carry bit 243 is stored in the MSB 111 of the vector register portion 110-1.
[0054] Figure 3 An exemplary configuration of the 1-bit VCP logic circuits 253-1 to 253-8 (collectively referred to as the 1-bit VCP logic circuit 253) shown in FIG. 2 is illustrated. The 1-bit VCP logic circuit 253 is an optimization that combines carry generation and propagation into a single (simple) logic block. The 1-bit VCP logic circuit 253 includes an OR logic gate 310 coupled to an AND logic gate 320. The OR logic gate 310 is configured to receive a carry bit 251 on its first input 311 and a carry bit 255 on its second input 312. The OR logic gate 310 outputs a carry bit 257 on its output 313.
[0055] The AND logic gate 320 is configured to receive the carry bit 257 from the output 313 of the OR logic gate 310 on its first input 321. The AND logic gate 320 is also configured to receive a binary value output by the AND logic gate 252 on its second input 322. The AND logic gate 320 is even further configured to output a carry bit 255.
[0056] Figure 4 Exemplary hardware components of the processing units 100, 200 shown in FIGS. 1 and 2 are illustrated, which are configured with another data path to perform a large number 512b addition. The carry bit generation circuits 150-1 to 150-8, 250-1 to 250-8 shown in FIGS. 1 and 2 respectively are excluded from the illustration for simplicity of explanation. Figure 4 Shown from FIGS. 1 and Figure 4The shared data path components, adder 140. After the carry bit has been generated by either of the processing units 100, 200 and stored in the vector register portions 110-1 to 110-8 of the first vector register 110, the addition of the large numbers A and B can be performed by adders 140-1 to 140-8.
[0057] The processing units 100, 200 utilize Figure 4 the data path shown to calculate the result of value A (e.g., a 512-bit large number binary value stored in the second vector register 120) + value B (e.g., another 512-bit large number binary value stored in the third vector register 130) + carry. The processing unit 100 first generates and propagates the carry bit. The processing unit 100 performs vector carry generation and propagation for this 512b addition operation. The instruction for generating / propagating the carry bit (generate_carry_512) has three input operands: the value stored at the first vector register 110 (carry, where only the most significant bit (MSB) 111 of the first vector register 110 is used and all other bits are ignored), the value stored at the second vector register 120 (the first 512-bit large number value A), and the value stored at the third vector register 130 zmm3 (storing the second 512-bit large number B). In some embodiments, some of the operands can be memory operands instead of register operands. In some embodiments, the generate_carry instruction can have a version with an implicit 0 carry value.
[0058] The destination register (the first vector register 110) of the processing units 100, 200 is written with: 1) the "carry bit" (i.e., the carry bit for updating the "vector sum" to the "real sum"), and 2) the "carry output" bit of the 512-bit addition operation, such that multiple 512b additions are chained to perform addition on a larger number (e.g., two 512b additions can be chained to perform a 1024b addition). These "carry bits" are respectively placed in the least significant bits 113-1 to 113-8 of the vector channels 101-1 to 101-8. The "carry output" bit or carry bit 143 is placed in the MSB 111 of the destination register (such as the first vector register 110), but in different embodiments, it can be placed in other unused bits.
[0059] The new ISA instruction added to the existing ISA instruction set to perform addition (complete_wide_add) with adders 140-1 to 140-8 has three input operands: carry bits 154-1 to 154-8, 257-1 to 257-8 stored in the least significant bits 113-1 to 113-8 of each vector channel 101-1 to 101-8, 201-1 to 201-8 by the first vector register 110, where the MSB 111 of the first vector register 110 for linking 512b is added during generation of the ignored carry bit, and the 512b value large number A is stored in the second vector register 120, and the 512b large value B is stored in the third vector register 130.
[0060] Adders 140-1 to 140-8 receive parts of large number A and large number B from vector channels 101-1 to 101-8, 201-1 to 201-8 respectively. The adders also receive the previously generated carry bits, such as the carry bits previously stored in the least significant bits 113-1 to 113-8 of each vector channel 101-1 to 101-8, 201-1 to 201-8. Then adders 140-1 to 140-8 add large number A and large number B respectively using the carry bits from each vector channel in vector channels 101-1 to 101-8, 201-1 to 201-8 to arrive at the appropriate parts of the sum of large number A and large number B. Then the processing units 100, 200 update the second vector register 120 with the correct sum of this operation. In some embodiments, the first vector register 110 should not be rewritten so that it can be saved to provide a "carry output" for linking subsequent additional 512b additions. Thus, to facilitate large number addition using existing 64b adders, for example in a single instruction, multiple data (SIMD) data path, vector channels 101-1 to 101-8, 201-1 to 201-8 are augmented with "carry input" carry bits by utilizing the bit feed of the first vector register 110.
[0061] The processing units 100, 200 are configured to generate / propagate carry bits for 512b large numbers. However, the generation / propagation of carry bits for adding large numbers discussed above can be applied to smaller large numbers, such as >64 bits but <512 bits. Figure 5A And Figure 5B (Collectively Figure 5) illustrates an exemplary processing unit 500 that can simultaneously perform carry generation / propagation for 2x256b large number addition.
[0062] The processing unit 500 includes a data path that generates / propagates carry bits for the result of two values of A + two values of B + carry bits for each of these two values. The processing unit 500 includes some of the same features shown in FIG. 1, which will not be discussed for the sake of brevity. The processing unit 500 utilizes vector channels 501-1 to 501-4 to calculate the first A + B + carry, and utilizes vector channels 501-5 to 501-8 to calculate the second A + B + carry, as discussed above in the singular of the processing unit 200. Thus, since the processing unit 500 can simultaneously generate / propagate carry bits for two additions, the only feature from the processing unit 200 is replicated.
[0063] To perform two carry bit propagations / generations simultaneously, the processing unit 500 includes a second copy of the carry bit generation circuit 250-1. This second copy of the carry bit generation circuit 250-1 is shown as the carry bit generation circuit 550-5 in the vector channel 501-5. Thus, the vector channels 501 to 504 generate carry bits for the first 256b large number addition, and the vector channels 501-5 to 501-8 generate carry bits for the second 256b large number addition, doubled as discussed above for FIG. 2, where the processing unit 200 generates carry bits for a single 512b large number addition. Similarly, instead of using a single MSB in the MSB 111 to generate the carry bits stored in the least significant bits 113-1 to 113-8 of the vector channels 101-1 to 101-8, the processing unit 500 uses the MSB 111 to generate the carry bits stored in the least significant bits 113-1 to 113-4 of the vector channels 501-1 to 501-4, and uses another significant bit 511 from the vector channel 511-5 to generate the carry bits stored in the least significant bits 113-1 to 113-4 of the vector channels 501-1 to 501-4.
[0064] This concept can be extended to even smaller large number additions. In some embodiments, the instructions generate_carry_256 and generate_carry_128 can perform concurrent 2x256b additions and 4x128b additions respectively. The hardware for generate_carry_256 is shown in FIG. 5. Similarly, generate_carry_128 will utilize two additional OR gates and two additional bits stored in the first vector register 110. In some embodiments, these 2x256b and 4x128b addition operations will not require different instructions to perform the final completion of the entire addition. After generate_carry_512 / 256 / 128, the same complete_wide_add instruction discussed above can be used.
[0065] The concepts for large number addition disclosed above can be extended to large number subtraction. Figure 6A and Figure 6B (collectively referred to as FIG. 6) illustrate an exemplary processing unit 600 that generates / propagates carry bits for integer large number subtraction (e.g., 512b). The processing unit 600 includes a data path that generates and propagates carry bits for subtracting a large number B stored in a third vector register 130 from a large number A stored in a second vector register 120, such as for a generate_sub_carry_512 instruction. The processing unit 600 takes advantage of the fact that A - B = A + (-B), and -B in two's complement can be calculated by ~B + 1 (bitwise inversion / negation of B, followed by increment). FIG. 6 shows the data path for the generate_sub_carry_512 instruction. In some embodiments, the instruction is used for the first 512b subtraction operation in a 512b subtraction operation chain, where the first vector register 110 is not used as an input. The initial carry 601 entering the first "1b VCP" block is set to 1. The complement of B (one's complement of 1) is taken. The initial carry bit is used to represent (-B) (two's complement).
[0066] The processing unit 600 includes all the hardware shown in FIG. 2, and these same features will not be discussed for the sake of brevity. To facilitate generating the complement of the large number B stored in the third vector register 130 so that the processing unit 600 can subtract the large number B from the large number A, the processing unit 600 further includes a plurality of bitwise inversion / NOT logics 610-1 to 610-8.
[0067] The processing unit 600 includes a plurality of bitwise inversion / NOT logic gates 610-1 to 610-8 for each of the vector channels 601-1 to 601-8 in the vector channels. The plurality of bitwise inversion / NOT logic gates 610-1 to 610-8 are coupled to the second vector register 120, specifically to the vector register portions 120-1 to 120-8. The plurality of bitwise inversion / NOT logic gates 610-1 to 610-8 are configured to receive 64b portions of the large number B from the vector register portions 120-1 to 120-8 of the second vector register 120 respectively. The plurality of bitwise inversion / NOT logic gates 610-1 to 610-8 perform bitwise inversion of the 64b portions of the large number B. The plurality of bitwise inversion / NOT logic gates 610-1 to 610-8 further output the 64b portions of the large number B to a plurality of carry bit generation circuits 250-1 to 250-8 respectively.
[0068] Figure 7A and Figure 7B(collectively referred to as FIG. 7) shows an exemplary processing unit 700 that generates / propagates carry bits for chaining multiple 512b subtraction operations. The processing unit 700 includes another data path 711 that is used to generate and propagate carry bits for such operations. The processing unit 700 includes all of the hardware shown in FIG. 2; these same features will not be discussed for the sake of brevity. To facilitate generating the complement of the large number B stored in the third vector register 130 such that the processing unit 700 can generate / propagate carry bits to subtract the large number B from the large number A, the processing unit 600 further includes a plurality of bitwise inversion / NOT logics 610-1 to 610-8.
[0069] For all subtractions other than the first 512b subtraction, a separate generate_sub_carry_chained_512 instruction can be used to generate additional carry bits. This first instruction does not use an initial carry bit of 1, but instead uses the carry bit set in the MSB 111 of the first vector register 110, which is filled by a previous generate_sub_carrry_(chained)_512 instruction. Thus, the carry bit generation circuit 250-8 receives the carry bit from this previous generate_sub_carrry_(chained)_512 instruction via the data path 711, and this carry bit comes from the MSB 111 of the first vector register 110. In some embodiments, these two instructions can be combined into one instruction that utilizes a separate static field, such as an immediate-encoded static field, that controls whether the carry input should be forced or whether the carry input should be taken from the first vector register 110.
[0070] Figure 8A and Figure 8B (collectively referred to as FIG. 8) shows an exemplary processing unit 800 that is capable of generating / propagating carry bits for both addition and subtraction of large numbers. The processing unit 800 includes all of the hardware shown in FIG. 6; these same features will not be discussed for the sake of brevity. The processing unit 800 includes a data path that further includes a plurality of multiplexers 810-1 to 810-8 for each of the vector channels 801-1 to 801-8. The multiplexers 810-1 to 810-8 are respectively coupled to the second vector register 120, specifically to the vector register portions 120-1 to 120-8. The plurality of multiplexers 810-1 to 810-8 are configured to respectively receive the bitwise-inverted X-bit binary portions from the plurality of bitwise inversion / NOT logic gates 610-1 to 610-8. The plurality of multiplexers 810-1 to 810-8 are configured to further respectively receive a plurality of second X-bit binary portions from the vector register portions 120-1 to 120-8 of the second vector register 120.
[0071] The plurality of multiplexers 810-1 to 810-8 further respectively receive binary values from adjacent multiplexers 810 and output the binary values to other adjacent multiplexers 810. For example, multiplexer 810-2 receives a binary value from multiplexer 810-3 and outputs the binary value to multiplexer 810-1, multiplexer 810-3 receives a binary value from multiplexer 810-4 and outputs the binary value to multiplexer 810-2, multiplexer 810-4 receives a binary value from multiplexer 810-5 and outputs the binary value to multiplexer 810-3, and so on. The multiplexers 810-1 to 810-8 also respectively output the multiplexed binary values to adders 140-1 to 140-8. Multiplexer 810-1 is different from the other multiplexers in that it only outputs to adder 140-1 and does not output to another multiplexer 810. Similarly, multiplexer 810-8 is different from the other multiplexers 810 in that multiplexer 810-8 receives a control bit 811 that controls whether the processing unit 800 is performing addition or subtraction.
[0072] The processing unit 800 even further includes another multiplexer, multiplexer 820. Multiplexer 820 receives a binary "1" on a first input (the MSB 111 from the first vector register 110). Multiplexer 820 further receives a control bit 821 that controls whether the processing unit 800 is processing an addition operation (no chaining) or a chaining operation. If the processing unit 800 is configured to perform a chaining operation, then multiplexer 820 processes the MSB111 from the first vector register 110; otherwise multiplexer 820 processes a binary "1" on its other input.
[0073] The processing unit 800 is configured to generate / propagate carry bits to add and subtract two 512b large numbers A+ / -B. However, the generation of carry bits for addition and subtraction of large numbers discussed above can be applied to smaller large numbers, e.g., >64 bits but <512 bits. Figure 9A and Figure 9B (collectively FIG. 9) illustrates an exemplary processing unit 900 that can perform carry generation / propagation for adding and subtracting 2x256b large numbers.
[0074] The processing unit 900 calculates the result of A + B for each of the two values, including the carry bit of the two. The processing unit 900 includes some of the same features shown in FIG. 8; these same features will not be discussed for the sake of brevity. The processing unit 900 includes a data path that calculates the first A + / - B + carry using vector channels 901-1 to 901-4 and calculates the second A + / - B + carry using vector channels 901-5 to 901-8, as discussed above in the context of the processing unit 800. Therefore, since the processing unit 900 can simultaneously perform the generation / propagation of carry bits for two addition / subtraction operations; thus the unique features from the processing unit 800 are replicated.
[0075] To simultaneously generate / propagate the carry bits for these two addition / subtraction operations, the processing unit 900 includes a second copy of the carry bit generation circuit 250-1. This second copy of the carry bit generation circuit 250-1 is shown in the vector channel 901-5 of the carry bit generation circuit 950-5. Thus, the vector channels 901 to 904 generate the carry bits for the first 256b large number addition / subtraction, and the vector channels 901-5 to 901-8 generate the carry bits for the second 256b large number addition / subtraction, doubled as discussed above for FIG. 8, where the processing unit 800 generates the carry bits for a single 512b large number addition / subtraction. Similarly, instead of using a single MSB in the MSB 111 to generate the carry bits stored in the least significant bits 113-1 to 113-8 of the vector channels 101-1 to 101-8, the processing unit 900 uses the MSB 111 to generate the carry bits stored in the least significant bits 113-1 to 113-4 of the vector channels 901-1 to 901-4, and uses another significant bit 911 from the vector channel 911-5 to generate the carry bits stored in the least significant bits 113-1 to 113-4 of the vector channels 901-1 to 901-4. To perform 128b or 256b subtraction, all the initial carry input bits of the large numbers are set to 1, as shown in the 2x256b subtraction example in FIG. 9.
[0076] Figure 10A and Figure 10B(Collectively referred to as FIG. 10) shows a processing unit 800 configured to perform large number 512b addition / subtraction as illustrated in FIG. 8. For simplicity of illustration, the carry bit generation circuits 250-1 to 250-8 shown in FIG. 8 are excluded from the illustration respectively. After the carry bits have been generated by the processing unit 800 and stored in the vector register portions 110-1 to 110-8 of the vector register 110, the processing unit 800 includes a data path through which the addition / subtraction of large number A and large number B can be performed by adders 140-1 to 140-8. The combination of multiplexers 810-1 to 810-8 and bitwise inversion / NOT logics 610-1 to 610-8 controls whether the bitwise inverted version of large number B is received by adders 140-1 to 140-8 to perform subtraction, while the non-bitwise inverted version of large number B performs addition.
[0077] The new ISA instruction (complete_wide_add / sub) added to the existing ISA instruction set to utilize adders 140-1 to 140-8 to perform addition / subtraction has three input operands: (1) the carry bits 257-1 to 257-8 stored in the least significant bits 113-1 to 113-8 of each vector channel 801-1 to 801-8 by the first vector register 110, where the MSB 111 of the first vector register 110 used to link 512b addition / subtraction during carry bit generation is ignored, (2) the 512b value large number A stored in the second vector register 120, and (3) the 512b large number value B stored in the third vector register 130.
[0078] Adders 140-1 to 140-8 receive the portions of large number A and large number B from vector channels 801-1 to 801-8 respectively. Adders 140-1 to 140-8 also receive the previously generated carry bits, such as the carry bits previously stored in the least significant bits 113-1 to 113-8 of each vector channel 801-1 to 801-8. Then adders 140-1 to 140-8 add / subtract large number A and large number B respectively using the carry bits from each of the vector channels 801-1 to 801-8 to arrive at the appropriate portions of the sum of large number A and large number B. Then the processing unit 800 updates the second vector register 120 with the correct sum / difference of this operation. In some embodiments, the first vector register 110 should not be rewritten so that it can be saved to provide a "carry output" for chaining subsequent additional 512b addition / subtractions. Therefore, to facilitate large number addition / subtraction using existing 64b adders, such as in a single instruction, multiple data (SIMD) data path, the existing 64b adders are augmented with a "carry input" or by using carry bits fed by the bits of the first vector register 110.
[0079] Figure 11A and Figure 11B (collectively referred to as FIG. 11) illustrate another configuration of the exemplary processing unit 1100, which is built on the configuration of the processing unit 200 shown in FIG. 2 to add two large binary values without first storing the carry in a vector register. If the MSB 111 from the first vector register 110 (“zmm1”) is not required to chain the large number calculation to a larger integer (i.e., when all operations in the calculation require <= 512b addition), then the first vector register 110 is not required to store the carry bit to the next stage. In this case, the processing unit 1100 can be used to include a data path that uses a single instruction (e.g., zmm1 = zmm2 + zmm3) to perform both carry generation and subsequent large number addition.
[0080] The processing unit 1100 includes all the hardware shown in FIG. 6; these same features will not be discussed for the sake of brevity. However, instead of the carry bit generation circuits 250-1 to 250-8 respectively outputting the carry bits to the first vector register 110, specifically the vector register portions 110-1 to 110-8, as described above for FIG. 2, the processing unit 1100 utilizes the carry bit generation circuits 250-1 to 250-8, which output the carry bits to adders, such as adders 1140-1 to 1140-8. The adders 1140-1 to 1140-8 receive the large number A, the large number B, and the carry bits from the carry bit generation circuits 250-1 to 250-8. The adders 1140-1 to 1140-8 can then use the carry bits received from the carry bit generation circuits 250-1 to 250-8 to add the large number A + the large number B to arrive at the correct sum of the large number A + the large number B.
[0081] The processing unit 1100 shows two adders 140-1 / 1140-1, 140-2 / 1140-2, 140-3 / 1140-3, 140-4 / 1140-4, 140-5 / 1140-5, 140-6 / 1140-6, 140-7 / 1140-7, 140-8 / 1140-8 (e.g., 64b+ adder blocks) respectively for each vector channel 1101-1 to 1101-8. In one embodiment, the adders 140-1 / 1140-1, 140-2 / 1140-2, 140-3 / 1140-3, 140-4 / 1140-4, 140-5 / 1140-5, 140-6 / 1140-6, 140-7 / 1140-7, 140-8 / 1140-8 are different from each other, where the processing unit 1100 utilizes two adders respectively for each vector channel 1101-1 to 1101-8. In another embodiment, the adder 140 is the same component as the adder 1140, where the adder 140 is reused in the pipeline manner in the instruction. Thus, in this embodiment, the same adder 140 / 1140 will be used to generate the carry bit and add the large numbers A+B both.
[0082] Figure 12 An exemplary critical path for chained 512b addition is shown. Even though the 512b addition operation uses two instructions, chaining them together to perform addition on larger integers can be faster in terms of addition / loop because the critical path for performing the calculation is only one instruction (generate_carry_512).
[0083] Figure 13 An exemplary method 1300 for generating a carry bit and adding a first binary value to a second binary value is illustrated. The method starts at block 1310. At block 1310, the first X-bit binary partial values of the first Y-bit binary value and the second Y-bit binary value and the second X-bit binary partial values are added, and a first carry bit is generated, where Y is a multiple of X. In some embodiments, the adders 140-1 to 140-8 are used to add the first X-bit binary partial values of the first Y-bit binary value and the second Y-bit binary value and the second X-bit binary partial values. According to the example given above, the first Y-bit binary value can be the 512b large number A and the second Y-bit binary value can be the 512b large number B. In some embodiments, the first carry bit can be the carry bits 151-1 to 151-8.
[0084] At block 1320, a second carry bit is generated based on the first carry bit. In some embodiments, the carry bit generation circuits 150-1 to 150-8, 250-1 to 250-8 are used to generate the second carry bit. In some embodiments, the second carry bit can be the carry bits 154-1 to 154-8, 257-1 to 257-8.
[0085] At block 1330, the first X-bit binary parts of the first Y-bit binary value and the second Y-bit binary value are added to the second X-bit binary part using the second carry bit respectively. In some embodiments, adders 140-1 to 140-8 use carry bits 154-1 to 154-8, 257-1 to 257-8 to add 64b parts of 512b large numbers A, B. Adders 140-1 to 140-8 can receive carry bits 154-1 to 154-8, 257-1 to 257-8 from vector register parts 110-1 to 110-8 of the first vector register 110 respectively.
[0086] Processing units 100 to 1100 can chain 512b operations to perform addition on large numbers even larger than 512b (e.g., 1024b). The following pseudocode can be used to chain two 512b additions to perform 1024b addition:
[0087] mov zmm1,[zeros]; set initial carry-in to zero
[0088] mov zmm2,[ALO]; lower 512b of 1024b operand A
[0089] mov zmm3,[BLO]; lower 512b of 1024b operand B
[0090] generate_carry_512zmm1,zmm2,zmm3; zmm1 = carry bits
[0091] complete_wide_add zmm2,zmm3,zmm1; zmm2 = sum
[0092] mov zmm4,[AHI]; upper 512b of 1024b operand A
[0093] mov zmm5,[BHI]; upper 512b of 1024b operand B
[0094] generate_carry_512zmm1,zmm4,zmm5; zmm1 = carry bits
[0095] complete_wide_add zmm4,zmm5,zmm1; zmm4 = sum
[0096] ; DONE: lower 512b sum in zmm2, upper 512b in zmm4
[0097] ; zmm1 holds carryout of 1024b add if addition additions are to be performed
[0098] In some embodiments, the above-described apparatus and techniques are implemented in a system that includes one or more integrated circuit (IC) devices (also referred to as integrated circuit packages or microchips), such as the processing units 100 to 1100 described above with reference to FIGS. 1 to 11. Electronic design automation (EDA) and computer-aided design (CAD) software tools can be used in the design and manufacture of these IC devices. These design tools are typically represented as one or more software programs. The one or more software programs include code executable by a computer system to manipulate the computer system to operate on code representing a circuit of one or more IC devices in order to perform at least a portion of a process for designing or tuning a manufacturing system to manufacture the circuit. The code can include instructions, data, or a combination of instructions and data. Software instructions representing design tools or manufacturing tools are typically stored in a computer-readable storage medium accessible by a computing system. Similarly, code representing one or more stages of the design or manufacture of an IC device can be stored in and accessed from the same computer-readable storage medium or different computer-readable storage media.
[0099] A computer-readable storage medium can include any non-transitory storage medium or combination of non-transitory storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media can include, but are not limited to, optical media (e.g., compact discs (CDs), digital versatile discs (DVDs), Blu-ray discs), magnetic media (e.g., floppy disks, magnetic tapes, or magnetic hard disk drives), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or microelectromechanical systems (MEMS)-based storage media. The computer-readable storage medium can be embedded in the computing system (e.g., system RAM or ROM), fixedly attached to the computing system (e.g., magnetic hard disk drive), removably attached to the computing system (e.g., optical disc or universal serial bus (USB)-based flash memory), or coupled to the computer system via a wired or wireless network (e.g., network-attached storage device (NAS)).
[0100] In some embodiments, certain aspects of the techniques described above can be implemented by one or more processors of a processing system executing software. The software includes one or more sets of executable instructions that are stored on or otherwise tangibly embodied in a non-transitory computer-readable storage medium. The software can include instructions and certain data that, when executed by one or more processors, manipulate the one or more processors to perform one or more aspects of the techniques described above. The non-transitory computer-readable storage medium can include, for example, a magnetic or optical disk storage device, a solid state storage device such as flash memory, a cache, a random access memory (RAM), or other one or more non-volatile memory devices. The executable instructions stored on the non-transitory computer-readable storage medium can be in source code, assembly language code, object code, or other instruction formats that are interpreted or otherwise executable by one or more processors.
[0101] Note that not all activities or elements described above in the general description are necessary, that a portion of a particular activity or device may not be necessary, and that one or more additional activities can be performed or elements can be included in addition to those described. Further, the order in which the listed activities are listed is not necessarily the order in which they are performed. Additionally, these concepts have been described with reference to specific embodiments. However, one of ordinary skill in the art understands that various modifications and changes can be made without departing from the scope of the present disclosure as set forth in the following claims. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive, and all such modifications are intended to be included within the scope of the present disclosure.
[0102] Benefits, other advantages, and solutions to problems have been described above with respect to specific embodiments. However, the benefits, advantages, solutions to problems, and any feature that may cause any benefit, advantage, or solution to occur or become more pronounced should not be construed as a critical, required, or essential feature of any or all of the claims. Further, the particular embodiments disclosed above are illustrative only, as the disclosed subject matter can be modified and practiced in different but equivalent manners apparent to those of ordinary skill in the art that benefit from the teachings herein. The details of the construction or design shown herein are not intended to be limiting other than as described in the following claims. Thus, it is evident that the specific embodiments disclosed above can be varied or modified, and all such variations are considered to be within the scope of the disclosed subject matter. Accordingly, the protection sought herein is as set forth in the following claims.
Claims
1. A processing unit, the processing unit comprises: a plurality of adders that add a first X-bit binary partial value and a second X-bit binary partial value of a first Y-bit binary value and a second Y-bit binary value to generate a first carry bit, where Y is a multiple of X; and a plurality of carry bit generation circuits respectively coupled to the plurality of adders to receive the first carry bit and generate a second carry bit based on the first carry bit; wherein the second carry bit is used to add the first X-bit binary part and the second X-bit binary part of the first Y-bit binary value and the second Y-bit binary value respectively.
2. The processing unit according to claim 1, wherein the plurality of carry bit generation circuits are configured to respectively receive the sum of the first X-bit binary part and the second X-bit binary part of the first Y-bit binary value and the second Y-bit binary value, and are configured to generate the second carry bit based on the sum of the first X-bit binary part and the second X-bit binary part of the first Y-bit binary value and the second Y-bit binary value.
3. The processing unit according to claim 1 or claim 2, wherein the first Y-bit binary value and the second Y-bit binary value are stored by at least one of a vector register, a memory operand, a double data rate (DDR) memory, a low power DDR (LPDDR), and Gen-Z.
4. The processing unit according to any one of claims 1 to 3, wherein a new instruction set architecture (ISA) instruction is added to an existing ISA instruction set to add the first X-bit binary part and the second X-bit binary part of the first Y-bit binary value and the second Y-bit binary value.
5. The processing unit according to any one of claims 1 to 4, wherein the plurality of carry bit generation circuits comprises: a plurality of AND logic gates respectively, a plurality of 1-bit full adders respectively; and a plurality of XOR logic gates respectively; wherein the plurality of AND logic gates are configured to: respectively receive the sum of the first X-bit binary part and the second X-bit binary part of the first Y-bit binary value and the second Y-bit binary value from the plurality of adders; and output binary values to the plurality of 1-bit full adders and the plurality of XOR logic gates respectively; the plurality of 1-bit full adders are configured to: receive the binary values output from the plurality of AND logic gates, the first carry bit from the plurality of adders, and binary values from adjacent 1-bit full adders; and output binary values; and the plurality of XOR logic gates are configured to: receive the binary values output by the plurality of AND logic gates and the binary values output by the plurality of 1-bit full adders; and output the second carry bit to a vector register.
6. The processing unit according to any one of claims 1 to 4, wherein the plurality of carry bit generation circuits comprises: a plurality of AND logic gates; and a plurality of 1-bit vector carry propagation (VCP) logic circuits; wherein The plurality of AND logic gates are configured to: receive the sum of the first X-bit binary portions and the second X-bit binary portions of the first Y-bit binary value and the second Y-bit binary value from the plurality of adders; and output binary values to the plurality of 1-bit vector carry propagation (VCP) logic circuits, respectively; and The plurality of 1-bit VCP logic circuits are configured to: receive the binary values from the plurality of AND logic gates, the first carry bit from the plurality of adders, and the binary values output by adjacent 1-bit VCP logic circuits among the plurality of 1-bit VCP logic circuits; output other binary values to other adjacent 1-bit VCP logic circuits among the plurality of 1-bit VCP logic circuits; and output the second carry bit to the vector register.
7. The processing unit according to claim 6, wherein the plurality of AND logic gates are a plurality of first AND logic gates, and the 1-bit VCP logic circuits comprise: a plurality of OR logic gates respectively; and a plurality of second AND logic gates respectively; wherein the plurality of OR logic gates are configured to respectively receive the binary values output from the plurality of first AND logic gates and the first carry bit from the plurality of adders; and wherein the plurality of second AND logic gates are configured to: receive the first carry bit from the plurality of adders and the binary values output by the plurality of OR logic gates; and output the second carry bit to the vector register respectively.
8. The processing unit according to claim 6, wherein the vector register is a first vector register, and the processing unit further comprises: a plurality of bitwise inversion / NOT logic gates, the plurality of bitwise inversion / NOT logic gates are coupled to a second vector register, and are respectively configured to receive the second X-bit binary portion, perform bitwise inversion on the second X-bit binary portion of the second Y-bit binary value, and output the bitwise-inverted X-bit binary portion of the second X-bit binary portion to the plurality of carry bit generation circuits.
9. The processing unit according to claim 8, wherein the plurality of carry bit generation circuits further comprise: a plurality of multiplexers, the plurality of multiplexers are coupled to a second vector register, and are configured to respectively receive the bitwise-inverted X-bit binary portion of the second Y-bit binary value from the plurality of bitwise inversion / NOT logic gates, receive the second X-bit binary portion, receive binary values from adjacent multiplexers, and output binary values to other adjacent multiplexers.
10. A system comprising the processing unit according to any one of claims 1 to 9, the system comprising another processing unit, the another processing unit comprising the plurality of adders and the plurality of carry bit generation circuits.
11. The processing unit according to claim 1, the processing unit further comprises: A vector register that stores the sum of the first X-bit binary parts and the second X-bit binary parts of the first Y-bit binary value and the second Y-bit binary value, respectively.
12. The processing unit according to claim 11, wherein the plurality of adders are a first plurality of adders, and the processing unit comprises: A second plurality of adders that respectively add the first Y-bit binary value and the second Y-bit binary value from the first vector register and the second vector register.
13. The processing unit according to claim 12, wherein the plurality of adders add the first Y-bit binary value and the second Y-bit binary value from the first vector register and the second vector register, respectively.
14. A method, the method comprises: Adding the first X-bit binary part value and the second X-bit binary part value of the first Y-bit binary value and the second Y-bit binary value by a plurality of adders to generate a first carry bit, where Y is a multiple of X; Generating a second carry bit based on the first carry bit by a plurality of carry bit generation circuits; and Using the second carry bit to add the first X-bit binary part and the second X-bit binary part of the first Y-bit binary value and the second Y-bit binary value, respectively.
15. The method according to claim 14, the method further comprises: Receiving, by the plurality of carry bit generation circuits, the sum of the first X-bit binary part and the second X-bit binary part of the first Y-bit binary value and the second Y-bit binary value, respectively; and Generating, by the plurality of carry bit generation circuits, the second carry bit based on the sum of the first X-bit binary part and the second X-bit binary part of the first Y-bit binary value and the second Y-bit binary value.
16. The method according to claim 14 or claim 15, the method further comprises storing the first Y-bit binary value and the second Y-bit binary value by at least one of a vector register, a memory operand, a double data rate (DDR) memory, a low power DDR (LPDDR), and Gen-Z.
17. The method according to any one of claims 14 to 16, wherein each carry bit generation circuit in the carry bit generation circuits comprises a plurality of AND logic gates; a plurality of 1-bit full adders; and a plurality of XOR logic gates, and the method further comprises: Receiving, by the plurality of AND logic gates, the sum of the first X-bit binary part and the second X-bit binary part of the first Y-bit binary value and the second Y-bit binary value from the plurality of adders, respectively; Outputting binary values to the plurality of 1-bit full adders and the plurality of XOR logic gates by the plurality of AND logic gates, respectively; Receiving, by the plurality of 1-bit full adders, the binary values output from the plurality of AND logic gates, the first carry bit from the plurality of adders, and the binary values from adjacent 1-bit full adders; Outputting binary values by the plurality of 1-bit full adders, respectively; The plurality of XOR logic gates respectively receive the binary values output by the plurality of AND logic gates and the binary values output by the plurality of one-bit full adders; and The plurality of XOR logic gates output the second carry bit to a vector register.
18. The method according to any one of claims 14 to 16, wherein each carry bit generation circuit in the carry bit generation circuit includes a plurality of AND logic gates and a plurality of one-bit vector carry propagation (VCP) logic circuits, and the method further includes: Receiving, by the plurality of AND logic gates respectively, the sum of the first X-bit binary parts and the second X-bit binary parts of the first Y-bit binary value and the second Y-bit binary value from the plurality of adders; Outputting, by the plurality of AND logic gates respectively, binary values to the plurality of one-bit vector carry propagation (VCP) logic circuits; Receiving, by the plurality of one-bit VCP logic circuits respectively, the binary values from the plurality of AND logic gates, the first carry bits from the plurality of adders, and the binary values output by adjacent one-bit VCP logic circuits among the plurality of one-bit VCP logic circuits; and Outputting, by the plurality of one-bit VCP logic circuits, other binary values to other adjacent one-bit VCP logic circuits among the plurality of one-bit VCP logic circuits, and outputting the second carry bit to a vector register.
19. The method according to claim 14, the method further includes: Storing, by a vector register respectively, the sum of the first X-bit binary parts and the second X-bit binary parts of the first Y-bit binary value and the second Y-bit binary value.
20. A processing unit, the processing unit includes: A first vector register that stores a first Y-bit binary value including a plurality of first X-bit binary parts, where Y is a multiple of X; A second vector register that stores a second Y-bit binary value including a plurality of second X-bit binary parts; A plurality of adders that add the first X-bit binary part values and the second X-bit binary part values of the first Y-bit binary value and the second Y-bit binary value to generate a first carry bit; and A plurality of carry bit generation circuits respectively coupled to the plurality of adders to receive the first carry bit and generate a second carry bit based on the first carry bit; wherein the plurality of adders respectively add the first X-bit binary parts and the second X-bit binary parts of the first Y-bit binary value and the second Y-bit binary value by using the second carry bit.