Reduction operation mapping system and method

By truncating the LSB of operands in the adder tree and introducing the trailing adder tree to process the truncated bits, the problem of low packaging efficiency on FPGA in the prior art is solved, and efficient and accurate arithmetic operations are achieved.

CN109254755BActive Publication Date: 2025-05-27ALTERA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201810612332.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-12-14
Filing Date
2018-06-14
Publication Date
2025-05-27
Estimated Expiration
2038-06-14

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently map the sum of multiple operands to programmable devices such as field programmable gate arrays (FPGAs), especially in the case of high precision and multiple operands, resulting in large area occupancy of integrated circuits and limited logical resources.

Method used

Improve the packaging efficiency of logical array blocks by building an adder tree and truncating the least significant bit (LSB) of operands at each stage node. At the same time, a trailing adder tree is introduced to process the truncated LSB, mitigating errors and improving accuracy.

Benefits of technology

It realizes more efficient packaging of arithmetic operations on integrated circuits, reduces the use of logical resources, and improves the accuracy and performance of arithmetic operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN109254755B_ABST
    Figure CN109254755B_ABST
Patent Text Reader

Abstract

An adder tree can be formed for efficient packing of arithmetic units into an integrated circuit. The operands of the tree can be truncated to pack an integer number of nodes in a logic array block. Thus, arithmetic operations can be more efficiently packed onto the integrated circuit, and increased precision and performance are provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is a non-provisional application claiming priority to U.S. Provisional Patent Application No. 62 / 532,871, filed on July 14, 2017, entitled “Reduction Operation Mapping Systems and Methods,” which is incorporated herein by reference. Technical Field

[0003] The present disclosure relates generally to integrated circuit devices, and more particularly to increasing the efficiency of mapping reduction operations (e.g., sums of multiple operands) onto programmable devices (e.g., field programmable gate array (FPGA) devices). In particular, the current disclosure relates to dot products based on small precision multiplications for machine learning operations. Background Art

[0004] This section is intended to introduce the reader to various aspects of the art that may be related to various aspects of the present disclosure, which are described and claimed below. This discussion is considered helpful in providing the reader with background information to facilitate a better understanding of various aspects of the present disclosure. Accordingly, it should be understood that these statements are to be read in this light, and not as admissions of prior art.

[0005] Machine learning is becoming an increasingly valuable application area. For example, it can be used in natural language processing, object recognition, bioinformatics, and economics, along with other fields and applications. The implementation of machine learning may involve large arithmetic operations, such as the summation of many operands. However, large arithmetic operations are difficult to fit into integrated circuits (e.g., FPGAs) (which can implement machine learning). It can be particularly difficult to install arithmetic operations on an integrated circuit, for example, when the operands have high precision, there are many operand requirements and / or there is a high percentage of logic used in the device for arithmetic operations. Therefore, the summation of many operands that may be involved in machine learning may involve a large part of the area of ​​the integrated circuit. For this reason, due to the physical layout and the way in which logic resources can be used in such designs, the available logic for dense arithmetic designs may be limited. For example, in some arithmetic designs, soft logic resources containing adder resources (e.g., adders) for performing arithmetic functions are often grouped together. Therefore, if the ripple carry adder for a specific node in the adder tree occupies more than half of the soft logic group, the remaining logic in the grouping may not be available for similar nodes. As a result, much of the logic in the integrated circuit may be inaccessible. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Various aspects of the disclosure may be better understood upon reading the following detailed description and upon reference to the accompanying drawings.

[0007] Figure 1 is a block diagram of a system for implementing arithmetic operations according to an embodiment;

[0008] Figure 2 is a block diagram of an integrated circuit in which arithmetic operations may be performed according to an embodiment;

[0009] Figure 3 is a block diagram of a data processing system in which an integrated circuit may be implemented according to an embodiment;

[0010] Figure 4 is a block diagram of an adder tree in which arithmetic operations may be performed according to an embodiment;

[0011] Figure 5 is a block diagram of a second embodiment of an adder tree;

[0012] Figure 6 According to an embodiment Figure 5 Block diagram of the adder tree and the trailing adder tree;

[0013] Figure 7 is a block diagram of a second embodiment of a trailing adder tree;

[0014] Figure 8 According to an embodiment Figure 5 The sum of the adder tree Figure 7 Block diagram of the sum of the trailing adder tree;

[0015] Fig. 9 is a block diagram of a third embodiment of an adder tree;

[0016] Fig.10 According to one embodiment, determining involves truncating Fig. 9 A flowchart of the total average truncated value of the operands in the adder tree;

[0017] Fig.11 According to one embodiment, Fig. 9 A diagram of the static distribution of bits truncated by operands in an adder tree of ;

[0018] Fig.12 According to an embodiment, Fig. 9 A block diagram of a method for determining a dynamic distribution of bits truncated from operands in an adder tree and a total average truncated value;

[0019] Fig.13 is a block diagram of an adder tree node according to an embodiment;

[0020] Fig.14 is a block diagram of an adder tree node implementing a compressor structure according to an embodiment;

[0021] Fig.15 is a block diagram of a block floating point tree according to an embodiment;

[0022] Fig.16 is a block diagram of a simplified block floating point tree according to one embodiment; and

[0023] Fig.17 is a block diagram of a block floating point combination tree according to an embodiment. DETAILED DESCRIPTION

[0024] One or more specific embodiments will be described below. In an effort to provide a brief description of these embodiments, not all features of an actual implementation are described in this specification. It should be understood that in the development of any such actual implementation, as in any engineering or design project, numerous implementation-specific decisions must be made to achieve the developer's specific goals, such as compliance with system-related and business-related constraints (which may vary from implementation to implementation). In addition, it should be understood that such development work may be complex and time-consuming, but is still a routine matter of design, fabrication, and manufacturing for those skilled in the art who benefit from this disclosure.

[0025] Machine learning becomes a valuable use case for integrated circuits (e.g., field programmable gate arrays, also known as FPGAs), and may utilize one or more arithmetic operations (e.g., reduction operations). To perform arithmetic operations, an integrated circuit may contain a logic array block (LAB), which may include multiple adaptive logic modules (ALMs) and / or other logic elements. The ALM may include resources (e.g., multiple lookup tables (LUTs), adders, carry chains, and the like) such that each ALM and subsequent LABs including the ALMs may be configured to implement arithmetic functions. Thus, machine learning implementations may, for example, utilize an FPGA with a LAB in order to perform multiple arithmetic operations. In such cases, the LAB may use its ALM resources to aggregate multiple operands. Furthermore, although the operand sizes involved in machine learning are generally relatively small, many parallel reduction operations may be implemented, which may utilize large portions of the integrated circuit.

[0026] Therefore, according to certain embodiments of the present disclosure, an adder tree may be constructed for efficient packing of arithmetic operators into an integrated circuit. The operands of the tree may be truncated (e.g., pruned) to pack an integer number of nodes per logic array block. In addition, the techniques described herein also provide a mechanism for determining possible errors, and a hardware structure for automatically mitigating errors (involving truncated operands). Therefore, arithmetic operations may be more efficiently packed onto an integrated circuit with increased precision and performance.

[0027] In view of the above, Figure 1A block diagram of a system 10 implementing arithmetic operations is shown. A designer may desire to implement functionality on an integrated circuit 12 (which may include, for example, an FPGA, an application specific integrated circuit (ASIC), a system on a chip (SoC), or the like). The designer may specify a program to be implemented, which may enable the designer to provide programming instructions to implement a circuit design for the integrated circuit 12. For example, the designer may specify programming instructions to configure or partially reconfigure a region of the integrated circuit 12.

[0028] The designer may use design software 14 (such as a version of Quartus made by Intel Corporation) to implement the level design. Design software 14 may use compiler 16 to convert the program into a low-level program. Compiler 16 may provide machine-readable instructions representing the program to host 18 and integrated circuit 12. In an example in which integrated circuit 12 includes an FPGA structure, integrated circuit 12 may receive one or more kernel programs 20, which describe the hardware implementation that should be programmed into the programmable structure of the integrated circuit. Host 18 may receive host program 22, which may be implemented by kernel program 20. In order to implement host program 22, host 18 may pass instructions from host program 22 to integrated circuit 12 via communication link 24, which may be, for example, direct memory access (DMA) communication or high-speed peripheral component interconnect (PCIe) communication. In some embodiments, kernel program 20 and host 18 may be able to implement the configuration of LAB 26 on integrated circuit 12. LAB 26 may include multiple ALMs and / or other logic elements, and may be configured to implement arithmetic functions.

[0029] exist Figure 2 In one example shown in FIG. 4 , the integrated circuit 12 may include a programmable logic device, such as a field programmable gate array (FPGA) 40. For the purpose of this example, the device is referred to as an FPGA 40, but it should be understood that the device may be any type of logic device (e.g., an application specific integrated circuit (ASIC) and / or an application specific standard product (ASSP)). As shown, the FPGA 40 may have an input / output circuit 42 for driving signals away from the device 40 and for receiving signals from other devices via input / output pins 44. Interconnection resources 46, such as global and local vertical and horizontal wires and buses, may be used to route signals on the FPGA 40. In addition, the interconnection resources 46 may include fixed interconnects (wires) and programmable interconnects (i.e., programmable connections between corresponding fixed interconnects). The programmable logic 48 may include combinational and sequential logic circuits. For example, the programmable logic 48 may include a lookup table, a register, and a multiplexer. In various embodiments, the programmable logic 48 may be configured to perform custom logic functions. The programmable interconnect associated with the interconnection resources may be understood as part of the programmable logic 48.

[0030] A programmable logic device, such as FPGA 40, may contain a programmable element 50 having programmable logic 48. For example, as discussed above, a designer (e.g., a customer) may program (e.g., configure) programmable logic 48 to perform one or more desired functions. By way of example, FPGA 40 may be programmed by configuring programmable element 50 using mask programming arrangements, which is performed during semiconductor manufacturing. In another example, FPGA 40 may be configured after semiconductor fabrication operations have been completed (e.g., by programming programmable element 50 using electrical programming or laser programming). In general, programmable element 50 may be based on any suitable programmable technology, such as fuses, anti-fuses, electrically programmable read-only memory technology, random access memory cells, mask-programmed elements, and the like.

[0031] FPGA 40 can be electrically programmable. Through the electrical programming arrangement, programmable element 50 can be formed by one or more memory cells. For example, during programming, configuration data is loaded into the memory cell using input / output pin 44 and input / output circuit 42. In one embodiment, the memory cell can be implemented as a random access memory (RAM) cell. The use of the memory cell based on RAM technology described herein is intended to be just an example. In addition, because these RAM cells are loaded with configuration data during programming, they are sometimes referred to as configuration RAM cells (CRAM). Each of these memory cells can provide a corresponding static control output signal, which controls the state of the associated logic component in the programmable logic 48. For example, in some embodiments, the output signal can be applied to the gate of the metal oxide semiconductor (MOS) transistor in the programmable logic 48.

[0032] The circuits of FPGA 40 may be organized using any suitable architecture. As an example, the logic of FPGA 40 may be organized as a series of rows and columns of larger programmable logic areas, each of which may contain multiple smaller logic areas. The logic resources of FPGA 40 may be interconnected by interconnection resources 46 (e.g., associated vertical and horizontal conductors). For example, in some embodiments, these conductors may include global wires that span substantially all of FPGA 40, fractional wires (such as half-wires or quarter-wires) that span a portion of FPGA 40, staggered wires of a specific length (e.g., sufficient to interconnect several logic areas), smaller local wires, or any other suitable interconnection resource arrangement. In addition, in other embodiments, the logic of FPGA 40 may be arranged in more levels or layers, where multiple large areas are interconnected to form larger portions of logic. Further, other device arrangements may use logic that is not arranged in a manner other than rows and columns.

[0033] Figure 3Data processing system 100 is shown, which may be an example of one of many electronic devices in which integrated circuit 12 (such as FPGA 40) may be used. Data processing system 100 may include processor 101, memory 102, input / output (I / O) ports 103, peripheral devices 104, and / or additional or fewer components. The components in data processing system 100 may be coupled together by system bus 105 and loaded on circuit board 106 (which may be contained in end-user system 107).

[0034] Data processing system 100 may be used in an array of applications. For example, it may be used in computer or data networking, instrumentation, video or digital signal processing, or in other applications where programmable logic may find advantages. Integrated circuit 12 may be used within data processing system 100 to perform logic functions. Integrated circuit 12 may be configured as a processor or controller, may cooperate with processor 101, and / or integrated circuit 12 may interface between processor 101 and other components in data processing system 100, among other examples.

[0035] Now go to Figure 4 In some embodiments, the ALM and / or LAB 26 may perform summation of multiple operands via the adder tree 200, which may aggregate the operands 201 at several nodes 207 and / or stages 208 until a final sum is generated. In the illustrated embodiment, for example, the first adder tree 210 may include input circuitry to receive four operands 201 in the first stage 208A, which ultimately aggregate into a final sum 206. The operands 201 may be paired together at two nodes 207 to form two sets of operands, where the sets of adders (e.g., adder circuits) may be individually aggregated into two intermediate results 203. Additional adders may aggregate the two intermediate results 203 of a single node 207 in the second stage 208B of the first adder tree 210 into the final sum 206 in the third stage 208C of the first adder tree 210 to complete the arithmetic operation of summing all four operands 201. Although an adder is not shown in the illustrated embodiment, it should be understood that the result of summing two or more operands 201 (eg, 203, 206) may be obtained using an adder (eg, an adder circuit).

[0036] The addition of two operands 201 may generate a result having more bits than any one of the operands 201. For example, if the addition involves a carry operation that affects the most significant bit (MSB) of the first 6-bit operand or the second 6-bit operand, the addition of the first 6-bit operand and the second 6-bit operand may generate a 7-bit result. Therefore, in some cases, each of the intermediate results 203 and the final sum 206 of the adder tree 200 may contain additional bits compared to the operands 201 in the previous stage (e.g., 208A, 208B). For example, an adder tree 200 with several stages may generate a final sum 206 having more bits than any one of the set of operands 201 input to the adder tree 200. Therefore, the growth of the result of the arithmetic operation in the adder tree may involve the use of additional resources and / or space on the integrated circuit 12, and may also adversely affect the packing efficiency of the integrated circuit 12.

[0037] Thus, in some embodiments, in order to control the growth of the results of arithmetic operations, and therefore control packing, the operands 201 may be truncated (e.g., compacted) at each stage (e.g., 208A-208C) or a subset of the stages 208. For example, by truncating the operands 201 in the first adder tree 210, the adder tree 200 may be more efficiently packed onto the integrated circuit 12. In the illustrated embodiment, for example, each operand 201 is 6 bits wide, and the 6-bit size is carried forward through the first adder tree 210. Thus, in the illustrated embodiment, each of the intermediate results 203 and the final sum 206 are also 6 bits wide. To facilitate a constant 6-bit size, the least significant bit (LSB) of each operand 201 (e.g., 202, 204) may be discarded. For example, soft logic circuitry (e.g., logic within a LAB) may shift the operand 201 right by one bit to truncate the LSB. That is, in Figure 4 In the example shown in , only the upper 5 bits of each operand 201 can be used, resulting in the addition of two 5-bit operands 201 at each level (e.g., 208A, 208B) of the first adder tree 210. Therefore, the intermediate result 203 in the second level 208B has a 6-bit size, because as discussed, the arithmetic operation can result in a result with a single bit increase. Therefore, the LSB 204 of each intermediate result 203 can also be truncated before the addition operation of the second level 208B occurs. Therefore, the final sum 206 can have a 6-bit width. Accordingly, the first adder tree 210 is an illustrative example of a single-bit reduction tree, because a single bit is removed (e.g., reduced) from each operand 201 at each level 208.

[0038] As an additional illustrative example, an 8-bit by 3-bit signed multiplier (which may be used in a machine learning implementation) may have a signed output range of 10 bits. Accordingly, the input precision of the adder tree used to generate the product of the 8-bit by 3-bit signed multiplier may be 10 bits. In some embodiments, because the 20 adder bits are a logical grouping in the integrated circuit 12, maintaining 10 bits per operand 201 at the adder node 207 may allow two nodes 207 to be packed into each routing grouping (e.g., a soft logic grouping). Therefore, in order to maintain 10 bits per operand 201 and to generate a final sum having a width of 10 bits, each operand 201 at each node 207 of the adder tree may be right shifted a single bit (e.g., truncated), which may account for the increase in the single bit in the sum output by the node 207.

[0039] In other embodiments, truncation may involve removal of LSB groups 305 (e.g., sets of two or more LSBs) instead of or in addition to removing a single LSB. For example, in some embodiments, the second adder tree 300 may include input circuitry to receive an operand 201 having many bits (e.g., a large operand), such as Figure 5 As can be shown. Therefore, in order to improve packing efficiency, the soft logic may, for example, truncate the LSB group 305 (e.g., 302-304) from each of the operands 201. For example, the second adder tree 300 may manipulate the addition of four 5-bit operands 201 instead of the four 8-bit operands 201 in the first stage 208A. Therefore, the intermediate result 203 may include 6 bits, so in order to generate a final sum 206 having a width of 6 bits, the soft logic may, for example, truncate a single LSB 204 from each of the intermediate results 203. Therefore, at each stage (e.g., 208A-208C) of the adder tree (e.g., the second adder tree 300), a different number of LSBs may be truncated as needed to improve packing efficiency. In addition, although Figure 4 and Figure 5 Truncation of a single LSB and / or a group of LSBs 305 (shown as a set of three LSBs (eg, 302 - 304 )) is shown, but any number of LSBs may be truncated from the operand 201 at any suitable stage 208 of the adder tree.

[0040] The final sum 206 may suffer from errors when compared to the actual (e.g., full precision) result of a corresponding complete, non-truncated adder tree. For example, because the truncated LSBs (e.g., 202, 204, and 302-304) are not included in the final sum 206 of the illustrated adder tree (e.g., 210 and / or 300, respectively), the final sum 206 may differ from the actual result of the corresponding complete adder tree. The final sum 206 may not differ significantly from the actual result of the corresponding complete adder tree, but in some embodiments, a more accurate adder tree sum is beneficial.

[0041] Therefore, in order to improve the accuracy of the final sum 206 while maintaining efficient packing in the integrated circuit 12, the adder tree 200 can be divided into multiple trees. Figure 6 An embodiment is shown in which the second adder tree 300 is divided into a main tree 308 (e.g., the second adder tree 300) corresponding to the summation of the truncated operand 201 and a first trailing adder tree 310A corresponding to the summation of the LSB group 305 truncated from the operand 201 in the first stage 208A of the main tree 308. Therefore, the LSB group 305 can be summed separately from the truncated operand 201. In addition, depending on the number of LSBs 302-304 truncated from the operand 201, the first trailing adder tree 310A may or may not implement further truncation. For example, in some embodiments, if the size and / or number of the truncated LSB group 305 is large enough to be efficiently packed into the integrated circuit 12, the first trailing adder tree 310A may not implement truncation of the LSB group 305. In other embodiments, many and / or large LSB groups 305 in the first trailing adder tree 310A may result in truncation of the LSBs from the LSB group 305 itself.

[0042] also, Figure 6 A first trailing adder tree 310A may be shown for summing the LSB group 305 from the first stage 208A of the main tree 308, but in some embodiments, the main tree 308 may truncate the operand 201 at more than one stage. For example, in the illustrated embodiment, the soft logic circuit may truncate the LSB 204 from each intermediate result 203 in the main tree 308. Thus, in some embodiments, the summation of the truncated LSBs at each stage 208 may be separated into separate trailing trees 310 corresponding to the stage 208. For example, Figure 7 The embodiment in FIG. 3 shows a first trailing adder tree 310A that can handle the sum of the LSBs 302-304 from the first stage 208A of the main tree 308, and a second trailing adder tree 310B that can handle the sum of the LSBs 204 from the second stage 208B of the main tree 308. In such embodiments, the sums of the first trailing adder tree 310A and the second trailing adder tree 310B can be aligned. For example, the LSB 204 can share the same bit position as the bit 351 of the intermediate result 203 in the second stage of the first trailing adder tree 310A. Therefore, the final sum 206 of the first trailing adder tree 310A can be aligned with the final sum 206 of the second trailing adder tree 310B to sum up to the final sum 206' generated by summing the first trailing adder tree 310A and the second trailing adder tree 310B.

[0043] Accordingly, if Figure 8, the final sum 320 of the main tree 308 may be summed with the final sum 206′ of the summed first trailing adder tree 310A and the second trailing adder tree 310B. In some embodiments, the MSB of the final sum 206′ may be aligned with the LSB of the final sum 206. For example, in the illustrated embodiment, bits 323 and 322 of the final sum 206 may be aligned with bits 380 and 381 of the final sum 206′, respectively.

[0044] Although the addition tree 200 (e.g., the second addition tree 300) is divided into separate trees (e.g., 308, 310A, 310B) may appear to be less efficiently packed than a single addition tree 200, the truncated main tree 308 can be packed to 100% in the current FPGA 40. In addition, although the trailing adder tree 310 is not streamlined in some cases, there are many possibilities that it can be effectively packed into the FPGA. For example, because the trailing adder tree 310 can only use a small portion of the soft logic group, multiple nodes of the trailing adder tree 310 can be packed into a single logic group. In addition, because the trailing adder tree 310 can handle smaller arithmetic operations (e.g., fewer operands, smaller operands, and / or the like) when compared to the main tree 308, the trailing adder tree 310 can be efficiently packed into the integrated circuit 12.

[0045] In some embodiments, different arithmetic structures are additionally or alternatively used to construct the trailing adder tree 310. For example, the trailing adder tree 310 may include a compressor. Although the compressor may utilize more logic per bit, the packing ratio may be high because the operation may be constructed more like a random logic problem rather than the limitations of the carry chain mapping of the arithmetic structure.

[0046] In some embodiments, building multiple trees may be expensive in terms of resources and area. However, in many cases, the precision of the final sum 206 of the sum is more valuable than the precision of the intermediate results 203. Therefore, in addition to manipulating the contribution of the truncated bits by adding the first trailing adder tree 310A to the final sum 206 of the main tree 308 or in its alternative, the contribution of the truncated bits can be taken into account by constants added to the main tree 308. For example, after streamlining or truncating the operands of the main tree 308, a constant or set of constants can be added to the main tree 308 based on an estimate of the value originally contributed by the truncated bits to the final sum 206 to improve the precision of the final sum 206.

[0047] Accordingly, Fig. 9A third adder tree 400 is shown that may include an input circuit configured to receive seven 13-bit operands 201 and may output a single 10-bit final sum 206. In the first stage 208A of the third adder tree 400, each 13-bit operand 201 is truncated to a 10-bit operand 201. For example, Fig. 9 , three LSBs are truncated from the operand 201. As discussed, the operand 201 may be truncated for more efficient packing into the integrated circuit 12. The truncated operands 201 are then added together using the first adder 401A. Because an odd number of operands 201 are involved in the first stage 208A of the adder tree 200, and because the first adder 401A may receive two operands 201, the no-operation block 402 (no-op) may receive an operand 201 that does not fit in the set of the first adder 401A. In addition, after adding the 10-bit operand 201 to the first adder 401A, the intermediate result 203 may include 11 bits due to the 1-bit increase.

[0048] Therefore, in the second stage 208B of the third adder tree 400, the LSB of the 11-bit intermediate result 203 formed by the first adder 401A can be truncated to form three 10-bit intermediate results 203. The LSB of the operand 201 manipulated by the no-op block 402 can also be truncated to form a fourth 10-bit intermediate result 203, and the alignment between the 10-bit intermediate results 203 is maintained. For example, the operand 201 manipulated by the no-op block 402 can be zero-filled and / or sign-extended before the LSB is truncated to create a 10-bit intermediate result 203 that is properly aligned with the other 10-bit intermediate results 203. Therefore, in the second stage 208B, the set of second adders 401B can add up two sets of two operands from the four 10-bit intermediate results 203. Therefore, the second stage 208B can output two 11-bit intermediate results 203.

[0049] The third stage 208C of the third adder tree 400 may receive two 11-bit intermediate results 203 and may truncate the LSB of each 11-bit intermediate result to form two 10-bit intermediate results. The third adder 401C may then sum these 10-bit intermediate results to form the 11-bit final sum 206. Accordingly, the fourth stage 208D may truncate the LSB of the 11-bit final sum 206 to form the 10-bit final sum 206. However, as discussed, because bits are truncated from the operands 201 and the intermediate results 203 at each stage (208A-208D) of the third adder tree 400, the 10-bit final sum 206 may suffer from errors when compared to the actual final result generated by the adder tree 200 without any truncation. Therefore, in some embodiments, a set of constants (eg, AF) may be added to the adder tree at a particular stage 208 in order to reduce the average relative error caused by truncating the LSB at any stage 208 in the third adder tree 400 .

[0050] In view of the above, Fig.10 A flow chart of a method 500 for determining a suitable set of constants (e.g., AF) to be added to a stage 208 of an adder tree 200 is shown in accordance with embodiments described herein. Although the following description of the method 500 is described in a specific order (which represents a specific embodiment), it should be noted that the method 500 may be performed in any suitable order. In addition, certain steps may be skipped altogether, and additional steps may be included in the method 500. In addition, although the following description of the method 500 is described as being performed by a processor 101 (which may include one or more processing systems), it should be noted that the method 500 may be performed by any suitable computing device.

[0051] At block 502, the processor 101 may receive and / or determine a plurality of inputs (as represented in the illustrated embodiment by wIn, wOut, and N). Thus, the processor 101 may receive a width of input data (wIn) (e.g., the width of operands 201), a width of output data (wOut) (e.g., the width of final sum 206), and a number (N) of operands 201 to be summed together in the adder tree 200. For example, referring to Fig. 9 For the third adder tree 400 , the processor 101 may receive as inputs wIn=13 (eg, a 13-bit operand 201 ), wOut=10 (eg, a 10-bit final sum 206 ), and N=7 (eg, 7 input operands 201 ).

[0052] At block 504, the processor 101 may then set the number of level inputs (LI) or the number of operands to be aggregated at a particular stage 208 of the adder tree 200 to N, since the first stage 208A of the adder tree 200 may receive all operands 201. Accordingly, for Fig. 9 For the third adder tree 400 , the processor 101 may set LI=7.

[0053] At block 506, the processor 101 may determine a truncation value or an average error introduced by truncation at the first stage 208A of the adder tree 200 based on the distribution of operands 201 (e.g., inputs) of the adder tree 200. In some embodiments, the processor 101 may determine the truncation value at the first stage of the adder tree 200 based on the assumption that the values ​​of the operands 201 are uniformly distributed. For example, referring to Fig. 9 , the LSB group 305 truncated at the first stage 208A may range in value from 0 (e.g., 000) to 7 (e.g., 111). With a uniform distribution of operands 201, the truncation value introduced by removing the LSB group 305 from the operands 201 is 3.5. Therefore, the sum of the truncated values ​​for each of the seven 13-bit operands 201 truncated to 10-bit operands 201 is 24.5.

[0054] At block 508, the processor 101 may then update the overall average truncation value for the adder tree 200. Since the overall average truncation value may be initialized to 0, after block 506 is completed, the overall average truncation value may be updated to match the value calculated at block 506.

[0055] At block 510, the processor 101 may then determine whether LI is greater than or equal to 2. Thus, the processor may determine whether it is manipulating calculations for the final stage 208 of the adder tree 200 or a previous stage of the tree. If LI is greater than or equal to 2, then at block 512, the processor may update LI to LI=ceil(LI / 2). Thus, the processor may round the operation LI / 2 to its ceiling. For example, Fig. 9 The third adder tree 400 receives seven 13-bit operands 201 at its first stage 208A. Therefore, the starting value of LI=7, which is greater than 2. Therefore, at block 512, the processor 101 may update according to (7 / 2), resulting in LI=4, which corresponds to the number of 11-bit intermediate results 203 received by the second stage 208B of the adder tree. In the next iteration, the processor 101 may update the value of LI to 2, which corresponds to the number of 11-bit intermediate results 203 received by the third stage 208C, and in the last iteration or at the final stage 208D, the value of LI will be 1, which corresponds to the final sum 206.

[0056] At block 514, the processor 101 may calculate an average truncation value for the level input of the next stage 208 based on the average truncation value. To do so, the processor 101 may update the stage 208 of the adder tree 200 that is checked after updating the value of LI, and similar to block 508, may determine an average truncation value resulting from truncating the intermediate result 203 at the checked stage 208. However, while the processor 101 may determine the truncation value at block 506 using distribution-based data (e.g., an assumption that the operands are uniformly distributed), the processor may determine the truncation value based on the average at block 508. For example, the processor 101 may calculate the truncation value for the first stage 208A of the third adder tree 400 at block 506, and then may determine the truncation value for the second stage 208B when LI=4, the truncation value for the third stage 208C when LI=2, and the truncation value for the fourth stage 208D when LI=1 at block 512. Because the processor 101 can determine the truncation value based on the average truncation value of each intermediate result 203 of the adder tree 200, the maximum value 8 (e.g., 2 3 ) to truncate a single bit will have an average value of 4. Therefore, across the four 11-bit intermediate results 203, the processor 101 can determine an average truncation value of 16 for the second stage 208B. The maximum value of the truncated LSB is 8 because, although it is a single bit, the LSB truncated from the 11-bit intermediate result 203 in the second stage 208B is the fourth bit when compared to the original 13-bit operand 201. In addition, for the third stage 208C, the processor 101 can calculate the average truncation value of 8 and the total truncation value of 16 across the two 11-bit intermediate results 201 because the maximum value of the truncated LSB (which is in the fifth bit position relative to the 13-bit operand 201) is 16. For the fourth stage, the processor 101 can calculate the average truncation value and the total truncation value of 16 because there is a single LSB truncated from the sixth bit position relative to the 13-bit operand 201 from the 11-bit final sum 206.

[0057] Accordingly, when the method 500 loops back, the processor may add the most recently calculated truncation value to the overall average truncation value at block 508 to account for the truncation value at each stage 208 of the adder tree 200. Thus, for the third adder tree 400, the processor 101 may iteratively add 24.5 to the average truncation value of the first stage 208A, 16 to the average truncation value of the second stage 208B, 16 to the truncation value of the third stage 208C, and 16 to the truncation value of the fourth stage 208D each time the block 508 is reached to obtain an overall average truncation value of 72.5.

[0058] After the processor 101 determines that LI is less than 2 at block 510 , the processor 101 may round the average truncated value to the nearest value representable in the adder tree at block 516 , which may be determined based on the location of the constant (eg, AF). Fig. 9 , for example, because constants AC each represent a single bit carried into each of the first adder 401A, and because constant AC is carried into the fourth bit position relative to the 13-bit operand 201, they each represent 8 or 0. Similarly, constant DE (which is carried into the set of second adders 410B in the fifth bit position) can represent 16 or 0, and constant F (which is carried into the third adder 401C in the sixth bit position) can represent 32 or 0. Therefore, A*8+B*8+C*8+D*16+E*16+F*32 (where AF is an integer value) is a typical equation of values ​​that can be represented in the adder tree 200. Therefore, the overall average truncated value of 72.5 can be approximated to 72 in the adder tree 200 by setting A=1, F=2, and all other constants to 0, as well as other combinations of constants.

[0059] After determining a value for the average truncated value that is representable in the adder tree 200, e.g., after determining a suitable combination of constants that most closely matches the average truncated value, the processor 101 may return the constant to be added to the appropriate stage of the adder tree 200. Accordingly, when the constant is generated in the integrated circuit 12, it may be integrated into the design of the adder tree 200.

[0060] Although the reference Fig. 9 The method 500 is described using the third adder tree 400 of FIG. 5 , but it should be understood that the method is applicable to any suitable adder tree having any suitable values ​​for wIn, wOut, and N.

[0061] Furthermore, in the embodiments described above, at block 506, the processor 101 may determine the truncation value at the first stage of the adder tree 200 based on the assumption that the values ​​of the operands 201 are uniformly distributed. Additionally or alternatively, the processor 101 may utilize data related to the operands 201, and / or determine that the values ​​of the operands 201 are not uniformly distributed. For example, referring to Fig.11, given an 8-bit by 5-bit unsigned multiplier, the graph 600 may capture the distribution of the 3 LSBs in the resulting 13-bit product, which may be input to the adder tree 200 as the 13-bit operand 201. For example, the graph 600 may show the distribution of each of the possible values ​​(which may vary discretely from 0 to 7) of the 3 LSBs in the product that may be generated from each of the possible inputs to the multiplier. As shown, the value of the 3 LSBs is most likely to be 0 for any given input to the multiplier. Accordingly, the weighted average of the values ​​of the 3 LSBs may fall at 2.375, as represented by line 602, rather than at the unweighted average of 3.5. To this end, if at block 506 the processor 101 determines, for example, that operand 201 is generated from an 8-bit by 5-bit unsigned multiplier, the processor 101 may use the distribution information provided in the graph 600 to determine that the truncated value for each of the seven operands 201 is 2.375, and the total value for operand 201 is 16.625. In such embodiments, the overall average truncated value for adder tree 200 determined by the processor may have improved precision, which may result in a set of constants that are more accurately corrected for truncation errors.

[0062] Furthermore, in some embodiments, the set of constants may be dynamically updated based on the current distribution of the LSBs of the operands. Fig. 9 and Fig.10 As discussed, although the processor 101 may determine a set of constants that may reduce truncation errors in the adder tree 200 before construction of the adder tree 200, in some embodiments, the data processing system 100 and / or the processor 101 may periodically update the values ​​of the set of constants after construction of the adder tree 200. Accordingly, Fig.12 An embodiment of an adder tree system 620 that can modify the values ​​of a set of constants (eg, AF) added to the adder tree 200 is shown.

[0063] In such embodiments, registers and / or locations in memory 102 may be mapped to each of the set of constants (e.g., AF) such that during execution of adder tree 200 of an arithmetic operation, the value of each of the set of constants (e.g., AF) may be retrieved and fed into adder tree 200. In addition, as shown, a processing system (such as data processing system 100) may receive data related to LSB 622 from adder tree 200. The data related to LSB 622 may include information such as the value of the LSB, and the stage 208 of adder tree 200 from which the LSB was received. For example, data processing system 100 may receive data related to LSB 622 for the LSB of any stage 208 within adder tree 200. In addition, data processing system 100 may include LSB distribution logic 624. A suitable combination of components of data processing system 100 (e.g., processor 101 and memory 102) may implement and / or facilitate LSB distribution logic 624. In any case, LSB distribution logic 624 may receive data associated with LSB 622 and may determine and / or update the distribution of LSBs for any appropriate stage 208 (e.g., the stage 208 from which data associated with LSB 622 was received). Thus, data processing system 100 may maintain one or more sets of LSB distributions, such as Fig.11 As shown in .

[0064] The calculation logic 626 may utilize one or more sets of LSB distributions maintained by the LSB distribution logic 624 to determine an appropriate set of constants (e.g., AF) to be fed into the adder tree 200 to reduce truncation error. To determine the appropriate set of constants (e.g., AF), as discussed above, the calculation logic 626 may aggregate the truncation error values ​​for each stage 208 of the adder tree 200 and may determine the closest value of the total truncation error representable by the constants fed into the adder tree 200 based on the location (e.g., stage 208) into which the constants are fed (e.g., based on an equation such as A*8+B*8+C*8+D*16+E*16+F*32). After the calculation logic 626 determines the appropriate set of constants, the data processing system 100 may transfer the appropriate set of constants to its respective registers to update the current value associated with each constant stored in the respective registers. Accordingly, the updated values ​​for the set of constants may be fed into the adder tree 200 via the respective set of registers.

[0065] In some embodiments, in addition to or in the alternative to the following method 500 for calculating a set of constants suitable for counteracting truncation errors, a set of pre-computed fixed adjustment values ​​can be used without analytical weights or applications. In one case, for example, these values ​​can be fixed to twice the number of input values ​​in the tree. In another case, heuristics can be used to determine the most likely adjustment values.

[0066] In any case, the constant can be added to the adder tree 200, which can be implemented in a number of suitable ways. A simple approach can involve adding the constant directly to the output of the adder tree 200. For example, the adder tree 200 can compute and round its intermediate result 203 and final sum 206 without any modification to the tree, and after the final sum 206 is generated, the constant can be added to the final sum 206. Thus, however, this approach can increase the latency of a cycle and additional soft logic adder resources. However, if there are unpaired tuples in the adder tree 200, the constant can be added to the unpaired tuples with a lower latency than adding the constant to the final sum 206 of the adder tree 200. Additionally, in some embodiments, by using soft logic associated with the embedded carry lookahead adders of modern FPGAs 40 to add the constant, no additional latency or area is available for correcting truncation errors.

[0067] Accordingly, Fig.13 An embodiment of an adder tree node 207 mapped to a 2-input adder 650 is shown. The 2-input adder 650 can receive a first 4-bit operand A (e.g., A1 - A4) and a second 4-bit operand B (e.g., B1 - B4), and can output a 4-bit result S (e.g., S1 - S4). To generate the 4-bit result S, the 2-input adder 650 can include a carry lookahead adder 654 for each bit in the first 4-bit operand A and / or the second 4-bit operand B. For example, the first bit (A1) of the first 4-bit operand A and the first bit (B1) of the second 4-bit operand B can be mapped to a first carry lookahead adder 654A, the second bit (A2) of the first 4-bit operand A and the second bit (B2) of the second 4-bit operand B can be mapped to a second carry lookahead adder 654B, the third bit (A3) of the first 4-bit operand A and the third bit (B3) of the second 4-bit operand B can be mapped to a third carry lookahead adder 654C, and the fourth bit (A4) of the first 4-bit operand A and the fourth bit (B4) of the second 4-bit operand B can be mapped to a fourth carry lookahead adder 654D. Thus, each carry lookahead adder 654 can output a single bit (e.g., S1 - S4) to form the 4-bit result S, and can pass a carry bit (which can be added to the sum computed by the carry lookahead adder 654) to the next carry lookahead adder 654. In some embodiments, although there is no carry lookahead adder 654 preceding the operation of the first carry lookahead adder 654A, the first carry lookahead adder 654A can receive a carry bit that can affect the sum generated in the 4-bit result S.

[0068] Furthermore, each operand bit (e.g., A1-A4 and / or B1-B4) may not be directly mapped to a corresponding adder (e.g., 654A-654D); instead, a set of soft logic blocks 652 associated with the adder 654 may process the operand bits and may output a result to the adder 654. In the illustrated embodiment, for example, the soft logic block 652 receives a first bit (A1) of a first 4-bit operand A and outputs a result based on the first bit (A1) of the first 4-bit operand A to a first ripple-carry adder 654A, and the soft logic block 652 receives a first bit (B1) of a second 4-bit operand B and outputs a result based on the first bit (B1) of the second 4-bit operand B. In some embodiments, the operand bits (e.g., A1-A4 and / or B1-B4) may pass through the soft logic block 652 without any modification before reaching the adder 654. For example, the 2-input adder 650 may operate as if the soft logic block 652 is not included between the operand bits (e.g., A1-A4 and / or B1-B4) and the adder 654, or as if the operand bits (e.g., A1-A4 and / or B1-B4) are directly coupled to the adder 654. However, in some embodiments, the soft logic block may affect a 4-bit result S of the 2-input adder 650.

[0069] Go to Fig.14 , the soft logic block 652 may emulate the constants of the 3-2 compressor structure 700 (e.g., constant compression). In such embodiments, additional connections between the soft logic block 652 and the adder 654 may be used to implement a 3-input adder by first generating the 3-2 compression. Accordingly, two operands A and B may be routed to multiple soft logic blocks 652 simultaneously. Since the integrated circuit 12 may contain many redundant connections available, this routing may not significantly burden local routing. The soft logic block 652 may directly encode the third operand (which may be a constant). Therefore, a constant (e.g., a third operand) may be added at any location in the tree without additional logic, routing, or latency impact.

[0070] Go to Fig.13, in some embodiments, the soft logic block 652 may contain a LUT. Thus, the sum bits (e.g., S1-S4) and the carry bits generated by the corresponding ripple-carry adders 654 may be determined based on the mapping in the LUT of the soft logic block 652. Further, in some embodiments, the soft logic block 652 may determine the carry bit received by the first ripple-carry adder 654A based at least in part on the LUT. In such embodiments, the soft logic block 652 may account for rounding errors associated with truncating the LSB. For example, the soft logic block 652 may determine the carry bit that it would have contributed to the 4-bit result S if the LSB had not been truncated, and may add the carry bit to the first ripple-carry adder 654A to reduce the rounding error caused without the contribution of the carry bit. To do so, the truncated bit (A0) from the first operand A and the sign bit (SA) from the first operand A may be fed into the first soft logic block 652. The first soft logic block 652 may contain a LUT that maps inputs A0 and SA to outputs (A0 XOR SA) or an exclusive OR of A0 and SA. In addition, the truncation bit (B0) from the second operand B and the sign bit (B0) from the second operand B may be fed into the second soft logic block 652. The second soft logic block may contain a LUT that maps inputs B0 and SB to outputs (B0 XOR SB) or an exclusive OR of B0 and SB. The outputs of the first soft logic block 652 and the second soft logic block 652 may be fed into the ripple carry adder 654. In some embodiments, the ripple carry adder 654 may also receive SA as a carry bit. Therefore, the ripple carry adder 654 may generate a sum bit and a carry output bit based on the addition of (A0 XOR SA), (B0 XOR SB), and SA (e.g., (A0 XOR SA)+(B0 XOR SB)+SA). Accordingly, Table 1 demonstrates possible combinations of SA, SB, A0, and B0 and the sums and carries resulting from each combination.

[0071]

[0072] Table 1.(A0 XOR SA)+(B0 XOR SB)+SA

[0073] Furthermore, as discussed, by utilizing the LUT, the soft logic block 652 may determine the carry bit that the addition of A0 and B0 would have contributed to the sum of the first 4-bit operand A and the second 4-bit operand B if they were not truncated. However, for two's complement operands, both SA and SB may be added to (A0 XOR SA)+(B0 XOR SB) in order to generate the appropriate carry bit. Thus, since the carry bit generated according to Table 1 (e.g., according to the results of the ripple-carry adder 654) lacks the contribution of SB (because, for example, the ripple-carry adder 654 may only receive a single carry input bit), the carry bit may be fed to the first ripple-carry adder 654A of the 2-input adder 650, and SB may be added to any suitable portion of the 2-input adder 650 and / or a later portion of the adder tree 200 at the appropriate bit position.

[0074] Furthermore, in some embodiments, due to the location of the ripple-carry adder 654, the SA bit may not be able to be fed into the ripple-carry adder 654. For example, the ripple-carry adder 654 may not receive a carry input bit. In such embodiments, the output of the adder may represent the sum of (A0 XOR SA) + (B0 XOR SB) without the contribution of SA. Thus, Table 2 demonstrates the results of the above summation without the contribution of SA for each combination of SA, SB, A0, and B0.

[0075]

[0076] Table 2. (A0 XOR SA) + (B0 XOR SB)

[0077] The bits marked with an asterisk (*) in the "Carry" column may represent a carry bit error due to the absence of an SA carry into the ripple carry adder 654. For example, compared to Table 1, the asterisked carries of Table 2 are incorrect for the same combination of SA, SB, A0, and B0 (e.g., {SA, SB, S0, A0} = {1, 0, 0, 0}, {1, 0, 1, 1}, {1, 1, 0, 1}, {1,1, 1, 0}). Therefore, to account for the missing SA carry input, the LUT and / or soft logic block 652 may contain additional logic to force the carry output value of 1 for the asterisked combination of SA, SB, A0, and B0.

[0078] In some embodiments, packing can be improved by removing the LSB from the ranks in the multiplier. The multiplier can calculate the product by generating a set of partial products (each at a different rank) before adding them together. For example, in the first rank of the multiplier, the partial products can be generated by multiplying the first bit of the multiplier in the multiplication operation with each bit of the multiplicand in the multiplication operation (e.g., logical AND). In the second rank, the partial products can be generated by multiplying the second bit of the multiplier with each bit of the multiplicand. Therefore, removing the LSB from the ranks in the multiplier can involve removing the LSB from the multiplier and multiplicand involved in the multiplication operation. Assuming that the multiplication operation is a signed magnitude operation (e.g., both the multiplier and the multiplicand are in a signed magnitude format), to take into account the error involved in truncating (e.g., removing) the LSB from the multiplier and the multiplicand, the carry that the multiplier and the multiplicand originally had at the rank where the LSB was removed can be calculated. To do so, the LSBs of the outputs (e.g., partial products) of the multipliers at the multiplier stage from which the LSB is removed and at the subsequent multiplier stage may be determined. For example, for a multiplier A having an LSB A0 truncated at a first multiplier stage and a multiplicand B having an SLB B0 truncated at the first multiplier stage, the first multiplier stage output LSBs may be determined (e.g., A0 AND B0), and the second multiplier stage output LSBs involving a second multiplier bit A1 and a second multiplicand bit B1 (e.g., A1 AND B1) may be determined. The contribution of LSBs A0 and B0 may then be determined by taking the exclusive OR (XOR) of the first multiplier stage output LSB and the sign bit (Sign1) of the first multiplier stage output (e.g., (A0 AND B0) XORSign1), by taking the XOR of the second multiplier stage output LSB and the sign bit (Sign2) of the second multiplier stage output (e.g., (A1 AND B1) XOR Sign2), and by adding the results of these two operations (e.g., ((A0 AND B0) XORSign1) + ((A1 AND B1) XOR Sign2)). To do this, the logic for the operations discussed above may be included in at least a portion of one or more ALMs. In some embodiments, the logic for calculating ((A1 AND B1) XOR Sign2) may be included in the MSB half or portion of the ALM because the A1 and B1 bits are in more significant bit positions than the A0 and B0 bits. The LSB half of the ALM may include logic to compute ((A0 AND B0) XOR Sign1) + (A0 ANDB0) XOR Sign1)), the result of which may force a carry into the MSB so that the contribution of A0 and B0 is applied to A1 and B1, even though A0 and B0 are truncated.In such embodiments, the multipliers may be packed more efficiently into the integrated circuit 12, and the final product computed by the multipliers may remain unchanged because the contribution of the truncated bits (eg, A0 and B0) is accounted for.

[0079] In addition, in some embodiments, packing can be improved by removing the MSB from the result of the first stage 208A of the adder tree 200. More specifically, in embodiments involving the addition of two or more signed magnitude operands 201 in the first stage 208A, the sign bit (e.g., MSB) of one or more results of the addition can be truncated. In such embodiments, a suitable sign bit can be included later, such as added to the final sum 206 of the adder tree 200. For example, when the first operand 201A with seven bits of signed magnitude is added to the second operand 201B with six bits of signed magnitude, the result of the addition includes seven data bits and an eighth bit for manipulating sign extension. Therefore, the MSB of the result is the sign bit. However, because when the sign of the first operand 201A and the sign of the second operand 201B are both 0 (e.g., the first operand 201A and the second operand 201B are positive), the sign bit of the result will also be 0. When the sign of the first operand 201A and the sign of the second operand 201B are both 1 (e.g., the first operand 201A and the second operand 201B are negative), the sign of the result will also be 1, and when the sign of the first operand 201A does not match the sign of the second operand 201B, the sign bit of the result will match the sign of the first operand 201A (e.g., the MSB of the first operand 201A). Accordingly, the MSB of the result can be encoded based on the sign bits of the first operand 201A and the second operand 201B. Therefore, the MSB of the result can be truncated and later decoded to be added back to the intermediate result 203 or the final sum 206 of the adder tree 200. To this end, because the MSB is not included in the result, the resources involved in maintaining the MSB at each stage 208 of the adder tree are removed (e.g., separated) from the adder tree, which can result in more efficient packing of the adder tree.

[0080] Now go to Fig.15, the block floating point tree 800 illustrates a method of increasing the precision of the final sum 206 using soft logic in a reduced adder tree 200 (e.g., an adder tree 200 with truncated operands 201). The block floating point tree 800 may include a set of XOR gates 802. Each XOR gate 802 may be configured to receive two MSBs 805 of an operand 201 input to the block floating point tree 800. At each XOR gate 802, the two MSBs 805 of each operand 201 are checked to predict a possible overflow in the next stage 208 of the block floating point tree 800. Thus, the XOR gates 802 are used to check the dynamic range of the operands 201 in the first stage 208A. The OR gate 804 then logically ORs the results from the XOR gates 802 together to generate a single deterministic factor for the entire first stage 208A. The deterministic factor may be fed as a selection signal to a set of multiplexers 806 (mux) for selecting between the operand 201 or the operand 201 shifted right by one bit (e.g., with the LSB 202 truncated therein), the operand 201 shifted right by one bit being fed to the mux 806 via the shift block 808. For example, when an XOR gate 802 coupled to any one of the operands 201 in the first stage 208A predicts a possible overflow due to the corresponding operand 201, an OR gate 804 may select the operand 201 shifted right by one bit from each of the muxes 806 to account for bit growth. On the other hand, when no XOR gate 802 predicts a possible overflow due to any one of the operands 201, an OR gate 804 may select the unchanged operand 201 from each of the muxes 806. In any case, the outputs of the muxes 806 may be added together in two sets of two to produce the intermediate result 203.

[0081] Similar to the first stage 208A of the block floating point tree 800, in the second stage 208B of the block floating point tree 800, the two MSBs 805 of the intermediate result 203 may be fed into a second set of XOR gates 802B. The second set of XOR gates 802B may predict a possible overflow of the intermediate result 203, and may be coupled to a second OR gate 804B, such that the second OR gate 804B may select between the intermediate result 203 and the intermediate result 203 shifted right by one bit via the shift block 808B based on the overflow prediction for the intermediate result 203. The second OR gate 804B may select the intermediate result 203 when there is no carry for any prediction of the intermediate result 203, and may select the intermediate result 203 shifted right by one bit when at least one intermediate result 203 is predicted to overflow.

[0082] In the third stage 208C of the block floating point tree 800, the outputs selected from the second set of muxes 806B are added together to generate a final sum 206 having a suitable number of bits. In addition, the outputs of the OR gate 804A and the second OR gate 804B are summed by an adder structure 810 (e.g., adder tree 200) to create a common block floating point exponent (e.g., normalization factor) for the block floating point tree 800; however, the adder structure 810 may sum the outputs of the OR gates (e.g., 804A-804B) at any suitable location in the block floating point tree 800 and is not limited to summing the outputs in the third stage 208C. In some embodiments, the adder structure 810 may manipulate any suitable number of inputs and / or sum the inputs in any suitable number of stages 208.

[0083] In some cases, the output of the adder structure 810 may be a block floating point exponent (e.g., a normalization factor), and the final sum 206 may represent a block floating point mantissa. In some embodiments, the block floating point representation may use an integer mantissa. In addition, the mantissa may or may not be normalized and may be represented in a signed magnitude or other format.

[0084] Although the method involved in addition using the block floating point tree 800 may provide increased accuracy compared to some of the other embodiments of the adder tree 200 described herein, large trees (e.g., dot products with many elements) may be difficult to implement directly using this method. For example, because there is a large fan-in from the second set of XOR gates 802A and XOR gates 802B to the results in the OR gates 804A and the second OR gates 804B, respectively, and because there is a large fan-out from the OR gates 804A and the second OR gates 804B to the select signals of the muxes 806A and 806B, respectively, the block floating point tree 800 may suffer from long delays, which may degrade performance.

[0085] Therefore, to address the complexity involved in these methods, in some embodiments, the level 208 of the block floating point tree 800 may be omitted. Fig.16A simplified block floating point tree 850 is shown, in which the operands 201 of the first stage 208A are not adjusted (e.g., selected) before being aggregated in the second stage 208B. Therefore, the result of the first OR gate 802A is directly passed to the adder structure 810. In addition, because the stage 208 is omitted, to properly adjust the intermediate result 203 in the second stage 208B, the output of the OR gate 804 can be selected between the intermediate result 203 and the intermediate result shifted right by two bits at the mux 806. The set of shift blocks 808 can shift the intermediate result 203 right by two bits to take into account the overflow of the first stage 208A and the second stage 208B at the same time. Therefore, to select the output of the mux 806, the OR gate 804 can receive input from the XOR gate 802. In such embodiments, each XOR gate 802 can receive the first three MSBs 805 from the corresponding operand 201 to enable the OR gate 804 to predict the overflow in the first stage 208A and / or the second stage 208B.

[0086] In some embodiments, the OR gate 804 may output a selection signal including a single bit or a set of bits to select the intermediate result 203 (e.g., shifted by 0 bits) or the intermediate result shifted right by 2 bits. For example, in some embodiments, the selection signal may match the value of the selected shift operation (e.g., 0 or 2), which may include the use of two bits, and in other embodiments, the selection signal may use a single bit, which is used to encode the shift of 0 or 2. In any case, the adder structure 810 can still add all the shifted values ​​together and adjust its output based on the type of selection signal used (e.g., a single bit or a set of bits). In some embodiments, the encoded selection signal can be taken into account at different locations in the integrated circuit 12.

[0087] In addition to or in alternative to omitting stage 208, such as Fig.16 As shown, reference Fig.15 The complexity of the described method can also be reduced by using a set of block floating point trees 800 and / or a set of simplified block floating point trees 850. For example, the set of operands 201 can be divided into subsets of operands 201, and each tree in the set of block floating point trees 800 and / or the set of simplified block floating point trees 850 can receive and aggregate different subsets of operands 201. Therefore, the number of operands 201 received by the block floating point trees 800 and / or the simplified block floating point trees 850 can be reduced, thereby reducing the fan-in and fan-out that can cause delays in the large block floating point trees 800. The results of each set in the set of block floating point trees 800 and / or the set of simplified block floating point trees 850 can then be combined together using a single block floating point representation.

[0088] Accordingly, Fig.17An embodiment of a block floating point combination tree 900 that can implement the summation of multiple block floating point trees 800 is shown. In the illustrated embodiment, the first block floating point tree 800A, the second block floating point tree 800B, and the third block floating point tree 800C each output both a block floating point exponent and a block floating point mantissa. The block floating point exponents from each block floating point tree (e.g., 800A-C) can be sorted by a circuit 902, which can select and output the maximum block floating point exponent received from the block floating point tree 800A-C. A set of subtractors 904 can be coupled to the circuit 902, and each block floating point exponent can then be subtracted from this maximum block floating point exponent. The output from the subtractor 904 can then be fed into a set of shifters 906A-906C to normalize the block floating point mantissa. Thus, each of the shifters 906A-906C may right shift the corresponding block floating point mantissa of the corresponding block floating point tree (e.g., 800A-800C) by a number of bits corresponding to the corresponding output of the subtractor 904. In some embodiments, for example, the first shifter 906A may not shift the block floating point mantissa of the first block floating point tree 800A at all, because the block floating point exponent of the first block floating point tree 800A may be the maximum block floating point exponent. In such cases, for example, if the result of the subtractor 904 of subtracting the block floating point exponent of the second block floating point tree 800B from the maximum block floating point exponent is three, the second shifter 906B may right shift the second block floating point mantissa of the second block floating point tree 800B by three bits. In addition, the third shifter 906C may right shift the third block floating point mantissa of the third block floating point tree 800C by two bits, for example, based on the output of the corresponding subtractor 904 input to the third shifter 906C. Shifters 906A-906C may be relatively small and / or shallow logic structures because the block floating point exponent is likely to be small. Fig.15 As described in , the exponent may typically be incremented by a maximum value of 1 at each level 208 in the block floating point tree 800, so the shift operation used to normalize the block floating point mantissa may also be small.

[0089] Accordingly, each block floating point mantissa of the block floating point mantissa output by the shifters 906A-906C may be normalized relative to the maximum block floating point exponent, and thus the block floating point mantissa may be fed into the final block floating point tree 800D to be summed. The final block floating point tree 800D may receive the block floating point mantissa as an operand, and may generate a final block floating point mantissa and a final block floating point exponent based on the summation of each tree of the block floating point trees 800A-800C.

[0090] Furthermore, in some embodiments, the block floating point combination tree 900 may include any suitable combination of the block floating point tree 800 and / or the simplified block floating point tree 850. That is, it should be understood that Fig.17 and the description thereof are intended to be illustrative only and not restrictive.

[0091] The present invention also discloses a set of technical solutions as follows:

[0092] 1. An integrated circuit having an adder tree, the adder tree configured to generate a sum based at least in part on output and appended values, the adder tree comprising:

[0093] a first input circuit configured to receive a first operand, wherein the first operand comprises a first plurality of bits;

[0094] a second input circuit configured to receive a second operand, wherein the second operand comprises a second plurality of bits;

[0095] a soft logic circuit configured to separate one or more bits from the first plurality of bits to generate a first subset operand and configured to separate additional one or more bits from the second plurality of bits to generate a second subset operand;

[0096] an adder circuit configured to generate the output based at least in part on the first subset operands and the second subset operands; and

[0097] Additional circuitry is configured to generate the additional value based at least in part on the one or more bits.

[0098] 2. An integrated circuit as described in technical solution 1, wherein the additional circuit includes a trailing adder tree, wherein the trailing adder tree includes additional adder circuits configured to generate the additional value based at least in part on the sum of the one or more bits and the additional one or more bits.

[0099] 3. An integrated circuit as described in technical solution 2, wherein the trailing adder tree is configured to separate bits from the one or more bits to generate a subset of the one or more bits, and wherein the additional adder circuit is configured to generate the additional value based at least in part on the sum of the subset of the one or more bits and the additional one or more bits.

[0100] 4. The integrated circuit of claim 1 , wherein the additional circuit is configured to generate the additional value based at least in part on a distribution of possible values ​​of the one or more bits.

[0101] 5. An integrated circuit as described in technical solution 1, wherein the soft logic circuit includes the additional circuit and is configured to generate the additional value by simulating constant compression of the additional value.

[0102] 6. An integrated circuit as described in technical solution 1, wherein the soft logic circuit includes a lookup table configured to generate the additional value based in part on the one or more bits.

[0103] 7. An integrated circuit as described in technical solution 1, wherein the one or more bits include one or more least significant bits.

[0104] 8. An integrated circuit as described in technical solution 1, wherein the adder tree is configured to append the additional value to the output, prepend the additional value to the output, or a combination thereof.

[0105] 9. An integrated circuit as described in technical solution 1, wherein the adder tree is configured to generate the sum based at least in part on the sum of the appended value and the output.

[0106] 10. An integrated circuit as described in technical solution 1, wherein the adder tree is configured to perform a multiplication operation of a multiplicand and a multiplier, wherein the first operand includes the multiplicand, wherein the second multiplier includes the multiplier, wherein the adder circuit is configured to generate the output based at least in part on a first partial product of the first subset operand and the second subset operand and a second partial product of the first subset operand and the second subset operand, and wherein the additional circuit is configured to generate the additional value based at least in part on a first least significant bit of a third partial product of the multiplicand and the multiplier and a second least significant bit of a fourth partial product of the multiplicand and the multiplier.

[0107] 11. An integrated circuit as described in technical solution 10, wherein the sum includes the product of the multiplication operation, and wherein the additional value includes a carry input value to the output.

[0108] 12. An integrated circuit as described in technical solution 1, wherein the one or more bits include one or more most significant bits.

[0109] 13. An integrated circuit as described in technical solution 12, wherein the one or more most significant bits include one or more sign bits, and wherein the additional value includes a sign bit decoded at least in part based on the one or more bits and the additional one or more bits.

[0110] 14. An integrated circuit having an adder tree stage, the adder tree stage comprising:

[0111] an adder circuit comprising an input and configured to generate an output and a normalization factor based at least in part on the input; and

[0112] Input circuit, configured as:

[0113] receiving an operand, wherein the operand comprises a plurality of bits;

[0114] determining a most significant bit of the plurality of bits; and

[0115] Based at least in part on the most significant bit, the operand or a subset operand is selectively routed to the input, wherein the subset operand includes a subset of the plurality of bits.

[0116] 15. An integrated circuit as described in technical solution 14, wherein when the operand is selectively routed to the input, the input circuit is configured to increment the normalization factor.

[0117] 16. An integrated circuit as described in technical solution 14, wherein the input circuit is configured to selectively route the operand or the subset operand to an additional input of an additional adder circuit, wherein the additional adder circuit is arranged within an additional adder tree level.

[0118] 17. An integrated circuit as described in technical solution 16, wherein the number of bits in the subset of the plurality of bits configures at least part of the input circuit to route the subset operand to the adder circuit or the additional adder circuit.

[0119] 18. A tangible, non-transitory machine-readable medium comprising machine-readable instructions that, when executed by one or more processors, cause the processors to:

[0120] determining the number of operands to be input to the adder tree;

[0121] determining the bit width of each of said operands;

[0122] determining a second bit width of an output of the adder tree;

[0123] determining a number of removable bits to separate from each of said operands based at least in part on said second bit width;

[0124] determining a value based at least in part on the value of the appended removable bits; and

[0125] Construct an adder tree, configured as:

[0126] receiving the operand as input;

[0127] separating said removable bits from each of said operands to generate a plurality of subset operands; and

[0128] An output is generated based at least in part on the sum of the subset operands and the value.

[0129] 19. A computer-readable medium as described in technical solution 18, wherein the machine-readable instructions, when executed by the one or more processors, cause the processor to determine the value based at least in part on the distribution of possible additional values ​​of the removable bits.

[0130] 20. A computer-readable medium as described in technical solution 18, wherein the value includes a carry input value for the sum.

[0131] Although the embodiments set forth in the present disclosure allow for various modifications and alternative forms, specific embodiments have been illustrated in the drawings by way of example and have been described in detail herein. However, it should be understood that the present disclosure is not intended to be limited to the specific forms disclosed. The present disclosure is intended to cover all modifications, equivalents and alternatives that fall within the spirit and scope of the present disclosure as defined by the following appended claims.

[0132] Embodiments of the present application

[0133] The following numbered items define embodiments of the current application.

[0134] Item A1. An integrated circuit having an adder tree configured to generate a sum based at least in part on output and appended values, the adder tree comprising:

[0135] a first input circuit configured to receive a first operand, wherein the first operand comprises a first plurality of bits;

[0136] a second input circuit configured to receive a second operand, wherein the second operand comprises a second plurality of bits;

[0137] a soft logic circuit configured to separate one or more bits from the first plurality of bits to generate a first subset operand and configured to separate additional one or more bits from the second plurality of bits to generate a second subset operand;

[0138] an adder circuit configured to generate the output based at least in part on the first subset operands and the second subset operands; and

[0139] Additional circuitry is configured to generate the additional value based at least in part on the one or more bits.

[0140] Item A2. The integrated circuit of item A1, wherein the additional circuitry comprises a trailing adder tree, wherein the trailing adder tree comprises additional adder circuitry configured to generate the additional value based at least in part on a summation of the one or more bits and the additional one or more bits.

[0141] Item A3. The integrated circuit of Item A2, wherein the trailing adder tree is configured to separate bits from the one or more bits to generate a subset of the one or more bits, and wherein the additional adder circuit is configured to generate the additional value based at least in part on a summation of the subset of the one or more bits and the additional one or more bits.

[0142] Item A4. The integrated circuit of any of items A1 or 2, wherein the additional circuitry is configured to generate the additional value based at least in part on a distribution of possible values ​​of the one or more bits.

[0143] Item A5. The integrated circuit of any of Items A1, 2, or 4, wherein the soft logic circuit includes the additional circuit and is configured to generate the additional value by simulating constant compression of the additional value.

[0144] Item A6. The integrated circuit of any of items A1, 2, 4, or 5, wherein the soft logic circuit comprises a lookup table configured to generate the additional value based in part on the one or more bits.

[0145] Item A7. The integrated circuit of any of items A1, 2, 4, 5, or 6, wherein the one or more bits include one or more least significant bits.

[0146] Item A8. The integrated circuit of any of items A1, 2, 4, 5, 6, or 7, wherein the adder tree is configured to append the appended value to the output, prepend the appended value to the output, or a combination thereof.

[0147] Item A9. The integrated circuit of any of items A1, 2, 4, 5, 6, 7, or 8, wherein the adder tree is configured to generate the sum based at least in part on a summation of the appended value and the output.

[0148] Item A10. The integrated circuit of any of items A1, 2, 4, 5, 6, 7, 8, or 9, wherein the adder tree is configured to perform a multiplication operation of a multiplicand and a multiplier, wherein the first operand comprises the multiplicand, wherein the second multiplier comprises the multiplier, wherein the adder circuit is configured to generate the output based at least in part on a first partial product of the first subset operand and the second subset operand and a second partial product of the first subset operand and the second subset operand, and wherein the additional circuit is configured to generate the additional value based at least in part on a first least significant bit of a third partial product of the multiplicand and the multiplier and a second least significant bit of a fourth partial product of the multiplicand and the multiplier.

[0149] Item A11. The integrated circuit of item A10, wherein the sum comprises a product of the multiplication operation, and wherein the appended value comprises a carry-in value to the output.

[0150] Item A12. The integrated circuit of any of items A1, 2, 4, 5, 6, 7, 8, 9, or 10, wherein the one or more bits include one or more most significant bits.

[0151] Item A13. The integrated circuit of item A12, wherein the one or more most significant bits include one or more sign bits, and wherein the additional value includes a sign bit decoded based at least in part on the one or more bits and the additional one or more bits.

[0152] Item A14. An integrated circuit having an adder tree stage, the adder tree stage comprising:

[0153] an adder circuit comprising an input and configured to generate an output and a normalization factor based at least in part on the input; and

[0154] Input circuit, configured as:

[0155] receiving an operand, wherein the operand comprises a plurality of bits;

[0156] determining a most significant bit of the plurality of bits; and

[0157] Based at least in part on the most significant bit, the operand or a subset operand is selectively routed to the input, wherein the subset operand includes a subset of the plurality of bits.

[0158] Item A15. The integrated circuit of item A14, wherein when the operand is selectively routed to the input, the input circuit is configured to increment the normalization factor.

[0159] Item A16. The integrated circuit of item A14 or 15, wherein the input circuit is configured to selectively route the operand or the subset operand to an additional input of an additional adder circuit, wherein the additional adder circuit is disposed within an additional adder tree level.

[0160] Item A17. The integrated circuit of item A16, wherein the number of bits in the subset of the plurality of bits determines at least in part whether the input circuit is configured to route the subset operand to the adder circuit or the additional adder circuit.

[0161] Item A18. A tangible, non-transitory machine-readable medium comprising machine-readable instructions that, when executed by one or more processors, cause the processors to:

[0162] determining the number of operands to be input to the adder tree;

[0163] determining a bit width of each of said operands;

[0164] determining a second bit width of an output of the adder tree;

[0165] determining a number of removable bits to separate from each of said operands based at least in part on said second bit width;

[0166] determining a value based at least in part on the value of the appended removable bits; and

[0167] Construct an adder tree, configured as:

[0168] receiving the operand as input;

[0169] separating said removable bits from each of said operands to generate a plurality of subset operands; and

[0170] An output is generated based at least in part on the sum of the subset operands and the value.

[0171] Item A19. The machine-readable medium of Item A18, wherein the machine-readable instructions, when executed by the one or more processors, cause the processors to determine the value based at least in part on a distribution of possible additional values ​​of the removable bits.

[0172] Item A20. The machine-readable medium of any of the preceding items, wherein the value comprises a carry input value to the sum.

[0173] Item B1. An integrated circuit having an adder tree configured to generate a sum based at least in part on an output and an appended value, the adder tree comprising:

[0174] a first input circuit configured to receive a first operand, wherein the first operand comprises a first plurality of bits;

[0175] a second input circuit configured to receive a second operand, wherein the second operand comprises a second plurality of bits;

[0176] a soft logic circuit configured to separate one or more bits from the first plurality of bits to generate a first subset operand and configured to separate additional one or more bits from the second plurality of bits to generate a second subset operand;

[0177] an adder circuit configured to generate the output based at least in part on the first subset operands and the second subset operands; and

[0178] Additional circuitry is configured to generate the additional value based at least in part on the one or more bits.

[0179] Item B2. The integrated circuit of item B1, wherein the additional circuitry comprises a trailing adder tree, wherein the trailing adder tree comprises additional adder circuitry configured to generate the additional value based at least in part on a summation of the one or more bits and the additional one or more bits.

[0180] Item B3. The integrated circuit of item B2, wherein the trailing adder tree is configured to separate bits from the one or more bits to generate a subset of the one or more bits, and wherein the additional adder circuit is configured to generate the additional value based at least in part on a summation of the subset of the one or more bits and the additional one or more bits.

[0181] Item B4. The integrated circuit of any of items Bl or 2, wherein the additional circuitry is configured to generate the additional value based at least in part on a distribution of possible values ​​of the one or more bits.

[0182] Item B5. The integrated circuit of any of Items B1, 2, or 4, wherein the soft logic circuit includes the additional circuit and is configured to generate the additional value by simulating constant compression of the additional value.

[0183] Item B6. The integrated circuit of any of Items Bl, 2, 4, or 5, wherein the soft logic circuit comprises a lookup table configured to generate the additional value based in part on the one or more bits.

[0184] Item B7. The integrated circuit of any of items Bl, 2, 4, 5, or 6, wherein the one or more bits include one or more least significant bits.

[0185] Item B8. The integrated circuit of any of items Bl, 2, 4, 5, 6, or 7, wherein the adder tree is configured to append the appended value to the output, prepend the appended value to the output, or a combination thereof.

[0186] Item B9. The integrated circuit of any of items Bl, 2, 4, 5, 6, 7, or 8, wherein the adder tree is configured to generate the sum based at least in part on a summation of the appended value and the output.

[0187] Item B10. The integrated circuit of any of items B1, 2, 4, 5, 6, 7, 8, or 9, wherein the adder tree is configured to perform a multiplication operation of a multiplicand and a multiplier, wherein the first operand comprises the multiplicand, wherein the second multiplier comprises the multiplier, wherein the adder circuit is configured to generate the output based at least in part on a first partial product of the first subset operand and the second subset operand and a second partial product of the first subset operand and the second subset operand, and wherein the additional circuit is configured to generate the additional value based at least in part on a first least significant bit of a third partial product of the multiplicand and the multiplier and a second least significant bit of a fourth partial product of the multiplicand and the multiplier.

[0188] Item B11. The integrated circuit of item B10, wherein the sum comprises a product of the multiplication operation, and wherein the appended value comprises a carry-in value to the output.

[0189] Item B12. An integrated circuit as described in any of items Bl, 2, 4, 5, 6, 7, 8, 9, or 10, wherein the one or more bits include one or more most significant bits.

[0190] Item B13. The integrated circuit of item B12, wherein the one or more most significant bits include one or more sign bits, and wherein the additional value includes a sign bit decoded based at least in part on the one or more bits and the additional one or more bits.

[0191] Item B14. An integrated circuit having an adder tree stage, the adder tree stage comprising:

[0192] an adder circuit comprising an input and configured to generate an output and a normalization factor based at least in part on the input; and

[0193] Input circuit, configured as:

[0194] receiving an operand, wherein the operand comprises a plurality of bits;

[0195] determining a most significant bit of the plurality of bits; and

[0196] Based at least in part on the most significant bit, the operand or a subset operand is selectively routed to the input, wherein the subset operand includes a subset of the plurality of bits.

[0197] Item B15. The integrated circuit of item B14, wherein when the operand is selectively routed to the input, the input circuit is configured to increment the normalization factor.

[0198] Item B16. The integrated circuit of item B14 or 15, wherein the input circuit is configured to selectively route the operand or the subset operand to an additional input of an additional adder circuit, wherein the additional adder circuit is disposed within an additional adder tree level.

[0199] Item B17. The integrated circuit of item B16, wherein the number of bits in the subset of the plurality of bits determines at least in part whether the input circuit is configured to route the subset operand to the adder circuit or the additional adder circuit.

[0200] Item B18. A method of constructing an adder tree, comprising:

[0201] determining the number of operands to be input to the adder tree;

[0202] determining the bit width of each of said operands;

[0203] determining a second bit width of an output of the adder tree;

[0204] determining a number of removable bits to separate from each of said operands based at least in part on said second bit width;

[0205] determining a value based at least in part on the value of the appended removable bits; and

[0206] The adder tree is implemented by a circuit configured to:

[0207] receiving the operand as input;

[0208] separating said removable bits from each of said operands to generate a plurality of subset operands; and

[0209] An output is generated based at least in part on the sum of the subset operands and the value.

[0210] Item B19. The method of item B18, wherein determining the value comprises determining the value based at least in part on a distribution of possible additional values ​​of the removable bits.

[0211] Item B20: The method of Item B18 or 19, wherein the value includes a carry input value to the sum.

[0212] Item B21. A tangible, non-transitory machine-readable medium comprising machine-readable instructions that, when executed by one or more processors, cause the processors to perform the method of Item B18, 19, or 20.

[0213] Item C1. An integrated circuit having an adder tree configured to generate a sum based at least in part on output and appended values, the adder tree comprising:

[0214] a first input circuit configured to receive a first operand, wherein the first operand comprises a first plurality of bits;

[0215] a second input circuit configured to receive a second operand, wherein the second operand comprises a second plurality of bits;

[0216] a soft logic circuit configured to separate one or more bits from the first plurality of bits to generate a first subset operand and configured to separate additional one or more bits from the second plurality of bits to generate a second subset operand;

[0217] an adder circuit configured to generate the output based at least in part on the first subset operands and the second subset operands; and

[0218] Additional circuitry is configured to generate the additional value based at least in part on the one or more bits.

[0219] Item C2. The integrated circuit of item C1, wherein the additional circuitry comprises a trailing adder tree, wherein the trailing adder tree comprises additional adder circuitry configured to generate the additional value based at least in part on a summation of the one or more bits and the additional one or more bits.

[0220] Item C3. The integrated circuit of any of items C1 or 2, wherein the additional circuit is configured to generate the additional value based at least in part on a distribution of possible values ​​of the one or more bits.

[0221] Item C4. The integrated circuit of any of Items C1, 2, or 3, wherein the soft logic circuit includes the additional circuit and is configured to generate the additional value by simulating constant compression of the additional value.

[0222] Item C5. The integrated circuit of any of items A1, 2, 3, or 4, wherein the soft logic circuit comprises a lookup table configured to generate the additional value based in part on the one or more bits.

[0223] Item C6. The integrated circuit of any of items C1, 2, 3, 4, and 5, wherein the adder tree is configured to append the appended value to the output, prepend the appended value to the output, or a combination thereof.

[0224] Item C7. The integrated circuit of any of items C1, 2, 3, 4, 5, or 6, wherein the adder tree is configured to generate the sum based at least in part on a summation of the appended value and the output.

[0225] Item C8. The integrated circuit of any of items C1, 2, 3, 4, 5, 6, or 7, wherein the adder tree is configured to perform a multiplication operation of a multiplicand and a multiplier, wherein the first operand comprises the multiplicand, wherein the second multiplier comprises the multiplier, wherein the adder circuit is configured to generate the output based at least in part on a first partial product of the first subset operand and the second subset operand and a second partial product of the first subset operand and the second subset operand, and wherein the additional circuit is configured to generate the additional value based at least in part on a first least significant bit of a third partial product of the multiplicand and the multiplier and a second least significant bit of a fourth partial product of the multiplicand and the multiplier.

[0226] Item C9. The integrated circuit of item C8, wherein the sum comprises a product of the multiplication operation, and wherein the appended value comprises a carry-in value to the output.

[0227] Item C10. The integrated circuit of any of items C1, 2, 3, 4, 5, 6, 7, or 8, wherein the one or more bits include one or more most significant bits, one or more least significant bits, or a combination thereof.

[0228] Item C11. The integrated circuit of item C10, wherein the one or more most significant bits include one or more sign bits, and wherein the additional value includes a sign bit decoded based at least in part on the one or more bits and the additional one or more bits.

[0229] Item C12. An integrated circuit having an adder tree stage, the adder tree stage comprising:

[0230] an adder circuit comprising an input and configured to generate an output and a normalization factor based at least in part on the input; and

[0231] Input circuit, configured as:

[0232] receiving an operand, wherein the operand comprises a plurality of bits;

[0233] determining a most significant bit of the plurality of bits; and

[0234] Based at least in part on the most significant bit, the operand or a subset operand is selectively routed to the input, wherein the subset operand includes a subset of the plurality of bits.

[0235] Item C13. The integrated circuit of item C12, wherein when the operand is selectively routed to the input, the input circuit is configured to increment the normalization factor.

[0236] Item C14. A tangible, non-transitory machine-readable medium comprising machine-readable instructions that, when executed by one or more processors, cause the processors to:

[0237] determining the number of operands to be input to the adder tree;

[0238] determining a bit width of each of said operands;

[0239] determining a second bit width of an output of the adder tree;

[0240] determining a number of removable bits to separate from each of said operands based at least in part on said second bit width;

[0241] determining a value based at least in part on the value of the appended removable bits; and

[0242] Construct an adder tree, configured as:

[0243] receiving the operand as input;

[0244] separating said removable bits from each of said operands to generate a plurality of subset operands; and

[0245] An output is generated based at least in part on the sum of the subset operands and the value.

[0246] Item C15. The machine-readable medium of item C14, wherein the value comprises a carry input value to the sum.

Claims

1. An integrated circuit having an adder tree, the adder tree configured to generate a sum at least in part based on an output and an additional value, the adder tree comprising: a first input circuit configured to receive a first operand, wherein the first operand includes a first plurality of bits; a second input circuit configured to receive a second operand, wherein the second operand includes a second plurality of bits; soft logic circuitry configured to separate one or more bits from the first plurality of bits to generate a first subset of operands, and configured to separate additional one or more bits from the second plurality of bits to generate a second subset of operands; an adder circuit configured to generate the output at least in part based on the first subset of operands and the second subset of operands; and additional circuitry configured to generate the additional value at least in part based on the one or more bits.

2. The integrated circuit of claim 1, wherein the additional circuitry includes a trailing adder tree, wherein the trailing adder tree includes an additional adder circuit configured to generate the additional value at least in part based on a sum of the one or more bits and the additional one or more bits.

3. The integrated circuit of claim 2, wherein the trailing adder tree is configured to separate bits from the one or more bits to generate a subset of the one or more bits, and wherein the additional adder circuit is configured to generate the additional value at least in part based on a sum of the subset of the one or more bits and the additional one or more bits.

4. The integrated circuit of any one of claims 1 or 2, wherein the additional circuitry is configured to generate the additional value at least in part based on a distribution of possible values of the one or more bits.

5. The integrated circuit of any one of claims 1 or 2, wherein the soft logic circuitry includes the additional circuitry and is configured to generate the additional value by simulating constant compression of the additional value.

6. The integrated circuit of any one of claims 1 or 2, wherein the soft logic circuitry includes a look-up table configured to generate the additional value in part based on the one or more bits.

7. The integrated circuit of any one of claims 1 or 2, wherein the one or more bits include one or more least significant bits.

8. The integrated circuit of any one of claims 1 or 2, wherein the adder tree is configured to append the additional value to the output, prepend the additional value to the output, or a combination thereof.

9. The integrated circuit of any one of claims 1 or 2, wherein the adder tree is configured to generate the sum at least in part based on a sum of the additional value and the output.

10. The integrated circuit according to any one of claims 1 or 2, wherein the adder tree is configured to perform a multiplication operation of a multiplicand and a multiplier, wherein the first operand includes the multiplicand, wherein the second operand includes the multiplier, wherein the adder circuit is configured to generate the output based at least in part on a first partial product of the first subset of operands and the second subset of operands and a second partial product of the first subset of operands and the second subset of operands, and wherein the additional circuit is configured to generate the additional value based at least in part on a first least significant bit of a third partial product of the multiplicand and the multiplier and a second least significant bit of a fourth partial product of the multiplicand and the multiplier.

11. The integrated circuit according to claim 10, wherein the sum includes the product of the multiplication operation, and wherein the additional value includes a carry input value to the output.

12. The integrated circuit according to any one of claims 1 or 2, wherein the one or more bits include one or more most significant bits.

13. The integrated circuit according to claim 12, wherein the one or more most significant bits include one or more sign bits, and wherein the additional value includes a sign bit decoded based at least in part on the one or more bits and the additional one or more bits.

Citation Information

Patent Citations

  • Multipurpose multiply-add functional unit

    CN101133389A

  • Floating point multiplier and adder unit with data forwarding structure

    CN101221490A