Dot product-based processing element
By roughening the dot product processing into large macros on the FPGA, and using the DSP column composed of adder and multiplier, the problems of high resource requirements and increased routing length are solved, and more efficient dot product calculation is achieved.
Patent Information
- Application Number
- CN202210320762.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2016-09-15
- Filing Date
- 2017-09-11
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2037-09-11
AI Technical Summary
When dot product processing is implemented on reconfigurable devices such as field programmable gate arrays (FPGAs), the prior art has problems such as high resource requirements and increased routing lengths, resulting in performance degradation.
By roughening the dot product processing into large macros, using DSP columns composed of adders and multipliers, the routing between DSP blocks is reduced, and fewer general routing resources are adopted to form continuous blocks to simplify placement and routing problems.
It improves the performance and area utilization efficiency of integrated circuits, reduces the routing length, and improves the speed and efficiency of dot product calculation.
Smart Images

Figure CN114943057B_ABST
Abstract
Description
[0001] This application is a divisional application, and its original application is the international patent application PCT / US2017 / 050989 which entered the Chinese national stage on February 14, 2019, with an international filing date of September 11, 2017. The Chinese national application number of this original application is 201780049809.5, and the invention title is "Dot Product Processing Element". Technical Field
[0002] The present disclosure generally relates to integrated circuits, such as field programmable gate arrays (FPGAs). More specifically, the present disclosure relates to dot product processing implemented on an integrated circuit. Background Art
[0003] This section is intended to introduce to the reader various aspects of the art that may be relevant to various aspects of the present disclosure, which are described and / or claimed hereinafter. Such a discussion is considered to be helpful in providing background information to the reader to facilitate a better understanding of various aspects of the present disclosure. Therefore, it should be understood that these written descriptions are to be read in this sense and not as an admission of prior art.
[0004] Vector dot product processing is often used in digital signal processing algorithms (such as audio / video codecs, video or audio processing, etc.). When implementing a digital signal processor (DSP) on an integrated circuit device including a reconfigurable device, such as a field programmable gate array (FPGA), the physical area and speed of the dot product processing structure are factors that ensure that the integrated circuit device is suitable for the task to be performed in terms of both size and speed. However, dot product calculations can use individual DSP and memory resources for each function, which increases the routing length, and thus may also increase the area and performance. Summary of the Invention
[0005] The following presents a summary of certain embodiments disclosed herein. It should be understood that these aspects are provided merely to give the reader a concise summary of these particular embodiments and are not intended to limit the scope of the present disclosure. Indeed, the present disclosure may cover aspects that may not be set forth below.
[0006] This embodiment relates to systems, methods, and devices for enhancing the performance of dot product processing using a reconfigurable device, such as a field programmable gate array (FPGA). Specifically, a macro of a roughened dot product processing unit can be used to efficiently utilize the space in the reconfigurable device while ensuring satisfactory performance. In addition, by organizing the reconfigurable device into units that perform dot product processing without using more general routing paths available in the integrated circuit, where different digital signal processing blocks are used independently, it is possible that numerous long paths may have a negative impact on the performance of the integrated circuit.
[0007] Various improvements to the above features may exist with respect to various aspects of the present invention. Other features may also be added to these various aspects. These improvements and other features may exist alone or in any combination. For example, the various features discussed below in connection with one or more of the illustrated embodiments may be incorporated into any one of the above aspects of the present disclosure alone or in any combination. Similarly, the above brief summary is only intended to acquaint the reader with specific aspects and contexts of the embodiments of the present invention and does not limit the subject matter claimed. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Aspects of the present disclosure may be better understood when the following detailed description is read in conjunction with the accompanying drawings, in which:
[0009] Figure 1 is a block diagram of a system that utilizes dot product processing in accordance with an embodiment;
[0010] Figure 2 is a block diagram of a programmable logic device that may include logic useful for implementing dot product processing in accordance with an embodiment;
[0011] Figure 3 is a block diagram showing a dot product processing circuit in accordance with an embodiment;
[0012] Figure 4 is a block diagram showing a dot product processing circuit implemented using individual digital signal processing units in accordance with an embodiment Figure 3 of;
[0013] Figure 5 is a block diagram showing a dot product processing circuit implemented using a coarsened dot product processing unit in accordance with an embodiment Figure 3 of;
[0014] Figure 6 is a block diagram showing a dot product processing circuit having an accumulator configured to generate a running sum of a dot product in accordance with an embodiment Figure 5 of;
[0015] Figure 7 is a block diagram showing a dot product processing circuit having a memory cache in accordance with an embodiment Figure 5 of;
[0016] Figure 8 is a process for calculating a dot product using a coarsened dot product processing unit in accordance with an embodiment. DETAILED DESCRIPTION
[0017] One or more specific embodiments will be described below. To provide a brief description of these embodiments, not all features of an actual implementation are described in this specification. It should be recognized that in the development of any such actual implementation, as in any engineering or design project, numerous implementation-specific decisions must be made to achieve the developer's specific goals, such as meeting system-related and transaction-related constraints that may vary with the implementation. In addition, it should be understood that such development work can be complex and time-consuming, but it will still be routine work for those of ordinary skill in the art who benefit from this disclosure in terms of design, fabrication, and manufacture.
[0018] This disclosure describes techniques for using enhanced dot product processing elements (PEs) on integrated circuits including reconfigurable devices. The resulting PE architectures are well-suited for associative computations such as matrix multiplication or convolution and can be chained together to implement systolic arrays. The techniques are also based on coarsening dot products to larger sizes for more efficient resource utilization. The techniques also assist in the overall computer-aided design (CAD) process because coarsened dot products can be placed as large macros, requiring fewer atoms to be placed and forming fewer routes. The coarsening techniques support data interleaving techniques to enable heavy pipelining and caching for data reuse. Coarsening also improves the mapping of dot products to reduce the total number of digital signal processing (DSP) units used. Additionally, the dot product coarsening size can be adjusted based on the matrix size to be implemented to obtain efficient results at least in part based on the matrix size. For example, for each element in a vector of a matrix being dot processed, a DSP unit can be included. Specifically, in some embodiments, a dot product processing macro can be used to implement a four-vector dot product, which includes four DSP units in a column within a single macro to perform dot product processing on four vectors.
[0019] Specifically, the dot product PEs utilize adder trees and multiplication inputs to achieve high-efficiency dot product calculations. Coarsening the dot product reduces resource requirements because dedicated high-speed routing between adjacent DSPs can be utilized. Coarsening the dot product can also simplify placement and routing issues because the dot product can be laid out as large contiguous blocks, thereby reducing the number of placement objects and general routes. Although the techniques of this disclosure are briefly described in the context of reconfigurable devices such as programmable logic devices having a field-programmable gate array (FPGA) fabric, this is intended to be illustrative and not limiting. In fact, the dot product circuits of this disclosure can be implemented in other integrated circuits. Other types of integrated circuits such as application-specific integrated circuits (ASICs), microprocessors, storage devices, transceivers, etc. can also use the dot product circuits of this disclosure.
[0020] In view of the above, Figure 1A block diagram of system 10 is shown, which includes a dot product processing operation that can reduce the logic area typically used for such implementations and / or increase the speed. As described above, a designer may wish to implement functions on an integrated circuit, such as a reconfigurable integrated circuit 12, such as a field programmable gate array (FPGA). The designer can use design software 14, such as a version of Quartus from Altera TM to implement the circuit design that will be programmed onto IC 12. The design software 14 can use a compiler 16 to generate a low-level circuit design kernel program 18 for programming the integrated circuit 12, sometimes referred to as a program object file or bitstream. That is, the compiler 16 can provide the IC 12 with machine-readable instructions representing the circuit design. For example, the IC 12 can receive one or more kernel programs 18 that describe the hardware implementation that should be stored in the IC. In some embodiments, a dot product processing operation 20 can be implemented on the integrated circuit 12. For example, the dot product processing operation 20 can operate on single-precision floating-point numbers, double-precision floating-point numbers, or other suitable objects for dot product processing. As will be described in more detail below, the dot product processing operation 20 can be coarsened into large macros, and the large macros can be inserted into the design using a lower number of general routing resources (or non-DSP block resources). For example, two DSP blocks for single-precision floating-point addition can be combined as described below to form a double-precision floating-point adder.
[0021] Now turning to a more detailed discussion of IC 12, Figure 2 An IC device 12 is shown, which can be a programmable logic device, such as a field programmable gate array (FPGA) 40. For the purposes of this example, the device 40 is referred to as an FPGA, but it should be understood that the device can be any type of reconfigurable device (e.g., an application specific integrated circuit and / or an application specific standard product). As shown, the FPGA 40 can have input / output circuitry 42 for driving signals out of the device 40 and receiving signals from other devices via input / output pins 44. Interconnect resources 46, such as global and local vertical and horizontal conductive lines and buses, can be used to route signals on the device 40. In addition, the interconnect resources 46 can include fixed interconnects (conductive lines) and programmable interconnects (i.e., programmable connections between corresponding fixed interconnects). The programmable logic 48 can include combinational and sequential logic circuits. For example, the programmable logic 48 can include look-up tables, registers, and multiplexers. In various embodiments, the programmable logic 48 can be configured to perform custom logic functions. The programmable interconnects associated with the interconnect resources can be considered part of the programmable logic 48. As described in more detail below, the FPGA 40 can include adjustable logic capable of partially reconfiguring the FPGA, such that kernels can be added, removed, and / or replaced during the runtime of the FPGA 40.
[0022] A programmable logic device such as FPGA 40 may include programmable elements 50 within programmable logic 48. For example, as described above, a designer (e.g., a customer) may program (e.g., configure) the programmable logic 48 to perform one or more desired functions. For example, some programmable logic devices may be programmed by using a mask programming arrangement to configure their programmable elements 50, which is performed during semiconductor manufacturing. After the semiconductor manufacturing operations have been completed, their programmable elements 50 are programmed, for example, using electrical programming or laser programming, so as to configure other programmable logic devices. Generally, the programmable elements 50 may be based on any suitable programmable technology, such as fuses, antifuses, electrically programmable read-only memory technology, random access memory cells, mask programming elements, etc.
[0023] Many programmable logic devices are programmed electrically. With an electrical programming arrangement, the programmable elements 50 may include one or more logic elements (wires, gates, registers, etc.). For example, during programming, configuration data is loaded into the memory 52 using the pins 44 and the input / output circuit 42. In some embodiments, the memory 52 may be implemented as a random access memory (RAM) cell. Describing the use of the memory 52 based on RAM technology here is only intended as an example. In addition, the memory 52 may be distributed throughout the device 40 (e.g., as RAM cells). Further, since configuration data is loaded into these RAM cells during programming, they are sometimes referred to as configuration RAM cells (CRAM). The memory 52 may provide corresponding static control output signals that control the states of associated logic components in the programmable logic 48. For example, in some embodiments, the output signals may be applied to the gates of metal oxide semiconductor (MOS) transistors within the programmable logic 48. In some embodiments, the programmable elements 50 may include DSP blocks that implement common operations, such as a dot product processing element implemented using a DSP block.
[0024] Any suitable architecture can be used to organize the circuitry of FPGA 40. For example, the logic of FPGA 40 can be organized in a series of rows and columns of a larger programmable logic region, each of which can contain multiple smaller logic regions. The logic resources of FPGA 40 can be interconnected by interconnect resources 46, such as associated vertical and horizontal conductors. For example, in some embodiments, these conductors can include global wires that span substantially all of FPGA 40, fractional lines that span portions of device 40, such as half-lines or quarter-lines, staggered lines of a particular length (e.g., long enough to interconnect several logic regions), smaller local wires, or any other suitable arrangement of interconnect resources. Additionally, in other embodiments, the logic of FPGA 40 can be arranged in more levels or layers, where multiple large regions are interconnected to form a larger portion of the logic. Furthermore, some device arrangements can use logic arranged in a manner other than rows and columns.
[0025] As described above, FPGA 40 can allow a designer to generate a custom design that can perform and customize functions. Each design can have its own hardware implementation realized on FPGA 40. These hardware implementations can include floating-point operations of DSP blocks using programmable elements 50.
[0026] The dot product can be algebraically defined as the sum of the products of the corresponding terms in the vectors for which the dot product is computed. For example, Equation 1 algebraically shows the dot product expression for two 4-vector dot products:
[0027]
[0028] where A1, A2, A3, and A4 are the elements in vector A, and B1, B2, B3, and B4 are the elements in vector B. For example, the elements in vector A can correspond to a time element and three spatial elements. The elements in vector B can be another vector, some scalar values equivalent to the elements of vector A, or some other operation, such as a Lorentz transformation.
[0029] Figure 3FIG. 0 shows a 4-vector dot product circuit 100. A first vector is submitted to the circuit 100 as its constituent elements 102, 104, 106, and 108, and a second vector is submitted as a group of constituent elements 110, 112, 114, and 116 for taking a dot product with the first vector. The second group of elements 110, 112, 114, and 116 can be different values or can be the same values. For example, the second group of elements 110, 112, 114, and 116 can be a second four-vector. Each pair of corresponding elements is submitted to a corresponding multiplier that multiplies the elements together. Specifically, element 102 and 110 are multiplied together in multiplier 118 to form a product; element 104 and 112 are multiplied together in multiplier 120 to form a product; element 106 and 114 are multiplied together in multiplier 122 to form a product; and element 108 and 116 are multiplied together in multiplier 124 to form a product. The products are then added together. In some embodiments, the products can be added together in a single 4-input adder. Additionally or alternatively, the products can be added together using sequential addition, which halves the number of products in each round of addition and may continue until a sum is found, the sum representing the cross product of the first vector and the second vector. For example, the product of element 102 and 110 and the product of element 104 and 112 are added together in adder 124; the product of element 106 and 114 and the product of element 108 and 116 are added together in adder 126. Then the sums from adder 124 and adder 126 are added together in another adder 128. In some embodiments, the circuit 100 includes a register 130 that is used to ensure that corresponding parts of the dot product calculation are substantially synchronized during processing. The register 130 also ensures that the data being calculated is transferred correctly. The output of adder 128 is the dot product 132.
[0030] In some embodiments, the dot product 132 can be submitted to an accumulator 133 to form a running sum. The accumulator 133 includes an adder 134 that adds the most recent dot product to the running sum 136. In other words, the accumulator 133 receives the dot product and adds it to all previous dot products, so that a running sum of the dot products can be calculated.
[0031] Figure 4 FIG. 7 shows a circuit 138 that can be implemented in a reconfigurable device using digital signal processing (DSP) blocks 140, 142, 144, and 146 Figure 3Dot product circuit 100. As shown, each of DSP blocks 140, 142, 144, and 146 includes an adder (e.g., adders 124, 126, 128, and 137) and a multiplier (e.g., multipliers 118, 120, 122, and 123), which are routed to provide 4-vector dot product processing. As shown, a portion (e.g., portion 148) of one or more DSP blocks (e.g., DSP block 144) remains unused. Additionally, the routing 150 between DSP blocks 140, 142, 144, and 146 may be long relative to the routing within DSP blocks 140, 142, 144, and 146. This may be due to the relatively dispersed locations of DSP blocks 140, 142, 144, and 146 within IC 12. Such routing may result in a decrease in the dot product calculation performance using circuit 138. To reduce such routing, the dot product processing can be coarsened into larger macros, also known as atoms, which can be designed as reconfigurable devices to reduce the routing between DSP blocks.
[0032] Figure 5 Circuit 158 is shown, which includes a coarsened 4-vector dot product processing macro that can be added to an integrated circuit, such as a reconfigurable device. Circuit 158 includes all the computations of circuit 100. However, in contrast to circuit 138, circuit 158 includes a DSP column 160, which includes DSP blocks 162, 164, 166, and 168. DSP blocks 162, 164, 166, and 168 are located in DSP column 160 to reduce the routing between DSP blocks using the general routing in the reconfigurable device, thereby increasing device performance and / or area consumption. Specifically, circuit 158 includes only a single general routing 170. Similar to DSP blocks 140, 142, 144, and 146, each of DSP blocks 162, 164, 166, and 168 includes an adder and a multiplier. In Figure 5 the specific embodiment of the 4-element vector dot product shown, a portion 172 of DSP block 162 is unused.
[0033] In some embodiments, this portion 172 can be used for computations external to the simple 4-vector dot product calculation. For example, Figure 6 circuit 173 is shown, and circuit 173 uses this portion 172 to implement an accumulator to find the running sum of the dot product. Specifically, adder 147 is used to add the previous dot product 132 to the product of elements 102 and 110 via accumulator routing 174. Then this sum of the previous dot product 132 is added to the product of elements 104 and 112. This additional addition in exchange for the running sum adds a greater degree of latency.
[0034] Figure 7A cache accumulator circuit 178 is shown, which includes an accumulator cache 180 that stores previous dot products. The accumulator cache 180 includes memory banks 182. The memory banks 182 include a plurality of memory blocks 184, 186, 188, and 190 collectively referred to as memory blocks 184-190. The memory blocks 184-190 can include any suitable memory blocks, for example, the company's M20K memory blocks. The memory blocks 184-190 can store a relatively large amount of sums. For example, the memory blocks 184-190 can store any exponent of two running sums, such as 128, 256, 512, 1024, 2048, etc. Each sum can increase the delay of the circuit 178 relative to Figure 5 the delay of the circuit 158 to exchange additional summations corresponding to previous dot products. For example, storing 2048 sums can increase the delay by 2048.
[0035] Figure 8 A process 200 for calculating a dot product using an integrated circuit is shown. The process 200 includes organizing the circuit into a dot product processing unit for dot product processing (block 202). The dot product processing unit includes two or more digital signal processing cores to reduce the amount of general routing used in dot product processing if individual digital signal processing blocks are used. The dot product processing unit is a coarsened dot product process that includes a number of digital signal processing units based on the size of the matrix (or individual vectors) for which the dot product is being processed. For example, it can include four digital signal processing units for four vectors (i.e., vectors having four constituent elements) and can include eight digital signal processing units for eight vectors. This dot product processing unit can be used to generate an integrated circuit (block 204), which can include, for example, programming a programmable logic device or an ASIC. As previously mentioned, coarsening the dot product process enables the use of less general routing than independent configured digital signal processing units used alone, because more routing and longer routing are used in independently configured digital signal processing units used alone. In other words, by creating a single chunk that can be sized according to the dot product matrix size, the performance of the integrated circuit device implementing the dot product processing can be improved.
[0036] Although the embodiments set forth in this disclosure may readily admit of various modifications and alternative forms, specific embodiments have been shown by way of example in the drawings and have been described in more detail herein. However, it should be understood that this disclosure is not intended to be limited to the particular forms disclosed. This disclosure is to cover all modifications, equivalent elements, and alternative elements falling within the spirit and scope of this disclosure as defined by the following appended claims.
Claims
1. An electronic device, comprising: A plurality of digital signal processing units, the plurality of digital signal processing units being configurable to be connected to perform various operations and being configured to work together to output a dot product, wherein at least one of the plurality of digital signal processing units includes: A plurality of input ports for receiving a plurality of inputs at a corresponding digital signal processing unit; A multiplier for generating a product based at least in part on the plurality of inputs; A dedicated routing for sending the product to or receiving an output from another digital signal processing unit of the plurality of digital signal processing units; and An adder for adding the product to an output of a multiplier of another digital signal processing unit of the plurality of digital signal processing units, wherein the dot product is based at least in part on the output of the adder.
2. The electronic device according to claim 1, wherein, The number of digital signal processing units for outputting the dot product corresponds to the size of the matrix to be processed.
3. The electronic device according to claim 2, wherein, The number of digital signal processing units is based at least in part on the number of elements to be processed in the dimensions of the matrix.
4. The electronic device according to claim 3, wherein, The number of elements includes the number of objects in a vector of the matrix.
5. The electronic device according to claim 1, wherein, At least one digital signal processing unit of the plurality of digital signal processing units includes: The dedicated routing for receiving the product from a first digital signal processing unit adjacent to the at least one digital signal processing unit; and An additional dedicated routing for sending an additional product to a second digital signal processing unit, wherein the second digital signal processing unit is adjacent to the at least one digital signal processing unit.
6. The electronic device according to claim 1, wherein, The plurality of digital signal processing units are arranged in columns of digital signal processing blocks.
7. The electronic device according to claim 5, wherein, The second digital signal processing unit includes: A second plurality of input ports for receiving a second plurality of inputs; A second multiplier for generating a second product based at least in part on the second plurality of inputs; A second dedicated routing for sending the second product to at least one digital signal processing unit of the plurality of digital signal processing units via the second dedicated routing.
8. The electronic device according to claim 7, wherein, The second digital signal processing unit includes an adder that does not increase the additional product before the additional product is output from the additional dedicated routing.
9. The electronic device according to claim 1, wherein, The second digital signal processing unit of the plurality of digital signal processing units includes: A second plurality of input ports for receiving a second plurality of inputs; A second multiplier for generating a second product based at least in part on the second plurality of inputs; A second dedicated routing for receiving the second product via the dedicated routing of one or more of the digital signal processing units; and A second adder for adding values based at least in part on the second product to form the dot product.
10. The electronic device according to claim 9, wherein, The second digital signal processing unit includes an output for outputting the dot product from the plurality of digital signal processing units.
11. An electronic device, comprising: A first digital signal processing unit, the first digital signal processing unit including: A first plurality of input ports that receive a plurality of inputs; and A first multiplier that generates a first product based at least in part on the first plurality of inputs; A second digital signal processing unit, the second digital signal processing unit comprising: A second plurality of input ports that receive a second plurality of inputs; A second multiplier that generates a second product based at least in part on the second plurality of inputs; and A first adder that adds the first product and the second product together to form an intermediate sum; A first hardwired routing for sending the first product from the first digital signal processing unit to the second digital signal processing unit; A third digital signal processing unit, the third digital signal processing unit comprising: A third plurality of input ports that receive a third plurality of inputs; A third multiplier that generates a third product based at least in part on the second plurality of inputs; Wherein the first digital signal processing unit, the second digital signal processing unit, and the third digital signal processing unit are configurable to be connected to perform a variety of operations including calculating a dot product; and A second adder that generates a dot product based at least in part on the intermediate sum and the third product; and A second hardwired routing that sends the second product from the second digital signal processing unit to the third digital signal processing unit.
12. The electronic device according to claim 11, wherein, The first digital signal processing unit, the second digital signal processing unit, and the third digital signal processing unit are arranged in a column of digital signal processing units.
13. The electronic device according to claim 11, wherein, The number of digital signal processing units for outputting the dot product corresponds to the size of the matrix to be processed.
14. The electronic device according to claim 13, wherein, The number of digital signal processing units is based on the amount of vectors to be calculated.
15. The electronic device according to claim 11, wherein, The first digital signal processing unit includes a third adder that does not add to the first product before passing the first product to the second digital signal processing unit.
16. The electronic device according to claim 15, wherein, The third adder receives the first product and adds the first product to zero.
17. The electronic device according to claim 11, wherein, The third digital signal processing unit includes an output configured to output the dot product as an output of the electronic device.
18. A method of generating a dot product calculation integrated circuit, comprising: Organizing a programmable circuit into a dot product processing configuration to process a dot product, wherein the dot product processing units in the dot product processing configuration include three or more digital signal processing blocks, and wherein each of the digital signal processing blocks includes a multiplier that generates a product based at least in part on corresponding inputs of corresponding digital signal processing units; and Generating the integrated circuit having the dot product processing units by positioning the dot product processing units in the integrated circuit, wherein generating the integrated circuit includes: Placing a first digital signal processing unit of the three or more digital signal processing units that receives a first plurality of inputs, multiplies the first plurality of inputs together as a first product, and outputs the first product; Place a second digital signal processing unit among the three or more digital signal processing units, the second digital signal processing unit receiving a second plurality of inputs, multiplying the second plurality of inputs together as a second product, summing the first product and the second product as a first sum using an adder of the second digital signal processing unit, and outputting the first sum, wherein the first digital signal processing unit transfers the first product to the second digital signal processing unit via a first dedicated route, the first dedicated route transferring the first product from the first digital signal processing unit to the second digital signal processing unit; and place a third digital signal processing unit among the three or more digital signal processing units, the third digital signal processing unit receiving a third plurality of inputs, multiplying the third plurality of inputs together as a third product, adding the first sum and a second value together to form the dot product, and outputting the dot product, wherein the second value is at least partially based on the third plurality of inputs, and the second digital signal processing unit transfers the second product to the third digital signal processing unit via a second dedicated route, the second dedicated route transferring the second product from the second digital signal processing unit to the third digital signal processing unit.
19. The method according to claim 18, wherein, Placing the first digital signal processing unit, the second digital signal processing unit, and the third digital signal processing unit includes placing the first digital signal processing unit, the second digital signal processing unit, and the third digital signal processing unit in a column, wherein the first digital signal processing unit and the second digital signal processing unit are adjacent to each other in the column, and the second digital signal processing unit and the third digital signal processing unit are adjacent to each other in the column.
20. The method according to claim 18, wherein The first product is added to a zero value in an adder of the first digital signal processing unit before being transferred to the second digital signal processing unit.
21. A method of configuring a series of digital signal processing units, comprising: configuring the digital signal processing units in the series of digital signal processing units to receive a plurality of inputs, wherein the digital signal processing units are configurable to perform multiple operations on the plurality of inputs; configuring at least one digital signal processing unit in the series of digital signal processing units to output results from an arithmetic circuit to a corresponding digital signal processing unit in the series of digital signal processing units using a corresponding dedicated route among a plurality of dedicated routes between the series of digital signal processing units; and configuring one of the digital signal processing units to output a dot product at least partially based on the results.
22. The method according to claim 21, wherein The results include a first product output from a multiplier, the arithmetic circuit includes the multiplier, and configuring the digital signal processing units includes configuring the multiplier in the series of digital signal processing units to multiply corresponding pluralities of inputs together.
23. The method according to claim 22, wherein, The result includes addition in an adder, the arithmetic circuit includes the adder, and configuring the digital signal processing unit includes configuring the adder in the addition digital signal processing unit of the series of digital signal processing units to receive the first product from the transmit digital signal processing unit of the series of digital signal processing units via a dedicated route among the plurality of dedicated routes and add the first product to a second product from a multiplier in the addition digital signal processing unit.
24. A computer-readable medium storing instructions that, when executed by a processor, are configured to cause the processor to configure a series of digital signal processing units, wherein, The instructions are configured to cause the processor to: Configure the digital signal processing units in the series of digital signal processing units to receive a plurality of inputs, wherein the digital signal processing units are configurable to perform a variety of operations on the plurality of inputs; Configure at least one digital signal processing unit in the series of digital signal processing units to output a result from the arithmetic circuit to a corresponding digital signal processing unit in the series of digital signal processing units using a corresponding dedicated route among the plurality of dedicated routes between the series of digital signal processing units; and Configure one of the digital signal processing units to output a dot product at least in part based on the result.
25. The computer-readable medium according to claim 24, wherein, The result includes a first product output from a multiplier, the arithmetic circuit includes the multiplier, and configuring the digital signal processing unit includes configuring the multiplier in the series of digital signal processing units to multiply corresponding pluralities of inputs together.
26. The computer-readable medium according to claim 25, wherein, The result includes addition in an adder, the arithmetic circuit includes the adder, and configuring the digital signal processing unit includes configuring the adder in the addition digital signal processing unit of the series of digital signal processing units to receive the first product from the transmit digital signal processing unit of the series of digital signal processing units via a dedicated route among the plurality of dedicated routes and add the first product to a second product from a multiplier in the addition digital signal processing unit.
27. A programmable logic device, comprising: A plurality of digital signal processing units, the plurality of digital signal processing units being configurable to be connected to perform a variety of operations and being configured to work together to output a dot product, wherein the plurality of digital signal processing units includes a first digital signal processing unit and a second digital signal processing unit, wherein the first digital signal processing unit includes: A first plurality of input ports for receiving a first plurality of inputs; A first multiplier for generating a first product at least in part based on the first plurality of inputs; and A first adder for generating a first sum by adding the first product to another value received from the second digital signal processing unit, wherein the another value is at least in part based on a second product generated by the second digital signal processing unit, wherein the dot product is at least in part based on the first sum; and The second digital signal processing unit includes: A second plurality of input ports for receiving a second plurality of inputs; A second multiplier for generating the second product at least in part based on the second plurality of inputs; A second adder, the second adder being configurable to generate a second sum, wherein the second sum is at least partially based on the second product; and Routing, the routing being configurable to output the second product to the first adder, wherein the routing bypasses the second adder.
28. The programmable logic device according to claim 27, wherein, The first digital signal processing unit and the second digital signal processing unit are adjacent to each other.
29. The programmable logic device according to claim 27, wherein, The plurality of digital signal processing units are arranged in columns of digital signal processing units.
30. The programmable logic device according to claim 27, wherein, The corresponding input in the first plurality of inputs received by the plurality of digital signal processing units is unique.
31. The programmable logic device according to claim 27, wherein, The first plurality of inputs and the second plurality of inputs include floating-point numbers.
32. The programmable logic device according to claim 31, wherein, The floating-point numbers include single-precision floating-point numbers, double-precision floating-point numbers, or both.
33. The programmable logic device according to claim 27, wherein, The plurality of digital signal processing units includes a third digital signal processing unit adjacent to at least one of the first digital signal processing unit and the second digital signal processing unit.
34. The programmable logic device according to claim 27, wherein, The plurality of digital signal processing units are communicatively coupled to a plurality of memory blocks; and The digital signal processing units among the plurality of digital signal processing units include adders configurable to receive values from a memory block among the plurality of memory blocks.
35. A method for generating a dot product, comprising:[[]]END] Receiving a first plurality of inputs via a first plurality of input ports of a first digital signal processing unit; Generating a first product at least partially based on the first plurality of inputs via a first multiplier of the first digital signal processing unit; Generating a first sum via a first adder of the first digital signal processing unit by adding the first product to another value, wherein the another value is at least partially based on a second product, wherein the dot product is at least partially based on the first sum; Receiving a second plurality of inputs via a second plurality of input ports of a second digital signal processing unit; Generating the second product at least partially based on the second plurality of inputs via a second multiplier of the second digital signal processing unit; and Generating a second sum via a second adder of the second digital signal processing unit, wherein the second sum is at least partially based on the second product, wherein the second digital signal processing unit includes a routing configurable to output the second product to the first adder, wherein the second adder is not disposed along the routing.
36. The method according to claim 35, wherein, The first digital signal processing unit and the second digital signal processing unit are adjacent to each other.
37. The method according to claim 35, wherein, The first plurality of inputs and the second plurality of inputs include floating-point numbers.
38. The method according to claim 37, wherein, The floating-point numbers include single-precision floating-point numbers.
39. The method according to claim 35, wherein: The first digital signal processing unit and the second digital signal processing unit are communicatively coupled to a plurality of memory blocks; and The digital signal processing units among the plurality of digital signal processing units include adders configurable to receive values from a memory block among the plurality of memory blocks.
40. The method according to claim 35, comprising:[[]]END] Receiving the dot product via an input port of a third digital signal processing unit; and Adding the dot product to a value stored in the third digital signal processing unit via a third adder of the third digital signal processing unit.
41. An integrated circuit device, comprising: A plurality of digital signal processing units, the plurality of digital signal processing units being configurable to be connected to perform various operations and being configured to work together to output a dot product, wherein the plurality of digital signal processing units includes a first digital signal processing unit, a second digital signal processing unit, and a third digital signal processing unit, wherein the first digital signal processing unit includes: A first plurality of input ports for receiving a first plurality of inputs; A first multiplier for generating a first product based at least in part on the first plurality of inputs; and A first adder for generating a first sum by adding the first product to another value received from the second digital signal processing unit, wherein the another value is at least in part based on a second product generated by the second digital signal processing unit, wherein the dot product is at least in part based on the first sum; and The second digital signal processing unit includes: A second plurality of input ports for receiving a second plurality of inputs; A second multiplier for generating the second product based at least in part on the second plurality of inputs; A second adder configurable to generate a second sum, wherein the second sum is at least in part based on the second product; and A routing configurable to output the second product to the first adder, wherein the routing bypasses the second adder; and The third digital signal processing unit includes: An input port configurable to receive the dot product; and A third adder configurable to generate a third sum by adding the dot product to a value stored in the third digital signal processing unit.
42. The integrated circuit device according to claim 41, wherein, The first digital signal processing unit and the second digital signal processing unit are adjacent to each other.
43. The integrated circuit device according to claim 41, wherein, The first digital signal processing unit, the second digital signal processing unit, and the third digital signal processing unit are communicatively coupled to a plurality of storage blocks; and The digital signal processing units in the plurality of digital signal processing units include adders configurable to receive values from storage blocks in the plurality of storage blocks.
44. The integrated circuit device according to claim 41, wherein, The first plurality of inputs, the second plurality of inputs, and the dot product include floating-point numbers.
45. The integrated circuit device according to claim 41, wherein, The integrated circuit device includes programmable logic.
46. A programmable logic device, comprising: Input / output circuitry for driving signals out of the programmable logic device and receiving signals from other devices via the input / output circuitry; Interconnect resources for routing signals; And Programmable logic including programmable elements and configured to perform one or more custom logic functions.
47. The programmable logic device according to claim 46, wherein, The interconnect resources include either fixed interconnects or programmable interconnects.
48. The programmable logic device according to claim 46, wherein, The programmable logic includes any one of look-up tables, registers, and multiplexers.
Citation Information
Patent Citations
Programmable device using fixed and configurable logic to implement floating-point rounding
US9348795B1