Dense digital arithmetic circuits for fixed-point machine learning utilize

By combining more than one quantity into operands in the neural network, the problem of throughput restriction caused by hardware limitation of integrated circuits is solved, and the effect of improving the execution speed and performance of neural networks is achieved.

CN113988285BActive Publication Date: 2025-05-30ALTERA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111331751.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-12-27
Filing Date
2018-03-21
Publication Date
2025-05-30
Estimated Expiration
2038-03-21

AI Technical Summary

Technical Problem

When performing mathematical operations, the throughput is limited due to the hardware limitation of the integrated circuit, which in turn affects the execution speed of the neural network.

Method used

The execution speed of MAC blocks in neural network applications is improved by combining more than one quantity into each operand of multiplication operation. The specific method includes packing two or more quantities in the operand and outputting multiple products in the multiplication operation to avoid overflow during the accumulation process.

Benefits of technology

Improves the execution speed of neural networks in integrated circuits, and improves overall performance by increasing throughput and reducing overflow risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113988285B_ABST
    Figure CN113988285B_ABST
Patent Text Reader

Abstract

Systems and methods relate to improving the throughput of neural networks in integrated circuits by combining values in operands to increase computational density. The system includes an integrated circuit (IC) having a multiplier circuit. The IC receives a first value and a second value in a first operand. The IC performs a multiplication operation on the first operand and a second operand via the multiplier circuit to produce a first multiplied product that is at least partially based on the first value and a second multiplied product that is at least partially based on the second value.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a patent application with the same name having an application number of 201810233635.4 and filed on March 21, 2018.

[0002] Cross - reference to related applications

[0003] This application is a non - provisional application claiming priority to U.S. Provisional Patent Application No. 62 / 488,636, filed on April 21, 2017, entitled "Lower Precision Neural Network Systems and Methods", the entire content of which is incorporated by reference for all purposes. Background of the disclosure

[0004] The present disclosure generally relates to the efficient utilization of arithmetic circuits (e.g., multiply - accumulate circuits and / or digital signal processor (DSP) circuits) in integrated circuits for machine learning.

[0005] This section is intended to introduce the reader to various aspects in the field that may be related to various aspects of the present disclosure described and / or claimed below. This discussion is considered to be helpful in providing background information to the reader to facilitate a better understanding of the various aspects of the present disclosure. Thus, it can be understood that these statements should be read in this context and should not be considered as prior art.

[0006] Integrated circuits such as field - programmable gate arrays (FPGAs) can include circuits for performing various mathematical operations. For example, deep - learning neural networks can be implemented in one or more integrated - circuit devices for machine - learning applications. The integrated - circuit devices can perform several operations to output results for the neural network. However, in some instances, the throughput of the mathematical operations in the neural network may be limited by the hardware of the integrated circuit. Due to these limitations, the neural network may execute at a slower rate than desired. Brief description of the drawings

[0007] Various aspects of the present disclosure can be better understood when reading the following detailed description and when referring to the accompanying drawings, in which:

[0008] Figure 1 is a block diagram of a data - processing system for performing machine learning via a machine - learning circuit according to an embodiment;

[0009] Figure 2 is according to an embodiment of Figure 1 the machine - learning circuit;

[0010] Figure 3 is according to an embodiment for via Figure 1The network diagram of a neural network for a machine learning circuit to perform a task;

[0011] Figure 4 is a flowchart of a process performed by a Figure 1 machine learning circuit according to an embodiment;

[0012] Figure 5 is a diagram of another neural network for a Figure 1 machine learning circuit to perform a task according to an embodiment;

[0013] Figure 6 is a block diagram of a data structure of a multiplication operation performed by a Figure 1 machine learning circuit according to an embodiment;

[0014] Figure 7 is a block diagram of a Figure 1 machine learning circuit for performing a multiplication operation according to an embodiment;

[0015] Figure 8 is a block diagram of a generalized data structure of a multiplication operation performed by a Figure 1 machine learning circuit according to an embodiment;

[0016] Figure 9 is a block diagram of a Figure 1 machine learning circuit for performing a generalized multiplication operation according to an embodiment;

[0017] Figure 10 is a block diagram of another data structure of a multiplication operation performed by a Figure 1 machine learning circuit according to an embodiment; and

[0018] Figure 11 is a flowchart of a process performed by a Figure 1 machine learning circuit according to an embodiment to perform a multiplication operation for a neural network. Detailed Description

[0019] One or more specific embodiments will be described below. To provide a concise description of these embodiments, not all features of the actual implementation are described in the specification. It will be appreciated that in the development of any such actual implementation (as in any engineering or design project), several implementation-specific decisions must be made to achieve the developer's specific goals, e.g., to comply with system-related constraints and business-related constraints, which may vary from one implementation to another. Additionally, it will be appreciated that such development efforts may be complex and time-consuming, but nonetheless will be routine tasks for those of ordinary skill in the art who benefit from the present disclosure in terms of design, fabrication, and manufacture.

[0020] Machine learning is used in various settings to perform tasks through the use of examples. For example, neural networks can be used to perform tasks without task-specific programming. That is, neural networks can be trained from previous data to classify or infer information from current data. For example, training data can be used to identify images containing an object by analyzing other images that include the object and images that do not include the object. While images are used as examples, this is merely illustrative, and any suitable neural network task can be performed in the embodiments described below.

[0021] Configurable devices such as programmable logic devices (PLDs) can perform one or more operations to perform tasks via machine learning. For example, an integrated circuit (IC) such as a field-programmable gate array (FPGA) can include one or more digital signal processing (DSP) blocks or DSP circuits that have one or more specialized processing blocks to perform arithmetic operations on data received by the DSP blocks. One type of specialized processing block in a DSP block can be a multiply-accumulate (MAC) block or MAC circuit that includes one or more multiplier circuits and / or one or more accumulator circuits. For example, in some FPGAs, the MAC block can be a hardened intellectual property (IP) block that has a specialized multiply circuit coupled to a specialized adder circuit. Examples of operations performed by the MAC block include dot products, vector multiplications, etc. As described below, one or more multipliers of the DSP block can be used to perform neural network arithmetic operations during a classification or inference phase. However, the throughput of a digital signal processor (DSP) can be limited by the hardware of the IC. For example, the number of MAC blocks can limit the performance (e.g., speed) of the IC when performing arithmetic operations of a neural network.

[0022] Some arithmetic operations in a neural network may not involve the same precision as that designed to be processed in a MAC block. For example, a MAC block can include circuitry for processing 18-bit operands, but a neural network may involve multiplying 6-bit operands of lower precision. The systems and methods described below improve the neural network performance in an IC by better utilizing the capacity of the operands in a multiplication operation. By combining more than one quantity into each operand of a multiplication operation, the speed of performing MAC operations (e.g., weighted sum) by the MAC blocks of the IC in a neural network application can be increased. For example, two or more quantities can be packed into a first operand received by a multiplier circuit. Two or more quantities can be packed into a second operand received by the multiplier circuit. Then, the multiplier circuit can perform a multiplication operation between the first operand and the second operand to determine the product between each of the corresponding quantities. Then, the multiplier circuit can output each of the products to be accumulated. To prevent multiplication overflow, there can be a gap between each of the quantities combined in the operands.

[0023] In addition, to reduce the likelihood of accumulation overflow, the accumulator circuit of the MAC block can be bypassed to a soft logic accumulator. That is, the multiplication of the MAC operation can be performed in a hardened multiplier dedicated to performing multiplication and accumulation, and the accumulation of the MAC operation can be performed in soft logic to prevent overflow due to the accumulation of several products output from the multiplication.

[0024] In view of the above, Figure 1 FIG. shows a block diagram of a data processing system 10 that can be used to perform one or more tasks via machine learning. The data processing system 10 may include a processor 10 operably coupled to a memory 14. The processor 10 may execute one or more instructions stored on the memory 14 to perform one or more tasks. The data processing system 10 may include a network interface 16 to send and / or receive data via a network to communicate with other electronic devices. The data processing system 10 may include one or more input / output (I / O) 18, which may be used to receive data via I / O devices such as a keyboard, mouse, display, buttons, or other controls. The data processing system 10 may include a machine learning circuit 20 that uses machine learning methods and techniques to perform one or more tasks. The machine learning circuit 20 may include a PLD, for example, an FPGA. Each of the processor 12, the memory 14, the network interface 16, the I / O 18, and the machine learning circuit 20 may be communicatively coupled to each other via an interconnect circuit 22 such as a communication bus.

[0025] The hardware of the machine learning circuit 20 can use neural networks 100 and 138 to perform one or more tasks. Now turning to a more detailed discussion of an example of the machine learning circuit 20, Figure 2Shown is an IC 30, which may be a programmable logic device, e.g., a field programmable gate array (FPGA) 32. For the purposes of this example, the device is referred to as IC 30, but it should be understood that the device may be any suitable type of device that can be used (e.g., an application specific standard product). As shown, the IC 30 may have input / output circuitry 34 for driving signals out of the IC 30 and receiving signals from other devices via input / output pins 36. Interconnect resources 38 (e.g., global and local, vertical and horizontal wires and buses) may be used to route signals on the IC 30. Additionally, the interconnect resources 38 may include fixed interconnects (wires) and programmable interconnects (i.e., programmable connections between corresponding fixed interconnects). The programmable logic 40 may include combinational and sequential logic circuits. For example, the programmable logic 40 may include look-up tables, registers, and multiplexers. In various embodiments, the programmable logic 40 may be configured to perform custom logic functions. The programmable interconnects associated with the interconnect resources may be considered part of the programmable logic 40. The IC 30 may include a programmable element 42 having the programmable logic 40. The programmable element 42 may be based on any suitable programmable technology, e.g., fuses, antifuses, electrically programmable read only memory technology, random access memory cells, mask-programmed elements, etc.

[0026] The circuitry of the IC 30 may be organized using any suitable architecture. As an example, the logic of the IC 30 may be organized in a series of rows and columns of larger programmable logic regions, each of the larger programmable logic regions may have multiple smaller logic regions. The logic resources of the IC 30 may be interconnected by the interconnect resources 38 such as associated vertical and horizontal conductors. For example, in some embodiments, these conductors may include global wires spanning substantially all of the IC 30, fractional lines such as half-lines or quarter-lines spanning portions of the IC 30, staggered lines of a particular length (e.g., sufficient to interconnect several logic regions), smaller local wires, or any other suitable arrangement of interconnect resources. Additionally, in further embodiments, the logic of the IC 30 may be arranged in more levels or layers, where multiple large regions are interconnected to form larger logic portions. Additionally, other device arrangements may use logic that is not arranged in a row and column fashion. As explained below, the machine learning circuitry 20 may use the hardware of the IC 30 to perform one or more tasks. For example, the machine learning circuitry 20 may utilize arithmetic logic circuitry to perform arithmetic operations used in machine learning methods and techniques.

[0027] Figure 3Is a network diagram of an example of a machine learning network (e.g., neural network 100) that can be used to perform one or more tasks on a machine learning circuit 20. Although the neural network 100 is described in detail as an example, any suitable machine learning methods and techniques can be used. The neural network 100 includes a set of inputs 102, 103, 104, and 106, a set of weights 108, 109, 110, and 112, a set of sums 114 and 115, and a result value 116. Each of the inputs 102, 103, 104, and 106 is weighted using the corresponding weight to determine the corresponding weighted values 108, 109, 110, and 112. The weighted values 108 and 109 can be summed at the sum 114, and the weighted values 110 and 112 can be summed at the sum 115. The result value 116 can be output from the sums 114 and 115 and used to perform one or more tasks from previous data. Although four inputs and two sums are shown, this is intended to be illustrative, and any suitable combination of inputs, weights, sums, and connections therebetween can be used.

[0028] Figure 4 Is a flowchart of a process 130 that can be performed on an IC 30 in conjunction with the neural network 100. At block 132, the IC 30 can perform training, where the weighted values 108, 109, 110, and 112 are determined and / or adjusted such that the weights applied to the inputs 102, 103, 104, and 106 indicate the likelihood that the corresponding inputs 102, 103, 104, and 106 predict the result value 116.

[0029] When training the neural network 100, at block 134, the IC 30 can perform inference and / or classification on new data. In an example involving, for example, image recognition, the neural network 100 can be trained using images of shapes (e.g., circles, triangles, squares), where the shapes in the images are known. Then, after adjusting the weights based on the training data, the IC 30 can use the neural network 100 to classify the shapes of new data. By adjusting the weights applied to the inputs 102, 103, 104, and 106 based on the training data, weights can be obtained that, when applied to a new image, reflect the likelihood that the corresponding inputs of the new image include a particular shape. In some embodiments, continuous learning may occur, where the new data is then verified and the weights are continuously adjusted. Each of the operations in block 134 can be performed via the machine learning circuit 20, and / or some of the operations in each of block 134 can be performed via the processor 12.

[0030] Figure 5It is a network diagram example of a neural network 138 having an input layer 140, more than one computing layer 142, and an output layer 144. The illustrated embodiment can be referred to as a deep neural network due to having more than one computing layer 142 (also referred to as a hidden layer). As the number of computing layers 142 increases, the complexity and processing of the input increase. In the illustrated embodiment, each input is weighted and summed at four sums, and then each corresponding sum is weighted and summed at three sums, and these three sums are then used to output the result value.

[0031] As explained below, the circuit of IC 30 can also include one or more DSP blocks. The DSP block can include one or more (multiply-accumulate) MAC blocks or MAC circuits. Each MAC block can include hardened circuits (e.g., a multiplier circuit and an accumulator circuit) that are designed and dedicated to performing multiplication and accumulation operations. Although the MAC block can include circuits that perform multiplication and accumulation on inputs with a specific amount of precision, the neural network 100 can have inputs 102, 103, 104, and 106 and weights with a lower precision than the circuits of the MAC block. For example, although the neural network 100 can utilize weights and inputs 102, 103, 104, and 106 with six-bit precision, the MAC block can include circuits designed to process 18-bit inputs. By combining more than one value from the neural network 100 into the same operand of the MAC block, each multiplication of the MAC block can process additional values associated with the neural network 100 to increase the throughput of the neural network 100.

[0032] Figure 6 It is an example of a set of data structures 200 of IC 30 that combines values in the same operand to allow IC 30 to process the values of the neural network 100 at a faster rate. IC 30 can combine a first value 204 and a second value 206 into a first operand 208. That is, IC 30 can pack each bit of the first value 204 and each bit of the second value 206 into the first operand 208. For example, the first component (e.g., the first set of bits) of the first operand 208 can represent the first value 204, and the second component (e.g., the second set of bits) of the first operand 208 can represent the second value 206. In addition, the operand 208 can include a gap between the first value 204 and the second value 206 to prevent overflow. For example, the gap 210 can be at least the number of bits of the first value 204 or the second value 206. The first value 204 can be the first input 102, and the second value 206 can be the second input 104.

[0033] Similarly, IC 30 can combine a third value 212 and a fourth value 214 into a second operand 216. The second operand 216 can include a gap 218 between the third value 212 and the fourth value 214 to prevent overflow. The gap 218 can be at least the number of bits of the third value 212 or the fourth value 214. In the example described above where the neural network 100 uses six-bit precision, the first value 204, the second value 206, the third value 212, the fourth value 214, and the gaps 210 and 218 can each be six bits. The third value 212 can be a first weight to be applied to the first input 102, and the fourth value 106 can be a second weight applied to the second input 104.

[0034] IC 30 can perform a multiplication operation on the first operand 208 and the second operand 216 such that the multiplied product 230 includes a first product 232 of the first value 204 multiplied by the third value 212 and a second product 234 of the second value 206 multiplied by the fourth value 214 from the same multiplication operation. That is, by combining or packing more than one value into each operand 208 and 216 with sufficient gaps 210 and 218 between the values, the multiplied product 230 can include each corresponding product without overflow. For example, in the neural network 100, the first product 232 can be a weighted value 108 from the first weight applied to the first input 102, and the second product 234 can be a second weighted value 110 from the second weight applied to the second input 104. By combining the values from the neural network 100 into each operand 208 and 216, the result value 116 can be determined at a faster rate due to increased throughput.

[0035] Each of the first product 232 and the second product 234 can then be split from the multiplied product 230 and accumulated. Since accumulation can be a faster operation than multiplication, the performance of the neural network 100 can be improved by determining more than one product from a single multiplication operation using more available precision in the hardened multiplier circuit of IC 30. Additionally, the hardened circuit of the MAC block in IC 30 can be dedicated to performing multiplication to determine the weighted values 108, 109, 110, and 112 at a faster rate due to the specialization of the hardened circuit compared to a circuit performing multiplication in soft logic.

[0036] Figure 7 is to perform with respect to Figure 6Block diagram of the circuit of the IC 30 for the described arithmetic operations. The IC 30 may include a DSP block 250 having a first input circuit 252 and a second input circuit 254 that respectively receive a first operand 208 and a second operand 216. The DSP block 250 may include a MAC block 260 having a multiplier circuit 262 that multiplies the first operand 208 by the second operand 216 and outputs the product. That is, the multiplier circuit 262 may be designed or hardened using a circuit that performs a multiplication operation on operands having a specific precision. By including more than one value having a precision lower than the precision of the designed operands in the operands before performing the multiplication operation, more than one product may be determined based on the multiplication operation.

[0037] In some embodiments, the MAC block 260 may include an adder circuit 264 that may add the products from the multiplier circuit 262. When the MAC operation is completed, the MAC block 260 may output the result via an output circuit 268. In the illustrated embodiment, the IC 30 may include more than one DSP block 250 (e.g., 2, 3, 4, 5, or more), and each DSP block may include more than one MAC block 252 (e.g., 2, 3, 4, 5, 10, 20, 50, or more).

[0038] The MAC block 252 may include a bypass circuit 270 (e.g., a multiplexer) to bypass the adder 264 and provide the multiplied product 230 to the soft logic 274 of the IC 30. Additionally, the IC 30 may then perform the sums 114 and 115 of the neural network in the soft logic 274 of the IC 30. The soft logic 274 may refer to programming instructions (e.g., code) stored in the memory on the IC 30 for performing the operations of the IC 30. The IC 30 may be programmed to execute instructions to split the first product 232 and the second product 234 from the multiplied product 230. The IC 30 may then execute instructions to accumulate the first product (e.g., the first weighted value 108) with one or more other products (e.g., weighted values 108) 276 to determine a total according to the sum 114. The IC 30 may execute instructions to accumulate the second product (e.g., weighted value 110) with one or more other products (e.g., weighted values 112) 278. For example, the first product (e.g., weighted value 108) may be retained at block 282. The IC 30 may then perform another multiplication to determine a third product and a fourth product (e.g., weighted values 109 and 112) by combining a fifth value and a sixth value (e.g., inputs 103 and 106) into a third operand and combining a seventh value and an eighth value (e.g., the weights of the corresponding inputs) into a fourth operand. The third product and the fourth product (e.g., weighted values 109 and 112) are then added to each respective total retained (282 and 284). By implementing an accumulator in the soft logic 274, more accumulation operations can be performed with a lower risk of overflow or no risk of overflow. By moving to a neural network with a lower precision than the six - bit example, the IC 30 can obtain additional products in each multiplication operation.

[0039] Figure 8 Is a more generalized example of the data structure 300 used when performing multiplication operations in the neural network 100 on the IC 30. The IC 30 may combine the values A[0] to A[n] into a first operand. Similarly, the IC 30 may combine the values B[0] to B[n] into a second operand. Each of the values 302 and 306 may be separated from other values 302 and 306 by gaps 304 and 308 to prevent overflow.

[0040] When performing a multiplication operation, the multiplied product 310 may include a set of multiplied values 312 C[0] to C[n] from multiplying each respective value of A to B. A lower precision level of the neural network may allow additional values to be included in each multiplication operation. For example, in an eighteen - bit multiplication operation, the following table may reflect the precision level regarding the number of values in each operand:

[0041] Precision (bits) Calculation improvement (low precision - 18 bits) 1 9x 2 5x 3 3x 4 2x 5 2x 6 2x

[0042] According to the following equation, this relationship can be generalized as follows:

[0043]

[0044] where num comp refers to the number of values that can be included in each operand, Width Mult refers to the number of bits in each operand, and precision refers to the number of bits used in the operations in neural network 100.

[0045] Figure 9 is a generalized block diagram of the circuit of IC 30 that performs arithmetic operations of a neural network using Figure 8 the data structure. IC 30 includes a circuit similar to the circuit described with respect to Figure 3 . In addition, additional accumulators (e.g., in code form) can be used for each of the values 312 in the product 310 of multiplication. Figure 7 described.

[0046] Figure 10 is a block diagram of another data structure 324 that can be used in combination with the circuit described with respect to Figure 9 . The data structure 324 includes a first operand having N values 326A[0] to A[n], with gaps 328 between each of the values in the values 326. The data structure 324 includes a second operand having a single value B[0] 330, and padding 332 is performed for the entire remainder of the first operand. The single value B[0] can have the same precision as each of the N values of the first operand. When performing a multiplication operation, IC 30 can determine a first product of multiplication 334 by multiplying the first value A[0] of the first operand by B[0] 330. IC 30 can determine a second product of multiplication 334 by multiplying the second value A[1] of the first operand by B[0] 330. That is, B[0] 330 can be multiplied by each of the N values 326 of the first operand.

[0047] Figure 11 is a flowchart of a process 340 performed by IC 30 in combination with Figure 5 and Figure 6 described by way of example to perform arithmetic operations of neural network 100 to output a result value 116. At block 342, IC 30 can combine (e.g., pack) a first value and a second value into a first operand. In some embodiments, IC 30 can combine (e.g., pack) a third value and a fourth value into a second operand. As mentioned above with respect to Figure 8 and Figure 9 , additional values can be included in each of the first operand and the second operand. In addition, with respect to Figure 10In a particular embodiment described, the IC 30 may have only a single value in the second operand. At block 344, the IC 30 may multiply the first operand by the second operand to determine a first multiplied product that is at least partially based on the first value and a second multiplied product that is at least partially based on the second value. In an example where the second operand includes a third value and a fourth value, for example, the first multiplied product may be the first value multiplied by the third value, and the second multiplied product may be the second value multiplied by the fourth value. In this way, more than one multiplied product can be determined according to the same multiplication operation performed by the hardened multiplier circuit. In some embodiments, the multiplication operation may be performed in the hardened multiplier circuit, and the multiplication result having both the first multiplied product and the second multiplied product may be output to the soft logic, where the first multiplied product and the second multiplied product may be split from each other.

[0048] In some embodiments, at block 346, the IC 30 may perform a correction operation (e.g., in the soft logic) to correct any conflicts in the multiplication operation. For example, if the first multiplied product overlaps with the second multiplied product due to an overflow, the IC 30 may perform an exclusive OR (XOR) operation, a masking operation, etc. to correct for the overlapping values. At block 348, the IC 30 may then add the first multiplied product, the second multiplied product, or both to at least one sum. For example, the first multiplied product and the second multiplied product may be added to the same sum 114, or each product may be separately added to different sums 114 and 115. At block 350, the IC 30 may output a result value that is at least partially based on the at least one sum. That is, the result value may be from the total of the two sums 114 and 115 (as in Figure 3 ), or the result value may be determined after several computational layers (as in Figure 5 ). The result value 116 may be output (e.g., via the I / O pin 36) to control the operation of the IC 30. In some embodiments, the result value 116 may be displayed to the user. In other embodiments, the result value 116 may be sent to another electronic device. By combining lower-precision values into operands having a design precision greater than that of the lower-precision values, the throughput through the neural network can be increased.

[0049] While the embodiments set forth in this disclosure may be susceptible to various modifications and alternative forms, specific embodiments have been shown by way of example in the drawings and have been described in detail herein. However, it is to be understood that the disclosure is not intended to be limited to the particular forms disclosed. The disclosure covers all modifications, equivalents, and alternatives falling within the spirit and scope of the disclosure as defined by the appended claims.

Claims

1. An integrated circuit device, comprising: a multiplier circuit; a first input circuit to the multiplier circuit, wherein the first input circuit is configurable to receive a first operand, wherein a first component of the first operand includes a first value, a second component of the first operand includes a second value, and a third component of the first operand includes a third value; a second input circuit to the multiplier circuit, wherein the second input circuit is configurable to receive a second operand, wherein a first component of the second operand includes a fourth value, a second component of the second operand includes a fifth value, and a third component of the second operand includes a sixth value, wherein the multiplier circuit is configurable to multiply the first operand by the second operand to produce a plurality of results corresponding to a plurality of different multiplication operations, wherein the plurality of results includes: a first result among the plurality of results, wherein the first result is equal to a first multiplication operation based on the first value and the fourth value; a second result among the plurality of results, wherein the second result is equal to a second multiplication operation based on the second value and the fifth value; and a third result among the plurality of results, wherein the third result is equal to a third multiplication operation based on the third value and the sixth value; and an accumulator circuit, wherein the accumulator circuit is configurable to perform separately adding each of the plurality of results to a previously generated corresponding result.

2. The integrated circuit device according to claim 1, wherein, the integrated circuit device includes a digital signal processing block, and the digital signal processing block includes the multiplier circuit, the first input circuit, and the second input circuit.

3. The integrated circuit device according to claim 1 or 2, wherein, the integrated circuit device includes an adder circuit.

4. The integrated circuit device according to claim 3, wherein, the adder circuit is configurable to receive one or more inputs including at least the first result, the second result, and the third result.

5. The integrated circuit device according to claim 1 or 2, wherein, each of the first value, the second value, and the third value has a precision between 1 bit and 6 bits.

6. The integrated circuit device according to claim 1 or 2, wherein, the integrated circuit device is configurable to perform arithmetic operations in a neural network.

7. A data processing system comprising the integrated circuit device according to any one of claims 1 to 6.

8. The data processing system according to claim 7, comprising: a processor; a memory coupled to the processor; and a network interface coupled to the memory and the processor, wherein the network interface is configured to send or receive data via a network to communicate with other electronic devices.

9. A method, comprising: receiving, via a first input circuit, a first component of a first operand including a first value, a second component of the first operand including a second value, and a third component of the first operand including a third value; Receiving, via a second input circuit, a first component of a second operand that includes a fourth value, a second component of the second operand that includes a fifth value, and a third component of the second operand that includes a sixth value; Multiplying the first operand by the second operand to produce a plurality of results corresponding to a plurality of different multiplication operations, wherein the plurality of results include: A first result among the plurality of results, wherein the first result is equal to a first multiplication operation based on the first value and the fourth value; A second result among the plurality of results, wherein the second result is equal to a second multiplication operation based on the second value and the fifth value; and A third result among the plurality of results, wherein the third result is equal to a third multiplication operation based on the third value and the sixth value; and Adding each result among the plurality of results separately to a previously generated corresponding result.

10. The method according to claim 9, wherein, The first value, the second value, and the third value each have a precision between 1 bit and 6 bits.

11. The method according to claim 9 or 10, wherein, The adder circuit can be configured to receive one or more inputs including at least the first result, the second result, and the third result.

12. The method according to claim 9 or 10, including performing arithmetic operations in a neural network.

13. A system, including an integrated circuit device, wherein, The integrated circuit device includes: A programmable logic circuit; A multiplier circuit; A first input circuit to the multiplier circuit, wherein the first input circuit can be configured to receive a first operand, wherein a first component of the first operand includes a first value, a second component of the first operand includes a second value, and a third component of the first operand includes a third value; and A second input circuit to the multiplier circuit, wherein the second input circuit can be configured to receive a second operand, wherein a first component of the second operand includes a fourth value, a second component of the second operand includes a fifth value, and a third component of the second operand includes a sixth value, wherein the multiplier circuit can be configured to multiply the first operand by the second operand to produce a plurality of results corresponding to a plurality of different multiplication operations, wherein the plurality of results include: A first result among the plurality of results, wherein the first result is equal to the product of a first multiplication operation based on the first value and the fourth value; A second result among the plurality of results, wherein the second result is equal to the product of a second multiplication operation based on the second value and the fifth value; and A third result among the plurality of results, wherein the third result is equal to the product of a third multiplication operation based on the third value and the sixth value; and An accumulator circuit, wherein the accumulator circuit can be configured to perform adding each result among the plurality of results separately to a previously generated corresponding result.

14. The system according to claim 13, wherein, The programmable logic circuit includes a multiplexer.

15. The system according to claim 13, wherein, the integrated circuit device includes a digital signal processing block, and the digital signal processing block includes the multiplier circuit, the first input circuit, and the second input circuit.

16. The system according to claim 13, wherein, the integrated circuit device can be configured to perform arithmetic operations in a neural network.

Citation Information

Patent Citations

  • Datapath circuit for digital signal processor

    CN103677736A

  • Multiplier operable to perform a variety of operations

    US7720901B1