A device for performing multiplication / accumulation
By optimizing GPU pipelines through value conversion and bit limitation techniques, the method enhances efficiency and reduces circuit size and power consumption in large-scale mathematical operations.
Patent Information
- Application Number
- JP2021094223
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-09
- Filing Date
- 2021-06-04
- Publication Date
- 2026-02-04
- Estimated Expiration
- 2041-06-04
AI Technical Summary
Existing GPUs and processing devices with multiple processing units face inefficiencies in performing large-scale mathematical operations, particularly in applications like artificial intelligence, due to limitations in circuit size and power usage.
A method and system that optimizes GPU pipelines by converting input values to unsigned format, limiting bit representation, and splitting signed values into multiple arguments for processing, allowing efficient multiplication and accumulation using smaller multipliers and adders, thereby reducing circuit area and power consumption.
This approach enables significant reductions in circuit area and power usage while maintaining accuracy, especially beneficial in GPUs with numerous computational units.
Smart Images

Figure 0007811090000002 
Figure 0007811090000003 
Figure 0007811090000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to a system and method for performing large amounts of mathematical operations.
[0002] One of the most common ways to increase execution speed is to perform operations in parallel, such as on multiple processor cores. This principle is exploited on a much larger scale by configuring graphics processing units (GPUs) with many (e.g., thousands) of processing pipelines, each of which can be configured to perform a mathematical function. In this way, large amounts of data can be processed in parallel. Although primarily used for graphics processing applications, GPUs are also often used for other applications, particularly artificial intelligence.
[0003] Improving the functionality of a GPU pipeline, or of any processing device that includes multiple processing units, would be an improvement in the art.
[0004] In order that the advantages of the invention may be readily understood, a more particular description of the invention briefly described above will be given by reference to specific embodiments which are illustrated in the accompanying drawings. The invention will be described and elucidated with additional specificity and detail through the use of the accompanying drawings, with the understanding that these drawings depict only typical embodiments of the invention and therefore should not be considered limiting of its scope. [Brief explanation of the drawings]
[0005] [Figure 1] FIG. 1 is a schematic block diagram of a computer system suitable for implementing methods according to embodiments of the invention. [Figure 2] FIG. 2 is a schematic block diagram of a multiply / accumulate circuit according to an embodiment of the present invention. [Figure 3] FIG. 2 is a process flow diagram of a method for processing input arguments in a multiply / accumulate circuit according to an embodiment of the present invention. [Figure 4]FIG. 1 is a process flow diagram of a method for post-processing products accumulated in a multiply / accumulate circuit according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0006] It will be readily understood that the components of the present invention, as generally described and illustrated in the Figures herein, could be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of embodiments of the invention, as represented in the Figures, is not intended to limit the scope of the invention as claimed, but is merely representative of certain examples of embodiments presently contemplated in accordance with the invention. The presently described embodiments will be best understood by reference to the drawings, in which like parts are designated with like numerals throughout.
[0007] Embodiments in accordance with the present invention may be embodied as an apparatus, a method, or a computer program product. Accordingly, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which may all be referred to generally herein as a "module" or a "system." Furthermore, the present invention may take the form of a computer program product embodied in any tangible medium of expression having computer-usable program code embodied in the medium.
[0008] Any combination of one or more computer usable or computer readable media, including non-transitory media, may be utilized. For example, computer readable media may include one or more of portable computer disks, hard disks, random access memory (RAM) devices, read-only memory (ROM) devices, erasable programmable read-only memory (EPROM or flash memory) devices, portable compact disk read-only memories (CDROMs), optical storage devices, and magnetic storage devices. In selected embodiments, computer readable media may include any non-transitory medium that may contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0009] Computer program code for carrying out operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, or C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code may execute entirely on a computer system as a standalone software package, on a standalone hardware unit, partially on a remote computer some distance from the computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., through the Internet using an Internet Service Provider).
[0010] The present invention is described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions or code. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, to produce a machine such that the instructions, executed via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart illustrations and / or block diagrams.
[0011] These computer program instructions may also be stored in a non-transitory computer-readable medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture that includes instruction means that implement the functions / acts identified in one or more blocks of the flowcharts and / or block diagrams.
[0012] Computer program instructions may also be loaded onto a computer or other programmable data processing device to cause a series of operational steps performed on the computer or other programmable device to produce a computer-implemented process, such that instructions executing on the computer or other programmable device provide a process for implementing the functions / acts identified in one or more blocks of the flowcharts and / or block diagrams.
[0013] 1 is a block diagram illustrating an exemplary computing device 100. Computing device 100 may be used to perform various procedures as discussed herein. Computing device 100 may function as a server, a client, or any other computing entity. The computing device may perform various monitoring functions as discussed herein and may execute one or more application programs, such as the application programs described herein. Computing device 100 may be any of a wide variety of computing devices, such as a desktop computer, a notebook computer, a server computer, a handheld computer, and a tablet computer.
[0014] Computing device 100 includes one or more processors 102, one or more memory devices 104, one or more interfaces 106, one or more mass storage devices 108, one or more input / output (I / O) devices 110, and a display device 130, all of which are coupled to a bus 112. Processor 102 includes one or more processors or controllers that execute instructions stored in memory device(s) 104 and / or mass storage device(s) 108. Processor 102 may also include various types of computer-readable media, such as cache memory.
[0015] The memory device 104 includes a variety of computer-readable media, such as volatile memory (e.g., random access memory (RAM) 114) and / or non-volatile memory (e.g., read-only memory (ROM) 116). The memory device 104 may also include re-writable ROM, such as flash memory.
[0016] Mass storage device 108 includes a variety of computer-readable media, such as magnetic tape, magnetic disks, optical disks, and solid-state memory (e.g., flash memory). As shown in Figure 1, a particular mass storage device is a hard disk drive 124. Mass storage device 108 may also include a variety of drives to enable reading from and / or writing to a variety of computer-readable media. Mass storage device 108 includes removable media 126 and / or non-removable media.
[0017] I / O devices 110 include various devices that allow data and / or other information to be entered into and / or retrieved from computing device 100. Exemplary I / O devices 110 include cursor control devices, keyboards, keypads, microphones, monitors or other display devices, speakers, printers, network interface cards, modems, lenses, and CCD or other image capture devices, etc.
[0018] Display device 130 includes any type of device capable of displaying information to one or more users of computing device 100. Examples of display device 130 include monitors, display terminals, video projection devices, and the like.
[0019] A graphics processing unit (GPU) 132 may be coupled to processor 102 and / or display device 130. The GPU may be operable to render computer-generated images and perform other graphics processing. The GPU may include some or all of the functionality of a general-purpose processor, such as processor 102. The GPU may also include additional functionality specific to graphics processing. The GPU may include hard-coded and / or hard-wired graphics functionality related to coordinate transformations, shading, texturing, rasterization, and other functions useful for rendering computer-generated images.
[0020] The interface 106 includes various interfaces that allow the computing device 100 to interact with other systems, devices, or computing environments. An exemplary interface 106 includes any number of different network interfaces 120, such as interfaces to a local area network (LAN), a wide area network (WAN), a wireless network, and the Internet. Other interfaces include a user interface 118 and a peripheral device interface 122. The interface 106 may also include one or more user interface elements 118. The interface 106 may also include one or more peripheral interfaces, such as interfaces for a printer, a pointing device (mouse, trackpad, etc.), a keyboard, etc.
[0021] Bus 112 allows processor 102, memory device 104, interface 106, mass storage device 108, and I / O device 110 to communicate with each other, along with other devices or components coupled to bus 112. Bus 112 represents one or more of several types of bus structures, such as a system bus, a PCI bus, an IEEE 1394 bus, and a USB bus.
[0022] In some embodiments, processor 102 may include cache 134, such as one or both of an L1 cache and an L2 cache. GPU 132 may also include cache 136, which may also include one or both of an L1 cache and an L2 cache.
[0023] For purposes of explanation, programs and other executable program components are illustrated herein as separate blocks, but it is understood that such programs and components may reside at various times in different storage components of computing device 100 and be executed by processor 102. Alternatively, the systems and procedures described herein may be implemented in hardware, or a combination of hardware, software, and / or firmware. For example, one or more application-specific integrated circuits (ASICs) may be programmed to execute one or more of the systems and procedures described herein.
[0024] Referring to FIG. 2, the GPU 132 or other components of the computing device 100 may include the components shown in FIG. 2. As shown, buffers 200, 202 may store arguments to be multiplied / accumulated. For example, one buffer 200 may store coefficients for implementing a graphics processing operation (e.g., a kernel), an artificial intelligence operation (e.g., as part of a convolutional neural network), or the like. The other buffer 202 may store values (often referred to as "active") to be multiplied by the coefficients. This is, of course, merely an example; any value may be loaded into the buffers 200, 202 and may be the subject of multiplication / accumulation. The buffers 200, 202 may be defined as part of memory (e.g., RAM 114) or as part of the cache 134 or 136.
[0025] Each value retrieved from buffers 200, 202 may be input to separator 204, which converts all values to unsigned values. For example, in some applications, values may be represented in the following format: [type][magnitude]. The [type] field indicates whether the bits in [magnitude] represent a signed or unsigned number, with 0 indicating unsigned and 1 indicating signed. If the [type] field indicates a signed value, the most significant bit (MSB) in the [magnitude] field will be 1 for negative numbers and 0 for positive numbers if two's complement representation is used.
[0026] The output of separator 204 will be sign 206 and magnitude 208 for values from buffer 200 and sign 210 and magnitude 212 for values from buffer 202 .
[0027] The signs 206, 210 and magnitudes 208, 212 may then be input to a checker 210. The checker 210 evaluates the magnitudes 208, 212 to detect certain cases requiring special handling. Specifically, to limit the size of the circuitry that performs the actual multiplication and addition of the multiplication / accumulation, the number of bits used to represent the magnitudes 208, 212 may be limited to a number of bits N. For example, if a value is defined as using the format [type][magnitude], the value of N may be the number of bits in the value actually input to the multiplication circuit, which may be less than the number of bits in the [magnitude] field. For example, if there are 9 bits in the input value, there will be 8 bits in the [magnitude] field. Thus, in some embodiments, the number of bits N input to the multiplication circuit per buffer 200, 202 may be N=7.
[0028] However, for signed values, 7 bits may be insufficient to represent the magnitude of the most negative number representable using 8 signed bits; for example, 7 unsigned bits may only represent 0 to 127, while 8 signed bits may represent -128 to 127. The largest positive number that can be represented by the [magnitude] field of a signed value is referred to herein as MaxSign and may be defined as 2^N-1, where the number of bits in the [magnitude] field is N+1.
[0029] Thus, the checker 214 may detect instances where the magnitude 208, 212 exceeds MaxSign and may make adjustments accordingly. The way this scenario is handled is described below with respect to FIG.
[0030] For unsigned values, the maximum value represented using N+1 bits is 2^(N+1)-1. Thus, values between 2^N and 2^(N+1)-1 cannot be represented using N bits. Checker 210 may similarly reason if magnitude 208, 212 for an unsigned value exceeds MaxSign, as described below with respect to FIG. 3, and may make adjustments accordingly.
[0031] The output of checker 214 is, for example, a pair of arguments for a pair of values from buffers 200, 202, and the output of checker 214 is one or more pairs of arguments that are input to sequencer 216. Sequencer 216 presents the argument pairs to computation unit 218. Specifically, there may be multiple computation units, for example, 8, 64, 1024, or any number of computation units. Sequencer 216 implements logic for presenting the arguments to the correct computation unit. Specifically, sequencer 216 ensures that the arguments for a pair of values from buffers 200, 202 are presented to computation unit 218, which accumulates the multiplication / addition result for that pair of values.
[0032] For example, in matrix multiplication, each value in the output matrix is the result of the dot product of a row of a first matrix and a column of a second matrix. Thus, in this example, sequencer 216 presents arguments to pairs of input values from buffers 200, 202 such that each computation unit 218 can accumulate the sum of the products of the elements of a particular row and the elements of that row's corresponding column. Of course, this is just one example, and sequencer 216 can be programmed to accumulate products according to any desired function.
[0033] Each computation unit 218 may include an N-bit multiplier 220 that takes as input a pair of arguments from sequencer 216, and an adder 222 that takes as input the product of multiplier 220 and the contents of accumulation buffer 224. The output of adder 222 is written back to accumulation buffer 224. The result of accumulation buffer 224 may be read by a controller of GPU 132, CPU 102, under application control, or according to any approach known to those skilled in the art for retrieving and processing the results of a multiplication / accumulation. As shown in FIG. 2, adder 222 may further take as input the sign of the input argument, as separated by separator 204.
[0034] 3, the illustrated method 300 may be performed by a checker 210 to determine whether to split an input magnitude 208, 212 into two arguments or to output a single argument containing the magnitude 208, 212. The method 300 may be performed for each input magnitude 208, 212, hereafter referred to as an "input magnitude."
[0035] The method 300 may include receiving 302 the magnitude and type of the input magnitude from the separator 204. If the type is found to be signed, the method 300 may include obtaining 306 the absolute value of the input magnitude.
[0036] Method 300 may then include evaluating 308 whether the absolute value is greater than MaxSign. If not, the absolute value may be input 314 as an argument (Arg) to sequencer 216. If so, method 300 may include splitting 310 the absolute value into two arguments (Arg_1, Arg_2). Specifically, for signed values, the value that is one greater than MaxSign possible is MaxSign+1, and Arg_1 and Arg_2 may be set equal to (MaxSign+1) / 2 accordingly. For example, for MaxSign=127, step 310 may include setting Arg_1=Arg_2=64.
[0037] The method 300 may then include inputting 312 Arg_1 and Arg_2 into the sequencer 216.
[0038] If the input magnitude is found not to be from a signed number, method 300 may include evaluating whether the input magnitude is greater than MaxSign 318. If so, two arguments are set according to steps 318 and 320: Arg_1 is set equal to the input argument minus MaxSign, and Arg_2 is set equal to MaxSign.
[0039] Arg_1 and Arg_2 are then input to the sequencer 322. If the input magnitude is not found to exceed MaxSign 316, it is input to the sequencer 216 as an argument (Arg).
[0040] The inputting of arguments in steps 312, 314, 322, and 324 may be performed in a coordinated manner. Specifically, one or more arguments as determined relative to the input magnitude for a value from buffer 200 may be input to sequencer 216 in coordination with one or more arguments as determined relative to the input magnitude for a corresponding value from buffer 202.
[0041] As described above, the first and second values to be multiplied together may be retrieved from buffers 200, 202, respectively, and processed by separator 204 and checker 210. The pairs of factors that may be input to sequencer 216 for various outcomes of method 300 for the first and second values are set forth in Table 1. Specifically, for the first value, the possible outcomes of method 300 are either a single output argument designated as Arg1 (step 312 or 324) or two output arguments designated here as Arg1_1 and Arg1_2 (step 312 or step 322). For the second value, the possible outcomes are designated as a single argument Arg2 (step 312 or 324) or two of the output arguments Arg2_1 and Arg2_2 (step 312 or step 322). In “Input to Sequencer String,” each pair in parentheses indicates a pair of arguments input to sequencer 216 that will be multiplied together and accumulated by calculation unit 218 . [Table 1]
[0042] The sequencer 216 may be programmed to input pairs of arguments to the same computation unit 218 that correspond to the first and second values. The sequencer 216 may also associate the sign of each argument with it. Specifically, if a signed value is split into two arguments, Arg_1 and Arg_2, a sign is associated by the sequencer 216 with both of the arguments, Arg_1 and Arg_2, in all argument pairs that include either of those arguments.
[0043] 4 illustrates a method 400 for performing a multiply-accumulate calculation on pairs of arguments input by sequencer 216 to computation units 218. Each pair of arguments input to sequencer 216 will be input to multiplier 220 of one of the computation units, which will then compute 402 the product P. Method 400 may further include evaluating one or both of the argument types and signs of the argument pairs. For example, for arguments obtained from unsigned values, the sign in step 404 may be assumed to be positive in all cases. For signed values, the sign will be the sign 206, 210 separated from the signed value by separator 204.
[0044] If only one of the arguments is found to have a negative sign, method 400 may include adjusting 408 the product P. If there is one negative argument, the sign of P is changed to negative, i.e., P is converted to a negative number, such as by the definition of two's complement. The negative product P is input to summer 222, which sums the negative product P with the current contents of accumulation buffer 224 and writes the resulting sum to accumulation buffer 224.
[0045] If none of the arguments are found to be negative, the product P is input 410 to the summer 222, which will then sum the product P with the current contents of the accumulation buffer 224 and write the result of the summation to the accumulation buffer 224.
[0046] As is clear from the above discussion, the multiplier 220 can be made much smaller while providing the same level of accuracy using the approach described in Figures 2-4. In applications such as GPUs where there are hundreds or thousands of computational units 218, this results in significant reductions in circuit area and power usage.
[0047] The present invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.
[0048] The requests are listed below:
Claims
1. configured to receive two input values, and for each of the two input values: if each input value has a type field set to indicate that it is a signed value, converting each input value to a sign value and a magnitude value for that input value; if each of the input values has the [type] field set to indicate that it is an unsigned value, setting the magnitude value of each of the input values to be the respective input value; a separator for a checker configured to convert the magnitude values of the two input values into one or more pairs of input arguments; a computation unit configured to perform an operation on each of said one or more pairs of input arguments and on any sign value of said two input values, and to produce an output in accordance with said operation; Including, an input argument of said one or more pairs of input arguments having N bits, where N is a predefined integer; The checker (a) for a first magnitude (M1) of the magnitude values of the two input values: expressing M1 using arguments Arg1_1=2^(N-1) and Arg1_2=2^(N-1), where M1 corresponds to a first of said two input values that is signed and is greater than 2^N-1; (b) for a second magnitude (M2) of the magnitude values of the two input values: If M2 corresponds to a second of the two input values that is signed and is greater than 2^N-1, then express M2 using arguments Arg2_1=2^(N-1) and Arg2_2=2^(N-1); and further configured to convert the magnitude values of the two input values to the one or more pairs of input arguments by The device, wherein the computation unit includes two or more multipliers having two N-1 bit multiplication inputs.
2. The device of claim 1 , wherein the computation unit is programmed to perform multiplication / accumulation.
3. 2. The device of claim 1, wherein the checker is configured to convert the magnitude values of the two input values of the one or more pairs of input arguments such that an input argument of the one or more pairs of input arguments has fewer bits than the magnitude values of the two input values.
4. 4. The device of claim 3, wherein the input argument of the one or more pairs of input arguments is one bit less than the magnitude values of the two input values.
5. The checker (c) for the first magnitude: where M1 corresponds to a first input value of the two input values that is unsigned, and if M1 is greater than 2^N-1, splitting M1 into arguments Arg1_1=M1-2^N+1 and Arg1_2=2^N-1; (d) for said second magnitude: M2 corresponds to the second input value of the two input values that are unsigned, and if M2 is greater than 2^N-1, splitting M2 into arguments Arg2_1=M2-2^N+1 and Arg2_2=2^N-1; 2. The device of claim 1, configured to convert the magnitude values of the two input values into the one pair of input arguments by:
6. The checker (e) if M1 is less than or equal to 2^N-1, setting argument Arg1 of said one or more pairs of arguments to be M1; (f) if M2 is less than or equal to 2^N-1, setting argument Arg2 of said one or more pairs of arguments to be M2; 6. The device of claim 5, configured to convert the magnitude values of the two input values into the one or more pairs of input arguments by:
7. The checker if the result of (a)-(f) is Arg1 for M1 and Arg2 for M2, outputting one pair of input arguments that is (Arg1, Arg2); if the result of (a)-(f) is Arg1 for M1 and Arg2_2 for M2, output two pairs of input arguments (Arg1,Arg2_1) and (Arg1,Arg2_2); if the results of (a) through (f) are Arg1_1 and Arg1_2 for M1 and Arg2_1 and Arg2_2 for M2, then output four pairs of input arguments (Arg1_1,Arg2_1), (Arg1_1,Arg2_2), (Arg1_2,Arg2_1), and (Arg1_2,Arg2_2); If the results of (a) through (f) are Arg1_1 and Arg1_2 for M1 and Arg2 for M2, then output two pairs of input arguments (Arg1_1, Arg2) and (Arg1_2, Arg2); 7. The device of claim 6, configured to convert the magnitude values of the two input values to the one or more pairs of input arguments by:
8. 8. The device of claim 7, further comprising a sequencer programmed to input the one or more pairs of input arguments to the computation unit, the computation unit being programmed to perform a multiply-accumulate operation.
9. The computation unit, for each pair of input arguments of the one or more pairs of input arguments: Computing the product P of each said pair of inputs; (g) if only one of the two input values from which each pair of input arguments is derived according to (a) through (f) is a negative-signed number; Setting P = -P; after performing (g), adding P to the contents of the accumulation buffer to obtain a sum, and writing said sum to said accumulation buffer; 9. The device of claim 8, programmed to:
10. The device of claim 1 , wherein the separator is configured to read the two input values from a coefficient buffer and an active buffer.
11. receiving a first input value and a second input value; converting the first input value and the second input value into one or more pairs of input arguments, each argument of the one or more pairs of input arguments having fewer bits than the first input value and the second input value such that each argument has no more bits than a multiplication input of a multiplier in a computation unit, the multiplication input having fewer bits than the first input value and the second input value; inputting said one or more pairs of input arguments into said computation unit; It is programmed to an input argument of said one or more pairs of input arguments having N bits, where N is a predefined integer; The device is (a) if the first input value is signed and the magnitude (M1) of the first input value is greater than 2^N-1, splitting M1 into arguments Arg1_1=2^(N-1) and Arg1_2=2^(N-1); (b) if the second input value is signed and the magnitude (M2) of the second input value is greater than 2^N-1, splitting M2 into arguments Arg2_1=2^(N-1) and Arg2_2=2^(N-1); the device is further configured to convert the first input value and the second input value into the one or more pairs of input arguments by
12. The device of claim 11 , wherein the computation unit performs multiply-accumulate operations.
13. 12. The device of claim 11, wherein N is one bit less than the number of bits in the magnitude values M1 and M2.
14. The device comprises: (c) if the first input value is unsigned and M1 is greater than 2^N-1, splitting M1 into arguments Arg1_1=M1-2^N+1 and Arg1_2=2^N-1; (d) if the second input value is unsigned and M2 is greater than 2^N-1, splitting M2 into arguments Arg2_1=M2-2^N+1 and Arg2_2=2^N-1; 12. The device of claim 11, further configured to convert the first input value and the second input value into the one or more pairs of input arguments by:
15. The device comprises: (e) if M1 is less than or equal to 2^N-1, setting argument Arg1 of said one or more pairs of arguments to be M1; (f) if M2 is less than or equal to 2^N-1, setting argument Arg2 of said one or more pairs of arguments to be M2; 15. The device of claim 14, further configured to convert the first input value and the second input value into the one or more pairs of input arguments by:
16. The device comprises: If the result of (a) through (f) is Arg1 for M1 and Arg2 for M2, converting the first input value and the second input value to a pair of input arguments that is (Arg1, Arg2); if the result of (a)-(f) is Arg1 for M1 and Arg2_1 and Arg2_2 for M2, converting the first input value and the second input value into two pairs of input arguments (Arg1, Arg2_1) and (Arg1, Arg2_2); if the results of (a) through (f) are Arg1_1 and Arg1_2 for M1 and Arg2_1 and Arg2_2 for M2, converting the first input value and the second input value into four pairs of input arguments: (Arg1_1,Arg2_1), (Arg1_1,Arg2_2), (Arg1_2,Arg2_1), and (Arg1_2,Arg2_2); if the results of (a) through (f) are Arg1_1 and Arg1_2 for M1 and Arg2 for M2, converting the first input value and the second input value into two pairs of input arguments (Arg1_1, Arg2) and (Arg1_2, Arg2); 16. The device of claim 15, further configured to convert the first input value and the second input value into the one or more pairs of input arguments by:
17. The device, for each pair of the one or more input argument pairs of inputs, computing a product P of each said pair of input arguments; (g) if only one of the first and second input values is a negative signed number; Setting P = -P; after performing (g), adding P to the contents of the accumulation buffer to obtain a sum, and writing said sum to said accumulation buffer; The device of claim 16 , further configured to:
18. The device of claim 11 , wherein the device is further configured to read the first input value from a coefficient buffer and the second input value from an active buffer.
Citation Information
Patent Citations
Device for instructing multiplication of floating point binary four times length word format
JP1999296346A
Neural network unit performing efficient three-dimensional convolution
JP2018092560A
Multiplication of the first and second operands using redundant representation
JP2019500673A
N bit by M bit multiplication of twos complement numbers using N / 2+1 X M / 2+1 bit multipliers
US6347326B1