Matrix multiplier implemented to perform concurrent storage and multiply-accumulate (MAC) operations
By using multiple accumulators to perform MAC operations and storage operations concurrently in the matrix multiplier engine, the problem of MAC operation stopping in the prior art is solved, and the operation efficiency and processing speed of the matrix multiplier are improved.
Patent Information
- Application Number
- CN202380077514.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-17
- Filing Date
- 2023-10-26
- Publication Date
- 2025-06-13
AI Technical Summary
When performing MAC operations, existing matrix multipliers need to stop the MAC operation to perform storage operations, resulting in inefficient operation.
By introducing multiple accumulators into the matrix multiplier engine and performing concurrent operations with the controller, we ensure that the MAC operation continues to be performed while storing the result value, avoiding the stop of the MAC operation.
The operation efficiency of the matrix multiplier is improved, the MAC operation pauses caused by storage operations are reduced, and the overall processing speed is improved.
Smart Images

Figure CN120153347A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This patent application claims the benefit of priority to pending U.S. non - provisional application No. 17 / 989,448, filed on November 17, 2022, and the application is assigned to the assignee of the present application and is hereby incorporated herein by reference in its entirety as if fully set forth below and for all applicable purposes. Technical field
[0003] Aspects of the present disclosure generally relate to matrix multipliers, and more particularly to matrix multipliers implemented to perform concurrent storage and multiply - accumulate (MAC) operations. Background art
[0004] Matrix multiplier processors or engines can be used to perform different types of matrix and / or vector multiplications for various applications. For example, a matrix multiplier engine can be employed to perform machine learning (ML) operations, image processing, facial feature extraction, object detection, speech - to - text processing, and so on. As these types of data - processing devices are continuously improved in terms of speed and operational efficiency, there is an interest in improving matrix multiplier engines to achieve higher speeds and operational benefits. Summary of the invention
[0005] The following presents a brief summary of one or more implementations in order to provide a basic understanding of such implementations. This summary is not an exhaustive overview of all contemplated implementations and is not intended to identify key or critical elements of all implementations nor to delineate the scope of any or all implementations. Its sole purpose is to present some concepts of one or more implementations in a simplified form as a prelude to the more detailed description that follows.
[0006] One aspect of the present disclosure relates to a device, comprising: a memory; a matrix multiplier engine coupled to the memory, the matrix multiplier engine including: an array of multiplier - accumulate units (MAUs) including: a first set of accumulators; and a second set of accumulators; and a controller coupled to the matrix multiplier engine and the memory, the controller being configured to concurrently perform the following operations: causing a first set of result values in the first set of accumulators to be transferred to the memory according to a first set of storage instructions, wherein the first set of result values is generated according to a first set of multiply - accumulate (MAC) operations performed by the set of multipliers and the first set of accumulators; and causing the set of multipliers and the second set of accumulators to perform a second set of MAC operations.
[0007] Another aspect of the present disclosure relates to a method of performing matrix multiplication. The method includes: transferring a first set of result values from a first set of accumulators to a memory, wherein the first set of result values is generated from a first set of multiply-accumulate (MAC) operations; and concurrently with transferring the first set of result values from the first set of accumulators to the memory, performing a second set of MAC operations using a second set of accumulators.
[0008] Another aspect of the present disclosure relates to an apparatus, including: a unit configured to transfer a first set of result values from a first set of accumulators to a memory, wherein the first set of result values is generated from a first set of multiply-accumulate (MAC) operations; and a unit configured to concurrently with transferring the first set of result values from the first set of accumulators to the memory, perform a second set of MAC operations using a second set of accumulators.
[0009] To achieve the above and related purposes, one or more implementations include the features described in detail below and particularly pointed out in the claims. The following description and the drawings set forth certain illustrative features of one or more implementations in detail. However, these aspects indicate only a few of the various ways in which the principles of the implementations may be employed, and the described implementations are intended to include all such aspects and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1A A block diagram of an example matrix multiplier engine in accordance with one aspect of the present disclosure is shown.
[0011] Figure 1B A block diagram of an example matrix multiplier system in accordance with another aspect of the present disclosure is shown.
[0012] Figure 1C A block diagram of an example multiplier-accumulator unit (MAU) in accordance with another aspect of the present disclosure is shown.
[0013] Figure 1D A sequence diagram of an example set of instructions for operating a matrix multiplier system in accordance with another aspect of the present disclosure is shown.
[0014] Figure 2A A block diagram of another example multiplier-accumulator unit (MAU) in accordance with another aspect of the present disclosure is shown.
[0015] Figure 2B A sequence diagram of another example set of instructions for operating a matrix multiplier system in accordance with another aspect of the present disclosure is shown.
[0016] Figure 2CA flowchart illustrating an example method of concurrently storing a set of values obtained from a previous set of multiply-accumulate (MAC) operations and performing a current set of MAC operations according to another aspect of the present disclosure.
[0017] Figure 3A A block diagram illustrating another example multiplier-accumulator unit (MAU) according to another aspect of the present disclosure.
[0018] Figure 3B A sequence diagram illustrating another example set of instructions for operating a matrix multiplier system according to another aspect of the present disclosure.
[0019] Figure 3C A flowchart illustrating another example method of concurrently storing a set of values obtained from a previous set of multiply-accumulate (MAC) operations and performing a current set of MAC operations according to another aspect of the present disclosure.
[0020] Figure 4 A block diagram illustrating an example multi-core integrated circuit (IC) including a matrix multiplier engine according to another aspect of the present disclosure.
[0021] Figure 5 A flowchart illustrating an example method of performing matrix multiplication according to another aspect of the present disclosure.
[0022] Figure 6 A block diagram illustrating an example apparatus for performing matrix multiplication according to another aspect of the present disclosure. Detailed Description
[0023] The detailed description set forth below in connection with the appended drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein may be practiced. The detailed description includes specific details for providing a thorough understanding of various concepts. However, it will be apparent to those skilled in the art that the concepts may be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form in order to avoid obscuring such concepts.
[0024] Figure 1AA block diagram showing an example matrix multiplier processor or engine 100 according to one aspect of the present disclosure. The matrix multiplier engine 100 includes a multi-dimensional (e.g., two-dimensional) array (e.g., a 16x16 MAU array, but can have different dimensions and can be a one-dimensional array or an array with more than two dimensions) of multiplier-accumulator units (MAUs) 110 / 210 / 310. The matrix multiplier engine 100 further includes a first set of one or more cascaded input registers 120 for providing a set of input vectors “A” to the rows of the two-dimensional array of MAUs 110 / 210 / 310. Similarly, the matrix multiplier engine 100 further includes a second set of one or more cascaded registers 130 for providing a set of input vectors “B” to the columns of the two-dimensional array of MAUs 110 / 210 / 310.
[0025] Figure 1B A block diagram showing an example matrix multiplier system 155 according to another aspect of the present disclosure. The matrix multiplier system 155 includes a matrix multiplier engine 100, an associated controller 140, and a memory 150. The matrix multiplier engine 100 is configured to sequentially receive a set of input vectors A and a set of input vectors B, and its MAU array 110 / 210 / 310 is configured to perform a set of MAC operations (via a set of corresponding multipliers 112 and accumulators 114) based on a set of MAC instructions executed by the controller 140. Once the MAU array 110 / 210 / 310 completes the set of MAC operations and the result values are held in the set of accumulators 114, the controller 140 can cause the result values in the set of accumulators 114 to be stored or transferred (e.g., row by row or column by column) to the memory 150 based on the execution of a set of store instructions.
[0026] Figure 1C A block diagram showing an example multiplier-accumulator unit (MAU) 110-ij according to another aspect of the present disclosure. The MAU 110-ij can be an example implementation of the MAU in the i-th row and j-th column of the two-dimensional array of MAUs 110. All other MAUs in the array 110 can be implemented similarly.
[0027] The MAU 110-ij includes a multiplier 112-ij and an accumulator 114-ij. The multiplier 112-ij includes an input for receiving an operand a i from the input vector A and an operand b j from the input vector B. The multiplier 112-ij is configured to multiply the operand a i and b j and to output the result value or product (a i x b j) are stored additively in accumulator 114-ij. Each of these operations may be referred to as a multiply-accumulate (MAC) operation or a matrix outer product (MOP) operation. This collection of multipliers and accumulators in the MAU array 110 may be collectively referred to as multiplier set 112 and accumulator set 114.
[0028] Figure 1D A sequence diagram showing an example instruction set 160 / 260 / 360 for operating the matrix multiplier engine 100 according to another aspect of the present disclosure is shown. The vertical axis of this sequence diagram represents time. The instruction set 160 is provided to the controller 140 for controlling the operations of the matrix multiplier engine 100 and the memory 150.
[0029] Specifically, the instruction set 160 includes a first set of MAC instructions (individually referred to as MAC-1). The first set of MAC instructions instructs at least a subset of the MAU array 110 to perform a first set of MAC operations based on a first set of input vectors A and B.
[0030] After the first set of MAC operations, the instruction set 160 includes a first set of store instructions (individually referred to as STORE-1). The first set of store instructions instructs the matrix multiplier engine 100 to store or transfer the result values (as a result of the first set of MAC operations) in the accumulator set 114 of the MAU array 110 to the memory 150 (e.g., row by row or column by column). Note that during this store operation, the MAU array 110 does not perform MAC operations; in other words, when the accumulator set 114 is being used for storage purposes, the MAC operations stop.
[0031] After the first set of store instructions, the instruction set 160 includes a zero instruction for zeroing, clearing, or resetting the accumulator set 114 in the MAU array 110. The zero instruction may indicate that subsequent MAC operations (MAC-2) performed by the MAU array 110 are independent of the MAC operations according to the first set of MAC instructions (MAC-1).
[0032] Then, the instruction set 160 includes a second set of MAC instructions (individually referred to as MAC-2). The second set of MAC instructions instructs at least another subset in the MAU array 110 to perform a second set of MAC operations based on a second set of input vectors A and B. As described above, the second set of MAC operations (MAC-2) may be independent of the first subset of MAC operations (MAC-1).
[0033] Similarly, after the second set of MAC operations (MAC-2), the instruction set 160 includes a subset of second store instructions (collectively referred to as STORE-2). The set of second store instructions (STORE-2) directs the matrix multiplier engine 100 to store or transfer the result values (as a result of the second set of MAC operations) in the accumulator set 114 of the MAU array 110 to the memory 150 (e.g., row-to-row or column-to-column). Again, during this store operation, the MAC operations performed by the MAU array 110 are stopped. After the set of second store instructions (STORE-2), the instruction set 160 includes another zero instruction to zero, clear, or reset the accumulator set 114 of the MAU array 110.
[0034] As shown, stopping the MAC operations during the store operation creates inefficiencies in the operation of the matrix multiplier engine 100.
[0035] Figure 2A A block diagram of another example multiplier-accumulator unit (MAU) 210-ij according to another aspect of the present disclosure is shown. Similarly, the MAU 210-ij can be an example implementation of the MAU in the i-th row and j-th column of a two-dimensional array 210 of newly designed MAUs for the matrix multiplier engine 100. All other MAUs in the array 210 can be implemented similarly.
[0036] The MAU 210-ij includes a multiplier 212-ij, a pair of accumulators 214-1ij(-1) and 214-2ij(-2) (e.g., two (2) in this example, but can be more as further discussed herein), and a demultiplexer 216-ij. The multiplier 212-ij includes inputs for receiving an operand a i from the input vector A and an operand b j from the input vector B. The multiplier 212-ij is configured to multiply the operand a i and b j , and store the result value or product (a i x b j ) additively in the "active" accumulator of the pair of accumulators 214-1ij and 214-2ij. Similarly, each of these operations can be referred to as a multiply-accumulate (MAC) operation or a matrix outer product (MOP) operation.
[0037] More specifically, the output of multiplier 212-ij is coupled to the input of demultiplexer 216-ij. Demultiplexer 216-ij includes first and second outputs respectively coupled to the pair of accumulators 214-1ij and 214-2ij. Demultiplexer 216-ij includes a select input configured to receive a control signal (SEL) from controller 140 to select which one of the pair of accumulators 214-1ij and 214-2ij is active. The set of multipliers, accumulators, and demultiplexers of MAU array 210 may be collectively referred to as multiplier set 212, accumulator set 214, and demultiplexer set 216.
[0038] As discussed in more detail below, by using multiple accumulators per MAU, the stalls of MAC operations as discussed with reference to instruction set 160 can be eliminated; thus, making the operation of matrix multiplier engine 100 having MAU array 210 substantially more efficient.
[0039] Figure 2B A sequence diagram of another example instruction set 260 for operating matrix multiplier system 100 in accordance with another aspect of the present disclosure is shown. Similarly, the vertical axis of this sequence diagram represents time. Further, the order of instruction set 260 is indicated vertically downwards and horizontally to the right. Instruction set 260 may include a first set of MAC instructions (individually referred to as MAC-1). The first set of MAC instructions (MAC-1) indicates that at least a subset of MAU array 210 performs a first set of MAC operations based on first input vector sets A and B. Note that in this example, the set of "active" accumulators for the first set of MAC operations is the accumulator-1 of the MAUs of array 210 (e.g., accumulator set 214-1) (as selected by controller 140 via the SEL control signal).
[0040] After the first set of MAC instructions (MAC-1), instruction set 260 includes a first set of store instructions (individually referred to as STORE-1). The first set of store instructions (STORE-1) indicates that matrix multiplier engine 100 stores or transfers the result values (as a result of the first set of MAC operations) in the first set of accumulators-1 214-1 of MAU array 210 (e.g., row by row or column by column) to memory 150.
[0041] Note that in this example, the controller 140 looks ahead to determine if there is a first zero instruction after the first set of store instructions (STORE-1), and uses it as a break point or instruction to zero, clear, or reset the "inactive" accumulator-2 (214-2 in this example), making the "inactive" accumulator-2 (214-2) the new "active" accumulator, and designating accumulator-1 (214-1) as the "inactive" or "store" accumulator. The above operations can be referred to as accumulator swapping or renaming. Subsequently, concurrently with the matrix multiplier engine 100 executing the first set of store instructions (STORE-1) to transfer the result values in the set of inactive group accumulators-1 (214-2) to the memory 150, the controller 140 can initialize at least a subset of the MAU array 210 to perform a second set of MAC operations (based on the second set of MAC instructions MAC-2).
[0042] Thus, in this example, the MAC operations do not stop, and the operation of the matrix multiplier engine 100 including the newly designed MAU 210 is significantly more efficient compared to the matrix multiplier engine including the MAU 110 discussed previously.
[0043] Figure 2C A flowchart of an example method 270 for concurrently storing a set of values obtained from a previous set of multiply-accumulate (MAC) operations and performing a current set of MAC operations is shown according to another aspect of the present disclosure. According to method 270, the controller 140 instructs / causes the matrix multiplier engine 100 to perform a first set of MAC operations (block 272) using a first set of accumulators (e.g., accumulator 214-1) based on a first set of MAC instructions.
[0044] Method 270 further includes: the controller 140 determines whether there is a breakpoint or an instruction (e.g., a zero instruction) after the first set of memory instructions (block 274). For example, the breakpoint or instruction can be sequentially located between the first set of memory instructions and the second set of MAC instructions. If, in block 276, the controller 140 determines that there is a breakpoint or an instruction after the first set of memory instructions (meaning that the subsequent MAC operations are independent of the previous MAC operations), then the controller 140 instructs / causes the matrix multiplier engine 100 to perform a second set of MAC operations based on the second set of MAC instructions using a second set of accumulators (e.g., accumulators 214-2) (block 278). Moreover, concurrently with the matrix multiplier engine 100 performing the second set of MAC operations according to block 278, the controller 140 instructs / causes the matrix multiplier engine 100 to store or transfer the result values in the first set of accumulators (e.g., accumulators 214-1) to the memory 150 based on the first set of memory instructions (block 280). Method 270 further includes: the controller 140 instructs / causes the matrix multiplier engine 100 to store or transfer the result values in the second set of accumulators (e.g., accumulators 214-2) to the memory 150 based on the second set of memory instructions (block 282).
[0045] If, in block 276, the controller 140 determines that there is no breakpoint after the first set of memory instructions (e.g., in other words, after the first set of memory instructions is the second set of MAC instructions without an intermediate zero instruction), then the second set of MAC instructions depends on the first set of MAC instructions. In this case, the controller 140 instructs / causes the matrix multiplier engine 100 to store or transfer the result values in the first set of accumulators (e.g., accumulators 214-1) to the memory 150 based on the first set of memory instructions (block 284). Then, the controller 140 instructs the matrix multiplier engine 100 to perform a second set of MAC operations using the first set of accumulators (e.g., accumulators 214-1) based on the second set of MAC instructions (block 286).
[0046] Figure 3A A block diagram of another exemplary multiplier-accumulator unit (MAU) 310-ij according to another aspect of the present disclosure is shown. Similarly, the MAU 310-ij can be an exemplary implementation of the MAU in the i-th row and j-th column of a two-dimensional array 310 of newly designed MAUs for the matrix multiplier engine 100. All other MAUs in the array 310 can be implemented similarly.
[0047] The MAU 310-ij includes a multiplier 312-ij, a set of accumulators 314-1ij(-1) to 314-Nij(-N) (where N is an integer greater than or equal to 2), and a demultiplexer 316-ij. The multiplier 312-ij includes means for receiving an operand a from the input vector Ai and receives operand b from input vector B j as input. Multiplier 312-ij is configured to multiply operand a i and b j and stores the resultant value or product (a i x b j ) additively in an “active” accumulator among accumulator sets 314-1ij through 314-Nij. Similarly, each of these operations can be referred to as a multiply-accumulate (MAC) operation or a matrix outer product (MOP) operation.
[0048] More specifically, the output of multiplier 312-ij is coupled to the input of demultiplexer 316-ij. Demultiplexer 316-ij includes an output set respectively coupled to accumulator sets 314-1ij through 314-Nij. Demultiplexer 316-ij includes a select input configured to receive a control signal (SEL) from controller 140 to control which of accumulator sets 314-1ij through 314-Nij is active. The set of multipliers, accumulators, and demultiplexers of MAU array 310 can be collectively referred to as multiplier set 312, accumulator set 314, and demultiplexer set 316.
[0049] Similarly, by using multiple accumulators per MAU, the stalls of MAC operations as discussed with reference to instruction set 160 can be eliminated; thus, making the operation of matrix multiplier engine 100 having MAU array 310 substantially more efficient.
[0050] Figure 3B A sequence diagram of another example instruction set 360 for operating matrix multiplier engine 100 in accordance with another aspect of the present disclosure is shown. Similarly, the vertical axis of this sequence diagram represents time. Further, instruction set 360 is indicated in a vertically downward and horizontally rightward order. Instruction set 360 may include a first set of MAC instructions (individually referred to as MAC-1). In response to the first set of MAC instructions, controller 140 instructs / causes at least a subset of MAU array 310 to perform a first set of MAC operations based on first input vector sets A and B. Note that in this example, the “active” accumulator in the first set of MAC operations is 314-1 (accumulator-1) of MAU array 310 (as selected by controller 140 via SEL control signal).
[0051] After the first set of MAC instructions (MAC-1), the instruction set 360 includes a first subset of store instructions (collectively referred to as STORE-1). In response to the first subset of store instructions (STORE-1), the controller 140 instructs / causes the matrix multiplier engine 100 to store or transfer the result values in the "store" accumulator set 314-1 of the MAU array 310 (e.g., row by row or column by column) to the memory 150.
[0052] Note that in this example, the controller 140 looks ahead to the first zero instruction and uses it as a breakpoint or instruction to zero, clear, or reset an inactive accumulator (e.g., accumulator-2 or 314-2), making the inactive accumulator (e.g., accumulator-2 or 314-2) the new "active" accumulator (and the previous active accumulator-1 or 314-1 the "store" accumulator). The above operations may be referred to as accumulator swapping or renaming. Concurrent with the matrix multiplier engine 100 executing the first set of store instructions (STORE-1) to transfer the values in the "store" accumulator-1 or 314-1 to the memory 150, the controller 140 then initializes at least a subset of the MAU array 310 to perform a second set of MAC operations (collectively referred to as MAC-2).
[0053] After the second set of MAC instructions (MAC-2), the instruction set 360 includes a second subset of store instructions (collectively referred to as STORE-2). The second subset of store instructions (STORE-2) instructs the matrix multiplier engine 100 to store or transfer the result values in the "store" accumulator set (e.g., accumulator-2 or 314-2) of the MAU array 310 (e.g., row by row or column by column) to the memory 150.
[0054] Note that in this example, the matrix multiplier engine 100 looks ahead to the second zero instruction and uses it as a breakpoint or instruction to zero, clear, or reset another inactive accumulator (e.g., accumulator-3 or 314-3), making the other inactive accumulator-3 or 314-3 the new "active" accumulator (and the previous active accumulator-2 or 314-2 the "store" accumulator). Similarly, the above operations may be referred to as accumulator swapping or renaming. Concurrent with the matrix multiplier engine 100 still executing the first subset of store instructions (STORE-1) to transfer the values in the "store" accumulator-1 or 314-1 to the memory 150, the controller 140 then initializes at least a subset of the MAU array 310 to perform a third set of MAC operations (collectively referred to as MAC-3).
[0055] This is why there may be more than two accumulators, because when the third MAC subset is executed, the first set of storage operations may not have been completed. Therefore, the first accumulator - 1 cannot be used for the third MAC subset because the first accumulator - 1 is being used for storage. Similarly, the second accumulator - 2 cannot be used because the second accumulator - 2 holds the value associated with the second MAC operation subset that is to be stored or transferred to the memory 150 after the completion of the first set of storage instructions.
[0056] Figure 3C A flowchart of an example method 370 is shown for concurrently storing a set of values obtained from a previous set of multiply - accumulate (MAC) operations and executing a current set of MAC operations according to another aspect of the present disclosure. According to method 370, the controller 140 instructs / causes the matrix multiplier engine 100 to execute the i - th set of MAC operations based on the i - th set of accumulators (e.g., accumulator - 1 or 314 - 1) using the i - th set of MAC instructions (block 372).
[0057] Method 370 further includes: the controller 140 determines whether there is a breakpoint or an instruction (e.g., a zero instruction) after the i - th set of storage instructions (block 374). For example, the breakpoint instruction may sequentially be between the i - th set of storage instructions and the (i + 1) - th set of MAC instructions. If, in block 376, the controller 140 determines that there is a breakpoint after the i - th set of storage instructions (meaning that subsequent MAC operations are independent of previous MAC operations), then the controller 140 instructs / causes the matrix multiplier engine 100 to execute the (i + 1) - th set of MAC operations based on the (i + 1) - th set of MAC instructions using the unused (i + 1) - th set of accumulators (e.g., accumulator - 2 or 314 - 2) (block 380). The newly selected (i + 1) - th set of accumulators is unused because it is not currently being used for MAC operations, for storage operations, or for holding values to be stored or transferred to the memory 150.
[0058] Moreover, concurrently with the matrix multiplier engine 100 executing the (i + 1) - th set of MAC operations according to block 380, the controller 140 instructs / causes the matrix multiplier engine 100 to store or transfer the result values in the i - th set of accumulators (e.g., accumulator - 1 or 314 - 1) to the memory 150 based on the i - th set of storage instructions (block 382). Method 370 further includes: the controller 140 instructs / causes the matrix multiplier engine 100 to store or transfer the result values in the (i + 1) - th set of accumulators to the memory 150 based on the (i + 1) - th set of storage instructions (block 384).
[0059] If, in block 376, the controller 140 determines that there is no demarcation point after the i-th set of memory instructions (e.g., in other words, after the i-th set of memory instructions is the (i + 1)-th set of MAC instructions without intermediate zero instructions), then the (i + 1)-th set of MAC instructions depends on the i-th set of MAC instructions. In such a case, the controller 140 instructs the matrix multiplier engine 100 to perform the associated MAC and memory operations (block 378).
[0060] Figure 4 FIG. shows a block diagram of an example multi-core integrated circuit (IC) 400 including a matrix multiplier engine 440 according to another aspect of the present disclosure. The IC 400 may be implemented as a system-on-chip (SOC) including a set of cores. The matrix multiplier engine 440 may be used in many different matrix multiplication applications.
[0061] Specifically, the multi-core IC 400 includes a set of processing cores 410-1 to 410-M, a memory 430, and a matrix multiplier engine 440, all coupled together via a data bus 420. The set of processing cores 410-1 to 410-M may include machine learning (ML) cores, image processing cores, face feature extraction cores, object detection cores, speech-to-text processing cores, and / or other or different sets of cores. These data processing cores 410-1 to 410-M may share the matrix multiplier engine 440 and the memory 430 for performing various matrix multiplication operations to facilitate their respective functions / operations.
[0062] Figures 5-6 FIG. shows a flowchart and a block diagram of an example method 500 and apparatus 600 for performing matrix multiplication. The method 500 includes: transferring a first set of result values from a first set of accumulators to a memory, where the first set of result values is generated from a first set of multiply-accumulate (MAC) operations (block 510). Examples of the unit 610 for transferring the first set of result values from the first set of accumulators to the memory include: the controller 140, the matrix multiplier engines 100 and 440, and the memories 150 and 430, where the first set of result values is generated from a first set of multiply-accumulate (MAC) operations.
[0063] The method 500 further includes: concurrently with transferring the first set of result values from the first set of accumulators to the memory, performing a second set of MAC operations using a second set of accumulators (block 520). Examples of the unit 620 for performing the second set of MAC operations using the second set of accumulators concurrently with transferring the first set of result values from the first set of accumulators to the memory include: the controller 140, the matrix multiplier engines 100 and 440, and the memories 150 and 430, the MAU arrays 210 and 310.
[0064] Some of the components described herein can be implemented using a processor. As used herein, a processor can be any dedicated circuit, processor-based hardware, a processing core of a system-on-chip (SOC), and so on. Examples of processor hardware can include: a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a state machine, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functions described throughout this disclosure.
[0065] The processor can be coupled to a memory (e.g., generally one or more computer-readable media), such as a magnetic storage device (e.g., a hard disk, a floppy disk, a magnetic stripe), an optical disk (e.g., a compact disc (CD) or a digital versatile disc (DVD)), a smart card, a flash memory device (e.g., a card, a stick, or a key drive), a random access memory (RAM), a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a register, a removable disk, and any other suitable medium for storing software and / or instructions that can be accessed and read by a computer. The memory can store computer-executable code (e.g., software). Whether referred to as software, firmware, middleware, microcode, a hardware description language, or otherwise, software should be broadly construed to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, processes / procedures, functions, and so on.
[0066] An overview of some aspects of the present disclosure is provided below:
[0067] Aspect 1: An apparatus, comprising: a memory; a matrix multiplier engine coupled to the memory, including: an array of multiplier-accumulator units (MAUs), including: a first set of accumulators; and a second set of accumulators; and a controller coupled to the matrix multiplier engine and the memory, the controller being configured to concurrently perform the following operations: causing a first set of result values in the first set of accumulators to be transferred to the memory according to a first set of storage instructions, wherein the first set of result values is generated according to a first set of multiply-accumulate (MAC) operations performed by the set of multipliers and the first set of accumulators; and causing the set of multipliers and the second set of accumulators to perform a second set of MAC operations.
[0068] Aspect 2: The apparatus according to aspect 1, wherein the first set of MAC operations is before the second set of MAC operations.
[0069] Aspect 3: The apparatus according to aspect 1 or 2, wherein the controller is configured to: in response to a demarcation instruction, select the second accumulator set for the second MAC operation set.
[0070] Aspect 4: The apparatus according to aspect 3, wherein the demarcation instruction includes an instruction for zeroing the first accumulator set.
[0071] Aspect 5: The apparatus according to aspect 3 or 4, wherein the MAU array includes: a demultiplexer set including a first input set coupled to the multiplier set, a first output set coupled to the first accumulator set, and a second output set coupled to the second accumulator set, wherein the controller is configured to: select the second accumulator set by respectively sending control signals to the selection input set of the demultiplexer set.
[0072] Aspect 6: The apparatus according to any one of aspects 3-5, wherein the demarcation instruction indicates that the second MAC operation set is independent of the first MAC operation set.
[0073] Aspect 7: The apparatus according to any one of aspects 3-6, wherein the controller is configured to: in response to a first set of MAC instructions, cause the multiplier set and the first accumulator set to perform the first set of MAC operations; and in response to a second set of MAC instructions, cause the multiplier set and the second accumulator set to perform the second set of MAC operations, wherein the demarcation instruction is sequentially located between the first set of storage instructions and the second set of MAC instructions.
[0074] Aspect 8: The apparatus according to aspect 7, wherein the controller is configured to: look ahead at the demarcation instruction to cause the transfer of the first set of result values to the memory and the second set of MAC operations to occur concurrently.
[0075] Aspect 9: The apparatus according to any one of aspects 1-8, wherein the second set of MAC operations respectively generate a second set of result values held in the second accumulator set.
[0076] Aspect 10: The apparatus according to aspect 9, wherein the second set of result values is generated before the first set of result values is completely transferred to the memory.
[0077] Aspect 11: The apparatus according to aspect 10, wherein the MAU array further includes a third accumulator set, and wherein the controller is configured to concurrently perform the following operations: continue to transfer the first set of result values to the memory; and cause the multiplier set and the third accumulator set to perform a third set of MAC operations.
[0078] Aspect 12: The apparatus according to aspect 11, wherein the controller is configured to: select the third accumulator set for the third set of MAC operations in response to a demarcation instruction.
[0079] Aspect 13: The apparatus according to aspect 12, wherein the demarcation instruction includes an instruction to zeroize the second accumulator set.
[0080] Aspect 14: The apparatus according to aspect 12 or 13, wherein the MAU array includes a demultiplexer set, which includes: a first input set coupled to the multiplier set, a first output set coupled to the first accumulator set, a second output set coupled to the second accumulator set, and a third output set coupled to the third accumulator set, and wherein the controller is configured to: select the third accumulator set by respectively sending control signals to the selection input set of the demultiplexer set.
[0081] Aspect 15: The apparatus according to any one of aspects 12 - 14, wherein the demarcation instruction indicates that the third set of MAC operations is independent of the second set of MAC operations.
[0082] Aspect 16: The apparatus according to any one of aspects 12 - 15, wherein the controller is configured to: in response to a second set of MAC instructions, cause the multiplier set and the second accumulator set to perform the second set of MAC operations; and in response to a third set of MAC instructions, cause the multiplier set and the third accumulator set to perform the third set of MAC operations, wherein the demarcation instruction is sequentially located between the second set of storage instructions and the third set of MAC instructions.
[0083] Aspect 17: The apparatus according to claim 11, wherein the controller is configured to concurrently perform the following operations: continue to cause the multiplier set and the third accumulator set to perform the third set of MAC operations; and cause the second set of result values to be transferred to the memory.
[0084] Aspect 18: The apparatus according to aspect 9, wherein the controller is configured to concurrently perform the following operations: cause the second set of result values to be transferred to the memory according to a second set of storage instructions; and cause the set of multipliers and the set of first accumulators to perform a third set of MAC operations.
[0085] Aspect 19: The apparatus according to aspect 18, wherein the second set of MAC operations is before the third set of MAC operations.
[0086] Aspect 20: The apparatus according to aspect 18 or 19, wherein the controller is configured to: in response to a demarcation instruction, select the set of first accumulators for the third set of MAC operations.
[0087] Aspect 21: The apparatus according to aspect 20, wherein the demarcation instruction includes an instruction for zeroing the set of second accumulators.
[0088] Aspect 22: The apparatus according to aspect 20 or 21, wherein the demarcation instruction indicates that the third set of MAC operations is independent of the second set of MAC operations.
[0089] Aspect 23: The apparatus according to any one of aspects 20-22, wherein the controller is configured to: in response to a third set of MAC instructions, cause the set of multipliers and the set of first accumulators to perform the third set of MAC operations, wherein the demarcation instruction is sequentially located between the second set of storage instructions and the third set of MAC instructions.
[0090] Aspect 24: The apparatus according to aspect 23, wherein the controller is configured to: look ahead at the demarcation instruction to cause the transfer of the second set of result values to the memory and the third set of MAC operations to occur concurrently.
[0091] Aspect 25: The apparatus according to any one of aspects 18-24, wherein the MAU array includes a set of demultiplexers, which includes: a first set of inputs coupled to the set of multipliers, a first set of outputs coupled to the set of first accumulators, and a second set of outputs coupled to the set of second accumulators, wherein the controller is configured to: select the set of first accumulators by sending control signals to the set of selection inputs of the demultiplexers respectively.
[0092] Aspect 26: A method for performing matrix multiplication, comprising: transferring a first set of result values from a first set of accumulators to a memory, wherein the first set of result values is generated from a first set of multiply-accumulate (MAC) operations; and concurrently with transferring the first set of result values from the first set of accumulators to the memory, performing a second set of MAC operations using a second set of accumulators.
[0093] Aspect 27: The method according to aspect 26, further comprising: selecting the second set of accumulators for performing the second set of MAC operations based on a demarcation instruction indicating that the second set of MAC operations is independent of the first set of MAC operations.
[0094] Aspect 28: The method according to aspect 26 or 27, wherein performing the second set of MAC operations respectively generates a second set of result values in the second set of accumulators.
[0095] Aspect 29: The method according to aspect 28, wherein the second set of result values is generated before completion of transferring the first set of result values to the memory, and the method further comprises: concurrently with continuing to transfer the first set of result values to the memory, performing a third set of MAC operations using a third set of accumulators.
[0096] Aspect 30: An apparatus, comprising: a unit for transferring a first set of result values from a first set of accumulators to a memory, wherein the first set of result values is generated from a first set of multiply-accumulate (MAC) operations; and a unit for concurrently with transferring the first set of result values from the first set of accumulators to the memory, performing a second set of MAC operations using a second set of accumulators.
[0097] The foregoing description of the disclosure is provided to enable a person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A device, which comprises: a memory; a matrix multiplier engine coupled to the memory, comprising: an array of multiplier-accumulator units (MAUs), which comprises: a set of multipliers; a first set of accumulators; and a second set of accumulators; and a controller coupled to the matrix multiplier engine and the memory, the controller being configured to concurrently perform the following operations: cause a first set of result values in the first set of accumulators to be transferred to the memory according to a first set of storage instructions, wherein the first set of result values is generated according to a first set of multiply-accumulate (MAC) operations performed by the set of multipliers and the first set of accumulators; and cause the set of multipliers and the second set of accumulators to perform a second set of MAC operations.
2. The device according to claim 1, wherein the first set of MAC operations is before the second set of MAC operations.
3. The device according to claim 1, wherein the controller is configured to: select the second set of accumulators for the second set of MAC operations in response to a demarcation instruction.
4. The device according to claim 3, wherein the demarcation instruction includes an instruction for zeroing the first set of accumulators.
5. The device according to claim 3, wherein the MAU array comprises: a set of demultiplexers, which comprises a first set of inputs coupled to the set of multipliers, a first set of outputs coupled to the first set of accumulators, and a second set of outputs coupled to the second set of accumulators, wherein the controller is configured to: select the second set of accumulators by respectively sending control signals to a set of selection inputs of the set of demultiplexers.
6. The device according to claim 3, wherein the demarcation instruction indicates that the second set of MAC operations is independent of the first set of MAC operations.
7. The device according to claim 3, wherein the controller is configured to: in response to a first set of MAC instructions, cause the set of multipliers and the first set of accumulators to perform the first set of MAC operations; and in response to a second set of MAC instructions, cause the set of multipliers and the second set of accumulators to perform the second set of MAC operations, wherein the demarcation instruction is sequentially located between the first set of storage instructions and the second set of MAC instructions.
8. The device according to claim 7, wherein the controller is configured to: look ahead at the demarcation instruction to cause the transfer of the first set of result values to the memory and the second set of MAC operations to be performed concurrently.
9. The device according to claim 1, wherein the second set of MAC operations respectively generates a second set of result values held in the second set of accumulators.
10. The device according to claim 9, wherein the second set of result values is generated before the first set of result values is completely transferred to the memory.
11. The device according to claim 10, wherein The MAU array further includes a third accumulator set, wherein the controller is configured to concurrently perform the following operations: Continue to transfer the first set of result values to the memory; and Cause the multiplier set and the third accumulator set to perform a third set of MAC operations.
12. The apparatus according to claim 11, wherein, The controller is configured to: in response to a demarcation instruction, select the third accumulator set for the third set of MAC operations.
13. The apparatus according to claim 12, wherein, The demarcation instruction includes an instruction for zeroing the second accumulator set.
14. The apparatus according to claim 12, wherein, The MAU array includes a demultiplexer set, the demultiplexer set including a first input set coupled to the multiplier set, a first output set coupled to the first accumulator set, a second output set coupled to the second accumulator set, and a third output set coupled to the third accumulator set, wherein the controller is configured to: select the third accumulator set by respectively sending control signals to the selection input set of the demultiplexer set.
15. The apparatus according to claim 12, wherein, The demarcation instruction indicates that the third set of MAC operations is independent of the second set of MAC operations.
16. The apparatus according to claim 12, wherein, The controller is configured to: In response to a second set of MAC instructions, cause the multiplier set and the second accumulator set to perform the second set of MAC operations; and In response to a third set of MAC instructions, cause the multiplier set and the third accumulator set to perform the third set of MAC operations, wherein the demarcation instruction is sequentially located between the second set of storage instructions and the third set of MAC instructions.
17. The apparatus according to claim 11, wherein, The controller is configured to concurrently perform the following operations: Continue to cause the multiplier set and the third accumulator set to perform the third set of MAC operations; and Cause the second set of result values to be transferred to the memory.
18. The apparatus according to claim 9, wherein, The controller is configured to concurrently perform the following operations: Cause the second set of result values to be transferred to the memory according to a second set of storage instructions; and Cause the multiplier set and the first accumulator set to perform a third set of MAC operations.
19. The apparatus according to claim 18, wherein, The second set of MAC operations is before the third set of MAC operations.
20. The apparatus according to claim 18, wherein, The controller is configured to: in response to a demarcation instruction, select the first accumulator set for the third set of MAC operations.
21. The apparatus according to claim 20, wherein, The demarcation instruction includes an instruction for zeroing the second accumulator set.
22. The apparatus according to claim 20, wherein, The demarcation instruction indicates that the third MAC operation set is independent of the second MAC operation set.
23. The apparatus according to claim 20, wherein, the controller is configured to: in response to a third set of MAC instructions, cause the multiplier set and the first accumulator set to perform the third MAC operation set, wherein the demarcation instruction is sequentially located between the second set of storage instructions and the third set of MAC instructions.
24. The apparatus according to claim 23, wherein, the controller is configured to: look ahead at the demarcation instruction to cause the transfer of the second set of result values to the memory and the third MAC operation set to occur concurrently.
25. The apparatus according to claim 18, wherein, the MAU array includes a set of demultiplexers, the set of demultiplexers including a first input set coupled to the multiplier set, a first output set coupled to the first accumulator set, and a second output set coupled to the second accumulator set, wherein the controller is configured to: select the first accumulator set by sending control signals to the selection input set of the set of demultiplexers, respectively.
26. A method for performing matrix multiplication, comprising: transferring a first set of result values from a first accumulator set to a memory, wherein the first set of result values is generated from a first multiplication-accumulation (MAC) operation set; and concurrently with transferring the first set of result values from the first accumulator set to the memory, performing a second MAC operation set using a second accumulator set.
27. The method according to claim 26, further comprising: selecting the second accumulator set for performing the second MAC operation set based on a demarcation instruction indicating that the second MAC operation set is independent of the first MAC operation set.
28. The method according to claim 26, wherein, performing the second MAC operation set generates a second set of result values in the second accumulator set, respectively.
29. The method according to claim 28, wherein, the second set of result values is generated before completing the transfer of the first set of result values to the memory, and the method further comprises: concurrently with continuing to transfer the first set of result values to the memory, performing a third MAC operation set using a third accumulator set.
30. An apparatus, which comprises: a unit for transferring a first set of result values from a first accumulator set to a memory, wherein the first set of result values is generated from a first multiplication-accumulation (MAC) operation set; and a unit for performing a second MAC operation set using a second accumulator set concurrently with transferring the first set of result values from the first accumulator set to the memory.