Multiply-accumulate circuit and in-memory processing device including the multiply-accumulate circuit.

By designing a MAC circuit that integrates memory and processor, efficient arithmetic operations were achieved, solving the problem of data communication limitations in artificial intelligence computing hardware and improving the data processing speed and computing performance of neural networks.

CN114756198BActive Publication Date: 2026-04-03SK HYNIX INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-07
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, the performance of artificial intelligence computing hardware is degraded due to the limited data communication between memory and processor, and the computational load increases exponentially when performing deep learning tasks, requiring efficient arithmetic operation solutions.

Method used

A multiply-accumulate (MAC) circuit was designed, which includes a MAC arithmetic unit and a data input circuit. It can selectively perform multiply-accumulate arithmetic operations of weighted data and vector data or element-wise multiplication arithmetic operations of weighted data and constant data. The arithmetic operations are performed directly in the semiconductor chip by integrating a processor and memory.

Benefits of technology

It improves the speed of neural network data processing, reduces the amount of data communication between memory and processor, and enhances the computing performance of artificial intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114756198B_ABST
    Figure CN114756198B_ABST
Patent Text Reader

Abstract

This application discloses a multiply-accumulate (MAC) circuit and an in-memory processing device including the same. The MAC circuit includes a MAC arithmetic unit and a data input circuit. The MAC arithmetic unit selectively performs MAC arithmetic operations on weighted data and vector data, or element-wise multiplication (EWM) arithmetic operations on the weighted data and constant data. The data input circuit provides the weighted data and the vector data to the MAC arithmetic unit when the MAC arithmetic unit performs the MAC arithmetic operation, and provides the weighted data and the constant data to the MAC arithmetic unit when the MAC arithmetic unit performs the EWM arithmetic operation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference of related applications

[0002] This application claims priority to Korean Patent Application No. 10-2021-0003632, filed on January 11, 2021, the entire contents of which are incorporated herein by reference. Technical Field

[0003] Various embodiments of this teaching relate to a multiply-accumulate (hereinafter referred to as "MAC") circuit and an in-memory processing (hereinafter referred to as "PIM") device including the MAC circuit, and more specifically, to a MAC circuit that performs element-wise multiplication (hereinafter referred to as "EWM") arithmetic operations between a matrix and a constant, and a PIM device including the MAC circuit. Background Technology

[0004] Recently, interest in artificial intelligence (AI) has been increasing not only in the information technology industry but also in the financial and healthcare sectors. Consequently, AI, or more precisely, the introduction of deep learning, has been considered and prototyped across various fields. Generally speaking, deep learning refers to techniques used to effectively learn deep neural networks (DNNs) or deep networks with increased layers compared to general neural networks for pattern recognition or reasoning.

[0005] One reason for the widespread attention is the improved performance of processors performing arithmetic operations. To improve the performance of artificial intelligence (AI), it may be necessary to increase the number of layers constituting the neural networks used to train the AI. This trend has continued in recent years, leading to an exponential increase in the computational demands of the hardware actually performing the calculations. Furthermore, if AI employs a general-purpose hardware system comprising separate memory and processors, its performance may be degraded due to limitations in the amount of data communication between memory and processor. To address this issue, PIM devices, where the processor and memory are integrated into a single semiconductor chip, have been used as neural network computing devices. Because PIM devices perform arithmetic operations directly internally, the data processing speed within the neural network can be improved. Summary of the Invention

[0006] According to one embodiment, a multiply-accumulate (MAC) circuit includes a MAC arithmetic unit and a data input circuit. The MAC arithmetic unit is configured to selectively perform MAC arithmetic operations on weighted data and vector data or element-wise multiplication (EWM) arithmetic operations on weighted data and constant data. The data input circuit is configured to provide weighted data and vector data to the MAC arithmetic unit when the MAC arithmetic unit performs MAC arithmetic operations, and is configured to provide weighted data and constant data to the MAC arithmetic unit when the MAC arithmetic unit performs EWM arithmetic operations. Attached Figure Description

[0007] Certain features of the disclosed technology are illustrated by way of various embodiments with reference to the accompanying drawings, wherein:

[0008] Figure 1 The configuration of a MAC circuit according to an embodiment of this teaching is shown;

[0009] Figure 2 It shows Figure 1 The configuration of the data selection circuit included in the MAC circuit shown;

[0010] Figure 3 It is shown Figure 2 The diagram shows the operation of the bit copying block included in the data selection circuit.

[0011] Figure 4 The illustration shows a MAC arithmetic operation performed by a MAC circuit, including matrix multiplication of a weight matrix and a vector matrix, according to an embodiment of this teaching.

[0012] Figure 5 EWM arithmetic operations, including the multiplication of a weight matrix and a constant matrix, performed by a MAC circuit according to an embodiment of this teaching are shown.

[0013] Figure 6 It is shown Figure 1 A block diagram of an example of a MAC arithmetic unit included in the MAC circuit shown;

[0014] Figure 7 It shows the result of Figure 1 The MAC circuit shown in the figure performs MAC arithmetic operations;

[0015] Figure 8 It shows the result of Figure 1 The MAC circuit shown in the figure performs EWM arithmetic operations;

[0016] Figure 9 It is shown Figure 1 A block diagram of another example of a MAC arithmetic unit included in a MAC circuit is shown below;

[0017] Figure 10It is shown Figure 1 A block diagram of yet another example of a MAC arithmetic unit included in a MAC circuit is shown below;

[0018] Figure 11 It is shown Figure 1 A block diagram of yet another example of a MAC arithmetic unit included in a MAC circuit is shown; and

[0019] Figure 12 This is a block diagram illustrating a PIM device according to an embodiment of the present teachings. Detailed Implementation

[0020] In the following description of the embodiments, it will be understood that the terms "first" and "second" are intended to identify elements, but are not used to limit a specific number or order of elements. Additionally, when an element is referred to as being "on," "above," "over," "below," or "under" another element, it is intended to indicate a relative positional relationship, but is not intended to limit the situation where an element directly contacts another element or where there is at least one intervening element between two elements. Therefore, terms such as "on," "above," "above," "below," "under," and "below" as used herein are for the purpose of describing particular embodiments only and are not intended to limit the scope of this disclosure. Furthermore, when an element is referred to as being "connected" or "coupled" to another element, the element may be directly electrically or mechanically connected or coupled to the other element, or may be indirectly electrically or mechanically connected or coupled to the other element using one or more additional elements between the two elements. Furthermore, when a parameter is referred to as "predetermined," it is intended to indicate that the value of the parameter is predetermined before it is used in a process or algorithm. The value of the parameter may be set at the start of the process or algorithm, or the value of the parameter may be set during the period in which the process or algorithm is executed. Logic "high" and logic "low" levels can be used to describe the logic levels of electrical signals. A signal with a logic "high" level can be distinguished from a signal with a logic "low" level. For example, when a signal with a first voltage corresponds to a signal with a logic "high" level, a signal with a second voltage can correspond to a signal with a logic "low" level. In one embodiment, a logic "high" level can be set to a voltage level higher than that of a logic "low" level. Moreover, the logic levels of signals can be set differently or oppositely depending on the embodiment. For example, a signal with a logic "high" level in one embodiment can be set to a logic "low" level in another embodiment.

[0021] Various embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. However, the embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the disclosure.

[0022] The various embodiments pertain to a MAC circuit and a PIM device that includes the MAC circuit.

[0023] Figure 1 A configuration of MAC circuit 10 according to an embodiment of this teaching is shown. Reference Figure 1 The MAC circuit 10 can perform MAC arithmetic operations and EWM arithmetic operations. MAC arithmetic operations can include matrix multiplication of a weight matrix and a vector matrix. EWM arithmetic operations can include matrix multiplication of a weight matrix and a constant matrix. Therefore, the MAC circuit 10 can receive weight data for the weight matrix and vector data for the vector matrix to perform MAC arithmetic operations, and the MAC circuit 10 can receive weight data for the weight matrix and a constant to perform EWM arithmetic operations.

[0024] To perform MAC arithmetic operations, MAC circuit 10 can receive weighted data DW<127:0> and vector data DV<127:0> from the storage area. MAC circuit 10 can receive a first latch control signal LATCH1, which controls the data input operation for MAC arithmetic operations. MAC circuit 10 can receive a second latch control signal LATCH2, which controls the data accumulation and output operations during MAC arithmetic operations. MAC circuit 10 can output MAC result data MAC_RST generated by the MAC arithmetic operation. MAC circuit 10 can receive a MAC result data output control signal MAC_RD_RST, which controls the output operation of the MAC result data MAC_RST.

[0025] To perform EWM arithmetic operations, MAC circuit 10 can receive weighted data DW<127:0> and constant data DC<15:0> from the storage area. MAC circuit 10 can receive a third latch control signal LATCH3, which controls the input operation of the constant data DC<15:0> used for EWM arithmetic operations. MAC circuit 10 can receive a fourth latch control signal LATCH4, which controls the data input operation used for EWM arithmetic operations. MAC circuit 10 can output EWM result data EWM_RST generated by the EWM arithmetic operations. MAC circuit 10 can receive an EWM result data output control signal EWM_RD_RST, which controls the output operation of the EWM result data EWM_RST.

[0026] Specifically, the MAC circuit 10 may include a data input circuit 100 and a MAC arithmetic unit 200. The data input circuit 100 may receive weighted data DW<127:0>, vector data DV<127:0>, and constant data DC<15:0>. The weighted data DW<127:0> may correspond to elements of a weight matrix. The vector data DV<127:0> may correspond to elements of a vector matrix. The constant data DC<15:0> may correspond to a constant with a specific value. In this embodiment, both the weighted data DW<127:0> and the vector data DV<127:0> may have a size of 128 bits, and the constant data DC<15:0> may have a size of 16 bits. However, this embodiment may only be an example of this disclosure. Therefore, the dimensions of the weighted data DW<127:0>, the vector data DV<127:0>, and the constant data DC<15:0> can vary depending on the dimensions of each element in the weighted matrix and the vector matrix, as well as the computational power of the MAC processor 200.

[0027] The data input circuit 100 may include a first input latch 110, a data selection circuit 120, a second input latch 130, and an OR gate 140. The first input latch 110 may receive weighted data DW<127:0>. The data selection circuit 120 may receive vector data DV<127:0>, constant data DC<15:0>, a flag signal FLAG, and a third latch control signal LATCH3. The OR gate 140 may receive the first latch control signal LATCH1 and the fourth latch control signal LATCH4. In one embodiment, when the MAC circuit 10 performs MAC arithmetic operations, the first latch control signal LATCH1 may have a logic "high" level, while when the MAC circuit 10 performs EWM arithmetic operations, the fourth latch control signal LATCH4 may have a logic "high" level.

[0028] The first input latch 110 can be synchronized with the output signal of the OR gate 140 to output the weighted data DW<127:0> to the MAC arithmetic unit 200. In one embodiment, the first input latch 110 can be implemented using a flip-flop. The data selection circuit 120 can receive and latch constant data DC<15:0> according to the activation of the third latch control signal LATCH3, and can selectively output vector data DV<127:0> or replicated constant data DC<127:0> according to the logic level of the flag signal FLAG. (See below for further details.) Figure 2The configuration of the data selection circuit 120 is described. The second input latch 130 can be synchronized with the output signal of the OR gate 140 to output vector data DV<127:0> or replicated constant data DC<127:0> selectively output from the data selection circuit 120 to the MAC arithmetic unit 200. The OR gate 140 can receive the first latch control signal LATCH1 and the fourth latch control signal LATCH4 to perform a logical OR operation on them. The OR gate 140 can transmit the signal generated by the logical OR operation to the clock terminals of the first input latch 110 and the second input latch 130.

[0029] Figure 2 It shows Figure 1 The configuration of the data selection circuit 120 included in the MAC circuit 10 shown, and Figure 3 It is shown Figure 2 This diagram illustrates the input and output data of the bit copy block 122 included in the data selection circuit 120. First, refer to... Figure 2 The data selection circuit 120 may include a third input latch 121, a bit copy block 122, and a data selection output circuit 123. The third input latch 121 may be synchronized with a third latch control signal LATCH3 to receive and output constant data DC<15:0>. The bit copy block 122 may copy the constant data DC<15:0> output from the third input latch 121 to generate and output replicated constant data DC<127:0> consisting of multiple copies of the constant data DC<15:0>. The bit copy block 122 may copy the constant data DC<15:0> such that the replicated constant data DC<127:0> has the same number of bits as the weighted data DW<127:0> or the vector data DV<127:0>. As a result, the bit copy block 122 can receive 16-bit constant data DC<15:0> to output 128-bit replicated constant data DC<127:0>.

[0030] like Figure 3As shown, it can be assumed that the constant data DC<15:0> input to the bit copy block 122 corresponds to the 16-bit binary stream "1000010100000001". The bit copy block 122 can expand the number of bits of the 16-bit constant data DC<15:0> to produce a 128-bit replicated constant data DC<127:0> with the same number of bits as the weight data DW<127:0>. In this case, the bit copy block 122 can repeatedly copy the 16-bit constant data DC<15:0> corresponding to the binary stream "1000010100000001" to array the repeatedly copied constant data DC<15:0>. In this embodiment, the replicated constant data DC<127:0> can be obtained by serially arraying the copied data of the 16-bit constant data DC<15:0> eight times. That is, the least significant bit (LSB) of one of the eight sets of constant data DC<15:0> that constitute the replicated constant data DC<127:0> can be positioned adjacent to the most significant bit (MSB) of the other set of constant data DC<15:0> in the two adjacent sets of constant data DC<15:0>.

[0031] Refer again Figure 2 The data selection output circuit 123 can receive copy-type constant data DC<127:0> from the bit copy block 122 through its first input terminal IN1. The data selection output circuit 123 can receive vector data DV<127:0> through its second input terminal IN2. The data selection output circuit 123 can output the copy-type constant data DC<127:0> input to the first input terminal IN1 or the vector data DV<127:0> input to the second input terminal IN2 through its output terminal OUT, depending on the logic level of the flag signal FLAG. For example, the flag signal FLAG with a logic "low" level can be transmitted to the data selection output circuit 123 to perform MAC arithmetic operations. In this case, the data selection output circuit 123 can output the vector data DV<127:0> input to the second input terminal IN2 through its output terminal OUT. Conversely, a flag signal with a logic "high" level (FLAG) can be sent to the data select output circuit 123 to perform EWM arithmetic operations. In this case, the data select output circuit 123 can output the duplicated constant data DC<127:0> input to the first input terminal IN1 through its output terminal OUT. In one embodiment, the data select output circuit 123 can be implemented using a 2-to-1 multiplexer.

[0032] Figure 4 The MAC arithmetic operation of the MAC circuit 10 according to an embodiment of this teaching is illustrated. Reference Figure 4 The MAC circuit 10 can perform MAC arithmetic operations, which produce a result matrix obtained by performing matrix multiplication of a weight matrix and a vector matrix. The weight matrix can have "M" rows and "N" columns. Each of the vector matrix and the result matrix can have "N" rows and one column. The number of rows "M" in the weight matrix can be set differently according to embodiments, and in the following description, it can be assumed that the number of rows "M" in the weight matrix is ​​512. Similarly, the number of columns "N" in the weight matrix can also be set differently according to embodiments, and in the following description, it can be assumed that the number of columns "N" in the weight matrix is ​​512. Therefore, the weight matrix can have 512 rows (i.e., rows 1 to 512, R1 to R512) and 512 columns (i.e., columns 1 to 512, C1 to C512), and can have 512×512 elements W1.1 to W1.512, ..., and W512.1 to W512.512. Similarly, the vector matrix can have 512 rows (i.e., rows 1 to 512, R1 to R512) and one column C1, and can have 512 elements V1.1 to V512.1.

[0033] In matrix multiplication of the weight matrix and the vector matrix, the amount of computation required to multiply the weight matrix and the vector matrix can be determined based on the size of each element (W1.1~W1.512, ..., W512.1~W512.512 and V1.1~V512.1) and the computational complexity of the MAC arithmetic operations that can be performed by the MAC arithmetic unit 200. Figure 1 The MAC processor 200 inputs the size of the weight data and the size of the vector data. For example, when each element in the weight matrix and vector matrix is ​​2 bytes in size and the MAC processor 200 has 8 multipliers, 128-bit weight data DW<127:0> and 128-bit vector data DV<127:0> can be input to the MAC processor 200 as in this embodiment. In this case, the 128-bit weight data DW<127:0> input to the MAC processor 200 can consist of 8 elements of the weight matrix, such as elements W1.1 to W1.8 located at the intersection of the first row R1 and the first to eighth columns C1 to C8 of the weight matrix. Similarly, the 128-bit vector data DV<127:0> input to the MAC processor 200 can consist of 8 elements of the vector matrix, such as elements V1.1 to W8.1 located at the intersection of the first to eighth rows R1 to R8 and the first column C1 of the vector matrix.

[0034] To obtain the MAC result data MAC_RST corresponding to the element MAC_RST1.1 located at the intersection of the first row Rl and the first column C1 of the result matrix through the MAC arithmetic operation of the MAC circuit 10, the MAC arithmetic operation must be performed iteratively 64 times. Because the MAC arithmetic operation is performed through matrix multiplication, in addition to multiplication, addition and accumulation can also be performed to obtain the MAC result data MAC_RST.

[0035] Figure 5 EWM arithmetic operations of a MAC circuit 10 according to an embodiment of this teaching are shown. Reference Figure 5 The MAC circuit 10 can perform EWM arithmetic operations, which produce a result matrix obtained by performing matrix multiplication on the weight matrix and a constant. In this embodiment, it can also be assumed that the weight matrix has 512 rows (i.e., rows 1 to 512 R1 to R512) and 512 columns (i.e., columns 1 to 512 C1 to C512). (See reference...) Figure 4 In the case where the weight matrix has 512×512 elements W1.1 to W1.512, ..., and W512.1 to W512.512, the constant can be a single data C. The result data (i.e., the result matrix) generated by the EWM arithmetic operation of the MAC circuit 10 can have the same size as the weight matrix. That is, the result matrix generated by the EWM arithmetic operation can also have 512 rows (i.e., rows 1 to 512 R1 to R512) and 512 columns (i.e., columns 1 to 512 C1 to C512). The elements of the result matrix can have values ​​obtained by multiplying the elements of the weight matrix by the constant data C. Thus, in order to perform the EWM arithmetic operation, it may be necessary to perform only multiplication calculations without any addition or accumulation calculations.

[0036] Figure 6 It is shown that... Figure 1 The block diagram showing the configuration of the MAC arithmetic unit 200 included in the MAC circuit 10 corresponds to the configuration of the MAC arithmetic unit 200A. (See reference) Figure 6 and Figure 7 The MAC arithmetic unit 200A may include a multiplication circuit 210, a data output selection circuit 220, an adder tree 230, an accumulator 240, and a data output circuit 250. (See reference...) Figure 4 , Figure 6 and Figure 7 As described, MAC arithmetic operations can be performed by executing multiplication, addition, and accumulation calculations. Therefore, multiplication circuit 210, adder tree 230, and accumulator 240 can all operate during MAC arithmetic operations. Conversely, as referenced... Figure 5As described, EWM arithmetic operations can be performed by running only multiplication calculations. Therefore, while multiplication circuit 210 operates during EWM arithmetic operations, adder tree 230 and accumulator 240 do not operate during EWM arithmetic operations. Consequently, data output selection circuit 220 includes multiple demultiplexers that send data to either MAC or EWM operations based on the FLAG signal.

[0037] Specifically, the multiplication circuit 210 may include multiple multipliers configured in parallel, for example, eight multipliers (i.e., first multipliers to eighth multipliers MUL0 to MUL7). Parallel configuration of multiple multipliers means that the multiple multipliers are configured such that the data input / output operations and multiplication calculations of the multiple multipliers are performed simultaneously and independently. The meaning of the term "parallel configuration" can be applied equivalently to all components disclosed in this application. Each of the first multipliers MUL0 to the eighth multiplier MUL7 can receive one of the first input data DA1_0 to DA1_7 (e.g., one of the elements W1.1 to W1.8 of the weight matrix) and one of the second input data DA2_0 to DA2_7 (e.g., one of the elements V1.1 to V8.1 of the vector matrix). Furthermore, the first multiplier MUL0 through the eighth multiplier MUL7 can perform multiplication calculations on the first input data DA1_0 to DA1_7 and the second input data DA2_0 to DA2_7 to output the first multiplication result data DM_0 to the eighth multiplication result data DM_7, respectively. For example, the first multiplier MUL0 can perform multiplication calculations on the first input data DA1_0 corresponding to the element W1.1 of the weight matrix and the second input data DA2_0 corresponding to the element V1.1 of the vector matrix to generate and output the first multiplication result data DM_0. In the same manner, the eighth multiplier MUL7 can perform multiplication calculations on the first input data DA1_7 corresponding to the element W1.8 of the weight matrix and the second input data DA2_7 corresponding to the element V8.1 of the vector matrix to generate and output the eighth multiplication result data DM_7.

[0038] The data output selection circuit 220 can output the first to eighth multiplication result data DM_0 to DM_7 generated by the multiplication circuit 210 via eight first output lines 261 or eight second output lines 262. The data output selection circuit 220 may include multiple demultiplexers arranged in parallel, such as first to eighth demultiplexers DEMUX0 to DEMUX7. Each of the first to eighth demultiplexers DEMUX0 to DEMUX7 can be implemented using a 1-to-2 demultiplexer having one input terminal and two output terminals. The number of demultiplexers constituting the data output selection circuit 220 can be equal to the number of multipliers included in the multiplication circuit 210. The input terminals of the first to eighth demultiplexers DEMUX0 to DEMUX7 can be coupled to the output terminals of the first to eighth demultiplexers MUL0 to MUL7, respectively. For example, the input terminal of the first demultiplexer DEMUX0 can be coupled to the output terminal of the first multiplier MUL0, and the input terminal of the second demultiplexer DEMUX1 can be coupled to the output terminal of the second multiplier MUL1. In the same manner, the input terminal of the eighth demultiplexer DEMUX7 can be coupled to the output terminal of the eighth multiplier MUL7.

[0039] In each of the demultiplexers DEMUX0 to DEMUX7, the selection of the output line for the multiplication result data output by it can be determined by a flag signal FLAG input to the data output selection circuit 220. For example, when the flag signal FLAG with a logic "low" level is input to the data output selection circuit 220, the demultiplexers DEMUX0 to DEMUX7 can output the first to eighth multiplication result data DM_0 to DM_7 from the multiplication circuit 210 through the first output line 261 of the demultiplexers DEMUX0 to DEMUX7. Conversely, when the flag signal FLAG with a logic "high" level is input to the data output selection circuit 220, the demultiplexers DEMUX0 to DEMUX7 can output the first to eighth multiplication result data DM_0 to DM_7 from the multiplication circuit 210 through the second output line 262 of the demultiplexers DEMUX0 to DEMUX7.

[0040] The first output line 261 of the demultiplexer DEMUX0_DEMUX7 can be coupled to the adder tree 230. Therefore, the multiplication result data DM_0 to DM_7 output from the demultiplexer DEMUX0_DEMUX7 via the first output line 261 can be transmitted to the adder tree 230. The second output line 262 of the demultiplexer DEMUX0 to DEMUX7 can be coupled to the data output circuit 250. The data output circuit 250 can output the multiplication result data DM_0 to DM_7 input via the second output line 262 of the demultiplexer DEMUX0 to DEMUX7 as the EWM result data EWM_RST in response to the EWM result data output control signal EWM_RD_RST.

[0041] Adder tree 230 may include multiple adders ADDER1, ADDER2, and ADDER3, which are arrayed in a hierarchical structure, such as a tree structure. In this embodiment, each of the multiple adders ADDER1, ADDER2, and ADDER3 constituting adder tree 230 may be implemented using a half adder. However, this embodiment including adder tree 230 implemented using half adders may be merely an example of this disclosure. That is, in some other embodiments, each of the multiple adders ADDER1, ADDER2, and ADDER3 constituting adder tree 230 may be implemented using a full adder. The highest level (i.e., first level ST1) of adder tree 230 may include four first adders ADDER1 arranged in parallel. The second level ST2 below the first level ST1 may include two second adders ADDER2 arranged in parallel. The third level ST3 corresponding to the lowest level of adder tree 230 may be located below the second level ST2 and may include a third adder ADDER3. When each of the multiple adders ADDER1, ADDER2, and ADDER3 is a half adder, the number of first adders ADDER1 can be half the number of multipliers MUL0 to MUL7, and the number of second adders ADDER2 can be half the number of first adders ADDER1. Additionally, the number of third adders ADDER3 can be half the number of second adders ADDER2.

[0042] The first and second input terminals of each first adder ADDER1 in the first stage ST1 can be coupled to the corresponding first output lines 261 of two demultiplexers DEMUX0 to DEMUX7 constituting the data output selection circuit 220. Therefore, each of the first adders ADDER1 can perform addition calculations on the output data (i.e., multiplication result data) of the two demultiplexers included in the data output selection circuit 220 to generate and output addition result data. Furthermore, each of the second adders ADDER2 in the second stage ST2 can perform addition calculations on the output data (i.e., addition result data) of the two first adders ADDER1 in the first stage ST1 to generate and output addition result data. Furthermore, the third adder ADDER3 in the third stage ST3 can perform addition calculations on the output data (i.e., addition result data) of the two second adders ADDER2 in the second stage ST2 to generate and output addition result data DMA.

[0043] Accumulator 240 may include an accumulator adder (ADDR_A) 241 and a latch circuit 242. Accumulator adder (ADDR_A) 241 performs addition on the addition result data DMA output from the third adder ADDER3 at the lowest level (i.e., the third level ST3) and the feedback data DF output from latch circuit 242 to generate and output the accumulated sum result data DMACC. In one embodiment, accumulator adder (ADDR_A) 241 may be implemented using a half adder. Latch circuit 242 receives the accumulated sum result data DMACC output from accumulator adder (ADDR_A) 241. Latch circuit 242 latches the accumulated sum result data DMACC to feed back the latched data of the accumulated sum result data DMACC corresponding to the feedback data DF to accumulator adder (ADDR_A) 241. In one embodiment, latch circuit 242 may include a flip-flop. When the matrix multiplication of a row in the weight matrix with a column in the vector matrix (i.e., MAC arithmetic operation) is completed (see reference) Figure 4 The summation result data DMACC latched by latch circuit 242 can be transmitted to data output circuit 250. Data output circuit 250 can output the summation result data DMACC generated by latch circuit 242 as MAC result data MAC_RST in response to the MAC result data output control signal MAC_RD_RST.

[0044] As described above, the MAC arithmetic unit 200A can perform both MAC arithmetic operations and EWM arithmetic operations. When the MAC arithmetic unit 200A performs MAC arithmetic operations, the output data of the demultiplexers DEMUX0 to DEMUX7 constituting the data output selection circuit 220 can be transmitted to the adder tree 230 through the first output line 261, and the addition result data DMA generated by the adder tree 230 can be transmitted to the accumulator 240. Therefore, multiplication, addition, and accumulation calculations for MAC arithmetic operations can be performed normally. When the MAC arithmetic unit 200A performs EWM arithmetic operations, the output data of the demultiplexers DEMUX0 to DEMUX7 constituting the data output selection circuit 220 can be output from the MAC arithmetic unit 200A through the second output line 262 and the data output circuit 250. Therefore, multiplication calculations for EWM arithmetic operations can be performed normally.

[0045] Figure 7 This illustrates when the MAC arithmetic unit 200A is used in the MAC circuit 10. Figure 1 The MAC circuit 10 shown demonstrates MAC arithmetic operations. Figure 7 In, with Figure 1 and Figure 6 The same reference numerals and reference symbols used can represent the same components. Therefore, references will be omitted below. Figure 1 and Figure 6 A detailed description of the same components. This will be combined via... Figure 4 This embodiment is described by performing MAC arithmetic operations on the matrix multiplication of the weight matrix and the vector matrix shown.

[0046] refer to Figure 7 The 128-bit weighted data DW<127:0> can be input to the first input latch 110 of the data input circuit 100. The 128-bit weighted data DW<127:0> can be composed of elements W1.1 to W1.8, which are located at the intersection of the first row R1 and the first to eighth columns C1 to C8 of the weight matrix. The 128-bit vector data DV<127:0> and the 16-bit constant data DC<15:0> can be input to the data selection circuit 120. The 128-bit vector data DV<127:0> can be composed of elements V1.1 to V8.1, which are located at the intersection of the first to eighth rows R1 to R8 and the column C1 of the vector matrix. Additionally, a flag signal FLAG with a logic "low (L)" level and a third latch control signal LATCH3 with a logic "low (L)" level can be input to the data selection circuit 120. (See reference...) Figure 2The data selection circuit 120 can output 128-bit vector data DV<127:0> based on a flag signal FLAG with a logic "low (L)" level. The 128-bit vector data DV<127:0> output from the data selection circuit 120 can be transmitted to the second input latch 130.

[0047] The OR gate 140 of the data input circuit 100 can receive a first latch control signal LATCH1 with a logic "high (H)" level and a fourth latch control signal LATCH4 with a logic "low (L)" level. The OR gate 140 can generate and output a signal with a logic "high (H)" level, which is transmitted to the first input latch 110 and the second input latch 130. The first input latch 110 can output 128-bit weighted data DW<127:0> to the MAC arithmetic unit 200A based on the signal with a logic "high (H)" level output from the OR gate 140, and the second input latch 130 can output 128-bit vector data DV<127:0> to the MAC arithmetic unit 200A based on the signal with a logic "high (H)" level output from the OR gate 140.

[0048] The 128-bit weighted data DW<127:0> input to the MAC arithmetic unit 200A can be divided into eight groups of element data, which are input to the corresponding multipliers of the first to eighth multipliers MUL0 to MUL7 in the multiplication circuit 210. That is, the first multiplier MUL0 can receive the first weighted data DW1.1<15:0> corresponding to the element W1.1 located at the intersection of the first row R1 and the first column C1 of the weight matrix. The first weighted data DW1.1<15:0> can be a binary stream consisting of the first to the sixteenth bits included in the 128-bit weighted data DW<127:0> output from the data input circuit 100. In addition, the second multiplier MUL1 can receive the second weighted data DW1.2<31:16> corresponding to the element W1.2 located at the intersection of the first row R1 and the second column C2 of the weight matrix. The second weight data DW1.2<31:16> can be a binary stream consisting of bits 17 through 32 of the 128-bit weight data DW<127:0> output from the data input circuit 100. Similarly, the eighth multiplier MUL7 can receive the eighth weight data DW1.8<127:112> corresponding to the element W1.8 located at the intersection of the first row R1 and the eighth column C8 of the weight matrix. The eighth weight data DW1.8<127:112> can be a binary stream consisting of bits 113 through 128 of the 128-bit weight data DW<127:0> output from the data input circuit 100.

[0049] The 128-bit vector data DV<127:0> input to the MAC arithmetic unit 200A can also be divided into eight groups of element data, which are input to the corresponding multipliers in the first to eighth multipliers MUL0 to MUL7 of the multiplication circuit 210. That is, the first multiplier MUL0 can receive the first vector data DV1.1<15:0> corresponding to the element V1.1 located at the intersection of the first row R1 and column C1 of the vector matrix. The first vector data DV1.1<15:0> can be a binary stream consisting of the first to sixteenth bits included in the 128-bit vector data DV<127:0> output from the data input circuit 100. In addition, the second multiplier MUL1 can receive the second vector data DV2.1<31:16> corresponding to the element V2.1 located at the intersection of the second row R2 and column C1 of the vector matrix. The second vector data DV2.1<31:16> can be a binary stream consisting of bits 17 through 32 of the 128-bit vector data DV<127:0> output from the data input circuit 100. Similarly, the eighth multiplier MUL7 can receive the eighth vector data DV8.1<127:112> corresponding to the element V8.1 located at the intersection of the eighth row R8 and column C1 of the vector matrix. The eighth vector data DV8.1<127:112> can be a binary stream consisting of bits 113 through 128 of the 128-bit vector data DV<127:0> output from the data input circuit 100.

[0050] The multipliers MUL0 to MUL7 in the multiplication circuit 210 can perform multiplication calculations of weighted data and vector data to generate and output first multiplication result data to eighth multiplication result data DWV1.1, DWV1.2, ..., and DWV1.8, respectively. The first multiplication result data to eighth multiplication result data DWV1.1, DWV1.2, ..., and DWV1.8 output from the corresponding multipliers MUL0 to MUL7 can be input to the first demultiplexers DEMUX0 to DEMUX7 in the data output selection circuit 220, respectively. The first demultiplexers DEMUX0 to DEMUX7 can respond to a flag signal FLAG with a logic "low (L)" level to output the first multiplication result data to eighth multiplication result data DWV1.1, DWV1.2, ..., and DWV1.8 through the first output line 261. The first multiplication result data to the eighth multiplication result data DWV1.1, DWV1.2, ... and DWV1.8 output from the first demultiplexer to the eighth demultiplexer DEMUX0 to DEMUX7 can be transmitted to the adder tree 230.

[0051] Adder tree 230 can perform addition calculations hierarchically to produce addition result data and output it to accumulator 240. Accumulator 240 can perform accumulation calculations in response to a second latch control signal LATCH2 with a logic "high" level. The data produced by the accumulation calculations of accumulator 240 can be latched in accumulator 240 to provide feedback data DF. Feedback data DF can also be transmitted to data output circuit 250. Because matrix multiplication for all elements included in the weight matrix and vector matrix is ​​not completed, the MAC result data output control signal MAC_RD_RST and the EWM result data output control signal EWM_RD_RST input to data output circuit 250 can have a logic "low (L)" level.

[0052] Figure 8 This illustrates when the MAC arithmetic unit 200A is used in the MAC circuit 10. Figure 1 The EWM arithmetic operation of the MAC circuit 10 shown. Figure 8 In, with Figure 1 and Figure 6 The same reference numerals and reference symbols used can represent the same components. Therefore, references will be omitted below. Figure 1 and Figure 6 A detailed description of the same components. This will be combined via... Figure 5 This embodiment is described by performing EWM arithmetic operations on matrix multiplication of the weight matrix and constants.

[0053] refer to Figure 8 The 128-bit weighted data DW<127:0> can be input to the first input latch 110 of the data input circuit 100. The 128-bit weighted data DW<127:0> can be composed of elements W1.1 to W1.8, which are located at the intersection of the first row R1 and the first to eighth columns C1 to C8 of the weight matrix. The 128-bit vector data DV<127:0> and the 16-bit constant data DC<15:0> can be input to the data selection circuit 120. The 128-bit vector data DV<127:0> can be composed of elements V1.1 to V8.1 located at the intersection of the first to eighth rows R1 to R8 and the column C1 of the vector matrix. In addition, a flag signal FLAG with a logic "high (H)" level and a third latch control signal LATCH3 with a logic "high (H)" level can be input to the data selection circuit 120. See reference Figure 2 and Figure 3As described, the third input latch 121 of the data selection circuit 120 can receive 16-bit constant data DC<15:0> and transmit it to the bit copy block 122 in response to a third latch control signal LATCH3 having a logic "high (H)" level. The bit copy block 122 can copy the 16-bit constant data DC<15:0> to generate 128-bit replicated constant data DC<127:0> and output it to the first input terminal IN1 of the data selection output circuit. Figure 2 (123). The data selection output circuit 123 can output 128-bit copy constant data DC<127:0> in response to a flag signal FLAG with a logic "high (H)" level. The 128-bit copy constant data DC<127:0> output from the data selection circuit 120 can be transmitted to the second input latch 130.

[0054] The OR gate 140 of the data input circuit 100 can receive a first latch control signal LATCH1 with a logic "low (L)" level and a fourth latch control signal LATCH4 with a logic "high (H)" level. The OR gate 140 can generate a signal with a logic "high (H)" level and output it to the first input latch 110 and the second input latch 130. The first input latch 110 can output 128-bit weighted data DW<127:0> to the MAC arithmetic unit 200A based on the signal with a logic "high (H)" level output from the OR gate 140, and the second input latch 130 can output 128-bit constant data DC<127:0> to the MAC arithmetic unit 200A based on the signal with a logic "high (H)" level output from the OR gate 140.

[0055] The 128-bit weighted data DW<127:0> input to the MAC processor 200A can be divided into eight groups of element data, which are respectively related to the reference... Figure 7 The same method is used to input the data into the corresponding multipliers of the first to eighth multipliers MUL0 to MUL7 of the multiplication circuit 210. The 128-bit replicated constant data DC<127:0> input to the MAC arithmetic unit 200A can be divided into eight groups of original constant data DC<15:0>, which are then input into the corresponding multipliers of the first to eighth multipliers MUL0 to MUL7 of the multiplication circuit 210.

[0056] The multipliers MUL0 to MUL7 in the multiplication circuit 210 can perform multiplication calculations of weighted data and constant data to generate and output first multiplication result data to eighth multiplication result data DWC1.1, DWC1.2, ..., and DWC1.8, respectively. The first multiplication result data to eighth multiplication result data DWC1.1, DWC1.2, ..., and DWC1.8 output from the corresponding multipliers MUL0 to MUL7 can be input to the first demultiplexer to the eighth demultiplexer DEMUX0 to DEMUX7 of the data output selection circuit 220, respectively. The first to eighth demultiplexers, DEMUX0 to DEMUX7, can respond to a flag signal FLAG with a logic "high (H)" level to output the first to eighth multiplication result data, DWC1.1, DWC1.2, ..., DWC1.8, via the second output line 262. The first to eighth multiplication result data, DWC1.1, DWC1.2, ..., DWC1.8, output from the first to eighth demultiplexers, DEMUX0 to DEMUX7, can be transmitted to the data output circuit 250.

[0057] In one embodiment, the data output circuit 250 may output the first multiplication result data to the eighth multiplication result data DWC1.1, DWC1.2, ..., DWC1.8 as the first EWM result data EWM_RST1 in response to the EWM result data output control signal EWM_RD_RST having a logic "high (H)" level. In this case, the first EWM result data EWM_RST1 may correspond to the data located at... Figure 5 The elements C·W1.1, C·W1.2, ..., and C·W1.8 at the intersection of the first row R1 and the first to eighth columns C1 to C8 of the result matrix shown. In another embodiment, since matrix multiplication for all elements and constants of the weight matrix has not been completed, the EWM result data output control signal EWM_RD_RST input to the data output circuit 250 can have a logic "low (L)" level. In this case, the data output circuit 250 can prevent the first EWM result data EWM_RST1 from being output from the data output circuit 250, and can keep the output of the first multiplication result data to the eighth multiplication result data DWC1.1, DWC1.2, ..., and DWC1.8 in a standby state.

[0058] Figure 9 It is shown that... Figure 1 The block diagram shows another example of the configuration of the MAC arithmetic unit 200 included in the MAC circuit 10, corresponding to the MAC arithmetic unit 200B. Figure 9 In, with Figure 6The same reference numerals and reference symbols used can represent the same components. Therefore, references will be omitted below. Figure 6 Detailed description of the same component. References Figure 9 ,and Figure 6 Compared to the MAC arithmetic unit 200A, the MAC arithmetic unit 200B may further include a post-processing circuit 310, which is coupled between the second output line 262 of the data output selection circuit 220 and the data output circuit 250. In one embodiment, the post-processing circuit 310 may include a normalizer 311.

[0059] Typically, when the first input data DA1 and the second input data DA2 input to the MAC arithmetic unit 200B are floating-point types represented by sign, exponent, and mantissa, the multiplication circuit 210 can apply normalization processing to the multiplication result data, which is used to shift the mantissa to the right or left and to increase or decrease the exponent based on the mantissa shift. However, when normalization processing is performed in the multiplication circuit 210, the efficiency of the layout area of ​​the MAC arithmetic unit 200B may be reduced. Therefore, normalization processing can be omitted in the multiplication circuit 210 and performed in the adder tree 230 or the accumulator 240. In this case, if performing EWM arithmetic operations causes the multiplication result data DM generated by the multiplication circuit 210 to be output through the second output line 262 of the data output selection circuit 220, normalization processing may not be applied to the multiplication result data DM. Therefore, according to this embodiment, the MAC arithmetic unit 200B can be designed to include a post-processing circuit 310, and even when performing EWM arithmetic operations, the normalizer 311 of the post-processing circuit 310 can apply normalization processing to the multiplication result data DM. The normalizer 311 can perform normalization processing on the multiplication result data DM to generate normalized multiplication result data DMN and output it to the data output circuit 250.

[0060] Figure 10 It shows the relationship with Figure 1 The block diagram shows another example of the configuration of the MAC arithmetic unit 200 included in the MAC circuit 10, corresponding to the MAC arithmetic unit 200C. Figure 10 In, with Figure 6 The same reference numerals and reference symbols used can represent the same components. (Reference) Figure 10 The MAC arithmetic unit 200C may include a multiplication circuit 210, a data output selection circuit 220, an accumulator circuit 420, an adder circuit 430, and a data output circuit 250. The multiplication circuit 210, the data output selection circuit 220, and the data output circuit 250 may have the same characteristics as the reference circuit. Figure 6The configuration described above is the same. The MAC arithmetic operation of the MAC arithmetic unit 200C can be implemented by performing multiple multiplication calculations and multiple accumulation calculations, and by performing multiple addition calculations on the accumulation result data.

[0061] Accumulation circuit 420 may include multiple accumulators arranged in parallel, such as first accumulators to eighth accumulators ACC0 to ACC7. First accumulators to eighth accumulators ACC0 to ACC7 may be respectively coupled to the first output lines 261 of first demultiplexers to eighth demultiplexers DEMUX0 to DEMUX7 in data output selection circuit 220. Therefore, first accumulator ACC0 can perform the accumulation calculation of first multiplication result data DM_0 output from first multiplier MUL0 and transmitted through the first output line 261 of first demultiplexer DEMUX0, and second accumulator ACC1 can perform the accumulation calculation of second multiplication result data DM_1 output from second multiplier MUL1 and transmitted through the first output line 261 of second demultiplexer DEMUX1. Similarly, each of the remaining accumulators (i.e., third accumulators to eighth accumulators ACC2 to ACC7) can perform the accumulation calculation. Each of first accumulators to eighth accumulators ACC0 to ACC7 may have a reference... Figure 6 The configuration of the accumulator 240 described is the same. In the following text, it will be combined with... Figure 4 The matrix multiplication of the weight matrix and the vector matrix shown is used to describe the MAC arithmetic operations of the MAC arithmetic unit 200C. Furthermore, it can be compared with the reference... Figure 8 The EWM arithmetic operations of the MAC arithmetic unit 200A described herein are performed in the same manner as the EWM arithmetic operations of the MAC arithmetic unit 200C. Therefore, the description of the EWM arithmetic operations of the MAC arithmetic unit 200C will be omitted below.

[0062] For reference Figure 7 As described, in the first MAC arithmetic operation, the first multiplier MUL0 of the MAC arithmetic unit 200C can receive weight data DW1.1 corresponding to the element W1.1 located at the intersection of the first row R1 and the first column C1 of the weight matrix (as...). Figure 10 The first input data DA1_0) and the vector data DV1.1 corresponding to the element V1.1 located at the intersection of the first row R1 and column C1 of the vector matrix (as... Figure 10 The second input data is DA2_0). The first multiplier MUL0 can perform multiplication of weight data DW1.1 and vector data DV1.1 to produce and output the first multiplication result data DWV1.1 of the first MAC arithmetic operation (as...). Figure 10The first multiplication result data DM_0). The first demultiplexer DEMUX0 of the data output selection circuit 220 can transmit the first multiplication result data DWV1.1 to the first accumulator ACC0 through the first output line 261 of the first demultiplexer DEMUX0 in response to a flag signal FLAG with a logic "low (L)" level. The first accumulator ACC0 can latch the first multiplication result data DWV1.1 of the first MAC arithmetic operation. In the same manner as described above, the remaining second to eighth accumulators ACC1, ..., and ACC7 can also latch the second to eighth multiplication result data of the first MAC arithmetic operation, respectively.

[0063] In the second MAC arithmetic operation, the first multiplier MUL0 of the MAC arithmetic unit 200C can receive weight data DW1.9 corresponding to the element W1.9 located at the intersection of the first row R1 and the ninth column C9 of the weight matrix (as...). Figure 10 The first input data DA1_0) and the vector data DV9.1 corresponding to the element V9.1 located at the intersection of the ninth row R9 and column C1 of the vector matrix (as... Figure 10 The second input data is DA2_0). The first multiplier MUL0 can perform multiplication of weight data DW1.9 and vector data DV9.1 to produce and output the first multiplication result data DWV1.9 of the second MAC arithmetic operation (as...). Figure 10 The first multiplication result data DM_0). The first demultiplexer DEMUX0 of the data output selection circuit 220 can transmit the first multiplication result data DWV1.9 to the first accumulator ACC0 via the first output line 261 of the first demultiplexer DEMUX0 in response to a flag signal FLAG with a logic "low (L)" level. The first accumulator ACC0 can add the first multiplication result data DWV1.9 to the first multiplication result data DWV1.1 to generate and latch the first accumulated data DMACC0. In the same manner as above, the remaining second to eighth accumulators ACC1, ..., and ACC7 can also latch the second to eighth multiplication result data of the second MAC arithmetic operation, respectively.

[0064] The third to the 64th MAC arithmetic operations can also be performed in the same manner as described above. The element MAC_RST1.1 located at the intersection of the first row R1 and column C1 of the result matrix (which is obtained as the result of matrix multiplication of elements W1.1 to W1.512 arrayed in the first row R1 of the weight matrix and elements V1.1 to V512.1 arrayed in the column C1 of the vector matrix through the first to the 64th MAC arithmetic operations) can be divided into 8 groups of data, and these 8 groups of data can be latched by the corresponding accumulators in the first to the eighth accumulators ACC0 to ACC7. For example, the first accumulator ACC0 of the accumulator circuit 420 can accumulate all the first multiplication result data generated during the first to the 64th MAC arithmetic operations to produce the first final multiplication result data DMACC0. Similarly, the second to eighth accumulators ACC1 to ACC7 of the accumulator circuit 420 can also accumulate all the second to eighth multiplication result data generated during the first MAC arithmetic operation to the 64th MAC arithmetic operation, respectively, to generate the second to eighth final multiplication result data DMACC1 to DMACC7.

[0065] The first to eighth final multiplication result data (DMACC0 to DMACC7) generated by the first to eighth accumulators (ACC0 to ACC7) of the accumulator circuit 420 can be transmitted to the adder circuit 430. The adder circuit 430 can add all the first to eighth final multiplication result data (DMACC0 to DMACC7) to generate and output the total accumulated result data (DMACCT). The total accumulated result data (DMACCT) output from the adder circuit 430 can correspond to the element MAC_RST1.1 located in the first row R1 and column C1 of the result matrix, which is obtained by matrix multiplication of the elements W1.1 to W1 arranged in the first row R1 of the weight matrix and the elements V1.1 to V512.1 arranged in the column C1 of the vector matrix. The data output circuit 250 can receive and output the total accumulated result data (DMACCT) as MAC result data (MAC_RST), which is transmitted to the external device of the MAC arithmetic unit 200C.

[0066] Figure 11 It is shown that... Figure 1 The block diagram shows another example of the configuration of the MAC arithmetic unit 200D included in the MAC circuit 10. Figure 11 In, with Figure 10 The same reference numerals and reference symbols used can represent the same components. (Reference) Figure 11 ,and Figure 10 Compared to the MAC arithmetic unit 200C shown, the MAC arithmetic unit 200D may further include a post-processing circuit 310 coupled between the second output line 262 of the data output selection circuit 220 and the data output circuit 250. In one embodiment, the post-processing circuit 310 may include a normalizer 311, as shown in the reference... Figure 9 As described. Therefore, the normalizer 311 of the post-processing circuit 310 can receive the multiplication result data DM output from the data output selection circuit 220 via the second output line 262. Furthermore, the normalizer 311 can perform normalization processing on the multiplication result data DM to generate normalized multiplication result data DMN and output it to the data output circuit 250. The data output circuit 250 can output the normalized multiplication result data DMN as EWM result data EWM_RST.

[0067] Figure 12 This is a block diagram illustrating a PIM device 500 according to an embodiment of the present teachings. Reference Figure 12 The PIM device 500 may include a command decoder 510, Figure 1 The diagram shows a MAC circuit 10, a first memory bank 521, a second memory bank 522, a global buffer 530, and a data input / output (I / O) circuit 540. The first memory bank 521 may include a first storage area that stores data related to... Figure 4 (or Figure 5 The weight data DW corresponds to the elements arranged in any row (e.g., the first row R1) of the weight matrix shown. The second memory bank 522 may include a second storage area that stores the weight data DW corresponding to the elements arranged in the matrix. Figure 4 The vector data DV corresponding to the elements V1.1 to V512.1 arranged in column C1 of the vector matrix shown. The global buffer 530 may include a third storage area that stores the vector data DV corresponding to the elements V1.1 to V512.1 arranged in the matrix. Figure 5 The constant data DC corresponding to the constant C shown.

[0068] MAC circuit 10 can receive weighted data DW and vector data DV from first memory bank 521 and second memory bank 522 to perform MAC arithmetic operations on the weighted data DW and vector data DV. In one embodiment, MAC circuit 10 can receive weighted data DW from first memory bank 521 via first memory bank data transmission line 551. Additionally, MAC circuit 10 can receive vector data DV from second memory bank 522 via second memory bank data transmission line 552. First memory bank data transmission line 551 provides a data transmission path between first memory bank 521 and MAC circuit 10. Second memory bank data transmission line 552 provides a data transmission path between second memory bank 522 and MAC circuit 10. First memory bank 521, MAC circuit 10, and second memory bank 522 can constitute a MAC unit. Although not shown in the figures, PIM device 500 may include multiple MAC units.

[0069] Alternatively, MAC circuit 10 can receive weighted data DW and constant data DC from first memory bank 521 and global buffer 530 to perform EWM arithmetic operations on weighted data DW and constant data DC. MAC circuit 10 can receive weighted data DW from first memory bank 521 via first memory bank data transmission line 551. Additionally, MAC circuit 10 can receive constant data DC from global buffer 530 via global data transmission line 553. Global data transmission line 553 can be used as a multi-purpose data transmission path in PIM device 500. In another embodiment, MAC circuit 10 can receive constant data DC directly from an external device (not shown) coupled to PIM device 500 via data I / O pin DQ of data I / O circuit 540. In this case, constant data DC input to data I / O circuit 540 can be transmitted to MAC circuit 10 via global data transmission line 553.

[0070] In the PIM device 500 according to this embodiment, the MAC circuit 10 constituting the MAC unit can correspond to the reference. Figures 1 to 11 The MAC circuit 10 is described above. Therefore, the MAC circuit 10 can selectively perform MAC arithmetic operations on weighted data DW and vector data DV, or EWM arithmetic operations on weighted data DW and constant data DC. The MAC arithmetic operations of the MAC circuit 10 can be controlled by various MAC control signals, which are generated by the command decoder 510 based on the MAC command MAC_CMD provided by an external device. The EWM arithmetic operations of the MAC circuit 10 can be controlled by various EWM control signals, which are generated by the command decoder 510 based on the EWM command EWM_CMD provided by an external device.

[0071] When the MAC command MAC_CMD is transmitted to the command decoder 510, the command decoder 510 can decode the MAC command MAC_CMD to generate a reference. Figure 7 The various MAC control signals mentioned above (e.g., MAC read control signal MAC_RD, first latch control signal LATCH1 with a logic "high (H)" level, second latch control signal LATCH2 with a logic "high (H)" level, third latch control signal LATCH3 with a logic "low (L)" level, fourth latch control signal LATCH4 with a logic "low (L)" level, flag signal FLAG with a logic "low (L)" level, and MAC result data output control signal MAC_RD_RST). The MAC read control signal MAC_RD can be transmitted to the first memory bank 521 and the second memory bank 522. The first memory bank 521 and the second memory bank 522 can output weighted data DW and vector data DV respectively in response to the MAC read control signal MAC_RD. If the MAC arithmetic operation of the MAC circuit 10 is completed, the MAC result data output control signal MAC_RD_RST, which controls the output operation of the MAC result data, can be transmitted from the command decoder 510 to the MAC circuit 10. The MAC arithmetic operations performed by the MAC circuit 10 based on the MAC control signal can be compared with the reference. Figure 7 The MAC arithmetic operations described are the same.

[0072] When the EWM command EWM_CMD is transmitted to the command decoder 510, the command decoder 510 can decode the EWM command EWM_CMD to generate a reference. Figure 8 The various EWM control signals mentioned (e.g., EWM read control signal EWM_RD, first latch control signal LATCH1 with logic "low (L)" level, second latch control signal LATCH2 with logic "low (L)" level, third latch control signal LATCH3 with logic "high (H)" level, fourth latch control signal LATCH4 with logic "high (H)" level, flag signal FLAG with logic "high (H)" level, and EWM result data output control signal EWM_RD_RST) are similar to... Figure 12As shown, the EWM read control signal EWM_RD can be transmitted to the first memory bank 521 and the global buffer 530. The first memory bank 521 and the global buffer 530 can output weighted data DW and constant data DC respectively in response to the EWM read control signal EWM_RD. If the EWM arithmetic operation of the MAC circuit 10 is completed, the EWM result data output control signal EWM_RD_RST, used to control the output operation of the EWM result data, can be transmitted from the command decoder 510 to the MAC circuit 10. The EWM arithmetic operation performed by the MAC circuit 10 based on the EWM control signal can be compared with a reference... Figure 8 The EWM arithmetic operations are the same.

[0073] For illustrative purposes, a limited number of possible embodiments of this teaching have been given above. Those skilled in the art will understand that various modifications, additions, and substitutions are possible. Although this patent document contains numerous details, these details should not be construed as limiting the scope of this teaching or the scope of the claims, but rather as descriptions of features particularly relevant to particular embodiments. Certain features described in this patent document as separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Moreover, although features may be described above as functioning in certain combinations and even as originally claimed, in some cases one or more features in the claimed combination may be separable from said combination, and the claimed combination may be for sub-combinations or variations thereof.

Claims

1. A multiply-accumulate MAC circuit, comprising: A MAC arithmetic unit is configured to selectively perform MAC arithmetic operations on weighted data and vector data or element-wise EWM arithmetic operations on said weighted data and constant data. as well as A data input circuit is configured to: provide the weight data and the vector data to the MAC arithmetic unit when the MAC arithmetic unit performs the MAC arithmetic operation; and is configured to: provide the weight data and the constant data to the MAC arithmetic unit when the MAC arithmetic unit performs the EWM arithmetic operation. The MAC arithmetic unit selectively performs the MAC arithmetic operation and the EWM arithmetic operation in response to the control signal provided to the MAC circuit. The control signals are jointly input to the MAC arithmetic unit and the data input circuit, and The MAC processor includes: A multiplication circuit comprising multiple multipliers arranged in parallel; A data output selection circuit is configured to receive the output data of the multiplication circuit and output the output data of the multiplication circuit through a first output line or a second output line of the data output selection circuit. An accumulator circuit, comprising multiple accumulators arranged in parallel; and An adder circuit configured to perform addition calculations on data output from the plurality of accumulators.

2. The MAC circuit according to claim 1, wherein, The first output line of the data output selection circuit is coupled to the accumulator circuit.

3. The MAC circuit according to claim 2, wherein, The MAC arithmetic unit also includes a data output circuit, which is directly coupled to the second output line of the data output selection circuit and the output terminal of the adder circuit.

4. The MAC circuit according to claim 1, in, The data output selection circuit includes multiple demultiplexers coupled to the respective multipliers in the plurality of multipliers; as well as The plurality of demultiplexers receive the output data of the plurality of multipliers and output the output data of the plurality of multipliers through the first output line or the second output line.

5. The MAC circuit according to claim 4, wherein, The plurality of demultiplexers are configured to transmit the output data of the plurality of multipliers to the accumulator circuit through the first output line when the MAC arithmetic unit performs the MAC arithmetic operation. And it is configured to: when the MAC arithmetic unit performs the EWM arithmetic operation, transmit the output data of the plurality of multipliers to the data output circuit through the second output line.

6. The MAC circuit according to claim 5, wherein, The data output circuit is configured to: output data received from the adder circuit in response to a MAC result data output control signal; and is configured to: output data received through the second output line in response to an EWM result data output control signal.

7. The MAC circuit according to claim 5, wherein, The MAC processor further includes a normalizer coupled between the second output line of the data output selection circuit and the data output circuit to perform normalization processing on the data transmitted through the second output line.

8. The MAC circuit according to claim 7, wherein, The data output circuit is configured to: output data received from the adder circuit in response to a MAC result data output control signal; and is configured to: output data received from the normalizer in response to an EWM result data output control signal.

9. The MAC circuit according to claim 5, wherein, Each of the plurality of accumulators includes: An accumulator adder configured to: perform an addition calculation to add feedback data to data output from one of the plurality of demultiplexers; and A latch circuit is configured to latch the output data of the accumulator to generate the feedback data, which is then transmitted to the accumulator.

10. The MAC circuit according to claim 1, wherein, The data input circuit includes: A first input latch is configured to latch the weight data and output the latched weight data to the MAC processor. A data selection circuit is configured to: receive the vector data and the constant data to selectively output either the vector data or the constant data; and The second input latch is configured to latch the output data of the data selection circuit so as to output the latched output data of the data selection circuit to the MAC arithmetic unit.

11. The MAC circuit according to claim 10, wherein, The data selection circuit includes: A third input latch is configured to latch and output the constant data; A bit copy block, configured to: copy constant data output from the third input latch to generate and output replicated constant data; and A data selection output circuit is configured to selectively output either the replicated constant data or the vector data.

12. The MAC circuit according to claim 11, wherein, The replicated constant data is generated with the same number of bits as the weight data.

13. The MAC circuit according to claim 11, wherein, The data input circuit further includes an OR gate, which performs a logical OR operation on the first latch control signal used for the MAC arithmetic operation and the fourth latch control signal used for the EWM arithmetic operation, so as to output the result signal of the logical OR operation to the clock terminal of the first input latch and the clock terminal of the second input latch.

14. The MAC circuit according to claim 13, in, The third input latch outputs the constant data in response to a third latch control signal having a first logic level; as well as Specifically, the output data of the data selection output circuit is selected based on the logic level of the flag signal.

Citation Information

Patent Citations

  • Method and system for providing a flexible and efficient processor for use in graphics processing

    US20040008201A1

  • Technologies for performing macro operations in memory

    US20190266219A1