Memory device

By designing memory units, buffer memory, input/output pads and MAC operators in memory devices, the problem that existing neural network engines are difficult to meet high bandwidth and low latency in convolution operations, and efficient convolution operations and low latency system performance are achieved.

CN111553472BActive Publication Date: 2025-07-01SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010043153.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-02-08
Filing Date
2020-01-15
Publication Date
2025-07-01
Estimated Expiration
2040-01-15

AI Technical Summary

Technical Problem

When existing neural network engines handle convolutional operations, it is difficult to meet the needs of high bandwidth and low latency, resulting in a decrease in system utilization.

Method used

A memory device is designed, including a memory unit, a buffer memory, an input/output pad and a MAC operator, to achieve efficient operations by performing convolution operations in the memory device and input and output data simultaneously in the MAC operator.

Benefits of technology

It realizes efficient execution of convolution operations in memory devices, meets the needs of low latency, and improves the utilization rate of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111553472B_ABST
    Figure CN111553472B_ABST
Patent Text Reader

Abstract

Provided is a memory device. The memory device includes: memory cells configured to store weight data; a buffer memory configured to read the weight data from the memory cells; input / output pads configured to receive input data; and a multiply-accumulate (MAC) arithmetic unit configured to: receive the weight data from the buffer memory and receive the input data from the input / output pads to perform a convolution operation of the weight data and the input data, wherein the input data is provided to the MAC arithmetic unit during a first period, and wherein the MAC arithmetic unit performs the convolution operation of the weight data and the input data during a second period overlapping with the first period.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority and all benefits arising therefrom to Korean Patent Application No. 10-2019-0014683, filed with the Korean Intellectual Property Office on February 8, 2019, the content of which is incorporated herein by reference in its entirety. Technical Field

[0002] The present disclosure relates to a memory device and a computing device using the memory device. Background Art

[0003] The basic algorithm of a neural network derives an output matrix through the operation of an input matrix and a convolution filter. Specifically, the output matrix can be determined through the convolution operation of the input matrix and the convolution filter.

[0004] The convolution operation includes a combination of multiplication operations and summation operations. With the recent explosive development of neural networks, a neural network engine with high bandwidth and low latency is required. Therefore, the size of the convolution filter increases, and the amount of weight data included in the convolution filter grows exponentially. Similarly, the amount of input data included in the input matrix also grows exponentially, so a very large number of multiplication operations and summation operations need to be performed to generate the output matrix.

[0005] In existing systems, it takes a long time to meet the increasing demand, which reduces the utilization rate of the system. Therefore, it is necessary to develop a neural network engine that maintains high bandwidth while meeting low latency. Summary of the Invention

[0006] Aspects of the present disclosure provide a memory device that performs a convolution operation in the memory device and performs simple and effective operations, and a computing device using the memory device.

[0007] Aspects of the present disclosure also provide a memory device that performs a convolution operation in the memory device and meets low latency, and a computing device using the memory device.

[0008] Aspects of the present disclosure also provide a memory device including a MAC arithmetic unit, and a computing device using the memory device, in which input data and output data are simultaneously input and output in the MAC arithmetic unit.

[0009] According to an embodiment of the present disclosure, a memory device is provided. The memory device includes: a memory cell configured to store weight data; a buffer memory configured to read the weight data from the memory cell; an input / output pad configured to receive input data; and a multiply-accumulate (MAC) arithmetic unit configured to: receive the weight data from the buffer memory and receive the input data from the input / output pad to perform a convolution operation of the weight data and the input data, wherein the input data is provided to the MAC arithmetic unit during a first period, and wherein the MAC arithmetic unit performs the convolution operation of the weight data and the input data during a second period overlapping with the first period.

[0010] According to the foregoing and other embodiments of the present disclosure, a memory device is provided. The memory device includes: a buffer memory configured to store weight data including a first weight bit and a second weight bit; an input / output pad configured to receive input data including a first input bit and a second input bit; and a MAC arithmetic unit including a first accumulator to a third accumulator and configured to: receive the weight data and the input data and perform a convolution operation of the weight data and the input data, wherein the step of performing the convolution operation of the weight data and the input data by the MAC arithmetic unit includes: calculating a first product of the first weight bit and the first input bit and providing the first product to the first accumulator; calculating a second product of the second weight bit and the first input bit and providing the second product to the second accumulator; calculating a third product of the first weight bit and the second input bit and providing the third product to the second accumulator; and calculating a fourth product of the second weight bit and the second input bit and providing the fourth product to the third accumulator.

[0011] According to the foregoing and other embodiments of the present disclosure, a memory device is provided. The memory device includes: a memory cell configured to store weight data; a buffer memory configured to read the weight data from the memory cell; an input / output pad configured to receive input data; and a MAC arithmetic unit configured to perform a convolution operation of the weight data and the input data, wherein the buffer memory reads the weight data from the memory cell before the input data is provided to the input / output pad, wherein the input data is provided to the MAC arithmetic unit from the input / output pad during a first period, and wherein the weight data is provided to the MAC arithmetic unit from the buffer memory during a second period overlapping with the first period.

[0012] According to the foregoing and other embodiments of the present disclosure, a computing device is provided. The computing device includes: a memory device including a MAC arithmetic unit and configured to store weight data; and a processor configured to provide input data to the memory device. Wherein, the MAC arithmetic unit receives the input data and the weight data and performs a convolution operation on the input data and the weight data, and wherein a first period in which the input data is provided to the MAC arithmetic unit overlaps with a second period in which the weight data is provided to the MAC arithmetic unit.

[0013] However, aspects of the present disclosure are not limited to those set forth herein. By referring to the specific embodiments of the present disclosure given below, the above and other aspects of the present disclosure will become clearer to those of ordinary skill in the art to which the present disclosure pertains. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The above and other aspects and features of the present disclosure will become clearer by referring to the exemplary embodiments of the present disclosure described in detail with reference to the accompanying drawings, wherein:

[0015] Figure 1 is an exemplary diagram illustrating a convolution operation.

[0016] Figure 2 is an exemplary block diagram showing a computing device according to some embodiments.

[0017] Figure 3 is an exemplary block diagram showing a non-volatile memory according to some embodiments.

[0018] Figure 4 is an exemplary diagram showing the operation of a computing device according to some embodiments.

[0019] Figure 5 is an exemplary diagram showing the operation of providing weight data from a memory cell to a buffer memory according to some embodiments.

[0020] Figure 6 is an exemplary diagram illustrating the operation of providing input data and weight data to a MAC arithmetic unit according to some embodiments.

[0021] Figure 7 is an exemplary diagram showing the operation of providing the operation result of a MAC arithmetic unit as output data according to some embodiments.

[0022] Figure 8 is an exemplary diagram showing the timing of data input / output according to some embodiments.

[0023] Figure 9 is an exemplary diagram showing the periods in which a MAC arithmetic unit receives input data and weight data according to some embodiments.

[0024] Figures 10 to 12 is an exemplary diagram showing the multiplication operation of input data and output data according to some embodiments.

[0025] Figure 13 is an exemplary block diagram showing a non - volatile memory according to some embodiments.

[0026] Figure 14 is an exemplary diagram showing the timing when data is input / output according to some embodiments. Detailed Description

[0027] Figure 1 is an exemplary diagram illustrating a convolution operation.

[0028] Referring to Figure 1 , an output matrix 3 can be generated by performing a convolution operation on an input matrix 1 and a convolution filter (or kernel) 2. For example, the input matrix 1 may include first input data X0 to twentieth input data X 19 . In addition, for example, the convolution filter 2 may include first weight data W0 to fourth weight data W3. In addition, for example, the output matrix 3 may include first output data S0 to twelfth output data S 11 . Some embodiments of the present disclosure are not limited to these terms, and those skilled in the art will clearly understand the following description. For the sake of brevity of description, although Figure 1 shows a case where the input matrix 1 is a 4×5 matrix, the convolution filter 2 is a 2×2 matrix, and the output matrix 3 is a 3×4 matrix, this is only exemplary and the embodiments are not limited thereto. The input matrix 1 and the convolution filter 2 may include more or fewer data, and the output matrix 3 can be determined according to the configurations of the input matrix 1 and the convolution filter 2.

[0029] The output matrix 3 can be determined by the multiplication operation and summation operation of input data I_Data and weight data W_Data. That is to say, the convolution operation can be a combination of multiplication operation and summation operation. For example, the first output data S0 and the second output data S1 can be determined by the following Equation 1 and Equation 2:

[0030] S0 = X0W0 + X1W1 + X5W2 + X6W3 …………………… Equation 1

[0031] S1 = X1W0 + X2W1 + X6W2 + X7W3 …………………… Equation 2

[0032] As represented in Equation 1, the first output data S0 can be determined by summing the product of the first input data X0 and the first weight data W0, the product of the second input data X1 and the second weight data W1, the product of the sixth input data X5 and the third weight data W2, and the product of the seventh input data X6 and the fourth weight data W3. In the same manner, as represented in Equation 2, the second output data S1 can be determined by summing the product of the second input data X1 and the first weight data W0, the product of the third input data X2 and the second weight data W1, the product of the seventh input data X6 and the third weight data W2, and the product of the eighth input data X7 and the fourth weight data W3.

[0033] Similarly, the third output data S2 to the twelfth output data S can be determined by performing multiplication operations and summation operations on the first input data X0 to the twentieth input data X 19 and the first weight data W0 to the fourth weight data W3. Hereinafter, the computing device 1000 that generates the output matrix 3 (i.e., the computing device 1000 that performs a convolution operation on the input data I_Data and the weight data W_Data) will be described. 11 is an exemplary block diagram showing a computing device according to some embodiments.

[0034] Figure 2 is an exemplary block diagram showing a computing device according to some embodiments.

[0035] Referring to Figure 2 According to some embodiments, the computing device 1000 may include an interface (I / F) 200, a processor 300, a cache memory 400, and a memory device 100.

[0036] The computing device 1000 according to some embodiments may include a personal computer (such as a desktop computer), a server computer, a portable computer (such as a laptop computer), and a portable device (such as a cellular phone, a smartphone, a tablet, an MP3, a portable multimedia player (PMP), a personal digital assistant (PDA), a digital camera, and a digital video camera). In addition, the computing device 1000 according to some embodiments may be a neural network-based processing device. For example, the computing device 1000 according to some embodiments may be used in a convolutional neural network (CNN)-based image processing device, an automatic steering device, a driving assistance device, etc. In addition, the computing device 1000 according to some embodiments may be used to perform digital signal processing (DSP). However, the technical concept of the present disclosure is not limited thereto, and the computing device 1000 according to some embodiments of the present disclosure may be used by those skilled in the art in various fields as needed.

[0037] The interface 200 can be used to input data into the computing device 1000 / output data from the computing device 1000. For example, referring toFigure 1 The described first input data X0 to the twentieth input data X 19 can be provided to the computing device 1000 through the interface 200, but the embodiments are not limited thereto. For example, the first input data X0 to the twentieth input data X 19 can be generated by specific components included in the computing device 1000.

[0038] The processor 300 can execute program code for controlling the computing device 1000. The processor 300 according to some embodiments may include a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), a microprocessor unit (MPU), etc., but the embodiments are not limited thereto.

[0039] The cache memory 400 can be a memory capable of temporarily storing data to prepare for future requests for accessing data at high speed. The data stored in the cache memory 400 can be the result of a previously executed operation. The cache memory 400 can be implemented as a static random access memory (SRAM), a fast static RAM (SRAM), and / or a dynamic RAM (DRAM), but the embodiments are not limited thereto. Although Figure 1 the cache memory 400 is shown separated from the processor 300, the embodiments are not limited thereto. For example, the cache memory 400 can be a tightly coupled memory (TCM) in the processor 300.

[0040] The memory device 100 may include a non-volatile memory 10 and a memory controller 20. The memory controller 20 can read or erase data stored in the non-volatile memory 10 or write data to the non-volatile memory 10 in response to a request from the processor 300. In addition, according to some embodiments, the memory controller 20 can receive a MAC command (MAC CMD) and control the non-volatile memory 10 to perform a convolution operation.

[0041] The non-volatile memory 10 can temporarily store data. For example, the non-volatile memory 10 can store the first weight data W0 to the fourth weight data W3. The non-volatile memory 10 according to some embodiments can perform a convolution operation in response to a request from the memory controller 20.

[0042] The non-volatile memory 10 may be a flash memory of a single-level cell (SLC) or a multi-level cell (MLC), but the embodiments are not limited thereto. For example, the non-volatile memory 10 may include a PC card (Personal Computer Memory Card International Association (PCMCIA)), a compact flash card (CF), a smart media card (SM, SMC), a memory stick, a multimedia card (MMC, RS-MMC, MMCmicro), an SD card (SD, Mini-SD, Micro-SD, SDHC), a universal flash storage (UFS), an embedded multimedia card (eMMC), a NAND flash memory, a NOR flash memory, and a vertical NAND flash memory.

[0043] Although not shown in the drawings, the memory controller 20 and / or the non-volatile memory 10 may be mounted using a package (such as a package-on-package (POP), a ball grid array (BGA), a chip-scale package (CSP), a plastic leaded chip carrier (PLCC), a plastic dual in-line package (PDIP), a waffle-packaged die, a wafer die, a chip-on-board (COB), a ceramic dual in-line package (CERDIP), a plastic metric quad flat package (MQFP), a thin quad flat package (TQFP), a small outline integrated circuit (SOIC), a shrink small outline package (SSOP), a thin small outline package (TSOP), a system-in-package (SIP), a multi-chip package (MCP), a wafer-level fabricating package (WFP), a wafer-level processing stacked package (WSP), etc.), but the embodiments are not limited thereto. A detailed description of the non-volatile memory 10 will be given with reference to Figure 3 A detailed description of the non-volatile memory 10 will be given.

[0044] Figure 3 is an exemplary block diagram showing a non-volatile memory according to some embodiments.

[0045] Referring to Figure 3 , the non-volatile memory 10 may include a storage area 10_S and a peripheral area 10_P. According to some embodiments, a plurality of memory cells 11 may be provided in the storage area 10_S. Each memory cell 11 may store data. For example, the memory cell 11 may store first weight data W0 to fourth weight data W3. For the sake of simplicity of description, the area other than the storage area 10_S in which the memory cells 11 are provided is defined as the peripheral area 10_P.

[0046] According to some embodiments, a buffer memory 12, a multiply-accumulate (MAC) arithmetic unit 13, a result output buffer 14, and an input / output (I / O) pad 15 may be provided in the peripheral area 10_P of the non-volatile memory 10.

[0047] The buffer memory 12 and the I / O pad 15 can provide data to the MAC arithmetic unit 13 respectively. For example, the buffer memory 12 can provide the weight data W_Data to the MAC arithmetic unit 13, and the I / O pad 15 can provide the input data I_Data to the MAC arithmetic unit 13.

[0048] The MAC arithmetic unit 13 can perform a convolution operation on the received weight data W_Data and input data I_Data. The MAC arithmetic unit 13 can provide the result of the convolution operation of the weight data W_Data and the input data I_Data to the result output buffer 14. For the sake of brevity of description, the data provided to the result output buffer 14 is defined as the result data R_Data. In some embodiments, the result data R_Data can be the intermediate result data of the convolution operation of the weight data W_Data and the input data I_Data. For example, the result data R_Data can be each of the first output data S0 to the twelfth output data S 11 among them. As another example, the result data R_Data can be the product W0X0 of the first weight data W0 and the first input data X0 to the product W3X 19 of the fourth weight data W3 and the twentieth input data X 19 among them. However, the embodiments are not limited thereto, and those skilled in the art can set the intermediate result of the convolution operation of the weight data W_Data and the input data I_Data as the result data R_Data.

[0049] The result output buffer 14 can store the result data R_Data. For example, each of the first output data S0 to the twelfth output data S 11 can be temporarily stored in the result output buffer 14. When all of the first output data S0 to the twelfth output data S 11 are stored in the result output buffer 14, the result output buffer 14 can provide the first output data S0 to the twelfth output data S 11 to the I / O pad 15.

[0050] The I / O pad 15 can receive the input data I_Data outside the non-volatile memory 10. The I / O pad 15 can provide the received input data I_Data to the MAC arithmetic unit 13. In addition, the I / O pad 15 can receive the data stored in the result output buffer 14 and provide it as the output data O_Data to the outside of the non-volatile memory 10. In some embodiments, the output data O_Data can be the intermediate result data or the final result data of the convolution operation of the weight data W_Data and the input data I_Data. For example, the output data O_Data can be the first output data S0 to the twelfth output data S 11。As another example, the output data O_Data may be the product W0X0 of the first weight data W0 and the first input data X0 to the product W3X 19 of the fourth weight data W3 and the twentieth input data X 19 among each of them.

[0051] Figure 4 is an exemplary diagram showing the operation of a computing device according to some embodiments.

[0052] Referring to Figure 4 , the processor 300 receives a request for a MAC operation. The processor 300 may provide a MAC command (MAC CMD) together with the input data I_Data to the memory controller 20.

[0053] In response to the received MAC command (MAC CMD), the memory controller 20 may provide a read command (Read CMD) for the weight data W_Data to the non-volatile memory 10. The non-volatile memory 10 may read the weight data W_Data stored in the storage area 10_S of the non-volatile memory 10 (e.g., stored in the memory cell 11) in response to the read command (Read CMD) for the weight data W_Data (S110). The read weight data W_Data may be provided to the buffer memory 12. An exemplary description will be given with reference to Figures 5 to 7 for an illustration.

[0054] Figure 5 is an exemplary diagram showing the operation of providing weight data from a memory cell to a buffer memory according to some embodiments. Figure 6 is an exemplary diagram illustrating the operation of providing input data and weight data to a MAC arithmetic unit according to some embodiments. Figure 7 is an exemplary diagram showing the operation in which the operation result of a MAC arithmetic unit is provided as output data according to some embodiments.

[0055] Referring to Figure 4 and Figure 5 , the weight data W_Data may be stored in at least a part of the plurality of memory cells 11. For example, the weight data W_Data may include the first weight data W0 to the fourth weight data W3. The memory controller 20 may provide a read command (Read CMD) for the weight data W_Data to the non-volatile memory 10 to provide the weight data W_Data stored in the memory cell 11 to the buffer memory 12. In other words, according to the command of the memory controller 20, the weight data W_Data may be latched from the memory cell 11 to the buffer memory 12.

[0056] That is, in response to a MAC command (MAC CMD), first, the memory controller 20 may control the weight data W_Data to be read from the memory cell 11 into the buffer memory 12. When the weight data W_Data has been read into the buffer memory 12, the non-volatile memory 10 may provide a read completion response to the memory controller 20.

[0057] Referring to Figure 4 and Figure 6 , the memory controller 20 may receive the read completion response. When the memory controller 20 receives the read completion response, the memory controller 20 may provide the input data I_Data to the non-volatile memory 10. For example, the memory controller 20 may provide the input data I_Data to the I / O pad 15.

[0058] The MAC arithmetic unit 13 may receive the input data I_Data via the I / O pad 15. For example, the MAC arithmetic unit 13 may receive the first input data X0 to the twentieth input data X 19 .

[0059] While the input data I_Data is provided to the MAC arithmetic unit 13, the weight data W_Data latched in the buffer memory 12 may also be provided to the MAC arithmetic unit 13. For example, while the MAC arithmetic unit 13 is provided with the first input data X0 to the twentieth input data X 19 , the first weight data W0 to the fourth weight data W3 latched in the buffer memory 12 may also be provided to the MAC arithmetic unit 13.

[0060] According to some embodiments, after the weight data W_Data is read from the memory cell 11 into the buffer memory 12, the MAC arithmetic unit 13 may receive the input data I_Data via the I / O pad 15. For example, before the MAC arithmetic unit 13 receives the first input data X0 to the twentieth input data X 19 , the first weight data W0 to the fourth weight data W3 stored in the memory cell 11 may be read into the buffer memory 12.

[0061] Referring to Figure 4 and Figure 7, the MAC arithmetic unit 13 can receive the input data I_Data and the weight data W_Data, and perform the convolution operation of the input data I_Data and the weight data W_Data (S120). The MAC arithmetic unit 13 can provide the result of the convolution operation of the input data I_Data and the weight data W_Data to the result output buffer 14. In other words, the result data R_Data generated in the MAC arithmetic unit 13 can be provided to the result output buffer 14. As described above, the result data R_Data can be the intermediate result data of the convolution operation of the weight data W_Data and the input data I_Data. For example, the result data R_Data can be each of the first output data S0 to the twelfth output data S 11 among them. As another example, the result data R_Data can be the product W0X0 of the first weight data W0 and the first input data X0 to the product W3X 19 of the fourth weight data W3 and the twentieth input data X 19 among them. According to some embodiments, the result data R_Data stored in the result output buffer 14 can be provided to the outside of the non-volatile memory 10 as the output data O_Data via the I / O pad 15.

[0062] Figure 8 is an exemplary diagram showing the timing of data input / output according to some embodiments.

[0063] will be described with reference to Figures 5 to 8 the timing of data input / output.

[0064] During the first period P1, the weight data W_Data can be latched into the buffer memory 12. That is, the weight data W_Data stored in the memory cell 11 of the non-volatile memory 10 can be provided to the buffer memory 12 during the first period P1. In other words, the buffer memory 12 can receive the weight data W_Data from the memory cell 11 and store the weight data W_Data during the first period P1.

[0065] During the second period P2, the buffer memory 12 can provide the latched weight data W_Data to the MAC arithmetic unit 13. In other words, the MAC arithmetic unit 13 can receive the weight data W_Data from the buffer memory 12 during the second period P2.

[0066] During the third period P3, the I / O pad 15 may provide input data I_Data to the MAC arithmetic unit 13. In other words, the MAC arithmetic unit 13 may receive the input data I_Data via the I / O pad 15 during the third period P3. According to some embodiments, the first period P1 may be earlier than the third period P3. In other words, the weight data W_Data may be read from the memory unit 11 into the buffer memory 12 before the MAC arithmetic unit 13 receives the input data I_Data.

[0067] According to some embodiments, the second period P2 and the third period P3 may overlap with each other. In other words, the MAC arithmetic unit 13 may receive the weight data W_Data and the input data I_Data simultaneously. In this specification, the term "simultaneously" does not refer to exactly the same time point. The term "simultaneously" means that two different events occur within the same period. In other words, the term "simultaneously" means that two events occur in parallel, rather than sequentially. For example, when the input data I_Data and the weight data W_Data are received within the same period, the input data I_Data and the weight data W_Data may be regarded as being received "simultaneously". As another example, when the MAC operation is performed during the period in which the input data I_Data is provided, the MAC operation may be regarded as being performed "simultaneously" when the input data I_Data is provided. Those skilled in the art can clearly understand the meaning of "simultaneously" used herein. Reference will be made to Figure 9 describe in more detail the periods during which the MAC arithmetic unit 13 receives the input data I_Data and the weight data W_Data.

[0068] Figure 9 is an exemplary diagram showing the periods during which a MAC arithmetic unit according to some embodiments receives input data and weight data. For the sake of brevity of description, repeated or similar descriptions will be given briefly or omitted.

[0069] Reference Figure 8 and Figure 9 , during the second period P2, the MAC arithmetic unit 13 may receive the weight data W_Data from the buffer memory 12. According to some embodiments, the second period P2 may include a first sub-period SP1 and a second sub-period SP2.

[0070] During the first sub-period SP1, the buffer memory 12 may provide the first weight data W0 to the MAC arithmetic unit 13. In other words, the MAC arithmetic unit 13 may receive the first weight data W0 from the buffer memory 12 during the first sub-period SP1.

[0071] During the second sub-period SP2, the buffer memory 12 may provide the second weight data W1 to the MAC arithmetic unit 13. In other words, the MAC arithmetic unit 13 may receive the second weight data W1 from the buffer memory 12 during the second sub-period SP2. According to some embodiments, the second sub-period SP2 may be arranged after the first sub-period SP1, but the embodiments are not limited thereto.

[0072] During the third period P3, the MAC arithmetic unit 13 may receive the input data I_Data via the I / O pad 15. According to some embodiments, the third period P3 may include a third sub-period SP3 and a fourth sub-period SP4.

[0073] During the third period P3, the I / O pad 15 may provide the first input data X0 to the MAC arithmetic unit 13. In other words, the MAC arithmetic unit 13 may receive the first input data X0 through the I / O pad 15 during the third sub-period SP3.

[0074] During the fourth sub-period SP4, the I / O pad 15 may provide the second input data X1 to the MAC arithmetic unit 13. In other words, the MAC arithmetic unit 13 may receive the second input data X1 through the I / O pad 15 during the fourth sub-period SP4. According to some embodiments, the fourth sub-period SP4 may be arranged after the third sub-period SP3, but the embodiments are not limited thereto.

[0075] According to some embodiments, the first sub-period SP1 and the third sub-period SP3 may overlap with each other. In addition, the second sub-period SP2 and the fourth sub-period SP4 may overlap with each other. In other words, according to some embodiments, the MAC arithmetic unit 13 may receive the first weight data W0 and the first input data X0 simultaneously. In addition, the MAC arithmetic unit 13 may receive the second weight data W1 and the second input data X1 simultaneously.

[0076] Referring back to Figures 5 to 8 , during the fourth period P4, the MAC arithmetic unit 13 may perform a convolution operation of the input data I_Data and the weight data W_Data. According to some embodiments, the fourth period P4 and the second period P2 may overlap with each other. In addition, according to some embodiments, the fourth period P4 and the third period P3 may overlap with each other. In other words, the MAC arithmetic unit 13 may perform a convolution operation of the input data I_Data and the weight data W_Data while receiving the input data I_Data and the weight data W_Data.

[0077] Although not shown, during the fourth period P4, the MAC arithmetic unit 13 may provide the intermediate result of the convolution operation of the input data I_Data and the weight data W_Data to the result output buffer 14. In other words, during the fourth period P4, the result output buffer 14 may be provided with the result data R_Data.

[0078] During the fifth period P5, the result output buffer 14 may provide the output data O_Data to the outside of the non-volatile memory 10 through the I / O pad 15. As described above, for example, the output data O_Data may be the first output data S0 to the twelfth output data S 11 , or the product W0X0 of the first weight data W0 and the first input data X0 to the product W3X 19 of the fourth weight data W3 and the twentieth input data X 19 .

[0079] According to some embodiments, from the time when the weight data W_Data is latched from the memory cell 11 to the buffer memory 12, the non-volatile memory 10 may remain in a busy state until the operation of the MAC arithmetic unit 13 terminates. In other words, when the internal operation of the non-volatile memory 10 is executed, the busy state signal RnBx may be at a logic low level (0).

[0080] According to some embodiments, the convolution operation of the input data I_Data and the weight data W_Data may be a combination of a multiplication operation and a summation operation. For example, referring to Equation 1 described above, the first output data S0 may be the sum of the product of the first input data X0 and the first weight data W0, the product of the second input data X1 and the second weight data W1, the product of the sixth input data X5 and the third weight data W2, and the product of the seventh input data X6 and the fourth weight data W3. The effective multiplication operation of the input data I_Data and the weight data W_Data will be described with reference to Figures 10 to 12 .

[0081] Figures 10 to 12 is an exemplary diagram showing the multiplication operation of the input data and the output data according to some embodiments. For the sake of brevity of description, Figures 10 to 12 shows the multiplication operation of the first input data X0 and the first weight data W0 as an example, but the embodiments are not limited thereto. In addition, for the sake of brevity of description, it is assumed that the first input data X0 is 3-bit data and the first weight data W0 is also 3-bit data, but the embodiments are not limited thereto. In Figures 10 to 12In this case, the first weight data W0 is defined as data whose most significant bit (MSB) is wb2, the second bit is wb1, and the least significant bit (LSB) is wb0. In addition, the first input data X0 is defined as data whose MSB is xb2, the second bit is xb1, and the LSB is xb0.

[0082] Referring to Figures 9 to 12 , the MAC arithmetic unit 13 may include a first multiplier M_1, a first accumulator AC_1, a second accumulator AC_2, a third accumulator AC_3, a fourth accumulator AC_4, and a fifth accumulator AC_5.

[0083] The MAC arithmetic unit 13 may receive the first weight data W0 during the first sub-period SP1 and may receive the first input data X0 during the third sub-period SP3. According to some embodiments, during the first sub-period SP1, all bits of the first weight data W0 may be provided and latched into the first multiplier M_1 simultaneously. In other words, the first weight data W0 may be the multiplicand of the first multiplier M_1. For example, during the first sub-period SP1, wb2, wb1, and wb0 may be provided and latched into the first multiplier M_1 simultaneously. On the other hand, during the third sub-period SP3, the first input data X0 may be sequentially provided to the first multiplier M_1. In other words, the first input data X0 may be the multiplier of the first multiplier M_1. For example, during the third sub-period SP3, xb2, xb1, and xb0 may be sequentially provided.

[0084] First, xb0 may be provided to the first multiplier M_1. At this time, the first multiplier M_1 may calculate wb0xb0, wb1xb0, and wb2xb0. The operations of wb0xb0, wb1xb0, and wb2xb0 may be performed in parallel in the first multiplier M_1. The first multiplier M_1 may provide wb0xb0 to the first accumulator AC_1, provide wb1xb0 to the second accumulator AC_2, and may provide wb2xb0 to the third accumulator AC_3.

[0085] Then, xb1 may be provided to the first multiplier M_1. At this time, the first multiplier M_1 may calculate wb0xb1, wb1xb1, and wb2xb1. The operations of wb0xb1, wb1xb1, and wb2xb1 may be performed in parallel in the first multiplier M_1. The first multiplier M_1 may provide wb0xb1 to the second accumulator AC_2, provide wb1xb1 to the third accumulator AC_3, and may provide wb2xb1 to the fourth accumulator AC_4.

[0086] Then, xb2 can be provided to the first multiplier M_1. At this time, the first multiplier M_1 can calculate wb0xb2, wb1xb2, and wb2xb2. The operations of wb0xb2, wb1xb2, and wb2xb2 can be executed in parallel in the first multiplier M_1. The first multiplier M_1 can provide wb0xb2 to the third accumulator AC_3, can provide wb1xb2 to the fourth accumulator AC_4, and can provide wb2xb2 to the fifth accumulator AC_5.

[0087] According to some embodiments, each of the outputs of the first accumulator AC_1 to the fifth accumulator AC_5 can be a bit corresponding to each digit of the product W0X0 of the first weight data W0 and the first input data X0. According to some embodiments, the output of the first accumulator AC_1 can be the least significant bit (LSB) of the product W0X0 of the first weight data W0 and the first input data X0, and the output of the fifth accumulator AC_5 can be the most significant bit (MSB) of the product W0X0 of the first weight data W0 and the first input data X0. The MAC operator 13 according to some embodiments can perform the multiplication operation of the weight data W_Data and the input data I_Data in a simple and efficient manner.

[0088] Although in Figures 10 to 12 the exemplary embodiment of, the first input data X0 is 3-bit data and the first weight data W0 is 3-bit data, the embodiments are not limited thereto. For example, the first input data X0 can be 2-bit data, 4-bit data, or more-bit data, and the first weight data W0 can be 2-bit data, 4-bit data, or more-bit data. In other words, the Figures 10 to 12 invention principle embodied in the regarding 3-bit input data and 3-bit weight data can also be applied to input data and weight data having various numbers of bits.

[0089] For example, in one embodiment, when the first input data X0 includes a first input bit and a second input bit, and the first weight data W0 includes a first weight bit and a second weight bit, the MAC operator can include a first multiplier, a first accumulator, a second accumulator, and a third accumulator. In this case, the first multiplier can perform the following operations: calculate a first product of the first weight bit and the first input bit and provide the first product to the first accumulator; calculate a second product of the second weight bit and the first input bit and provide the second product to the second accumulator; calculate a third product of the first weight bit and the second input bit and provide the third product to the second accumulator; and calculate a fourth product of the second weight bit and the second input bit and provide the fourth product to the third accumulator. The output of the first accumulator is the least significant bit (LSB) of the product of the weight data and the input data. The second accumulator outputs the sum of the second product and the third product.

[0090] AlthoughFigures 10 to 12 The first multiplier M_1 is shown as one component, but the embodiments are not limited thereto. Embodiments of the present disclosure can be implemented by those skilled in the art using multiple multipliers without undue experimentation.

[0091] Figure 13 is an exemplary block diagram showing a non-volatile memory according to some embodiments. For the sake of brevity of description, repeated or similar descriptions will be given briefly or omitted.

[0092] Referring to Figure 13 , in the non-volatile memory 10, memory cells 11 may be disposed in the storage area 10_S. In addition, a buffer memory 12, a MAC arithmetic unit 13, a result output pad 16, and an I / O pad 15 may be disposed in the peripheral area 10_P of the non-volatile memory 10. In other words, the non-volatile memory 10 according to some embodiments may be substantially the same as the non-volatile memory 10 described with reference to Figure 3 , except that the non-volatile memory 10 includes a result output pad 16 instead of a result output buffer 14.

[0093] The MAC arithmetic unit 13 may generate result data R_Data by performing a convolution operation on weight data W_Data and input data I_Data. The result data R_Data generated in the MAC arithmetic unit 13 may be provided to the result output pad 16. As described above, the result data R_Data may be intermediate result data of the convolution operation of the weight data W_Data and the input data I_Data. For example, the result data R_Data may be each of the first output data S0 to the twelfth output data S 11 . As another example, the result data R_Data may be the product W0X0 of the first weight data W0 and the first input data X0 to the product W3X 19 of the fourth weight data W3 and the twentieth input data X 19 .

[0094] The result data R_Data provided to the result output pad 16 may be provided to the outside of the non-volatile memory 10 as output data O_Data. According to some embodiments, the output data O_Data may be the same data as the result data R_Data.

[0095] According to some embodiments, the result output pad 16 may be separately configured from the I / O pad 15 to which the input data I_Data is provided. Thus, while the input data I_Data is provided to the MAC arithmetic unit 13 via the I / O pad 15, the output data O_Data may be provided to the outside of the non-volatile memory 10 through the result output pad 16. An exemplary description will be given with reference to Figure 14 .

[0096] Figure 14 is an exemplary diagram showing the timing of data input / output according to some embodiments. For the sake of brevity in description, repeated or similar descriptions will be given briefly or omitted.

[0097] Referring to Figure 13 and Figure 14 , during a first period P1, the buffer memory 12 may latch the weight data W_Data. That is, the weight data W_Data stored in the memory cells 11 of the non-volatile memory 10 may be provided to the buffer memory 12 during the first period P1.

[0098] During a second period P2, the buffer memory 12 may provide the latched weight data W_Data to the MAC arithmetic unit 13.

[0099] During a third period P3, the I / O pad 15 may provide the input data I_Data to the MAC arithmetic unit 13. According to some embodiments, the second period P2 and the third period P3 may overlap with each other. In other words, the MAC arithmetic unit 13 may receive the weight data W_Data and the input data I_Data simultaneously. According to some embodiments, the first period P1 may be earlier than the third period P3.

[0100] During a fourth period P4, the MAC arithmetic unit 13 may perform a convolution operation of the input data I_Data and the weight data W_Data. According to some embodiments, the fourth period P4 and the second period P2 may overlap with each other. In addition, according to some embodiments, the fourth period P4 and the third period P3 may overlap with each other. In other words, the MAC arithmetic unit 13 may perform the convolution operation of the input data I_Data and the weight data W_Data while receiving the input data I_Data and the weight data W_Data.

[0101] During a fifth period P5, the MAC arithmetic unit 13 may provide the result data R_Data to the result output pad 16. The result output pad 16 that has received the result data R_Data may provide the result data R_Data as output data O_Data to the outside of the non-volatile memory 10. According to some embodiments, the fifth period P5 may at least partially overlap with the second period P2. In addition, the fifth period P5 may at least partially overlap with the third period P3. In addition, the fifth period P5 may at least partially overlap with the fourth period P4. In other words, in at least a partial period, the MAC arithmetic unit 13 may simultaneously provide the output data O_Data to the outside of the non-volatile memory 10 through the result output pad 16 while receiving the input data I_Data through the I / O pad 15. For example, the output data O_Data may be the first output data S0 to the twelfth output data S 11Each of those in is either the product W0X0 of the first weight data W0 and the first input data X0 to the product W3X of the fourth weight data W3 and the twentieth input data X 19 of those in. 19 Each of those in.

[0102] In summary of the specific embodiments, those skilled in the art will understand that many variations and modifications can be made to the preferred embodiments without substantially departing from the principles of the present invention. Therefore, the preferred embodiments of the present invention disclosed are used only in a general and descriptive sense and not for limiting purposes.

Claims

1. A memory device, comprising: Memory cells configured to store weight data; A buffer memory configured to read the weight data from the memory cells; Input / output pads configured to receive input data; And A multiply-accumulate (MAC) arithmetic unit configured to: receive the weight data from the buffer memory and receive the input data from the input / output pads to perform a convolution operation of the weight data and the input data, wherein the input data is provided to the MAC arithmetic unit during a first time period, and wherein the MAC arithmetic unit performs the convolution operation of the weight data and the input data during a second time period overlapping with the first time period, wherein the weight data includes a first weight bit and a second weight bit, wherein the input data includes a first input bit and a second input bit, wherein the MAC arithmetic unit includes a first multiplier, a first accumulator, a second accumulator, and a third accumulator, wherein the process of performing the convolution operation by the MAC arithmetic unit includes: performing a multiplication operation of the weight data and the input data by the first multiplier, and wherein the process of performing the multiplication operation by the first multiplier includes: by the first multiplier: Calculating a first product of the first weight bit and the first input bit and providing the first product to the first accumulator, Calculating a second product of the second weight bit and the first input bit and providing the second product to the second accumulator, Calculating a third product of the first weight bit and the second input bit and providing the third product to the second accumulator, and Calculating a fourth product of the second weight bit and the second input bit and providing the fourth product to the third accumulator.

2. The memory device according to claim 1, wherein, The weight data is provided to the MAC arithmetic unit during a third time period overlapping with the first time period.

3. The memory device according to claim 1, wherein, Before the input data is provided to the MAC arithmetic unit, the buffer memory reads the weight data from the memory cells.

4. The memory device according to claim 1, wherein, The input data includes first input data and second input data, wherein the weight data includes first weight data and second weight data, wherein the first input data and the second input data are respectively provided to the MAC arithmetic unit during a first sub-time period and a second sub-time period, wherein the first weight data and the second weight data are respectively provided to the MAC arithmetic unit during a third sub-time period and a fourth sub-time period, and wherein the first sub-time period overlaps with the third sub-time period, and the second sub-time period overlaps with the fourth sub-time period.

5. The memory device according to claim 1, wherein, The output of the first accumulator is the least significant bit of the product of the weight data and the input data.

6. The memory device according to claim 1, wherein, The second accumulator outputs the sum of the second product and the third product.

7. The memory device according to claim 1, further comprising: A result output buffer configured to store the result of the convolution operation of the weight data and the input data.

8. The memory device according to claim 7, wherein, The result of the convolution operation stored in the result output buffer is output through the input / output pads.

9. The memory device according to claim 1, further comprising: A result output pad that outputs the result of the convolution operation of the weight data and the input data and is different from the input / output pads.

10. The memory device according to claim 9, wherein, The MAC arithmetic unit provides the result of the convolution operation to the result output pad during a fourth time period overlapping with the second time period.

11. A memory device, comprising: A buffer memory configured to store weight data including a first weight bit and a second weight bit; Input / output pads configured to receive input data including a first input bit and a second input bit; And A multiply-accumulate MAC arithmetic unit, including a first accumulator, a second accumulator, and a third accumulator, and configured to: receive weight data and input data, and perform a convolution operation of the weight data and the input data, Wherein, the process of performing the convolution operation of the weight data and the input data by the MAC arithmetic unit includes: Calculating a first product of a first weight bit and a first input bit, and providing the first product to the first accumulator, Calculating a second product of a second weight bit and the first input bit, and providing the second product to the second accumulator, Calculating a third product of the first weight bit and a second input bit, and providing the third product to the second accumulator, and Calculating a fourth product of the second weight bit and the second input bit, and providing the fourth product to the third accumulator.

12. The memory device according to claim 11, wherein, Performing the calculation of the first product and the calculation of the second product in parallel, and Wherein, the calculation of the third product and the calculation of the fourth product are performed in parallel.

13. The memory device according to claim 11, wherein, The process of performing the convolution operation by the MAC arithmetic unit includes: performing a multiplication operation of the weight data and the input data by the MAC arithmetic unit, and Wherein, the output of the first accumulator is the least significant bit of the product of the weight data and the input data.

14. The memory device according to claim 11, wherein, The input data is provided to the MAC arithmetic unit during a first time period, and Wherein, the MAC arithmetic unit performs the convolution operation during a second time period overlapping with the first time period.

15. The memory device according to claim 11, wherein, The input data is provided to the MAC arithmetic unit during a first time period, and Wherein, the weight data is provided to the MAC arithmetic unit during a third time period overlapping with the first time period.

16. The memory device according to claim 11, further comprising: A memory unit, configured to store weight data, Wherein, the weight data is read from the memory unit and stored in the buffer memory.

17. The memory device according to claim 16, wherein, Before the MAC arithmetic unit receives the input data, the weight data is read from the memory unit into the buffer memory.

18. The memory device according to claim 11, further comprising: A result output buffer, configured to store the convolution operation result of the weight data and the input data.

19. A memory device, including: A memory unit, configured to store weight data; A buffer memory, configured to read weight data from the memory unit; An input / output pad, configured to receive input data; And A multiply-accumulate MAC arithmetic unit, configured to perform a convolution operation of the weight data and the input data, Wherein, the buffer memory reads the weight data from the memory unit before the input data is provided to the input / output pad, Wherein, the input data is provided to the MAC arithmetic unit from the input / output pad during a first time period, and Wherein, the weight data is provided to the MAC arithmetic unit from the buffer memory during a second time period overlapping with the first time period, Wherein, the weight data includes a first weight bit and a second weight bit, Wherein, the input data includes a first input bit and a second input bit, Wherein, the MAC arithmetic unit includes a first multiplier, a first accumulator, a second accumulator, and a third accumulator, Wherein, the process of performing the convolution operation by the MAC arithmetic unit includes: performing a multiplication operation of the weight data and the input data by the first multiplier, and Wherein, the process of performing the multiplication operation by the first multiplier includes: by the first multiplier: Calculating a first product of a first weight bit and a first input bit, and providing the first product to the first accumulator, Calculate a second product of a second weight bit and a first input bit, and provide the second product to a second accumulator, calculate a third product of a first weight bit and a second input bit, and provide the third product to the second accumulator, and calculate a fourth product of the second weight bit and the second input bit, and provide the fourth product to a third accumulator.

Citation Information

Patent Citations

  • Painting method for vehicle parts and vehicle parts using the same

    KR1020190014683A

  • Apparatuses and methods for data movement

    US20170278559A1

  • Convolution operation circuit and object recognition apparatus

    WO2010064728A1