Memory device, memory system, and method for computing data using memory device
By introducing processing units into the peripheral circuit of the memory device, the memory bottleneck problem is solved, efficient computing of the AI system is achieved, and performance and speed are improved.
Patent Information
- Application Number
- CN202380013166.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-08-29
AI Technical Summary
Large converter models have high power consumption and limited computing performance problems caused by memory bottlenecks in AI computing, especially when the memory access speed lags behind the processor's computing speed, forming a memory wall, hindering the progress of high-performance computing.
The processing unit is introduced into the peripheral circuit of the memory device, and the calculation is performed under the control of the control logic unit, and the calculation task is distributed to the memory device to complete, reducing dependence on the processor.
By completing some computing tasks in the memory device, the computing speed and efficiency of the AI system are improved, the data transmission needs for the processor are reduced, and the overall performance is improved.
Smart Images

Figure CN120569779A_ABST
Abstract
Description
Background Art
[0001] The present disclosure relates to a memory device, a memory system, and a method for performing data calculation using the memory device.
[0002] Generative artificial intelligence (AI) reasoning involves AI computation. For example, transformer models typically use tensor processing units (TPUs) and memory for computation. Large transformer models require large amounts of data and computation, which necessitates high power consumption and sufficient memory. When memory access speed lags behind processor computation speed, memory bottlenecks prevent high-performance processors from operating efficiently and significantly constrain high-performance computing (HPC). This problem is known as the memory wall. Breaking down the memory wall is a promising approach to further improve the performance of AI systems. Summary of the Invention
[0003] In one aspect, a memory device is provided, comprising a group of memory cells and a peripheral circuit coupled to the group of memory cells. The peripheral circuit comprises a control logic unit configured to program first data and second data into different groups of the memory cells; and at least one processing unit coupled to the group of memory cells via a data path bus of the peripheral circuit and configured to perform a calculation based on the first data and the second data.
[0004] In some implementations, the first data includes at least one row. The control logic unit is configured to receive the first data from the data interface and program each row of the first data into one of the groups of memory cells based on a first data pattern.
[0005] In some embodiments, the first data pattern includes N first data segments having equal data lengths, where N is a positive integer and N≥2. The sequence of the N first data segments of the first data pattern is the same as the sequence of the first data.
[0006] In some implementations, a data length of each first data segment is less than or equal to a bandwidth of a data path bus.
[0007] In some embodiments, the second data includes M columns, where M is a positive integer and M≥2. The control logic unit is configured to program each of the M columns into M groups of memory cells in the group of memory cells based on the second data pattern, the number of the groups of memory cells being greater than M.
[0008] In some embodiments, the second data pattern includes N data groups, each of the N data groups having M second data segments with equal data lengths from M columns of the second data, and the first data segment and the second data segment are configured to share equal data lengths.
[0009] In some embodiments, each of the M second data segments of each data group of the second data is assigned an error checking and correcting (ECC) code.
[0010] In some implementations, a data length of each second data segment is less than or equal to a bandwidth of the data path bus.
[0011] In some embodiments, the control logic unit is configured to receive second data from a data interface of the memory device based on a second data pattern.
[0012] In some embodiments, each processing unit in at least one processing unit includes M processing elements, configured to perform a convolution operation based on the i-th first data segment in N first data segments and the M second data segments in the i-th data group in N data groups, where i is a positive integer and N≥i≥1.
[0013] In some embodiments, the control logic unit is configured to control one of the groups of memory cells to send the i-th first data segment of the first data to each of the M processing elements. The control logic unit is further configured to control the M groups of memory cells to send the M second data segments to the M processing elements.
[0014] In some embodiments, each processing unit of the at least one processing unit includes a control element configured to distribute the M second data segments to the M processing elements accordingly based on the sequence of the second data.
[0015] In some embodiments, the control logic unit is configured to output the calculation result to a data interface of a peripheral circuit of the memory device.
[0016] In some embodiments, the control logic unit is configured to output the calculation results to the group of memory units.
[0017] In some embodiments, the number of the at least one processing unit is equal to the number of the groups of memory units, and each processing unit corresponds to a corresponding group of the groups of memory units.
[0018] In some embodiments, the number of at least one processing unit is less than the number of groups of memory units.
[0019] In some embodiments, the number of at least one processing unit is half the number of the groups of memory cells, and one processing unit corresponds to two groups of memory cells, respectively.
[0020] In some embodiments, the number of the at least one processing unit is one fourth the number of the groups of memory cells, and one processing unit corresponds to four groups of memory cells, respectively.
[0021] In some embodiments, the number of the at least one processing unit is one, and one processing unit corresponds to the group of memory units.
[0022] In some implementations, the memory device includes dynamic random access memory (DRAM).
[0023] In another aspect, a method for performing data computation using a memory device is provided, the memory device including a group of memory cells and a peripheral circuit coupled to the group of memory cells. The method includes: obtaining, by a control logic unit of the peripheral circuit, first data and second data from a data interface of the memory device via a data path bus of the peripheral circuit; programming the first data and second data into the group of memory cells; and performing computation based on the first data and second data by at least one processing unit of the peripheral circuit.
[0024] In some embodiments, the first data includes at least one row.Programming the first data and the second data into the groups of memory cells includes programming each row of the first data into one of the groups of memory cells based on the first data pattern.
[0025] In some embodiments, the first data pattern includes N first data segments having equal data lengths, where N is a positive integer and N≥2. The sequence of the N first data segments of the first data pattern is the same as the sequence of the first data.
[0026] In some implementations, a data length of each first data segment is less than or equal to a bandwidth of a data path bus.
[0027] In some embodiments, the second data includes M columns, where M is a positive integer and M≥2, and obtaining the second data from the data interface of the memory device includes: programming each of the M columns into M groups of memory cells in a group of memory cells based on the second data pattern, the number of groups of memory cells being greater than M.
[0028] In some embodiments, the second data pattern includes N data groups, each of the N data groups having M second data segments with equal data lengths from M columns of the second data, and the first data segment and the second data segment are configured to share equal data lengths.
[0029] In some implementations, obtaining the second data from the data interface of the memory device includes assigning an error checking and correcting (ECC) code to each of the M second data segments of each data group of the second data.
[0030] In some implementations, a data length of each second data segment is less than or equal to a bandwidth of the data path bus.
[0031] In some embodiments, performing calculations based on the first data and the second data includes: performing a convolution operation based on the i-th first data segment among N first data segments and the M second data segments of the i-th data group among N data groups by each M processing elements of at least one processing unit, where i is a positive integer and N≥i≥1.
[0032] In some embodiments, performing a computation based on the first data and the second data includes: sending an i-th first data segment of the first data from a memory group to each of M processing elements; and sending M second data segments from the M groups of memory units to the M processing elements.
[0033] In some embodiments, the method further includes outputting the calculation result to a group of memory cells of a data interface of a peripheral circuit of the memory device.
[0034] In another aspect, a system is provided that includes a memory device and a controller. The memory device includes a group of memory cells and peripheral circuitry coupled to the memory cells. The peripheral circuitry includes a control logic unit configured to program first and second data into the memory group; and at least one processing unit coupled to the group of memory cells via a data path bus of the peripheral circuitry and configured to perform a calculation based on the first and second data. The controller is coupled to the memory device and configured to transfer the first data to the memory device and receive a result of the calculation from the memory device.
[0035] In some embodiments, the controller is further configured to transfer the second data to the memory device.
[0036] In some implementations, the memory device includes dynamic random access memory (DRAM). BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate aspects of the disclosure and, together with the description, serve to further explain the disclosure and enable a person skilled in the relevant art to make and use the disclosure.
[0038] Figure 1A A block diagram of a system having a memory device according to some aspects of the present disclosure is shown.
[0039] Figure 1B A diagram of a memory card having a memory device according to some aspects of the present disclosure is shown.
[0040] Figure 1C A diagram of a solid-state drive (SSD) having a memory device according to some aspects of the present disclosure is shown.
[0041] Figure 1D A schematic diagram of a memory device including peripheral circuits according to some aspects of the present disclosure is shown.
[0042] Figure 1E A block diagram of a memory device including a memory cell array and peripheral circuits according to some aspects of the present disclosure is shown.
[0043] Figure 1F A block diagram of a memory cell array and at least one processing unit according to some aspects of the present disclosure is shown.
[0044] Figure 1G A block diagram of at least one processing unit according to some aspects of the present disclosure is shown.
[0045] Figure 2A First data and second data processed by a memory device according to some aspects of the present disclosure are shown.
[0046] Figure 2B According to some aspects of the present disclosure Figure 2A The data shapes of the first data and the second data in .
[0047] Figure 2C According to some aspects of the present disclosure, Figure 2A The first data pattern and the second data pattern of the first data and the second data in the embodiment.
[0048] Figure 2D According to some aspects of the present disclosure, Figure 2C The second data mode in Figure 2A The storage mapping of the second data in.
[0049] Figure 3A Data flow in a processing unit of a memory device according to some aspects of the present disclosure is shown.
[0050] Figure 3B Data flow in a processing unit of a memory device according to some aspects of the present disclosure is shown.
[0051] Figure 4 According to some aspects of the present disclosure, a method for Figure 2C The first data mode and the second data mode are processed Figure 2A The process of first data and second data in.
[0052] Figure 5An operational pipeline of a memory device according to some aspects of the present disclosure is shown.
[0053] Figure 6 A flowchart of a method for performing data calculation using a memory device according to some aspects of the present disclosure is shown. DETAILED DESCRIPTION
[0054] Generally, terms can be understood, at least in part, from usage in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Similarly, depending at least in part on the context, terms such as "a" or "the" can also be understood to convey singular usage or to convey plural usage. Additionally, also depending at least in part on the context, the term "based on" can be understood as not necessarily intended to convey an exclusive set of factors, but rather can allow for the presence of additional factors that are not necessarily explicitly described.
[0055] Generative artificial intelligence (AI) reasoning involves AI computing. For example, transformer models, a common model in AI systems, typically use tensor processing units (TPUs) and memory for computation. Large transformer models require large amounts of data and computation, which necessitates high power consumption and sufficient memory. When memory access speed lags behind processor computation speed, memory bottlenecks prevent high-performance processors from operating efficiently and significantly restrict high-performance computing (HPC). This problem is known as the memory wall.
[0056] In order to solve one or more of the aforementioned problems and break the memory wall, the present disclosure introduces a solution, which provides a memory device and a method for performing calculations using the memory device. A processing unit is provided in the peripheral circuit of the memory device to perform calculations under the control of the control logic unit of the peripheral circuit. In this way, part of the computing tasks of the AI system can be distributed to the memory device of the AI system, especially tasks that require a large data width. Without transferring a large amount of data from the memory device to the processor of the AI system to perform calculations, the computing tasks are completed within the memory device, and the processor can handle other calculations. Therefore, the computing speed of the AI system is effectively improved by introducing the processing unit in the memory device. Figure 1AA block diagram of a system 10 having a host computer 20 and a memory system 30 according to some aspects of the present disclosure is shown. The system 10 may be a mobile phone, a desktop computer, a laptop computer, a tablet computer, a car computer, a game controller, a printer, a positioning device, a wearable electronic device, a smart sensor, a virtual reality (VR) device, an augmented reality (AR) device, an artificial intelligence (AI) device, or any other suitable electronic device having a storage device therein. Figure 1A As shown, the system 10 may include a host 20 and a memory system 30 having one or more non-volatile memory devices 34 (e.g., Figure 1A NAND flash memory in ), one or more volatile memory devices 36 (e.g., Figure 1A DRAM in), and a memory controller 32.
[0057] The memory system 30 can be configured to sense, read, program, and store data under the control of the host 20. The memory controller 32 can provide a physical connection between the host 20 and the memory system 30. That is, the memory controller 32 can provide a data interface between the host and the memory system 30 according to the format of the host's data bus. The memory controller 32 can decode instructions provided by the host 20 and access one or more non-volatile memory devices 34. One or more volatile memory devices 36 can be configured as a cache to temporarily store programming data provided by the host or data read from the non-volatile memory devices 34. When a read request is sent from the host 20, if the requested data in the non-volatile memory device 34 is cached in the volatile memory device 36, the volatile memory device 36 can directly send the cached data to the host 20. The data transmission speed between the volatile memory device 36 and the host 20 via the host 20's data bus is much higher than the data transmission speed between the non-volatile memory device 34 and the host 20. By introducing the volatile memory device 36, the performance degradation of the system 10 caused by the speed difference between the host 20 and the non-volatile memory device 34 can be minimized. In some embodiments, the volatile memory device 36 can also be configured to store a mapping table between the logical address and the physical address of the data stored in the non-volatile memory device 34. In some embodiments, the memory controller 32 can communicate with the volatile memory device 36 using at least one communication protocol or technical standard (e.g., associated with dual in-line memory modules (DIMMs), registered DIMMs (RDIMMs), load-reduced DIMMs (LRDIMMs), unregistered DIMMs (UDIMMs), etc.).
[0058] In some embodiments, the host 20 may be a processor (e.g., a tensor processing unit (TPU), a central processing unit (CPU)), or a system on chip (SoC) (e.g., an application processor (AP)) of an electronic device. The host 20 may be configured to send data to or receive data from the memory system 30. The non-volatile memory device 34 may include, but is not limited to, NAND flash memory, resistive random access memory (RRAM), nano random access memory (NRAM), phase change random access memory (PCRAM), ferroelectric random access memory (FRAM), magnetoresistive random access memory (MRAM), etc. The volatile memory device 36 may include, but is not limited to, dynamic random access memory (DRAM), static random access memory (SRAM), etc.
[0059] According to some embodiments, the memory controller 32 is coupled to the non-volatile memory device 34 and the host 20 and is configured to control the non-volatile memory device 34. The memory controller 32 can manage data stored in the non-volatile memory device 34 and communicate with the host 20. In some embodiments, the memory controller 32 is designed to operate in a low duty cycle environment, such as a secure digital (SD) card, a compact flash (CF) card, a universal serial bus (USB) flash drive, or other media used in electronic devices such as personal computers, digital cameras, mobile phones, etc. In some embodiments, the memory controller 32 is designed to operate in a high duty cycle environment, such as an SSD or an embedded multimedia card (eMMC), which is used as a data storage device for mobile devices such as smartphones, tablets, laptops, etc., as well as enterprise storage arrays. The memory controller 32 can be configured to control the operations of the non-volatile memory device 34 (e.g., read operations, erase operations, and program operations). The memory controller 32 may also be configured to manage various functions related to data stored or to be stored in the non-volatile memory device 34, including, but not limited to, bad block management, garbage collection, logical-to-physical address translation, wear leveling, and the like. In some embodiments, the memory controller 32 may also be configured to process error checking and correction (ECC) codes for data read from or written to the non-volatile memory device 34. The memory controller 32 may also perform any other appropriate functions, such as formatting the non-volatile memory device 34. The memory controller 32 may communicate with an external device (e.g., the host 20) according to a specific communication protocol. For example, the memory controller 32 may communicate with the external device using at least one of various interface protocols, such as a USB protocol, an MMC protocol, a Peripheral Component Interconnect (PCI) protocol, a PCI-Express (PCI-E) protocol, an Advanced Technology Attachment (ATA) protocol, a Serial ATA protocol, a Parallel ATA protocol, a Small Computer Small Interface (SCSI) protocol, an Enhanced Small Disk Interface (ESDI) protocol, an Integrated Drive Electronics (IDE) protocol, a FireWire protocol, and the like.
[0060] The memory controller 32 and the one or more non-volatile memory devices 34 can be integrated into various types of storage devices, for example, included in the same package (e.g., a Universal Flash Storage (UFS) package or an eMMC package). That is, the memory system 30 can be implemented and packaged into different types of terminal electronic products. Figure 1BIn one example shown, the memory controller 32 and the single volatile memory device 36 may be integrated into a memory card 40. The memory card 40 may include a PC card (PCMCIA, Personal Computer Memory Card International Association), a CF card, a Smart Media (SM) card, a memory stick, a multimedia card (MMC, RS-MMC, MMCmicro), an SD card (SD, miniSD, microSD, SDHC), UFS, etc. The memory card 40 may also include a computer that connects the memory card 40 to a host (e.g., Figure 1A The host computer 20) is coupled to the memory card connector 42. Figure 1C In another example shown, the memory controller 32 and the plurality of volatile memory devices 36 may be integrated into the SSD 50. The SSD 50 may also include a processor that interfaces the SSD 50 with a host (e.g., Figure 1A In some embodiments, the storage capacity and / or operating speed of the SSD 50 is greater than the storage capacity and / or operating speed of the memory card 40.
[0061] Figure 1D A schematic diagram of a memory device 60 is shown, including a memory cell array 62 and peripheral circuitry 64 coupled to the memory cell array 62. The memory cell array 62 may include groups 66 of memory cells. Each group 66 of memory cells may include a memory cell 622. Each memory cell 622 includes a transistor 624 and a storage element 626 coupled to the vertical transistor 624. In some embodiments, the memory cell array 62 is a DRAM cell array, and the storage element 626 is a capacitor for storing charge as binary information stored by the corresponding DRAM cell. In some embodiments, the memory cell array 62 is a PCM cell array, and the storage element 626 is a PCM element (e.g., comprising a chalcogenide alloy) for storing binary information for the corresponding PCM cell based on the different resistivity of the PCM element in the amorphous and crystalline phases. In some embodiments, the memory cell array 62 is a FRAM cell array, and the storage element 626 is a ferroelectric capacitor for storing binary information for the corresponding FRAM cell based on the switching of the ferroelectric material between two polarization states under an external electric field.
[0062] like Figure 1DAs shown, the memory cells 622 can be arranged in a two-dimensional (2D) array having rows and columns. The memory device 60 may include: word lines 627, which couple the peripheral circuit 64 with the memory cell array 62 for controlling the switching of the transistors 624 in the memory cells 622 located in a row, and bit lines 629, which couple the peripheral circuit 64 with the memory cell array 62 for sending data to the memory cells 622 located in a column and / or receiving data from the memory cells 622 located in a column. That is, each word line 627 is coupled to the memory cells 622 in a corresponding row, and each bit line 629 is coupled to the memory cells 622 in a corresponding column.
[0063] The storage element 626 may include any device capable of storing binary data (e.g., 0s and 1s), including, but not limited to, capacitors for DRAM cells and FRAM cells, and PCM elements for PCM cells. In some embodiments, the transistor 624 controls the selection and / or state switching of the corresponding storage element 626 coupled to the transistor 624. The peripheral circuit 64 may be coupled to the memory cell array 62 via the bit lines 629, the word lines 627, and any other suitable metal connections. As described above, the peripheral circuit 64 may include any suitable circuitry for applying voltage and / or current signals to each memory cell 622 via the word lines 627 and the bit lines 629, and sensing voltage and / or current signals from each memory cell 622 to facilitate the operation of the memory cell array 62. The peripheral circuit 64 may include any suitable analog, digital, or mixed-signal circuitry for facilitating the associated operation of the memory cell array by applying voltage and / or current signals to each target memory cell and sensing voltage and / or current signals from each target memory cell. In addition, the peripheral circuit 64 may include various types of peripheral circuitry formed using metal oxide semiconductor (MOS) technology.
[0064] refer to Figure 1E The peripheral circuit 64 includes a sense amplifier 71, a column decoder / bit line driver 72, a row decoder / word line driver 73, a voltage generator 74, a control logic unit 75, an address register 76, a data register 77, a data interface 79, a processing unit 80 and a data path bus 81. It should be understood that the peripheral circuit 70 can be connected to Figure 1D The peripheral circuit 64 is the same as that in FIG. 1 , and in some other examples, the peripheral circuit 70 may also include Figure 1E No additional peripheral circuits are shown.
[0065] The sense amplifier 71 may be configured to read data from the memory cell array 62 according to a control signal from the control logic unit 75. The column decoder / bit line driver 72 may be configured to be controlled by the control logic unit 75 and select one or more memory cells by applying a bit line voltage generated from the voltage generator 74.
[0066] The row decoder / word line driver 73 may be configured to be controlled by the control logic unit 75 and to select / deselect the group 66 of memory cells of the memory cell array 62 and to select / deselect the word lines of the group 66 of memory cells. The row decoder / word line driver 73 may also be configured to drive the word lines using a word line voltage generated from the voltage generator 74. As described in detail below, the row decoder / word line driver 73 is configured to apply a read voltage to the selected word line in a read operation on the memory cells coupled to the selected word line.
[0067] The voltage generator 74 may be configured to be controlled by the control logic unit 75 and generate word line voltages (eg, read voltage, program voltage, refresh voltage, etc.), bit line voltages, and source line voltages to supply to the memory cell array 62 .
[0068] The control logic unit 75 can be coupled to each peripheral circuit described above and is configured to control the operation of each peripheral circuit. The address register 76 and the data register 77 can be coupled to the control logic unit 75 and are configured to store status information, command operation code (OP code) and command address for controlling the operation of each peripheral circuit. The data interface 79 can be coupled to the control logic unit 75 via the data path bus 81 and act as a control buffer to buffer control commands received from a host (not shown) and relay them to the control logic unit 75, as well as buffer status information received from the control logic unit 75 and relay it to the host. The data interface 79 can also be coupled to the column decoder / bit line driver 72 and act as a data input / output (I / O) interface and a data buffer to buffer and relay data to and from the memory cell array 62.
[0069] like Figure 1F As shown, the memory cell array 62 includes a group 66 of memory cells, which are coupled to at least one processing unit 80 via a data register 77 and a data path bus 81 of the peripheral circuit 64. For example, in this embodiment, Figure 1F As shown, each memory cell array 62 may include sixteen memory cell groups 66. The plurality of memory cell groups 66 may be coupled to the peripheral circuit 64 via a data path bus 81. Before performing a calculation, first data and second data may be programmed into the plurality of memory cell groups 66.
[0070] In some embodiments, the number of at least one processing unit 80 is equal to the number of groups 66 of memory cells, and each processing unit 80 corresponds to a corresponding group of the groups of memory cells, for example, Figure 1F In some embodiments, the number of processing units 80 may be sixteen. In some embodiments, the number of at least one processing unit 80 is less than the number of groups 66 of memory cells. In some embodiments, the number of at least one processing unit 80 is half the number of groups of memory cells, and each processing unit 80 corresponds to two corresponding groups of memory cells. For example, in Figure 1F In some embodiments, the number of processing units 80 is one-fourth the number of groups of memory cells, and each processing unit 80 corresponds to four corresponding groups 66 of memory cells. For example, in Figure 1F In the embodiment, the number of the processing units 80 may be four. The number of the processing units 80 may be designed as needed in practice, and the above embodiments in the present disclosure are intended to be illustrative and should not be interpreted as limiting the present disclosure.
[0071] refer to Figure 1G , shows a processing unit 80, which includes a plurality of processing elements 82, a control element 84, a plurality of first registers 86, and a plurality of second registers 88. The processing unit 80 can be coupled to the memory cell array 62 through the first register 86 and the second register 88 to receive first data and second data from a plurality of groups of memory cells in the memory cell array 62. The plurality of processing elements 82 are coupled to the first register 86 and the second register 88, and are configured to perform convolution calculations based on the first data and the second data. In some embodiments, each processing element 82 can be provided with a corresponding result register, which is configured to store the calculation results of the processing element 82. The number of processing elements 82 is equal to the number of second registers 88 and the number of columns of the second data. In this embodiment, as Figure 1G As shown, each processing unit includes six processing elements 82 and six second registers 88. The control element 84 is configured to distribute the first data and the second data to the plurality of processing elements 82 according to a predetermined data pattern. In some embodiments, the number of at least one first register 86 is equal to the number of rows of the first data. For example, in this embodiment, each processing unit 80 includes one first register 86. In some embodiments, the first register 86 and the second register 88 are first-in, first-out (FIFO) registers.
[0072] AI systems are mainly used in two aspects: training and reasoning, and the present disclosure can be mainly used for AI reasoning, where data is input into a trained AI module for recognition and analysis to obtain the expected results of the input data. In AI reasoning, calculations are performed based on the input data and data pre-stored in the AI system to confirm one or more properties of the input data. In AI reasoning, such as Figure 2A As shown, in many cases, the input data may be one-dimensional data and the reference data may be two-dimensional data, wherein the first data is a one-dimensional vector and the second data is a two-dimensional matrix. In the AI system, three modules are provided to perform data calculations. The first module is calculating near the memory device, wherein the calculation is performed outside the memory device. The second module is calculating in the memory unit, wherein the calculation is performed by the memory unit of the memory device. The third module is processing in the memory unit, wherein the calculation is performed by an additional processing unit of the memory device. The third module (i.e., processing in the memory unit) is adopted in the embodiment of the present disclosure.
[0073] Figure 2B Shown Figure 2A In some embodiments, the one-dimensional first data may be equivalent to a row of data having a length of a, and the two-dimensional second data (i.e., an a×b matrix) may be equivalent to b columns, each having a length of a. In some embodiments, the first data may be a two-dimensional matrix including more than one row of equal data length, and dimensionality reduction may be performed on the more than one row of first data to decompose the first data into a plurality of single rows, thereby applying the present disclosure.
[0074] In some embodiments, the first data and the second data can be obtained from the data interface 79 of the memory device and programmed into the plurality of groups of memory cells in the memory cell array 62. For example, the first data can be stored in one group 66 of the plurality of groups 66 of memory cells in the memory cell array 62, and the second data can be stored in the other group 66 of memory cells in the plurality of groups 66 of memory cells in the memory cell array 62. The first data can be updated after each calculation. The second data can be stored in the group 66 of memory cells for multiple calculations using different first data and can be updated according to instructions from the host 20. In some embodiments, the first data can be updated based on the first data. Figure 2C The first data pattern and the second data pattern shown in FIG. 8 program the first data and the second data into the group of memory cells.
[0075] In some embodiments, reference Figure 1A, the second data may be stored in the nonvolatile memory device 34. The memory controller 32 may read the second data from the nonvolatile memory device 34 and program the second data to the volatile memory device 36 before performing the calculation. The second data may be stored in the nonvolatile memory device 34 according to the second data pattern or programmed into the volatile memory device 36 (i.e., a group of memory cells) according to the second data pattern.
[0076] In some embodiments, the first data comprises a row, and the control logic unit 75 is configured to program the row of first data into one of the plurality of groups of memory cells based on the first data pattern. Figure 1F As shown, the first data can be programmed into group 1 of multiple groups of memory cells. The first data pattern includes N first data segments with equal data length, where N is a positive integer and N≥2. In this embodiment, N=4 is used as an example to illustrate the present disclosure. The first data includes four first data segments based on the first data pattern, namely, the first data segment S1-0, the first data segment S1-1, the first data segment S1-2 and the first data segment S1-3. The sequence of the four first data segments of the first data pattern is the same as the sequence of the first data. In some embodiments, the data length of each first data segment is less than or equal to the bandwidth of the data path bus. In some embodiments, an error checking and correction (ECC) code is assigned to each first data segment for verification. The ECC code can also be used as an identifier for each first data segment, which is configured to identify that the first data pattern is applied to the first data.
[0077] In some embodiments, the second data includes M columns, where M is a positive integer and M ≥ 2. The control logic unit 75 is configured to program each of the M columns into M groups of memory cells in the plurality of groups of memory cells based on the second data pattern, the number of the plurality of groups of memory cells being greater than M. In some embodiments, as Figure 1F As shown, the number of groups of memory cells of the memory cell array 62 is sixteen, and the number of columns of the second data is six (as shown in FIG. Figure 2B ). The six columns of second data can then be programmed into six groups of memory cells in the sixteen memory groups. For example, column 0 can be programmed into group 2, column 1 can be programmed into group 3, column 2 can be programmed into group 4, column 3 can be programmed into group 5, column 4 can be programmed into group 6, and column 5 can be programmed into group 7.
[0078] The second data pattern includes N data groups, each of the N data groups has M second data segments with equal data lengths respectively from M columns of the second data, and the first data segment and the second data segment are configured to share equal data lengths. In some embodiments, N=4 and M=6 are used as examples to illustrate the present disclosure. Figure 2C As shown, the second data includes six columns; each of the six columns includes four second data segments, namely, the second data segment S2-0, the second data segment S2-1, the second data segment S2-2 and the second data segment S2-3. Figure 2C , the six second data segments S2-0 are regrouped into a first data group of a second data pattern, the six second data segments S2-1 are regrouped into a second data group of the second data pattern, the six second data segments S2-2 are regrouped into a third data group of the second data pattern, and the six second data segments S2-3 are regrouped into a fourth data group of the second data pattern. Each second data segment of a data group is programmed into a different memory group. For example, the first data group includes the second data segment S2-0 from column 0 programmed into group 2, the second data segment S2-0 from column 1 programmed into group 3, the second data segment S2-0 from column 3 programmed into group 5, the second data segment S2-0 from column 4 programmed into group 6, and the second data segment S2-0 from column 5 programmed into group 7. In some embodiments, an error checking and correction (ECC) code is assigned to each second data segment for verification. The ECC code can also be used as an identifier for each second data segment, which is configured to identify that the second data pattern is applied to the second data.
[0079] In some embodiments, the control logic unit 75 is further configured to Figure 2C The first data is sent from the plurality of groups of memory units to at least one processing unit 80. For example, Figure 3A As shown, the first data segment S1-0 is first sent to the processing unit 80, and as shown Figure 3B As shown, the first data segment S1-1 is sent to the processing unit 80 continuously after the first data segment S1-0. The first data segment S1-2 and the first data segment S1-3 are sent to the processing unit 80 continuously after the first data segment S1-1 (not shown). Taking the first data segment S1-0 as an example, in some embodiments, the first data segment S1-0 is sent to the first register 86 and is cached in the first register 86. In some embodiments, as shown Figure 3A As shown, each processing unit 80 includes a control element 84 configured to distribute the first data segment S1 - 1 to each processing element 82 .
[0080] In some embodiments, the control logic unit 75 is further configured to Figure 2C The data sequence of the second data pattern in the memory unit sends N data groups of the second data from the plurality of groups of the memory unit to at least one processing unit 80. For example, Figure 3A As shown, the six second data segments S2-0 of the first data group are first sensed and sent to the processing unit 80, and as shown in FIG. Figure 3B As shown, six second data segments S2-1 of the second data group are continuously sensed after the first data group and sent to the processing unit 80. Six second data segments S2-2 of the third data group of the second data are continuously sensed after the second data group and sent to the processing unit 80, and six second data segments S2-3 of the fourth data group of the second data are continuously sensed after the third group (not shown) and sent to the processing unit 80. Taking the six second data segments S2-0 as an example, in some embodiments, the six second data segments S2-0 are sent to the six second registers 88 and cached in the six second registers 88. In some embodiments, as Figure 3A As shown, the control element 84 is configured to distribute the six second data segments S2.1 to the six processing elements accordingly based on the sequence of the six second data segments S2.1.
[0081] In some embodiments, the M processing elements 82 of each processing unit 80 are configured to perform a convolution operation based on the i-th first data segment of the N first data segments and the M second data segments of the i-th data group of the N data groups, where i is a positive integer and N≥i≥1. Figure 3A and Figure 3B In this embodiment, the first data segment S1-0 is sent to each of the six processing elements 82, and the six second data segments S2-0 of the first data group of the second data are correspondingly sent to the six processing elements 82. The six processing elements 82 then perform a convolution operation on the first data segment S1-0 and the six second data segments S2-0, obtaining a first calculation result. The control logic unit 75 then sends the first calculation result to the data interface 79 for further calculation. The first data segment S1-1 and the six second data segments S2-1 are then sent to the six processing elements 82 to perform a convolution operation and generate a second calculation result. The control logic unit 75 then continuously sends the second calculation result to the data interface 79. In some embodiments, each of the calculation results can be stored in a result register of the corresponding processing element, a data register 77 of the peripheral circuit 64, or in one or more memory cell groups 66 of the memory cell array 62.
[0082] Figure 4A computational principle of at least one processing unit is provided, wherein the i-th first data segment is multiplied by the M second data segments of the i-th data group of N data groups to obtain the i-th result. The N i-th results are accumulated to obtain a convolution result. In some embodiments, the peripheral circuit 70 includes a processing unit 80, and the convolution operation between the N first data segments and the N data groups of the second data is then performed sequentially by the processing unit 80. In some embodiments, the peripheral circuit 70 includes more than one processing unit 80, and the convolution operation between the N first data segments and the N data groups of the second data is then performed simultaneously by different processing units 80.
[0083] At least one processing unit 80 is independently provided within the peripheral circuit 70 and is a separate module. As the number of at least one processing unit 80 within the peripheral circuit 70 increases, the computing speed of the peripheral circuit 70 improves, but at the same time, a larger area of the peripheral circuit 70 is required. There is a tradeoff between computing speed and the area of the peripheral circuit 70. In some embodiments, the memory cell array 62 is divided into more than one plane of memory cells, each plane including multiple memory cells. The number of at least one processing unit 80 is equal to the number of memory cell groups, meaning that at least one processing unit 80 corresponds to each of the multiple memory cell groups. For example, the memory cell array 62 is divided into 128 memory cell groups, and the number of at least one processing unit 80 is also 128. In some embodiments, the number of at least one processing unit 80 is less than the number of memory cell groups. For example, the memory cell array 62 is divided into 128 memory cell groups, and the number of at least one processing unit 80 can be 100, 64, 50, or another number less than 128. In some embodiments, the number of at least one processing unit 80 is half the number of groups of memory cells, and one processing unit corresponds to two groups of memory cells, respectively. For example, the memory cell array 62 is divided into 128 groups of memory cells, and the number of at least one processing unit 80 is 64. In some embodiments, the number of at least one processing unit 80 is one-quarter the number of groups of memory cells, and one processing unit corresponds to four groups of memory cells, respectively. For example, the memory cell array 62 is divided into 128 groups of memory cells, and the number of at least one processing unit 80 is 32. The number of at least one processing unit 80 can be set and adjusted based on the needs of the AI system. The embodiments of the present disclosure are intended to illustrate the present disclosure and should not be interpreted as limiting.
[0084] In another aspect of the present disclosure, the control logic unit 75 of the peripheral circuit 70 is configured to send the first data and the second data from the plurality of groups of memory cells to the at least one processing unit 80 .
[0085] refer to Figure 5 , shows the operation pipeline of a processing unit 80. Figure 5 As shown, N first data segments of the first data can be continuously sent to the first register 86. The sequence of the N first data segments of the first data pattern is the same as the sequence of the first data. The data length of each first data segment is less than or equal to the bandwidth of the data path bus 81. Figure 5 As shown, N data groups of second data can be continuously sent to the second register 88. The sequence of the N data groups of the second data pattern is the same as the sequence of the second data. The M second data segments within the same data group are arranged according to the order of the M columns of the second data. The first data segment and the second data segment are configured to share equal data length.
[0086] In some embodiments, at least one processing element 82 of each processing unit 80 performs Figure 5 In some embodiments, the number of processing elements 82 is equal to the number M of columns of the second data. Figure 5 As shown, the control logic unit 75 is configured to send the i-th first data segment to each processing element 82, and correspondingly send M second data segments of the i-th data group of the second data to the M processing elements 82. Each of the at least one processing unit 80 includes a control element 84, and the control element 84 is configured to distribute the M second data segments to the M processing elements accordingly based on the sequence of the M second data segments.
[0087] refer to Figure 3A and Figure 5 , shows the inputs of the first register 86, the second register 88, and the M processing elements 82. The first data segment S1-0, the first data segment S1-1, ..., and the first data segment S1-(N-1) are successively sent to the first register 86, and the M second data segments S2-0 of the first data group, the M second data segments S2-1 of the second data group, ..., and the M second data segments S2-(N-1) of the Nth data group are correspondingly sent to the second register 88. Figure 5 As shown, the first data segment S1-0 is distributed to each of the M processing elements, and the M second data segments S2-0 are distributed to the corresponding processing elements under the control of the control element 84. Figure 5As shown, the second data segment S2-0 from column 0 is assigned to the first processing element, the second data segment S2-0 from column 1 is assigned to the second processing element, and the second data segment S2-0 from column (M-1) is assigned to the Mth processing element. The M processing elements 82 then perform a first calculation and generate a first result based on the first data segment S1-0 and the M second data segments S2-0. The first result is copied and sent to the data register 77 by the control logic unit 75.
[0088] refer to Figure 3B and Figure 5 , the first data segment S1-1 is continuously sent to the first register 86, and the M second data segments S2-1 of the first data group are correspondingly sent to the second register 88. Then the first data segment S1-1 can be distributed to each of the M processing elements, and the M second data segments S2-1 can be distributed to the corresponding processing elements under the control of the control element 84. For example, Figure 5 As shown, the second data segment S2-1 from column 0 is distributed to the first processing element, the second data segment S2-1 from column 1 is distributed to the second processing element, and the second data segment S2-1 from column (M-1) is distributed to the Mth processing element. The M processing elements 82 perform a second calculation based on the first data segment S1-1, the M second data segments S2-1, and the first result to generate a second result. The first result is copied and sent to the data register 77 by the control logic unit 75.
[0089] refer to Figure 5 , the first data segment S1-(N-1) is continuously sent to the first register 86, and the first group of M second data segments S2-(N-1) is continuously sent to the second register 88. The first data segment S1-(N-1) can then be distributed to each of the M processing elements, and the M second data segments S2-(N-1) can be distributed to the corresponding processing elements under the control of the control element 84. For example, Figure 5 As shown, the second data segment S2-(N-1) from column 0 is distributed to the first processing element, the second data segment S2-(N-1) from column 1 is distributed to the second processing element, and the second data segment S2-(N-1) from column (M-1) is distributed to the (M-1)th processing element. The M processing elements 82 perform the (N-1)th calculation based on the first data segment S1-(N-1), the M second data segments S2-(N-1), and the (N-2)th result to generate the Nth result. The Nth result is copied by the control logic unit 75 and sent to the data register 77.
[0090] A system is provided, comprising a memory device and a controller. The memory device includes a plurality of memory cell groups and peripheral circuitry coupled to the memory cells. The peripheral circuitry includes a control logic unit configured to program first and second data into the plurality of memory cells; at least one processing unit configured to perform a calculation based on the first and second data; and a data path bus coupled to the control logic unit and the at least one processing unit for transmitting the first and second data. The controller is coupled to the memory device and configured to transmit the first data to the memory device and receive a result of the calculation from the memory device.
[0091] In some embodiments, the system can be any electrical system to which an AI system is applied, such as a computer, a digital camera, a mobile phone, a smart appliance, the Internet of Things (IoT), a server, a base station, etc. In the present disclosure, data processing and calculations of the AI system can be performed by the processing unit 80 of the peripheral circuit of the memory device. In some embodiments, by adding at least one processing unit to the memory device, computing tasks that consume a lot of resources can be distributed to the memory device instead of the TPU or the graphics processing unit (GPU), thereby improving the performance of the AI system. The number of processing units can be designed based on the needs of the AI system. By integrating more processing units into the memory device, the AI system will be more efficient.
[0092] refer to Figure 6 , Figure 6 A flow chart of a method 600 for performing data calculations using a memory device is shown. The memory device includes an array of memory cell arrays 62 and peripheral circuits 70 coupled to the memory cell arrays 62. The memory device can be the same as described above and will not be repeated herein. It should be understood that the operations shown in method 600 are not exhaustive and that other operations may be performed before, after, or between any of the operations shown. Furthermore, some of these operations may be performed simultaneously or in parallel. Figure 6 The order shown in the following example is different from the order in which they are executed.
[0093] like Figure 6 As shown, method 600 may begin at operation 602, where a control logic unit of a peripheral circuit obtains first data and second data from a data interface of a memory device. Method 600 then proceeds to operation 604, where the first data and second data are programmed into memory cell array 62 based on a preset data pattern.
[0094] In some embodiments, as Figure 2AAs shown, in many cases, the first data may be one-dimensional data and the second data may be two-dimensional data, wherein the first data is a one-dimensional vector and the second data is a two-dimensional matrix. In some embodiments, the one-dimensional first data may be equivalent to a row of data having a length of a, and the two-dimensional second data (i.e., an a×b matrix) may be equivalent to b columns, each having a length of a. In some embodiments, the first data may be a two-dimensional matrix including more than one row of equal data length, and dimensionality reduction may be performed on more than one row of first data to decompose the first data into multiple single rows, thereby applying the present disclosure.
[0095] In some embodiments, the Figure 2C The first data pattern and the second data pattern shown in FIG. 6A program the first data and the second data into the memory cell array 62 .
[0096] In some embodiments, the first data includes a row. The row of first data is programmed into one of a plurality of groups of memory cells based on a first data pattern. For example, Figure 1F As shown, the first data can be programmed into group 0 of multiple groups of memory cells. The first data pattern includes N first data segments with equal data length, where N is a positive integer and N≥2. In this embodiment, N=4 is used as an example to illustrate the present disclosure. The first data includes four first data segments based on the first data pattern, namely, the first data segment S1-0, the first data segment S1-1, the first data segment S1-2 and the first data segment S1-3. The four first data segments are programmed into group 1 continuously. The sequence of the four first data segments of the first data pattern is the same as the sequence of the first data. In some embodiments, the data length of each first data segment is less than or equal to the bandwidth of the data path bus. In some embodiments, an error checking and correction (ECC) code is assigned to each first data segment for verification. The ECC code can also be used as an identifier for each first data segment, which is configured to identify that the first data pattern is applied to the first data.
[0097] In some embodiments, the second data includes M columns, where M is a positive integer and M ≥ 2. Each of the M columns is programmed into M groups of memory cells in a plurality of groups of memory cells based on the second data pattern, where the number of the plurality of groups of memory cells is greater than M. In some embodiments, Figure 1F As shown, the number of groups of memory cells of the memory cell array 62 is sixteen, and the number of columns of the second data is six (as shown in FIG. Figure 2B). The six columns of second data can then be programmed into six groups of memory cells in the sixteen memory groups. For example, column 0 can be programmed into group 1, column 1 can be programmed into group 2, column 2 can be programmed into group 3, column 3 can be programmed into group 4, column 4 can be programmed into group 5, and column 5 can be programmed into group 6.
[0098] The second data pattern includes N data groups, each of the N data groups has M second data segments with equal data lengths respectively from M columns of the second data, and the first data segment and the second data segment are configured to share equal data lengths. In some embodiments, N=4 and M=6 are used as examples to illustrate the present disclosure. Figure 2C As shown, the second data includes six columns; each of the six columns includes four second data segments, namely, the second data segment S2-0, the second data segment S2-1, the second data segment S2-2 and the second data segment S2-3. Figure 2C , six second data segments S2-0 are regrouped into a first data group of a second data pattern, six second data segments S2-1 are regrouped into a second data group of the second data pattern, six second data segments S2-2 are regrouped into a third data group of the second data pattern, and six second data segments S2-3 are regrouped into a fourth data group of the second data pattern. Each second data segment of a data group is programmed into a different memory group. For example, the first data group includes a second data segment S2-0 from column 0 programmed into group 2, a second data segment S2-0 from column 1 programmed into group 3, a second data segment S2-0 from column 3 programmed into group 5, a second data segment S2-0 from column 4 programmed into group 6, and a second data segment S2-0 from column 5 programmed into group 7. In some embodiments, an error checking and correction (ECC) code is assigned to each second data segment for verification. The ECC code can also be used as an identifier for each second data segment, which is configured to identify that the second data pattern is applied to the second data.
[0099] like Figure 6 As shown, method 600 may begin at operation 606 , where at least one processing unit of the peripheral circuit performs a calculation based on first data and second data.
[0100] In some embodiments, based on Figure 2C The first data pattern in the data sequence senses each row of the first data and sends it to at least one processing unit 80. For example, Figure 3A As shown, the first data segment S1-0 is first sent to the processing unit 80, and as shown in FIG. Figure 3BAs shown, the first data segment S1-1 is sent to the processing unit 80 after the first data segment S1-0. The first data segment S1-2 and the first data segment S1-3 are sent to the processing unit 80 after the first data segment S1-1 (not shown). Taking the first data segment S1-0 as an example, in some embodiments, the first data segment S1-0 is sent to the first register 86 and cached in the first register 86. In some embodiments, as shown Figure 3A As shown, the first data segment S1-1 is distributed to each processing element 82 by the control element 84 of each processing unit.
[0101] In some embodiments, based on Figure 2C The data sequence of the second data pattern in the second data pattern senses the data group of the second data pattern and sends it to at least one processing unit 80. For example, Figure 3A As shown, first, the six second data segments S2-0 of the first data group are sent to the processing unit 80, and as shown in FIG. Figure 3B As shown, six second data segments S2-1 of the second data group are continuously sent to the processing unit 80 after the first data group. Six second data segments S2-2 of the third data group of the second data are continuously sent to the processing unit 80 after the second data group, and six second data segments S2-3 of the fourth data group of the second data are continuously sent to the processing unit 80 after the third data group (not shown). Taking the six second data segments S2-0 as an example, in some embodiments, the six second data segments S2-0 are sent to the six second registers 88 and cached in the six second registers 88. In some embodiments, as Figure 3A As shown, each processing unit includes a control element configured to distribute six second data segments S2 - 0 to six processing elements accordingly based on the sequence of the six second data segments S2 - 0 .
[0102] In some embodiments, operation 606 includes: performing, by the M processing elements 82 of each processing unit 80, a convolution operation based on the i-th first data segment of the N first data segments and the M second data segments of the i-th data group of the N data groups, where i is a positive integer and N≥i≥1. Figure 3A and Figure 3BIn this embodiment, the first data segment S1-0 is sent to each of the six processing elements 82, and the six second data segments S2-0 of the first data group of the second data are correspondingly sent to the six processing elements 82. The six processing elements 82 then perform a convolution operation on the first data segment S1-0 and the six second data segments S2-0 to obtain a first calculation result. The control logic unit 75 then sends the first calculation result to the data interface 79. The first data segment S1-1 and the six second data segments S2-1 are then sent to the six processing elements 82 to perform the convolution operation and generate a second calculation result. The control logic unit 75 then continuously sends the second calculation result to the data interface 79. In some embodiments, each of the calculation results can be stored in a result register of the corresponding processing element, a data register 77 of the peripheral circuit 64, or in one or more memory cell groups 66 of the memory cell array 62.
[0103] exist Figure 5 A computational principle of at least one processing unit is provided, wherein the i-th first data segment is multiplied by the M second data segments of the i-th data group of N data groups to obtain the i-th result. The N i-th results are accumulated to obtain a convolution result. In some embodiments, the peripheral circuit 70 includes a processing unit 80, and the convolution operation between the N first data segments and the N data groups of the second data is then performed continuously by the processing unit 80. In some embodiments, the peripheral circuit 70 includes more than one processing unit 80, and the convolution operation between the N first data segments and the N data groups of the second data is then performed simultaneously by different processing units 80.
[0104] The foregoing description of specific embodiments can be readily modified and / or adapted for various applications. Therefore, based on the teaching and guidance provided herein, such adaptations and modifications are intended to be within the meaning and range of equivalents of the disclosed embodiments.
[0105] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.
[0106] Although specific configurations and arrangements have been discussed, it should be understood that this is for illustrative purposes only. Therefore, other configurations and arrangements may be used without departing from the scope of this disclosure. In addition, the subject matter as described in this disclosure may also be used in a variety of other applications. The functions and structural features as described in this disclosure may be combined, adjusted, modified, and rearranged with one another in a manner consistent with the scope of this disclosure.
Claims
1. A memory device comprising: groups of memory cells; as well as a peripheral circuit coupled to the group of memory cells and comprising: a control logic unit configured to program first data and second data into different ones of the groups of memory cells; At least one processing unit is coupled to the group of memory cells via a data path bus of the peripheral circuit and is configured to perform a calculation based on the first data and the second data.
2. The memory device according to claim 1, wherein The first data includes at least one row; and The control logic unit is configured to: receiving the first data from a data interface; and Each row of the first data is programmed into one of the groups of memory cells based on a first data pattern.
3. The memory device of claim 2, wherein: The first data pattern includes N first data segments having equal data lengths, wherein N is a positive integer and N≥2; and The sequence of the N first data segments of the first data pattern is the same as the sequence of the first data.
4. The memory device according to claim 3, wherein The data length of each first data segment is less than or equal to the bandwidth of the data path bus.
5. The memory device according to claim 3, wherein The second data includes M columns, where M is a positive integer and M≥2; and The control logic unit is configured to program each of the M columns into M groups of memory cells in the groups of memory cells based on a second data pattern, the number of the groups of memory cells being greater than M. The memory device according to claim 5 , wherein: The second data pattern includes N data groups, each of the N data groups having M second data segments with equal data lengths respectively from the M columns of the second data; and The first data segment and the second data segment are configured to share equal data length.
7. The memory device according to claim 6, wherein: Each of the M second data segments of each data group of the second data is assigned an error checking and correcting (ECC) code.
8. The memory device according to claim 6, wherein The data length of each second data segment is less than or equal to the bandwidth of the data path bus.
9. The memory device according to claim 6, wherein: The control logic unit is configured to receive the second data from the data interface of the memory device based on the second data pattern.
10. The memory device according to claim 6, wherein Each processing unit of the at least one processing unit includes M processing elements and is configured to: A convolution operation is performed based on an i-th first data segment among the N first data segments and the M second data segments of an i-th data group among the N data groups, where i is a positive integer and N≥i≥1. The memory device according to claim 10 , wherein: The control logic unit is configured to: controlling the one of the groups of memory units to send the i-th first data segment of the first data to each of the M processing elements; as well as The M groups of memory units are controlled to send the M second data segments to the M processing elements.
12. The memory device according to claim 11, wherein Each of the at least one processing unit comprises a control element configured to: The M second data segments are correspondingly allocated to the M processing elements based on the sequence of the second data.
13. The memory device according to claim 12, wherein: The control logic unit is configured to: The calculation result is output to a data interface of the peripheral circuit of the memory device.
14. The memory device according to claim 13, wherein: The control logic unit is configured to: The calculation results are output to the group of memory cells.
15. The memory device according to claim 13, wherein: The number of the at least one processing unit is equal to the number of the groups of memory units; and Each processing unit corresponds to a corresponding group of the groups of memory units.
16. The memory device according to claim 13, wherein The number of the at least one processing unit is smaller than the number of the groups of the memory units.
17. The memory device according to claim 16, wherein: The number of the at least one processing unit is half the number of the groups of the memory units; and Each processing unit corresponds to two corresponding groups of memory units.
18. The memory device according to claim 16, wherein: The number of the at least one processing unit is one quarter the number of the groups of the memory units; and Each processing unit corresponds to four corresponding groups of memory units.
19. The memory device of claim 16, wherein: The number of the at least one processing unit is one; and The one processing unit corresponds to the group of memory units.
20. The memory device of claim 1, wherein: The memory device includes dynamic random access memory (DRAM).
21. A method for performing data calculation using a memory device, the memory device comprising a group of memory cells and a peripheral circuit coupled to the group of memory cells, the method comprising: obtaining, by a control logic unit of the peripheral circuit, first data and second data from a data interface of the memory device via a data path bus of the peripheral circuit; programming the first data and the second data into the group of memory cells; as well as A calculation is performed by at least one processing unit of the peripheral circuit based on the first data and the second data.
22. The method according to claim 21, wherein The first data includes at least one row; and Programming the first data and the second data into the group of memory cells includes: Each row of the first data is programmed into one of the groups of memory cells based on a first data pattern.
23. The method according to claim 22, wherein The first data pattern includes N first data segments having equal data lengths, wherein N is a positive integer and N≥2; and The sequence of the N first data segments of the first data pattern is the same as the sequence of the first data.
24. The method according to claim 23, wherein The data length of each first data segment is less than or equal to the bandwidth of the data path bus.
25. The method according to claim 23, wherein The second data includes M columns, where M is a positive integer and M≥2, and obtaining the second data from the data interface of the memory device includes: Each of the M columns is programmed into M groups of memory cells in the groups of memory cells based on a second data pattern, the number of groups of memory cells being greater than M.
26. The method according to claim 25, wherein The second data pattern includes N data groups, each of the N data groups having M second data segments with equal data lengths respectively from the M columns of the second data; and The first data segment and the second data segment are configured to share equal data length.
27. The method according to claim 26, wherein Obtaining the second data from the data interface of the memory device includes: An error checking and correction (ECC) code is assigned to each of the M second data segments of each data group of the second data.
28. The method according to claim 26, wherein The data length of each second data segment is less than or equal to the bandwidth of the data path bus.
29. The method according to claim 26, wherein Performing a calculation based on the first data and the second data includes: Each M processing element of the at least one processing unit performs a convolution operation based on the i-th first data segment among the N first data segments and the M second data segments of the i-th data group among the N data groups, where i is a positive integer and N≥i≥1.
30. The method according to claim 29, wherein Performing a calculation based on the first data and the second data includes: sending the i-th first data segment of the first data from the one memory bank to each processing element of the M processing elements; and The M second data segments are sent from the M groups of memory units to the M processing elements.
31. The method of claim 21 , further comprising: The calculation result is output to a group of memory cells or a data interface of the peripheral circuit of the memory device.
32. A system comprising: A memory device, the memory device comprising: a group of memory cells in the memory cell; and a peripheral circuit coupled to the group of memory cells and comprising: a control logic unit configured to program first data and second data into the memory group; at least one processing unit coupled to the group of memory cells via a data path bus of the peripheral circuit and configured to perform calculations based on the first data and the second data; and A controller is coupled to the memory device and configured to transmit the first data to the memory device and receive a result of the calculation from the memory device.
33. The system of claim 32, wherein: The controller is further configured to transfer the second data to the memory device.
34. The system of claim 32, wherein: The memory device includes dynamic random access memory (DRAM).