Techniques for data copy and data move operations in low power double data rate (LPDDR) memory devices
Patent Information
- Application Number
- PCT/CN2025/085732
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025085732_01102026_PF_FP_ABST
Abstract
Description
TECHNIQUES FOR DATA COPY AND DATA MOVE OPERATIONS IN LOW POWER DOUBLE DATA RATE (LPDDR) MEMORY DEVICESBACKGROUNDField
[0001] Aspects of the present disclosure relate to computing devices, and more specifically to techniques for performing data move and data copy operations in low power double data rate (LPDDR) dynamic random access memory (DRAM) devices. Background
[0002] In computing systems, memory may be volatile memory or nonvolatile memory. Volatile memory, such as random access memory (RAM) , specifies continuous power to maintain stored data. Volatile memory is primarily used for temporary storage while a processor is performing computations. Non-volatile memory, such as negative OR (NOR) memory and negative AND (NAND) memory used in solid state drives (SSDs) and hard drives, retains data even when powered off, making non-volatile memory preferable for long-term data storage.
[0003] A host device, such as a system-on-a-chip (SoC) , may be coupled to memory via command buses, data buses, and clock buses. Command buses transmit operational instructions from the SoC to memory. The SoC may use a command bus to direct how data should be handled, whether stored, erased, or modified. Data buses facilitate the transfer of data back and forth between the SoC and the memory, enabling the system to access and utilize the data as specified for processing tasks. Clock buses provide timing signals to synchronize data transfers and operations across the system.
[0004] Communications between the host device and the memory are slower than on-chip communications. Moreover, bandwidth restrictions may further limit the throughput of communications between the host and the memory. It would be desirable to improve processing involving memory, such as dynamic random access memory (DRAM) , by reducing communications between the host and the memory.SUMMARY
[0005] Aspects of the present disclosure are directed to an apparatus. The apparatus includes a first memory subchannel. The first memory channel includes a first set of memory banks. The first memory channel also includes a first set of row buffers. Each of the row buffers is coupled with a respective memory bank of the first set of memory banks. The first memory channel further includes a first data bus coupled to the first set of row buffers. The first memory channel still further includes a first data mover (DM) coupled to the first data bus. The first memory channel also includes a first set of meta data registers coupled to the first data bus.
[0006] In other aspects of the present disclosure, a method for moving data within a low power double data rate (LPDDR) memory device includes transferring the data from a source memory address of a source location within the LPDDR memory device to a temporary buffer located within the LPDDR memory device. The source location initially stores the data to be transferred. The method also includes transferring the data from the temporary buffer to a target memory address of a target location within the LPDDR memory device such that the data does not leave the LPDDR memory device during the transfer.
[0007] This has outlined, rather broadly, the features and technical advantages of the present disclosure in order that the detailed description that follows may be better understood. Additional features and advantages of the present disclosure will be described below. It should be appreciated by those skilled in the art that this present disclosure may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. It should also be realized by those skilled in the art that such equivalent constructions do not depart from the teachings of the present disclosure as set forth in the appended claims. The novel features, which are believed to be characteristic of the present disclosure, both as to its organization and method of operation, together with further objects and advantages, will be better understood from the following description when considered in connection with the accompanying figures. It is to be expressly understood, however, that each of the figures is provided for the purpose of illustration and description only and is not intended as a definition of the limits of the present disclosure.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] For a more complete understanding of the present disclosure, reference is now made to the following description taken in conjunction with the accompanying drawings.
[0009] FIGURE 1 illustrates an example implementation of a host system-on-a-chip (SoC) 100, that communicates with dynamic random access memory (DRAM) having data move / copy capabilities, in accordance with aspects of the present disclosure.
[0010] FIGURE 2 is a block diagram illustrating an SoC in conventional communication with memory, in accordance with various aspects of the present disclosure.
[0011] FIGURE 3 is a block diagram illustrating a configuration of low power double data rate (LPDDR) dynamic random access memory (DRAM) , in accordance with various aspects of the present disclosure.
[0012] FIGURE 4 is a block diagram illustrating another configuration of low power double data rate (LPDDR) dynamic random access memory (DRAM) , in accordance with various aspects of the present disclosure.
[0013] FIGURE 5 is a block diagram illustrating a meta data register configuration in a subchannel input / output (IO) block, in accordance with various aspects of the present disclosure.
[0014] FIGURE 6 is a block diagram illustrating a processor in memory (PIM) architecture, in accordance with various aspects of the present disclosure.
[0015] FIGURE 7 is a block diagram illustrating a DRAM bank, in accordance with various aspects of the present disclosure.
[0016] FIGURE 8 is a table illustrating low power double data rate version six (LPDDR6) addressing.
[0017] FIGURE 9 is a block diagram illustrating an exemplary architecture for internal per subchannel move / copy operations, in accordance with various aspects of the present disclosure.
[0018] FIGURE 10 is a block diagram illustrating an exemplary architecture for internal move / copy operations in two subchannels, in accordance with various aspects of the present disclosure.
[0019] FIGURE 11 is a table illustrating parameters for a move command, in accordance with various aspects of the present disclosure.
[0020] FIGURE 12 is a flow diagram illustrating an example process performed, for example, by a computing device, in accordance with various aspects of the present disclosure.
[0021] FIGURE 13 is a block diagram showing an exemplary wireless communications system in which a configuration of the present disclosure may be advantageously employed.
[0022] FIGURE 14 is a block diagram illustrating a design workstation used for circuit, layout, and logic design of components, in accordance with various aspects of the present disclosure.DETAILED DESCRIPTION
[0023] The detailed description set forth below, in connection with the appended drawings, is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of the various concepts. It will be apparent, however, to those skilled in the art that these concepts may be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form in order to avoid obscuring such concepts.
[0024] As described, the use of the term “and / or” is intended to represent an “inclusive OR, ” and the use of the term “or” is intended to represent an “exclusive OR. ” As described, the term “exemplary” used throughout this description means “serving as an example, instance, or illustration, ” and should not necessarily be construed as preferred or advantageous over other exemplary configurations. As described, the term “coupled” used throughout this description means “connected, whether directly or indirectly through intervening connections (e.g., a switch) , electrical, mechanical, or otherwise, ” and is not necessarily limited to physical connections. Additionally, the connections can be such that the objects are permanently connected or releasably connected. The connections can be through switches. As described, the term “proximate” used throughout this description means “adjacent, very near, next to, or close to. ” As described, the term “on” used throughout this description means “directly on” in some configurations, and “indirectly on” in other configurations.
[0025] With existing dynamic random access memory (DRAM) configurations, data movement, especially between a system-on-a-chip (SoC) and low power double data rate (LPDDR) DRAM (e.g., off-chip to on-chip) , is very expensive in terms of bandwidth, energy, and latency. A data move or data copy operation is a common operation in computing systems. For example, a data move from a source in memory to a destination in memory requires the SoC to read from the source inside the memory and write back to the destination within the memory. These operations consume significant power and increase latency on the memory channel. A data copy operation specifies a read from the source and a write to the destination. It would be desirable to reduce movement between the memory and the SoC.
[0026] According to aspects of the present disclosure, data move and data copy operations are performed within DRAM, without using input / output (IO) communications to and from the SoC. For a move / copy operation within a same DRAM bank, initially, a row buffer connection is opened to a source location being read. Data from the source location is then written into the row buffer. Next, a newly added command: move / copy data, is issued from a starting column (C) address with a transfer size. The move / copy data command may move data from the source location into meta data registers, row buffers, a processor in memory (PIM) register file (PRF) , or any other designated temporary buffer. Next, a row buffer connection is opened to a destination location, and data from the destination location is written into the row buffer. A move / copy data command is then issued to move or copy data from the temporary buffer into the destination location. Additional aspects relate to move / copy operations between different banks.
[0027] Particular aspects of the subject matter described in this disclosure can be implemented to realize one or more of the following potential advantages. In some examples, the described techniques, such as moving data within an LPDDR memory device, significantly reduce power consumption and latencies for copying / moving data from one memory location to another location.
[0028] FIGURE 1 illustrates an example implementation of a host system-on-a-chip (SoC) 100, that communicates with dynamic random access memory (DRAM) having data move / copy capabilities, in accordance with aspects of the present disclosure. The host SoC 100 includes processing blocks tailored to specific functions, such as a connectivity block 110. The connectivity block 110 may include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity, universal serial bus (USB) connectivity, connectivity, Secure Digital (SD) connectivity, and the like.
[0029] In this configuration, the host SoC 100 includes various processing units that support multi-threaded operation. For the configuration shown in FIGURE 1, the host SoC 100 includes a multi-core central processing unit (CPU) 102, a graphics processor unit (GPU) 104, a digital signal processor (DSP) 106, and a neural processor unit (NPU) 108. The host SoC 100 may also include a sensor processor 114, image signal processors (ISPs) 116, a navigation module 120, which may include a global positioning system (GPS) , and a memory 118. The multi-core CPU 102, the GPU 104, the DSP 106, the NPU 108, and the multi-media engine 112 support various functions such as video, audio, graphics, gaming, artificial networks, and the like. Each processor core of the multi-core CPU 102 may be a reduced instruction set computing (RISC) machine, an advanced RISC machine (ARM) , a microprocessor, or some other type of processor. The NPU 108 may be based on an ARM instruction set.
[0030] FIGURE 2 is a block diagram illustrating an SoC in conventional communication with memory, in accordance with various aspects of the present disclosure. As shown in FIGURE 2, an SoC 200 is coupled to serial negative OR (NOR) flash memory 202 via a first data bus 204. The SoC 200 is also coupled to low power double data rate (LPDDR) memory 206 via a second data bus 208. The “6” appended to the “LPDDR” label in FIGURE 2 indicates that the LPDDR memory 206 may be version six or later of LPDDR memory.
[0031] FIGURE 3 is a block diagram illustrating a configuration of low power double data rate (LPDDR) dynamic random access memory (DRAM) , in accordance with various aspects of the present disclosure. In the example of FIGURE 3, an SoC 300 communicates with LPDDR memory 306 including two subchannels SC0 and SC1. Each subchannel SC0, SC1 includes 16 banks of memory. A subchannel physical layer (SC PHY) is allocated within the SoC 300 for each subchannel SC0, SC1. One data bus (X12) DQ, one command bus (x4) CA, and one clock line CK are provided for each subchannel SC0, SC1.
[0032] With the existing dynamic random access memory (DRAM) configurations, data movement, especially between the SoC 300 and the LPDDR DRAM 306 (e.g., off-chip to on-chip) , is expensive in terms of bandwidth, energy, and latency. Such data movement is more expensive than computation. A data move or data copy operation is a common operation in computing systems. For example, a data move from a source in memory to a destination in memory requires the SoC 300 to read from the source inside the memory 306 and write back to the destination within the memory 306. These operations consume significant power and increase latency on the memory channel. A copy operation requires a read from the source and a write to the destination. For LPDDR6, reading / writing 256 bits requires fourteen clock cycles. Copying one row (e.g., 16, 384 bits) takes an additional (2*14*16, 384 / 256) 1, 792 clock cycles on the memory bus. It would be desirable to reduce movement between the memory 306 and the SoC 300.
[0033] FIGURE 4 is a block diagram illustrating another configuration of low power double data rate (LPDDR) dynamic random access memory (DRAM) , in accordance with various aspects of the present disclosure. In the example of FIGURE 4, four bank groups BG0, BG1, BG2, BG3 are provided within each subchannel SC0, SC1. Each bank group BG0, BG1, BG2, BG3 includes four memory banks BK0, BK1, BK2, BK3. An IO block 402 for each subchannel SC0, SC1 includes a write clock line WCK, twelve data lines X12DQs, a read data strobe line RDQS, a chip select line CS, four command address line 4CAs, and a clock line CK. The ‘_t / c’ in FIGURE 4 represents true (_t) and complement (_c) of a differential pair.
[0034] FIGURE 5 is a block diagram illustrating a meta data register configuration in a subchannel input / output (IO) block, in accordance with various aspects of the present disclosure. In the example of FIGURE 5, each LPDDR6 dynamic random access memory (DRAM) subchannel IO block (e.g., for the X12 DQs) includes a total of sixteen meta data registers MDR0 to MDR15. One meta data register is assigned to each bank and designated to carve out column addresses (0x3C, 0x3D, 0x3E, 0x3F) for meta data. Each 32-byte meta data register stores sixteen two-byte pieces of meta data from associated sixteen column addresses except address 0x3F. It would be desirable to repurpose the meta data registers as intermediate buffers when moving and copying data.
[0035] FIGURE 6 is a block diagram illustrating a processor in memory (PIM) architecture, in accordance with various aspects of the present disclosure. The example of FIGURE 6 shows a host SoC communicating with a set of LPDDR6 dies, each including a PIM instruction table (PIT) and multiple DRAM banks. The PIT is a software programmed table inside memory that translates from protocol commands to PIM instructions. Memory internal PIM instructions may include MAC, MULT, MOV, NOP, register specifiers, precisions, etc. The MAC operation is a multiply and accumulate operation, the MULT operation is a multiply operation, the MOV operation is a move operation, NOP represents no operation, register specifiers indicate an address or location of a register, and precisions indicates fixed or floating point precision.
[0036] An LPDDR6 memory interface is provided for PIM commands, such as PIM compute commands, and commands for reading and writing PIM registers.
[0037] FIGURE 7 is a block diagram illustrating a DRAM bank, in accordance with various aspects of the present disclosure. The example of FIGURE 7 shows an LPDDR6-PIM architecture with a processor in memory (PIM) execution unit (PEU) for each DRAM bank. A single PEU comprises a PIM arithmetic logic unit (ALU) and a PIM register file (PRF) . A minimum of sixteen 256-byte registers are specified in the PRF. The DRAM bank also includes DRAM cells with a row buffer / sense amplifier. Although FIGURE 7 illustrates one PEU per bank, in alternate architectures one PEU is provided for each bank group or for each subchannel.
[0038] FIGURE 8 is a table illustrating LPDDR6 addressing, in accordance with various aspects of the present disclosure. As seen in the table of FIGURE 8, the architecture includes four bank groups, with four banks in every bank group (BG) . There are 64 columns per row with 32 bytes per column. The number of rows for each bank varies based on the memory size.
[0039] According to aspects of the present disclosure, data move and data copy operations are performed within DRAM, without using IO communications to and from the SoC. FIGURE 9 is a block diagram illustrating an exemplary architecture for internal per subchannel move / copy operations, in accordance with various aspects of the present disclosure. In the example of FIGURE 9, a subchannel 900 includes N+1 DRAM banks 902 (Bank 0 to Bank N) , along with corresponding row buffers 904. A PEU 906 may optionally be provided for each DRAM bank 902. In an alternative configuration, a single PEU 908 is provided for the entire subchannel 900. The PEU (s) 906, 908 include a PRF, which may potentially be used as an intermediate buffer for data storage during data move / copy operations. The PEU (s) 906, 908 are each coupled to a data bus, which may be a 32-byte data bus. N+1 meta data registers 912 (Meta Data Register 0 to Meta Data Register N) are also coupled to the data bus. The meta data registers 912 may also be potential intermediate buffers for data storage during data move / copy operations. A data mover 910 is coupled to the data bus and enables internal data move / copy operations. The move / copy operation may be performed either with or without the PEU (s) 906, 908.
[0040] FIGURE 10 is a block diagram illustrating an exemplary architecture for internal move / copy operations in two subchannels, in accordance with various aspects of the present disclosure. In the example of FIGURE 10, a first subchannel 1000 includes N+1 DRAM banks 1002 (Bank 0 to Bank N) , along with corresponding row buffers 1004. A PEU 1006 may optionally be provided for each DRAM bank 1002. In an alternative configuration, a single PEU 1008 is provided for the entire subchannel 1000. The PEU (s) 1006, 1008 include a PRF, which may potentially be used as an intermediate buffer for data storage during data move / copy operations. The PEU (s) 1006, 1008 are each coupled to a data bus, which may be a 32 byte data bus. N+1 meta data registers 1012 (Meta Data Register 0 to Meta Data Register N) are also coupled to the data bus. The meta data registers 1012 may also be potential intermediate buffers for data storage during data move / copy operations. A second subchannel 1050 is similarly configured.
[0041] A switch 1020 selectively couples a corresponding data bus of each subchannel 1000, 1050 to a set of data movers 1010. The set of data movers 1010 includes one independent data mover (SC0 DM, SC1 DM) for each subchannel 1000, 1050. As a result of the independent data movers 1010, each subchannel 1000, 1050 can perform parallel data movements independently.
[0042] The data movers 910, 1010 of FIGURES 9 and 10 enable internal data move / copy operations as will now be described in more detail in which the move / copy operations may be performed either with or without the PEU 906, 908, 1006, 1008.
[0043] For a move / copy operation within a same memory bank, initially, a row buffer connection is opened to a source location being read. The connection may be opened by issuing ACTIVATE-1 and ACTIVATE-2 commands specifying a bank group (BG) , a bank (BA) , and a row (R) address for the source location. Data from the source location is thus written into the row buffer. Next, a newly added command: move / copy data, is issued from a starting column (C) address with a transfer size. The column address and transfer size may be indicated by the command or mode registers. The move / copy data command may be issued into meta data registers, row buffers, the PRF, or any other designated temporary buffer. A PRECHARGE command then closes the row buffer connection to the source BG, BA, and R. By closing the row buffer connection, any data currently in the row buffer is written to the source location.
[0044] Next, a row buffer connection is opened to a destination location. The connection may be opened by issuing ACTIVATE-1 and ACTIVATE-2 commands specifying a bank group (BG) , a bank (BA) , and a row (R) address for the destination location. The BG and BA values are provided if the destination location is a different location than the source location. Data from the destination location is thus written into the row buffer, overwriting the previous data stored in the row buffer. Next, a move / copy data command is issued to move or copy data from the meta data registers, row buffers, the PRF, or any designated temporary buffer that was used when moving / copying data from the source location into the destination location. The move / copy command indicates a starting column (C) address, a transfer size, the BG, BA, and R address of the destination location. The column address and transfer size may be indicated by the command or mode registers. The column address and transfer size can be specified in the Move / Copy command directly as described, or specified in the mode registers.
[0045] The data moves to the row buffer corresponding to the destination address with the move / copy command. Similar to the source operation, the system knows which row is opened. When the row is closed, the content of the row buffer is written back to the same row location. A PRECHARGE command then closes the row buffer connection to the destination bank group, bank, and row. By closing the row buffer connection, any data currently in the row buffer is written to the destination location. If the size of the move / copy operation is greater than the temporary buffer size, (e.g., the number of meta data registers) , then the column address may be incremented and the steps repeated until the move / copy operation completes.
[0046] A data move / copy operation between different banks is now described. First, a row buffer connection is opened to a source location being read. The connection may be opened by issuing ACTIVATE-1 and ACTIVATE-2 commands specifying a bank group (BG) , a bank (BA) , and a row (R) address for the source location. Next, a row buffer connection is opened to a destination location. The connection may be opened by issuing ACTIVATE-1 and ACTIVATE-2 commands specifying a bank group (BG) , a bank (BA) , and a row (R) address for the destination location. The BG and BA values are provided if the destination location is a different location than the source location.
[0047] A data move / copy command is then issued from a starting column (C) address with a transfer size. The column address and transfer size may be indicated by the command or mode registers. The move / copy data command may be issued into meta data registers, row buffers, the PRF, or any other designated temporary buffer such that the source data transfers into the selected temporary buffer. Alternatively, the move / copy command is issued directly from a source row buffer to a destination row buffer, such that data is transferred directly between these row buffers.
[0048] Next, a move / copy data command is issued to move or copy data from the meta data registers, row buffers, the PRF, or any designated temporary buffer that was used when moving / copying data from the source destination. The move / copy command indicates a starting column (C) address, a transfer size, the BG, BA, and R address of the destination location. The column address and transfer size may be indicated by the command or mode registers. This step may not be necessary if row buffer to row buffer transfer is allowed. The switch 1020 (FIGURE 10) between subchannels switch is enabled if an inter-subchannel data move / copy occurs.
[0049] After moving the data, a PRECHARGE command closes the row buffer connections to the source and destination BG, BA, and R address. By closing the row buffer connections, the data currently in the row buffers is written to the source and destination locations. If the size of the move / copy operation is greater than the temporary buffer size (e.g., the number of meta data registers) , then the column address increments and the issuing of move / copy commands for the source and destination locations repeats until the move / copy operation completes.
[0050] Aspects of the present disclosure introduce data move / copy commands. Data move / copy blocks are also included, for example, a single integrated data move / copy block (see data mover 910 of FIGURE 9) or two separate data movers where each data mover handles one subchannel (see data movers 1010 of FIGURE 10) . The source address, destination address, and transfer size are programmable through SoC program mode registers before the operation occurs. The source / destination address include the BG, the BA, row address, and column address. The source and destination addresses can start from any column.
[0051] Additionally, data move / copy operations may occur across subchannels. Meta data registers, row buffers, and PIM register files (PRFs) may be employed, as needed, to buffer data. Each subchannel has 16 meta data registers that can support a transfer size of up to 16 columns. To move / copy one row (64 columns) , four operations are specified. Employing all meta registers (32) from both subchannels allows for a transfer size of up to 32 columns. In this scenario, two operations move / copy one row (64 columns) . A switch to connect two subchannel data buses is included. A mode register controls switch enablement.
[0052] According to aspects of the present disclosure, a computing device includes a DRAM device. The DRAM device may include means for transferring, means for opening, means for repeating, and means for closing. In one configuration, the transferring means, the opening means, the repeating means, and the closing means may be the data movers 910, 1010, meta data registers 912, 1012, row buffers 904, 1004, and / or PEUs 906, 908, 1006, 1008 as shown in FIGURES 9 and 10. In other aspects, the aforementioned means may be any structure or any material configured to perform the functions recited by the aforementioned means.
[0053] FIGURE 11 is a table illustrating parameters for a move command, in accordance with various aspects of the present disclosure. In the table of FIGURE 11, the DDR command pins: command address 0 (CA0) , command address 1 (CA1) , and command address 2 (CA2) are respectively set to low (L) , high (H) , and high (H) to trigger the command. Although a particular configuration is shown in FIGURE 11, any reserved for future use command may be used instead for the move operation. A copy operation may be similar. One possible pin assignment is illustrated in the table of FIGURE 11, where D in the command address 3 (CA3) pin indicates the direction between the row buffer and temporary buffer (e.g., meta data registers) . The values C5: C0 indicate the column address, and the lines S5: S0 indicate the transfer size in combination with the chip select (CS) pin, where X is ‘don’ t care’ .
[0054] FIGURE 12 is a flow diagram illustrating an example process 1200 performed, for example, by a computing device, in accordance with various aspects of the present disclosure. The example process 1200 is an example of performing data move and data copy operations in low power double data rate (LPDDR) dynamic random access memory (DRAM) devices.
[0055] As shown in FIGURE 12, in some aspects, the process 1200 may include transferring the data from a source memory address of a source location within the LPDDR memory device to a temporary buffer located within the LPDDR memory device. The source location initially stores the data to be transferred (block 1202) . The source memory address may comprise a bank group, a bank, and a row address of the source location.
[0056] The process 1200 may include transferring the data from the temporary buffer to a target memory address of a target location within the LPDDR memory device such that the data does not leave the LPDDR memory device during the transfer (block 1204) . In some aspects, transferring the data from the source location into the temporary buffer within the LPDDR memory device comprises transferring the data from a row buffer into one of a meta data register or a processor in memory (PIM) register file (PRF) . In some aspects, the source location and the target location are within a same memory bank, and the process further comprises opening a first row buffer connection to the source location in the LPDDR memory device, transferring the data from the source location into the temporary buffer within the LPDDR memory device, opening a second row buffer connection to the target location in the LPDDR memory device, and transferring the data from the temporary buffer to the target location. In other aspects, the source location and the target location are within a same memory bank, and the process further comprises opening a first row buffer connection to the source location in the LPDDR memory device, transferring the data from the source location into the temporary buffer within the LPDDR memory device, opening a second row buffer connection to the target location in the LPDDR memory device, and transferring the data from the temporary buffer to the target location.
[0057] FIGURE 13 is a block diagram showing an exemplary wireless communications system 1300, in which an aspect of the present disclosure may be advantageously employed. For purposes of illustration, FIGURE 13 shows three remote units 1320, 1330, and 1350, and two base stations 1340. It will be recognized that wireless communications systems may have many more remote units and base stations. Remote units 1320, 1330, and 1350 include integrated circuit (IC) devices 1325A, 1325B, and 1325C that include the disclosed LPDDR DRAM device. It will be recognized that other devices may also include the disclosed LPDDR DRAM device, such as the base stations, switching devices, and network equipment. FIGURE 13 shows forward link signals 1380 from the base stations 1340 to the remote units 1320, 1330, and 1350, and reverse link signals 1390 from the remote units 1320, 1330, and 1350 to the base stations 1340.
[0058] In FIGURE 13, remote unit 1320 is shown as a mobile telephone, remote unit 1330 is shown as a portable computer, and remote unit 1350 is shown as a fixed location remote unit in a wireless local loop system. For example, the remote units may be a mobile phone, a hand-held personal communication systems (PCS) unit, a portable data unit, such as a personal data assistant, a GPS enabled device, a navigation device, a set top box, a music player, a video player, an entertainment unit, a fixed location data unit, such as meter reading equipment, or other device that stores or retrieves data or computer instructions, or combinations thereof. Although FIGURE 13 illustrates remote units according to the aspects of the present disclosure, the disclosure is not limited to these exemplary illustrated units. Aspects of the present disclosure may be suitably employed in many devices, which include the disclosed LPDDR DRAM device.
[0059] FIGURE 14 is a block diagram illustrating a design workstation 1400 used for circuit, layout, and logic design of a semiconductor component, such as the LPDDR DRAM device disclosed above. The design workstation 1400 includes a hard disk 1401 containing operating system software, support files, and design software such as Cadence or OrCAD. The design workstation 1400 also includes a display 1402 to facilitate design of a circuit 1410 or a semiconductor component 1412, such as the LPDDR DRAM device. A storage medium 1404 is provided for tangibly storing the design of the circuit 1410 or the semiconductor component 1412 (e.g., the PLD) . The design of the circuit 1410 or the semiconductor component 1412 may be stored on the storage medium 1404 in a file format such as GDSII or GERBER. The storage medium 1404 may be a CD-ROM, DVD, hard disk, flash memory, or other appropriate device. Furthermore, the design workstation 1400 includes a drive apparatus 1403 for accepting input from or writing output to the storage medium 1404.
[0060] Data recorded on the storage medium 1404 may specify logic circuit configurations, pattern data for photolithography masks, or mask pattern data for serial write tools such as electron beam lithography. The data may further include logic verification data such as timing diagrams or net circuits associated with logic simulations. Providing data on the storage medium 1404 facilitates the design of the circuit 1410 or the semiconductor component 1412 by decreasing the number of processes for designing semiconductor wafers. Example Aspects
[0061] Aspect 1: A memory apparatus, comprising: a first memory subchannel comprising: a first plurality of memory banks; a first plurality of row buffers, each row buffer of the first plurality of row buffers coupled with a respective memory bank of the first plurality of memory banks; a first data bus coupled to the first plurality of row buffers; a first data mover (DM) coupled to the first data bus; and a first plurality of meta data registers coupled to the first data bus.
[0062] Aspect 2: The memory apparatus of Aspect 1, further comprising at least one processor in memory (PIM) execution unit (PEU) including a first PIM register file (PRF) , the at least one PEU coupled to the data bus of the first memory subchannel.
[0063] Aspect 3: The memory apparatus of Aspect 1 or 2, further comprising: a second memory subchannel comprising: a second plurality of memory banks; a second plurality of row buffers, each row buffer of the second plurality of row buffers coupled with a respective memory bank of the second plurality of memory banks; a second data bus coupled to the second plurality of row buffers; a second DM coupled to the second data bus; and a second plurality of meta data registers coupled to the second data bus; and a switch selectively coupled to the first data bus, the second data bus, the first DM, and the second DM.
[0064] Aspect 4: The memory apparatus of any of the preceding Aspects, further comprising at least one processor in memory (PIM) execution unit (PEU) , the at least one PEU including a PIM register file (PRF) the at least one PEU coupled to the data bus of the second memory subchannel.
[0065] Aspect 5: A method of moving data within a low power double data rate (LPDDR) memory device, comprising: transferring the data from a source memory address of a source location within the LPDDR memory device to a temporary buffer located within the LPDDR memory device, the source location initially storing the data to be transferred; and transferring the data from the temporary buffer to a target memory address of a target location within the LPDDR memory device such that the data does not leave the LPDDR memory device during the transfer.
[0066] Aspect 6: The method of Aspect 5, in which the source location and the target location are within a same memory bank, and the method further comprises: opening a first row buffer connection to the source location in the LPDDR memory device; transferring the data from the source location into the temporary buffer within the LPDDR memory device; opening a second row buffer connection to the target location in the LPDDR memory device; and transferring the data from the temporary buffer to the target location.
[0067] Aspect 7: The method of Aspect 5 or 6, further comprising repeating the opening the first row buffer connection, transferring the data to the source location, opening the second row buffer connection, and transferring the data from the temporary buffer in response to a first size of the data being larger than a second size of the temporary buffer.
[0068] Aspect 8: The method of any of the Aspects 5-7, in which the temporary buffer comprises a row buffer.
[0069] Aspect 9: The method of any of the Aspects 5-8, in which the temporary buffer comprises a meta data register.
[0070] Aspect 10: The method of any of the Aspects 5-9, in which the temporary buffer comprises a processor in memory (PIM) register file (PRF) .
[0071] Aspect 11: The method of any of the Aspects 5-10, further comprising: closing the first row buffer connection after transferring the data from the source location; and closing the second row buffer connection after transferring the data from the temporary buffer.
[0072] Aspect 12: The method of any of the Aspects 5-11, in which the source memory address comprises a bank group, a bank, and a row address of the source location.
[0073] Aspect 13: The method of any of the Aspects 5-12, in which the transferring the data from the source location into the temporary buffer within the LPDDR memory device comprises transferring the data from a row buffer into one of a meta data register or a processor in memory (PIM) register file (PRF) .
[0074] Aspect 14: The method of Aspect 5, in which the source location and the target location are in different memory banks, and the method further comprises: opening a first row buffer connection to the source location in the LPDDR memory device; opening a second row buffer connection to the target location in the LPDDR memory device; transferring the data from the source location into the temporary buffer within the LPDDR memory device; and transferring the data from the temporary buffer to the target location.
[0075] Aspect 15: The method of Aspects 5 or 14, further comprising repeating the opening the first row buffer connection, opening the second row buffer connection, transferring the data to the source location, and transferring the data from the temporary buffer in response to a size of the data being larger than the temporary buffer.
[0076] Aspect 16: The method of any of the Aspects 5 and 14-15, in which the temporary buffer comprises a row buffer.
[0077] Aspect 17: The method of any of the Aspects 5 and 14-16, in which the temporary buffer comprises a meta data register.
[0078] Aspect 18: The method of any of the Aspects 5 and 14-17, in which the temporary buffer comprises a processor in memory (PIM) register file (PRF) .
[0079] Aspect 19: The method of any of the Aspects 5 and 14-18, further comprising: closing the first row buffer connection after transferring the data from the source location; and closing the second row buffer connection after transferring the data from the temporary buffer.
[0080] Aspect 20: The method of any of the Aspects 5 and 14-19, in which the source memory address comprises a bank group, a bank, and a row address of the source location.
[0081] For a firmware and / or software implementation, the methodologies may be implemented with modules (e.g., procedures, functions, and so on) that perform the functions described. A machine-readable medium tangibly embodying instructions may be used in implementing the methodologies described. For example, software codes may be stored in a memory and executed by a processor unit. Memory may be implemented within the processor unit or external to the processor unit. As used, the term “memory” refers to types of long term, short term, volatile, non-volatile, or other memory and is not limited to a particular type of memory or number of memories, or type of media upon which memory is stored.
[0082] If implemented in firmware and / or software, the functions may be stored as one or more instructions or code on a computer-readable medium. Examples include computer-readable media encoded with a data structure and computer-readable media encoded with a computer program. Computer-readable media includes physical computer storage media. A storage medium may be an available medium that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can include random access memory (RAM) , read-only memory (ROM) , electrically erasable read-only memory (EEPROM) , compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage devices, or other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Disk and disc, as used, include compact disc (CD) , laser disc, optical disc, digital versatile disc (DVD) , floppy disk, and disc, where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0083] In addition to storage on computer-readable medium, instructions and / or data may be provided as signals on transmission media included in a communications apparatus. For example, a communications apparatus may include a transceiver having signals indicative of instructions and data. The instructions and data are configured to cause one or more processors to implement the functions outlined in the claims.
[0084] Although the present disclosure and its advantages have been described in detail, it should be understood that various changes, substitutions, and alterations can be made without departing from the technology of the disclosure as defined by the appended claims. For example, relational terms, such as “above” and “below” are used with respect to a substrate or electronic device. Of course, if the substrate or electronic device is inverted, above becomes below, and vice versa. Additionally, if oriented sideways, above, and below may refer to sides of a substrate or electronic device. Moreover, the scope of the present disclosure is not intended to be limited to the particular configurations of the process, machine, manufacture, composition of matter, means, methods, and steps described in the specification. As one of ordinary skill in the art will readily appreciate from the present disclosure, processes, machines, manufacture, compositions of matter, means, methods, or steps, presently existing or later to be developed that perform substantially the same function or achieve substantially the same result as the corresponding configurations described may be utilized according to the present disclosure. Accordingly, the appended claims are intended to include within their scope such processes, machines, manufacture, compositions of matter, means, methods, or steps.
[0085] Those of skill would further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the present disclosure may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0086] The various illustrative logical blocks, modules, and circuits described in connection with the disclosure may be implemented or performed with a general-purpose processor, a digital signal processor (DSP) , an application-specific integrated circuit (ASIC) , a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described. A general-purpose processor may be a microprocessor, but, in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0087] The steps of a method or algorithm described in connection with the present disclosure may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module may reside in RAM, flash memory, ROM, erasable programmable read-only memory (EPROM) , EEPROM, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0088] The previous description of the present disclosure is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the generic principles defined may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the present disclosure is not intended to be limited to the examples and designs described, but is to be accorded the widest scope consistent with the principles and novel features disclosed.
Claims
1.A memory apparatus, comprising:a first memory subchannel comprising:a first plurality of memory banks;a first plurality of row buffers, each row buffer of the first plurality of row buffers coupled with a respective memory bank of the first plurality of memory banks;a first data bus coupled to the first plurality of row buffers;a first data mover (DM) coupled to the first data bus; anda first plurality of meta data registers coupled to the first data bus.2.The memory apparatus of claim 1, further comprising at least one processor in memory (PIM) execution unit (PEU) including a first PIM register file (PRF) , the at least one PEU coupled to the data bus of the first memory subchannel.3.The memory apparatus of claim 1, further comprising:a second memory subchannel comprising:a second plurality of memory banks;a second plurality of row buffers, each row buffer of the second plurality of row buffers coupled with a respective memory bank of the second plurality of memory banks;a second data bus coupled to the second plurality of row buffers;a second DM coupled to the second data bus; anda second plurality of meta data registers coupled to the second data bus; anda switch selectively coupled to the first data bus, the second data bus, the first DM, and the second DM.4.The memory apparatus of claim 3, further comprising at least one processor in memory (PIM) execution unit (PEU) , the at least one PEU including a PIM register file (PRF) the at least one PEU coupled to the data bus of the second memory subchannel.5.A method of moving data within a low power double data rate (LPDDR) memory device, comprising:transferring the data from a source memory address of a source location within the LPDDR memory device to a temporary buffer located within the LPDDR memory device, the source location initially storing the data to be transferred; andtransferring the data from the temporary buffer to a target memory address of a target location within the LPDDR memory device such that the data does not leave the LPDDR memory device during the transfer.6.The method of claim 5, in which the source location and the target location are within a same memory bank, and the method further comprises:opening a first row buffer connection to the source location in the LPDDR memory device;transferring the data from the source location into the temporary buffer within the LPDDR memory device;opening a second row buffer connection to the target location in the LPDDR memory device; andtransferring the data from the temporary buffer to the target location.7.The method of claim 6, further comprising repeating the opening the first row buffer connection, transferring the data to the source location, opening the second row buffer connection, and transferring the data from the temporary buffer in response to a first size of the data being larger than a second size of the temporary buffer.8.The method of claim 6, in which the temporary buffer comprises a row buffer.9.The method of claim 6, in which the temporary buffer comprises a meta data register.10.The method of claim 6, in which the temporary buffer comprises a processor in memory (PIM) register file (PRF) .11.The method of claim 6, further comprising:closing the first row buffer connection after transferring the data from the source location; andclosing the second row buffer connection after transferring the data from the temporary buffer.12.The method of claim 6, in which the source memory address comprises a bank group, a bank, and a row address of the source location.13.The method of claim 6, in which the transferring the data from the source location into the temporary buffer within the LPDDR memory device comprises transferring the data from a row buffer into one of a meta data register or a processor in memory (PIM) register file (PRF) .14.The method of claim 5, in which the source location and the target location are in different memory banks, and the method further comprises:opening a first row buffer connection to the source location in the LPDDR memory device;opening a second row buffer connection to the target location in the LPDDR memory device;transferring the data from the source location into the temporary buffer within the LPDDR memory device; andtransferring the data from the temporary buffer to the target location.15.The method of claim 14, further comprising repeating the opening the first row buffer connection, opening the second row buffer connection, transferring the data to the source location, and transferring the data from the temporary buffer in response to a size of the data being larger than the temporary buffer.16.The method of claim 14, in which the temporary buffer comprises a row buffer.17.The method of claim 14, in which the temporary buffer comprises a meta data register.18.The method of claim 14, in which the temporary buffer comprises a processor in memory (PIM) register file (PRF) .19.The method of claim 14, further comprising:closing the first row buffer connection after transferring the data from the source location; andclosing the second row buffer connection after transferring the data from the temporary buffer.20.The method of claim 14, in which the source memory address comprises a bank group, a bank, and a row address of the source location.