Scheduling memory requests for linked memory devices
By adopting a multi-channel architecture and address generator method in the memory controller, the problems of memory access time and power consumption of computing system are solved, and efficient memory access is achieved.
Patent Information
- Application Number
- CN201880087934.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-12-21
- Filing Date
- 2018-09-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2038-09-19
AI Technical Summary
The prior art present challenges in reducing memory access time in computing system and reducing power consumption, especially in heterogeneous systems that deal with multiple types of IC integration.
By introducing a multi-channel architecture and an address generator in the memory controller, two addresses are generated for accessing non-continuous data in the same page, thereby simultaneously transmitting 64 bytes of data to improve memory access efficiency.
It realizes efficient memory access to the computing system, reduces the pause time of system memory, and reduces power consumption to interfaces and protocols.
Smart Images

Figure CN111656322B_ABST
Abstract
Description
Background of the Invention Background Art
[0002] Maintaining a relatively high level of performance generally requires fast access to stored data. Several types of data-intensive applications rely on fast access to data storage to provide reliable high performance to several local and remote programs and their users. A variety of computing devices utilize heterogeneous integration of integrated ICs of various types to provide system functionality. Various functions include audio / video (A / V) data processing, other high data parallel applications for medical and commercial fields, processing instructions of general instruction set architectures (ISAs), digital, analog, mixed signal, and radio frequency (RF) functions, etc. There are multiple options for placing processing nodes in system packages to integrate various types of ICs. Some examples are system on chip (SOC), multi-chip module (MCM), and system-in-package (SiP).
[0003] Regardless of which system package is chosen, the performance of one or more computing systems may depend on the processing node in several uses. In one example, the processing node is used in a mobile computing device that runs several different types of applications and may relay information to multiple users (both local and remote) at one time. In another example, the processing node is used in a desktop. In yet another example, the processing node is one of multiple processing nodes in a multi-socket server socket. The server is used to provide services to other computer programs in remote computing devices as well as computer programs in the server.
[0004] The memory hierarchy in each of the various computing systems described above transitions from relatively fast volatile memory, such as registers on the processor die and caches located on or connected to the processor die, to non-volatile and relatively slower memory, such as magnetic hard disks. The memory hierarchy presents challenges to maintaining high performance to meet the demands of running computer programs for fast access. One challenge is to reduce the amount of time in system memory, which is random access memory (RAM) located outside the cache subsystem but does not include non-volatile disk storage. Synchronous dynamic RAM (SDRAM) and other conventional memory technologies reduce system memory stall times due to limited bandwidth, but access latency is not improved using these technologies. In addition, a large amount of on-die area and power consumption is used to support interfaces and protocols to access data stored in system memory.
[0005] In view of the above, efficient methods and systems for performing efficient memory access for a computing system are desired. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Advantages of the methods and mechanisms described herein may be better understood by referring to the following description in conjunction with the accompanying drawings, in which:
[0007] Figure 1 is a block diagram of one embodiment of a computing system.
[0008] Figure 2 is a block diagram of one implementation of a memory controller.
[0009] Figure 3 is a flow chart of one embodiment of a method for performing efficient memory access for a computing system.
[0010] Figure 4 is a flow chart of one embodiment of a method for performing efficient memory access for a computing system.
[0011] Although the present invention is susceptible to various modifications and alternative forms, specific embodiments are shown in the drawings by way of example and are described in detail herein. It should be understood that the drawings and their detailed description are not intended to limit the invention to the specific forms disclosed, but on the contrary, the invention will cover all modifications, equivalents and alternatives falling within the scope of the invention as defined by the appended claims. DETAILED DESCRIPTION
[0012] In the following description, many specific details are set forth to provide a thorough understanding of the methods and mechanisms presented herein. However, it will be appreciated by those skilled in the art that various embodiments may be implemented without these specific details. In some cases, well-known structures, components, signals, computer program instructions, and techniques are not shown in detail to avoid confusing the methods described herein. It should be understood that in order to make the description clear and simple, the elements shown in the drawings may not be drawn to scale. For example, the size of some of the elements may be enlarged relative to other elements.
[0013] Various systems, devices, methods and computer-readable media for performing efficient memory access on a computing system are disclosed. In various embodiments, a computing system includes one or more clients for processing an application. An example of a client is a general-purpose central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), an input / output (I / O) device, etc. A memory controller is configured to transmit traffic between the memory controller and two channels connected to each memory device. In some embodiments, one or more of the two memory devices is one of a variety of random access memories (RAMs) on a dual in-line memory module (DIMM). In other embodiments, one or more of the two memory devices is a planar mounted RAM device that is inserted into or soldered to a motherboard. In yet other embodiments, one or more of the two memory devices is a three-dimensional integrated circuit (3DIC). In an embodiment, a command processor in a memory controller converts a memory request received from a client into a command to be processed by a memory device of the selected type.
[0014] In an embodiment, the client sends a 64-byte memory request with an indication, and the indication specifies that there are two 32-byte requests with non-continuous data as the target in the same page. The memory controller generates two addresses. The memory controller sends a single command and two addresses to two channels to access the data in the same page. In one embodiment, the memory controller sends two addresses generated separately or a part thereof to two channels. In some embodiments, an address is offset relative to another address in the two generated addresses. In some embodiments, a single command with two addresses accesses non-continuous data in the same page. In other embodiments, a single command with two addresses accesses continuous data in the same page. Therefore, adjacent data (in the same page) as continuous data or non-continuous data is accessed simultaneously. Therefore, the memory controller will not transmit 64 bytes for a single 32-byte memory request, and discard 32 bytes in 64 bytes, which is inefficient. Instead, the memory controller transmits 64 bytes for two 32-byte memory requests for accessing data in the memory address range (such as in a memory page).
[0015] refer to Figure 1, a generalized block diagram of one embodiment of a computing system 100 is shown. As shown, clients 110 and 112 send memory requests to memory controllers 130A and 130B via data structure 120. As shown, each memory controller has a single memory channel capable of sending two addresses. For example, memory controller 130A includes memory channel 140A having address generator 142A and address generator 144A. Similarly, memory controller 130B includes memory channel 140B having address generator 142B and address generator 144B. Memory controller 130A transmits commands, addresses, and data on channels 152A and 154A to memory devices 160A and 162A. Memory controller 130B transmits commands, addresses, and data on channels 152B and 154B to memory devices 160B and 162B.
[0016] For ease of illustration, input / output (I / O) interfaces for I / O devices, power managers, and any links and interfaces for network connections are not shown in computing system 100. In some embodiments, a component of computing system 100 is a single die on an integrated circuit (IC), such as a system on a chip (SOC). In other embodiments, the component is a single die in a system-in-package (SiP) or a multi-chip module (MCM). In some embodiments, clients 110 and 112 include one or more of a central processing unit (CPU), a graphics processing unit (GPU), a multimedia engine hub, and the like. Each of clients 110 and 112 is one of a variety of computing resources capable of processing applications and generating memory requests.
[0017] When one of the clients 110-112 is a central processing unit (CPU), in some embodiments, each of the one or more processor cores in the CPU includes a circuit for executing instructions according to a given selected instruction set architecture (ISA). In various embodiments, each of the processor cores in the CPU includes a superscalar, multi-threaded micro-architecture for processing instructions of a given ISA. In an embodiment, when one of the clients 110-112 is a graphics processing unit (GPU), it includes a high parallel data micro-architecture with a large number of parallel execution paths. In one embodiment, the micro-architecture uses a single instruction multiple data (SIMD) pipeline for parallel execution paths. When one of the clients 110-112 is a multimedia engine, it includes a processor for processing audio data and visual data of a multimedia application. Other instances of the processing unit generating memory requests for the clients 110-112 are possible and conceivable.
[0018] In various embodiments, communication fabric 120 transfers traffic back and forth between clients 110 and 112 and memory controllers 130A and 130B. Data fabric 120 includes interfaces for supporting corresponding communication protocols. In some embodiments, communication fabric 120 includes queues for storing requests and responses, selection logic for arbitrating between received requests before sending them across the internal network, logic for building and decoding packets, and logic for selecting routes for packets.
[0019] In various embodiments, the memory controllers 130A-130B receive memory requests from the clients 110-112 via the communication fabric 120, convert the memory requests into commands, and send the commands to one or more of off-chip disk storage (not shown) and system memory, which is implemented as one of a variety of random access memories (RAMs) in the memory devices 160A, 162A, 160B, and 162B. The memory controller 130 also receives responses from the memory devices 160A, 162A, 160B, and 162B and the disk storage, and sends the responses to the corresponding sources of the clients 110-112.
[0020] In some embodiments, the address space of the computing system 100 is divided between at least the client 110-112 and one or more other components (such as input / output peripherals (not shown) and other types of computing resources). A memory mapping is maintained to determine which address is mapped to which component, and therefore a memory request for a specific address should be routed to which of the clients 110-112. One or more of the clients 110-112 include a cache memory subsystem to reduce the memory latency of the corresponding processor core. In addition, in some embodiments, before accessing the memory devices 160A, 162A, 160B, and 162B, a shared cache memory subsystem is used by the processor core as a last-level cache (LLC). As used herein, the term "memory access" refers to performing a memory read request or a memory write request operation, and if the request data of the corresponding request address resides in the cache, the memory read request or the memory write request operation causes a cache hit. Alternatively, if the requested data does not reside in the cache, the memory access request causes a cache miss.
[0021] In various embodiments, the system memory includes a multi-channel memory architecture. This type of architecture increases the speed at which data is transferred to memory controllers 130A and 130B by adding more communication channels between communication channels (such as channels 152A, 154A, 152B, and 154B). In an embodiment, the multi-channel architecture utilizes multiple memory modules and a motherboard and / or card capable of supporting multiple channels.
[0022] In some embodiments, the computing system 100 utilizes one of a variety of dynamic RAMs (DRAMs) to provide system memory. In other embodiments, the computing system 100 utilizes a three-dimensional integrated circuit (3D IC) to provide system memory. In such embodiments, 3D integrated DRAM provides low latency interconnects and additional on-chip memory storage to reduce off-chip memory accesses. Other memory technologies for system memory using row-based access schemes including one or more row buffers or other equivalent structures are possible and may be conceivable. Examples of other memory technologies include phase change memory, spin torque transfer resistive memory, memristors, etc.
[0023] In various embodiments, the components within memory controller 130B have the same functions as the components in memory controller 130A. In some embodiments, control units 132A and 132B within memory controllers 130A and 130B convert received memory requests into transactions (such as read / write transactions and activation and precharge transactions). As used herein, "transactions" are also referred to as "commands". In various embodiments, each of channels 152A, 154A, 152B, and 154B is a link that includes a command bus, an address bus, and a data bus for a plurality of memory banks within a corresponding one of memory devices 160A, 162A, 160B, and 162B.
[0024] In various embodiments, memory devices 160A, 162A, 160B, and 162B include multiple memory banks, each of which has multiple memory array memory banks. Each of the memory banks includes multiple rows and a row buffer. Each row buffer stores data corresponding to the rows accessed in the multiple rows in the memory array memory bank. The accessed rows are identified by the DRAM address in the received memory request. Typically, each row stores a data page. The size of the page is selected based on design considerations. This page size can be one kilobyte (1KB), four kilobytes (4KB), or any other size.
[0025] Memory channels 140A and 140B are interfaced with PHY 150A and 150B. In some embodiments, each of the physical interfaces PHY 150A and 150B transmits a command stream from memory controllers 130A and 130B to memory devices 160A, 162A, 160B, and 162B at a given timing. The protocol determines the values used for information transmission, such as the number of data transmissions per clock cycle, signal voltage levels, signal timing, signal and clock phases, and clock frequencies. In some embodiments, each of PHY 150A and 150B includes a state machine for the initialization and calibration sequences specified in the protocol.
[0026] In addition, in an embodiment, each of PHYs 150A and 150B includes self-test, diagnostic, and error detection and correction hardware. Protocol examples for the respective interfaces between PHYs 150A and 150B and memory devices 160A, 162A, 160B, and 162B include DDR2 SDRAM, DDR3 SDRAM, GDDR4 (Graphics Double Data Rate, version 4) SDRAM, GDDR5 SDRAM, and GDDR6 SDRAM.
[0027] As shown, memory channel 140A includes address generators 142A and 144A, and memory channel 140B includes address generators 142B and 144B. In various embodiments, address generators 142A and 144A convert memory request addresses received by memory controller 130A into values identifying a given rank, a given bank, and a given row in one of memory devices 160A and 162A. Although two address generators are shown, in other embodiments, another number of address generators are included in memory controller 130A.
[0028] In some embodiments, the address generator 144A generates the second address as an offset relative to the first address generated by the address generator 142A. In one embodiment, the address generator 144A uses the same identifier as the first address generated by the address generator 142A in the second address to identify a given rank and a given storage body and a given row in one of the memory devices 160A and 162A. In addition, in the embodiment, the first address identifies the starting byte of the requested first data in the identified row, and the second address identifies the starting byte of the requested second data that does not overlap with the first data. In the embodiment, the second data is continuous with the first data in the identified row. In other embodiments, the second data is discontinuous with the first data in the identified row. Therefore, the single memory controller 130A transmits data and commands to two channels 152A and 154A, while also supporting simultaneous access to the data in the same row for two different requests.
[0029] In various embodiments, when control unit 132A determines that each of the first memory request and the second memory request targets data within a given memory address range, control unit 132A stores an indication that the given memory access command services each of the first memory request and the second memory request that is different from the first memory request. In an embodiment, the given memory address range is a range of memory pages in one of memory devices 160A and 162A. In some embodiments, in response to determining that the given memory access command has been completed, control unit 130A marks each of the first memory request and the second memory request as completed.
[0030] refer to Figure 2 , a generalized block diagram of one embodiment of a memory controller 200 is shown. In the illustrated embodiment, the memory controller 200 includes: an interface 210, the interface 210 is used to connect with a computing resource through a communication structure; a queue 220, the queue 220 is used to store received memory access requests and received responses; a control unit 250; and an interface 280, the interface 280 is used to connect with a memory device through at least one physical interface and at least two channels. Each of the interfaces 210 and 280 supports a corresponding communication protocol.
[0031] In an embodiment, queue 220 includes a read queue 232 for storing received read requests and a separate write queue 234 for storing received write requests. In other embodiments, queue 220 includes a unified queue for storing both memory read requests and memory write requests. In one embodiment, queue 220 includes a queue 236 for storing the scheduled memory access requests selected from read queue 232, write queue 234 or unified queue (if one queue is used). Queue 236 is also referred to as pending queue 236. In some embodiments, control register 270 stores an indication of the current mode. For example, an off-chip memory data bus and a memory device support a read mode or a write mode at a given time. Therefore, traffic is routed in a given single direction during the current mode and changes direction when the current mode ends.
[0032] In some embodiments, read scheduler 252 includes arbitration logic for selecting read requests out of order from read queue 232. Read scheduler 252 schedules out-of-order issuance of requests stored in read queue 232 to memory devices based on quality of service (QoS) or other priority information, age, process or thread identifier (ID), and relationship to other stored requests (such as targeting the same memory channel, targeting the same rank, targeting the same memory bank, and / or targeting the same page). Write scheduler 254 includes similar selection logic for write queue 234. In an embodiment, response scheduler 256 includes similar logic for issuing responses received from memory devices out of order to computing resources based on priority.
[0033] In various embodiments, command processor 272 converts received memory requests into one or more transactions (or commands) (such as read / write transactions and activation and precharge transactions). In some embodiments, commands are stored in queues 232-236. In other embodiments, a single set of queues is used. As shown, control unit 250 includes address generators 260 and 262. In various embodiments, address generators 260 and 262 convert memory request addresses received by memory controller 130A into values identifying a given rank, a given storage body, and a given row in one of the memory devices connected to memory controller 200. Although two address generators are shown, in other embodiments, another number of address generators are included in control unit 250.
[0034] In some embodiments, address generator 262 generates the second address as an offset relative to the first address generated by address generator 260. In one embodiment, address generator 262 uses the same identifier in the second address as the first address for identifying a given rank and a given bank and a given row within one of the memory devices. In addition, in an embodiment, the first address identifies the starting byte of the requested data in the identified row, and the second address identifies the starting byte of the requested data that does not overlap with the first data and is not continuous with the first data in the identified row. Thus, a single memory controller 200 transmits data and commands to at least two channels while also supporting simultaneous access to non-contiguous data.
[0035] In various embodiments, when control unit 250 determines that each of the first memory request and the second memory request targets data within a given memory address range, control unit 250 stores an indication that a given memory access command stored in queue 220 services each of the first memory request and a second memory request different from the first memory request stored in queue 220. In an embodiment, the given memory address range is an address range of memory pages in one of the memory devices. In some embodiments, in response to determining that the given memory access command has been completed, control unit 250 marks each of the first memory request and the second memory request as completed.
[0036] In some embodiments, control register 270 stores an indication of the current mode. For example, the memory data bus and the memory device support read mode or write mode at a given time. Therefore, traffic is routed along a given single direction during the current mode and changes direction when the current mode changes after the data bus inversion delay. In various embodiments, control register 270 stores a threshold number of read requests (read burst length) to be sent during the read mode. In some embodiments, control register 270 stores the weight of the standard used by the selection algorithm in read scheduler 252 and write scheduler 254 for selecting the request to be issued in queue 232-236.
[0037] Similar to computing system 100, connecting two memory channels to memory controller 200 is called "ganging". Each of the at least two channels connected to memory controller 200 through a physical interface receives the same command to access data in the same page within the selected memory device. In addition, each channel has its own address. For example, a first channel receives a first address from address generator 260, and a second channel different from the first channel receives a second address from address generator 262. In an embodiment, the address generated by address generators 260 and 262 is a column address for DRAM. In various embodiments, memory controller 200 accesses non-contiguous data simultaneously.
[0038] In some embodiments, the memory controller 200 supports the GDDR6 DRAM protocol. In such embodiments, the interface 280 supports read transactions and write transactions for each channel (in two channels) with a width of 16 bits (2 bytes) and a burst length of 16. Two linked 16-bit wide channels provide equivalent to 32-bit (4-byte) wide channels. For 64-byte requests, a 32-bit (4-byte) wide equivalent channel provided by two channels and with a burst length of 16 will transmit 64 bytes for servicing 64-byte memory requests. The two channels are linked, and the memory controller 200 manages two 16-bit wide interfaces.
[0039] In the embodiment using the GDDR6 protocol, the control unit 250 manages the 64-byte interface as two independent 32-byte interfaces for 32-byte requests. In the embodiment, the control unit 250 sends a command to open the same page simultaneously across two 16-bit channels. For example, the control unit 250 issues an activation command to each of the two channels simultaneously, and issues a memory access command to each of the two channels simultaneously, but the control unit 250 sends two different addresses to access the opened page simultaneously and independently by address generators 260 and 262. Simultaneous access is also used as adjacent data (in the same page) of non-continuous data. Therefore, the memory controller 200 will not transmit 64 bytes for a single 32-byte memory request, and discards 32 bytes in 64 bytes, which is inefficient. Instead, the memory controller 200 transmits 64 bytes for two 32-byte memory requests for non-continuous data in the access memory address range (such as in a memory page).
[0040] In some embodiments, control unit 250 determines when two 32-byte memory requests access non-contiguous data within the same page in one of the memory devices. In other embodiments, a client (such as a GPU) determines when two 32-byte memory requests access the same page in one of the memory devices. The client sends a 64-byte memory request with an indication that there are two 32-byte requests targeting non-contiguous data within the same page. In an embodiment, when control unit 250 issues a 64-byte command, the address from address generator 262 is ignored.
[0041] Reference Figure 3 , showing an embodiment of a method 300 for performing efficient memory access on a computing system. For the purpose of discussion, the following embodiments (and Figure 4 300. However, it should be noted that in various embodiments of the described method, one or more of the described elements are performed concurrently in a different order than shown, or are omitted entirely. Other additional elements may also be performed as needed. Any of the various systems or devices described herein is configured to implement method 300.
[0042] One or more clients execute computer programs or software applications. The client determines within the cache memory subsystem that a given memory access request misses, and sends the memory access request to the system memory through the memory controller. The memory request is stored when the memory request is received (frame 302). If the received memory request does not request data with a data size less than a size threshold (the "no" branch of conditional frame 304), the memory request is converted into a command (frame 310). In some embodiments, the memory request requests data with a size of 64 bytes and 32 bytes. In an embodiment, the size threshold is set to 64 bytes. Therefore, a memory request requesting data with a data size of 64 bytes does not request data with a data size less than the size threshold.
[0043] In various embodiments, a memory request (such as a memory read request) is converted into one or more commands based on the memory being accessed. For example, control logic within the DRAM performs complex transactions (such as an activate (open) transaction and precharges data and control lines within the DRAM), once to access the identified row and once to put the modified contents stored in the row buffer back to the identified row during a close transaction. Each of the different DRAM transactions (such as activate / open, column access, read access, write access, and precharge / close) has a different corresponding delay.
[0044] Memory access commands are scheduled to be issued to service the memory requests (block 312). In some embodiments, the memory access commands are marked as out-of-order issue based at least on the priority and target of the corresponding memory request. In other embodiments, the memory requests are scheduled prior to conversion to commands. Thus, the memory controller supports out-of-order issue of memory requests.
[0045] If the data size of the received memory request is less than the size threshold (the "yes" branch of conditional block 304), and the first memory request and the second memory request do not target the same given address range (the "no" branch of conditional block 306), the method 300 moves to block 310, where the memory request is converted to a command. However, if the data size of the received memory request is less than the size of the memory data bus (the "yes" branch of conditional block 304), and the first memory request and the second memory request target the same given address range (the "yes" branch of conditional block 306), an indication that the given memory access command services each of the first memory request and the second memory request is stored (block 308). Thereafter, the method 300 moves to block 310, where the memory request is converted to a command.
[0046] Steering Figure 4, one embodiment of a method 400 for performing efficient memory access on a computing system is shown. An indication that a given memory access command services each of a first memory request and a second memory request is detected (block 402). The given memory access command is sent to a memory device (block 404). For example, scheduling logic in a memory controller selects the given memory access command to be issued to the memory device based on priority level, age, etc.
[0047] The memory controller sends a first address to the memory device that points to a first location in the memory device storing the first data (block 406). The memory controller sends a second address to the memory device that points to a second location in the memory device storing second data that is non-contiguous with the first data (block 408). In response to determining that the given memory access command has been completed, each of the first memory request and the second memory request is marked as completed (block 410).
[0048] In various embodiments, the previously described methods and / or mechanisms are implemented using program instructions of a software application. The program instructions describe the behavior of the hardware in a high-level programming language (such as C). Alternatively, a hardware design language (HDL) (such as Verilog) can be used. The program instructions are stored on a non-transitory computer-readable storage medium. Many types of storage media are available. During use, a computing system can access the storage medium to provide the program instructions and accompanying data to the computing system for program execution. The computing system includes at least one or more memories and one or more processors configured to execute the program instructions.
[0049] It should be emphasized that the above embodiments are only non-limiting examples of implementation. Once the above disclosure is fully understood, many changes and modifications will become apparent to those skilled in the art. It is intended that the appended claims be interpreted as including all such changes and modifications.
Claims
1. A memory controller comprising: a first interface for receiving a memory request, the memory request comprising a single client memory request from a client, the single client memory request including an indication that the single client memory request targets two non-contiguous data addresses, the two non-contiguous data addresses corresponding to a first memory request and a second memory request; A second interface, wherein the second interface comprises: a command bus, the command bus being used to send a memory access command corresponding to the memory request to a memory device; a first address bus for sending addresses to the memory device; and a second address bus, the second address bus being used to send an address to the memory device; control logic comprising circuitry, wherein in response to determining that a given memory access command corresponding to the single client memory request is scheduled to be issued at a given point in time and detecting an indication that the memory access command services both the first memory request and the second memory request, the control logic is configured to, at the given point in time: sending the memory access command to the memory device via the command bus; sending a first address corresponding to the first memory request on the first address bus, wherein the first address points to a first location in the memory device storing first data; and A second address corresponding to the second memory request is sent on the second address bus, wherein the second address points to a second location in the memory device where second data is stored.
2. The memory controller of claim 1, wherein the control logic is further configured to store the indication in response to determining that each of the first memory request and the second memory request targets data within a given memory address range.
3. The memory controller of claim 2, wherein the given memory address range is a range of memory pages.
4. The memory controller of claim 1, wherein the first memory request targets data that is non-contiguous with data targeted by the second memory request.
5. The memory controller of claim 1, wherein: The second address is offset relative to the first address; and The memory access command accesses each of the first data and the second data. 6 . The memory controller of claim 1 , wherein each of the first memory request and the second memory request targets data having a same size.
7. The memory controller of claim 1, wherein the second interface further comprises a data bus for transmitting data between the memory controller and the memory device, and wherein each of the first data and the second data has a size less than a size threshold.
8. The memory controller of claim 7, wherein the first data and the second data are transmitted simultaneously on the data bus.
9. The memory controller of claim 1, wherein the control logic is further configured to mark each of the first memory request and the second memory request as completed in response to determining that the memory access command is completed.
10. A method comprising: receiving, by a first interface comprising circuitry, one or more client memory requests, the one or more client memory requests comprising a single client memory request from a client, the single client memory request including an indication that the single client memory request targets two non-contiguous data addresses, the two non-contiguous data addresses corresponding to a first memory request and a second memory request; In response to determining that a given memory access command corresponding to the single client memory request is scheduled to be issued at a given point in time and detecting an indication that the memory access command services both the first memory request and the second memory request, by a control unit at the given point in time: sending the memory access command to the memory device via a command bus; sending a first address corresponding to the first memory request on a first address bus, wherein the first address points to a first location in the memory device storing first data; and A second address corresponding to the second memory request is sent on a second address bus, wherein the second address points to a second location in the memory device where second data is stored.
11. The method of claim 10, further comprising storing the indication in response to determining that each of the first memory request and the second memory request targets data within a given memory address range.
12. The method of claim 11, wherein the given memory address range is a range of memory pages.
13. The method of claim 10, wherein: The second address is offset relative to the first address; and The memory access command accesses each of the first data and the second data.
14. The method of claim 10, wherein each of the first memory request and the second memory request targets data having a same size.
15. The method of claim 10, wherein the first interface is included in a memory controller, wherein the memory controller further includes a second interface, wherein the second interface includes a data bus for transmitting data between the memory controller and the memory device, and wherein each of the first data and the second data has a size less than a size threshold.
16. The method of claim 10, further comprising marking each of the first memory request and the second memory request as completed in response to determining that the given memory access command has completed.
17. A computing system comprising: a processor configured to generate the single client memory request; as well as A memory controller as claimed in any one of claims 1 to 9.
Citation Information
Patent Citations
Apparatus, system, and method for coalescing parallel memory requests
US7492368B1