Dynamic Memory Reconfiguration

By dynamically allocating memory pages across a configurable subset of channels based on an allocation mode, the method addresses inefficiencies in conventional multi-channel memory systems, improving processing efficiency and reducing latency.

JP2025521383APending Publication Date: 2025-07-10ADVANCED MICRO DEVICES INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024535691
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-30
Filing Date
2023-06-30
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Conventional multi-channel memory systems utilize fixed channel configurations for data interleaving, which may not optimize performance for diverse memory access patterns, leading to inefficiencies and increased latency.

Method used

A method for selectively allocating memory pages across a configurable subset of channels based on an allocation mode, allowing for both unified memory architecture shared by multiple parallel processing unit chips and private non-uniform memory architecture for single chips, using a kernel mode driver to manage memory access.

Benefits of technology

This approach enhances processing efficiency by reducing interference between memory pools and shortening latency through flexible memory allocation, optimizing performance for different types of memory access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025521383000001_ABST
    Figure 2025521383000001_ABST
Patent Text Reader

Abstract

A processing system is disclosed that includes a parallel processing unit that selectively allocates pages of memory to interleave across a configurable subset of channels based on an allocation mode. In some embodiments, in a first mode, pages of memory are allocated to and interleaved across a plurality of channels, and in a second mode, pages of memory are allocated to and interleaved across a subset of the plurality of channels.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] To improve overall processing efficiency, a processing system typically employs a multi-channel high-bandwidth memory such as a multi-channel Dynamic Random Access Memory (DRAM). For example, such multi-channel memories are often implemented within a processing system such that multiple memory dies are accessible in parallel by a host processor within the system. This multi-channel parallel access typically increases the amount of data that the system can read or write within a given period of time, and in turn, enables reduced processing latency and increased overall system performance as well as increased total DRAM capacity.

[0002] A multi-channel memory system is typically configured to store data across multiple memory devices according to an interleaving pattern. Some conventional multi-channel memory systems utilize only a fixed number of channels, within which data is interleaved based on an interleaving pattern, and according to the interleaving pattern, data is sequentially stored across the memory devices of the multi-channel memory system.

Summary of the Invention

Means for Solving the Problems

[0003] In the embodiments described herein, techniques are provided for selectively allocating memory pages to interleave across a configurable subset of channels of a processing system including a parallel processing unit based on an allocation mode. In one exemplary embodiment, a method executed by a computer may include allocating memory pages to be interleaved across a plurality of channels in a first mode and allocating memory pages to be interleaved across a subset of the plurality of channels in a second mode.

[0004] In some embodiments, the first mode includes allocating memory pages to an integrated memory architecture shared by a plurality of parallel processing unit chips, and the second mode includes allocating memory pages to a heterogeneous memory architecture that is private to a single parallel processing unit chip. The method may further include selecting the first mode in response to the memory pages being shared by a plurality of parallel processing unit chips, and selecting the second mode in response to the memory pages being accessed by a single parallel processing unit chip. Selecting the first mode may be in response to the memory pages including one or more of a texture and a vertex buffer. Selecting the second mode may be in response to the memory pages including one or more of a thread private memory and a render target.

[0005] In some embodiments, the method further includes indicating the use of the first mode in a first physical address aperture and indicating the use of the second mode in a second physical address aperture. The plurality of channels corresponds to a plurality of memory devices shared by a plurality of parallel processing unit chips, and a subset of the plurality of channels may correspond to memory devices that are private to a parallel processing unit chip.

[0006] In another exemplary embodiment, the apparatus includes a plurality of memory devices and a plurality of parallel processing unit chips configured to access data stored in the plurality of memory devices via a plurality of channels. The apparatus further includes a kernel mode driver configured to allocate memory pages to be interleaved across the plurality of channels in a first mode and to allocate memory pages to be interleaved across a subset of the plurality of channels in a second mode.

[0007] In some embodiments, the first mode includes allocating a memory page to an integrated memory architecture shared by a plurality of parallel processing unit chips, and the second mode includes allocating a memory page to a heterogeneous memory architecture that is private to a single parallel processing unit chip. In some embodiments, the kernel mode driver is further configured to select the first mode in response to the memory page being shared by the parallel processing unit chips, and to select the second mode in response to the memory page being accessed by a single parallel processing unit chip.

[0008] The kernel mode driver may select the first mode in response to the memory page including one or more of a texture and a vertex buffer. In some embodiments, the kernel mode driver may select the second mode in response to the memory page including one or more of thread private memory and a render target. In some embodiments, the kernel mode driver may indicate the use of the first mode in a first physical address aperture and the use of the second mode in a second physical address aperture.

[0009] In another exemplary embodiment, the method includes, in a first mode, interleaving a first memory page across a plurality of channels shared by a plurality of parallel processing unit chips, and, in a second mode, interleaving a second memory page across a subset of a plurality of channels shared by a subset of the plurality of parallel processing unit chips. In some embodiments, the first mode includes allocating the first memory page to an integrated memory architecture shared by a plurality of parallel processing unit chips, and the second mode includes allocating the second memory page to a heterogeneous memory architecture that is private to a subset of the plurality of parallel processing unit chips.

[0010] This disclosure can be better understood by referring to the accompanying drawings, and its various features and advantages can become apparent to those skilled in the art. The use of the same reference numerals in different drawings indicates similar or identical items.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Modes for Carrying Out the Invention

[0012] Parallel processing units such as graphics processing units (GPUs) include front end (FE) hardware for processing tasks such as fetching commands, processing jitter, performing geometry operations, and performing ray tracing. FE hardware typically includes a command fetcher such as a first-in-first-out (FIFO) buffer for holding fetched commands, a queue, and a scheduler for scheduling commands from the command buffer for execution on a shader engine within the GPU. The shader engine is executed using one or more processors and one or more arithmetic logic units (ALUs) and executes the commands provided by the FE hardware. Results generated by the shader, such as the values of shaded pixels, are output to one or more caches that store frequently used information, which is also stored in the corresponding memory. Thus, the GPU includes memory or registers, caches, ports, and interfaces for holding results between these entities. Information flows from the FE hardware to the memory via a path that includes a command bus for carrying commands from the FE hardware to the shader engine, a cache for storing the output from the shader engine, and a memory channel for transferring the cached information to the memory.

[0013] Some GPUs include multiple instances of FE hardware and corresponding shader engines (each such instance is referred to herein as a "partition") that access memory collectively via multiple memory channels. In some cases, the partitions are implemented on separate dies and are referred to herein as parallel processing unit chips. Certain data, such as texture and vertex buffers, is shared by all partitions, while other data, such as render targets and thread private memory, is used by only a single partition.

[0014] Figures 1-7 illustrate techniques for selectively allocating pages of memory to interleave across a configurable subset of channels of a processing system that includes a parallel processing unit, based on an allocation mode. In some embodiments, in a first mode, pages of memory are allocated across and interleaved among a plurality of channels, and in a second mode, pages of memory are allocated across and interleaved among a subset of the plurality of channels. For example, in some embodiments, a parallel processing unit accesses memory via 48 channels that are divided into three partitions each accessing memory via 16 channels. In the first mode, a kernel mode driver of the processing system allocates pages of memory that are interleaved across all 48 channels of the parallel processing unit. In the second mode, the kernel mode driver allocates a page of memory to 16 channels corresponding to a single partition and interleaves that page across those 16 channels.

[0015] In some embodiments, the first mode specifies allocating memory pages to a unified memory architecture (UMA) shared by all partitions of the parallel processing unit, and the second mode specifies allocating memory pages to a private non-uniform memory architecture (NUMA) of a single partition of the parallel processing unit. By flexibly allocating and interleaving memory pages across a subset of channels, the processing system maintains separate memory pools that do not interfere with each other. Further, by allocating memory pages only to partitions that need access, the kernel mode driver reduces transfers from one partition to the next and shortens latency.

[0016] In some embodiments, the kernel mode driver selects a first mode in response to a memory page being shared by all partitions of the parallel processing unit. For example, in some embodiments, the kernel mode driver selects the first mode in response to the memory page including a texture or a vertex buffer. In other embodiments, the kernel mode driver selects the first mode in response to the memory page including one or more of a microcode engine in a central processing unit (CPU) that provides instructions to the parallel processing unit, context save and restore data, and a platform security processor (PSP) state. Conversely, the kernel mode driver selects a second mode in response to a memory page being accessed by a single partition of the parallel processing unit. For example, the kernel mode driver selects the second mode in response to the memory page including thread private memory or a render target.

[0017] In some embodiments, the physical address of the memory page includes additional address bits (referred to as an aperture) that are used to indicate the mode selected by the kernel mode driver. For example, in some embodiments, the kernel mode driver indicates the use of the first mode in a first physical address aperture and the use of the second mode in a second physical address aperture.

[0018] FIG. 1 is a block diagram of a processing system 100 that performs dynamic memory reconfiguration for a parallel processing unit such as a graphics processing unit (GPU) 102 according to some embodiments. In various embodiments, the GPU 102 is associated with resources such as a conventional CPU, a conventional graphics processing unit (GPU), and combinations thereof, and performs functions and computations associated with accelerating graphics processing tasks, data parallel tasks, and nested data parallel tasks in an accelerated manner. It is a parallel processor that includes any cooperating collection of hardware and / or software.

[0019] Processing system 100 includes one or more central processing units (CPUs) 150. Although one CPU 150 is shown in FIG. 1, some embodiments of processing system 100 include more CPUs. A scalable data fabric (SDF) 170 supports data flow between endpoints within processing system 100. Some embodiments of SDF 170 may be implemented as a peripheral component interface (PCI) bus, a PCI-E bus, or a peripheral component interface (PCI) physical layer, a memory controller, a universal serial bus (USB) hub, a GPU 102, and a computing and execution unit including CPU 150, and other types of buses that support data flow between connection points such as other endpoints. The components of processing system 100 may be implemented as hardware, firmware, software, or any combination thereof. It should be understood that processing system 100 may include one or more software components, hardware components, and firmware components in addition to or different from those shown in FIG. 1. For example, processing system 100 may further include one or more input interfaces, non-volatile storage, one or more output interfaces, a network interface, and one or more displays or display interfaces. Processing system 100 includes, for example, servers, desktop computers, laptop computers, tablet computers, mobile phones, game consoles, and the like.

[0020] In various embodiments, the CPU 150 is connected to a system memory 135, such as a dynamic random access memory (DRAM), that includes a plurality of memory devices 135-0, 135-1, 135-2, ..., 135-N, each of which is accessed via a corresponding channel 165-0, 165-1, 165-2, ..., 165-N through the SDF 170. In various embodiments, the system memory 135 may be implemented using other types of memory, including static random access memory (SRAM), non-volatile RAM, and the like. In the illustrated embodiment, the CPU 150 communicates with the system memory 135 (also referred to as the memory 135) via the SDF 170 and communicates with the GPU 102. However, some embodiments of the processing system 100 include a GPU 102 that communicates with the CPU 102 directly or via a dedicated bus, bridge, switch, router, or the like.

[0021] SDF170 processes (services) the memory access requests provided by GPU102, provides read / write access to memory 135, and converts the physical memory address provided in the memory access request to the physical memory location (e.g., memory block) of one or more corresponding memory devices 135-0, 135-1, 135-2, ..., 135-N of memory 135 via channels 165-0, 165-1, 165-2, ..., 165-N. To convert the physical memory address provided in such a memory access request, SDF170 refers to the allocation mode selected by kernel mode driver 145 to determine to which of channels 165-0, 165-1, 165-2, ..., 165-N the physical memory address is mapped. As used herein, physical address mapping refers not to the mapping between virtual and physical addresses, but rather to the mapping between the physical addresses of a given processing system and channels and / or physical memory locations. In this specification, the terms "physical address" and "physical memory address" are used interchangeably to refer to a specific physical memory location of memory device 135 or an address otherwise associated therewith.

[0022] As shown, CPU 150 includes several processes such as executing one or more applications 155 for generating graphics commands. In various embodiments, the one or more applications 155 include applications that utilize the functions of GPU 102, such as applications that generate work in the processing system 100 or an operating system (OS). In some embodiments, the application 155 includes one or more graphics commands that direct the GPU 102 to render a graphical user interface (GUI) and / or a graphics scene. For example, in some embodiments, the graphics commands include commands that define a set of one or more graphics primitives to be rendered by the GPU 102.

[0023] In some embodiments, application 155 utilizes a graphics application programming interface (API) 160 to call a user-mode driver (not shown) (or a similar GPU driver). The user-mode driver issues one or more commands to GPU 102 to render one or more graphics primitives into a displayable graphics image. Based on the graphics instructions issued by application 155 to the user-mode driver, the user-mode driver generates one or more graphics commands that specify one or more operations of GPU 102 to perform the rendering of the graphics. In some embodiments, the user-mode driver is part of application 155 running on CPU 150. For example, in some embodiments, the user-mode driver is part of a game application running on CPU 150. Similarly, in some embodiments, kernel-mode driver 145 generates one or more graphics commands as part of the operating system running on CPU 150, either alone or in combination with the user-mode driver.

[0024] The GPU 102 includes three partitions 104, 106, and 108. Each partition 104, 106, 108 includes a set of shader engines (SE) 105 that are used to receive and execute commands simultaneously or in parallel. In some embodiments, each SE 105 is implemented as an individual die that includes a configurable number of shader engines, each shader engine includes a configurable number of workgroup processors, and each workgroup processor includes a configurable number of compute units. Some embodiments of the SE 105 are configured to shade the vertices of primitives that represent a model of a scene using information within a draw call received from any of the CPUs 150. Also, the SE 105 shades the pixels generated based on the shaded primitives and provides the shaded pixels to a display for presentation to a user, for example, via an I / O hub (not shown). In some embodiments, the GPU 102 further includes a display engine and a PCIe interface. Although three SE 105 are shown for each partition 104, 106, 108 as a total of nine SE 105 are shown in FIG. 1, some embodiments of the GPU 102 include more or fewer partitions, and some embodiments of the partitions 104, 106, 108 include more or fewer SE 105.

[0025] Each set of SE105 within partitions 104, 106, 108 is connected to a front end (e.g., front end 0 (FE-0) 110, front end 1 (FE-1) 120, and front end 2 (FE-2) 130) that fetches and schedules commands for processing a graphics workload received and executed by the shader engine of SE105. In some embodiments, SE105 of partitions 104, 106, 108 is vertically stacked on top of the corresponding front ends FE-0 110, FE-1 120, FE-2 130 of partitions 104, 106, 108. In some embodiments, each of FE-0 110, FE-1 120, FE-2 130 is implemented as an individual die. In some embodiments, each of the front end dies FE-0 110, FE-1 120, FE-2 130 includes a graphics L2 cache (not shown) that stores frequently used data and instructions. In some embodiments, the L2 cache is connected to one or more L1 caches implemented in SE105 and one or more L3 caches (or other last-level caches) implemented in the processing system 100. The caches collectively form a cache hierarchy.

[0026] Each of the front ends FE-0 110, FE-1 120, and FE-2 130 within GPU102 fetches primitives of the graphics workload, performs scheduling of the graphics workload for execution on SE105, and in some cases, handles serial synchronization, state updates, draw calls, cache activities, and primitive tessellation. Each of FE-0 110, FE-1 120, and FE-2 130 within GPU102 includes a command processor (not shown) that receives command buffers for execution on SE105. Also, each of FE-0 110, FE-1 120, and FE-2 130 includes a graphics register bus manager (GRBM) (not shown) that acts as a hub for register read and write operations that support multiple masters and multiple slaves. Thus, FE-0 110, FE-1 120, and FE-2 130 fetch commands for processing the respective sets of graphics workloads of SE105. Each of SE105 includes a shader engine configured to receive and execute commands from the respective FE-0 110, FE-1 120, and FE-2 130. In some embodiments, GPU102 operates in two modes (independent of the memory channel allocation mode described herein), namely, a first "single partition" mode in which FE-1 120 drives all of SE105 while FE-1 110 and FE-2 130 are inactive so that software interacts with only one FE, and a second "triple partition" mode in which each of FE-0 110, FE-1 120, and FE-2 130 drives its local SE105 such that each guest is assigned to one FE.

[0027] The kernel mode driver 145 includes a mode selector 175 that selects an assignment mode for each memory page. In a first mode, the kernel mode driver 145 assigns memory pages across all memory channels 165-0, 165-1, 165-2, ..., 165-N such that the memory pages are accessible by each of the partitions 104, 106, 108. In a second mode, the kernel mode driver 145 assigns memory pages across a subset (such as one) of the memory channels 165-0, 165-1, 165-2, ..., 165-N such that the memory pages are accessible by only a subset of the partitions 104, 106, 108. For example, in some embodiments, the processing system includes 48 memory channels 165-0, 165-1, 165-2, ..., 165-N. If the mode selector 175 determines that a memory page is accessible by all of the partitions 104, 106, 108, the mode selector 175 selects the first mode and the memory page is assigned and interleaved across all 48 memory channels 165-0, 165-1, 165-2, ..., 165-N. Conversely, if the mode selector 175 determines that a memory page is accessible by only one of the partitions (e.g., partition 104), the mode selector 175 selects the second mode and the memory page is assigned and interleaved across only a subset (e.g., 16) of the channels 165-0, 165-1, 165-2, ..., 165-N corresponding to partition 104.

[0028] When a memory page is assigned to all of memory channels 165-0, 165-1, 165-2, ..., 165-N, the memory page is interleaved (or "striped") across all of the channels, so that, for example, the first portion of the memory page is transmitted via the first channel 165-0, the second portion is transmitted via the second channel 165-1, the third portion is transmitted via the third channel 165-2, (and so on) until the Nth portion is transmitted via the Nth channel 165-N, at which point the N+1th portion is transmitted via the first channel 165-0 (and so on). In some embodiments, the kernel mode driver assigns the memory page across all of the channels, but executes a different interleaving pattern.

[0029] When the kernel mode driver 145 assigns a memory page across a subset of memory channels 165-0, 165-1, 165-2, ..., 165-N, the memory page is interleaved only across the subset. For example, if the subset includes three memory channels, the memory page is interleaved such that the first portion of the memory page is transmitted via the first channel 165-0, the second portion is transmitted via the second channel 165-1, the third portion is transmitted via the third channel 165-2, and the fourth portion is transmitted via the first channel 165-0 (and so on).

[0030] Figure 2 is a block diagram 200 of memory allocation across channels based on mode, according to some embodiments. In a first mode 202 shown by solid lines, a memory page (not shown) is allocated and interleaved across all of a memory channel 206 between a memory 135 and an SDF 170. In some embodiments, a mode selector 175 selects the first mode 202 in response to identifying that the memory page includes data that should be accessible by all of parallel processor unit partitions 104, 106, 108. For example, the mode selector 175 selects the first mode 202 in response to identifying that the memory page includes a texture or a vertex buffer, and allocates the memory page across all of the memory channel 206.

[0031] In a second mode 204 shown by dashed lines, a memory page is allocated and interleaved across a subset 208 of a memory channel. In some embodiments, a mode selector 175 selects the second mode 204 in response to identifying that the memory page includes data that should be accessible by all of parallel processor unit partitions 104, 106, 108. For example, the mode selector 175 selects the second mode 204 in response to identifying that the memory page includes thread private memory or a render target, and allocates the memory page across only a subset 208 of the memory channel 206.

[0032] FIG. 3 is a block diagram 300 showing the interleaving of memory page 302 across various numbers of channels based on mode, according to some embodiments. In the first mode 202, the memory page 302 is assigned to and interleaved across all of the memory channels 165-0 to 165-N, where in the illustrated example, N = 15. In the illustrated example, the memory page 302 is divided into 32 parts for interleaving, and the first 16 parts are assigned to any of the corresponding ones of the memory channels 0 to 15, at which point the interleaving wraps around so that the 17th part is assigned to memory channel 0, and the interleaving continues in this way.

[0033] In contrast, in the second mode 204, the memory page 302 is assigned to and interleaved across only a subset of the memory channels 165-0 to 165-7 in the illustrated example. The memory page 302 is divided into the same number of parts as the assignment according to the first mode 202 such that the first 8 parts are assigned to any of the corresponding ones of the memory channels 0 to 7, at which point the interleaving wraps around so that the 9th part is assigned to memory channel 0. The interleaving continues in this way until all 32 parts of the memory page 302 are assigned to any of the 8 channels of the subset of memory channels.

[0034] FIG. 4 is a diagram showing the representation of an indicator 400 of the use of the first mode 202 for assigning a memory page to the memory channels of a processing system. For ease of explanation, this example is described with respect to an exemplary execution in the processing system 100 of FIG. 1 and its constituent components and modules.

[0035] The physical memory address 402 is shown here as including an array of bits 404, each associated with a respective index. In some embodiments, the physical memory address 402 is mapped to all of channels 165-0, ..., 165-N based on the respective binary values in an index 406, herein referred to as a first aperture 406, of a given physical memory address. That is, the memory devices 135-0, ..., 13-N used to store and retrieve data associated with channel 165 and thus the physical memory address 402 are selected by the value of the bits in index 206. In some embodiments, the value of the bits in index 206 is used to select a channel identifier (channel ID) number associated with a given channel among channels 165-0, ..., 165-N.

[0036] In this example, a first interleaved configuration 200, which may be shown as [11, 10, 9, 8], causes SDF170 to map physical memory address 402 across all of channels 165-0 through 165-N based on the binary numbers at indices 11, 10, 9, and 8 of physical memory address 402. That is, the bit values of physical memory address 402 at indices 11, 10, 9, and 8 are used by SDF170 to determine the channel ID number corresponding to any one of channels 165-1 through 165-N to which physical memory address 402 is to be mapped. As shown, the least significant 16 bits of physical memory address 402 are indexed as [15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, 0]. Here, the use of 4 bits of physical memory address 402 to select a channel ID number, as defined by the first mode 202, enables a physical memory address such as physical memory address 402 to be mapped across up to 16 channels (i.e., 2^4 channels, since 4 bits at bit indices 11, 10, 9, and 8 of the group of bits of the first aperture 406 are used by SDF170 to determine the channel ID number). In some embodiments, the number of physical memory address bits within a group of bits 408 (including the group of bits to the right of the group of bits of the first aperture 406) determines the size of each set of consecutive physical memory addresses to be mapped to a given channel. This size may also be referred to as the "interleave granularity" and may be characterized as the maximum number of consecutive bytes to be stored in each channel before switching to the next channel. Continuing the example, the number of bits included in the group of bits 208, shown here as [7, 6, 5, 4, 3, 2, 1, 0], determines the byte-unit interleave granularity of the first mode 202.In this example, the interleaving granularity is 256B (i.e., since there are 8 bits within the group of bits 208 to the right of the group of bits 206, it is 2^8B, enabling 256 combinations of those 8 bits corresponding to 256 consecutive physical memory addresses, and 1B of data can be stored at each physical memory address).

[0037] FIG. 5 is a diagram showing the representation of an indicator 500 of the use of a second mode 204 for allocating memory pages to the memory channels of a processing system. For ease of explanation, this example is described with respect to an exemplary execution in the processing system 100 of FIG. 1 and its constituent components and modules.

[0038] In this example, the change in the memory channel allocation of the processing system 100 from the first mode 202 to the second mode 204 is performed by changing the group of bits of the physical memory address 402 used by the SDF 170 to map the physical memory address 402 to a specific subset of channels 165-1,..., 165-N (e.g., from the group of bits of the first aperture 406 to the group of bits 506-1, 506-2).

[0039] For example, in response to identifying that the memory page 302 requires access only by one of the partitions 102, 104, 106, the mode selector 175 selects the second mode 204 and changes the group of bits used to determine the channel ID numbers of channels 165-1,..., 165-N to which the physical memory address 402 should be mapped from the group of bits of the first aperture 406 to the group of bits of the second apertures 506-1, 506-2, indicating that the second mode 204 is in use. In the illustrated example, the indices 9, 8 indicate to which of the local channels 165-1,..., 165-N the memory page 302 is allocated, and the indices 15, 14 indicate the chip select.

[0040] FIG. 6 is a block diagram of a physical address map 600 of DRAM dynamically assigned to a subset of memory channels, according to some embodiments. In the illustrated example, 128GB of DRAM is divided into two apertures, i.e., the first 64GB is assigned to the first aperture 602 and the second 64GB is assigned to the second aperture 604. Most of the first 64GB assigned to the first aperture 602 is interleaved across all memory channels 206, but a portion of the memory assigned to the first aperture 602 is reserved for physical function resources.

[0041] The second aperture 604 corresponds to portions of DRAM interleaved across the first, second, and third subsets 208 of memory channels. Similar to the first aperture 602, a portion of the memory assigned to the second aperture 604 is reserved for physical function resources.

[0042] FIG. 7 is a flowchart showing a method 700 for allocating memory pages across various numbers of channels, according to some embodiments. For ease of explanation, method 600 is described with respect to an exemplary execution in the processing system 100 of FIG. 1 and its constituent components and modules.

[0043] In block 702, mode selector 175 determines whether memory page 302 should be shared by a plurality of partitions 104, 106, 108. If in block 702 mode selector 175 determines that memory page 302 should be shared across partitions 104, 106, 108, the method flow continues to block 704. In block 704, kernel mode driver 145 selects a first mode 202 and allocates memory page 302 for interleaving across a plurality of channels 165-0 to 165-N. In block 706, kernel mode driver 145 indicates that memory page 302 is allocated according to the first mode 202 at a first aperture 406 of the physical address 402 of memory page 302.

[0044] If in block 702 mode selector 175 determines that memory page 302 should not be shared across partitions 104, 106, 108, the method flow continues to block 708. In block 708, kernel mode driver 145 selects a second mode 204 and allocates memory page 302 for interleaving across a subset 208 of the plurality of channels 165-0 to 165-N. In block 710, kernel mode driver 145 indicates that memory page 302 is allocated according to the second mode 204 at second apertures 506-1, 506-2 of the physical address 402 of memory page 302.

[0045] In some embodiments, the above-described apparatus and techniques are implemented in a system that includes one or more integrated circuit (IC) devices (also referred to as integrated circuit packages or microchips), such as the processing system described above with reference to FIGS. 1-7. Electronic design automation (EDA) and computer aided design (CAD) software tools can be used to design and manufacture these IC devices. These design tools are typically represented as one or more software programs. The one or more software programs operate a computer system to act on code representing the circuits of the one or more IC devices to perform at least a portion of the process for designing or adapting a manufacturing system for fabricating the circuits. This code can include instructions, data, or a combination of instructions and data. Software instructions representing design tools or manufacturing tools are typically stored on a computer-readable storage medium accessible to a computing system. Similarly, code representing one or more stages of the design or manufacture of an IC device is stored on and accessed from the same or a different computer-readable storage medium.

[0046] A computer-readable storage medium includes any non-transitory storage medium or combination of non-transitory storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media include, but are not limited to, optical media (e.g., compact disc (CD), digital versatile disc (DVD), Blu-ray (registered trademark) disc), magnetic media (e.g., floppy (registered trademark) disc, magnetic tape, magnetic hard drive), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or microelectromechanical systems (MEMS)-based storage media. A computer-readable storage medium (e.g., system RAM or ROM) may be built into the computing system, a computer-readable storage medium (e.g., magnetic hard drive) may be fixedly attached to the computing system, a computer-readable storage medium (e.g., optical disc or universal serial bus (USB)-based flash memory) may be removably attached to the computing system, or a computer-readable storage medium (e.g., network-accessible storage (NAS)) may be coupled to the computer system via a wired or wireless network.

[0047] In some embodiments, certain aspects of the techniques described above are implemented by one or more processors of a processing system that executes software. The software includes one or more sets of executable instructions stored on a non-transitory computer-readable storage medium or otherwise tangibly embodied. The software may include instructions and certain data, and when executed by one or more processors, the instructions and certain data operate the one or more processors to perform one or more aspects of the techniques described above. Non-transitory computer-readable storage media can include, for example, magnetic or optical disk storage devices, solid state storage devices such as flash memory, cache, random access memory (RAM), or other non-volatile memory device(s). The executable instructions stored on the non-transitory computer-readable storage medium can be implemented in source code, assembly language code, object code, or other instruction formats interpretable or otherwise executable by one or more processors.

[0048] In addition to the above, it should be noted that not all activities or elements described in the general description are required, and some activities or parts of a particular device may not be required, and one or more additional activities may be performed and one or more additional elements may be included. Further, the order in which activities are recited is not necessarily the order in which they are performed. Also, the concepts have been described with reference to specific embodiments. However, those skilled in the art will understand that various changes and modifications can be made without departing from the scope of the invention as set forth in the claims. Accordingly, the specification and drawings are to be considered in an illustrative rather than a limiting sense, and all such modifications are intended to be included within the scope of the invention.

[0049] Benefits, other advantages, and solutions to problems have been described above with respect to specific embodiments. However, benefits, advantages, solutions to problems, and features that may give rise to or manifest any benefit, advantage, or solution are not to be construed as important, essential, or indispensable features of any or all of the claims. Further, the disclosed invention may be modified and practiced in different but similar ways that will be apparent to those skilled in the art having the benefit of the teachings herein, so the specific embodiments described above are merely illustrative. There is no limitation as to the details of construction or design shown herein other than as described in the appended claims. Accordingly, the specific embodiments described above may be changed or modified, and it is clear that all such variations are considered to be within the scope of the disclosed invention. Accordingly, the protection sought herein is set forth in the appended claims.

Claims

1. A method comprising: In a first mode, allocating memory pages to be interleaved across a plurality of channels; and In a second mode, allocating the memory pages to be interleaved across a subset of the plurality of channels, The method.

2. The first mode includes allocating the memory pages to an integrated memory architecture shared by a plurality of parallel processing unit chips, and the second mode includes allocating the memory pages to a private heterogeneous memory architecture of a single parallel processing unit chip, The method of Claim 1.

3. Selecting the first mode in response to the memory pages being shared by a plurality of parallel processing unit chips; and Selecting the second mode in response to the memory pages being accessed by a single parallel processing unit chip, The method of Claim 1 or 2.

4. Selecting the first mode is performed in response to the memory pages including one or more of a texture and a vertex buffer, The method of Claim 3.

5. Selecting the second mode is performed in response to the memory pages including one or more of a thread private memory and a render target, The method of Claim 3.

6. Indicating the use of the first mode in a first physical address architecture; and Indicating the use of the second mode in a second physical address architecture, The method of Claim 1.

7. The plurality of channels correspond to a plurality of memory devices shared by a plurality of parallel processing unit chips, and the subset of the plurality of channels corresponds to a memory device private to a parallel processing unit chip, The method of Claim 1.

8. An apparatus comprising: A plurality of memory devices; A plurality of parallel processing unit chips configured to access data stored in the plurality of memory devices via a plurality of channels; and A kernel mode driver, The kernel mode driver: In a first mode, allocating memory pages to be interleaved across the plurality of channels; In a second mode, allocate the memory page so as to be interleaved across a subset of the plurality of channels; configured to perform, an apparatus. **Claim 9** The first mode includes allocating the memory page to an integrated memory architecture shared by the plurality of parallel processing unit chips, and the second mode includes a private heterogeneous memory architecture for a subset of the plurality of parallel processing unit chips. The apparatus of claim 8. **Claim 10** The kernel mode driver, select the first mode in response to the memory page being shared by a plurality of parallel processing unit chips; select the second mode in response to the memory page being accessed by a single parallel processing unit chip; configured to perform, the apparatus according to claim 8 or 9. **Claim 11** The kernel mode driver selects the first mode in response to the memory page including one or more of a texture and a vertex buffer. The apparatus of claim 10. **Claim 12** The kernel mode driver selects the second mode in response to the memory page including one or more of thread private memory and a render target. The apparatus of claim 10. **Claim 13** The kernel mode driver, indicates use of the first mode in a first physical address aperture; indicates use of the second mode in a second physical address aperture. The apparatus of claim 8. **Claim 14** A method comprising: in a first mode, interleaving a first memory page across a plurality of channels shared by a plurality of parallel processing unit chips; in a second mode, interleaving a second memory page across a subset of the plurality of channels shared by a subset of the plurality of parallel processing unit chips. A method. **Claim 15** The first mode includes allocating the first memory page to an integrated memory architecture shared by the plurality of parallel processing unit chips, and the second mode includes allocating the second memory page to a private heterogeneous memory architecture for a subset of the plurality of parallel processing unit chips. The method of claim 14.