Unified Kernel Virtual Address Space for Heterogeneous Computing

By using an IOMMU to create a unified kernel virtual address space, the challenge of inefficient memory buffer sharing between computing devices is addressed, enabling effective buffer pointer sharing and improved memory management in heterogeneous systems.

JP7682148B2Active Publication Date: 2025-05-23ATI TECHNOLOGIES ULC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022503804
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-07-23
Filing Date
2020-07-22
Publication Date
2025-05-23
Estimated Expiration
2040-07-22

AI Technical Summary

Technical Problem

Sharing memory buffers between computing devices is inefficient due to the need for pointer translation across different virtual address spaces, making it difficult to share pointers or structures containing pointers.

Method used

Implementing a unified kernel virtual address space through an Input/Output Memory Management Unit (IOMMU) that allows different subsystems with different kernel address spaces to share memory buffers by generating mappings from logical addresses in each kernel address space to a shared physical memory block.

Benefits of technology

Enables efficient sharing of buffer pointers between subsystems by creating a unified address space, allowing direct pointer exchange and improving memory management efficiency across heterogeneous computing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682148000001
    Figure 0007682148000001
  • Figure 0007682148000002
    Figure 0007682148000002
  • Figure 0007682148000003
    Figure 0007682148000003
Patent Text Reader

Abstract

A system, apparatus, and method for implementing a unified kernel virtual address space for heterogeneous computing are disclosed. The system includes at least a first subsystem that executes a first kernel, an input / output memory management unit (IOMMU), and a second subsystem that executes a second kernel. To share a memory buffer between the two subsystems, the first subsystem allocates a memory block in a portion of system memory controlled by the first subsystem. A first mapping is created from a first logical address in the first subsystem's kernel address space to the memory block. The IOMMU then creates a second mapping to map from a second logical address in the second subsystem's kernel address space to the physical address of the memory block. These mappings allow the first subsystem and the second subsystem to share buffer pointers that reference the memory block.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Description of Related Art Generally, sharing a memory buffer between computing devices requires the devices to exchange pointers or handles in the device or physical address space.

[0002] In most cases, before using this pointer, the computing device translates it into the local virtual address space with a memory management unit (MMU) page table mapping, which makes sharing the pointer itself or structures that contain the pointer inefficient and difficult.

[0003] Advantages of the methods and mechanisms described herein may be better understood by referring to the following description in conjunction with the accompanying drawings. [Brief description of the drawings]

[0004] [Figure 1] FIG. 1 is a block diagram of one embodiment of a computing system. [Diagram 2] FIG. 1 illustrates an embodiment of creating a unified kernel virtual address space for heterogeneous computing. [Diagram 3] FIG. 1 illustrates an embodiment of sharing a buffer between two separate subsystems. [Figure 4] FIG. 2 illustrates one embodiment of mapping a region of memory into multiple kernel address spaces. [Diagram 5] FIG. 2 is a generalized flow diagram illustrating one embodiment of a method for creating a unified kernel virtual address space. [Figure 6] FIG. 2 is a generalized flow diagram illustrating one embodiment of a method for enabling a common shared region in multiple kernel address spaces. [Figure 7] FIG. 2 is a generalized flow diagram illustrating one embodiment of a method for enabling buffer pointer sharing between two different subsystems. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0005] In the following description, numerous specific details are set forth to provide a thorough understanding of the methods and mechanisms presented herein. However, those skilled in the art will recognize that various embodiments may be practiced without these specific details. In some instances, well-known structures, components, signals, computer program instructions, and techniques have not been shown in detail so as not to obscure the approaches described herein. For simplicity and clarity of illustration, it should be understood that elements illustrated in the figures have not necessarily been drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements.

[0006] Disclosed herein are systems, apparatus, and methods for implementing a unified kernel virtual address space for heterogeneous computing. In one embodiment, the system includes at least a first subsystem that executes a first kernel, an input / output memory management unit (IOMMU), and a second subsystem that executes a second kernel. In one embodiment, the IOMMU generates a unified kernel address space that allows the first subsystem and the second subsystem to share memory buffers at the kernel level. To share the memory buffer between the two subsystems, the first subsystem allocates a memory block in a portion of the system memory controlled by the first subsystem. A first mapping is generated from a first logical address in a first kernel address space of the first subsystem to the memory block. The IOMMU then generates a second mapping to map a physical address of the memory block from a second logical address in a second kernel address space of the second subsystem. These mappings allow the first subsystem and the second subsystem to share buffer pointers in the kernel address space that reference the memory block.

[0007] 1, a block diagram of one embodiment of a computing system 100 is shown. In one embodiment, computing system 100 includes at least a first subsystem 110, a second subsystem 115, input / output (I / O) interface(s) 120, an input / output memory management unit (IOMMU) 125, a memory subsystem 130, and peripheral device(s) 135. In other embodiments, computing system 100 may include other components and / or computing system 100 may be configured differently.

[0008] In one embodiment, the first and second subsystems 110, 115 have different kernel address spaces, but a unified kernel address space is generated by the IOMMU 125 for the first and second subsystems 110, 115. The unified kernel address space allows the first and second subsystems 110, 115 to pass pointers to each other and share buffers. For the first and second subsystems 110, 115, each kernel address space includes a kernel logical address and a kernel virtual address. In some architectures, a kernel's logical address and its associated physical address differ by a constant offset. Kernel virtual addresses do not necessarily have a linear one-to-one mapping to the physical addresses that characterize kernel logical addresses. All kernel logical addresses are kernel virtual addresses, but kernel virtual addresses do not necessarily have to be kernel logical addresses.

[0009] In one embodiment, each of the first subsystem 110 and the second subsystem 115 includes one or more processors that execute an operating system. The processor(s) also execute one or more software programs in various embodiments. The processor(s) of the first subsystem 110 and the second subsystem 115 may include any number and type of processing units (e.g., central processing units (CPUs), graphic processing units (GPUs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs)). The first subsystem 110 also includes a memory management unit (MMU) 112, and the second subsystem 115 includes an MMU 117, with each MMU handling virtual to physical address translation for the corresponding subsystem. Although the first and second subsystems 110, 115 have different kernel address spaces, a unified kernel address space is created by the IOMMU 125 for the first and second subsystems 110, 115, allowing the first and second subsystems 110, 115 to pass pointers and share buffers within the memory subsystem 130.

[0010] Memory subsystem 130 includes any number and type of memory devices. For example, the types of memory in memory subsystem 130 may include high bandwidth memory (HBM), non-volatile memory (NVM), dynamic random access memory (DRAM), static random access memory (SRAM), NAND flash memory, NOR flash memory, ferroelectric random access memory (FeRAM), or other memories. I / O interface 120 represents any number and type of I / O interface (e.g., Peripheral Component Interconnect (PCI) bus, PCI-Extended (PCI-X), PCI Express (PCIE) bus, Gigabit Ethernet (GBE) bus, Universal Serial Bus (USB)). Various types of peripheral devices 135 may be coupled to I / O interface 120. Such peripheral devices 135 include, but are not limited to, displays, keyboards, mice, printers, scanners, joysticks or other types of game controllers, media recording devices, external storage devices, network interface cards, and the like.

[0011] In one embodiment, to create a memory space shared between the first subsystem 110 and the second subsystem 115, a memory block is allocated in a portion of the system memory managed by the first subsystem 110. After the initial block of memory is allocated, an appropriate I / O virtual address (VA) is assigned to the second subsystem 115. In one embodiment, an IOMMU mapping is created from the kernel address space of the second subsystem 115 to the physical address of the memory block. In this embodiment, the IOMMU 125 performs the virtual address mapping of the second subsystem 115 to the memory block. Then, when additional memory is allocated, a heap allocate function is called and the address is mapped based on the same I / O VA address previously generated. A message is then sent to the second subsystem 115 informing the second subsystem 115 of the unified address.

[0012] In various embodiments, computing system 100 is a computer, a laptop, a mobile device, a game console, a server, a streaming device, a wearable device, or any other variety of computing system or device. It should be noted that the number of components of computing system 100 may vary from embodiment to embodiment. In other embodiments, there may be more or fewer components than those shown in FIG. 1. It should also be noted that in other embodiments, computing system 100 includes other components not shown in FIG. 1. Additionally, in other embodiments, computing system 100 is structured in a manner other than that shown in FIG. 1.

[0013] Referring to Figure 2, a diagram of one embodiment of creating a unified kernel virtual address space for heterogeneous computing is shown. The vertical dashed lines from the top of Figure 2 represent, from left to right, a first subsystem 202, an MMU 204 of the first subsystem 202, a device driver 206 running on the first subsystem 202, a shared region of memory 208, an IOMMU 211, an MMU 212 of a second subsystem 214, and a second subsystem 214. For purposes of explanation, it is assumed that the first subsystem 202 has a first operating system and the second subsystem 214 has a second operating system, the second operating system being different from the first operating system. Also, for purposes of explanation, it is assumed that the first subsystem 202 and the second subsystem 214 are part of a heterogeneous computing system. In one embodiment, for this heterogeneous computing system, the first subsystem 202 and the second subsystem 214 each execute a portion of the workload, and to execute the workload, the first subsystem 202 and the second subsystem 214 share buffer pointers and buffers with each other. The diagram of Figure 2 shows an example of allocating memory shared between the first subsystem 202 and the second subsystem 214.

[0014] In one embodiment, a first step 205 is performed by the device driver 206 to create a heap (i.e., a shared memory region). As used herein, the term "heap" is defined as a virtual memory pool that is mapped to a physical memory pool. Next, in step 210, the desired size of the heap is allocated to the physical memory subsystem. In step 215, the kernel to heap mapping is created. In one embodiment, when the carveout heap is first created, a new flag indicates whether the heap is taken from the kernel logical address space. For example, the kmalloc function returns memory in the kernel logical address space. In one embodiment, for a Linux operating system, the memory manager uses the Linux genpool library to manage the buffers allocated to the heap. In this embodiment, when a buffer is allocated, it is marked in the internal pool and the physical address is returned. The carveout heap then wraps this physical address with the buffer and the sg_table descriptor. In one embodiment, once a buffer is mapped in the kernel address space, the heap_map_kernel function uses kmap instead of vmap to map the buffer in step 220. The function kmap maps the buffer to a given virtual address based on the logical mapping. The choice of kmap or vmap is controlled by a new flag provided during the creation of the carveout heap. Alternatively, the heap is created with the kernel mapping using the genpool library, the carveout application programming interface (API), and a new flag. In this case, the kernel map returns the pre-mapped address. The user mode mapping is still applied on the fly. In the case of the second subsystem 214, the shared memory buffers allocated by the first subsystem 202 are managed similarly to the carveout heap wrapped genpool. A new API allows adding external buffers to the carveout heap.Local tasks on the second subsystem 214 use the same API to allocate from this carveout heap. After step 220, the allocation of the shared memory region is complete.

[0015] Next, in step 225, an input / output (I / O) virtual address is assigned to the shared memory region by the IOMMU 211. Next, in step 230, a contiguous block of DMA address space is reserved for the shared memory region. Next, in step 235, the shared memory region is mapped by the IOMMU 211 to the kernel address space of the second subsystem 214. Note that the IOMMU mapping should not be released until the device driver 206 is shut down. The mapping from the kernel address space of the first subsystem 202 to the shared memory region is invalidated in step 240. Next, in step 245, the kernel address space mapping is released by executing a memory release function.

[0016] Referring to Figure 3, a diagram of one embodiment of sharing buffers between two separate subsystems is shown. The vertical dashed lines extending from the top to the bottom of Figure 3 represent, from left to right, a first subsystem 302, an MMU 304 of the first subsystem 302, a device driver 306 running on the first subsystem 302, a shared area of ​​memory 308, an IOMMU 310, an MMU 312 of a second subsystem 314, and a second subsystem 314. For purposes of explanation, it is assumed that the first subsystem 302 has a first operating system and the second subsystem 314 has a second operating system, the second operating system being different from the first operating system.

[0017] For loop exchange 305, a memory block is allocated and a mapping from the kernel address space of the first subsystem 302 to the physical address of the memory block is generated by the device driver 306. The mapping is then maintained by the MMU 304. The memory block may be allocated exclusively to the second subsystem 314, may be shared between the first subsystem 302 and the second subsystem 314, or may be allocated exclusively to the first subsystem 302. A message is then sent from the first subsystem 302 to the second subsystem 314 with the address and size of the block of memory. In one embodiment, the message may be sent out of band. After the message is received by the second subsystem 314, a mapping from the kernel address space of the second subsystem 314 to the physical address of the memory block is generated and maintained by the MMU 312. For loop exchange 315, data is exchanged between the first subsystem 302 and the second subsystem 314 using the shared region 308. Since the kernel virtual addresses are unified, the buffer pointer 1 st _SS_buf and 2 nd The _SS_buf is the same and can be freely exchanged between the first subsystem 302 and the second subsystem 314, and can be further split using the genpool library.

[0018] Referring to Figure 4, a diagram of one embodiment of mapping memory regions to multiple kernel address spaces is shown. In one embodiment, a heterogeneous computing system (e.g., system 100 of Figure 1) includes multiple different subsystems with their own independent operating systems. The vertical rectangular blocks shown in Figure 4 represent the address spaces of the different components of the heterogeneous computing system. The address spaces shown in Figure 4 are, from left to right, a first subsystem virtual address space 402, a physical memory space 404, a device memory space 406, and a second subsystem virtual address space 408.

[0019] In one embodiment, first subsystem virtual address space 402 includes shared region 420 that is mapped to memory blocks 425 and 430 of physical memory space 404. In one embodiment, the mapping of shared region 420 to memory blocks 425 and 430 is generated and maintained by first subsystem MMU 412. To enable shared region 420 to be shared with the second subsystem, memory blocks 425 and 430 are mapped by IOMMU 414 to shared region 435 of device memory space 406. Shared region 435 is then mapped by second subsystem MMU 416 to shared region 440 of second subsystem virtual address space 408. Through this mapping scheme, the first subsystem and second subsystem can share buffer pointers with each other in their own kernel address spaces.

[0020] Referring to FIG. 5, one embodiment of a method 500 for generating a unified kernel virtual address space is shown. For purposes of illustration, the steps in this embodiment and in FIGS. 6-7 are shown in sequence. However, it should be noted that in various embodiments of the described method, one or more of the described elements may be performed simultaneously, in a different order than shown, or may be omitted entirely. Other additional elements may also be performed as desired. Any of the various systems or devices described herein may be configured to perform the method 500.

[0021] The first subsystem allocates a memory block at a first physical address in a physical address space corresponding to the memory subsystem (block 505). In one embodiment, the first subsystem executes a first operating system having a first kernel address space. The first subsystem then generates a mapping of a first logical address to a first physical address in the first kernel address space (block 510). Note that the first logical address space of the first kernel address space is a first linear offset from the first physical address. The IOMMU then generates an IOMMU mapping of a second logical address to a first physical address in a second kernel address space (block 515). In one embodiment, the second kernel address space is associated with a second subsystem. Note that the second logical address space of the second kernel address space is a second linear offset from the first physical address.

[0022] The first subsystem then communicates the buffer pointer to the second subsystem, where the buffer pointer points to the first logical address (block 520). The second subsystem then generates an access request using the buffer pointer and communicates the access request to the IOMMU (block 525). The IOMMU then translates the virtual address of the buffer pointer to the first physical address using the previously generated IOMMU mapping (block 530). The second subsystem then accesses the memory block at the first physical address (block 535). After block 535, the method 500 ends.

[0023] Referring to FIG. 6, one embodiment of a method 600 for enabling a common shared region in multiple kernel address spaces is illustrated. A first subsystem allocates a memory block at a first physical address in a physical address space corresponding to the memory subsystem (block 605). For purposes of illustration, it is assumed that the system described by the method 600 includes a first subsystem and a second subsystem. It is also assumed that the first subsystem executes a first operating system having a first kernel address space, and the second subsystem executes a second operating system having a second kernel address space. The first subsystem generates a mapping of a first logical address in the first kernel address space to a first physical address, where the first logical address in the first kernel address space is a first offset away from the first physical address (block 610).

[0024] The IOMMU then selects a device address that is a second offset away from the first logical address in the second kernel address space (block 615). The IOMMU then creates an IOMMU mapping of the selected device address in the device address space to the first physical address (block 620). The IOMMU mapping enables a common shared region of both the first kernel address space and the second kernel address space to be used by both the first subsystem and the second subsystem (block 625). After block 625, the method 600 ends.

[0025] Referring to FIG. 7, one embodiment of a method 700 for enabling sharing of buffer pointers between two different subsystems is shown. A first subsystem maps a first region of a first kernel address space to a second region of a physical address space (block 705). In one embodiment, the physical address space corresponds to a system memory controlled by the first subsystem. An IOMMU then maps a third region of a second kernel address space to a fourth region in a device address space (block 710). For purposes of illustration, it is assumed that the second kernel address space corresponds to a second subsystem that is different from the first subsystem. The IOMMU then maps the fourth region in the device address space to a second region in the physical address space such that addresses in the first and third regions point to matching (i.e., identical) addresses in the second region (block 715). The first subsystem and the second subsystem can share buffer pointers in the first or second kernel address space that reference the second region in the physical address space (block 720). After block 720, the method 700 ends.

[0026] In various embodiments, program instructions of a software application are used to implement the methods and / or mechanisms described herein. For example, program instructions executable by a general-purpose processor or a special-purpose processor are contemplated. In various embodiments, such program instructions may be expressed by a high-level programming language. In other embodiments, the program instructions may be compiled from the high-level programming language into a binary, intermediate, or other form. Alternatively, program instructions describing the operation or design of hardware may be written. Such program instructions may be expressed by a high-level programming language such as C. Alternatively, a hardware design language (HDL) such as Verilog may be used. In various embodiments, the program instructions are stored in any of a variety of non-transitory computer-readable storage media. The storage medium is accessible by the computing system during use to provide the program instructions to the computing system for program execution. Generally, such a computing system includes at least one or more memories and one or more processors configured to execute the program instructions.

[0027] It should be emphasized that the above embodiments are merely non-limiting examples of embodiments. Numerous variations and modifications will become apparent to those skilled in the art once the above disclosure is fully appreciated. It is intended that the following claims be interpreted to include all such variations and modifications.

Claims

1. A method comprising: a first subsystem mapping a first region of a first kernel address space to a first region of a first physical address space; a second subsystem different from the first subsystem mapping a first region of a second kernel address space into a first region of the first physical address space and into a device memory space; enabling a common shared region in both the first kernel address space and the second kernel address space. method.

2. the common shared area can be shared by both a first subsystem and a second subsystem different from the first subsystem; 2. The method of claim 1.

3. an address in a first region of the first kernel address space is mapped to the same location in the first physical address space as an address in the first region of the second kernel address space; 2. The method of claim 1.

4. Mapping the first region of the first kernel address space to the first region of the first physical address space is performed by a first subsystem.

2. The method of claim 1.

5. the mapping of the first region of the second kernel address space to the first region of the first physical address space is performed by a second subsystem distinct from the first subsystem. The method of claim 4.

6. the second subsystem being an input / output memory management unit; The method of claim 5.

7. 1. A system comprising: a first subsystem configured to map a first region of a first kernel address space to a first region of a first physical address space; a second subsystem, distinct from the first subsystem, configured to map a first region of a second kernel address space to a first region of the first physical address space and to a device memory space; the system is configured to enable a common shared region in both the first kernel address space and the second kernel address space; system.

8. an address in a first region of the first kernel address space is mapped to the same location in the first physical address space as an address in the first region of the second kernel address space; The system of claim 7.

9. the second subsystem being an input / output memory management unit; The system of claim 7.

10. the first subsystem is configured to map the first region to the first physical address space; The system of claim 7.

11. the first subsystem is configured to communicate a pointer to a location within a first region of the first physical address space to the second subsystem via the common shared region; The system of claim 8.

12. the second subsystem comprising an input / output memory management unit configured to map a first region of the second kernel address space to a first region of the first physical address space; The system of claim 11.

13. 1. A system comprising: A memory device; a first processing device configured to map a first region of a first virtual address space to a region within the memory device; a second processing device configured to map a first region of a second virtual address space to a region in the memory device such that the region in the memory device is shared by the first processing device and the second processing device; addresses in a first region of the first virtual address space are mapped to the same locations in the memory device as addresses in the first region of the second virtual address space; system.

14. the second processing device is configured to map a first region of the second virtual address space to a region in the memory device, the first region of the second virtual address space being mapped to a device memory space that is mapped to a region in the memory device; The system of claim 13.

15. the second processing device comprising an input / output memory management unit configured to map the device memory space to a region within the memory device; 15. The system of claim 14.

16. the first processing device is configured to allocate the region to the memory device; The system of claim 13.

17. the first processing device is configured to communicate a pointer to a location within a region of the memory device to the second processing device via the location within the memory device; The system of claim 13.

Citation Information

Patent Citations

  • Input / Output memory management unit with a protection mode that prevents I / O devices from accessing memory.

    JP2014531672A

  • Efficient memory and resource management

    JP2015500524A

  • Dynamic Address Negotiation for Shared Memory Regions in Heterogeneous Multiprocessor Systems

    JP2016532959A

  • Method for processing input and output on multi kernel system and apparatus for the same

    US20190114193A1