Method and apparatus for managing LLC resources

By partitioning SRAM blocks in the Last Cache (LLC) and managing them using CSR, the problem of frequent access to main memory by user-mode software is solved, resulting in higher system efficiency and performance.

CN122045083APending Publication Date: 2026-05-15INTEL CHINA RES CENT CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTEL CHINA RES CENT CO LTD
Filing Date
2026-04-14
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, user-mode software such as binary translation (BT) frequently accesses main memory when storing latency-sensitive data, leading to performance bottlenecks and limiting the overall system throughput and efficiency.

Method used

By configuring the control status register, a portion of the final cache (LLC) is divided into static random access memory (SRAM), and this SRAM block is managed by the CSR to store the virtual memory addresses of user-mode software, thus eliminating frequent accesses to main memory.

Benefits of technology

It effectively reduces the latency-sensitive data storage overhead of user-mode software and improves system efficiency, especially achieving a performance improvement of approximately 8.9% on the RISC-V architecture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045083A_ABST
    Figure CN122045083A_ABST
Patent Text Reader

Abstract

The invention relates to a method and a device for managing LLC (Logical Link Control) resources. There is provided a method for managing LLC resources, comprising: receiving a memory mapping request from user mode software, the memory mapping request requesting to map a virtual memory address of the user mode software to a physical address of an LLC SRAM block; in response to the memory mapping request, allocating at least one LLC SRAM block from an LLC SRAM resource pool, where the LLC SRAM resource pool is a storage resource dynamically configured in an SRAM mode by configuring a first control state register in the LLC; establishing mapping from a virtual memory address of the user mode software to a physical address of the allocated LLC SRAM block, and configuring a page table item attribute of the mapping as a non-cache mode; and enabling the allocated LLC SRAM block by a second control state register.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more specifically, to methods and apparatus for managing last-level cache (LLC) resources. Background Technology

[0002] User-mode software (e.g., binary translation (BT)) requires storage resources to store latency-sensitive data. While storing relevant data in main memory is common practice for user-mode software, the inherently high access latency of main memory can pose a significant performance bottleneck for latency-sensitive data requiring high-frequency access. Frequent main memory accesses become a major factor limiting overall system throughput and efficiency, thus restricting the higher performance that such software can achieve. Summary of the Invention

[0003] According to embodiments of this application, a method for managing last-level cache LLC resources is provided, comprising: receiving a memory mapping request from user-mode software, the memory mapping request requesting that a virtual memory address of the user-mode software be mapped to a physical address of an LLC static random access memory (SRAM) block; in response to the memory mapping request, allocating at least one LLC SRAM block from an LLC SRAM resource pool, wherein the LLC SRAM resource pool is a storage resource in the LLC that is dynamically configured to SRAM mode by configuring a first control status register; establishing a mapping from the virtual memory address of the user-mode software to the physical address of the allocated LLC SRAM block, and configuring the page table entry attribute of the mapping to a non-cached mode; and enabling the allocated LLC SRAM block via a second control status register.

[0004] According to an embodiment of this application, an apparatus for managing LLC resources is provided, wherein the apparatus includes a processor circuit configured to perform the above-described method for managing LLC resources.

[0005] According to an embodiment of this application, a computer-readable storage medium is provided, on which instructions are stored, wherein the instructions, when executed by a processor, cause the processor to perform the above-described method for managing LLC resources.

[0006] According to an embodiment of this application, a computer program product is provided, including instructions, wherein when executed by a processor, the instructions cause the processor to perform the above-described method for managing LLC resources. Attached Figure Description

[0007] Embodiments of this application will be described by way of example rather than limitation in conjunction with the accompanying drawings, wherein similar reference numerals denote similar elements, and wherein: Figure 1 The diagram illustrates a block diagram of an example processor and / or SoC, which may have one or more cores and an integrated memory controller.

[0008] Figure 2 An example of flexibly allocating host hardware registers is shown.

[0009] Figure 3 An example of a fixed mapping to host hardware registers is shown.

[0010] Figure 4 A flowchart of a method for managing LLC resources according to an embodiment of this application is shown.

[0011] Figure 5 An example workflow for using LLC SRAM resources based on CSR configuration according to an embodiment of this application is shown.

[0012] Figure 6 An example of the bcr register according to an embodiment of this application is shown.

[0013] Figure 7 An example of a bamr register according to an embodiment of this application is shown.

[0014] Figure 8 An example of a ber register according to an embodiment of this application is shown.

[0015] Figure 9 An example of LLC SRAM software management according to an embodiment of this application is shown.

[0016] Figure 10 An example of an LLC SRAM mapping process according to an embodiment of this application is shown.

[0017] Figure 11 This is a block diagram illustrating components capable of reading instructions from a machine-readable or computer-readable medium (e.g., a non-transitory machine-readable storage medium) and performing any one or more of the methods discussed herein, according to some example embodiments. Detailed Implementation

[0018] The features and exemplary embodiments of various aspects of this application will now be described in detail. Numerous specific details are set forth in the following detailed description to provide a comprehensive understanding of this application. However, it will be apparent to those skilled in the art that this application can be implemented without requiring some of these specific details. The following description of embodiments is merely intended to provide a better understanding of this application by illustrating examples. This application is by no means limited to any specific configuration presented below, but covers any modifications, substitutions, and improvements to elements, components, and algorithms without departing from the spirit of this application. Well-known structures and techniques are not shown in the accompanying drawings and the following description in order to avoid unnecessary obfuscation of this application.

[0019] Furthermore, the various operations will be described as multiple discrete operations in a manner most conducive to understanding the illustrative embodiments; however, the order of description should not be construed as implying that these operations must depend on the order. In particular, these operations do not need to be performed in the order presented.

[0020] The phrases “in an embodiment,” “in one embodiment,” and “in some embodiments” are used repeatedly throughout this document. These phrases do not typically refer to the same embodiment; however, they may refer to the same embodiment. Unless the context otherwise specifies, the terms “comprising,” “having,” and “including” are synonyms. The phrases “A or B” and “A / B” mean “(A), (B) or (A and B).”

[0021] Figure 1 A block diagram of an example processor and / or system-on-a-chip (SoC) 100 is illustrated, which may have one or more cores and an integrated memory controller. The processor 100 illustrated by solid-line boxes has a single core 102(A), system proxy unit circuitry 110, and a set of one or more interface controller unit circuitry 116, while alternative processors 100 can be illustrated by dashed-line boxes having multiple cores 102(A)-(N), a set of one or more integrated memory control unit circuitry 114 in the system proxy unit circuitry 110, dedicated logic 108, and a set of one or more interface controller unit circuitry 116.

[0022] Different implementations of processor 100 may include: 1) a CPU, where dedicated logic 108 is integrated graphics and / or scientific (throughput) logic (which may include one or more cores, not shown), and cores 102(A)-(N) are one or more general-purpose cores (e.g., general-purpose ordered cores, general-purpose out-of-order cores, or a combination of both); 2) a coprocessor, where cores 102(A)-(N) are a large number of dedicated cores primarily for graphics and / or scientific (throughput) purposes; and 3) a coprocessor, where cores 102(A)-(N) are a large number of general-purpose ordered cores. Thus, processor 100 may be a general-purpose processor, a coprocessor, or a dedicated processor, such as a network or communication processor, a compression engine, a graphics processor, a GPGPU (General-Purpose Graphics Processing Unit), a high-throughput integrated many-core (MIC) coprocessor (including 30 or more cores), an embedded processor, etc. The processor may be implemented on one or more chips. The processor 100 may be part of one or more substrates and / or may be implemented on one or more substrates using any of a variety of process technologies, such as complementary metal oxide semiconductor (CMOS), bipolar CMOS (BiCMOS), P-type metal oxide semiconductor (PMOS), or N-type metal oxide semiconductor (NMOS).

[0023] The memory hierarchy includes one or more levels of cache cell circuitry 104(A)-(N) within cores 102(A)-(N), a group of one or more shared cache cell circuitry 106, and external memory (not shown) coupled to the group of integrated memory controller cell circuitry 114. The group of one or more shared cache cell circuitry 106 may include one or more intermediate level caches, such as level 2 (L2), level 3 (L3), level 4 (4), or other levels of cache, such as the last-level cache (LLC), and / or combinations thereof. While in some examples interface network circuitry 112 (e.g., a ring interconnect) provides an interface to dedicated logic 108 (e.g., integrated graphics logic), the group of shared cache cell circuitry 106, and system agent cell circuitry 110, alternative examples use any number of known techniques to interface to these units. In some examples, one or more circuits in the shared cache cell circuitry 106 maintain consistency with cores 102(A)-(N). In some examples, the interface controller unit circuit 116 couples these cores to one or more other devices 118, such as one or more I / O devices, storage devices, one or more communication devices (e.g., wireless networks, wired networks, etc.).

[0024] In some examples, one or more of cores 102(A)-(N) have multi-threading capabilities. System agent unit circuitry 110 includes those components that coordinate and operate cores 102(A)-(N). System agent unit circuitry 110 may include, for example, power control unit (PCU) circuitry and / or display unit circuitry (not shown). The PCU may be (or may include) the logic and components required to regulate the power state of cores 102(A)-(N) and / or dedicated logic 108 (e.g., integrated graphics logic). Display unit circuitry is used to drive one or more externally connected displays.

[0025] Core 102(A)-(N) can be homogeneous in terms of instruction set architecture (ISA). Alternatively, core 102(A)-(N) can also be heterogeneous in terms of ISA; that is, a subset of core 102(A)-(N) may be able to execute one ISA, while other cores may be able to execute only a subset of that ISA or be able to execute another ISA. In one embodiment, processor core 102(A)-(N) may employ all or part of the Reduced Instruction Set Computer (RISC-V) instruction set architecture. In another embodiment, processor core 102(A)-(N) may employ various other instruction set architectures, such as x86, ARM, MIPS, etc.

[0026] Some user-mode (U-mode) software (e.g., binary translation, BT) requires storage resources to store data. BT is a technology that directly translates executable binaries, enabling the translation of binaries from one processor to another for execution on another, thus facilitating easy portability between different processors and expanding the applicability of hardware and software. For example, BT technology needs to manage thread-local and frequently accessed latency-sensitive data (e.g., BT thread-local metadata), such as saving or retrieving values ​​from host physical registers into or from the target architecture's data structures. In some examples, BT thread-local metadata includes the target ISA's virtual CPU (vCPU) context, external function call pointers, simulated stack information, jump table entries, preset immediate values, temporary registers, etc. Maintaining this data is costly; for example, maintaining the vCPU context typically accounts for 25-30% of the total translated code. Since binary translation technology is considered one of the key technologies in the RISC-V software ecosystem, the increasing complexity and fragmentation of the RISC-V ISA further exacerbates the overhead of managing this data, including vCPU contexts. Similarly, other user-mode software also requires storage resources to store data, including but not limited to artificial intelligence (AI) operator acceleration, video and audio processing, and software or applications related to encryption scenarios.

[0027] The following description uses BT software as an example to illustrate some embodiments of this application. However, the principles of these embodiments are not limited to BT technology, but are applicable to any user-mode software, and this application makes no limitation thereto. Additionally, the following description uses vCPU context as an example to illustrate some embodiments of this application. However, the principles of these embodiments are not limited to BT thread-local metadata such as vCPU context data, but are applicable to any thread-local and frequently accessed latency-sensitive data, and this application makes no limitation thereto. Furthermore, the following description uses RISC-V architecture as an example to illustrate some embodiments of this application. However, the principles of these embodiments are not limited to RISC-V architecture, but are applicable to any host architecture, and this application makes no limitation thereto.

[0028] Translators (BTs) on RISC-V face architectural and fragmentation challenges. On one hand, the various optional extensions to the RISC-V ISA lead to fragmentation of the hardware register set. Register sets can differ between different RISC-V variants. Because BT software struggles to stably map target vCPU registers to hardware registers, it frequently uses data structures in main memory to store vCPU values. On RISC-V architectures, the vCPU context maintenance overhead of BTs becomes even more challenging. On the other hand, in addition to general-purpose registers (GPRs), the translator also handles special state registers such as flag registers (e.g., EFLAGS for x86) and vector registers (e.g., AVX / AVX512 for x86, SVE for ARM, or RVV for RISC-V). Simulating or mapping these complex states incurs additional overhead. Current RISC-V binary translation solutions (e.g., Box64) consistently manage these complex states within a single large context structure in slow main memory, and their performance is inevitably affected by cache misses.

[0029] In some embodiments, the BT software reduces the overhead of emulating vCPU register states by utilizing host hardware registers. Figure 2 An example of flexible allocation of host hardware registers is shown. As illustrated, guest register state is retained in available host registers. However, when host registers are scarce or during frequent function calls and context switches, CPU state may need to be saved to a structure in main memory. This frequent access to main memory for saving and restoring state creates a major performance bottleneck.

[0030] Figure 3 An example of fixed mapping to host hardware registers is shown. This fixed mapping directly assigns some client registers to host registers. This avoids the overhead of flexible on-demand allocation, but it requires addressing the issue of register number mismatches between different ISAs. For example, ARM has 31 general-purpose registers, while x86 only has 16, resulting in a large number of registers that cannot be directly mapped. This forces the translator to frequently overflow unmapped registers into main memory. Typically, a large portion (e.g., 20–30%) of the overhead in commercially available fixed-mapping BT software is dedicated to this management.

[0031] These two techniques utilizing host hardware registers face a common core contradiction: the conflict between the high-speed access advantage of hardware registers and their limited number. Flexible allocation strategies are constrained by context instability, while fixed mapping strategies are constrained by resource mismatch. Neither can fundamentally avoid frequent accesses to slow main memory.

[0032] In this application, a method is proposed to allocate a portion of the final level cache (LLC) as static random access memory (SRAM) on a host architecture (e.g., RISC-V architecture) for user-mode software such as BT software to maintain data such as vCPU context.

[0033] Figure 4 A flowchart of a method 400 for managing LLC resources according to an embodiment of this application is shown. In some embodiments, method 400 may be executed by a RISC-V processor. In another embodiment, method 400 may be executed by a processor based on any other instruction set architecture, such as x86, ARM, MIPS, etc. Method 400 may include operations 402, 404, 406, and 408. However, in some embodiments, method 400 may include more or fewer different steps, which is not limited by this application.

[0034] At position 402, a memory mapping request is received from the user-mode software. This memory mapping request is used to request that the virtual memory address of the user-mode software be mapped to the physical address of the LLC SRAM block.

[0035] At 404, in response to a memory mapping request, at least one LLC SRAM block is allocated from the LLC SRAM resource pool. This LLC SRAM resource pool is a storage resource in the LLC that is dynamically configured to SRAM mode by configuring the first control status register (CSR).

[0036] At 406, a mapping is established from the virtual memory address of the user-mode software to the physical address of the allocated LLC SRAM block, and the page table entry attribute of the mapping is configured to non-cached mode.

[0037] At 408, the allocated LLC SRAM block is enabled via the second CSR.

[0038] Method 400 dynamically uses a portion of the LLC as SRAM to store data related to user-mode software (e.g., BT software) by configuring a CSR. This method includes support for several CSRs and corresponding software workflows. This method eliminates the main memory access overhead present in previous solutions because all vCPU contexts can be stored in low-latency LLC SRAM. This avoids the overhead of accessing main memory and improves the efficiency of user-mode software on host architectures (e.g., RISC-V architecture).

[0039] In some embodiments, the first CSR includes a block control register (bcr) and a set of block address mapping registers (bamr) (including one or more bamr registers). The bcr can be used to define the number of LLC SRAM blocks in the LLC SRAM resource pool and the size of each LLC SRAM block. The bamr in this set of bamr can be used to define the physical memory address of the corresponding LLC SRAM block in the LLC SRAM resource pool. In other embodiments, the first CSR includes more or fewer or different CSRs, which is not limited in this application. For example, in some embodiments, the first CSR includes a block number register and a block size register to define the number and size of LLC SRAM blocks, respectively. As another example, in some embodiments, the first CSR includes a single block address mapping memory to define the physical memory addresses of multiple LLC SRAM blocks.

[0040] In some embodiments, the second CSR may be a block enable register (ber), which is a bitmap register in which the state of each bit corresponds to the operating mode of the corresponding LLC SRAM block in the LLC SRAM resource pool. For example, if a bit is set, it indicates that the operating mode of the corresponding LLC SRAM block is SRAM mode; if the bit is cleared, it indicates that the operating mode of the corresponding LLC SRAM block is normal LLC mode, and vice versa.

[0041] In some embodiments, a memory mapping request is initiated by user-mode software via a system call with predetermined flags. In one example, the memory mapping request can be initiated using the mmap() system call with predetermined flags (e.g., MAP_LLC). In another example, the memory mapping request can be initiated using the madvise() system call with predetermined flags (MADV_LLC). In yet another example, other system calls (e.g., mremap or the sysfs interface) can also be used to initiate the memory mapping request, and this application is not limited thereto. In some embodiments, the memory mapping request may include the amount of storage resources desired by the user-mode software.

[0042] Figure 5 An example workflow for using LLC SRAM resources based on CSR configuration according to an embodiment of this application is shown. Figure 5In the example, LLC management and BT vCPU context management are separated into two parts running in different CPU modes. LLC management is handled by the operating system kernel's memory management module, which configures the CSR bcr and bamr, which are only accessible in privileged mode. Once the CSR bcr and bamr are correctly configured, the operating system kernel can switch the corresponding LLC block to SRAM mode by setting the CSR ber. BT software running in user mode (U-mode) requests the kernel's memory management module to map the LLC SRAM block into its own process memory space by calling system calls such as mmap() or madvise(). The BT software then maintains the vCPU context in the allocated virtual memory.

[0043] exist Figure 5 In one embodiment, the kernel configures and initializes the CSRs (bcr, bamr, and ber). In other embodiments, bcr and bamr may be configured and / or initialized by firmware (machine-mode (M-mode) firmware), as detailed later. Figure 5 In one embodiment, the kernel configures the page table of the memory management unit (MMU) to establish a "path" so that when the binary translator accesses a virtual address, the MMU can automatically redirect it to the physical address of a specific block (e.g., block 0) in the LLC (for SRAM purposes).

[0044] In some embodiments, these CSRs are configured on a per-processor-core basis. In other embodiments, these CSRs are configured to be shared across multiple or all processor cores. This application is not limited in this respect.

[0045] In this document, the terms "bcr register", "CSR bcr", "bcr CSR", and "bcr" are used interchangeably. Similarly, the terms "bamr register", "CSR bamr", "bamr CSR", and "bamr" are used interchangeably. Likewise, the terms "ber register", "CSR ber", "ber CSR", and "ber" are used interchangeably.

[0046] Figure 5 An example of a process is provided for allocating a portion of LLC as SRAM for user-mode software such as BT software to maintain local metadata such as BT thread metadata. This approach proposed in this application will be described in more detail below.

[0047] First, a programming interface was introduced for host architectures such as RISC-V, which partitions a portion of the LLC into SRAM blocks. This interface includes (for example, added in S-mode) three new CSRs: bcr, bamr, and ber. These CSRs for each RISC-V core in the system are mapped to the same system-level hardware control block, which manages the configuration of the LLC controller and the on-chip network (NoC). Therefore, writes to these CSRs on any core are visible to all hardware threads. The operating system kernel does not need to save or restore these CSRs during thread context switching.

[0048] In some embodiments, the bcr register may be a read / write register with a user-mode extended length (UXLEN) bit width, used to store configuration information of the LLC SRAM block, such as the number of blocks (COUNT), the operating mode (MODE), and / or the block size (SIZE). UXLEN may represent the integer bit width in the current mode, such as 32 bits or 64 bits.

[0049] Figure 6 An example of the bcr register according to an embodiment of this application is shown. The COUNT field in the bcr register can be used to store the total number of valid blocks in the LLC that can be configured as SRAM. This COUNT field can be a WriteAny Read Legal (WARL) field, allowing any value to be written, but the value read must be a hardware-supported and legal value to ensure software compatibility.

[0050] In some embodiments, the encoding of the MODE field in the bcr register is shown in Table 1. When MODE=0, the bcr and bamr CSRs are read-only for S-mode; when MODE=1, they are WARL for S-mode. In one embodiment, the bits of this field are set by the M-mode firmware according to platform design requirements. This is essentially a permission control switch, allowing the M-mode firmware to decide whether to delegate LLC configuration rights to the operating system kernel.

[0051] Encoding of the MODE field in Table 1bcr In some embodiments, the SIZE field in the bcr register is encoded as shown in Table 2. This field stores the effective size of the LLC configurable SRAM block. Its granularity depends on the specific platform and is related to the capabilities of the LLC controller IP, as well as the power consumption and frequency considerations of the SoC platform design. For example, 0 represents 64KB, and 1 represents 2MB.

[0052] The encoding of the SIZE field in Table 2bcr In some embodiments, the bcr register is set to its default value after reset, for example, all bits are set to 0 by default, all bits are set to 1 by default, and so on.

[0053] Figure 7 An example of a bamr register according to an embodiment of this application is shown. In some embodiments, the bamr register is a WARL register used to store the valid physical address of the corresponding LLC SRAM block. These SRAM physical memory addresses in the bamr registers can be reserved by the address space allocation of a host system such as a RISC-V platform. Each bamr register in the bamr register group corresponds to a corresponding SRAM block defined in the bcr (e.g., a specific SRAM block), storing the physical address of that SRAM block in the system's physical memory address space. This ensures that the operating system can access these SRAM blocks through the MMU as if they were ordinary memory regions.

[0054] In some embodiments, the bamr register is set to its default value after reset, for example, all bits are set to 0 by default, all bits are set to 1 by default, and so on.

[0055] In some embodiments, bcr and bamr can be configured by the M-mode Supervised Binary Interface (SBI) firmware or the S-mode OS kernel according to platform-specific pre-configurations. For the S-mode OS kernel, this can be achieved by reading these CSRs or obtaining these pre-configuration values ​​from the Advanced Configuration and Power Interface (ACPI) or Device Tree Source (DTS) configuration tables provided by the Basic Input / Output System (BIOS). If the M-mode SBI firmware wishes to take over configuration control of bcr and bamr, it can set the MODE bit of bcr to 0, thus preventing the S-mode kernel from modifying them.

[0056] Figure 8 An example of a ber register according to an embodiment of this application is shown. In some embodiments, the ber register stores a bitmap for enabling or disabling SRAM blocks in a reserved LLC portion. Each bit may correspond to a specific block in the bcr register. The ber register can be configured by the S-mode operating system kernel. For example, the ber register is configured when the S-mode operating system kernel needs to switch certain LLC blocks to SRAM mode and map them to the BT software address space in user space.

[0057] In some embodiments, when a block is enabled via ber (e.g., a bit in ber changes from 0 to 1, or vice versa), the hardware control logic may (e.g., must) first automatically refresh the corresponding LLC block. This ensures the consistency of cached data and prevents old data from contaminating new SRAM space. In some embodiments, the data of the allocated LLC SRAM block is automatically written back to main memory before or simultaneously with enabling the allocated LLC SRAM block via ber.

[0058] In some embodiments, the ber register is set to its default value after reset, for example, all bits are set to 0 by default, all bits are set to 1 by default, and so on.

[0059] In some embodiments, the hardware control logic related to bcr, bamr, and ber CSR that is responsible for configuring the LLC controller and NoC controller can also be implemented through M-mode SBI firmware, for example, by monitoring write accesses to the ber register.

[0060] The above provides some examples for the bcr, bamr, and ber registers. The following details the software workflow for managing LLC SRAM blocks and the operating system kernel application programming interface (API).

[0061] Figure 9 An example of LLC SRAM software management according to an embodiment of this application is shown. Figure 9 As shown, in order to efficiently utilize LLC SRAM blocks, this application introduces three key software workflows: LLC SRAM initialization process, LLC SRAM mapping process, and LLC SRAM allocation process.

[0062] During the LLC SRAM initialization process, the M-mode firmware or S-mode operating system kernel can configure the aforementioned bcr, bamr, and ber registers, and subsequently perform read or write operations on them.

[0063] In some embodiments, the M-mode SBI firmware can first set the bcr CSR based on the system's pre-configuration to define the reserved LLC SRAM block size and number of blocks, while setting the bamr CSR group using the physical memory address for each SRAM block specified in the bcr CSR.

[0064] In some embodiments, the block size and number of blocks in the bcr CSR can depend on the capabilities of the LLC controller and the design of the RISC-V SoC; while the corresponding SRAM physical memory address in the bamr CSR is reserved by the address space allocation of the RISC-V platform.

[0065] In some embodiments, during the S-mode operating system kernel boot process, the kernel's LLC SRAM management module can first check the MODE field / bit of the CSR bcr. If MODE is set to 0, it means that the CSR bcr and bamr have been initialized by the SBI firmware, and the kernel can initialize the LLC SRAM pool metadata based on the initial values ​​of the CSR bcr and bamr; otherwise, if the MODE bit is set to 1, the kernel module needs to initialize the CSR bcr and bamr according to the ACPI table provided by the platform BIOS or the configuration in the DTS. Thus, the MODE field / bit provides two flexible boot paths: hardware preset and kernel dynamic configuration, adapting to different platform strategies.

[0066] In the LLC SRAM mapping process, in some embodiments, the BT software can interact with the operating system kernel through the system call mmap() with the added mapping flag MAP_LLC or the system call madvise() with the added mapping flag MADV_LLC to request the kernel to establish a memory mapping from the thread virtual memory to the reserved LLC SRAM block.

[0067] In some embodiments, upon receiving such a request, the operating system kernel may perform the following workflow: a) first allocate a physical SRAM block from the LLC SRAM pool; b) then map the virtual memory to the physical address of the allocated SRAM block and set the Page-Based Memory Type (pbmt) field of the page table entry (PTE) to the non-cached (UC) mode of the main memory with non-cached, idempotent, and weak-order attributes; and c) finally set the corresponding bit in the ber register to enable the corresponding LLC SRAM block.

[0068] Figure 10 An example of an LLC SRAM mapping process according to an embodiment of this application is shown. Figure 10 As shown, once the kernel completes the mapping process, the corresponding LLC block will function as low-latency SRAM.

[0069] In some embodiments of the LLC SRAM allocation process, as the mmap() / madvise() system call returns successfully, the BT software can access these mapped virtual memories as if they were low-latency memories, and put latency-sensitive metadata such as the vCPU context into the allocated SRAM blocks.

[0070] For user-mode software such as BT software, these allocated virtual memories are large and contiguous memory spaces. In some embodiments, when user-mode software such as BT software exits or wishes to release these SRAMs back to the kernel, it can invoke system calls (e.g., munmap()) to return these memories to the kernel.

[0071] In some embodiments, the operating system kernel can prepare a dedicated LLC SRAM pool during kernel initialization to manage reserved LLC SRAM. Once these LLC blocks need to be mapped into the address space of the BT software, they can first be allocated from the kernel's LLC SRAM pool. In some embodiments, the kernel can manage this pool based on the physical addresses of the SRAM, regardless of whether these addresses are contiguous. For example, when the BT software requests the allocation of LLC SRAM blocks, the operating system kernel can utilize the MMU's page tables to construct contiguous virtual memory for the BT software's use. When the BT software releases the SRAM back to the operating system kernel, the kernel can put them back into the LLC SRAM pool and reconfigure these LLC SRAMs as normal LLCs via CSR ber. This achieves dynamic resource reuse and avoids idle waste.

[0072] The management solution for LLC resources proposed in this application addresses at least the following challenges.

[0073] Challenge 1 involves providing a platform-independent programming interface to partition and manage LLCs for use as SRAM blocks. M-mode firmware or an S-mode operating system kernel (e.g., in S-mode) can configure bcr CSR and bamr CSR groups. This can trigger hardware control logic or M-mode firmware to configure the LLC controller and NoC controller, thereby allocating a portion of the LLC for SRAM use. The block size and number of blocks in the bcr CSR depend on the capabilities of the LLC controller and the design considerations of the target architecture (e.g., RISC-V) SoC; while the corresponding SRAM physical memory addresses in the bamr CSR can be reserved by the address space allocation of the target architecture (e.g., RISC-V platform). These configuration values ​​can be provided by the platform vendor through platform-specific ACPI tables or DTS configuration. For the operating system kernel, the programming interface can be one of the three CSRs for RISC-V: bcr, bamr, and ber.

[0074] Challenge 2 involves user-mode software, such as BT software, seamlessly accessing these SRAM blocks as low-latency memory space to maintain latency-sensitive data such as vCPU context. The operating system kernel manages these SRAM blocks as physical memory segments and marks their region attributes as write-through. When user-mode software, such as BT software, calls, for example, the mmap() system call with the MAP_LLC flag or the madvise() system call with the MADV_LLC flag to request the kernel to establish a memory mapping, the kernel sets the pbmt field of the PTE in the page table to UC mode (uncached, idempotent, weakly ordered (RVWMO) main memory). User-mode software, such as BT software, accesses these virtual memories as low-latency memory and then manages them according to the memory management design within the user-mode software, such as BT software.

[0075] Challenge three involves efficiently utilizing LLCs to achieve better performance. When no user-mode software, such as BT software, is running, the reserved LLC portion operates as a normal LLC. Once the operating system kernel receives mmap() / madvise() requests from user-mode software such as BT software, the kernel can allocate some SRAM blocks from the LLC based on the configuration of the bcr and bamr CSRs and the requested LLC SRAM size. Even if the SRAM physical addresses specified in the bamr CSR are not contiguous, the operating system kernel can still use the MMU's page tables to construct a large, contiguous virtual memory space for the BT software to use. When the BT software exits or releases these SRAMs back to the kernel via munmap(), the operating system kernel can reconfigure these LLC blocks into normal mode via the bamr CSR. The kernel can establish a dedicated LLC SRAM pool to manage these LLC SRAM resources.

[0076] This application defines a method for allocating a portion of LLC as SRAM for user-mode software, such as BT software, on a host such as a RISC-V architecture to maintain data such as vCPU context. The method includes support for several CSRs (e.g., by adding bcr CSR, bamr CSR, and ber CSR, for example, in S-mode) and corresponding software workflows to eliminate main memory access overhead. For example, software workflow and API examples include: the M-mode SBI firmware or the S-mode operating system kernel sets the bcr CSR based on the system's preset configuration to define the size and number of reserved LLC SRAM blocks, and sets the bamr CSR group using the physical memory address of each SRAM block specified in the bcr CSR. The BT software interacts with the operating system kernel via the mmap() or madvise() system calls by adding the mapping flag MAP_LLC or MADV_LLC to request the kernel to establish a memory mapping from thread virtual memory to the reserved LLC SRAM blocks. Upon receiving the request, the kernel first allocates a physical SRAM block and maps the virtual memory to the physical address of the allocated SRAM block. Finally, it sets the corresponding bit in ber to enable the LLC SRAM block.

[0077] Based on the technical solution of this application, a performance improvement of approximately 8.9% or even higher can be achieved for the vCPU context in typical use cases. This performance gain is expected to be further optimized to exceed 10%, because the accelerated objects of SRAM are not limited to vCPU state; it can also accelerate any thread-local and frequently accessed data, such as "BT thread-local metadata." Below are some examples of BT thread-local metadata that can also be accelerated using the technical solution of this application.

[0078] For example, storing frequently used external function call pointers in SRAM ensures fast lookup when the translator needs to call functions in the host native library.

[0079] For example, for emulating stack information, binary translators typically manage a shadow stack for the client program. Storing critical stack pointers and return addresses in SRAM enables extremely low-latency access, which is crucial for the core of any program's execution.

[0080] For example, for jump table entries, BT needs to find the correct jump target for switch-case statements or virtual method dispatches. Storing these entries in SRAM ensures minimal latency for branch jumps.

[0081] For example, for preset immediate values, the fixed 32-bit instruction size of RISC-V typically requires multiple instructions to generate a single large constant. Using SRAM, these frequently used immediate values ​​can be stored in high-speed on-chip memory, allowing translated code to be executed in a single, fast load, rather than using multiple instructions.

[0082] For example, binary translators often face the problem of insufficient host registers for temporary registers, forcing temporary values ​​to overflow into main memory. By using SRAM, the translator can offload these temporary registers to high-speed on-chip locations, thereby effectively expanding the available register pool.

[0083] This application eliminates slow memory accesses generated by user-mode software such as BT software when saving / restoring data such as vCPU context, and significantly reduces ISA emulation overhead. Using the technical solution of this application, user-mode software such as BT software can achieve a performance improvement of approximately 8.9% or even higher in typical usage scenarios.

[0084] Figure 11 This is a block diagram illustrating components capable of reading instructions from a machine-readable or computer-readable medium (e.g., a non-transitory machine-readable storage medium) and performing any one or more of the methods discussed herein, according to some example embodiments. Specifically, Figure 11 A schematic diagram of hardware resource 1100 is shown, which includes one or more processors (or processor cores) 1110, one or more memory / storage devices 1120, and one or more communication resources 1130, wherein each of these processors, memory / storage devices, and communication resources can be communicatively coupled via bus 1140 or other interface circuitry. For embodiments utilizing node virtualization (e.g., Network Functions Virtualization (NFV)), a hypervisor 1102 can be executed to provide an execution environment for one or more network slices / subslices, thereby utilizing hardware resource 1100.

[0085] Processor 1110 may include, for example, a first processor 1112 and a second processor 1114. Processor 1110 may be, for example, a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP) such as a baseband processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a radio frequency integrated circuit (RFIC), another processor (including those discussed herein), or any suitable combination thereof.

[0086] The memory / storage device 1120 may include main memory, disk storage devices, or any suitable combination thereof. The memory / storage device 1120 may include, but is not limited to, any type of volatile, non-volatile, or semi-volatile memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, solid-state memory, etc.

[0087] Communication resource 1130 may include interconnect or network interface controllers, components, or other suitable devices for communicating with one or more peripheral devices 1104 or one or more databases 1106 or other network elements via network 1108. For example, communication resource 1130 may include wired communication components (e.g., for coupling via USB, Ethernet, etc.), cellular communication components, near field communication (NFC) components, Bluetooth® (or Bluetooth® Low Energy) components, Wi-Fi® components, and other communication components.

[0088] Instructions 1150 may include software, programs, application programs, applets, or other executable code for causing at least any one of processors 1110 to perform any one or more of the methods discussed herein. Instructions 1150 may reside wholly or partially within processor 1110 (e.g., in the processor's cache), memory / storage device 1120, or any suitable combination thereof. Furthermore, any portion of instructions 1150 may be transferred from any combination of peripheral device 1104 or database 1106 to hardware resource 1100. Therefore, the memory of processor 1110, memory / storage device 1120, peripheral device 1104, and database 1106 are examples of computer-readable and machine-readable media.

[0089] Some examples may be implemented or be implemented as an article of art or at least a computer-readable medium. The computer-readable medium may include a non-transitory storage medium for storing logic. In some examples, the non-transitory storage medium may include one or more types of computer-readable storage media capable of storing electronic data, including volatile or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, and so on. In some examples, the logic may include various software elements, such as software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, APIs, instruction sets, computational code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof.

[0090] According to some examples, computer-readable media may include non-transitory storage media to store or maintain instructions that, when executed by a machine, computing device, or system, cause that machine, computing device, or system to perform methods and / or operations according to the described examples. Instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, etc. Instructions may be implemented according to a predetermined computer language, manner, or syntax to instruct a machine, computing device, or system to perform specific functions. Instructions may be implemented using any suitable high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language.

[0091] One or more aspects of at least one example can be implemented by representative instructions representing various logic within a processor, stored on at least one machine-readable medium, which, when read by a machine, computing device, or system, cause the machine, computing device, or system to manufacture logic to perform the techniques described herein. This representation, referred to as an "IP core," can be stored on a tangible machine-readable medium and provided to various customer or manufacturing facilities for loading into the manufacturing machine that actually manufactures the logic or processor.

[0092] The phrase "an example" or "an example" does not necessarily refer to the same example or embodiment. Any aspect described herein may be combined with any other aspect or similar aspect described herein, whether or not these aspects are described with reference to the same drawings or elements. The division, omission, or inclusion of block functions depicted in the drawings does not imply that hardware components, circuits, software, and / or elements used to implement these functions will necessarily be divided, omitted, or included in the embodiments.

[0093] Examples can be described using the terms “coupling” and “connection” and their derivatives. These terms are not necessarily intended to be synonyms. For example, a description using the terms “connection” and / or “coupling” may indicate that two or more elements are in direct physical or electrical contact with each other. However, the term “coupling” may also refer to two or more elements that are not in direct contact with each other but still cooperate or interact with each other.

[0094] The terms “first,” “second,” and the like are used herein not to indicate any order, quantity, or importance, but rather to distinguish one element from another. The term “a” herein does not imply a limitation on quantity, but rather indicates the presence of at least one mentioned item. The term “assertion” as used herein when referring to a signal refers to a state in which the signal is valid and can be achieved by applying any logic level (whether logic 0 or logic 1) to the signal. The terms “subsequently” or “after” can mean immediately following or following one or more other events. According to alternative embodiments, other sequences of steps may also be performed. Furthermore, depending on the specific application, additional steps may be added or removed. Any combination of variations can be used, and many variations, modifications, and alternative embodiments will be understood by those skilled in the art who benefit from this application.

[0095] Unless otherwise specifically stated, disjunctive language such as the phrase "at least one of X, Y, or Z" is understood in context to generally state that an item, term, etc., can be X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is generally not intended, nor should it imply, that certain embodiments require the presence of each of at least one X, at least one Y, or at least one Z. Furthermore, unless otherwise specifically stated, connective language such as the phrase "at least one of X, Y, and Z" should also be understood to refer to X, Y, Z, or any combination thereof, including "X, Y, and / or Z".

Claims

1. A method for managing last-level cache LLC resources, characterized in that, The method includes: Receive a memory mapping request from user-mode software, the memory mapping request being used to request that the virtual memory address of the user-mode software be mapped to the physical address of an LLC static random access memory (SRAM) block; In response to the memory mapping request, at least one LLC SRAM block is allocated from the LLC SRAM resource pool, wherein the LLC SRAM resource pool is a storage resource in the LLC that is dynamically configured to SRAM mode by configuring a first control status register; Establish a mapping from the virtual memory addresses of the user-mode software to the physical addresses of the allocated LLC SRAM blocks, and configure the page table entry attributes of this mapping to non-cached mode; and The allocated LLC SRAM block is enabled via the second control status register.

2. The method according to claim 1, characterized in that, The first control status register includes a block control register and a set of block address mapping registers, wherein: The block control register is used to define the number of LLC SRAM blocks in the LLC SRAM resource pool and the size of each LLC SRAM block. The block address mapping register in the set of block address mapping registers is used to define the physical memory address of the corresponding LLC SRAM block in the LLC SRAM resource pool.

3. The method according to claim 2, characterized in that, The block control register and the set of block address mapping registers are configured by the machine-mode firmware or supervisory operating system kernel of the RISC-V architecture.

4. The method according to claim 2, characterized in that, The block control register and the set of block address mapping registers are configured according to the configuration information provided by the system.

5. The method according to claim 2, characterized in that, The block control register includes a mode field, wherein: When the mode field is the first value, the regulatory mode operating system kernel is prohibited from writing to the block control register and the set of block address mapping registers; When the mode field is the second value, the regulatory mode operating system kernel is allowed to write to the block control register and the set of block address mapping registers.

6. The method according to claim 1, characterized in that, The memory mapping request is initiated by user-mode software through a system call with a predetermined flag.

7. The method according to claim 1, characterized in that, The non-cached mode includes the non-cached main memory mode defined in the RISC-V architecture that has non-cached, idempotent, and weak-order properties.

8. The method according to claim 1, characterized in that, The second control status register is a bitmap register, and the state of each bit of the bitmap register corresponds to the working mode of the corresponding LLC SRAM block in the LLC SRAM resource pool.

9. The method according to claim 8, characterized in that, Before or simultaneously with enabling the allocated LLC SRAM block via the second control status register, the data of the LLC SRAM block is automatically written back to main memory.

10. The method according to claim 1, characterized in that, The method further includes: When the user-mode software releases the allocated LLC SRAM block, it disables the LLC SRAM block via the second control status register to release it back to the LLC.

11. The method according to claim 1, characterized in that, The method further includes: Establish a contiguous virtual address space mapping for at least one allocated LLC SRAM block, regardless of whether the physical addresses of the at least one allocated LLC SRAM block are contiguous.

12. The method according to claim 1, characterized in that, The user-mode software includes binary translation software.

13. The method according to claim 1, characterized in that, The allocated LLC SRAM block is used to store latency-sensitive data.

14. The method according to claim 13, characterized in that, The latency-sensitive data includes local metadata of the binary translation thread.

15. An apparatus for managing last-level cache LLC resources, characterized in that, The apparatus includes a processor circuit configured to perform the method of any one of claims 1-14.

16. A computer-readable storage medium having instructions stored thereon, characterized in that, When executed by a processor, the instructions cause the processor to perform the method according to any one of claims 1-14.

17. A computer program product comprising instructions, characterized in that, When executed by a processor, the instructions cause the processor to perform the method according to any one of claims 1-14.