Techniques for generating system cache partitioning policies

CN116303127BActive Publication Date: 2026-09-25NVIDIA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211503595.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-12-21
Filing Date
2022-11-28
Publication Date
2026-09-25
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

因此,不同处理单元对系统高速缓存的低效共享也会导致功耗增加

Benefits of technology

[0009]所公开的技术相对于现有技术的至少一个技术优势在于,利用所公开的技术,对系统高速缓存的访问在多个处理单元之间更有效地共享。特别是,系统高速缓存的不同部分被分配给每个处理单元。每个处理单元都对系统高速缓存的其他部分具有读取访问权限,但仅将数据写入其分配的系统高速缓存的部分。因此,每个处理单元不会将数据写入系统高速缓存的被分配给其他处理单元的部分或覆盖存储在系统高速缓存的被分配给其他处理单元的部分中的数据。因此,与以前的方法相比,缓存的数据不太频繁地移除或覆盖,这与以前的方法相比也导致了功耗降低。此外,使用所公开的技术,确定分配给每个处理单元的系统高速缓存的部分的最佳尺寸,使得系统的整体功耗最小化。这些技术优势提供了优于现有技术方法的一个或更多个技术进步。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116303127B_ABST
    Figure CN116303127B_ABST
Patent Text Reader

Abstract

The present disclosure relates to techniques for generating a system cache partitioning policy. In various embodiments, a computing system includes a plurality of processing units that share access to a system cache, for example. A cache management application receives resource savings information for each processing unit, for example. The resource savings information indicates an amount of resource (e.g., power) saved when different units of the system cache are allocated to the processing units, for example. The cache management application determines a number of units of the system cache to allocate to each processing unit based on the received resource savings information, for example.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure generally relate to computer memory and processing, and more specifically, to techniques for generating system cache partitioning strategies. Background Technology

[0002] Various processing units in a computer system, such as the Central Processing Unit (CPU) or Graphics Processing Unit (GPU), can access cache memory used to store temporary copies of data. Applications executing on these processing units can cache data in the corresponding cache memory for future access. For example, an application can read data from a storage device and store a copy of that data in the cache memory. Subsequently, if the application needs the same data, it can read the cached data from the cache memory instead of reading it from the storage device. Because accessing cache memory is faster than reading data from storage or recalculating values, using cached data allows hardware units to increase processing throughput. When the cache memory is full, some existing data is removed from the cache memory and replaced with the newly added data to cache additional data.

[0003] One type of computer system is a System-on-a-Chip (SoC), which integrates multiple different processing units, such as one or more CPUs, GPUs, and / or other types of data processing units, on a single chip. An SoC may include a system cache accessible to most (if not all) of these different processing units. Each processing unit can cache data in the system cache and read cached data stored by other processing units. For example, a GPU can cache the results of its computations in the system cache, and then the CPU can read the results of its computations instead of waiting for the GPU to transfer them. Therefore, the system cache also allows processing units to efficiently transfer data to other processing units.

[0004] However, one problem with sharing the system cache across multiple processing units is that the total cache size required by these multiple processing units often exceeds the system cache size. Therefore, when an application caches data in the system cache, that data is often quickly removed from the cache by other applications that also cache data in the system cache. However, if the data is lost from the cache when an application needs it, the application must retrieve the data again, for example, by reading it from a different storage device or storage location. Thus, different processing units on a system-on-a-chip cannot effectively utilize the system cache to improve processing speed.

[0005] Furthermore, System-on-a-Chip (SoC) is frequently used in battery-powered devices where low power consumption is desired, such as mobile phones or laptops. Reading data from storage devices or recalculating data uses more power than reading data from the system cache. Therefore, inefficient sharing of the system cache by different processing units also leads to increased power consumption.

[0006] Furthermore, in some battery-powered and / or non-battery-powered devices, SOC consumption is not allowed to exceed a fixed power budget. In this case, reducing the amount of power consumed by different processing units at a given performance level allows them to operate at higher performance levels, while increasing the amount of power consumed by different processing units results in them operating at lower performance levels, such as at lower processing speeds.

[0007] As mentioned above, what is needed in the art is a more efficient way to manage system caches shared by multiple processing units. Summary of the Invention

[0008] One embodiment of this disclosure illustrates a computer-implemented method for managing a system cache shared by multiple processing units. The method includes receiving resource-saving information for each of the multiple processing units, the resource-saving information specifying the amount of corresponding resources saved when different system cache units are allocated to the processing unit. The method also includes, for each of the multiple processing units, calculating, at least based on the resource-saving information associated with the multiple processing units, the number of system cache units to be allocated to the processing unit.

[0009] At least one technical advantage of the disclosed technique over the prior art lies in the more efficient sharing of access to the system cache among multiple processing units. Specifically, different portions of the system cache are allocated to each processing unit. Each processing unit has read access to other portions of the system cache but only writes data to its allocated portion. Therefore, each processing unit does not write data to or overwrite data stored in portions of the system cache allocated to other processing units. Consequently, cached data is removed or overwritten less frequently compared to previous methods, resulting in reduced power consumption. Furthermore, using the disclosed technique, the optimal size of the portion of the system cache allocated to each processing unit is determined, minimizing the overall power consumption of the system. These technical advantages provide one or more technical advancements superior to existing methods. Attached Figure Description

[0010] To gain a more detailed understanding of the features of the various embodiments described above, the briefly summarized inventive concept can be described in more specific terms with reference to various embodiments (some of which are illustrated in the accompanying drawings). However, it should be noted that the accompanying drawings illustrate only typical embodiments of the inventive concept and are therefore not intended to limit the scope in any way; other equally effective embodiments exist.

[0011] Figure 1 This is a block diagram illustrating a computer system configured to implement one or more aspects of various embodiments;

[0012] Figure 2 It is integrated into a system-on-a-chip (SOC) according to various embodiments. Figure 1 A block diagram of a CPU and one or more processing units;

[0013] Figure 3 According to various embodiments Figure 1 A more detailed illustration of a cache management application and one or more drives;

[0014] Figure 4A The diagram illustrates various embodiments of the relationship between the two. Figure 2 Example bandwidth saving graph associated with the processing unit;

[0015] Figure 4B The diagram illustrates various embodiments of the relationship between the two. Figure 2 An exemplary power-saving graph associated with the processing unit; and

[0016] Figure 5 This is a flowchart of method steps for automatically generating system cache partitioning strategies according to various embodiments. Detailed Implementation

[0017] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the various embodiments. However, it will be apparent to those skilled in the art that the inventive concept can be practiced without one or more of these specific details.

[0018] System Overview

[0019] Figure 1 This is a block diagram illustrating a computer system 100 configured to implement one or more aspects of the present invention, according to one embodiment. As shown, the computer system 100 includes, but is not limited to, a central processing unit (CPU) 102 and a system memory 104, which are coupled to one or more processing units 112 via a memory bridge 105 and a communication path 113. The memory bridge 105 is further coupled to an I / O (input / output) bridge 107 via a communication path 106, and the I / O bridge 107 is in turn coupled to a switch 116.

[0020] In operation, I / O bridge 107 is configured to receive user input from input device 108 (such as a keyboard or mouse) and forward the input to CPU 102 for processing via communication path 106 and memory bridge 105. Switch 116 is configured to provide connectivity between I / O bridge 107 and other components of computer system 100, such as network adapter 118 and various add-on cards 120 and 121.

[0021] As also shown in the figure, I / O bridge 107 is coupled to system disk 114, which can be configured to store content, applications, and data for use by CPU 102 and one or more processing units 112. Generally, system disk 114 provides non-transitory storage for applications and data and may include fixed or removable hard disk drives, flash memory devices, and CD-ROMs (optical disc read-only memory), DVD-ROMs (digital versatile discs), Blu-ray, HD-DVDs (high-definition DVDs), or other magnetic, optical, or solid-state storage devices. Finally, although not explicitly shown, other components (such as universal serial bus or other port connections, optical disc drives, digital versatile disc drives, film recording devices, etc.) may also be connected to I / O bridge 107.

[0022] In various embodiments, memory bridge 105 may be a northbridge chip, and I / O bridge 107 may be a southbridge chip. Furthermore, communication paths 106 and 113, as well as other communication paths, can be implemented within computer system 100 using any technically suitable protocol (including but not limited to AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol known in the art).

[0023] In some embodiments, one or more processing units 112 include one or more graphics processing units that feed pixels to a display device 110, which may be any conventional cathode ray tube, liquid crystal display, light-emitting diode display, etc. The one or more graphics processing units incorporate circuitry optimized for graphics and video processing, including, for example, video output circuitry. Such circuitry may be incorporated across one or more parallel processing units (PPUs) included in the one or more graphics processing units. In some embodiments, one or more processing units 112 include processing units that incorporate circuitry optimized for general and / or computational processing. Again, such circuitry may be incorporated across one or more PPUs included in the processing unit, which are configured to perform such general and / or computational operations. In some embodiments, the one or more PPUs included in the processing unit 112 may be configured to perform graphics processing, general processing, and computational processing operations.

[0024] System memory 104 includes one or more device drivers 103 configured to manage processing operations within one or more processing units 112 and / or one or more PPUs. System memory 104 also includes multiple software applications executing on CPU 102, such as cache management application 125. As explained in further detail below, cache management application 125 analyzes bandwidth saving and / or power saving information received from one or more device drivers 103 and generates a system cache partitioning strategy.

[0025] In various embodiments, one or more processing units 112 may be connected with Figure 1 One or more other components can be integrated to form a single system. For example, as discussed in further detail below, one or more processing units 112 can be integrated with CPU 102 and other interconnect circuitry on a single chip to form a system-on-a-chip (SoC).

[0026] It should be understood that the system illustrated herein is illustrative, and variations and modifications are possible. The connection topology (including the number and arrangement of bridges, the number of CPUs 102, and the number of parallel processing subsystems 112) can be modified as needed. For example, in some embodiments, system memory 104 may be directly connected to CPU 102 instead of being connected to CPU 102 via memory bridge 105, and other devices will communicate with system memory 104 via memory bridge 105 and CPU 102. In other alternative topologies, one or more processing units 112 may be connected to I / O bridge 107 or directly to CPU 102 instead of being connected to memory bridge 105. In other embodiments, I / O bridge 107 and memory bridge 105 may be integrated into a single chip rather than existing as one or more discrete devices.

[0027] System Cache Overview

[0028] Figure 2 It includes various embodiments. Figure 1 A block diagram of a system-on-a-chip (SOC) 200 comprising a CPU 102 and one or more processing units 112. The CPU 102 and one or more processing units 112(1)-(N) are integrated on a single chip. As shown, one or more processing units 112 include an additional CPU 112(1), a graphics processing unit (GPU) 112(2), a deep learning accelerator (DLA) 112(3), an encoding accelerator (ENC) 112(4), and a decoding accelerator (DEC) 112(5). Although Figure 2N processing units 112 are depicted, but the SOC 200 may include any number of processing units 112, including more or fewer processing units. Additionally, the SOC 200 may include... Figure 2 One or more processing units of different types, not shown in the figure.

[0029] In some embodiments, CPU 102 processes data asynchronously from one or more processing units 112. Additionally, each of the one or more processing units 112(1)-(N) may asynchronously process data from CPU 102 and other processing units 112. In some embodiments, each processing unit 112 is configured for general-purpose processing, computational processing, and / or special-purpose processing. For example, GPU 112(2) may be configured to implement a graphics rendering pipeline to perform various operations related to generating pixel data based on graphics data provided by CPU 102, one or more other processing units 112, and / or system memory 104.

[0030] like Figure 2 As shown, the SOC 200 includes a system cache 202 and is coupled to dynamic random access memory (DRAM) 210. The CPU 102 and one or more processing units 112(1)-(N) can access the system cache 202 and DRAM 210. The system cache 202 and DRAM 210 can be used to store and update data generated by each of the CPU 102 and one or more processing units 112(1)-(N). Additionally, as shown, the system cache 202 is divided into multiple partitions 202(1)-(N). Each partition is assigned to a corresponding processing unit, such as the CPU 102 or one or more processing units 112. Although... Figure 2 The same number of partitions as the processing units are depicted, but the system cache 202 can be divided into any number of partitions and each partition can be allocated to any number of processing units.

[0031] In some embodiments, allocating a partition to a processing unit includes granting the processing unit read and write access to a portion of the system cache included in the partition. Allocating a partition to a processing unit may also include denying write access to other processing units. For example, partition 204(1) is allocated to CPU 102. CPU 102 has read and write access to portion 204(1). One or more processing units 112, such as CPU 112(1), GPU 112(2), etc., do not have write access to portion 204(1).

[0032] In some embodiments, one or more processing units not assigned to a partition retain read access to portions of the system cache included in that partition. For example, partition 204(3) is assigned to GPU 112(2). GPU 112(2) can store output data in partition 204(3). CPU 102 or other processing units 112 can read the output data generated by GPU 112(2) from partition 204(3). Similarly, GPU 112(2) can read the output data generated by CPU 1092 or other processing units 112 from the partitions to which they are assigned. In this way, CPU 102 and one or more processing units 112 can transfer data to each other via system cache 202.

[0033] In some embodiments, if data cached in system cache 202 is evicted from system cache 202, the data is stored in DRAM 210. When CPU 102 or processing unit 112 needs data, CPU 102 or processing unit 112 first determines whether the data is in system cache 202. If the data is cached in system cache 202, CPU 102 or processing unit 112 reads the data from system cache 202. If the data is not cached in system cache 202, CPU 102 or processing unit 112 determines whether the data is in DRAM 210. If the data is stored in DRAM 210, CPU 102 or processing unit 112 reads the data from DRAM 210.

[0034] As discussed in further detail below, reading data from DRAM 210 consumes more power than reading data from system cache 202. Therefore, reducing the number of times CPU 102 and one or more processing units 112 read data from DRAM 210 instead of from system cache 202 reduces the amount of power consumed during operation of system 100.

[0035] System cache partitioning strategy

[0036] Figure 3 According to various embodiments Figure 1 A more detailed illustration of the cache management application 125 and one or more drives 103. One or more drives include one or more drives corresponding to the CPU 102 and one or more processing units 112. (See attached diagram.) Figure 3As shown, one or more drivers 103 include a CPU driver 103(1), a CPU driver 103(2), a GPU driver 103(3), a DLA driver 103(4), an ENC driver 103(5), a DEC driver 103(6), and a processing unit driver 103(N) corresponding to a CPU 102, a CPU 112(1), a GPU 112(2), a DLA 112(3), an ENC 112(4), a DEC 112(5), and a processing unit 112(N), respectively. Each driver is configured to manage the processing operations of the corresponding processing unit.

[0037] In some embodiments, each drive is configured to calculate and / or store resource saving information 302 associated with the corresponding processing unit. Resource saving information 302 indicates the amount of resources saved for each unit (e.g., megabytes) of the system cache allocated to the processing unit. Resource saving information 302 includes, for example, power saving information and / or bandwidth saving information. Bandwidth saving information for the corresponding processing unit indicates the amount of DRAM bandwidth saved for each unit of the system cache allocated to the processing unit. Power saving information for the corresponding processing unit indicates the amount of power saved by the processing unit itself for each unit of the system cache allocated to that processing unit.

[0038] In some embodiments, bandwidth saving information includes one or more functions indicating the amount of DRAM bandwidth saved for different cells of the system cache allocated to a processing unit. For example, bandwidth saving information may include functions that map different cells of the system cache allocated to a processing unit to different amounts of DRAM bandwidth used by the processing unit. As another example, bandwidth saving information may include functions that map different amounts of system cache allocation, increasing or decreasing the DRAM bandwidth saved by the processing unit or the additional DRAM bandwidth used by the processing unit, respectively. For a given cache size allocated to a processing unit, these functions can be used to calculate the corresponding amount of DRAM bandwidth used. Therefore, the function can also be used to calculate differences in DRAM bandwidth used, for example, the amount of bandwidth saved or the additional bandwidth required for two different system cache sizes.

[0039] In some embodiments, power saving information includes one or more functions that indicate the amount of power saved for different cells of the system cache allocated to a processing unit. For example, the power saving information may include a function that maps different cells of the system cache allocated to the processing unit to different amounts of power (e.g., milliwatts or watts) consumed by the processing unit for a given processing unit performance level. For a given cache size allocated to a processing unit, this function can be used to calculate the corresponding amount of power consumed to maintain the current performance level of the processing unit. As another example, the power saving information may include multiple functions, each mapping different amounts of power consumption to different performance levels for a given system cache size allocated to a processing unit. For a given performance level, these functions can be used to calculate the power consumption difference for two different system cache sizes allocated to the processing unit. Therefore, these functions can also be used to calculate the power difference used for two different system cache sizes, such as power savings or additional power required.

[0040] Figure 4A The diagram illustrates various embodiments of the relationship between the two. Figure 2 Example bandwidth saving graph 400 associated with the processing unit. Bandwidth saving graph 400 is a graph corresponding to the bandwidth saving function of the processing unit. Figure 4A As shown, line 402 represents the amount of corresponding DRAM bandwidth saved by the processing unit for different numbers of system caches allocated to the processing unit.

[0041] Figure 4B The diagram illustrates various embodiments of the relationship between the two. Figure 2 An exemplary power-saving graph 410 associated with the processing unit. Power-saving graph 400 is a graph corresponding to the power-saving function of the processing unit. For example... Figure 4B As shown, line 412 represents the amount of power saved by the processing unit for different units of the system cache allocated to the processing unit. In some embodiments, the power saving graph 400 corresponds to a specific performance level of the processing unit. The metrics used to determine the performance level may vary depending on the type of processing unit. For example, the performance level of GPU 112(2) may be frames per second (FPS).

[0042] like Figure 4A and 4B As shown, the bandwidth saving function and power saving function are monotonically non-decreasing functions. As the size of the system cache allocated to the processing unit increases, the amount of bandwidth saved and the amount of power saved also increase. However, for each additional system cache unit, the amount of increase by the amount of bandwidth saved and / or power saved can vary. Figure 4A and 4BAs shown, the slopes of lines 402 and 412 vary at different points on the bandwidth saving graph 400 and the power saving graph 410, respectively.

[0043] In some embodiments, the resource saving information 302 calculated and / or stored by the processing unit is based on the unit type of the processing unit. For example, the amount of total power saved by a processing unit of a first unit type may be associated with the amount of DRAM bandwidth saved by allocating additional units to the processing unit via system cache; the amount of total power saved by a processing unit of a second unit type may be associated with the amount of processing time and / or processing power saved by allocating additional units to the processing unit via system cache; and the amount of total power saved by a processing unit of a third unit type may be associated with the amount of DRAM bandwidth saved by allocating additional units to the processing unit via system cache, as well as the amount of processing time and / or processing power saved. A processing unit of the first unit type may generate only bandwidth saving information, a processing unit of the second unit type may generate only power saving information, and a processing unit of the third unit type may generate both bandwidth saving information and power saving information simultaneously.

[0044] For example, refer to Figure 2 Each of CPU 102, CPU 112(1), GPU 112(2), and DLA 112(3) may be a processing unit type, wherein the amount of total power saved is associated with the amount of DRAM bandwidth saved by allocating additional units of the system cache to the processing unit, as well as the amount of processing time and / or power. ENC 112(4) and DEC112(5) are processing unit types, wherein the amount of total power saved is associated with the amount of DRAM bandwidth saved by allocating additional units of the system cache to the processing unit. The computational and / or storage resource saving information 302 for each of CPU 102, CPU 112(1), GPU 112(2), and DLA 112(3) may include bandwidth saving information and power saving information, while the computational and / or storage resource saving information 302 for each of ENC 112(4) and DEC 112(5) may include only bandwidth saving information.

[0045] like Figure 3 As shown, each processing unit driver 103 transmits resource-saving information 302 to the cache management application 125. The cache management application 125 receives the resource-saving information 302 from each processing unit and generates a cache partitioning policy 304. The cache partitioning policy 304 indicates the number of units of the system cache 202 to be allocated to each processing unit (e.g., CPU 102 and one or more processing units 112).

[0046] In some embodiments, to generate cache partitioning policy 304, cache management application 125 determines for each processing unit the amount of total power saved by allocating different units of the system cache to the processing unit. In some embodiments, cache management application 125 generates a function that maps different units of the allocated system cache to different amounts of power saved. In some embodiments, the amount of bandwidth saved increases with each additional unit of the system cache allocated to the processing unit. Similarly, the amount of power saved increases with each additional unit of the system cache allocated to the processing unit. In this case, the function corresponding to the total power saved for each unit of the system cache allocated to the processing unit is a monotonically non-decreasing function. That is, as the number of units of the system cache allocated to the processing unit increases, the amount of total power saved also increases.

[0047] In some embodiments, the amount of total power saved by a processing unit when allocating additional units to the system cache is based on the unit type of the processing unit. As described above with respect to resource saving information 302, the amount of total power saved by a processing unit may be associated with the amount of DRAM bandwidth saved by allocating additional units to the processing unit, the amount of processing time saved, and / or with the processing capacity saved by allocating additional units to the processing unit, or both, depending on the unit type.

[0048] In some embodiments, the cache management application 125 determines the total power saved for the processing unit based on the unit type of the processing unit. For example, if the processing unit is a first unit type, the cache management application 125 may determine the total power saved for the processing unit based on the bandwidth saving information of the processing unit. If the processing unit is a second unit type, the cache management application 125 may determine the total power saved for the processing unit based on the power saving information of the processing unit.

[0049] In some embodiments, the amount of total power saved by the processing unit when allocating additional units to the system cache corresponds to the amount of power used to read data from DRAM 210 instead of from system cache 202. The processing unit driver 103 generates and / or stores bandwidth saving information and sends it to the cache management application 125. The cache management application 125 calculates the amount of total power saved for the processing unit based on the bandwidth saving information received from the processing unit driver 103. In some embodiments, the bandwidth saving information includes a bandwidth saving function that associates the saved bandwidth with different amounts of system cache allocated to the processing unit. The cache management application 125 generates a function corresponding to the amount of total power saved for the processing unit by multiplying the bandwidth saving function by the power saved per unit of DRAM bandwidth. In some embodiments, the power saved per unit of DRAM bandwidth is based on the operating frequency and voltage of the system cache 202 and DRAM 210. In some embodiments, the cache management application 125 stores a mapping that indicates the amount of power saved and / or consumed per unit of DRAM bandwidth. This mapping can be associated with a specific DRAM 210 included in system 100 and derived when designing a specific DRAM 210 and / or by simulating the operation of DRAM 210. The cache management application 125 determines, for a given operating frequency and voltage, the amount of power consumed by each of the system cache 202 and DRAM 210 at different bandwidths.

[0050] In some embodiments, the amount of total power saved by the processing unit when allocating additional units to the system cache corresponds to the amount of power used by the processing unit, such as for performing calculations or waiting to access data in system cache 202 or DRAM 210. The processing unit driver 103 generates and / or stores power-saving information and sends it to the cache management application 125. The cache management application 125 calculates the amount of total power saved for the processing unit based on the power-saving information received from the processing unit driver 103. In some embodiments, the power-saving information includes a power-saving function that associates the power saved for the processing unit with different amounts of system cache allocated to the processing unit. The cache management application 125 uses the power-saving function as a function corresponding to the amount of total power.

[0051] In some embodiments, the amount of total power saved by the processing unit when allocating additional units to the system cache corresponds to the amount of power used to read data from DRAM 210 and the amount of power used by the processing unit, for example, for performing calculations. The processing unit driver 103 generates and / or stores bandwidth saving information and power saving information, and sends the bandwidth saving information 302 and power saving information to the cache management application 125. The cache management application 125 calculates the amount of total power saved for the processing unit based on both the bandwidth saving information and the power saving information received from the processing unit driver 103. In some embodiments, the bandwidth saving information includes a bandwidth saving function that associates the saved bandwidth with different amounts of system cache allocated to the processing unit, and the power saving information includes a power saving function that associates the power saved by the processing unit with different amounts of system cache allocated to the processing unit. As described above, the cache management application 125 generates a function corresponding to the amount of total power saved for the processing unit by multiplying the bandwidth saving function by the bandwidth per unit of DRAM bandwidth and adding the power saving function to the result.

[0052] After determining the power saved for CPU 102 and one or more processing units 112, cache management application 125 generates cache partitioning strategy 304 based on the total power saved for each processing unit. Cache management application 125 determines the number of units of the system cache allocated to each processing unit that maximizes the amount of total power saved across CPU 102 and one or more processing units 112. In some embodiments, cache management application 125 determines the amount of total power saved for each processing unit by generating a function corresponding to the amount of total power saved for different system cache sizes available for the processing unit. Cache management application 125 formulates an optimization function based on the function corresponding to the amount of total power saved for different processing units. Cache management application 125 solves the optimization function under the constraint that the sum of the system cache allocated to each processing unit equals the total system cache size and the system-size cache allocated to each processing unit is greater than or equal to 0. Example equations representing the optimization problem and constraints are given by equations (1a) and (1b):

[0053] Maximum size O∑ s∈S U s (x s (1a) Subject to ∑ s∈S x s =Cand

[0054] In equations (1a) and (1b), S represents a set of processing units, O represents the total power saved by the system in this set of processing units, and Us (x s ) represents the allocation of the system cache size x s The total power saved by unit s is given by C, where C represents the total system cache size available on the SOC200. As shown in equations (1a) and (1b), the total power saved by the system on a set of processing units should be maximized, subject to the constraint that the sum of the system cache sizes across the set of processing units is equal to the total system cache size available on the SOC200 and the system cache size allocated to each processing unit s is greater than or equal to 0.

[0055] In some embodiments, the optimization problem described by equations (1a) and (1b) is solved by formulating and solving the optimization function using the Kuhn-Tucker optimization method. Equation (2) gives an example optimization function:

[0056] L=∑ s∈S U s (x s )+λ(∑ s∈S x s -C)-∑ s∈S μ s x s (2)

[0057] In equation (2), L represents the Lagrange function, U s (x s ) represents the allocation of the system cache size x s The total power saved by unit s, where C represents the total system cache size available on the SOC 200. Furthermore, in equation (2), λ and μ... s This represents the Kuhn-Tucker multiplier of the Lagrange function L. The cache management application 125 solves the optimization function to find each x. s The value of . The system of equations used to solve the optimization function shown in equation (2) is given by equations (3A) and (3B):

[0058]

[0059] -μ s x s =0(3B)

[0060] The system of equations shown in equations (3A) and (3B) is solved for each processing unit s, for example, for each of CPU 102 and one or more processing units 112, to determine the system cache partition size for each processing unit that maximizes the total power savings across all processing units s∈S. The cache management application 125 can solve the system of equations using any technically feasible simultaneous equation solving algorithm. According to the Kuhn-Tucker method, if the inequality constraints are invalid (i.e., the optimal solution has a system cache size greater than 0 allocated to each processing unit), then μ s The value is equal to 0. For a single λ, solving the system of equations for all processing units means determining the equal gradient point on the total power saving function for each processing unit. Because the total power saving function is monotonically non-decreasing, allocating more system cache units to a given processing unit will save less or equal power than the amount added to other processing units.

[0061] After generating the cache partitioning policy 304, the cache management application 125 divides the system cache 202 into multiple partitions according to the cache partitioning policy 304. Each partition among the multiple partitions is assigned to a processing unit indicated by the cache partitioning policy 304. In some embodiments, the cache management application 125 partitions the system cache 202 and assigns each partition to a different processing unit. In some embodiments, the cache management application 125 sends the cache partitioning policy 304 to another application that partitions the system cache 202 and / or assigns partitions to different processing units.

[0062] In some embodiments, the cache management application 125 is configured to generate an updated cache partitioning policy 304 when the characteristics of the CPU 102, processing unit 112, SOC 200, and / or DRAM 210 change. For example, when the cache management application 125 determines that at least one of the following has changed: the frequency of DRAM 210, the frequency of the memory where the system cache 202 resides, the voltage of DRAM 210, the voltage of the memory where the system cache 202 resides, the system temperature, or the workload allocated to one of the CPU 102 or processing unit 112, the cache management application 125 may request updated resource-saving information 302 from the driver 103. In some embodiments, the processing unit driver 103 determines that the workload allocated to the corresponding processing unit has changed and transmits the resource-saving information 302 to the cache management application 125. In response to receiving the resource-saving information 302 from one or more processing unit drivers 103, the cache management application 125 generates an updated cache partitioning policy 304. In some embodiments, the cache management application 125 generates and solves an updated optimization function.

[0063] Typical methods for partitioning cache memory include dividing the cache evenly among various applications sharing the cache or based on one or more predetermined criteria. One problem with these methods is that, for a system cache such as system cache 202, if the system cache is not effectively partitioned, different processing units still cannot efficiently utilize the system cache. For example, if a partition is small compared to the amount of cache data generated by a processing unit, that processing unit will frequently fill the partition and have to remove and replace cache data. As another example, if a partition is large compared to the amount of cache data generated by a processing unit, portions of the system cache will remain unused or infrequently used, even if other processing units would benefit from the extra space in the system cache. As mentioned above, inefficient sharing of the system cache by different processing units also leads to increased power consumption. If the computer system is a power-constrained system, inefficient sharing of the system cache by different processing units results in reduced performance and / or processing speed for the different processing units. In contrast, using the disclosed techniques, the size of the partitions allocated to different processing units is chosen such that the overall power consumption of the system, such as battery power, is reduced. Therefore, using the disclosed technology, the computer system uses less battery power and can maintain a longer battery life compared to existing systems. Similarly, if the computer system is a power-constrained system, the processing speed and / or performance of the computer system will be increased.

[0064] Furthermore, by using mathematical optimization (e.g., as shown in Equations 1-3B) to determine the size of the partitions allocated to different processing units, the optimal partition size for each partition is identified to minimize the total power consumption of the system. That is, choosing other sizes for the partitions allocated to different processing units will result in a higher total power consumption of the system compared to the total power consumption of the system when the partition sizes are selected using the disclosed techniques.

[0065] Figure 5 This is a flowchart of method steps for automatically generating a system cache partitioning strategy 304 according to various embodiments. Although references... Figure 1-3 The system described herein describes the method steps, but those skilled in the art will understand that any system configured to implement the method steps in any order falls within the scope of this invention.

[0066] like Figure 5 As shown, method 500 begins at step 502, where cache management application 125 receives resource saving information 302 from one or more processing unit drivers (e.g., processing unit drivers 103(1)-(N)). In various embodiments, resource saving information 302 includes bandwidth saving information and / or power saving information. The bandwidth saving information for the processing unit indicates the amount of DRAM bandwidth saved for different amounts of system cache allocated to the processing unit. In some embodiments, the bandwidth saving information includes a function that maps different units of the system cache allocated to the processing unit to different amounts of DRAM bandwidth saved by the processing unit. The power saving information for the processing unit indicates the amount of power saved for different amounts of system cache allocated to the processing unit. In some embodiments, the power saving information includes a function that maps different units of the system cache allocated to the processing unit to different amounts of power bandwidth saved by the processing unit.

[0067] In some embodiments, the type of information received from the processing unit driver is based on the unit type of the corresponding processing unit. If the total power saved by allocating additional units to the system cache for the processing unit is related to the amount of DRAM bandwidth saved by allocating additional units to the system cache, then the cache management application 125 receives bandwidth saving information from the processor's driver. If the total power saved by allocating additional units to the system cache for the processing unit is related to the amount of power saved by allocating additional units to the system cache, then the cache management application 125 receives power saving information from the processing unit's driver. Therefore, when transmitting resource saving information 302, some processing unit drivers may transmit only bandwidth saving information, some processing unit drivers may transmit only power saving information, and some processing unit drivers may transmit both bandwidth saving information and power saving information simultaneously.

[0068] In step 504, the cache management application 125 determines the total power saved for each processing unit based on the resource saving information 302 (e.g., bandwidth saving information and / or power saving information) received from the corresponding processing unit driver. The determination of the total power saved for each processing unit is performed in a manner similar to that discussed above regarding the cache management application 125.

[0069] In some embodiments, determining the total power saved for each processing unit includes generating a total power saving function that maps different units of the system cache allocated to the processing unit to different amounts of total power saved by the processing unit. If the cache management application 125 receives bandwidth saving information from the processing unit, the cache management application 125 generates the total power saving function based on the bandwidth saving information. If the cache management application 125 receives power saving information from the processing unit, the cache management application 125 generates the total power saving function based on the power saving information. If the cache management application 125 receives both bandwidth saving information and power saving information from the processing unit, the cache management application 125 generates the total power saving function based on both the bandwidth saving information and the power saving information.

[0070] In step 506, the cache management application 125 generates an initial system cache partitioning strategy 304 based on the total power saved for each processing unit. The generation of the initial system cache partitioning strategy 304 is performed in a manner similar to that discussed above regarding the cache management application 125. In some embodiments, the cache management application 125 determines the amount of total power saved for each processing unit by generating a function corresponding to the amount of total power saved for different system cache sizes available for the processing unit. The cache management application 125 formulates an optimization function based on the function corresponding to the amount of total power saved for different processing units. The cache management application 125 obtains values ​​corresponding to the system cache sizes allocated to different processing units by solving the optimization function under the constraint that the sum of the system cache sizes allocated to each processing unit equals the total system cache size. A set of example equations for formulating and solving the optimization function are given in equations (1A)-(3B) above.

[0071] In step 508, the cache management application 125 divides the system cache into multiple partitions 202(1)-(N) based on the initial system cache partitioning policy, and each partition 202 is assigned to a corresponding processing unit. In some embodiments, the cache management application 125 partitions the system cache 202 and assigns each partition to a different processing unit. In some embodiments, the cache management application 125 sends the cache partitioning policy 304 to another application that partitions the system cache 202 and / or assigns the partitions to different processing units.

[0072] In step 510, the cache management application 125 determines whether the system 100 has changed. Determining whether the system 100 has changed includes, for example, determining whether at least one of the following has changed: the frequency of DRAM 210, the frequency of the memory where the system cache 202 is located, the voltage of DRAM 210, the voltage of the memory where the system cache 202 is located, the system temperature, or the workload assigned to one of the CPU 102 or the processing unit 112.

[0073] If the system has changed, in step 512, the cache management application 125 receives updated resource saving information 302 from one or more processing unit drivers. The method returns to step 504, where the cache management application 125 determines the updated total power saving per processing unit based on the updated resource saving information 302. The cache management application 125 determines an updated system cache partitioning strategy 304 based on the updated total power saving per processing unit.

[0074] In summary, a cache management application partitions the system cache across multiple different processing units of a computer system. The cache management application receives resource-saving information from each processing unit, indicating one or more benefits associated with each additional unit of the system cache allocated to that processing unit. In some embodiments, this information indicates the amount of bandwidth saved by each additional unit allocated to the cache by the processing unit. In some embodiments, this information indicates the amount of power saved by each additional unit allocated to the cache by the processing unit. The amount of bandwidth saved and / or power saved is determined based on a specific performance point of the processing unit, such as the target frames per second for a GPU. The cache management application generates a partitioning strategy for the system cache based on the information received from the different processing units. The partitioning strategy indicates the number of units of the system cache allocated to each processing unit. The cache management application determines the number of units of the system cache allocated to each processing unit that maximizes the amount of bandwidth and / or maximizes the amount of power saved for the computer system.

[0075] At least one technical advantage of the disclosed technique over the prior art lies in the more efficient sharing of access to the system cache among multiple processing units. Specifically, different portions of the system cache are allocated to each processing unit. Each processing unit has read access to other portions of the system cache but only writes data to its allocated portion. Therefore, each processing unit does not write data to or overwrite data stored in portions of the system cache allocated to other processing units. Consequently, cached data is removed or overwritten less frequently compared to previous methods, resulting in reduced power consumption. Furthermore, using the disclosed technique, the optimal size of the portion of the system cache allocated to each processing unit is determined, minimizing the overall power consumption of the system. These technical advantages provide one or more technological advancements superior to existing methods.

[0076] Any and all combinations of any claim element recited in any way in any claim and / or any element described in this application fall within the scope of this invention and protection.

[0077] Various embodiments have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.

[0078] Aspects of this embodiment may be embodied as a system, method, or computer program product. Therefore, aspects of this disclosure may take the form of a completely hardware embodiment, a completely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all collectively referred to herein as a “module,” “system,” or “computer.” Furthermore, any hardware and / or software technology, process, function, component, engine, module, or system described in this disclosure may be implemented as a circuit or a set of circuits. Additionally, aspects of this disclosure may take the form of a computer program product embodied in one or more computer-readable media containing computer-readable program code.

[0079] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium includes, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination thereof. More specific examples (not an exhaustive list) of computer-readable storage media may include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the context of this document, a computer-readable storage medium can be any tangible medium that may include or store a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0080] Aspects of this disclosure have been described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine. When executed by a processor of a computer or other programmable data processing apparatus, the instructions enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such processors can be, but are not limited to, general-purpose processors, special-purpose processors, application-specific processors, or field-programmable gate arrays.

[0081] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code comprising one or more executable instructions for implementing one or more specified logical functions. It should also be noted that in some alternative embodiments, the functions indicated in the blocks may not occur in the order shown in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action or a combination of dedicated hardware and computer instructions.

[0082] While the foregoing description is directed to embodiments of this disclosure, other and further embodiments of this disclosure may be designed without departing from its essential scope, as defined by the appended claims.

Claims

1. A computer-implemented method for managing a system cache shared by multiple processing units, the method comprising: For each of the plurality of processing units, resource saving information is received, wherein the resource saving information specifies the amount of corresponding resources saved when different units of the system cache are allocated to the processing unit; For each of the plurality of processing units, determine the amount of power saved by that processing unit when different units of the system cache are allocated to that processing unit according to the resource saving information; as well as For each of the plurality of processing units, the number of units to be allocated to the system cache of the processing unit is calculated based at least on the amount of corresponding power saved by the plurality of processing units.

2. The computer-implemented method of claim 1, wherein the resource saving information includes bandwidth saving information, the bandwidth saving information specifying for different units of the system cache the amount of corresponding dynamic random access memory (DRAM) bandwidth saved when different units of the system cache are allocated to the processing unit.

3. The computer-implemented method of claim 1, wherein the resource saving information includes power saving information, the power saving information specifying for different units of the system cache the amount of corresponding power saved by the processing unit when different units of the system cache are allocated to the processing unit.

4. The computer-implemented method as described in claim 1, further comprising: For each of the plurality of processing units, a total power saving function based on the resource saving information is generated, wherein the total power saving function maps different units of the system cache to the amount of corresponding system power saved when different units of the system cache are allocated to the processing unit; The number of units of the system cache to be allocated to each processing unit is further calculated based on the total power saving function of the processing unit.

5. The computer-implemented method of claim 4, wherein the resource saving information comprises one or more functions that map different units of the system cache to the amount of corresponding resources saved when different units of the system cache are allocated to the processing unit, and wherein the total power saving function is generated based on the one or more functions.

6. The computer-implemented method of claim 4, wherein calculating the number of units to be allocated to each processing unit's system cache comprises: An optimization function is generated based on multiple total power saving functions of the multiple processing units; as well as For each of the plurality of processing units, the optimization function is solved to obtain the number of units to be allocated to the system cache of the processing unit.

7. The computer-implemented method as described in claim 1, further comprising: Receive second resource saving information from at least one of the plurality of processing units; as well as For each of the plurality of processing units, a second number of units to be allocated to the system cache of the processing unit is calculated, based at least on the second resource saving information.

8. The computer-implemented method of claim 7, wherein receiving the second resource-saving information is in response to a change in the workload of the at least one processing unit.

9. The computer-implemented method of claim 1, further comprising dividing the system cache into a plurality of partitions, wherein each partition among the plurality of partitions corresponds to a different processing unit among the plurality of processing units, and the partition includes a number of units of the system cache to be allocated to the corresponding processing unit.

10. The computer-implemented method of claim 1, wherein the plurality of processing units are integrated on a system-on-a-chip.

11. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: For each of the multiple processing units, resource saving information is received, wherein the resource saving information specifies the amount of corresponding resources saved when different units of the system cache are allocated to the processing unit; For each of the plurality of processing units, determine the amount of power saved by that processing unit when different units of the system cache are allocated to that processing unit according to the resource saving information; as well as For each of the plurality of processing units, the number of units to be allocated to the system cache of the processing unit is calculated based at least on the amount of corresponding power saved by the plurality of processing units.

12. The non-transitory computer-readable medium of claim 11, further comprising the following steps: Generate a total power saving function based on the resource saving information for each of the plurality of processing units, wherein the total power saving function maps different units of the system cache to the amount of corresponding system power saved when different units of the system cache are allocated to the processing unit; The number of units of the system cache to be allocated to each processing unit is further calculated based on the total power saving function of the processing unit.

13. The non-transitory computer-readable medium of claim 12, wherein the resource saving information includes bandwidth saving information, the bandwidth saving information specifying for different units of the system cache the amount of corresponding dynamic random access memory (DRAM) bandwidth saved when different units of the system cache are allocated to the processing unit.

14. The non-transitory computer-readable medium of claim 13, wherein generating the total power saving function comprises: For different units of the system cache, the amount of corresponding power saved is determined based on the amount of corresponding DRAM bandwidth saved.

15. The non-transitory computer-readable medium of claim 12, wherein the resource-saving information includes power-saving information, the power-saving information being different performance levels of the processing unit and different units of the system cache, specifying the amount of corresponding power saved by the processing unit when different units of the system cache are allocated to the processing unit at the different performance levels.

16. The non-transitory computer-readable medium of claim 15, wherein generating the total power saving function comprises: Determine the current performance level of the processing unit; as well as For different units of the system cache, determine the amount of power saved by the processing unit at the current performance level of the processing unit.

17. The non-transitory computer-readable medium of claim 12, wherein calculating the number of units to be allocated to the system cache for each processing unit comprises: An optimization function is generated based on multiple total power saving functions of the multiple processing units; as well as For each of the plurality of processing units, the optimization function is solved to obtain the number of units to be allocated to the system cache of the processing unit.

18. The non-transitory computer-readable medium of claim 11, further comprising the following steps: Receive second resource saving information from at least one of the plurality of processing units; as well as For each of the plurality of processing units, a second number of units to be allocated to the system cache of the processing unit is calculated, based at least on the second resource saving information.

19. The non-transitory computer-readable medium of claim 11, further comprising the step of dividing the system cache into a plurality of partitions, wherein each partition among the plurality of partitions corresponds to a different processing unit among the plurality of processing units, and the partition includes a number of units of the system cache to be allocated to the corresponding processing unit.

20. A system comprising: Multiple processing units; The system cache is shared by the multiple processing units; One or more memories storing instructions that, when executed by at least one of the plurality of processing units, cause the processing unit to perform the following steps: For each of the plurality of processing units, resource saving information is received, wherein the resource saving information specifies the amount of corresponding resources saved when different units of the system cache are allocated to the processing unit; For each of the plurality of processing units, determine the amount of power saved by that processing unit when different units of the system cache are allocated to that processing unit according to the resource saving information; as well as For each of the plurality of processing units, the number of units to be allocated to the system cache of the processing unit is calculated based at least on the amount of corresponding power saved by the plurality of processing units.

Citation Information

Patent Citations

  • Application aware SOC memory cache partitioning

    US20210034527A1