Graphics Processing Unit with Selective Two-Level Binning

The selective two-level binning architecture in GPUs adapts rendering modes based on device conditions, addressing the trade-off between performance and battery life/temperature by optimizing power consumption and thermal management.

JP7765456B2Active Publication Date: 2025-11-06ADVANCED MICRO DEVICES INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023506484
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-25
Filing Date
2021-08-05
Publication Date
2025-11-06
Estimated Expiration
2041-08-05

AI Technical Summary

Technical Problem

In battery-powered devices, there is a trade-off between GPU performance and battery life/temperature, with existing binning modes not adequately addressing the need for adaptive performance based on device conditions.

Method used

A selective two-level binning architecture in GPUs that dynamically selects rendering modes based on performance criteria such as thermal and power characteristics, using command buffer patching to modify workloads for efficient power consumption and temperature management.

Benefits of technology

Enhances user experience by optimizing GPU performance while reducing power consumption and heat, thus extending battery life and maintaining device efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007765456000001
    Figure 0007765456000001
  • Figure 0007765456000002
    Figure 0007765456000002
  • Figure 0007765456000003
    Figure 0007765456000003
Patent Text Reader

Abstract

A system and method are provided for runtime selection of a rendering mode for executing a command buffer using a graphics processing unit (GPU) of a device based on performance data corresponding to the device. A user mode driver (UMD) or kernel mode driver (KMD) executing on a central processing unit (CPU) selects a binning mode based on whether performance data, including sensor data or performance counter data, indicates that an associated binning or override condition is met. The UMD or KMD patches pending command buffers to execute in the selected binning mode based on whether the binning mode is enabled or disabled.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Computer processing systems generally use graphics processing units (GPUs) to perform graphics operations such as texture mapping, rendering, and vertex transformations. Performance requirements or specifications for a GPU can vary depending on the type of associated electronic device. For example, GPUs used in mobile devices or other battery-powered devices have characteristics and requirements that can differ significantly from other non-battery-powered platforms. Performance, battery life, and heat are generally important criteria for battery-powered device platforms, with sustained performance and low idle power consumption and temperature being desirable. However, there is generally a trade-off between GPU performance and battery life / temperature in battery-powered devices.

[0002] The present disclosure may be better understood, and its numerous features and advantages made apparent to those skilled in the art by referencing the accompanying drawings. The use of the same reference numbers in different drawings indicates similar or identical items. [Brief explanation of the drawings]

[0003] [Figure 1] FIG. 1 is a block diagram of an example device that selectively sets primitive binning modes based on GPU performance data, according to some embodiments. [Figure 2] FIG. 1 is a flow diagram illustrating a method for selectively setting a binning mode based on GPU performance data, according to some embodiments. [Figure 3] FIG. 10 is a flow diagram illustrating a method for selectively patching pending command buffers before submitting the command buffers to a GPU based on a determined binning mode, according to some embodiments. [Figure 4]FIG. 10 is a flow diagram illustrating a method for selectively executing a command buffer workload in a two-level binning mode or a non-two-level binning mode after a command buffer is submitted to a GPU and before execution of the command buffer by the GPU, based on whether two-level binning mode is enabled, in accordance with some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0004] Using the techniques described herein, a GPU selects a primitive binning mode (sometimes referred to herein as a "binning mode") based on the performance characteristics of the GPU. The binning mode determines how an image frame is divided into regions and how primitives are assigned to bins corresponding to each region. By selecting a binning mode based on performance characteristics, the GPU adapts the binning process according to the operating conditions of the electronic device, thereby improving the user experience.

[0005] To illustrate, converting information about three-dimensional (3D) objects into a two-dimensional (2D) image frame that can be displayed is known as rendering, and in some cases requires the device performing the rendering to utilize significant processing power and memory resources. Pixels in the image frame are generated by rendering graphical objects to determine color values ​​for each pixel. Exemplary graphical objects include points, lines, polygons, and higher-order three-dimensional (3D) surfaces. Points, lines, and polygons represent rendering primitives that are the basis of most 3D rendering instructions. More complex structures, such as 3D objects, are formed from combinations or meshes of such primitives. To display a particular scene using conventional rendering techniques, a GPU renders primitives with potential contributing pixels associated with the scene individually for each primitive by determining the pixels that fall within the edge of each primitive and obtaining the primitive's attributes corresponding to each of those pixels.

[0006] In other cases, the GPU renders primitives using a binning process, in which the GPU divides an image frame into regions, identifies primitives that intersect with the predetermined region, and places the identified primitives into bins that correspond to the predetermined region. Thus, each region of the frame is associated with a corresponding bin, and the bin contains primitives or portions of primitives that intersect with the associated bin. The GPU renders the frame on a bin-by-bin basis by rendering pixels of primitives that intersect with the region of the frame that corresponds to the bin. This allows the GPU to render frames more efficiently, at least in some cases, by requiring fewer memory accesses, increasing cache usage, etc.

[0007] One example of a binning process is primitive batch binning (PBB), in which a GPU receives a sequence of primitives and advantageously segments the primitives into temporally related primitive batches. Sequential primitives are captured until a predetermined condition, such as a batch full condition, a state storage full condition, or a dependency on a previously rendered primitive is determined. When performing PBB, the screen space displaying the rendered primitives is divided into several blocks. Each block of screen space is associated with a respective bin. Each primitive in the received sequence of primitives in the batch intersects one or more bins. For each received primitive in the batch, an initial bin intercept is calculated, which is the top-leftmost bin on the screen that the primitive intersects. After the batch is closed, the first bin to be processed is identified. Primitives that intercept the identified bin are processed. For each identified primitive that intercepts a bin, a next bin intercept is identified, and pixels contained in the primitives enclosed by the identified bins are sent for detailed rasterization. The next bin intercept is the next upper-most left bin in raster order that the processed primitive intersects.

[0008] In some embodiments, the GPU implements different binning techniques, referred to herein as binning modes or primitive binning modes, where different binning modes correspond to different region sizes for each bin, different numbers of binning levels, etc. For example, in some embodiments, the GPU includes both a single-level binning mode and a two-level binning mode. In single-level binning mode, also referred to as primitive batch binning (PBB) mode, the GPU divides the image frame into a specified number of regions and renders each region as described above.

[0009] In the two-level binning mode, two types of binning are performed: coarse level binning and fine level binning. In some embodiments, coarse level binning uses large bins (e.g., a total of 32 bins to cover the entire display area), which reduces binning overhead. Visibility information for each coarse bin is generated during the rendering of the first coarse bin (i.e., coarse bin 0) and used to render the other coarse bins. After coarse level binning, fine level binning is performed sequentially for each coarse bin. In some embodiments, fine level binning involves dividing each coarse bin into smaller “fine” bins, such as by performing primitive batch binning (PBB) to further bin each coarse bin into a 64×64 array of fine bins during PBB-based fine level binning. Each fine bin is then rendered using rendering information, such as primitive visibility information, generated for the corresponding coarse bin. In some embodiments, two-level binning occurs at the top of the graphics processing pipeline (e.g., before vertex processing and rasterization), as opposed to a single-level PBB-only binning mode that occurs in the middle of the graphics processing pipeline (e.g., after vertex processing but before pixel shading).

[0010] In some cases, different binning modes are preferred under different device conditions. For example, under some conditions, a single-level or PBB binning mode (where only PBB is used without the combination of coarse and fine level binning described above) provides better performance than two-level binning, but at the expense of increased device power consumption and higher operating temperature. In contrast, in some cases, two-level binning supports reduced power consumption at the expense of some performance.

[0011] To adapt the binning mode according to device conditions, in some embodiments, the GPU employs a selective two-level binning architecture that supports runtime selection of rendering modes. For example, in some embodiments, a device implementing the selective two-level binning architecture implements runtime selection between a two-level binning mode and a default rendering mode, such as a PBB rendering mode in which only PBB is executed. The selection of the binning mode is based on any of several performance criteria, such as thermal characteristics, power characteristics (e.g., battery life), etc. For example, in some embodiments, a driver, such as a user-mode driver (UMD) or a kernel-mode driver (KMD), receives performance data, such as sensor data and performance counter data, and selects the binning mode based on the performance data.

[0012] The techniques described herein enable such pending command buffers to be modified via command buffer patching, such that one or more workloads in the command buffer are configured to execute according to the current two-level binning mode or a non-two-level binning mode, such as a PBB rendering mode, depending on whether the two-level binning mode is enabled. As used herein, command buffer patching refers to modifying data in a command buffer by a driver or other module executed by the CPU or GPU, and may be performed on either the CPU or the GPU.

[0013] 1 illustrates an example of a device 100 including parallel processors, specifically a GPU 102, according to some embodiments. The device 100 implements a two-level binning architecture that enables runtime selection of a rendering mode for rendering image data. In addition to the GPU 102, the device 100 includes a CPU 104, a memory controller 105, a system memory 106, sensors 108, and a battery 111. In some embodiments, the GPU 102, the CPU 104, the memory controller 105, and the sensors 108 are communicatively coupled to each other via a bus 126. The memory controller 105 manages memory access requests provided by the GPU 102, the CPU 104, and the sensors 108 to access the system memory 106.

[0014] During graphics processing operations, applications in system memory 106 generate commands to instruct GPU 102 to render image data at defined locations in system memory 106 for subsequent display in image frames on device 100's electronic display (not shown). Commands output by the applications are recorded in command buffer 114 by UMD 110 executing on CPU 104. A given command buffer 114 contains commands for one or more workloads, each workload configured to run in two-level binning mode, non-two-level binning mode, or capable of running in either mode. Once UMD 110 has completed recording the commands in command buffer 114, KMD 112 submits command buffer 114 to GPU 102, where it is loaded into ring buffer 120 of GPU 102. One or more command processors 122 of the GPU 102 retrieve commands corresponding to particular command buffers 114 from the ring buffer 120 and execute those commands, for example, by retrieving image data from the system memory 106 and instructing shaders, compute units, and other graphics processing circuitry (not shown) to render the retrieved image data. In the example of FIG. 1, the GPU 102 selects between a two-level binning mode and a single-level binning mode, such as a PBB mode (sometimes referred to herein as a “single-level PBB mode” or a “single-level PBB only binning mode”), when executing the workload of each command buffer 114. In some embodiments, the GPU 102 selects which binning mode to use to execute the workload of each command buffer 114 based on one or more status bits stored in the system memory 106 or based on one or more patch valid bits (described below) stored in the GPU memory 124.In some embodiments, the GPU 102 selects a binning mode based on the status bits or the patch valid bits when executing corresponding logic in the command buffer 114 that causes the GPU 102 to check the status bits or the patch valid bits to determine which binning mode to use for execution of the corresponding workload in the command buffer 114. In some embodiments, the CPU 104 selectively sets the values ​​of the status bits and the patch valid bits based on identified performance characteristics (sometimes referred to herein as “performance data”) of the device 100. In some embodiments, a single command buffer 114 is configured such that some workloads in the command buffer 114 are executed in one-level binning mode and other workloads in the same command buffer 114 are executed in two-level binning mode. For example, some workloads are executable only using one-level binning, and therefore the command buffer 114 is configured such that such workloads are executed in one-level binning mode even when two-level binning conditions are met and two-level binning mode is enabled. In some embodiments, workloads are recorded in the command buffer 114 by the UMD 110 so that they can be executed in one-level binning mode and two-level binning mode, and the mode in which these workloads are executed is then selected by the GPU when the workload is executed.

[0015] Generally, drivers within an operating system run in either user mode or kernel mode. UMDs, such as UMD 110, run in a non-privileged processor mode where other application code, including protected subsystem code, executes. UMDs cannot gain access to system data or hardware except by calling application programming interfaces (APIs) that invoke system services. KMDs, such as KMD 112, run as part of the operating system and support one or more protected subsystems. UMDs and KMDs have different structures, entry points, and system interfaces. KMDs can perform certain protected operations and access system structures that UMDs cannot access. In one example, primitives generated by an application are recorded by UMD 110 into one or more command buffers, and KMD 112 submits the one or more command buffers to GPU 102 for subsequent rendering of the primitives or other commands stored in one or more command buffers 114. The command processor 122 causes the image data to be rendered according to a particular rendering mode, such as a two-level binning mode or a non-two-level binning mode, such as a PBB rendering mode. In some embodiments, the command processor 122 selects which rendering mode to use to render the image data associated with a particular command buffer 114 by determining whether the two-level binning mode is enabled or disabled. In some embodiments, the command processor 122 determines whether the two-level binning mode is enabled or disabled by checking one or more status bits stored in the GPU memory 124 or the system memory 106.

[0016] In some embodiments, the CPU 104 enables or disables the two-level binning mode based on performance data including performance counter data received from performance counters 116 stored in the system memory 106, sensor data 118 stored in the system memory 106 by the sensors 108, or both. In some embodiments, the UMD 110 or KMD 112 of the CPU 104 receives the performance data and processes the performance data to determine whether to enable or disable the two-level binning mode.

[0017] In some embodiments, the sensor data 118 generated by the sensor 108 includes one or more temperature measurements, voltage measurements, current measurements, instantaneous power measurements, peak power measurements, or other applicable sensor data. In some embodiments, the sensor 108 includes one or more temperature sensors, current sensors, voltage sensors, or power sensors.

[0018] In some embodiments, performance counters 116 track activity in various modules of the device, such as battery 111, CPU 104, ring buffer 120, level 1 (L1) cache, level 2 (L2) cache, or shaders of GPU 102. In some embodiments, performance counter data includes a respective quantity of one or more of cache accesses, cache hit rate, cache miss rate, memory accesses, GPU 102 utilization, CPU 104 utilization, current supplied to GPU 102, current supplied to CPU 104, voltage at GPU 102, voltage at CPU 104, frequency of GPU 102, and / or frequency of CPU 104.

[0019] In some embodiments, the performance data includes one or more parameters derived from sensor data 118 or performance counter data generated by performance counter 116, such as the average temperature of device 100, the rate of change (RoC) of the average temperature of device 100, the peak instantaneous power consumption of device 100 over a predetermined period of time, the average power consumption of device 100 over a predetermined period of time, the RoC of the average power consumption of device 100, or the state of charge (SoC) of battery 111 (i.e., the remaining charge of battery 111 expressed as a percentage of the charge capacity of battery 111). As used herein, according to various embodiments, the “average temperature” of device 100 refers to the mean, median, or mode of instantaneous temperatures measured at various locations on the device (e.g., CPU 104, GPU 102, battery 111, or a combination thereof), the mean, median, or mode of temperatures measured at various locations on the device over a defined period of time, or the mean, median, or mode of an estimated temperature of device 100 derived from estimated power consumption based on performance counter data generated by performance counters 116 over a defined period of time. As used herein, the “average power consumption” of device 100 refers, according to various embodiments, to the mean, median, or mode of instantaneous power consumption measured at battery 111 over a defined period of time, or the mean, median, or mode of estimated instantaneous power consumption based on performance counter data generated by performance counters 116 over a defined period of time.

[0020] The UMD 110 or KMD 112 monitors the performance data to determine whether one or more predetermined conditions for enabling the two-level binning mode have occurred, sometimes referred to herein as “two-level binning conditions.” In some embodiments, enabling or disabling the two-level binning mode involves the UMD 110 or KMD 112 setting the value of one or more status bits in the system memory 106 or GPU 102 that indicate whether the two-level binning mode is enabled. In some embodiments, the two-level binning conditions include one or more of: an average temperature of the device above a predetermined temperature threshold; an RoC of the average temperature of the device above a predetermined RoC threshold; a local temperature at a defined location of the device above a predetermined temperature threshold; an RoC of the local temperature above a predetermined RoC threshold; a peak instantaneous power consumption of the device above a predetermined threshold; an average power consumption of the device above a predetermined threshold; an RoC of the average power consumption of the device above a predetermined threshold; a battery SoC below a predetermined SoC threshold; or a combination of these conditions. It should be understood that in some embodiments, if a two-level binning condition is met and two-level binning mode is enabled by the UMD 110 or KMD 112, and then it is subsequently determined based on changes in performance data that the two-level binning condition is no longer met, the device disables the two-level binning mode. However, in some embodiments, other detectable conditions, sometimes referred to herein as “override conditions,” override the detection of the two-level binning condition. For example, if the device 100 is determined by the UMD 110 or KMD 112 to meet the two-level binning condition but is determined to meet the override condition of being plugged in (e.g., the battery is determined to be in a “charging” state), the two-level binning mode is disabled. In some embodiments, alternative or additional override conditions are set, such as determining that the average power consumption of the device 100 has fallen below a threshold, or determining that the GPU 102 or CPU 104 is no longer thermally throttled (e.g., this may be determined based on the clock frequency of the GPU 102 or CPU 104 exceeding a threshold).

[0021] In some embodiments, the UMD 110 records a given workload in the command buffer 114 differently depending on whether two-level binning mode is enabled or disabled based on a corresponding status bit stored in the system memory 106 when recording the command buffer 114. For example, when enabling two-level binning mode, the UMD 110 records all subsequent command buffers 114 to be executable according to the two-level binning mode, at least until the two-level binning mode is disabled again. In some embodiments, when disabling two-level binning mode, the UMD 110 records all subsequent command buffers to be executable according to a single-level binning mode, such as a non-two-level or single-level PBB mode. In some embodiments, the UMD 110 individually determines the binning mode for each of multiple workloads in a given command buffer 114 based on whether two-level binning mode is enabled when each workload is recorded by the UMD 110, and, in some examples, based on whether the given workload can be executed in two-level binning mode.

[0022] For example, in some embodiments, UMD 110 is configured to record workloads in command buffer 114 in one-level binning mode by default, and is configured to modify one or more workloads in pending command buffer 114 to run in two-level binning mode before submission to GPU 102 if two-level binning conditions are met. In other embodiments, UMD 110 is configured to record workloads in command buffer 114 in two-level binning mode by default, and is configured to modify one or more workloads in pending command buffer 114 to run in one-level binning mode before submission to GPU 102 if two-level binning conditions are not met.

[0023] In some cases, the state of the two-level binning mode (i.e., enabled / disabled) changes after the UMD 110 records or begins recording one or more of the command buffers 114, referred to in such examples as “pending command buffers,” but before the pending command buffers are executed by the GPU 102. In some embodiments, such pending command buffers are modified via command buffer patching to execute according to a two-level binning mode or a non-two-level binning mode, depending on whether the two-level binning mode is enabled. As used herein, command buffer patching refers to modifying data in a command buffer by a driver or other module executed by the CPU 104 or the GPU 102, and may be performed in either the CPU 104 or the GPU 102.

[0024] In one example, pending command buffer 114 that was recorded when two-level binning mode was enabled is modified by CPU 104 or GPU 102 via command buffer patching to perform according to a non-two-level binning mode in response to a determination by CPU 104 or GPU 102 that two-level binning mode has been disabled since the start of recording of pending command buffer 114. As another example, pending command buffer 114 that was recorded when two-level binning mode was disabled is modified by CPU 104 or GPU 102 via command buffer patching to perform according to two-level binning mode in response to a determination by CPU 104 or GPU 102 that two-level binning mode has been enabled since the start of recording of pending command buffer 114.

[0025] In some embodiments where command buffer patching is performed on CPU 104, UMD 110 performs command buffer patching near the end of the command buffer recording process. In some embodiments, when command buffer patching is performed on CPU 104, UMD performs command buffer patching immediately before submitting command buffer 114 to GPU 102, except when a pending command buffer 114 is configured to be executed more than once simultaneously (a predetermined condition that becomes known when recording command buffer 114).

[0026] In some embodiments involving CPU-side command buffer patching, the UMD 110 stores metadata for each workload, where a workload is defined as a set of work or graphics drawing for a given set of render targets, depth stencil targets, or both. In some embodiments, the metadata stored for each workload includes one or more tokens and one or more offsets. Each offset defines a location in the command buffer 114 that needs to be modified when two-level binning mode is enabled. Each token defines how the code in the command buffer 114 at the location defined in the corresponding offset should be modified when two-level binning mode is enabled. In one example, the metadata tokens cause the UMD 110 to modify code in the command buffer 114 that describes the visibility of primitives. In some embodiments, command buffer patching is required only if two-level binning mode is enabled, and UMD 110 initially (i.e., by default) records each workload in command buffer 114 to run in a non-two-level binning mode in such embodiments, and then makes a determination as to whether to patch command buffer 114 to be executable in two-level binning mode at the end of the recording process or immediately prior to submitting command buffer 114. As noted above, in some embodiments, UMD 110 alternatively records each workload in command buffer 114 to run in two-level binning mode, at least for workloads that can run in two-level binning mode, and then determines whether to modify one or more workloads to instead run in a non-two-level binning mode based on whether two-level binning mode is enabled and, in some examples, whether predetermined override conditions are met.

[0027] In the GPU 102, command buffer patching is performed based on a value or group of values, referred to herein as a “patch effect value,” stored in the GPU memory 124. In some embodiments, each patch effect value is a single Boolean value corresponding to a respective pending command buffer 114. In some embodiments, the KMD 112 determines whether two-level binning mode is enabled based on a corresponding status bit stored in the system memory 106 or based on an analysis of performance data, and then the KMD 112 causes the patch effect value to be set according to whether two-level binning mode is enabled prior to execution of the command buffer 114 by the GPU 102. In some embodiments in which GPU-side command buffer patching is performed, the UMD 110 must record the command buffer 114 to be capable of execution in both two-level binning mode and non-two-level binning mode, and the command processor 122 determines in which mode to execute the command buffer 114 based on the corresponding patch effect value.

[0028] In some embodiments, the patch valid values ​​may instead be command buffer-based patch valid values ​​stored in the command buffer 114 by the UMD 110 during recording, such that the GPU 102 checks one or more patch valid values ​​for each workload when executing a given command buffer 114. In such embodiments, the command buffer 114 self-modifies one or more workloads in the command buffer 114 to be executable in two-level binning mode or non-two-level binning mode based on the command buffer-based patch valid values ​​during execution of the command buffer 114 by the GPU 102. For example, in some embodiments, the command processor of the GPU 102 or a shader core of the GPU 102 modifies the command buffer 114 if GPU-side command buffer patching is required, which is determined based on a patch valid bit stored in the GPU memory 124 or a status bit stored in the system memory 106, as described above.

[0029] 2 illustrates an exemplary process flow for a method 200 for selectively executing a command buffer in a first binning mode or a second binning mode based on performance data obtained from performance counters or sensors, according to some embodiments. In some embodiments, the first binning mode is a two-level binning mode and the second binning mode is a single-level binning mode, such as a PBB mode. Method 200 is described with reference to an exemplary implementation in device 100 of FIG. 1 and its constituent components and modules.

[0030] At block 202, the UMD 110 or the KMD 112 receives performance data. In some embodiments, the performance data includes sensor data 118 generated by the sensors 108. In some embodiments, the data includes performance counter data generated by the performance counters 116. In some embodiments, the performance data includes both performance counter data and sensor data 118. In some embodiments, the sensor data 118 generated by the sensors 108 includes one or more temperature measurements, voltage measurements, current measurements, instantaneous power measurements, peak power measurements, or other applicable sensor data. In some embodiments, the performance counter data includes one or more respective quantities of cache accesses, cache hit rates, cache miss rates, memory accesses, GPU 102 utilization, CPU 104 utilization, current supplied to GPU 102, current supplied to CPU 104, voltage at GPU 102, voltage at CPU 104, frequency of GPU 102, and / or frequency of CPU 104, each corresponding to activity occurring in one or more modules of device 100, such as battery 111, CPU 104, ring buffer 120, level 1 (L1) cache, level 2 (L2) cache, or shaders of GPU 102. In some embodiments, the performance data includes one or more parameters derived from sensor data 118 or performance counter data generated by performance counters 116, such as the average temperature of the device, the rate of change (RoC) of the average temperature of the device, the peak instantaneous power consumption of the device during a predetermined period of time, the average power consumption of the device over a predetermined period of time, the RoC of the average power consumption of the device, or the state of charge (SoC) of the battery (i.e., the remaining charge of the battery, in some embodiments expressed as a percentage of the battery's charge capacity), etc. In some embodiments, the derived parameters are calculated by UMD 110 or KMD 112.

[0031] At block 204, the UMD 110 or KMD 112 determines whether the binning conditions are met based on the performance data. For example, in some embodiments, the binning conditions include one or more two-level binning conditions including one or more of: an average temperature of the device above a predetermined temperature threshold; an RoC of the average temperature of the device above a predetermined RoC threshold; a local temperature at a defined location of the device above a predetermined temperature threshold; an RoC of such local temperature above a predetermined RoC threshold; a peak instantaneous power consumption of the device above a predetermined threshold; an average power consumption of the device above a predetermined threshold; an RoC of the average power consumption of the device above a predetermined threshold; a battery SoC below a predetermined SoC threshold; or a combination of these conditions. If the UMD 110 or KMD 112 determines that the binning conditions are met, method 200 proceeds to block 206. Otherwise, if the UMD 110 or KMD 112 determines that the binning conditions are not met, method 200 proceeds to block 214.

[0032] At block 206, the UMD 110 or KMD 112 determines whether an override condition is met based on the performance data. For example, the override condition may include one or more of: the device 100 entering a charging condition in which the battery 111 is charging; the average temperature of the device 100 falling below a predetermined threshold; the RoC of the average temperature of the device 100 falling below a predetermined threshold; or a combination thereof. If the UMD 110 or KMD 112 determines that the override condition is not met, the method 200 proceeds to block 208. Otherwise, if the UMD 110 or KMD 112 determines that the override condition is met, the method 200 proceeds to block 214.

[0033] In block 208, the UMD 110 or KMD 112 enables a first binning mode for the newly created command buffer. In some embodiments, the first binning mode is a two-level binning mode. For example, to enable the first binning mode, the UMD 110 or KMD 112 sets a status bit value in the system memory 106 to indicate that the first binning mode is enabled. In some embodiments, when recording a subsequent command buffer, the UMD 110 checks the status bit value and determines that the command buffer should be recorded to run in the first binning mode.

[0034] At block 210, CPU 104 or GPU 102 patches pending command buffer workloads to enable those workloads to execute in the first binning mode. In some embodiments, UMD 110 patches a given pending command buffer workload on CPU 104 to execute in the first binning mode upon completion of the recording process for the pending command buffer workload. In some embodiments, UMD 110 patches a given pending command buffer workload on CPU 104 to execute in the first binning mode after completion of recording and before (e.g., immediately before) submitting the command buffer to GPU 102. In some embodiments, GPU 102 patches a given pending command buffer workload to execute in the first binning mode before (e.g., immediately before) execution of the pending command buffer based on one or more patch enable values ​​stored in GPU memory 124.

[0035] In block 212, the command buffer is executed on the GPU 102 in a first binning mode.

[0036] In block 214, the UMD 110 or KMD 112 disables the first binning mode for the newly created command buffer. For example, to disable the first binning mode, the UMD 110 or KMD 112 sets a status bit value in the system memory 106 to indicate that the first binning mode is disabled. In some embodiments, the UMD 110 checks the status bit value when recording a workload for a subsequent command buffer and, if applicable, determines that the workload for the command buffer should be recorded to run in the second binning mode. In some embodiments, the second binning mode is PBB mode.

[0037] In block 216, the UMD 110 or KMD 112 disables the first binning mode for the pending command buffer. In embodiments in which the UMD 110 records the workload of the command buffer to run in the second binning mode by default, no further action is required to disable the first binning mode other than changing the status bit in block 214, and block 216 is skipped. In some embodiments in which GPU-side command buffer patching is performed, the KMD 112 disables the first binning mode for the pending command buffer by setting one or more patch enable values ​​in GPU memory 124 to indicate that the first binning mode is disabled.

[0038] In block 218, the command buffer is executed on the GPU 102 in the second binning mode.

[0039] 3 illustrates an exemplary process flow for a method 300 for selectively patching a command buffer in a CPU to be able to execute in a two-level binning mode or a non-two-level binning mode, according to some embodiments. Method 300 is described with respect to an exemplary implementation in device 100 of FIG. 1 and its components and modules. In some embodiments, method 300 is performed in conjunction with block 210 of FIG. 2.

[0040] At block 302, the UMD 110 collects metadata for each workload for a given command buffer 114 (i.e., "per-workload metadata") when recording the workload to the command buffer 114. In some embodiments, the metadata stored for each workload includes one or more tokens and one or more offsets. Each offset defines a location in the command buffer 114 that needs to be modified if two-level binning mode is enabled to execute the corresponding workload of the command buffer 114. Each token defines how the code in the command buffer 114 at the location defined at the corresponding offset should be modified if two-level binning mode is enabled. In one example, the metadata tokens cause the UMD 110 to modify code in the command buffer 114 that describes the visibility of primitives when two-level binning mode is enabled.

[0041] At block 304, the UMD 110 determines whether two-level binning mode is enabled at or near the end of the recording process for the command buffer 114. In some embodiments, the UMD 110 checks the values ​​of one or more status bits stored in the system memory 106 to determine whether two-level binning mode is enabled. If it is determined that two-level binning mode is enabled, the method proceeds to block 310. If it is determined that two-level binning mode is disabled, the method 300 proceeds to block 306.

[0042] At block 306, the UMD 110 determines whether two-level binning mode is enabled after recording the command buffer 114 and before (e.g., immediately before) submitting the command buffer 114 to the GPU 102. In some embodiments, the UMD 110 checks the values ​​of one or more status bits stored in the system memory 106 to determine whether two-level binning mode is enabled. If it is determined that two-level binning mode is enabled, the method proceeds to block 312. If it is determined that two-level binning mode is disabled, the method 300 proceeds to block 308.

[0043] In block 308 , the KMD 112 submits the command buffer 114 to the GPU 104 .

[0044] At block 310, while recording the command buffer 114 (e.g., near the end of the recording process), the UMD 110 patches the command buffer 114 based on per-workload metadata so that one or more workloads in the command buffer 114 are executable in two-level binning mode. Generally, the manner in which the command buffer 114 is patched by the UMD 110 depends on the hardware implementation of the device 100.

[0045] In one example, two-level binning essentially uses visibility information in a buffer (i.e., a “visibility information buffer”) as a basis for determining which primitives are visible in which bins. In this example, UMD 110 records a workload in command buffer 114 to be executed using one-level binning, and the UMD does not include commands for the GPU to bind such visibility information buffers; however, if the workload were recorded by UMD 110 to be executed using two-level binning, the workload would need to include such commands. Thus, UMD 110 generates metadata that includes a token and offset indicating a location within command buffer 114, and a command to bind a visibility information buffer would need to be included for the workload if executed in two-level binning mode. In this manner, if two-level binning mode is enabled before submitting command buffer 114 to GPU 102, UMD 110 or KMD 112 patches the workload in command buffer 114 to include a command to bind a visibility information buffer at the location indicated in the metadata.

[0046] In another example, GPU 102 generally needs to receive bin information indicating how many bins there are, the size of those bins, and / or the order in which the bins should be processed when executing in two-level binning mode. In this example, UMD 110 generates metadata for each workload recorded in command buffer 114, including binning information indicating the number of bins, the size of each bin, and the order in which the bins should be processed, which is required to execute that workload in two-level binning mode. In this manner, if two-level binning mode is enabled before submitting command buffer 114 to GPU 102, UMD 110 or KMD 112 patches the workload in command buffer 114 to include the binning information indicated in the metadata.

[0047] In block 312, after recording the command buffer 114 and before submitting the command buffer 114 to the GPU 102, the UMD applies a patch to the command buffer 114 to enable one or more workloads in the command buffer 114 to be executed in two-level binning mode based on metadata for each workload.

[0048] 4 illustrates an exemplary process flow for a method 400 of selectively patching a workload of a command buffer in a GPU in a two-level binning mode or a non-two-level binning mode, according to some embodiments. In some embodiments, the GPU performs command buffer patching of the command buffer to enable execution of the workload in either a two-level binning mode or a non-two-level binning mode based on corresponding metadata generated by the UMD. Method 400 is described with respect to an exemplary implementation in device 100 of FIG. 1 and its constituent components and modules. In some embodiments, method 400 is performed in conjunction with block 210 of FIG. 2.

[0049] In block 402, the UMD 110 records the command buffer 114 to include one or more workloads. In some embodiments, the UMD 110 records the workloads so that they can be executed in either two-level binning mode or non-two-level binning mode without patching. In some other embodiments, the UMD 110 records workloads that run in non-two-level binning mode by default, and generates metadata that enables the GPU 102 to change the workload to run in two-level binning mode if necessary (i.e., if two-level binning mode is enabled after the workload is recorded in the command buffer 114 and before it is executed by the GPU 102).

[0050] In one example, the UMD 110 records the command buffer 114 to include a conditional statement for one or more workloads in the command buffer 114, the conditional statement causing the GPU 102 to check a patch valid value stored in a register in the GPU memory 124 and execute the one or more workloads in a two-level binning mode or a non-two-level binning mode depending on the value of the patch valid value. In some embodiments, the patch valid value is a Boolean value stored in a single bit of a register in the GPU memory 124. In some embodiments, the patch valid value is set by the UMD 110 or the KMD 112.

[0051] In block 404, the KMD 112 submits the command buffer 114 to the GPU 104. In some embodiments, once submitted to the GPU 104, the command buffer 114 is added to the ring buffer 120.

[0052] At block 406, the GPU 102 determines whether the two-level binning mode is enabled. In some embodiments, the GPU 102 checks one or more patch valid values ​​stored in the GPU memory 124 to determine whether the two-level binning mode is enabled. In some embodiments, the KMD 112 determines whether the two-level binning mode is enabled based on corresponding performance data and sets the patch valid values ​​in the GPU memory 124 accordingly. If the GPU 102 determines that the two-level binning mode is enabled, the method 400 proceeds to block 408. Otherwise, if the GPU 102 determines that the two-level binning mode is not enabled, the method 400 proceeds to block 410.

[0053] At block 408, the GPU 102 executes one or more workloads in the command buffer 114 in two-level binning mode. In some embodiments, the GPU 102 utilizes metadata generated by the UMD 110 during recording of the command buffer 114, as described above, to patch one or more workloads in the command buffer 114 to execute in two-level binning mode in response to determining that the patch effect value indicates that those workloads should be executed in two-level binning mode. In some other embodiments, the UMD 110 records each workload that may be executed in two-level binning mode to be executable in either two-level binning mode or a non-two-level binning mode, and the GPU 102 is configured to execute those workloads in a selected one of the two-level binning mode or the non-two-level binning mode based on the patch effect value.

[0054] In block 410, the GPU 102 executes one or more workloads in a non-two-level binning mode in the command buffer 114. In some embodiments, the non-two-level binning mode is a PBB rendering mode.

[0055] As disclosed herein, in some embodiments, a method includes determining that a binning condition is met based on performance data; and, in response to determining that the binning condition is met, patching a pending command buffer to enable execution in a binning mode associated with the binning condition; and executing the pending command buffer in the binning mode using a graphics processing unit (GPU). In one aspect, the binning condition is a two-level binning condition and the binning mode is a two-level binning mode. In another aspect, the two-level binning condition includes one or more of an average temperature of a device including a GPU exceeding a first predetermined temperature threshold, a first rate of change of the average temperature of the device exceeding a second predetermined threshold, a local temperature at a defined location of the device exceeding a third predetermined threshold, a second rate of change of the local temperature exceeding a fourth predetermined threshold, an average power consumption of the device exceeding a fifth predetermined threshold, a peak instantaneous power consumption of the device exceeding a sixth predetermined threshold, or a battery state of charge below a seventh predetermined threshold.

[0056] In one aspect, the performance data includes one or more of an average temperature of a device including a GPU, a first rate of change of the average temperature of the device, an average power consumption of the device, a second rate of change of the average power consumption, a peak instantaneous power consumption of the device, or a state of charge of a battery of the device. In another aspect, the method includes collecting per-workload metadata for the pending command buffers while recording the pending command buffers, and patching the pending command buffers is performed based on the per-workload metadata. In yet another aspect, determining that a two-level binning condition has been met based on the performance data occurs while recording the pending command buffers, and patching the pending command buffers occurs while recording the pending command buffers. In yet another aspect, determining that a two-level binning condition has been met based on the performance data occurs after recording the pending command buffers and before submitting the pending command buffers to the GPU, and patching the pending command buffers occurs after recording the pending command buffers and before submitting the pending command buffers to the GPU.

[0057] In one aspect, patching the pending command buffer to be executable in a two-level binning mode includes patching, with a GPU, the pending command buffer to be executable in the two-level binning mode based on at least one patch valid value stored in a memory of the GPU. In another aspect, executing the pending command buffer in a two-level binning mode includes dividing an image frame associated with the pending command buffer into a plurality of coarse bins, and for each of the plurality of coarse bins, dividing the coarse bin into a plurality of fine bins, segmenting a plurality of primitives associated with the pending command buffer into temporally related primitive batches, and for each coarse bin, rendering primitives of the temporally related primitive batch based on how the primitives intercept the plurality of fine bins of the coarse bin.

[0058] In some embodiments, the device includes a central processing unit (CPU) configured to determine that a binning condition is met based on performance data indicative of a temperature or power consumption of the device, and to patch the pending command buffer so that it can execute in a first binning mode associated with the binning condition in response to determining that the binning condition is met; and a graphics processing unit (GPU) configured to execute the pending command buffer in a selected one of the first binning mode and a second binning mode. In one aspect, the binning condition is a two-level binning condition, the first binning mode is a two-level binning mode, and the second binning mode is a primitive batch binning mode. In another aspect, the CPU is configured to collect per-workload metadata for the pending command buffer during recording of the pending command buffer, and patching the pending command buffer is performed based on the per-workload metadata, the per-workload metadata including an offset that identifies a location of code in the command buffer that is changed via applying the patch in the two-level binning mode and a token that identifies how the code is modified in the two-level binning mode.

[0059] In one aspect, the CPU is configured to determine that a two-level binning condition has been met based on the performance data while recording the pending command buffer and to apply a patch to the pending command buffer while recording the pending command buffer. In another aspect, the CPU is configured to determine that a two-level binning condition has been met based on the performance data after recording the pending command buffer and before submitting the pending command buffer to the GPU and to apply a patch to the pending command buffer after recording the pending command buffer and before submitting the pending command buffer to the GPU. In yet another aspect, the GPU is configured to execute the pending command buffer in a two-level binning mode by dividing an image frame associated with the pending command buffer into a plurality of coarse bins, and for each of the plurality of coarse bins, dividing the coarse bin into a plurality of fine bins, segmenting a plurality of primitives associated with the pending command buffer into temporally related primitive batches, and for each coarse bin, rendering primitives of the temporally related primitive batch based on how the primitives intercept the plurality of fine bins of that coarse bin.

[0060] In some embodiments, the device includes a central processing unit (CPU) configured to determine that a binning condition is met based on performance data indicative of a temperature or power consumption of the device, and to set one or more patch enable values ​​to indicate that a binning mode associated with the binning condition is enabled in response to determining that the binning condition is met; and a graphics processing unit (GPU) configured to determine that the binning mode is enabled based on the patch enable values, patch pending command buffers to enable execution in the binning mode, and execute the pending command buffers in the binning mode. In one aspect, the binning condition is a two-level binning condition, and the binning mode is a two-level binning mode. In another aspect, the device includes a first sensor configured to generate temperature data indicative of a temperature of the device and a second sensor configured to generate power consumption data indicative of power consumption of the device, and the CPU is configured to calculate the performance data based on at least one of the temperature data or the power consumption data.

[0061] In one aspect, the device includes at least one performance counter configured to generate performance counter data indicative of activity occurring at the device, and the CPU is configured to calculate the performance data based on the performance counter data. In another aspect, the GPU is configured to execute the pending command buffer in a two-level binning mode by dividing an image frame associated with the pending command buffer into a plurality of coarse bins, and for each of the plurality of coarse bins, dividing the coarse bin into a plurality of fine bins, segmenting a plurality of primitives associated with the pending command buffer into temporally related primitive batches, and, for each coarse bin, rendering primitives of the temporally related primitive batch based on how the primitives intercept the plurality of fine bins of the coarse bin.

[0062] In some embodiments, the above-described apparatus and techniques are implemented in a system that includes one or more integrated circuit (IC) devices (also called integrated circuit packages or microchips), such as device 100 including GPU 102, CPU 104, and system memory 106 described above with reference to FIG. 1 . Electronic design automation (EDA) and computer-aided design (CAD) software tools can be used in the design and manufacture of these IC devices. These design tools are typically represented as one or more software programs. The one or more software programs include code executable by a computer system for operating the computer system to operate on code representing the circuits of one or more IC devices to perform at least a portion of a process for designing or adapting a manufacturing system for producing the circuits. This code may include instructions, data, or a combination of instructions and data. The software instructions representing the design or manufacturing tools are typically stored on a computer-readable storage medium accessible to the computing system. Similarly, code representing one or more stages of the design or manufacture of the IC devices is stored on and accessed from the same or a different computer-readable storage medium.

[0063] A computer-readable storage medium includes any non-transitory storage medium or combination of non-transitory storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media may include, but are not limited to, optical media (e.g., compact discs (CDs), digital versatile discs (DVDs), Blu-ray discs), magnetic media (e.g., floppy disks, magnetic tape, magnetic hard drives), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or micro-electromechanical systems (MEMS)-based storage media. The computer-readable storage medium (e.g., system RAM or ROM) may be internal to the computing system, the computer-readable storage medium (e.g., a magnetic hard drive) may be permanently attached to the computing system, the computer-readable storage medium (e.g., an optical disk or Universal Serial Bus (USB)-based flash memory) may be removably attached to the computing system, or the computer-readable storage medium (e.g., network-accessible storage (NAS)) may be coupled to the computer system via a wired or wireless network.

[0064] In some embodiments, certain aspects of the techniques described above are implemented by one or more processors of a processing system executing software. The software includes one or more sets of executable instructions stored or otherwise tangibly embodied in a non-transitory computer-readable storage medium. The software may include instructions and specific data that, when executed by one or more processors, operate the one or more processors to perform one or more aspects of the techniques described above. The non-transitory computer-readable storage medium may include, for example, a magnetic or optical disk storage device, a solid-state storage device such as flash memory, a cache, a random access memory (RAM), or other non-volatile memory device(s). The executable instructions stored on the non-transitory computer-readable storage medium may be implemented as source code, assembly language code, object code, or other form of instructions that can be interpreted or otherwise executed by one or more processors.

[0065] In addition to the above, it should be noted that not all activities or elements described in the summary description are required, that some of the particular activities or devices may not be required, that one or more additional activities may be performed, and that one or more additional elements may be included. Furthermore, the order in which the activities are listed is not necessarily the order in which they are performed. Also, the concepts have been described with reference to specific embodiments. However, those skilled in the art will recognize that various modifications and variations can be made without departing from the scope of the invention as set forth in the claims. Accordingly, the specification and drawings should be regarded in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of the present invention.

[0066] Benefits, other advantages, and solutions to problems have been described above with regard to specific embodiments. However, the benefits, advantages, solutions to problems, and features from which any benefit, advantage, or solution may arise or be manifested are not construed as critical, essential, or essential features of any or all claims. Moreover, the specific embodiments described above are illustrative only, since the disclosed invention may be modified and practiced in different, but similar manners apparent to those skilled in the art having the benefit of the teachings herein. No limitations are intended to the details of construction or design herein shown, other than as described in the appended claims. It is therefore apparent that the specific embodiments described above may be altered or modified, and that all such variations are considered within the scope of the disclosed invention. Accordingly, the protection sought herein is set forth in the appended claims.

Claims

1. 1. A method comprising: determining that a binning condition is met based on the performance data; and In response to determining that the binning condition is satisfied, patching the pending command buffer to be executable in a first binning mode associated with the binning condition, the first binning mode being one of a plurality of binning modes executable by a graphics processing unit (GPU), each of the plurality of binning modes including dividing an image frame associated with the pending command buffer into a plurality of bins; executing the pending command buffer in the first binning mode using the GPU. method.

2. The first binning mode is a two-level binning mode, the two-level binning mode including dividing an image frame associated with the pending command buffer into a plurality of coarse bins and dividing each of the plurality of coarse bins into fine bins.

10. The method of claim 1.

3. the binning conditions include one or more of: an average temperature of a device including the GPU exceeds a first predetermined temperature threshold; a first rate of change of the average temperature of the device exceeds a second predetermined threshold; a local temperature at a defined location of the device exceeds a third predetermined threshold; a second rate of change of the local temperature exceeds a fourth predetermined threshold; an average power consumption of the device exceeds a fifth predetermined threshold; a peak instantaneous power consumption of the device exceeds a sixth predetermined threshold; or a battery state of charge below a seventh predetermined threshold. The method of claim 2.

4. the performance data includes one or more of an average temperature of a device including the GPU, a first rate of change of the average temperature of the device, an average power consumption of the device, a second rate of change of the average power consumption, a peak instantaneous power consumption of the device, or a state of charge of a battery of the device; The method of claim 2.

5. and collecting per-workload metadata for the pending command buffers during recording of the pending command buffers, wherein patching the pending command buffers is performed based on the per-workload metadata. The method of claim 2.

6. determining that the binning condition is met based on the performance data occurs while recording the pending command buffer, and applying a patch to the pending command buffer occurs while recording the pending command buffer. The method of claim 5.

7. determining that the binning condition is met based on the performance data occurs after recording the pending command buffer but before the pending command buffer is submitted to the GPU, and patching the pending command buffer occurs after recording the pending command buffer but before the pending command buffer is submitted to the GPU. The method of claim 5.

8. patching the pending command buffer to be executable in the two-level binning mode, patching, using the GPU, the pending command buffer to be executable in the two-level binning mode based on at least one patch valid value stored in a memory of the GPU. The method of claim 2.

9. Executing the pending command buffer in the two-level binning mode includes: segmenting a plurality of primitives associated with the pending command buffer into temporally related batches of primitives; for each coarse bin, rendering a primitive of the temporally related batch of primitives based on how the primitive intercepts the plurality of fine bins of the coarse bin; The method of claim 2.

10. A device, a central processing unit (CPU); a graphics processing unit (GPU), The CPU determining that a binning condition is met based on performance data indicative of a temperature or power consumption of the device; responsive to determining that the binning condition is satisfied, patching the pending command buffer to be executable in a first binning mode associated with the binning condition, the first binning mode being one of a plurality of binning modes, each of the plurality of binning modes including dividing an image frame associated with the pending command buffer into a plurality of bins; and The GPU is configurable to execute each of the plurality of binning modes and configured to execute the pending command buffer in the first binning mode; device.

11. The first binning mode is a two-level binning mode, and the plurality of binning modes include a second binning mode including a primitive batch binning mode, the two-level binning mode including dividing image frames associated with the pending command buffer into a plurality of coarse bins and dividing each of the plurality of coarse bins into fine bins, and the primitive batch binning mode including dividing image frames associated with the pending command buffer into a single level of bins. The device of claim 10.

12. the CPU is configured to collect per-workload metadata for the pending command buffer during recording of the pending command buffer, and patching the pending command buffer is performed based on the per-workload metadata, the per-workload metadata including: an offset identifying a location of code in the command buffer that is modified via patching in the two-level binning mode; and a token identifying how the code is modified in the two-level binning mode. The device of claim 11.

13. the CPU is configured to determine that the binning condition is met based on the performance data while recording the pending command buffer, and to apply a patch to the pending command buffer while recording the pending command buffer. The device of claim 11.

14. the CPU is configured to: determine, after recording the pending command buffer and before submitting the pending command buffer to the GPU, that the binning condition is met based on the performance data; and apply a patch to the pending command buffer, after recording the pending command buffer and before submitting the pending command buffer to the GPU. The device of claim 11.

15. The GPU is segmenting a plurality of primitives associated with the pending command buffer into temporally related batches of primitives; for each coarse bin, rendering primitives of the temporally related batch of primitives based on how the primitives intercept the plurality of fine bins of the coarse bin; and executing the pending command buffer in the two-level binning mode by The device of claim 11.

Citation Information

Patent Citations

  • Hierarchical tile-based rasterization algorithm

    JP2008117384A

  • Switching between direct rendering and binning in graphics processing using an overdraw tracker.

    JP2015506017A

  • Optimized multi-pass rendering on tile-based architectures

    JP2017505476A

  • Flex rendering based on render targets in graphics processing

    JP2017516207A

  • Systems and methods for rendering multiple levels of detail

    JP2019510993A