Techniques for optimizing power and performance for XR workloads
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2023-05-09
- Publication Date
- 2026-05-07
AI Technical Summary
Current techniques fail to address performance degradation caused by inter-frame power collapse (IFPC) exit delays in GPUs when processing fixed periodic workloads.
A method where a graphics processing unit (GPU) receives an indication of a timer duration associated with exiting an IFPC state, processes predefined workloads upon timer trigger, enters the IFPC state upon workload completion, and exits the IFPC state upon timer expiration, thereby avoiding unnecessary hysteresis timeouts and reducing power consumption.
This approach eliminates delays in processing commands by advancing the GPU wake-up timeline, reduces power consumption by avoiding redundant hysteresis timeouts, and improves performance by ensuring the GPU is ready to process commands immediately upon receiving an IPCC interrupt.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. patent application Ser. No. 17 / 663,637, filed May 16, 2022, entitled "TECHNIQUE TO OPTIMIZE POWER AND PERFORMANCE OF XR WORKLOAD," which is expressly incorporated by reference in its entirety.
[0002] The present disclosure relates generally to processing systems, and more particularly to one or more techniques for graphics processing.
[0003] introduction Computing devices often perform graphics and / or display processing (e.g., utilizing a graphics processing unit (GPU), a central processing unit (CPU), a display processor, etc.) to render and display visual content. Such computing devices may include, for example, computer workstations, mobile phones such as smartphones, embedded systems, personal computers, tablet computers, and video game consoles. A GPU is configured to execute a graphics processing pipeline that includes one or more processing stages that work together to execute graphics processing commands and output frames. A central processing unit (CPU) may control the operation of the GPU by issuing one or more graphics processing commands to the GPU. Modern CPUs are typically capable of simultaneously executing multiple applications, each of which may need to utilize the GPU during execution. The display processor may be configured to convert digital information received from the CPU to analog values and may issue commands to a display panel to display the visual content. A device that provides content for visual presentation on a display may utilize a CPU, a GPU, and / or a display processor.
[0004] Current techniques may not address performance degradation associated with inter-frame power collapse (IFPC) exit delays in GPUs when the GPU processes fixed periodic workloads. Improved power collapse techniques are needed. Summary of the Invention
[0005] SUMMARY OF THE DISCLOSURE The following presents a simplified summary of one or more aspects in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects, nor is it intended to identify key or critical elements of all aspects or to delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that is presented later.
[0006] In one aspect of the disclosure, a method, a computer-readable medium, and an apparatus are provided. The apparatus may receive an indication of a duration of a timer associated with exiting an inter-frame power collapse (IFPC) state from an application. The apparatus may process one or more predefined workloads upon triggering the timer associated with exiting the IFPC state. The apparatus may enter the IFPC state upon finishing processing the one or more predefined workloads. The apparatus may exit the IFPC state upon detecting expiration of the timer.
[0007] To the accomplishment of the foregoing and related ends, the one or more aspects include the features hereinafter fully described and particularly pointed out in the claims. The following description and the annexed drawings set forth in detail certain illustrative features of the one or more aspects. These features are indicative, however, of only a few of the various ways in which the principles of the various aspects may be employed and the description is intended to include all such aspects and their equivalents. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram illustrating an example content generation system in accordance with one or more techniques of this disclosure. [Diagram 2] FIG. 1 illustrates an example GPU in accordance with one or more techniques of this disclosure. [Diagram 3] FIG. 1 is a block diagram illustrating an exemplary environment in which aspects of the present disclosure may be implemented. [Figure 4] FIG. 2 illustrates an example GPU state timeline associated with an IFPC, according to one or more aspects. [Diagram 5] FIG. 2 illustrates an example GPU state timeline associated with an IFPC, according to one or more aspects. [Figure 6] 1 is a call flow diagram illustrating example communications between an application, a first component, and a GPU in accordance with one or more techniques of this disclosure. [Figure 7] 1 is a flowchart of an exemplary method of graphics processing in accordance with one or more techniques of this disclosure. [Figure 8] 1 is a flowchart of an exemplary method of graphics processing in accordance with one or more techniques of this disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Various aspects of the system, device, computer program product, and method are more fully described below with reference to the accompanying drawings. However, the present disclosure may be embodied in many different forms and should not be construed as limited to any specific structure or function presented throughout the present disclosure. Rather, these aspects are provided so that the disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art. Based on the teachings herein, those skilled in the art should understand that the scope of the present disclosure encompasses any aspect of the system, device, computer program product, and method disclosed herein, whether implemented independently of or in combination with other aspects of the present disclosure. For example, an apparatus can be implemented or a method can be performed using any number of the aspects described herein. It is intended that the scope of the present disclosure encompasses such an apparatus or method that is practiced using other structure, functionality, or structure and functionality in addition to or other than the various aspects of the present disclosure described herein. Any aspect disclosed herein may be embodied by one or more elements of a claim.
[0010] Although various aspects are described herein, many variations and permutations of these aspects fall within the scope of the disclosure. Although some potential benefits and advantages of aspects of the disclosure are described, the scope of the disclosure is not limited to any particular benefit, use, or purpose. Rather, aspects of the disclosure are broadly applicable to different wireless technologies, system configurations, processing systems, networks, and transmission protocols, some of which are illustrated by way of example in the figures and the following description. The detailed description and drawings are merely illustrative of the disclosure, rather than limiting, the scope of the disclosure being defined by the appended claims and their equivalents.
[0011] Certain aspects are presented with reference to various apparatus and methods that are described in the following detailed description and illustrated in the accompanying drawings by various blocks, components, circuits, processes, algorithms, and the like (collectively referred to as "elements"). These elements may be implemented using electronic hardware, computer software, or any combination thereof. Whether such elements are implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system.
[0012] As an example, the elements, or any portion of the elements, or any combination of the elements, may be implemented as a "processing system" (sometimes referred to as a processing unit) that includes one or more processors. Examples of processors include microprocessors, microcontrollers, graphics processing units (GPUs), general purpose GPUs (GPGPUs), central processing units (CPUs), application processors, digital signal processors (DSPs), reduced instruction set computing (RISC) processors, systems-on-chip (SOCs), baseband processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gate logic, discrete hardware circuits, and other suitable hardware configured to perform various functions described throughout this disclosure. The one or more processors in the processing system may execute software. Software may be interpreted broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software components, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0013] The term application may refer to software. As described herein, one or more techniques may refer to an application (e.g., software) being configured to perform one or more functions. In such examples, the application may be stored in a memory, such as an on-chip memory of a processor, a system memory, or any other memory. Hardware described herein, such as a processor, may be configured to execute the application. For example, an application may be described as including code that, when executed by the hardware, causes the hardware to execute one or more techniques described herein. As an example, the hardware may access code from a memory and execute the code accessed from the memory to execute one or more techniques described herein. In some examples, components are identified in this disclosure. In such examples, the components may be hardware, software, or a combination thereof. The components may be separate components or subcomponents of a single component.
[0014] In one or more examples described herein, the functions described may be implemented in hardware, software, or any combination thereof. If implemented in software, the functions may be stored or encoded as one or more instructions or code on a computer-readable medium. Computer-readable media includes computer storage media. A storage medium may be any available medium that can be accessed by a computer. By way of example and not limitation, such computer-readable media may comprise random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), optical disk storage, magnetic disk storage, other magnetic storage devices, combinations of the above types of computer-readable media, or any other medium that can be used to store computer-executable code in the form of instructions or data structures that can be accessed by a computer.
[0015] As used herein, the term "content" may refer to "graphical content," "images," and the like, regardless of whether those terms are used as adjectives, nouns, or other parts of speech. In some examples, the term "graphical content" as used herein may refer to content produced by one or more processes of a graphics processing pipeline. In further examples, the term "graphical content" as used herein may refer to content produced by a processing unit configured to perform graphics processing. In further examples, the term "graphical content" as used herein may refer to content produced by a graphics processing unit.
[0016] When IFPC (e.g., power collapsing the GPU during command submission when the GPU is idle) is utilized in the GPU, IFPC exit latency may cause unnecessary performance degradation when the GPU operates as a fixed function block and processes fixed periodic workloads. Furthermore, a hysteresis timeout associated with IFPC may be redundant when the GPU processes such fixed periodic workloads. The redundant hysteresis timeout may be associated with unnecessary power consumption. According to one or more aspects, a hint regarding a timer value may be provided to a graphics management unit (GMU) firmware. As a result, unnecessary hysteresis timeouts for fixed periodic workloads may be avoided. Furthermore, a timeline associated with GPU wake-up may be advanced based on the timer such that a delay between receipt of an inter-processor communication controller (IPCC) interrupt and the time when the GPU is awake and ready to process commands may be eliminated. Elimination of the delay may result in performance benefits.
[0017] FIG. 1 is a block diagram illustrating an example content generation system 100 configured to implement one or more techniques of the present disclosure. The content generation system 100 includes a device 104. The device 104 may include one or more components or circuits for performing various functions described herein. In some examples, one or more components of the device 104 may be components of a SOC. The device 104 may include one or more components configured to perform one or more techniques of the present disclosure. In the illustrated example, the device 104 may include a processing unit 120, a content encoder / decoder 122, and a system memory 124. In some aspects, the device 104 may include several components (e.g., a communication interface 126, a transceiver 132, a receiver 128, a transmitter 130, a display processor 127, one or more displays 131). The display(s) 131 may refer to one or more displays 131. For example, display 131 may include a single display or multiple displays, which may include a first display and a second display. The first display may be a left-eye display and the second display may be a right-eye display. In some examples, the first display and the second display may receive different frames for presentation thereon. In other examples, the first display and the second display may receive the same frames for presentation thereon. In further examples, the results of the graphics processing may not be displayed on the device, e.g., the first display and the second display may not receive any frames for presentation thereon. Instead, the frames or the graphics processing results may be forwarded to another device. In some aspects, this may be referred to as split rendering.
[0018] The processing unit 120 may include an internal memory 121. The processing unit 120 may be configured to perform graphics processing using the graphics processing pipeline 107. The content encoder / decoder 122 may include an internal memory 123. In some examples, the device 104 may include a processor that may be configured to perform one or more display processing techniques on one or more frames generated by the processing unit 120 before the frames are displayed by the one or more displays 131. Although the processor in the example content generation system 100 is configured as a display processor 127, it should be understood that the display processor 127 is one example of a processor and that other types of processors, controllers, etc. may be used in place of the display processor 127. The display processor 127 may be configured to perform display processing. For example, the display processor 127 may be configured to perform one or more display processing techniques on one or more frames generated by the processing unit 120. The one or more displays 131 may be configured to display or otherwise present the frames processed by the display processor 127. In some examples, the one or more displays 131 may include one or more of a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, a projection display device, an augmented reality display device, a virtual reality display device, a head-mounted display, or any other type of display device.
[0019] Memory external to the processing unit 120 and the content encoder / decoder 122, such as a system memory 124, may be accessible to the processing unit 120 and the content encoder / decoder 122. For example, the processing unit 120 and the content encoder / decoder 122 may be configured to read from and / or write to an external memory, such as the system memory 124. The processing unit 120 may be communicatively coupled to the system memory 124 via a bus. In some examples, the processing unit 120 and the content encoder / decoder 122 may be communicatively coupled to an internal memory 121 via a bus or via a different connection.
[0020] The content encoder / decoder 122 may be configured to receive graphical content from any source, such as the system memory 124 and / or the communications interface 126. The system memory 124 may be configured to store the received encoded or decoded graphical content. The content encoder / decoder 122 may be configured to receive the encoded or decoded graphical content, for example, in the form of encoded pixel data, from the system memory 124 and / or the communications interface 126. The content encoder / decoder 122 may be configured to encode or decode any graphical content.
[0021] The internal memory 121 or the system memory 124 may include one or more volatile or non-volatile memories or storage devices. In some examples, the internal memory 121 or the system memory 124 may include static random access memory (RAM), dynamic random access memory (DRAM), erasable programmable ROM (EPROM), EEPROM, flash memory, magnetic or optical data medium, or any other type of memory. The internal memory 121 or the system memory 124 may be a non-transitory storage medium, according to some examples. The term "non-transitory" may indicate that the storage medium is not embodied in a carrier wave or a propagated signal. However, the term "non-transitory" should not be interpreted to mean that the internal memory 121 or the system memory 124 is non-movable or that its contents are static. As an example, the system memory 124 may be removed from the device 104 and moved to another device. As another example, the system memory 124 may not be removable from the device 104.
[0022] The processing unit 120 may be a CPU, GPU, GPGPU, or any other processing unit that may be configured to perform graphics processing. In some examples, the processing unit 120 may be integrated into the motherboard of the device 104. In further examples, the processing unit 120 may be on a graphics card installed in a port on the motherboard of the device 104, or may be otherwise integrated into a peripheral device configured to interoperate with the device 104. The processing unit 120 may include one or more processors, such as one or more microprocessors, GPUs, ASICs, FPGAs, arithmetic logic units (ALUs), DSPs, discrete logic, software, hardware, firmware, other equivalent integrated or discrete logic circuits, or any combination thereof. If the techniques are implemented in part in software, the processing unit 120 may store instructions for the software in a suitable non-transitory computer-readable storage medium, such as the internal memory 121, and may execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Any of the above, including hardware, software, combinations of hardware and software, etc., may be considered to be one or more processors.
[0023] The content encoder / decoder 122 may be any processing unit configured to perform content decoding. In some examples, the content encoder / decoder 122 may be integrated into the motherboard of the device 104. The content encoder / decoder 122 may include one or more processors, such as one or more microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), arithmetic logic units (ALUs), digital signal processors (DSPs), video processors, discrete logic, software, hardware, firmware, other equivalent integrated or discrete logic circuits, or any combination thereof. If the techniques are implemented partially in software, the content encoder / decoder 122 may store instructions for the software in a suitable non-transitory computer-readable storage medium, such as the internal memory 123, and may execute the instructions in the hardware using one or more processors to perform the techniques of this disclosure. Any of the above, including hardware, software, combinations of hardware and software, etc., may be considered to be one or more processors.
[0024] In some aspects, the content generation system 100 may include a communication interface 126. The communication interface 126 may include a receiver 128 and a transmitter 130. The receiver 128 may be configured to perform any receiving function described herein with respect to the device 104. Additionally, the receiver 128 may be configured to receive information from another device, such as eye or head position information, rendering commands, and / or location information. The transmitter 130 may be configured to perform any transmitting function described herein with respect to the device 104. For example, the transmitter 130 may be configured to transmit information to another device, which may include a request for content. The receiver 128 and the transmitter 130 may be combined into a transceiver 132. In such an example, the transceiver 132 may be configured to perform any receiving and / or transmitting functions described herein with respect to the device 104.
[0025] Referring again to FIG. 1, in certain aspects, processing unit 120 may include a power collapse scheduler 198 configured to receive an indication of a period of a timer associated with exiting an IFPC state from an application. Power collapse scheduler 198 may be configured to process one or more predefined workloads upon triggering a timer associated with exiting an IFPC state. Power collapse scheduler 198 may be configured to enter an IFPC state upon completion of processing of one or more predefined workloads. Power collapse scheduler 198 may be configured to exit an IFPC state upon detecting expiration of the timer. The following description may focus on graphics processing, although the concepts described herein may be applicable to other similar processing techniques.
[0026] A device, such as device 104, may refer to any device, apparatus, or system configured to perform one or more techniques described herein. For example, a device may be a server, a base station, a user equipment, a client device, a station, an access point, a computer, such as a personal computer, a desktop computer, a laptop computer, a tablet computer, a computer workstation, or a mainframe computer, an end product, an apparatus, a phone, a smartphone, a server, a video game platform or console, a handheld device, such as a portable video game device or a personal digital assistant (PDA), a wearable computing device, such as a smart watch, an augmented reality device, or a virtual reality device, a non-wearable device, a display or display device, a television, a television set-top box, an intermediate network device, a digital media player, a video streaming device, a content streaming device, an in-vehicle computer, any mobile device, any device configured to generate graphical content, or any device configured to perform one or more techniques described herein. Although the processes herein may be described as being performed by a particular component (e.g., a GPU), in alternative embodiments, they may be performed using other components (e.g., a CPU) consistent with the disclosed embodiments.
[0027] A GPU can process multiple types of data or data packets in a GPU pipeline. For example, in some aspects, a GPU can process two types of data or data packets, e.g., context register packets and draw call data. A context register packet can be global state information, e.g., information about global registers, shading programs, or a set of constant data, that can adjust how a graphics context is to be processed. For example, a context register packet can include information about a color format. In some aspects of a context register packet, there can be a bit that indicates which workload belongs to the context register. There can also be multiple functions or programming running simultaneously and / or in parallel. For example, a function or programming can represent a certain operation, e.g., a color mode or color format. Thus, a context register can define multiple states of a GPU.
[0028] The context state may be utilized to determine how individual processing units function, e.g., vertex fetcher (VFD), vertex shader (VS), shader processor, or geometry processor, and / or in what mode the processing units function. To do so, the GPU may use context registers and programming data. In some aspects, the GPU may generate workloads, e.g., vertex or pixel workloads, in a pipeline based on the context register definition of a mode or state. Some processing units, e.g., VFDs, may use these states to determine how some functions, e.g., vertices, are assembled. Because these modes or states may change, the GPU may need to modify the corresponding context. Furthermore, the workloads corresponding to the modes or states may follow the changing mode or state.
[0029] FIG 2 illustrates an example GPU 200 in accordance with one or more techniques of this disclosure. As illustrated in FIG 2, GPU 200 includes a command processor (CP) 210, a draw call packet 212, a VFD 220, a VS 222, a vertex cache (VPC) 224, a triangle setup engine (TSE) 226, a rasterizer (RAS) 228, a Z process engine (ZPE) 230, a pixel interpolator (PI) 232, a fragment shader (FS) 234, a render backend (RB) 236, an L2 cache (UCHE) 238, and a system memory 240. Although FIG 2 illustrates GPU 200 including processing units 220-238, GPU 200 may include several additional processing units. Additionally, processing units 220-238 are merely examples and any combination or order of processing units may be used by a GPU in accordance with this disclosure. GPU 200 also includes a command buffer 250, a context register packet 260, and a context state 261.
[0030] 2, the GPU can utilize a CP, e.g., CP 210, or a hardware accelerator, to parse the command buffer into context register packets, e.g., context register packet 260, and / or draw call data packets, e.g., draw call packet 212. CP 210 can then send the context register packet 260 or draw call packet 212 through separate paths to a processing unit or block in the GPU. Furthermore, command buffer 250 can alternate between different states of context registers and draw calls. For example, a command buffer can be constructed as follows: context register for context N, draw call(s) for context N, context register for context N+1, and draw call(s) for context N+1.
[0031] In an extended reality (XR) pipeline, a complete data path may include two SOCs associated with two devices. The companion device may generate visual content and may send the visual content to the XR device. The XR device may then perform operations such as late stage reprojection (LSR) for final display based on the user's latest head pose. In particular, LSR may be a feature that may ensure the responsiveness of the XR headset to the user's movements. LSR may help reduce the perceived input lag and improve the user experience. As part of LSR, previously rendered frames may be reprojected or warped to a prediction of what a successfully rendered frame would look like using newer motion information from the headset sensors. In particular, the GPU in the XR device may be used to generate a motion vector (MV) grid using one or more of the depth, rendering pose, or latest head pose details.
[0032] The XR pipeline may be used to handle head movements (e.g., translation and rotation) or to perform optical correction. In one or more examples below, references to XR may also include references to augmented reality (AR) or virtual reality (VR).
[0033] FIG. 3 is a block diagram 300 illustrating an example environment in which aspects of the disclosure may be implemented. In particular, an example XR pipeline is illustrated in FIG. 3. In some configurations, an XR application 302 can generate commands associated with MV grid generation using a graphics application programming interface (API) 304. A graphics driver 310 (e.g., a graphics kernel driver or kernel graphics support layer (KGSL)) can receive the commands and communicate with an enhanced visual analytics (EVA) driver 306 to exchange appropriate data and / or commands associated with the XR pipeline. Additionally, an EVA firmware 308 can provide depth buffer details to a GPU 312 (e.g., via a host firmware interface (HFI) queue 316) and can trigger an inter-processor communication controller (IPCC) interrupt (the IPCC can be a centralized block for managing inter-processor interrupts at the SoC level) at regular intervals in the GPU 312 via an IPCC 318 when an LSR workload is ready to be processed by the GPU 312.
[0034] In an LSR use case, GPU 312 may be reserved and may operate as a fixed functional block. Additionally, in the LSR context, a graphics management unit (GMU) 314 within GPU 312 may always be active and may monitor IPCC interrupts from EVA firmware 308 (in other words, GMU 314 and EVA firmware 308 may communicate using IPCC interrupts).
[0035] There may be performance goals or targets associated with the XR pipeline. For example, the motion-render-photon ("photon" may refer to the corresponding change on a display, such as a head-mounted display (HMD)) latency (i.e., the latency from the companion device to the XR device) may be approximately 50-55 ms. Furthermore, the motion-to-photon latency may be less than 9 ms. Therefore, it may be important to meet performance goals and simultaneously reduce power consumption.
[0036] The GMU 314 may constantly monitor the IPCC interrupt from the EVA firmware 308 so that the graphics driver 310 cannot disable the clocks / regulators of the GMU 314 to put the GMU 314 in a slumbering state. To take advantage of another potential power saving opportunity, the GMU 314 may power collapse the GPU 312 during command submission (workload submission) when the GPU 312 is idle. This may be referred to as IFPC. In particular, IFPC may be a power saving feature where the GPU may be switched off between frames. IFPC may be controlled by the GMU 314 firmware. Based on IFPC, the GMU 314 firmware may switch off the GPU even if it is idle for a short period of time.
[0037] 4 is a diagram 400 illustrating an example GPU state timeline associated with an IFPC, according to one or more aspects. When an IFPC is enabled, the GPU may be in one of five possible states at any given time: an active state (also referred to as the A state), a hysteresis timeout state (also referred to as the B state), an IFPC entry state (also referred to as the C state), an IFPC state (also referred to as the D state) (when there is no workload for the GPU, the GMU 314 may switch off the GPU's clocks and regulators, and the GPU may be completely off when in the IFPC state), and an IFPC exit state (also referred to as the E state) (when a new workload is submitted while the GPU is in the IFPC state, the GMU 314 may switch on the GPU's clocks and regulators, and the IFPC exit state may be a transition state corresponding to a transition from the IFPC state to the active state). In particular, when in the active (A) state, the GPU may process a command submission corresponding to the current sample. The hysteresis timeout (B) state may be a timeout period after the GPU becomes idle before entering the IFPC entry (C) state. The IFPC entry (C) state may correspond to the time it takes for the GMU to switch off the GPU's clocks and regulators. When in the IFPC (D) state, the GPU may be completely off. Additionally, the IFPC exit (E) state may correspond to the time it takes for the GMU to turn on the GPU's clocks and regulators. In other words, when IFPC is enabled, there may be a latency associated with entry into and exit from the IFPC (D) state.
[0038] In one example, as shown in FIG. 4, when IFPC is enabled, upon receiving an IPCC interrupt 402 from the EVA firmware, the GMU firmware can place the GPU in an IFPC done (E) state to wake up the GPU from the IFPC (D) state. The IFPC done (E) state can thus represent a delay between receiving the IFPC interrupt 402 and the time when the GPU is woken up and ready to process commands. Once the GPU is ready and in the active (A) state, the GPU can process the command associated with the current sample. Once the GPU has completed processing the command, the GPU can provide a command complete interrupt to the GMU. The GMU can then notify the EVA firmware that the MV grid for the current sample is ready by triggering a reverse IPCC interrupt in the EVA firmware.
[0039] The hysteresis timeout (B) state may begin at the same time that the GPU completes processing of a command. When the hysteresis timeout (B) state expires, the GMU may power collapse the GPU by first placing it in the IFPC entry (C) state and then in the IFPC (D) state.
[0040] The hysteresis timeout (B) state can help to avoid unnecessary sequences of IFPC entries and exits if there is some immediate additional workload after the GPU has completed processing a command. This can be useful, for example, when the GPU receives an unpredictable workload from the CPU.
[0041] In an illustrative example, there may be 480 samples per second for the GPU to process. In other words, the interval between two adjacent IPCC interrupts 402 may be approximately 2.08 ms. Based on the projection, it may take the GPU 0.22 ms to complete the MV grid generation for each sample. In other words, for each sample, the GPU may be in the active (A) state for approximately 0.22 ms. Furthermore, as shown in FIG. 4, the total duration between two adjacent IPCC interrupts 402 may be equal to the sum of the durations associated with all five GPU states, and since it is known that 1) the durations of the hysteresis timeout (B) states may be approximately 0.3 ms each, 2) the durations of the IFPC entry (C) states may be approximately 0.1 ms each, and 3) the durations of the IFPC exit (E) states may be approximately 0.08 ms each, it may be calculated that the duration of each instance of the IFPC (D) state in this example may be approximately 1.38 ms. In other words, the total GPU rail active duration may be approximately 0.7 ms for each interval between two adjacent IPCC interrupts 402.
[0042] FIG. 5 is a diagram 500 illustrating an example GPU state timeline associated with an IFPC, according to one or more aspects. In one or more configurations, since the XR workload may be of a persistent type that occurs at fixed intervals across the LSR context, additional adaptations, as described in more detail below, may be employed to further conserve power while the XR pipeline performance goals may continue to be met. In particular, referring back to FIG. 3, in one configuration, the XR application 302 may provide a hint to the GMU 314 firmware corresponding to a timer value (e.g., T1). In another configuration, the hint may be provided to the GMU 314 firmware by the EVA firmware 308 during the LSR context setup. In yet another configuration, the graphics driver 310 or the GMU 314 firmware may derive the hint based on machine learning techniques.
[0043] The timer value T1 may be related to controlling the flow between the EVA and the GMU and may correspond to the interval between two adjacent IPCC interrupts 502 sent by the EVA firmware to the GMU firmware. Thus, in one or more configurations, based on the latency associated with the IFPC exit (E) state, the GMU firmware may trigger or reset a timer (e.g., Tg) upon receiving the IPCC interrupt 502 from the EVA firmware. The value of the timer Tg may be calculated by subtracting the latency associated with the IFPC exit (E) state from the timer value T1, i.e., Tg=T1-duration per instance of the E state.
[0044] Thus, the GMU firmware can initiate the wake-up of the GPU upon expiration of timer Tg rather than upon receipt of a subsequent IPCC interrupt 502', thereby advancing the timeline for waking up the GPU so that the GPU can be ready in an active (A) state to process commands approximately when the GMU receives the subsequent IPCC interrupt 502'. Thus, the delay between receipt of an IPCC interrupt 502' and the time the GPU is awake and ready to process commands can be eliminated, or at least significantly reduced, and the GPU can begin fetching and processing commands for the current sample immediately after receiving the corresponding IPCC interrupt 502'.
[0045] Additionally, once the timer value T1 is obtained, the GMU may also remove the hysteresis timeout (B) condition (i.e., set the hysteresis timeout duration to 0) because it may be known in the LSR context that timer Tg has expired and there may be no further immediate GPU workload until the next IPCC interrupt is received.
[0046] The total duration between two adjacent IPCC interrupts 502 may be equal to the sum of the durations associated with all five GPU states as shown in Figure 5, and since it is known that 1) the durations of the hysteresis timeout (B) states may each be 0 ms, 2) the durations of the IFPC entry (C) states may each be approximately 0.1 ms, and 3) the durations of the IFPC exit (E) states may each be approximately 0.08 ms, it may be calculated that the duration of the IFPC (D) state in this example may be approximately 1.68 ms. In other words, the total GPU rail active duration may be approximately 0.4 ms for each interval between two adjacent IPCC interrupts 502. Thus, compared to the timeline shown in Figure 4, the total GPU rail active duration in Figure 5 may be reduced by approximately 42%, which may be associated with a corresponding power savings.
[0047] Thus, according to one or more aspects, at least one of the XR application, the EVA driver, or the graphics driver (e.g., the graphics kernel driver) can provide a hint regarding the timer value T1 to the GMU firmware. As a result, a hysteresis timeout that is unnecessary for fixed periodic workloads in the LSR context can be avoided. In other words, the GPU can enter the IFPC(D) state immediately after completing processing of the command. Avoiding the hysteresis timeout can save power. Furthermore, the wakeup of the GPU can be initiated before the IPCC interrupt and the corresponding workload are actually received. Thus, the delay in processing the command associated with the delay between the receipt of the IPCC interrupt and the time when the GPU is awake and ready to process the command can be eliminated. Eliminating the delay can provide performance benefits.
[0048] In one or more configurations, the hint regarding the timer value T1 may be implemented as an extension within a graphics API to allow an application (e.g., an XR / AR / VR application) to pass the timer value T1 (e.g., the interval between workload submissions to the GPU) to a graphics driver (e.g., a graphics kernel driver).
[0049] In one or more configurations, in addition to the GMU / GPU, the techniques described above may be similarly applied to other intellectual property (IP) blocks (e.g., video, EVA, etc.) to improve power collapse operations in the respective IP blocks.
[0050] 6 is a call flow diagram 600 illustrating example communications between an application 602 (e.g., XR application 302), a first component 604 (e.g., EVA firmware 308), and a GPU 606 (including a GMU within the GPU 606) in accordance with one or more techniques of this disclosure. At 608, the GPU 606 can receive an indication from the application 602 of the duration of a timer associated with exiting an IFPC state.
[0051] In one configuration, the duration of the timer may further be based at least in part on the IFPC termination latency.
[0052] At 610, the GPU 606 may receive a first instruction to begin processing one or more predefined workloads. The user space may submit the one or more predefined workloads to the GPU scheduler (GMU) once. Additionally, the GPU scheduler (GMU) may repeatedly submit the one or more predefined workloads to the GPU at regular intervals upon an event such as an IPCC interrupt.
[0053] In one configuration, the one or more predefined workloads may be one or more LSR workloads (the LSR workload may be a predefined workload for generating an MV grid based on a depth buffer and head pose), in a further configuration, the one or more predefined workloads may be any workload that may be repeatedly submitted to the GPU.
[0054] In one configuration, the first indication may be an IPCC interrupt.
[0055] In one configuration, the first indication may be received from at least one of a scheduler, an application, or a service layer.
[0056] In one configuration, the one or more predefined workloads may be associated with at least one of an XR application, an AR application, or a VR application.
[0057] At 612, the GPU 606 may trigger a timer upon receiving the first instruction.
[0058] At 614, the GPU 606 may process one or more predefined workloads upon triggering a timer associated with the exit of the IFPC state.
[0059] At 616, the GPU 606 may enter an IFPC state upon finishing processing one or more predefined workloads.
[0060] At 618, the GPU 606 may detect the expiration of the timer.
[0061] At 620, upon detecting the expiration of the timer, the GPU 606 may exit the IFPC state.
[0062] In one configuration, the hysteresis timeout within the first period associated with the timer is zero.
[0063] At 622, the GPU 606 may receive a second instruction to begin processing one or more predefined workloads.
[0064] 7 is a flowchart 700 of an exemplary method of graphics processing in accordance with one or more techniques of this disclosure. The method may be performed by an apparatus for graphics processing, a GPU, a CPU, a wireless communication device, etc., such as those used in connection with the embodiments of FIGS. 1-6.
[0065] At 702, the apparatus may receive from an application an indication of a duration of a timer associated with exiting an IFPC state. For example, with reference to FIG. 6, at 608, GPU 606 may receive from application 602 an indication of a duration of a timer associated with exiting an IFPC state. Further, 702 may be executed by processing unit 120.
[0066] At 704, the device may process one or more predefined workloads upon triggering a timer associated with exiting the IFPC state. For example, with reference to FIG. 6, at 614, GPU 606 may process one or more predefined workloads upon triggering a timer associated with exiting the IFPC state. Additionally, 704 may be executed by processing unit 120.
[0067] At 706, the apparatus may enter the IFPC state upon completion of processing of the one or more predefined workloads. For example, with reference to FIG. 6, at 616, the GPU 606 may enter the IFPC state upon completion of processing of the one or more predefined workloads. Additionally, 706 may be executed by the processing unit 120.
[0068] At 708, the device may exit the IFPC state upon detecting expiration of the timer. For example, with reference to FIG. 6, at 620, the GPU 606 may exit the IFPC state upon detecting expiration of the timer. Additionally, 708 may be performed by the processing unit 120.
[0069] 8 is a flowchart 800 of an exemplary method of graphics processing in accordance with one or more techniques of this disclosure. The method may be performed by an apparatus for graphics processing, a GPU, a CPU, a wireless communication device, etc., such as those used in connection with the embodiments of FIGS. 1-6.
[0070] At 802, the apparatus may receive from an application an indication of a duration of a timer associated with exiting an IFPC state. For example, with reference to FIG. 6, at 608, GPU 606 may receive from application 602 an indication of a duration of a timer associated with exiting an IFPC state. Further, 802 may be executed by processing unit 120.
[0071] At 808, the apparatus may process one or more predefined workloads upon triggering a timer associated with exiting the IFPC state. For example, with reference to FIG. 6, at 614, the GPU 606 may process one or more predefined workloads upon triggering a timer associated with exiting the IFPC state. Additionally, 808 may be executed by the processing unit 120.
[0072] At 810, the apparatus may enter the IFPC state upon finishing processing one or more predefined workloads. For example, referring to FIG. 6, at 616, the GPU 606 may enter the IFPC state upon finishing processing one or more predefined workloads. Additionally, 810 may be executed by the processing unit 120.
[0073] At 814, the device may exit the IFPC state upon detecting expiration of the timer. For example, with reference to FIG. 6, at 620, GPU 606 may exit the IFPC state upon detecting expiration of the timer. Further, 814 may be executed by processing unit 120.
[0074] In one configuration, the apparatus may receive a first instruction to begin processing one or more predefined workloads at 804. For example, with reference to FIG. 6, at 610, the GPU 606 may receive a first instruction to begin processing one or more predefined workloads. Further, 804 may be executed by the processing unit 120.
[0075] In 806, the apparatus may trigger a timer upon receiving the first instruction. For example, referring to FIG. 6, in 612, the GPU 606 may trigger a timer upon receiving the first instruction. Further, 806 may be executed by the processing unit 120.
[0076] The apparatus may detect expiration of the timer at 812. For example, with reference to FIG. 6, GPU 606 may detect expiration of the timer at 618. Additionally, 812 may be executed by processing unit 120.
[0077] In one configuration, the one or more predefined workloads may be one or more LSR workloads.
[0078] In one configuration, the first indication may be an IPCC interrupt.
[0079] In one configuration, the first indication may be received from at least one of a scheduler, an application, or a service layer.
[0080] In one configuration, the one or more predefined workloads may be associated with at least one of an XR application, an AR application, or a VR application.
[0081] In one configuration, the apparatus may receive a second instruction to begin processing the one or more predefined workloads at 816. For example, with reference to FIG. 6, at 622, GPU 606 may receive a second instruction to begin processing the one or more predefined workloads. Further, 816 may be executed by processing unit 120.
[0082] In one configuration, with reference to FIG. 6, exiting the IFPC state upon detecting expiration of the timer may include exiting the IFPC state in the GPU 606.
[0083] In one configuration, the duration of the timer may further be based at least in part on the IFPC termination latency.
[0084] In one configuration, the hysteresis timeout within the first period associated with the timer may be zero.
[0085] In an arrangement, a method or apparatus for graphics processing is provided. The apparatus may be a GPU, a CPU, or some other processor capable of performing graphics processing. In an aspect, the apparatus may be a processing unit 120 in the device 104, or some other hardware in the device 104 or another device. The apparatus may include means for receiving an indication of a duration of a timer associated with exiting an IFPC state from an application. The apparatus may further include means for processing one or more predefined workloads upon triggering a timer associated with exiting an IFPC state. The apparatus may further include means for entering an IFPC state upon finishing processing the one or more predefined workloads. The apparatus may further include means for exiting an IFPC state upon detecting expiration of the timer.
[0086] In one configuration, the apparatus may further include means for receiving a first indication to start processing the one or more predefined workloads. The apparatus may further include means for triggering a timer upon receiving the first indication. The apparatus may further include means for detecting expiration of the timer. In one configuration, the one or more predefined workloads may be one or more LSR workloads. In one configuration, the first indication may be an IPCC interrupt. In one configuration, the first indication may be received from at least one of a scheduler, an application, or a service layer. In one configuration, the one or more predefined workloads may be associated with at least one of an XR application, an AR application, or a VR application. In one configuration, the apparatus may further include means for receiving a second indication to start processing the one or more predefined workloads. In one configuration, exiting the IFPC state upon detecting expiration of the timer may include exiting the IFPC state in the GPU. In one configuration, the duration of the timer may further be based at least in part on an IFPC exit latency. In one configuration, the hysteresis timeout within the first period associated with the timer may be zero.
[0087] It should be understood that the particular order or hierarchy of the blocks / steps in the processes, flowcharts, and / or call flow diagrams disclosed herein is an illustration of an example approach. Based on design preferences, it should be understood that the particular order or hierarchy of the blocks / steps in the processes, flowcharts, and / or call flow diagrams may be rearranged. Furthermore, some blocks / steps may be combined and / or omitted. Other blocks / steps may be added. The accompanying method claims present elements of the various blocks / steps in an example order, and are not meant to be limited to the particular order or hierarchy presented.
[0088] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Thus, the claims are not limited to the aspects set forth herein but are to be accorded the full scope consistent with the claim language, and reference to an element in the singular does not mean "the only one" unless so expressly stated, but rather means "one or more." The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" should not necessarily be construed as preferred or advantageous over other aspects.
[0089] Unless otherwise specified, the term "some" refers to one or more, and the term "or" may be interpreted as "and / or" unless the context dictates otherwise. Combinations such as "at least one of A, B, or C," "one or more of A, B, or C," "at least one of A, B, and C," "one or more of A, B, and C," and "A, B, C, or any combination thereof" include any combination of A, B, and / or C, and may include multiple As, multiple Bs, or multiple Cs. Specifically, combinations such as "at least one of A, B, or C," "one or more of A, B, or C," "at least one of A, B, and C," "one or more of A, B, and C," and "A, B, C, or any combination thereof" may be A only, B only, C only, A and B, A and C, B and C, or A and B and C, and any such combination may include one or more elements of A, B, or C. All structural and functional equivalents of the elements of the various aspects described throughout this disclosure that are known or that later become known to those of skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be made public, regardless of whether such disclosure is expressly recited in the claims. Words such as "module," "mechanism," "element," "device," and the like may not be substitutes for the word "means." Therefore, no element of the claims should be construed as a means-plus-function unless the element is expressly recited using the phrase "means for."
[0090] In one or more examples, the functions described herein may be implemented in hardware, software, firmware, or any combination thereof. For example, although the term "processing unit" is used throughout this disclosure, such a processing unit may be implemented in hardware, software, firmware, or any combination thereof. If any function, processing unit, technique, or other module described herein is implemented in software, the function, processing unit, technique, or other module described herein may be stored on or transmitted over as one or more instructions or code on a computer-readable medium.
[0091] Computer-readable media may include computer data storage media or communication media, including any medium that facilitates transfer of a computer program from one place to another. In this manner, computer-readable media may generally correspond to (1) tangible computer-readable storage media that are non-transitory, or (2) communication media, such as signals or carrier waves. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. By way of example and not limitation, such computer-readable media may comprise RAM, ROM, EEPROM, compact disc-read only memory (CD-ROM) or other optical disk storage, magnetic disk storage, or other magnetic storage devices. As used herein, disk and disc include compact disc (CD), laser disc, optical disk, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks usually reproduce data magnetically while discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media. A computer program product may include a computer-readable medium.
[0092] The techniques of the present disclosure may be implemented in a wide variety of devices or implementations, including wireless handsets, integrated circuits (ICs) or sets of ICs, e.g., chipsets. In this disclosure, various components, modules or units have been described to highlight functional aspects of devices configured to implement the disclosed techniques, but the components, modules or units do not necessarily require implementation by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit or may be provided by a collection of interoperable hardware units including one or more processors as described above, together with appropriate software and / or firmware. Thus, the term "processor" as used herein may refer to any of the above structures or any other structures suitable for implementing the techniques described herein. Also, the techniques may be fully implemented in one or more circuits or logic elements.
[0093] The following aspects are exemplary only and can be combined with other aspects or teachings described herein without limitation.
[0094] Aspect 1 is a method of graphics processing, comprising the steps of: receiving from an application an indication of a duration of a timer associated with exiting the IFPC state; processing one or more predefined workloads upon triggering the timer associated with exiting the IFPC state; entering an IFPC state upon finishing processing the one or more predefined workloads; and exiting the IFPC state upon detecting expiration of the timer; A method comprising:
[0095] Aspect 2 may be combined with aspect 1 and includes receiving a first instruction to start processing one or more predefined workloads; triggering a timer upon receiving the first instruction; and detecting expiration of the timer; Further includes.
[0096] Aspect 3 may be combined with aspect 2, including where the one or more predefined workloads are one or more LSR workloads.
[0097] Aspect 4 may be combined with any of aspects 2 and 3, and includes the first indication being an IPCC interrupt.
[0098] The fifth aspect may be combined with any of the second to fourth aspects, and includes that the first indication is received from at least one of a scheduler, an application, or a service layer.
[0099] Example 6 may be combined with any of Examples 2 to 5, and includes the one or more predefined workloads being associated with at least one of an XR application, an AR application, or a VR application.
[0100] Aspect 7 may be combined with any of aspects 2-6 and further includes receiving a second instruction to start processing the one or more predefined workloads.
[0101] Example 8 may be combined with any of Examples 1-7, and includes where exiting the IFPC state upon detecting expiration of the timer includes exiting the IFPC state in the GPU.
[0102] Aspect 9 may be combined with any of aspects 1-8, including the timer period being further based at least in part on the IFPC termination latency.
[0103] Aspect 10 may be combined with any of aspects 1 to 9 and includes a hysteresis timeout within a first period associated with the timer being 0.
[0104] Example 11 is an apparatus for graphics processing including at least one processor coupled to a memory and configured to perform a method according to any of examples 1 to 10.
[0105] Aspect 12 may be combined with aspect 11, including where the apparatus is a wireless communication device.
[0106] A thirteenth aspect is an apparatus for graphics processing, comprising means for carrying out the method according to any one of the first to tenth aspects.
[0107] Aspect 14 is a non-transitory computer-readable medium storing computer-executable code that, when executed by at least one processor, causes the at least one processor to perform a method according to any of aspects 1 to 10.
[0108] Various aspects have been described herein. These and other aspects are within the scope of the following claims.
Claims
1. A device for graphics processing, Memory and At least one processor coupled to the memory, The at least one processor is The application receives an instruction for the duration of the timer associated with ending the inter-frame power collapse (IFPC) state. When the timer associated with terminating the IFPC state is triggered, one or more default workloads are processed, When processing of one or more of the aforementioned default workloads is completed, the IFPC state is started, When the timer expires, the IFPC state is terminated. It is structured in such a way. Device.
2. The aforementioned at least one processor, Upon receiving a first instruction to begin processing one or more of the aforementioned default workloads, Upon receiving the first instruction, the timer is triggered, The timer's expiration is detected. It is further structured in such a way. The apparatus according to claim 1.
3. The apparatus according to claim 2, wherein the one or more default workloads are one or more late reprojection (LSR) workloads.
4. The apparatus according to claim 2, wherein the first instruction is an inter-processor communication controller (IPCC) interrupt.
5. The apparatus according to claim 2, wherein the first instruction is received from at least one of the scheduler, the application, or the service layer.
6. The apparatus according to claim 2, wherein the one or more default workloads are associated with at least one of an augmented reality (XR) application, an augmented reality (AR) application, or a virtual reality (VR) application.
7. The aforementioned at least one processor, Receiving a second instruction to start processing one or more of the aforementioned default workloads, It is further structured in such a way. The apparatus according to claim 2.
8. The apparatus according to claim 1, wherein the at least one processor is further configured to terminate the IFPC state in a graphics processing unit (GPU).
9. The apparatus according to claim 1, wherein the period of the timer is at least partially based on the IFPC termination latency.
10. The apparatus according to claim 1, wherein the hysteresis timeout within a first period associated with the timer is 0.
11. The apparatus according to claim 1, wherein the apparatus is a wireless communication device.
12. A method of graphics processing, Receiving instructions from the application for the duration of a timer associated with ending an inter-frame power collapse (IFPC) state, When the timer associated with terminating the IFPC state is triggered, one or more default workloads are processed, When one or more of the aforementioned default workloads are processed, the IFPC state is initiated, When the expiration of the timer is detected, the IFPC state is terminated, Methods that include...
13. A computer-readable medium for storing computer executable code, wherein when the code is executed by at least one processor, the at least one processor is configured to: The application receives an instruction for the duration of the timer associated with ending the inter-frame power collapse (IFPC) state. When the timer associated with terminating the IFPC state is triggered, one or more default workloads are processed. When processing of one or more of the aforementioned default workloads is completed, the IFPC state is initiated. When the timer expires, the IFPC state is terminated. Computer-readable media.