A system and method for performing efficient multi-GPU rendering of geometry by performing a pre-test on an interleaved screen area before rendering
The method of dividing GPU rendering responsibilities based on interleaved screen regions and performing pre-tests addresses inefficiencies in multi-GPU rendering, optimizing workload distribution and enhancing performance by ensuring each GPU renders only necessary geometry.
Patent Information
- Application Number
- JP2024167072
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-02-03
- Filing Date
- 2024-09-26
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2041-02-01
AI Technical Summary
Existing technologies face challenges in efficiently utilizing multiple GPUs for rendering complex scenes or images, as they struggle to evenly distribute the workload and support increased screen pixels and geometry density, leading to inefficiencies and imbalances in processing.
A method for multi-GPU rendering that involves dividing the responsibility for rendering geometry among multiple GPUs based on interleaved screen regions, performing a pre-test to determine the relationship of geometry pieces to these regions, and using the generated information to skip unnecessary rendering, thereby optimizing workload distribution.
This approach allows for more efficient rendering of complex scenes by balancing the workload among GPUs, reducing processing time, and ensuring that each GPU renders only the necessary geometry, thereby improving overall performance and efficiency.
Smart Images

Figure 0007702056000001 
Figure 0007702056000002 
Figure 0007702056000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to graphics processing, and more specifically, to multi-GPU cooperation when rendering an image for an application.
Background Art
[0002] In recent years, there has been an ongoing effort for online services that enable online or cloud gaming in a streaming format between a cloud gaming server and a client connected through a network. The streaming format is becoming increasingly popular. This is because game titles are available on demand, more complex games can be executed, network connections can be made between players in the case of multiplayer gaming, assets can be shared between players, instantaneous experiences can be shared between players and / or spectators, friends can watch friends play video games, and friends can participate in a friend's ongoing game play, among other things.
[0003] A cloud gaming server may be configured to provide resources to one or more clients and / or applications. That is, the cloud gaming server may be configured with resources that enable high throughput. For example, there are limits to the performance that an individual Graphics Processing Unit (GPU) can achieve. To render more complex scenes or to use more complex algorithms (such as materials, lighting, etc.) when generating a scene, it may be desirable to render a single image using multiple GPUs. However, it is difficult to use these graphics processing units evenly. Further, even when there are multiple GPUs for processing images for an application using conventional techniques, it is not possible to support a corresponding increase in both the number of screen pixels and the geometry density (it is impossible to write four times as many pixels to an image and / or process four times as many vertices or primitives with four GPUs).
[0004] Embodiments of the present disclosure have been made under such a background.
SUMMARY OF THE INVENTION
[0005] Embodiments of the present disclosure relate to rendering a single image using multiple GPUs in cooperation, for example, performing multi-GPU rendering of geometry for an application by performing a pre-test on a screen area (which may be interleaved) before rendering.
[0006] Embodiments of the present disclosure disclose a method for performing graphics processing. The method includes rendering graphics for an application using a plurality of graphics processing units (GPUs). The method divides the responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, and each GPU has a corresponding division of the responsibility known to the plurality of GPUs. The screen regions are interleaved. The method includes allocating a plurality of pieces of the geometry of an image frame to the plurality of GPUs for geometry testing. The method includes allocating to a GPU a piece of the geometry of an image frame generated by an application for geometry testing. The method includes performing a geometry test in the GPU to generate information regarding the piece of the geometry and its relationship to each of the plurality of screen regions. The method includes rendering the piece of the geometry using the information in each of the plurality of GPUs, and using the information can include, for example, completely skipping the rendering if it is determined that the piece of the geometry does not overlap any of the screen regions assigned to a given GPU.
[0007] In another embodiment, a non-transitory computer-readable medium for performing the method is disclosed. The computer-readable medium includes program instructions for rendering graphics for an application using a plurality of graphics processing units (GPUs). The computer-readable medium includes program instructions for dividing the responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, where each GPU has a corresponding division of the responsibility known to the plurality of GPUs, and the screen regions in the plurality of screen regions are interleaved. The computer-readable medium includes program instructions for allocating a piece of the geometry of an image frame generated by the application to the GPUs for pre-geometry testing. The computer-readable medium includes program instructions for performing pre-geometry testing in the GPUs to generate information regarding the piece of geometry and its relationship to each of the plurality of screen regions. The computer-readable medium includes program instructions for using the information in each of the plurality of GPUs when rendering the image frame.
[0008] In yet another embodiment, a computer system is disclosed. The computer system includes a processor and a memory coupled to the processor and storing instructions that, when executed by the computer system, cause the computer system to execute a method for performing graphics processing. The method includes rendering graphics for an application using a plurality of graphics processing units (GPUs). The method includes dividing the responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, each GPU having a corresponding division of the responsibility known to the plurality of GPUs, and the screen regions in the plurality of screen regions being interleaved. The method includes allocating to the GPUs pieces of the geometry of an image frame generated by the application for pre-geometry testing. The method includes performing pre-geometry testing in the GPUs to generate information regarding the pieces of the geometry and its relationship to each of the plurality of screen regions. The method includes using the information in each of the plurality of GPUs when rendering the image frame.
[0009] Embodiments of the present disclosure disclose a method for performing graphics processing. The method includes rendering graphics for an application using a plurality of graphics processing units (GPUs). The method includes dividing the responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, each GPU having a corresponding division of the responsibility known to the plurality of GPUs. The method includes performing a geometry test in a GPU pre-test on a plurality of pieces of the geometry of an image frame generated by the application to generate information regarding the relationship of each piece of the geometry to each of the plurality of screen regions. The method includes rendering the plurality of pieces of the geometry in each of the plurality of GPUs using the information generated for each of the plurality of pieces of the geometry, where using the information includes, for example, completely skipping rendering if a piece of the geometry is determined not to overlap any of the screen regions assigned to a given GPU.
[0010] In another embodiment, a non-transitory computer-readable medium for performing the method is disclosed. The computer-readable medium includes program instructions for rendering graphics for an application using a plurality of graphics processing units (GPUs). The computer-readable medium includes program instructions for dividing the responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, wherein each GPU has a corresponding division of the responsibility known to the plurality of GPUs. The computer-readable medium includes program instructions for performing a geometry test in a GPU pre-test on a plurality of pieces of the geometry of an image frame generated by an application to generate information regarding the relationship of each piece of the geometry to each of the plurality of screen regions. The computer-readable medium includes program instructions for rendering a plurality of pieces of the geometry in each of the plurality of GPUs using the information generated for each of the plurality of pieces of the geometry, wherein using the information includes, for example, completely skipping the rendering if it is determined that a piece of the geometry does not overlap with any of the screen regions assigned to a given GPU.
[0011] In yet another embodiment, a computer system is disclosed. The computer system includes a processor and a memory coupled to the processor and storing instructions that, when executed by the computer system, cause the computer system to execute a method for performing graphics processing. The method includes rendering graphics for an application using a plurality of graphics processing units (GPUs). The method includes dividing responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, each GPU having a corresponding division of the responsibility known to the plurality of GPUs. The method includes performing a geometry test in a GPU pre-test on a plurality of pieces of the geometry of an image frame generated by the application to generate information regarding the relationship of each piece of the geometry to each of the plurality of screen regions. The method includes using the information generated for each of the plurality of pieces of the geometry to render the plurality of pieces of the geometry in each of the plurality of GPUs, and when using the information, for example, completely skipping rendering if it is determined that a piece of the geometry does not overlap any of the screen regions assigned to a given GPU.
[0012] In an embodiment of the present disclosure, a method for performing graphics processing is disclosed. The method includes rendering graphics for an application using a plurality of graphics processing units (GPUs). The method includes dividing the responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, each GPU having a corresponding division of the responsibility known to the plurality of GPUs. The method includes rendering a first plurality of pieces of geometry in the plurality of GPUs during a rendering phase of a previous image frame generated by the application. The method includes generating statistical values for the rendering of the previous image frame. The method includes, based on the statistical values, allocating a second plurality of pieces of geometry of a current image frame generated by the application to the plurality of GPUs for geometry testing. The method includes performing geometry testing on the second plurality of pieces of geometry in the current image frame to generate information regarding each piece of the second plurality of pieces of geometry and its relationship to each of the plurality of screen regions, the geometry testing being performed in each of the plurality of GPUs based on the allocation. The method includes using the information generated for each of the second plurality of pieces of geometry to render the second plurality of pieces of geometry in each of the plurality of GPUs, and when using the information, completely skipping the rendering if, for example, a piece of geometry is determined not to overlap any of the screen regions allocated to a given GPU.
[0013] In another embodiment, a non-transitory computer-readable medium for performing a method is disclosed. The computer-readable medium includes program instructions for rendering graphics for an application using a plurality of graphics processing units (GPUs). The computer-readable medium includes program instructions for dividing the responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, where each GPU has a corresponding division of the responsibility known to the plurality of GPUs. The computer-readable medium includes program instructions for rendering a first plurality of pieces of geometry in the plurality of GPUs during a rendering phase of a previous image frame generated by the application. The computer-readable medium includes program instructions for generating statistical values for the rendering of the previous image frame. The computer-readable medium includes program instructions for allocating a second plurality of pieces of geometry of the current image frame generated by the application to the plurality of GPUs for geometry testing based on the statistical values. The computer-readable medium has program instructions for performing a geometry test on the second plurality of pieces of geometry in the current image frame to generate information regarding the relationship of each piece of the second plurality of pieces of geometry to each of the plurality of screen regions, where the geometry test is performed in each of the plurality of GPUs based on the allocation. The computer-readable medium has program instructions for rendering the second plurality of pieces of geometry in each of the plurality of GPUs using the information generated for each piece of the second plurality of pieces of geometry, and when using the information, for example, rendering can be completely skipped if it is determined that a piece of geometry does not overlap with any of the screen regions allocated to a given GPU.
[0014] In yet another embodiment, a computer system is disclosed. The computer system includes a processor and a memory coupled to the processor and storing instructions that, when executed by the computer system, cause the computer system to execute a method for performing graphics processing. The method includes rendering graphics for an application using a plurality of graphics processing units (GPUs). The method includes dividing responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, with each GPU having a corresponding division of the responsibility known to the plurality of GPUs. The method includes rendering a first plurality of pieces of geometry in the plurality of GPUs during a rendering phase of a previous image frame generated by the application. The method includes generating statistical values for the rendering of the previous image frame. The method includes, based on the statistical values, allocating a second plurality of pieces of geometry of a current image frame generated by the application to the plurality of GPUs for geometry testing. The method includes performing geometry testing on the second plurality of pieces of geometry in the current image frame to generate information regarding each piece of the second plurality of pieces of geometry and its relationship to each of the plurality of screen regions, the geometry testing being performed in each of the plurality of GPUs based on the allocation. The method includes using the information generated for each of the second plurality of pieces of geometry to render the second plurality of pieces of geometry in each of the plurality of GPUs, and when using the information, for example, completely skipping the rendering if it is determined that a piece of geometry does not overlap any of the screen regions allocated to a given GPU.
[0015] Embodiments of the present disclosure disclose a method for performing graphics processing. The method includes rendering graphics for an application using a plurality of graphics processing units (GPUs). The method divides the responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, and each GPU has a corresponding division of the responsibility known to the plurality of GPUs. The method allocates a plurality of pieces of the geometry of the image frame to the plurality of GPUs for geometry testing. The method includes setting a first state that configures one or more shaders to perform a geometry test. The method includes performing a geometry test on the plurality of pieces of geometry in the plurality of GPUs to generate information regarding the relationship of each piece of geometry to each of the plurality of screen regions. The method includes setting a second state that configures one or more shaders to perform rendering. The method includes rendering the plurality of pieces of geometry in each of the plurality of GPUs using the information generated for each of the plurality of pieces of geometry, and when using the information, completely skipping the rendering if, for example, a piece of geometry is determined not to overlap any of the screen regions assigned to a given GPU.
[0016] In another embodiment, a non-transitory computer-readable medium for performing a method is disclosed. The computer-readable medium includes program instructions for rendering graphics for an application using a plurality of graphics processing units (GPUs). The computer-readable medium has program instructions for dividing the responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, each GPU having a corresponding division of the responsibility known to the plurality of GPUs. The computer-readable medium includes program instructions for allocating a plurality of pieces of the geometry of an image frame to the plurality of GPUs for geometry testing. The computer-readable medium includes program instructions for setting a first state that configures one or more shaders to perform geometry testing. The computer-readable medium includes program instructions for performing geometry testing on the plurality of pieces of the geometry in the plurality of GPUs to generate information regarding the relationship of each piece of the geometry to each of the plurality of screen regions. The computer-readable medium includes program instructions for setting a second state that configures one or more shaders to perform rendering. The computer-readable medium has program instructions for rendering the plurality of pieces of the geometry in each of the plurality of GPUs using the information generated for each of the plurality of pieces of the geometry, and when using the information, for example, completely skipping rendering if a piece of the geometry is determined not to overlap any of the screen regions assigned to a given GPU.
[0017] In yet another embodiment, a computer system is disclosed. The computer system includes a processor and a memory coupled to the processor and storing instructions that, when executed by the computer system, cause the computer system to execute a method for performing graphics processing. The method includes rendering graphics for an application using a plurality of graphics processing units (GPUs). The method includes dividing responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, each GPU having a corresponding division of the responsibility known to the plurality of GPUs. The method includes allocating a plurality of pieces of the geometry of an image frame to the plurality of GPUs for geometry testing. The method includes setting a first state that configures one or more shaders to perform geometry testing. The method includes performing geometry testing on the plurality of pieces of geometry in the plurality of GPUs to generate information regarding the relationship of each piece of geometry to each of the plurality of screen regions. The method includes setting a second state that configures one or more shaders to perform rendering. The method includes rendering the plurality of pieces of geometry in each of the plurality of GPUs using the information generated for each of the plurality of pieces of geometry, and completely skipping rendering if, for example, a piece of geometry is determined not to overlap any of the screen regions assigned to a given GPU when using the information.
[0018] Embodiments of the present disclosure disclose a method for performing graphics processing. The method includes rendering graphics for an application using a plurality of graphics processing units (GPUs). The method divides the responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, and each GPU has a corresponding division of the responsibility known to the plurality of GPUs. The method includes allocating a plurality of pieces of the geometry of an image frame to the plurality of GPUs for geometry testing. The method includes interleaving a first set of shaders that perform geometry testing and rendering on a first set of pieces of the geometry and a second set of shaders that perform geometry testing and rendering on a second set of pieces of the geometry. The geometry testing generates corresponding information regarding each piece of the geometry within the first set or the second set and its relationship to each of the plurality of screen regions. The plurality of GPUs use the corresponding information to render each piece of the geometry within the first set or the second set. When using the information, for example, rendering is completely skipped if it is determined that a piece of the geometry does not overlap any of the screen regions assigned to a given GPU.
[0019] In another embodiment, a non-transitory computer-readable medium for performing the method is disclosed. The computer-readable medium includes program instructions for rendering graphics for an application using a plurality of graphics processing units (GPUs). The computer-readable medium has program instructions for dividing the responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, each GPU having a corresponding division of the responsibility known to the plurality of GPUs. The computer-readable medium includes program instructions for allocating a plurality of pieces of the geometry of an image frame to the plurality of GPUs for geometry testing. The computer-readable medium includes programming instructions for interleaving a first set of shaders that perform geometry testing and rendering on a first set of pieces of the geometry and a second set of shaders that perform geometry testing and rendering on a second set of pieces of the geometry. The geometry testing generates corresponding information regarding each piece of the geometry within the first or second set and its relationship to each of the plurality of screen regions. The plurality of GPUs use the corresponding information to render each piece of the geometry within the first or second set. When using the information, for example, if it is determined that a piece of the geometry does not overlap any of the screen regions assigned to a given GPU, the rendering is skipped entirely.
[0020] In yet another embodiment, a computer system is disclosed. The computer system includes a processor and a memory coupled to the processor and storing instructions that, when executed by the computer system, cause the computer system to execute a method for performing graphics processing. The method includes rendering graphics for an application using a plurality of graphics processing units (GPUs). The method includes dividing the responsibility for rendering the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions, with each GPU having a corresponding division of the responsibility known to the plurality of GPUs. The method includes allocating a plurality of pieces of the geometry of the image frame to the plurality of GPUs for geometry testing. The method includes interleaving a first set of shaders that perform geometry testing and rendering on a first set of pieces of the geometry and a second set of shaders that perform geometry testing and rendering on a second set of pieces of the geometry. The geometry testing generates corresponding information regarding the relationship of each piece of the geometry within the first or second set to each of the plurality of screen regions. The plurality of GPUs use the corresponding information to render each piece of the geometry within the first or second set. When using the information, for example, rendering is completely skipped if it is determined that a piece of the geometry does not overlap with any of the screen regions assigned to a given GPU.
[0021] Other aspects of the disclosure will become apparent from the accompanying drawings, which illustrate, by way of example, the principles of the disclosure in conjunction with the following detailed description.
[0022] The present disclosure may be best understood by reference to the accompanying drawings in conjunction with the following description.
Brief Description of the Drawings
[0023]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 7A
Figure 7B-1
Figure 7B-2
Figure 7C
Figure 8A
Figure 8B
Figure 9
Figure 10
Figure 11A
Figure 11B
Figure 12A
Figure 12B
Figure 13A
Figure 13B
Figure 14
Embodiments for Carrying Out the Invention
[0024] The following detailed description includes many specific details for purposes of explanation, but as will be apparent to those skilled in the art, many variations and modifications to the following details are within the scope of the present disclosure. Accordingly, the aspects of the present disclosure described below are presented without any loss of generality with respect to the claims that follow this description and without imposing any limitation on the claims.
[0025] Generally speaking, there are limits to the performance that an individual GPU can achieve, which are derived from, for example, limits on how large a GPU can be. In embodiments of the present disclosure, it is desirable to use multiple GPUs to render a single image in order to render more complex scenes or to use more complex algorithms (such as materials, lighting, etc.). Specifically, in various embodiments of the present disclosure, methods and systems are described that are configured to perform multi-GPU rendering of geometry for an application by performing a pre-test of the geometry against a screen region (which can be interleaved). The multiple GPUs cooperate to generate an image. The responsibility or obligation or responsiveness for rendering is divided among the multiple GPUs based on the screen region. Before rendering the geometry, the GPU generates information regarding the relationship of the geometry to the screen region. This allows the GPU to render the geometry more efficiently or to avoid rendering altogether. For example, this has the advantage that multiple GPUs can render more complex scenes and / or images in the same amount of time.
[0026] Based on the foregoing general understanding of the various embodiments, the detailed examples of the embodiments will now be described with reference to the various drawings.
[0027] Throughout the specification, when reference is made to an "application" or a "game" or a "video game" or a "gaming application", it is intended to represent any type of interactive application that is directed through the execution of input commands. For purposes of illustration only, interactive applications include applications for gaming, document processing, video processing, video game processing, and the like. Further, the terms introduced above are interchangeable.
[0028] Throughout the specification, various embodiments of the present disclosure are described with respect to performing multi-GPU processing or rendering of geometry for an application using a typical architecture having four GPUs. However, of course, any number of GPUs (e.g., two or more GPUs) may cooperate when rendering the geometry for an application.
[0029] FIG. 1 is a diagram of a system for performing multi-GPU processing when rendering an image (e.g., an image frame) for an application, according to an embodiment of the present disclosure. According to an embodiment of the present disclosure, the system is configured to provide gaming via a network among one or more cloud gaming servers, and more specifically, is configured to cooperate multiple GPUs to render a single image of an application. Cloud gaming includes executing a video game at a server to generate a game-rendered video frame. This is then sent to and displayed on a client. Specifically, system 100 is configured to perform efficient multi-GPU rendering of the geometry for an application by performing a pre-test on screen regions (which may be interleaved or arranged alternately) prior to rendering.
[0030] FIG. 1 illustrates an embodiment of multi-GPU rendering of geometry among one or more cloud gaming servers of a cloud gaming system. However, in other embodiments of the present disclosure, efficient multi-GPU rendering of geometry for an application is provided by performing a region test during rendering within a stand-alone system (e.g., a personal computer or gaming console having a high-end graphics card with multiple GPUs).
[0031] Naturally, multi-GPU rendering of geometry can also be performed using physical GPUs, virtual GPUs, or a combination of both in various embodiments (e.g., in a cloud gaming environment or within a stand-alone system). For example, a virtual machine (e.g., an instance) can be formed using a hypervisor of host hardware (e.g., located in a data center) that uses one or more components of the hardware layer (e.g., multiple CPUs, memory modules, GPUs, network interfaces, communication components, etc.). These physical resources can be arranged within racks (e.g., a rack of CPUs, a rack of GPUs, a rack of memory, etc.). The physical resources within the rack can be accessed using a top-of-rack switch, which facilitates the structure for assembling and accessing the components to be used for an instance (e.g., when constructing the virtualized components of an instance). Generally, a hypervisor can represent multiple guest operating systems of multiple instances configured using virtual resources. That is, each operating system can be configured using a corresponding set of virtualized resources supported by one or more hardware resources (e.g., located in a corresponding data center). For example, each operating system can be supported by virtual CPUs, multiple virtual GPUs, virtual memory, virtualized communication components, etc. Further, the configuration of an instance can be transferred from one data center to another to reduce latency. When saving a user's gaming session, the GPU utilization rate defined for the user or the game can be used. The GPU utilization rate can include any number of configurations described herein for optimizing the high-speed rendering of video frames for a gaming session. In one embodiment, the GPU utilization rate defined for a game or user can be transferred between data centers as a configurable setting.When a user connects to play a game from different geolocations, the GPU operating rate setting can be transferred, enabling efficient transfer of game play from one data center to another data center.
[0032] According to one embodiment of the present disclosure, system 100 provides gaming via cloud gaming network 190. The game is executed remotely from the corresponding user's client device 110 (e.g., sync client) that is playing the game. System 100 may provide gaming control for one or more users playing one or more games through cloud gaming network 190 via network 150 in either single-player mode or multiplayer mode. In some embodiments, cloud gaming network 190 may include a plurality of virtual machines (VMs) running on a hypervisor of a host machine. One or more virtual machines are configured to execute a game processor module that utilizes the hardware resources available to the host's hypervisor. Network 150 may include one or more communication technologies. In some embodiments, network 150 may include fifth generation (5G) network technology with an advanced wireless communication system.
[0033] In some embodiments, wireless technology may be used to facilitate communication. Such technology may include, for example, 5G wireless communication technology. 5G is the fifth generation of cellular network technology. A 5G network is a digital cellular network, where the service area covered by a provider is divided into small geographical areas called cells. Analog signals representing sound and images are digitized within a telephone and converted by an analog-to-digital converter and transmitted as a stream of bits. All 5G wireless devices within a cell communicate via radio waves using a local antenna array and a low-power automated transceiver (transmitter and receiver) within the cell, and this communication occurs on a frequency channel assigned by the transceiver from a pool of frequencies that are reused in other cells. The local antenna is connected to the telephone network and the Internet via high-bandwidth optical fiber or a wireless backhaul connection. As with other cell networks, when a mobile device crosses from one cell to another, it is automatically transferred to the new cell. Of course, the 5G network is simply an example type of communication network, and in embodiments of the present disclosure, previous generation wireless or wired communication, as well as later generation wired or wireless technology coming after 5G, may be used.
[0034] As shown, cloud game network 190 includes game server 160 that provides access to multiple video games. Game server 160 may be any type of server computing device available within the cloud and may be configured as one or more virtual machines running on one or more hosts. For example, game server 160 may manage a virtual machine that supports a game processor that instantiates an instance of a game for a user. Accordingly, the multiple game processors of game server 160 associated with multiple virtual machines are configured to execute multiple instances of one or more games associated with the game play of multiple users. In this way, backend server support provides streaming of the media (e.g., video, audio, etc.) of the game play of multiple gaming applications to multiple corresponding users. That is, game server 160 is configured to stream data (e.g., the rendered image and / or frame of the corresponding game play) back to the corresponding client device 110 through network 150 by streaming. In this way, a computationally complex gaming application may be executed on the backend server in response to controller input received and forwarded by client device 110. Each server can render images and / or frames, which are then encoded (e.g., compressed) and streamed to the corresponding client device for display.
[0035] For example, multiple users may access a cloud game network 190 via a communication network 150 using corresponding client devices 110 configured to receive streaming media. In one embodiment, the client device 110 may be configured as a thin client to provide interconnection with a backend server (e.g., the cloud game network 190) configured to provide computing capabilities (e.g., including a game title processing engine 111). In another embodiment, the client device 110 may be configured using a game title processing engine and game logic for performing at least some local processing of a video game, and may further be used to receive streaming content generated by a video game executed on the backend server or for other content provided by backend server support. For local processing, the game title processing engine includes basic processor-based functions for executing video games and services associated with the video game. In that case, the game logic may be stored on the local client device 110 and used to execute the video game.
[0036] Each client device 110 may request access to different games from the cloud game network. For example, the cloud game network 190 may execute one or more game logics built on the game title processing engine 111 to be executed using the CPU resources 163 and GPU resources 365 of the game server 160. For example, the game logic 115a may cooperate with the game title processing engine 111 and be executed on the game server 160 for one client, the game logic 115b may cooperate with the game title processing engine 111 and be executed on the game server 160 for a second client, ··· and the game logic 115n may cooperate with the game title processing engine 111 and be executed on the game server 160 for the nth client.
[0037] Specifically, the client device 110 of the corresponding user (not shown) is configured to request access to the game via a communication network 150 (e.g., the Internet), and to render a display image (e.g., an image frame) generated by a video game executed by the game server 160. The encoded image is sent to the client device 110 and is displayed in relation to the corresponding user. For example, the user may interact with an instance of the video game running on the game processor of the game server 160 through the client device 110. More specifically, the instance of the video game is executed by the game title processing engine 111. The corresponding game logic (e.g., executable code) 115 for implementing the video game is stored and accessible through a data store (not shown) and is used to execute the video game. The game title processing engine 111 can support multiple video games using multiple game logics (e.g., gaming applications), and each game logic is selectable by the user.
[0038] For example, the client device 110 is configured to interact with a game title processing engine 111 related to the game play of a corresponding user through, for example, input commands used to drive the game play. Specifically, the client device 110 may receive input from various types of input devices, such as game controllers, tablet computers, keyboards, gestures captured by video cameras, mice, touch pads, and the like. The client device 110 can be any type of computing device having at least a memory and a processor module capable of connecting to the game server 160 via the network 150. The backend game title processing engine 111 is configured to generate a rendered image. The rendered image is sent via the network 150 and displayed on a corresponding display related to the client device 110. For example, through a cloud-based service, the game rendering image may be sent by an instance of a corresponding game (e.g., game logic) being executed on the game execution engine 111 of the game server 160. That is, the client device 110 is configured to receive an encoded image (e.g., encoded from a game rendering image generated through the execution of a video game) and to display the rendered image on the display 11. In one embodiment, the display 11 includes an HMD (e.g., for displaying VR content). In some embodiments, the rendered image may be streamed to a smartphone or tablet wirelessly or wired directly from a cloud-based service or via the client device 110 (e.g., PlayStation (registered trademark) Remote Play).
[0039] In one embodiment, the game server 160 and / or the game title processing engine 111 include basic processor-based functions for executing games and services associated with gaming applications. For example, the game server 160 includes central processing unit (CPU) resources 163 and graphics processing unit (GPU) resources 365 configured to perform processor-based functions (such as 2D or 3D rendering, physical simulation, scripting, audio, animation, graphics processing, lighting, shading, rasterization, ray tracing, shadowing, culling, transformation, artificial intelligence, etc.). Additionally, the CPU and GPU groups may execute services for gaming applications (partially including memory management, multithreading management, quality of service (QoS), bandwidth testing, social networking, social friend management, communication with friends' social networks, communication channels, texting, instant messaging, chat support, etc.). In one embodiment, one or more applications share specific GPU resources. In one embodiment, multiple GPU devices may be combined to perform graphics processing for a single application running on the corresponding CPU.
[0040] In one embodiment, the cloud game network 190 is a distributed game server system and / or architecture. Specifically, the distributed game engine that executes game logic is configured as a corresponding instance of the corresponding game. Generally, the distributed game engine takes each of the functions of the game engine and distributes those functions for execution by a number of processing entities. The individual functions can be further distributed across one or more processing entities. The processing entities may be configured in different configurations (e.g., physical hardware), and / or as virtual components or virtual machines, and / or as virtual containers. Containers are different from virtual machines in that they virtualize an instance of a gaming application running on a virtualized operating system. The processing entities may use and / or rely on the servers of the cloud game network 190 and the underlying hardware (on one or more servers (compute nodes)). The servers may be arranged in one or more racks. The coordination, allocation, and management of the execution of these functions for the various processing entities are performed by the distributed synchronization layer. In this way, the distributed synchronization layer controls the execution of these functions to enable the generation of media (e.g., video frames, audio, etc.) for the gaming application in response to controller input by the player. The distributed synchronization layer can execute these functions efficiently across the distributed processing entities (e.g., through load balancing), distribute and reassemble important game engine components / functions, and enable more efficient processing to be performed.
[0041] FIG. 2 is a diagram of a typical multi-GPU architecture 200 in which a plurality of GPUs cooperate to render a single image of an application according to an embodiment of the present disclosure. Of course, in various embodiments of the present disclosure, many architectures in which a plurality of GPUs cooperate to render a single image are possible, but will not be explicitly described or illustrated. For example, multi-GPU rendering of geometry for an application by performing area testing during rendering may be performed between one or more cloud gaming servers of a cloud gaming system, or within a stand-alone system (e.g., a personal computer or gaming console that includes a high-end graphics card having a plurality of GPUs).
[0042] The multi-GPU architecture 200 includes a CPU 163 and a plurality of GPUs configured to perform multi-GPU rendering of a single image for an application and / or each image within an image sequence for the application. Specifically, the CPU 163 and the GPU resources 365 are configured to perform processor-based functions (e.g., 2D or 3D rendering, physical simulation, scripting, audio, animation, graphics processing, lighting, shading, rasterization, ray tracing, shadowing, culling, transformation, artificial intelligence, etc.) as described above.
[0043] For example, although four GPUs are shown in the GPU resources 365 of the multi-GPU architecture 200, any number of GPUs may be used when rendering an image for an application. Each GPU is connected to a corresponding dedicated memory (e.g., random access memory (RAM)) via a high-speed bus 220. Specifically, GPU-A is connected to memory 210A (e.g., RAM) via bus 220, GPU-B is connected to memory 210B (e.g., RAM) via bus 220, GPU-C is connected to memory 210C (e.g., RAM) via bus 220, and GPU-D is connected to memory 210D (e.g., RAM) via bus 220.
[0044] Furthermore, each GPU is connected to each other via bus 240. Bus 240 may be approximately equal to or slower than bus 220 used for communication between the corresponding GPU and its corresponding memory, depending on the architecture. For example, GPU-A is connected to each of GPU-B, GPU-C, and GPU-D via bus 240. Also, GPU-B is connected to each of GPU-A, GPU-C, and GPU-D via bus 240. In addition, GPU-C is connected to each of GPU-A, GPU-B, and GPU-D via bus 240. Furthermore, GPU-D is connected to each of GPU-A, GPU-B, and GPU-C via bus 240.
[0045] CPU 163 is connected to each of the GPUs via a slower bus 230 (e.g., bus 230 is slower than bus 220 used for communication between the corresponding GPU and its corresponding memory). Specifically, CPU 163 is connected to each of GPU-A, GPU-B, GPU-C, and GPU-D.
[0046] FIG. 3 is a diagram of a graphics processing unit resource 365 configured to perform multi-GPU rendering of geometry on an image frame generated by an application by performing a pre-test on a screen region (which may be interleaved) prior to rendering, according to one embodiment of the present disclosure. For example, the game server 160 may be configured to include the GPU resources 365 within the cloud game network 190 of FIG. 1. As shown, the GPU resources 365 include a plurality of GPUs (e.g., GPU 365a, GPU 365b ··· GPU 365n). As described above, various architectures may include multiple GPUs cooperating to render a single image by performing multi-GPU rendering of geometry on an application through region testing during rendering. For example, performing multi-GPU rendering of geometry between one or more cloud gaming servers of a cloud gaming system, or performing multi-GPU rendering of geometry within a stand-alone system (e.g., a personal computer or gaming console that includes a high-end graphics card having multiple GPUs).
[0047] Specifically, in one embodiment, the game server 160 is configured to perform multi-GPU processing when rendering a single image of an application, where multiple GPUs cooperate to render a single image and / or render each of one or more images in a sequence of images when executing the application. For example, in one embodiment, the game server 160 may include a CPU and a GPU group configured to perform multi-GPU rendering of each of one or more images within a sequence of images of an application. Here, one CPU and GPU group can execute graphics and / or render a pipeline for an application. The CPU and GPU group can be configured as one or more processing devices. As described above, the GPU and GPU group may include the CPU 163 and the GPU resource 365. The CPU 163 and the GPU resource 365 are configured to perform processor-based functions (e.g., 2D or 3D rendering, physical simulation, scripting, audio, animation, graphics processing, lighting, shading, rasterization, ray tracing, shadowing, culling, transformation, artificial intelligence, etc.).
[0048] The GPU resource 365 is responsible for and / or configured to perform object rendering (e.g., writing color or normal vector values for object pixels to multiple rendering targets - MRT) and execution of synchronous compute kernels (e.g., full-screen effects on the resulting MRT). The synchronous computations to be executed and the objects to be rendered are specified by commands included in the rendering command buffer 325 that the GPU executes. Specifically, the GPU resource 365 is configured to render an object and perform synchronous computations (e.g., while executing a synchronous compute kernel) when executing commands from the rendering command buffer 325, and the commands and / or operations may depend on other operations such that they are performed sequentially.
[0049] For example, the GPU resource 365 is configured to perform synchronous calculations and / or to render objects using one or more rendering command buffers 325 (e.g., rendering command buffers 325a, 325b... 325n). In one embodiment, each GPU in the GPU resource 365 may have its own command buffer. Alternatively, when substantially the same set of objects is being rendered by each GPU (e.g., because the size of the area is small), the GPUs in the GPU resource 365 may use the same command buffer or the same set of command buffers. Further, each of the GPUs in the GPU resource 365 may support that a command can be executed by one GPU but not by another. For example, if there is a flag on a draw command or predication within the rendering command buffer, a single GPU can execute one or more commands within the corresponding command buffer, while other GPUs ignore the commands. For example, the rendering command buffer 325a may support a flag 330a, the rendering command buffer 325b may support a flag 330b... the rendering command buffer 325n may support a flag 330n.
[0050] Performing synchronous calculations (e.g., executing a synchronous calculation kernel) and rendering of objects are part of the overall rendering. For example, when a video game is running at 60 Hz (e.g., 60 frames per second), all object rendering and execution of the synchronous calculation kernel for an image frame typically must be completed within approximately 16.67 ms (e.g., 1 frame at 60 Hz). As described above, the operations performed when rendering an object and / or executing a synchronous calculation kernel are ordered, and the operations may depend on other operations (e.g., a command in a rendering command buffer may need to complete execution before other commands in that rendering command buffer can be executed).
[0051] Specifically, each of the rendering command buffers 325 contains various types of commands (e.g., commands that affect the corresponding GPU configuration (e.g., commands that specify the location and format of a rendering target), as well as commands to render an object and / or execute a synchronous calculation kernel). For purposes of illustration, the synchronous calculations performed when executing a synchronous calculation kernel may include performing a full-screen effect when all objects are rendered to one or more corresponding multiple render targets (MRTs).
[0052] In addition, when the GPU resource 365 renders an object for an image frame and / or executes a synchronous computing kernel when generating an image frame, the GPU resource 365 is configured via the registers of each GPU 365a, 365b... 365n. For example, GPU 365a is configured to perform its rendering or calculate kernel execution in a specific manner via its registers 340 (e.g., registers 340a, 340b... 340n). That is, the values stored in the registers 340 define the hardware context (e.g., GPU configuration or GPU state) for GPU 365a when executing the commands in the rendering command buffer 325 used to render an object for an image frame and / or execute a synchronous computing kernel. Each of the GPUs in the GPU resource 365 can be configured similarly, such that GPU 365b is configured to perform its rendering or calculate kernel execution in a specific manner via its registers 350 (e.g., registers 350a, 350b... 350n),... and GPU 365n is configured to perform its rendering or calculate kernel execution in a specific manner via its registers 370 (e.g., registers 370a, 370b... 370n).
[0053] Some examples of GPU configurations include the location and format of the rendering target (e.g., MRT). Other examples of GPU configurations include the operation procedure. For example, when rendering an object, the Z value of each pixel of the object can be compared with the Z buffer in various ways. For example, an object pixel can be written only when the object Z value matches the value in the Z buffer. Alternatively, an object pixel can be written only when the object Z value is the same as or less than the value in the Z buffer. The type of test to be performed is defined within the GPU configuration.
[0054] FIG. 4 is a schematic diagram of a rendering architecture implementing a graphics pipeline 400 configured for multi-GPU processing such that a plurality of GPUs cooperate to render a single image, according to an embodiment of the present disclosure. The graphics pipeline 400 illustrates a general process for rendering an image using a 3D (three-dimensional) polygon rendering process. The graphics pipeline 400 outputs corresponding color information for each pixel in the display for the rendered image. The color information may represent textures and shading (e.g., color, shadowing, etc.). The graphics pipeline 400 may be implemented within the client device 110, game server 160, game title processing engine 111, and / or GPU resource 365 of FIGS. 1 and 3. That is, various architectures may include multiple GPUs cooperating to render a single image by performing multi-GPU rendering of geometry for an application through an area test during rendering. For example, performing multi-GPU rendering of geometry among one or more cloud gaming servers of a cloud gaming system, or performing multi-GPU rendering of geometry within a stand-alone system (e.g., a personal computer or gaming console including a high-end graphics card having multiple GPUs).
[0055] As shown, the graphics pipeline receives an input geometry 405. For example, a geometry processing stage 410 receives the input geometry 405. For example, the input geometry 405 may include vertices within a 3D gaming world and information corresponding to each vertex. A given object within the gaming world can be represented using polygons (e.g., triangles) defined by vertices. Next, the surface of the corresponding polygon is processed through the graphics pipeline 400 to achieve a final effect (e.g., color, texture, etc.). Vertex attributes may include normals (e.g., which direction is perpendicular to the geometry at that location), colors (e.g., RGB - three colors of red, green, and blue, etc.), and texture coordinates / mapping information.
[0056] The geometry processing stage 410 is responsible for (and can perform) both vertex processing (e.g., via a vertex shader) and primitive processing. Specifically, the geometry processing stage 410 may output a set of vertices that define a primitive and send it to the next stage of the graphics pipeline 400, as well as the positions (specifically, homogeneous coordinates) and various other parameters for those vertices. The positions are placed in a position cache 450 for access by later shader stages. The other parameters are also placed in a parameter cache 460 for access by later shader stages.
[0057] Various operations may be performed by the geometry processing stage 410. For example, performing lighting and shadowing calculations on primitives and / or polygons. In one embodiment, since the geometry stage can process primitives, it can perform backface culling and / or clipping (e.g., tests against a frustum), resulting in a reduced load on downstream stages (e.g., a rasterization stage 420, etc.). In another embodiment, the geometry stage may generate primitives (e.g., through a function equivalent to a conventional geometry shader).
[0058] The primitive output by the geometry processing stage 410 is supplied into the rasterization stage 420, where the primitive is converted into a raster image consisting of pixels. Specifically, the rasterization stage 420 is configured to project the objects in the scene onto a two-dimensional (2D) image plane defined by a viewing location (e.g., camera location, user's eye location, etc.) within the 3D gaming world. At a simplified level, the rasterization stage 420 looks at each primitive and determines which pixels are affected by the corresponding primitive. Specifically, the rasterizer 420 divides the primitive into pixel-sized fragments. Each fragment corresponds to a pixel in the display. It is important to note that when displaying the image, one or more fragments can contribute to the color of the corresponding pixel.
[0059] As described above, further operations may be performed by the rasterization stage 420. For example, clipping with respect to the viewing location (identifying and ignoring fragments outside the frustum) and culling (ignoring fragments hidden by closer objects). In connection with clipping, the geometry processing stage 410 and / or the rasterization stage 420 may be configured to identify and ignore primitives outside the frustum defined by the viewing location within the gaming world.
[0060] The pixel processing stage 430 may generate values such as the color resulting from the pixel using the parameters (as well as other data) formed by the geometry processing stage. Specifically, the pixel processing stage 430 fundamentally performs a shading operation on the fragment to determine how the color and luminance of the primitive differ depending on the available lighting. For example, the pixel processing stage 430 may determine the depth, color, normal, and texture coordinates (e.g., texture details) for each fragment, and may further determine the appropriate levels of light, darkness, and color for the fragment. Specifically, the pixel processing stage 430 calculates the characteristics of each fragment. For example, color and other attributes (e.g., z-depth relative to the viewing location and alpha value for transparency). In addition, the pixel processing stage 430 applies a lighting effect to the fragment based on the available lighting that affects the corresponding fragment. Further, the pixel processing stage 430 may apply a shadowing effect to each fragment.
[0061] The output of the pixel processing stage 430 includes the processed fragments (e.g., texture and shading information) and is sent to the 440 output merger stage, which is the next stage in the graphics pipeline 400. The output merger stage 440 generates the final color for the pixel using the output of the pixel processing stage 430, as well as other data (e.g., values already in memory). For example, the output merger stage 440 may perform optional blending of values between the fragment and / or pixel determined from the pixel processing stage 430 and the value already written to the MRT for that pixel.
[0062] Color values for each pixel in the display may be stored in a frame buffer (not shown). These values are scanned to the corresponding pixels when displaying the corresponding image of the scene. Specifically, the display reads the color values from the frame buffer for each pixel, row by row, from left to right or right to left, from top to bottom or bottom to top, or in any other pattern, and illuminates the pixels that use these pixel values when displaying the image.
[0063] With the detailed description of the cloud game network 190 (e.g., within the game server 160) and the GPU resources 365 in FIGS. 1-3, FIG. 5's flowchart 500, according to one embodiment of the present disclosure, performs a pre-test of the geometry on the interleaved screen area before rendering, to illustrate a method for performing graphics processing when performing multi-GPU rendering of the geometry for the image frames generated by the application. In this way, multiple GPU resources are used to efficiently perform the rendering of objects when running an application. As described above, various architectures may include multiple GPUs cooperating to render a single image by performing multi-GPU rendering of the geometry for the application through area tests during rendering. Rendering may be performed, for example, within one or more cloud game servers of a cloud gaming system, or within a stand-alone system (e.g., a personal computer or a gaming console including a high-end graphics card having multiple GPUs), etc.
[0064] At 510, the method includes rendering the graphics for an application using a plurality of graphics processing units (GPUs) that cooperate to generate an image. Specifically, multi-GPU processing is performed when rendering each of one or more image frames of a single image frame and / or a sequence of image frames for a real-time application.
[0065] At 520, the method includes dividing the responsibility of rendering the graphics geometry among multiple GPUs based on multiple screen regions. That is, each GPU has a corresponding division or portion of the responsibility (e.g., the corresponding screen region) known to all GPUs. More specifically, each GPU has the responsibility of rendering the geometry within the corresponding set of screen regions out of the multiple screen regions. The corresponding set of screen regions includes one or more screen regions. For example, the first GPU has a first division of the responsibility for rendering the objects within the first set of screen regions. Also, the second GPU has a second division of the responsibility for rendering the objects within the second set of screen regions. This is repeatedly applied to the remaining GPUs.
[0066] At 530, the method includes allocating, for the geometry test, to the first GPU a first piece or fragment of the geometry of an image frame generated during the execution of the application. For example, the image frame may include one or more objects. Each object may be defined by one or more pieces of geometry. That is, in one embodiment, the geometry pre - test and rendering are performed on pieces of geometry that are the entire object. In other embodiments, the geometry pre - test and rendering are performed on pieces of geometry that are part of the entire object.
[0067] For example, allocate each of a plurality of GPUs to corresponding portions of the geometry associated with an image frame. Specifically, for the purpose of geometry pre-testing, allocate all portions of the geometry to the corresponding GPUs. In one embodiment, the geometry may be evenly allocated among the plurality of GPUs. For example, if there are four GPUs in the plurality, each GPU may process one quarter of the geometry within the image frame. In other embodiments, the geometry may be unevenly allocated among the plurality of GPUs. For example, in an example using four GPUs for multi-GPU rendering of an image frame, the geometry of the image frame processed by one GPU may be more than that of another GPU.
[0068] At 540, the method includes performing a geometry pre-test at a first GPU to generate information on how pieces of the geometry relate to a plurality of screen regions. Specifically, the first GPU generates information on a piece of the geometry and how it relates to each of the plurality of screen regions. For example, the geometry pre-test by the first GPU may determine whether a piece of the geometry overlaps a particular screen region assigned to a corresponding GPU for object rendering. The first piece of the geometry may overlap a screen region for which another GPU has the responsibility of performing object rendering and / or a screen region for which the first GPU has the responsibility of performing object rendering. In one embodiment, before any of the plurality of GPUs performs rendering of the geometry, a shader in a corresponding command buffer performed by the first GPU performs a geometry test. In other embodiments, the geometry test is performed by hardware, for example, at a rasterization stage 420 of a graphics pipeline 400.
[0069] In embodiments, the geometry pre-test is typically performed simultaneously by a plurality of GPUs on all the geometry of the corresponding image frame. That is, each GPU performs a geometry pre-test on that portion of the geometry of the corresponding image frame. In this way, by performing the geometry pre-test, each GPU can know which piece of the geometry to render and which piece of the geometry to skip. Specifically, when the corresponding GPU performs the geometry pre-test, the corresponding GPU tests that portion of the geometry against the screen regions of each of the plurality of GPUs used to render the image frame. For example, if there are four GPUs, and especially for the purpose of the geometry test, if the geometry is evenly assigned to the GPUs, each GPU may perform the geometry test on one-fourth of the geometry of the image frame. Thus, even though each GPU performs the geometry pre-test only on that portion of the geometry of the corresponding image frame, in embodiments, the geometry pre-test is typically performed simultaneously on all the geometry of the image frame across the plurality of GPUs, so the information generated indicates how all the geometry (e.g., pieces of the geometry) in the image frame relates to the screen regions of all the GPUs. Each screen region is assigned to the corresponding GPU for object rendering, and / or the rendering may be performed on a piece of the geometry (e.g., the whole object or a part of the object).
[0070] At 550, the method includes using information at each of a plurality of GPUs when rendering a piece of geometry (e.g., to include fully rendering the piece of geometry or skipping rendering of the piece of geometry). That is, a piece of geometry is rendered using information at each of the plurality of GPUs. Geometry test results (e.g., information) are sent to other GPUs so that the information is known to each GPU. For example, geometry (e.g., a piece of geometry) within an image frame is typically rendered simultaneously by a plurality of GPUs in an embodiment. Specifically, when a piece of geometry overlaps any screen region assigned to the corresponding GPU for object rendering, the GPU renders that piece of geometry based on the information. On the other hand, when a piece of geometry does not overlap any screen region assigned to the corresponding GPU for object rendering, the GPU can skip rendering of that piece of geometry based on the information. Thus, with the information, all GPUs can render the geometry within the image frame more efficiently and / or completely avoid rendering of the geometry. For example, rendering may be performed by a shader within a corresponding command buffer to be executed by a plurality of GPUs. As will be more fully described below in FIGS. 7A, 12A, and 13A, the shader may be configured to perform one or both of geometry testing and / or rendering based on the corresponding GPU configuration.
[0071] According to one embodiment of the present disclosure, in some architectures, if the corresponding rendering GPU receives the corresponding information in time for it to be used, the GPU uses that information when determining which geometry in the corresponding image should be rendered. That is, the information can be taken as a hint. Otherwise, the rendering GPU processes the piece of geometry as it normally would. Using an example where the information can indicate whether the geometry overlaps any screen region assigned to the rendering GPU (e.g., a second GPU), if the information indicates that there is no overlap of the geometry, the rendering GPU may completely skip rendering the geometry. Also, if only a piece of the geometry does not overlap, the second GPU may skip rendering the piece of the geometry that does not overlap any of the screen regions assigned to the second GPU for object rendering. On the other hand, the information may indicate that there is an overlap for the geometry, in which case the second or rendering GPU renders the geometry. Also, the information may indicate that a particular piece of the geometry overlaps any screen region assigned to the second or rendering GPU for object rendering. In that case, the second or rendering GPU renders only the piece of the geometry that overlaps. In yet other embodiments, if there is no information, or if the generation or receipt of the information is not in time, the second GPU performs the rendering as normal (e.g., renders the geometry). Thus, the information provided as a hint can increase the overall efficiency of the graphics processing system if received in time. If the information is not received in time, the graphics processing system still operates properly even without such information.
[0072] In one embodiment, a certain GPU (e.g., a pre-test GPU) is dedicated to performing geometry pre-tests to generate information. That is, the dedicated GPU is not used for rendering objects (e.g., pieces of geometry) within the corresponding image frame. Specifically, as described above, graphics for an application are rendered using multiple GPUs. The responsibility for rendering the geometry of the graphics is divided among multiple GPUs based on multiple screen regions (which may be interleaved). Each GPU has a corresponding division of the responsibilities known to the multiple GPUs. The geometry test is performed in the pre-test GPU on multiple pieces of the geometry of the image frame generated by the application to generate information regarding the relationship of each piece of the geometry to each of the multiple screen regions. The multiple pieces of the geometry are rendered in each of the multiple GPUs using the information generated for each respective piece of the geometry. That is, the information is used when rendering each piece of the geometry by the corresponding rendering GPU from the GPUs used to render the image frame.
[0073] Figures 6A - 6B show, for purely illustrative purposes, rendering for a screen divided into regions and sub-regions. Of course, the number of regions and sub-regions to be divided can be selected for efficient multi-GPU processing of one or more each of the images and / or image sequences. That is, the screen may be divided into two or more regions, and each region may be further divided into sub-regions. In one embodiment of the present disclosure, as shown in FIG. 6A, the screen is divided into four quadrants. In another embodiment of the present disclosure, as shown in FIG. 6B, the screen is divided into a larger number of interleaved regions. The following description of FIGS. 6A - 6B is intended to illustrate the inefficiencies that occur when performing multi-GPU rendering on multiple screen regions assigned to multiple GPUs. FIGS. 7A - 7C and FIGS. 8A - 8B show more efficient rendering according to some embodiments of the present invention.
[0074] Specifically, FIG. 6A is a diagram of a screen 610A that is divided into quadrants (e.g., four regions) when performing multi-GPU rendering. As shown, screen 610A is divided into four quadrants (e.g., A, B, C, and D). Each quadrant is assigned to one of four GPUs [GPU-A, GPU-B, GPU-C, and GPU-D] in a one-to-one relationship. For example, GPU-A is assigned to quadrant A, GPU-B is assigned to quadrant B, GPU-C is assigned to quadrant C, and GPU-D is assigned to quadrant D.
[0075] Geometry can be culled. For example, CPU 163 can check the bounding box for the frustum of each quadrant and request each GPU to render only the objects that overlap its corresponding frustum. As a result, each GPU is responsible for rendering only a portion of the geometry. For illustrative purposes, screen 610 shows pieces of geometry, each piece being a corresponding object, and screen 610 shows objects 611 - 617 (e.g., pieces of geometry). Since there are no objects that overlap quadrant A, GPU-A does not render any objects. GPU-B renders objects 615 and 616 (since a portion of object 615 is within quadrant B, the CPU's culling test correctly concludes that GPU-B must render it). GPU-C renders objects 611 and 612. GPU-D renders objects 612, 613, 614, 615, and 617.
[0076] In FIG. 6A, when the screen 610A is divided into quadrants A - D, the amount of work that each GPU has to perform can be very different. This is because, in some cases, an unequal amount of geometry can be in one quadrant. For example, there are no pieces of geometry in quadrant A, but there are 5 pieces of geometry or at least a portion of at least 5 pieces of geometry in quadrant D. Thus, GPU - A assigned to quadrant A is not used, while GPU - D assigned to quadrant D is unfairly busy when rendering objects in the corresponding image.
[0077] FIG. 6B illustrates another approach when subdividing the screen into regions. Specifically, when performing multi - GPU rendering of each image within a single image or image sequence, instead of dividing into quadrants, the screen 610B is divided into a plurality of interleaved regions. In that case, the screen 610B is divided into a larger number of interleaved regions (e.g., more than 4 quadrants), while using the same number of GPUs for rendering (e.g., 4). The objects (611 - 617) shown in screen 610A are also shown in the same corresponding locations of screen 610B.
[0078] Specifically, 4 GPUs (e.g., GPU - A, GPU - B, GPU - C, and GPU - D) are used to render an image for the corresponding application. Each GPU is responsible for rendering the geometry that overlaps with the corresponding region. That is, each GPU is assigned to a corresponding set of regions. For example, GPU - A is responsible for each region labeled A in the corresponding set, GPU - B is responsible for each region labeled B in the corresponding set, GPU - C is responsible for each region labeled C in the corresponding set, and GPU - D is responsible for each region labeled D in the corresponding set.
[0079] Furthermore, the regions are interleaved in a specific pattern. To interleave the regions (and have a greater number of regions), the amount of work that each GPU has to perform can be much more balanced. For example, a pattern for interleaving screen 610B includes alternating rows (e.g., regions A - B - A - B, etc., and regions C - D - C - D, etc.). In embodiments of the present disclosure, other patterns for interleaving regions are also supported. For example, the pattern may include regions of repeating sequences, regions of uniform distribution, regions of non - uniform distribution, regions of arrays of repeatable rows, regions of random sequences, regions of arrays of random rows, etc.
[0080] It is important to choose the number of regions. For example, if the region distribution is too fine - grained (e.g., the number of regions is too large and not optimal), each GPU still has to process most or all of the geometry. For example, it may be difficult for a GPU to check the bounding boxes of objects for all regions for which it has responsibility. Also, even if the bounding boxes can be checked in a timely manner, due to the small region size, as a result, each GPU may have to process most of the geometry. This is because all objects in the image overlap at least one region of each set of regions assigned to each GPU (e.g., a GPU processes the entire object even if only a part of the object overlaps at least one region within the set of regions assigned to that GPU).
[0081] As a result, it is important to select the number of regions, the interleaving pattern, etc. Selecting too few or too many regions, or selecting too few or too many regions for interleaving, or selecting an inefficient pattern for interleaving can lead to inefficiencies when performing GPU processing (e.g., each GPU processes most or all of the geometry). In such cases, even when there are multiple GPUs for rendering an image, due to the inefficiency of the GPUs, it is not possible to support the corresponding increase in both the number of screen pixels and the geometry density (i.e., four GPUs cannot write four times as many pixels and process four times as many vertices or primitives). In the following embodiments, improvements are targeted, among other things, at culling strategies (Figs. 7A - 7C) and culling granularity (Figs. 8A - 8B).
[0082] Figs. 7A - 7C are diagrams illustrating, in embodiments of the present disclosure, rendering at least one or more respective images within a single image and / or an image sequence using multiple GPUs. The selection of four GPUs is merely for the purpose of simply illustrating multi - GPU rendering when rendering an image while running an application, and of course, any number of GPUs may be used for multi - GPU rendering in various embodiments.
[0083] Specifically, FIG. 7A is a diagram of a rendering command buffer 700A shared by a plurality of GPUs that cooperate to render a single image frame according to an embodiment of the present disclosure. That is, in this embodiment, each of the plurality of GPUs uses the same rendering command buffer (e.g., buffer 700A), and each GPU executes all commands in the rendering command buffer. A plurality of commands (a complete set) are loaded into the rendering command buffer 700A and used to render the corresponding image frame. Of course, one or more rendering command buffers may be used to generate the corresponding image frame. In one example, the CPU generates one or more draw calls for the image frame. The draw calls include commands placed in one or more rendering command buffers to be executed by one or more of the GPU resources 365 of FIG. 3 when performing multi-GPU rendering of the corresponding image. In some embodiments, the CPU 163 may request one or more GPUs to generate all or part of the draw calls used to render the corresponding image. Further, FIG. 7A may show the entire set of commands included in the rendering command buffer 700A, or FIG. 7A may show a part of the entire set of commands included in the rendering command buffer 700A.
[0084] In an embodiment, a GPU typically renders images simultaneously when performing multi-GPU rendering of one or more images within an image or image sequence. Rendering of an image can be decomposed into multiple phases. In each phase, the GPUs need to be synchronized, and the faster GPUs have to wait until the slower GPUs are complete. The commands shown in FIG. 7A for the rendering command buffer 700A represent one phase. In FIG. 7A, only the commands for one phase are shown, but the rendering command buffer 700A may contain commands for one or more phases when rendering an image. In FIG. 7A, only a portion of all the commands are shown, and the commands for other phases are not shown. In the piece of the rendering command buffer 700A shown in FIG. 7A that illustrates one phase, there are four objects to be rendered (e.g., object 0, object 1, object 2, and object 3). This is shown in FIG. 7B-1.
[0085] As shown, the piece of the rendering command buffer 700A shown in FIG. 7A includes commands for a geometry test, commands for rendering an object (e.g., a piece of geometry), and commands for configuring the state of one or more rendering GPUs that execute the commands from the rendering command buffer 700A. For illustrative purposes only, the piece of the rendering command buffer 700A shown in FIG. 7A includes commands (710-728) for a geometry pre-test, rendering of an object, and / or execution of a synchronous compute kernel when rendering the corresponding image for the corresponding application. In some embodiments, the geometry pre-test, and rendering of the object for that image, and / or execution of the synchronous compute kernel must be performed within a frame period. Two processing sections are shown within the rendering command buffer 700A. Specifically, processing section 1 includes a pre-test or geometry test 701, and section 2 includes rendering 702.
[0086] Section 1 includes performing a geometry test 701 on an object within an image frame. Each object can be defined by one or more pieces of geometry. The pre-test or geometry test 701 can be performed by one or more shaders. For example, in multi-GPU rendering of a corresponding image frame, each GPU used is assigned a portion of the geometry of the image frame to execute the geometry test. In one embodiment, all portions may be assigned for the pre-test. The assigned portions may include one or more pieces of geometry. Each piece may include the entire object or a part of the object (e.g., a vertex, a primitive, etc.). Specifically, the geometry test is performed on the pieces of geometry to generate information about how that piece of geometry relates to each of a plurality of screen regions. For example, the geometry test may determine whether a piece of geometry overlaps a particular screen region assigned to the corresponding GPU for object rendering.
[0087] As shown in FIG. 7A, the geometry test 701 (e.g., pre-geometry test) of section 1 includes commands for configuring the state of one or more GPUs to execute commands from the rendering command buffer 700A, and commands for performing the geometry test. Specifically, the GPU state of each GPU is configured before the GPU executes the geometry test on the corresponding object. For example, commands 710, 713, and 715 are each used to configure the GPU state of one or more GPUs for the purpose of executing commands for the geometry test. As shown, command 710 configures the GPU state so that the geometry test commands 711-712 can be properly executed. Command 711 executes a geometry test on object 0, and command 712 executes a geometry test on object 1. Similarly, command 713 configures the GPU state so that the geometry test command 714 can execute a geometry test on object 2. Also, command 715 configures the GPU state so that the geometry test command 716 can execute a geometry test on object 3. Naturally, the GPU state may be configured for one or more geometry test commands (e.g., test commands 711 and 712).
[0088] As described above, when executing commands within the rendering command buffer 700A used for geometry testing and / or rendering of objects and / or execution of synchronous compute kernels on corresponding images, the values stored in the registers define the hardware context (e.g., GPU configuration) for the corresponding GPU. As shown, the GPU state may be changed throughout the processing of the commands within the rendering command buffer 700A. Each subsequent section of the commands may be used to configure the GPU state. As applied to FIG. 7A, and throughout the specification, when referring to the setting of the GPU state, the GPU state may be set in various ways. For example, the CPU or GPU can set values in random access memory (RAM). The GPU checks the values in the RAM. In another example, the state may be internal to the GPU, which is the case, for example, when the command buffer is called twice as a subroutine and the internal GPU state is different between the two subroutine calls.
[0089] Section 2 includes performing rendering 702 of an object within an image frame. (A piece of geometry is rendered.) Rendering 702 can be performed by one or more shaders within command buffer 700A. As shown in FIG. 7A, rendering 702 of Section 2 includes commands for configuring the state of one or more GPUs to execute commands from rendering command buffer 700A and commands for performing the rendering. Specifically, the GPU state of each GPU is configured before the corresponding object (e.g., a piece of geometry) is rendered. For example, commands 721, 723, 725, and 727 are each used to configure the GPU state of one or more GPUs for the purpose of executing commands for rendering. As shown, command 721 configures the GPU state so that rendering command 722 can render object 0. Command 723 configures the GPU state so that rendering command 724 can render object 1. Command 725 configures the GPU state so that rendering command 726 can render object 2. Also, command 727 configures the GPU state so that rendering command 728 can render object 3. In FIG. 7A, it is shown that the GPU state is configured for each rendering command (e.g., rendering object 0, etc.), but of course, the GPU state may be configured for one or more rendering commands.
[0090] As described above, each GPU used in the multi-GPU rendering of corresponding image frames renders the corresponding piece of geometry based on the information generated during the geometry pre-test. Specifically, the information known to each GPU provides the relationship between the object and the screen region. When rendering the corresponding piece of geometry, the GPU can use that information when received in a timely manner for the purpose of efficiently rendering those pieces of geometry. Specifically, as the information indicates, when a piece of geometry overlaps any screen region or regions (plural) assigned to the corresponding GPU for object rendering, the GPU performs rendering on that piece of geometry. On the other hand, the information may indicate that the first GPU has to completely skip rendering a piece of geometry (for example, if the piece of geometry does not overlap any screen region for which the first GPU is assigned the responsibility of object rendering). Thus, each GPU renders only the pieces of geometry that overlap the screen region or regions (plural) for which it has the responsibility of performing object rendering. Therefore, the information is provided as a hint to each GPU and is considered by each GPU performing rendering of a piece of geometry if received before rendering begins. In one embodiment, rendering proceeds normally if the information is not received in time. For example, regardless of whether the corresponding piece of geometry overlaps any screen region assigned to the GPU for object rendering, that piece of geometry is completely rendered by the corresponding GPU.
[0091] For illustrative purposes only, four GPUs divide the corresponding screens into areas between them. As described above, each GPU has the responsibility of rendering objects in the corresponding set of areas. The corresponding set includes one or more areas. In one embodiment, the rendering command buffer 700A is shared by multiple GPUs that cooperate to render a single image. That is, the GPUs used for multi-GPU rendering of a single image or each of one or more images within an image sequence share a common command buffer. In another embodiment, each GPU may have its own command buffer.
[0092] Alternatively, in still other embodiments, the GPUs may each render somewhat different sets of objects. This can hold true when it can be determined that a particular GPU does not need to render a particular object, for example because it does not overlap its corresponding screen area in the corresponding set. As described above, as long as the command buffer supports that a command can be executed by one GPU but not by another, multiple GPUs can still use the same command buffer (e.g., share one command buffer). For example, the execution of commands within the shared rendering command buffer 700A may be limited to one of the rendering GPUs. This can be achieved in various ways. In another example, a flag may be used on the corresponding command to indicate which GPU should execute it. Also, predicates may be executed within the rendering command buffer using bits that indicate what each GPU does under what conditions. An example of a predicate is "if this is GPU-A, skip the next X commands".
[0093] In yet another embodiment, since substantially the same set of objects is being rendered by each GPU, multiple GPUs can still use the same command buffer. For example, as described above, when the area is relatively small, each GPU may render all of the objects.
[0094] Screen 700B is illustrated in FIG. 7B-1. Screen 700B shows an image including four objects that are rendered by multiple GPUs using the rendering command buffer 700A of FIG. 7A, according to one embodiment of the present disclosure. According to one embodiment of the present disclosure, multi-GPU rendering of geometry is performed on an application by pre-testing the geometry against screen regions (which may be interleaved) before rendering pieces of geometry corresponding to objects within an image frame.
[0095] Specifically, the responsibility for rendering the geometry is divided by screen regions among multiple GPUs. The multiple screen regions are configured to reduce the imbalance in rendering time among the multiple GPUs. For example, Screen 700B shows the screen region responsibility for each GPU when rendering the objects of the image. Four GPUs (GPU-A, GPU-B, GPU-C, and GPU-D) are used to render the objects within the image shown in Screen 700B. To balance the pixel and vertex loads among the GPUs, Screen 700B is divided more finely than the quadrants shown in FIG. 6A. In addition, Screen 700B is divided into regions that may be interleaved. For example, the interleaving includes multiple rows of regions. In rows 731 and 733, region A appears alternately with region B. In rows 732 and 734, region C appears alternately with region D. More specifically, within the pattern, the rows including regions A and B appear alternately with the rows including regions C and D.
[0096] As described above, various techniques may be used when dividing the screen into regions to achieve GPU processing efficiency. For example, increasing or decreasing the number of regions (e.g., to select the exact amount of regions), interleaving the regions, increasing or decreasing the number of regions to select a specific pattern when interleaving regions and / or sub-regions, and so on. In one embodiment, the plurality of screen regions are each of a uniform size. In one embodiment, the plurality of screen regions are each of a non-uniform size. In yet other embodiments, the number and sizing of the plurality of screen regions vary dynamically.
[0097] Each GPU has the responsibility for rendering the objects within the corresponding set of regions. Each set may include one or more regions. Thus, GPU-A has the responsibility for rendering the objects within each A-region in the corresponding set, GPU-B has the responsibility for rendering the objects within each B-region in the corresponding set, GPU-C has the responsibility for rendering the objects within each C-region in the corresponding set, and GPU-D has the responsibility for rendering the objects within each D-region in the corresponding set. There may be other GPUs with other responsibilities, and they may not perform rendering (e.g., performing asynchronous type calculation kernels executed over multiple frame periods, performing culling on the rendering GPU, etc.).
[0098] The amount of rendering to be performed varies for each GPU. Figure 7B-2 illustrates, according to one embodiment of the present disclosure, a table showing the rendering performed by each GPU when rendering the four objects of Figure 7B-1. As shown in the table, after the geometry pre-test, it may be determined that object 0 is being rendered by GPU-B, object 1 is being rendered by GPU-C and GPU-D, object 2 is being rendered by GPU-A, GPU-B, and GPU-D, and object 3 is being rendered by GPU-B, GPU-C, and GPU-D. Since GPU A only needs to render object 2 and GPU D needs to render objects 1, 2, and 3, there may still be some unbalanced rendering. However, overall, through the interleaving of screen regions, the rendering of objects in the image is reasonably balanced among the multiple GPUs used for multi-GPU rendering of the image or rendering of each of one or more images within the image sequence.
[0099] Figure 7C is a diagram illustrating, according to one embodiment of the present disclosure, the rendering of each object performed by each GPU when multiple GPUs cooperate to render a single image frame (for example, the image frame 700B shown in Figure 7B-1). Specifically, Figure 7C shows the rendering processes of objects 0 to 3 performed by each of the four GPUs (for example, GPU-A, GPU-B, GPU-C, and GPU-D) using the shared rendering command buffer 700A of Figure 7A.
[0100] Specifically, two rendering timing diagrams are shown with respect to the time axis 740. The rendering timing diagram 700C-1 shows the multi-GPU rendering of objects 0 to 3 of the corresponding image in one phase of rendering. Each GPU performs rendering when there is no information regarding the overlap between objects 0 to 3 and the screen area. The rendering timing diagram 700C-2 shows the multi-GPU rendering of objects 0 to 3 of the corresponding image in the same phase of rendering. Information generated during the geometry test of the screen area (e.g., performed before rendering) is shared by each GPU and used to render objects 0 to 3 through the corresponding GPU pipeline. The rendering timing diagrams 700C-1 and 700C-2 each show the time required for each GPU to process each piece of geometry (e.g., perform geometry tests and rendering). In one embodiment, a piece of geometry is the entire object. In another embodiment, a piece of geometry may be a part of an object. For illustrative purposes, the example of FIG. 7C shows the rendering of a piece of geometry. Each piece of geometry corresponds to an object (e.g., in its entirety). In each of the rendering timing diagrams 700C-1 and 700C-2, an object (e.g., a piece of geometry) that has no geometry overlapping at least one screen area of the corresponding GPU (e.g., within the corresponding set of areas) is represented by a box drawn with a dashed line. On the other hand, an object having geometry overlapping at least one screen area of the corresponding GPU (e.g., within the corresponding set of areas) is represented by a box drawn with a solid line.
[0101] Rendering timing diagram 700C-1 shows the rendering of objects 0 to 3 using four GPUs (for example, GPU-A, GPU-B, GPU-C, and GPU-D). In rendering timing diagram 700C-1, vertical line 755a indicates the start of the rendering phase for an object, and vertical line 755b indicates the end of the rendering phase for an object. The start and end points along time axis 740 for the illustrated rendering phase represent synchronization points. The four GPUs are synchronized when each executes its corresponding GPU pipeline. For example, at vertical line 755b indicating the end of the rendering phase, all GPUs must wait for the slowest GPU (for example, GPU-B) to finish rendering objects 0 to 3 through the corresponding graphics pipeline before moving on to the next phase of rendering.
[0102] In rendering timing diagram 700C-1, no geometry pre-test is performed. Therefore, each GPU must process each object through its corresponding graphics pipeline. If there are no pixels to be drawn for an object in any area assigned to the corresponding GPU for object rendering (for example, within the corresponding group), the GPU may not fully render the object through the graphics pipeline. For example, when objects do not overlap, only the geometry processing stage of the graphics pipeline is executed. However, this still takes some time for processing.
[0103] Specifically, GPU-A does not fully render objects 0, 1, and 3. This is because they do not overlap with any of the screen areas (e.g., in the corresponding set) assigned to GPU-A for object rendering. The rendering of these three objects is shown within the box with the dashed line. This indicates that at least the geometry processing stage is being performed, but the graphics pipeline is not being fully executed. GPU-A fully renders object 2. This is because that object overlaps with at least one of the screen areas assigned to GPU-A for rendering. The rendering of object 2 is shown within the box with the solid line. This indicates that all stages of the corresponding graphics pipeline are being performed. Similarly, GPU-B does not fully render object 1 (shown by the box with the dashed line) (i.e., at least the geometry processing stage is performed), but fully renders objects 0, 2, and 3 (shown by the boxes with the solid lines). This is because these objects overlap with at least one of the screen areas (e.g., in the corresponding set) assigned to GPU-B for rendering. Also, GPU-C does not fully render objects 0 and 2 (shown by the boxes with the dashed lines) (i.e., at least the geometry processing stage is performed), but fully renders the object (shown by the box with the solid line). This is because these objects overlap with at least one of the screen areas (e.g., in the corresponding set) assigned to GPU-C for rendering. Further, GPU-D does not fully render object 0 (shown by the box with the dashed line) (i.e., at least the geometry processing stage is performed), but fully renders objects 1, 2, and 3 (shown by the boxes with the solid lines). This is because these objects overlap with at least one of the screen areas (e.g., in the corresponding set) assigned to GPU-D for rendering.
[0104] Rendering timing diagram 700C-2 shows the geometry pre-test 701’ and rendering 702’ of objects 0 to 3 using multiple GPUs. In rendering timing diagram 700C-2, vertical line 750a indicates the start of the rendering phase for an object (including, for example, geometry pre-test and rendering), and vertical line 750b indicates the end of the rendering phase for an object. The start and end points along the time axis 740 for the rendering phase shown in timing diagram 700C-2 represent synchronization points. As described above, the four GPUs are synchronized when each executes the corresponding GPU pipeline. For example, at vertical line 750b indicating the end of the rendering phase, all GPUs must wait for the slowest GPU (e.g., GPU-B) to finish rendering objects 0 to 3 through the corresponding graphics pipeline before moving on to the next rendering phase.
[0105] First, the GPU performs a geometry pre-test 701'. Each GPU executes a geometry pre-test for a sub-set of the geometry of the image frame for all screen regions. Each screen region is assigned to the corresponding GPU for object rendering. As described above, each GPU is assigned to the corresponding portion of the geometry associated with the image frame. The geometry pre-test generates information about how each piece of the geometry relates to each of the screen regions (e.g., whether a piece of the geometry overlaps any screen region assigned to the corresponding GPU for object rendering (e.g., in the corresponding set)). This information is shared by each GPU used to render the image frame. For example, the 701' geometry pre-test shown in FIG. 7C includes causing GPU-A to perform a geometry pre-test for object 0, causing GPU-B to perform a geometry pre-test for object 1, causing GPU-C to perform a geometry pre-test for object 2, and causing GPU-D to perform a geometry pre-test for object 3. Depending on the object being tested, the time taken to perform the geometry pre-test can vary. For example, the time taken for the geometry pre-test for object 0 is shorter than when performing a geometry pre-test for object 1. This can be due to factors such as object sizing, the number of overlapping screen regions, etc.
[0106] After the geometry pre-test, each GPU performs rendering on all objects or pieces of geometry that intersect its screen area. In one embodiment, each GPU begins rendering that piece of geometry as soon as the geometry test is complete. That is, there is no synchronization point between the geometry test and the rendering. This is possible because the generated geometry test information is treated as a hint rather than being hardware-dependent. For example, GPU-A begins rendering object 2 before GPU-B finishes the geometry pre-test of object 1, and thus before GPU-B begins rendering objects 0, 2, and 3.
[0107] Vertical line 750a is aligned with vertical line 755a, and rendering timing diagrams 700C-1 and 700C-2 each start simultaneously and are configured to render objects 0-1. However, the rendering of objects 0-3 shown in rendering timing diagram 700C-2 is performed in a shorter time than the rendering shown in rendering timing diagram 700C-1. That is, vertical line 750b indicating the end of the rendering phase for the lower timing diagram 700C-2 appears earlier than vertical line 755b indicating the end of the rendering phase for the upper timing diagram 700C-1. Specifically, when performing multi-GPU rendering of the geometry of an image for an application (including a pre-test of the geometry against the screen area before rendering) and providing the results of the geometry pre-test as information (e.g., a hint), a speed increase 745 is achieved when rendering objects 0-3. As shown, speed increase 745 is the time difference between vertical line 750b of timing diagram 700C-2 and vertical line 755b of timing diagram 700C-1.
[0108] The speed increase is achieved through the generation and sharing of information generated during the geometry pre - test. For example, during the geometry pre - test, GPU - A generates information indicating that it only needs to render object 0 on GPU - B. Thus, GPU - B is notified that it should render object 0, and other GPUs (e.g., GPU - A, GPU - C, and GPU - D) can completely skip rendering object 0. This is because object 0 does not overlap with any area (e.g., in the corresponding set) allocated to these GPUs for object rendering. For example, these GPUs do not need to execute the geometry processing stage. On the other hand, as shown in timing diagram 700C - 1, even if these GPUs do not completely render object 0, this stage is processed without a geometry pre - test. Also, during the geometry pre - test, GPU - B generates information indicating that object 1 should be rendered by GPU - C and GPU - D, and GPU - A and GPU - B may completely skip rendering object 1. This is because object 1 does not overlap with any area (e.g., in the individual corresponding sets) allocated to GPU - A or GPU - B for object rendering. Also, during the geometry pre - test, GPU - C generates information indicating that object 2 should be rendered by GPU - A, GPU - B, and GPU - D, and GPU - C may completely skip rendering object 2. This is because object 2 does not overlap with any area (e.g., in the corresponding set) allocated to GPU - C for object rendering. Further, during the geometry pre - test, GPU - D generates information indicating that object 3 should be rendered by GPU - B, GPU - C, and GPU - D, and GPU - A may completely skip rendering object 3.This is because object 3 does not overlap with any of the areas (e.g., in the corresponding set) assigned to GPU-A for object rendering.
[0109] Since the information generated from the geometry pre-test is shared among the GPUs, each GPU can determine which objects to render. Thus, after performing the geometry pre-test and the results from the test are shared by all GPUs, each GPU has information on which object or piece of geometry needs to be rendered by the corresponding GPU. For example, GPU-A renders object 2, GPU-B renders objects 0, 2, and 3, GPU-C renders objects 1 and 3, and GPU-D renders objects 1, 2, and 3.
[0110] Specifically, GPU A performs geometry processing on object 1 and determines that object 1 can be skipped by GPU-B. This is because object 1 does not overlap with any of the areas (e.g., in the corresponding set) assigned to GPU-B for object rendering. Additionally, object 1 is not fully rendered by GPU-A because object 1 does not overlap with any of the areas (e.g., in the corresponding set) assigned to GPU-A for object rendering. The determination that object 1 does not overlap with any of the areas assigned to GPU-B is made before GPU-B starts geometry processing on object 1, so GPU-B skips rendering object 1.
[0111] Figures 8A-8B show object tests for screen regions 820A and 820B. The screen regions may be interleaved (e.g., screen regions 820A and 820B represent a part of the display). Specifically, multi-GPU rendering of an object is performed by executing a geometry test on an object in the screen for each single image frame or one or more of the image frames within a sequence of image frames before rendering the object in the screen. As shown, GPU-A is assigned the responsibility of rendering the objects within screen region 820A. GPU-B is assigned the responsibility of rendering the objects within screen region 820B. Information for "pieces of geometry" is generated. The piece of geometry can be the entire object or a part of the object. For example, the piece of geometry can be object 810 or a part of object 810.
[0112] Figure 8A is a diagram illustrating an object test for a screen region when multiple GPUs cooperate to render a single image according to an embodiment of the present disclosure. As described above, the piece of geometry can be an object, and the piece corresponds to the geometry used or generated by the corresponding draw call. During the geometry pre-test, it can be determined that object 810 overlaps region 820A. That is, part 810A of object 810 overlaps region 820A. In that case, GPU-A is tasked with rendering object 810. Also, during the geometry pre-test, it can be determined that object 810 overlaps region 820B. That is, part 810B of object 810 overlaps region 820B. In that case, GPU-B is also tasked with rendering object 810.
[0113] FIG. 8B is a diagram illustrating a test of a part of an object with respect to a screen region and / or a screen sub-region when a plurality of GPUs cooperate to render a single image frame according to an embodiment of the present disclosure. That is, a piece of geometry can be taken as a part of an object. For example, the object 810 may be divided into pieces, and the geometry used or generated by a draw call is subdivided into smaller pieces of geometry. In one embodiment, the pieces of geometry are each approximately the size to which a position cache and / or a parameter cache is assigned. In that case, information (e.g., hint or hints) is generated for those smaller pieces of geometry during the geometry test. As described above, the information is used by the rendering GPU.
[0114] For example, the object 810 is divided into smaller objects. The pieces of geometry used for the area test correspond to these smaller objects. As shown, the object 810 is divided into pieces of geometry "a", "b", "c", "d", "e", and "f". After the geometry pre-test, GPU-A renders only the pieces of geometry "a", "b", "c", "d", and "e". That is, GPU-A can skip rendering the piece of geometry "f". Also, after the geometry pre-test, GPU-B renders only the pieces of geometry "d", "e", and "f". That is, GPU-B can skip rendering the pieces of geometry "a", "b", and "c".
[0115] In one embodiment, since the geometry processing stage is configured to perform both vertex processing and primitive processing, geometry pre-tests can be performed on pieces of geometry using a shader in the geometry processing stage. For example, the geometry processing stage can generate information (e.g., hints) by testing the frustum of a cone for the geometry against the GPU screen area (which can be done by software shader operations). In one embodiment, this test is accelerated by using dedicated instructions or instructions through hardware, and as a result, a software / hardware solution is implemented. That is, dedicated instructions or instructions are used to accelerate the generation of information regarding a piece of geometry and its relationship to the screen area. For example, the homogeneous coordinates of the vertices of the primitive of a piece of geometry are provided as input to the instructions for the geometry pre-test in the geometry processing stage. The test may generate a boolean return value indicating whether the primitive overlaps any screen area (e.g., in the corresponding set) assigned to that GPU for object rendering for each GPU. Thus, the information (e.g., hints) generated during the geometry pre-test regarding the corresponding piece of geometry and its relationship to the screen area is generated by the shader in the geometry processing stage.
[0116] In another embodiment, the geometry pre-test for a piece of geometry can be performed in the hardware rasterization stage. For example, a hardware scan converter can be configured to perform the geometry pre-test so that the scan converter generates information regarding all screen areas assigned to multiple GPUs for object rendering of the corresponding image frame.
[0117] In yet other embodiments, the geometry pieces can be primitives. That is, a portion of the object used for geometry pre-testing may be a primitive. Thus, information (e.g., hints) generated by a GPU during geometry pre-testing indicates whether individual triangles (e.g., representing primitives) need to be rendered by another rendering GPU.
[0118] In one embodiment, information generated during geometry pre-testing and shared by the GPU used for rendering includes the number of primitives (e.g., the number of remaining primitives) that overlap with any screen area (e.g., in the corresponding set) assigned to the corresponding GPU for object rendering. The information may also include the number of vertices used to construct or define these primitives. That is, the information includes the number of remaining vertices. Thus, when rendering, the corresponding rendering GPU may allocate space in the position cache and parameter cache using the supplied number of vertices. For example, in one embodiment, since no space is allocated to unnecessary vertices, the efficiency of rendering can be increased.
[0119] In other embodiments, information (e.g., hints) generated during geometry pre-testing includes specific primitives (e.g., primitives remaining as exact matches) that overlap with any screen area (e.g., in the corresponding set) assigned to the corresponding GPU for object rendering. That is, the information generated for the rendering GPU includes a specific set of primitives for rendering. The information may also include specific vertices used to construct or define these primitives. That is, the information generated for the rendering GPU includes a specific set of vertices for rendering. This information can save other rendering GPU time, for example, during that geometry processing stage when rendering a piece of geometry.
[0120] In yet other embodiments, there may be a processing overhead (either software or hardware) associated with the generation of information during the geometry test. In that case, it may be useful to skip the generation of information for certain pieces of geometry. That is, the information provided as a hint is generated for certain objects but not for others. For example, a piece of geometry (e.g., an object or a piece of an object) representing a skybox or a large terrain piece may include large triangles. In that case, each GPU used for multi-GPU rendering of one or more respective image frames within an image frame or a sequence of image frames may need to render these pieces of geometry. That is, depending on the characteristics of the corresponding piece of geometry, information may or may not be generated.
[0121] Figures 9A - 9C illustrate various strategies for allocating screen regions to corresponding GPUs when multiple GPUs cooperate to render a single image according to an embodiment of the present disclosure. To achieve GPU processing efficiency, various techniques may be used when dividing the screen into regions. For example, increasing or decreasing the number of regions (e.g., to select an exact amount of regions), interleaving the regions, increasing or decreasing the number of regions to select a specific pattern when interleaving the regions, etc. For example, multiple GPUs are configured to perform a pre - test of the geometry on interleaved screen regions before rendering the objects in the corresponding image by performing multi - GPU rendering of the geometry for the image frames generated by the application. The screen region configurations in Figures 9A - 9C are designed to reduce even a slight imbalance in rendering times among multiple GPUs. The complexity of the test (e.g., overlapping the corresponding screen region) varies depending on how the screen regions are allocated to the GPUs. As shown in the figures shown in Figures 9A - 9C, the bold box 910 is the outline of the corresponding screen or display used when rendering the image.
[0122] In one embodiment, the multiple screen regions or multiple regions are each of uniform size. In one embodiment, the multiple screen regions are each of non - uniform size. In still other embodiments, the number and sizing of the screen regions among the multiple screen regions change dynamically.
[0123] Specifically, FIG. 9A illustrates a simple pattern 900A for screen 910. The screen regions are each of a uniform size. For example, the size of each region may be a rectangle with a dimension that is a power of 2 pixels. For example, each region may be 256×256 pixels in size. As shown, the region assignment is a checkerboard pattern, with a row of A and B regions alternating with another row of B and C regions. Pattern 900A can be easily tested during a geometry pre-test. However, there may be some rendering inefficiencies. For example, the screen area assigned to each GPU is substantially different (i.e., the coverage for screen regions C and D within screen 910 is smaller). Therefore, it can lead to an imbalance in the rendering time for each GPU.
[0124] FIG. 9B illustrates a pattern 900B of screen regions for screen 910. The screens or sub-regions are each of a uniform size. The screen regions are assigned and distributed so as to reduce the imbalance in rendering time between GPUs. For example, when assigning GPUs to screen regions with pattern 900B, the number of screen pixels assigned to each GPU across screen 910 is approximately equal. That is, the screen regions are assigned to the GPUs so that the screen area or coverage within screen 910 is equal. For example, if each region can be 256×256 pixels in size, the coverage of each region within screen 910 is approximately the same. Specifically, the set of screen regions A covers an area of 6×256×256 pixels in size, the set of screen regions B covers an area of 5.75×256×256 pixels in size, the set of screen regions C covers an area of 5.5×256×256 pixels in size, and the set of screen regions D covers an area of 5.5×256×256 pixels in size.
[0125] FIG. 9C illustrates a pattern 900C of screen areas for a screen 910. The screen areas are not of uniform size. That is, the screen areas to which the responsibility of rendering an object is assigned to the GPU may not be of uniform size. Specifically, the screen 910 is divided such that each GPU is assigned the same number of pixels. For example, when a 4K display (3840×2160) is equally divided into four regions vertically, each region has a height of 520 pixels. However, usually, the GPU performs many operations in 32×32 pixel blocks, and 520 pixels is not a multiple of 32 pixels. Thus, in one embodiment, the pattern 900C may include blocks with a height of 512 pixels (a multiple of 32) and other blocks with a height of 544 pixels (also a multiple of 32). In other embodiments, blocks of different sizes may be used. The pattern 900C shows how an equal amount of screen pixels are assigned to each GPU by using non-uniform screen areas.
[0126] In still other embodiments, the needs of the application when rendering an image change over time, and the screen area is dynamically selected. For example, if it is found that most of the rendering time is spent on the lower half of the screen, it is convenient to allocate regions such that approximately equal amounts of screen pixels in the lower half of the display are assigned to each GPU used to render the corresponding image. That is, the regions assigned to each GPU used to render the corresponding image may be changed dynamically. For example, the change may be applied based on the game mode, different games, the size of the screen, the pattern selected for the region, and the like.
[0127] FIG. 10 is a diagram illustrating various allocations of GPU assignments to pieces of geometry for the purpose of performing a geometry pre-test according to an embodiment of the present disclosure. That is, FIG. 10 shows the allocation of responsibilities for generating information while performing a geometry pre-test among a plurality of GPUs. As described above, each GPU is assigned to the corresponding portion of the geometry of an image frame. That portion can be further divided into objects, parts of objects, geometries, pieces of geometry, and the like. The geometry pre-test includes determining whether a particular piece of geometry overlaps any screen region or regions of the screen assigned to the corresponding GPU for object rendering. The geometry pre-test is typically performed simultaneously by the GPUs for all of the geometry (e.g., all pieces of geometry) of the corresponding image frame in an embodiment. In this way, the geometry tests are performed cooperatively by the GPUs, whereby each GPU can know which piece of geometry to render and which piece of geometry rendering to skip, as described above.
[0128] As shown in FIG. 10, each piece of geometry can be an object, a part of an object, etc. For example, as described above, a piece of geometry can be a part of an object (e.g., a piece that is approximately the size to which a location and / or parameter cache is assigned). For purely illustrative purposes, object 0 (e.g., specified to be rendered by command 722 in rendering command buffer 700A) is divided into pieces "a", "b", "c", "d", "e", and "f" (e.g., object 810 in FIG. 8B). Also, object 1 (e.g., specified to be rendered by command 724 in rendering command buffer 700A) is divided into pieces "g", "h", and "i". Further, object 2 (e.g., specified to be rendered by command 724 in rendering command buffer 700A) is divided into pieces "j", "k", "l", "m", "n", and "o". For the purpose of distributing the responsibility for geometry tests to the GPU, the pieces may be ordered (e.g., a~o).
[0129] Allocation 1010 (e.g., ABCDABCDABCD... rows) indicates an even distribution of the responsibility for the geometry test among multiple GPUs. Specifically, rather than having a certain GPU take the first quarter of the geometry (e.g., in a block, for example GPU A takes the first of approximately 16 pieces that include "a", "b", "c", and "d" for the geometry test), and having the second GPU take the second quarter, etc., the assignments to the GPUs are interleaved. That is, consecutive pieces of the geometry are assigned to different GPUs. For example, piece "a" is assigned to GPU-A, piece "b" is assigned to GPU-B, piece "c" is assigned to GPU-C, piece "d" is assigned to GPU-D, piece "e" is assigned to GPU-A, piece "f" is assigned to GPU-B, piece "g" is assigned to GPU-C, etc. As a result, the processing of the geometry test is roughly balanced among the GPUs (e.g., GPU-A, GPU-B, GPU-C, and GPU-D).
[0130] Assignment 1020 (e.g., the ABBCDABBCDABBCD... row) indicates an asymmetric assignment of the responsibility for the geometry test among multiple GPUs. Asymmetric assignment can be advantageous when a particular GPU spends more time executing the geometry test than other GPUs when rendering the corresponding image frame. For example, a certain GPU may finish rendering an object for a previous frame or frames of a scene earlier than other GPUs, and thus (since this frame is also expected to finish earlier), more pieces of geometry for performing the geometry test can be assigned to that GPU. Also in this case, the assignments to the GPUs are interleaved. As shown, more pieces of geometry for the pre-geometry test are assigned to GPU-B than to other GPUs. For illustration, piece "a" is assigned to GPU-A, piece "b" is assigned to GPU-B, piece "c" is also assigned to GPU-B, piece "d" is assigned to GPU-C, piece "e" is assigned to GPU-D, piece "f" is assigned to GPU-A, piece "g" is assigned to GPU-B, piece "h" is also assigned to GPU-B, piece "i" is assigned to GPU-C, etc. The assignment of the geometry test to the GPUs may not be balanced, but the combined processing of the complete phase (e.g., the pre-geometry test and the rendering of the geometry) may turn out to be roughly balanced (e.g., each GPU spends approximately the same amount of time executing the pre-geometry test and the rendering of the geometry).
[0131] Figures 11A - 11B illustrate using statistical values for one or more image frames when assigning the responsibility for the geometry test among multiple GPUs. For example, based on the statistical values, some GPUs may process more or fewer pieces of geometry during the geometry test to generate useful information when rendering.
[0132] Specifically, FIG. 11A illustrates, according to one embodiment of the present disclosure, the pre-testing and rendering of the geometry of previous image frames by a plurality of GPUs and using the statistical values collected during rendering to affect the assignment of the pre-testing of the geometry of the current image frame to the plurality of GPUs in the current image frame. For purely illustrative purposes, in the second frame 1100B of FIG. 11A, GPU-B processes twice as many pieces of geometry (e.g., during the pre-test) as the other GPUs (e.g., GPU-A, GPU-C, and GPU-D). Assigning and allocating more pieces of geometry to GPU-B to perform the geometry pre-test in the current image frame is based on the statistical values collected during the rendering of the previous image frame or previous image frames.
[0133] For example, timing diagram 1100A shows the geometry pre-test 701A and rendering 702A for a previous image frame. For both processes, four GPUs (e.g., GPU-A, GPU-B, GPU-C, and GPU-D) are used. The assignment of the geometry of the previous image frame (e.g., pieces of geometry) is evenly distributed among the GPUs. This is shown by the roughly balanced performance of the geometry pre-test 701A by each GPU.
[0134] Based on rendering statistical values collected from one or more image frames, it may be determined how to perform the geometry test and rendering of the current image frame. That is, the statistical values may be provided as information to be used when performing the geometry test and rendering of subsequent image frames (e.g., the current image frame). For example, the statistical values collected during the rendering of an object (e.g., a piece of geometry) in a previous image frame may indicate that GPU-B finished rendering earlier than other GPUs. Specifically, GPU-B has an idle time 1130A after rendering that part of the geometry that overlaps any screen area (e.g., in the corresponding set) assigned to GPU-B for object rendering. The other GPUs, GPU-A, GPU-C, and GPU-D, each execute the rendering until approximately the end 710 of the corresponding frame period of the previous image frame.
[0135] Previous image frames and current image frames can be generated for a particular scene when the application is run. Thus, objects from scene to scene can be approximately the same in number and location. In that case, the time for performing geometry pre-tests and rendering is similar for the GPU across multiple image frames in the image frame sequence. That is, based on statistical values, it is reasonable to assume that GPUB also has idle time when performing geometry tests and rendering in the current image frame. Thus, more pieces of geometry for geometry pre-tests in the current frame may be allocated to GPUB. For example, as a result of having GPUB process more pieces of geometry during geometry pre-tests, GPUB finishes at approximately the same time as the other GPUs after rendering the objects in the current image frame. That is, GPUs A, B, C, and D each perform rendering until approximately the end 711 of the corresponding frame period of the current image frame. In one embodiment, the total time for rendering the current image frame is reduced, and the time for rendering the current image frame is shorter when using rendering statistical values. Thus, statistical values for rendering previous frames and / or previous frame(s) may be used to adjust geometry pre-tests (e.g., the distribution of the allocation of geometry (e.g., pieces of geometry) among the GPUs in the current image frame).
[0136] FIG. 11B is a flowchart 1100B that illustrates, according to one embodiment of the present disclosure, a method for performing graphics processing, including pre-testing and rendering the geometry of previous image frames by a plurality of GPUs, and using statistical values collected during rendering to affect the assignment of pre-testing the geometry of the current image frame to a plurality of GPUs in the current image frame. The diagram of FIG. 11A illustrates determining the distribution of the assignment of geometry (e.g., pieces of geometry) among GPUs for an image frame using statistical values in the method of flowchart 1100B. As described above, various architectures may include multiple GPUs cooperating to render a single image by performing multi-GPU rendering of geometry for an application. For example, within one or more cloud gaming servers of a cloud gaming system, or within a stand-alone system (e.g., a personal computer or gaming console that includes a high-end graphics card having multiple GPUs).
[0137] Specifically, at 1110, the method includes rendering graphics for an application using a plurality of GPUs, as described above. At 1120, the method includes dividing the responsibility for rendering the geometry of the graphics among a plurality of GPUs based on a plurality of screen regions. Each GPU has a corresponding division of the responsibility known to the plurality of GPUs. More specifically, as described above, each GPU has the responsibility of rendering the geometry within a corresponding set of screen regions of the plurality of screen regions. The corresponding set of screen regions includes one or more screen regions. In one embodiment, the screen regions are interleaved (e.g., when the display is divided into sets of screen regions for geometry pre-testing and rendering).
[0138] At 1130, the method includes rendering a first plurality of pieces of geometry in a plurality of GPUs for previous image frames generated by an application. For example, timing diagram 1100A illustrates the timing for performing a geometry test on pieces of geometry in a previous image frame and rendering an object (e.g., a piece of geometry). At 1140, the method includes generating statistical values for the rendering of the previous image frame. That is, statistical values may be collected when rendering the previous image frame.
[0139] At 1150, the method includes, based on the statistical values, allocating a second plurality of pieces of geometry of a current image frame generated by the application to the plurality of GPUs for geometry testing. That is, using these statistical values, pieces of geometry for geometry testing may be allocated to specific GPUs the same, less, or more when rendering the next or current image frame. In some cases, the statistical values may indicate that pieces within the second plurality of pieces of geometry must be evenly allocated to the plurality of GPUs when performing the geometry test.
[0140] In other cases, the statistical values may indicate that when performing a geometry test, the pieces within the second plurality of pieces of geometry have to be unevenly allocated to multiple GPUs. For example, as shown on the time axis 1100A, the statistical values may indicate that in a previous image frame, GPU-B finishes rendering earlier than any of the other GPUs. Specifically, it may be determined that the first GPU (e.g., GPU-B) finishes rendering the first plurality of pieces of geometry before the second GPU (e.g., GPU-A) finishes rendering the first plurality of pieces of geometry. As described above, the first GPU (e.g., GPU-B) renders one or more pieces out of the first plurality of pieces of geometry that overlap with any screen area allocated to the first GPU for object rendering, and the second GPU (e.g., GPU-A) renders one or more pieces out of the first plurality of pieces of geometry that overlap with any screen area allocated to the second GPU for object rendering. Therefore, based on the statistical values, it is expected that the time required for the first GPU (e.g., GPU-B) to render the second plurality of pieces of geometry is shorter than that of the second GPU (e.g., GPU-A). Thus, when rendering the current image frame, more pieces of geometry may be allocated to the first GPU for the geometry pre-test. For example, the first number of the second plurality of pieces of geometry may be allocated to the first GPU (e.g., GPU-B) for the geometry test, and the second number of the second plurality of pieces of geometry may be allocated to the second GPU (e.g., GPU-A) for the geometry test. The first number is larger than the second number (if the time imbalance is large enough, no pieces need to be allocated to GPU-A). In this way, GPU-B processes more pieces of geometry than GPU-A during the geometry test.For example, timing diagram 1100B shows that GPU-B has more geometry pieces allocated and spends more time performing geometry tests than other GPUs.
[0141] At 1160, the method includes performing a geometry pre-test on a second plurality of pieces of geometry in the current image frame to generate information regarding the relationship of each piece of the second plurality of pieces of geometry to each of the plurality of screen regions. The geometry pre-test is performed at each of the plurality of GPUs based on the allocation. The geometry pre-test is performed at the pre-test GP on a plurality of pieces of geometry of an image frame generated by the application to generate information regarding the relationship of each piece of geometry to each of the plurality of screen regions.
[0142] At 1170, the method includes using the information generated for each of the second plurality of pieces of geometry to render the plurality of pieces of geometry during the rendering phase (e.g., including fully rendering a piece of geometry or skipping the rendering of that piece of geometry at the corresponding GPU). Rendering is typically performed simultaneously at each GPU in an embodiment. Specifically, the plurality of pieces of geometry of the current image frame are rendered at each of the plurality of GPUs using the information generated for each piece of geometry.
[0143] In other embodiments, the distribution of geometry pieces to the GPU for information generation is dynamically adjusted. That is, the assignment of geometry pieces to the current image frame for performing a geometry pre-test may be dynamically adjusted during the rendering of the current image frame. For example, in the example of timing diagram 1100B, it may be determined that GPU-A is performing the geometry pre-test on its assigned piece of geometry at a rate slower than expected. Thus, the piece of geometry assigned to GPU-A for the geometry pre-test can be reassigned on the fly (e.g., reassign the piece of geometry from GPU-A to GPU-B), and GPU-B is then tasked with performing the geometry pre-test on that piece of geometry during the frame period used to render the current image frame.
[0144] Figures 12A - 12B illustrate another approach for processing a rendering command buffer. Previously, the approach related to Figures 7A - 7C was described. Here, the command buffer contains commands for performing a geometry pre-test on an object (e.g., a piece of geometry), followed by commands for rendering the object (e.g., a piece of geometry). Figures 12A - 12B show a geometry pre-test and rendering approach that uses a shader capable of performing either operation depending on the GPU configuration.
[0145] Specifically, Figure 12A, according to one embodiment of the present disclosure, illustrates using a shader configured to perform both a geometry pre-test and rendering of an image frame in two passes through a portion of command buffer 1200A. That is, the shader used to execute the commands within command buffer 1200A may be configured to perform a geometry pre-test when appropriately configured, or to perform rendering when appropriately configured.
[0146] As shown, a part of the command buffer 1200A shown in FIG. 12A is executed twice, and different operations result from each execution. The first execution results in a geometry pre-test, and the second execution results in the rendering of the geometry. This can be achieved in various ways. For example, a part of the command buffer shown in 1200A can be explicitly called twice as a subroutine. Before each call, different states (e.g., register settings or values in RAM) are explicitly set to different values. Alternatively, a part of the command buffer shown in 1200A can be implicitly executed twice, for example, by using special commands to mark the start and end of that part and executing it twice, and implicitly setting different configurations (e.g., register settings) for the first and second executions of that part of the command buffer. When commands (e.g., commands to set states or commands to execute shaders) in a part of the command buffer 1200A are executed, based on the GPU state, the results of the commands are different (e.g., performing a geometry pre-test versus performing rendering). That is, the commands in the command buffer 1200A may be configured for a geometry pre-test or rendering. Specifically, a part of the command buffer 1200A includes commands for configuring the state of one or more GPUs for executing commands from the rendering command buffer 1200A and commands for executing a shader that performs either a geometry pre-test or rendering depending on the state. For example, commands 1210, 1212, 1214, and 1216 are each used to configure the state of one or more GPUs for the purpose of executing a shader that performs either a geometry pre-test or rendering depending on the state. As shown, command 1210 configures the GPU state so that shader 0 can be executed via command 1211 to perform either a geometry pre-test or rendering. Also, command 1212 configures the GPU state so that shader 1 can be executed via command 1213 to perform a geometry pre-test or rendering.In addition, command 1214 configures the GPU state so that shader 2 can execute either a geometry pretest or rendering via command 1215. Finally, command 1216 configures the GPU state so that shader 3 can execute either a geometry pretest or rendering via command 1217.
[0147] In a first traversal 1291 through command buffer 1200A, corresponding shaders execute a geometry pretest based on the GPU state that is set explicitly or implicitly as described above, and the GPU state configured by commands 1210, 1212, 1214, and 1216. For example, shader 0 is configured to execute a geometry pretest on object 0 (e.g., a piece of geometry) (e.g., based on the object shown in FIG. 7B-1), shader 1 is configured to execute a geometry pretest on object 1, shader 2 is configured to execute a geometry pretest on object 2, and shader 3 is configured to execute a geometry pretest on object 3.
[0148] In one embodiment, based on the GPU state, commands may be skipped or interpreted differently. For example, certain commands for setting the state (the portions of 1210, 1212, 1214, and 1216) may be skipped based on the GPU state set explicitly or implicitly as described above. For example, when configuring shader 0 executed via command 1210, if the GPU state that needs to be configured for geometry pre - testing is less than when configuring for geometry rendering, it may be useful to skip setting the unnecessary portions of the GPU state. This is because setting the GPU state can have overhead. To show another example, certain commands for setting the state (a part of 1210, 1212, 1214, and 1216) may be interpreted differently based on the GPU state set explicitly or implicitly as described above. For example, when shader 0 executed via command 1210 needs to be configured for different GPU states for geometry pre - testing than for geometry rendering, or when shader 0 executed via command 1210 requires different inputs for geometry pre - testing and geometry rendering.
[0149] In one embodiment, a shader configured for geometry pre - testing does not allocate space in the position and parameter cache as described above. In another embodiment, a single shader is used to perform either pre - testing or rendering. This can be done in various ways. For example, it can be done via an external hardware state that the shader can check (e.g., as set explicitly or implicitly as described above), or via the input to the shader (e.g., as set by commands interpreted differently in the first and second passes through the command buffer).
[0150] In a second traverse 1292 through command buffer 1200A, based on the GPU state set explicitly or implicitly as described above, and the GPU state constituted by commands 1210, 1212, 1214, and 1216, the corresponding shader executes rendering of a piece of geometry for the corresponding image frame. For example, shader 0 is configured to execute rendering of object 0 (for example, a piece of geometry) (for example, based on the object shown in FIG. 7B-1). Also, shader 1 is configured to execute rendering of object 1, shader 2 is configured to execute rendering of object 2, and shader 3 is configured to execute rendering of object 3.
[0151] FIG. 12B is a flowchart 1200B illustrating a method for performing graphics processing including performing both pre-testing and rendering of the geometry of an image frame using the same set of shaders in two passes through a portion of a command buffer, according to an embodiment of the present disclosure. As described above, various architectures may include multiple GPUs collaborating to render a single image by performing multi-GPU rendering of the geometry for an application. For example, within one or more cloud gaming servers of a cloud gaming system, or within a stand-alone system (for example, a personal computer or a gaming console including a high-end graphics card having multiple GPUs), and the like.
[0152] Specifically, at 1210, the method includes rendering graphics for an application using a plurality of GPUs, as described above. At 1220, the method includes dividing the responsibility for rendering the geometry of the graphics among a plurality of GPUs based on a plurality of screen regions. Each GPU has a corresponding division of the responsibility known to the plurality of GPUs. More specifically, as described above, each of the GPUs is responsible for rendering the geometry within a corresponding set of screen regions out of the plurality of screen regions. The corresponding set of screen regions includes one or more screen regions. In one embodiment, the screen regions are interleaved (e.g., when the display is divided into sets of screen regions for geometry pre-testing and rendering).
[0153] At 1230, the method includes allocating a plurality of pieces of the geometry of an image frame to a plurality of GPUs for geometry testing. Specifically, each of the plurality of GPUs is allocated a corresponding portion of the geometry associated with the image frame for the purpose of performing a geometry test. As described above, the allocation of the pieces of geometry may be distributed evenly or unevenly. In an embodiment, each portion includes one or more pieces of geometry or potentially no pieces of geometry at all.
[0154] At 1240, the method includes loading a first GPU state that configures one or more shaders to perform a geometry pretest. For example, depending on the GPU state, the corresponding shader may be configured to perform different operations. Thus, the first GPU state configures the corresponding shader to perform a geometry pretest. In the example of FIG. 12A, this can be set in various ways. For example, as described above, it can be done by explicitly or implicitly setting the state externally to a part of the command buffer shown at 1200A. Specifically, the GPU state may be set in various ways. For example, the CPU or GPU can set a value in random access memory (RAM). The GPU checks the value in the RAM. In another example, the state may be internal to the GPU. For example, the command buffer is called twice as a subroutine, and the internal GPU state is different between the two subroutine calls. Alternatively, command 1210 of FIG. 12A can be interpreted differently or skipped based on the state that is explicitly or implicitly set, as described above. Based on this first GPU state, shader 0 executed by command 1211 is configured to perform a geometry pretest.
[0155] At 1250, the method includes performing a geometry pre-test on multiple pieces of geometry in multiple GPUs to generate information regarding the relationship of each piece of geometry to each of the multiple screen regions. As described above, the geometry pre-test may determine whether a piece of geometry overlaps any of the screen regions (e.g., in the corresponding set) assigned to the corresponding GPU for object rendering. The geometry pre-test is typically performed simultaneously by the GPU on all of the geometry for the corresponding image frame in an embodiment, so each GPU can know which pieces of geometry to render and which pieces of geometry to skip. This completes the first pass through the command buffer. The shader may be configured to perform each of the geometry pre-test and / or rendering according to the GPU state.
[0156] At 1260, the method includes loading a second GPU state that configures one or more shaders to perform rendering. As described above, depending on the GPU state, the corresponding shader may be configured to perform different operations. Thus, the second GPU state configures the corresponding shader (the same shader previously used to perform the geometry pre-test) to perform rendering. In the example of FIG. 12A, based on this second GPU state, shader 0 executed by command 1211 is configured to perform rendering.
[0157] At 1270, the method includes using, at each of a plurality of GPUs, information generated for each of a plurality of pieces of geometry when rendering the plurality of pieces of geometry (e.g., to include fully rendering a piece of geometry at the corresponding GPU or skipping rendering of that piece of geometry). As described above, the information may indicate whether a piece of geometry overlaps any screen area (e.g., in the corresponding set) assigned to the corresponding GPU for object rendering. This information may be used to render each of the plurality of pieces of geometry at each of the plurality of GPUs, and each GPU can efficiently render only the pieces of geometry that overlap at least one screen (e.g., in the corresponding set) assigned to that corresponding GPU for object rendering. Thereby, a second traversal through the command buffer ends. The shader may be configured to perform each of the geometry pretests and / or rendering according to the GPU state.
[0158] Figures 13A - 13B illustrate another approach for processing a rendering command buffer. Previously, an approach related to Figures 7A - 7C was described. Here, the command buffer holds commands for a geometry pretest of an object (e.g., a piece of geometry), followed by commands for rendering the object (e.g., the piece of geometry). Also in Figures 12A - 12B, another approach was described that uses a shader capable of performing any operation according to the GPU configuration. Figures 13A - 13B show a geometry test and rendering approach that uses a shader capable of performing either a geometry pretest or rendering. Here, according to an embodiment of the present disclosure, the processes of geometry pretesting and rendering are interleaved for different sets of pieces of geometry.
[0159] Specifically, FIG. 13A is a diagram illustrating the use of a shader configured to perform both geometry pre - testing and rendering. According to one embodiment of the present disclosure, the geometry pre - testing and rendering performed on different sets of pieces of geometry are interleaved using separate portions of the corresponding command buffer 1300A. That is, rather than executing a part of the command buffer 1300A from start to end, the command buffer 1300A is dynamically configured and executed so that the geometry pre - testing and rendering are interleaved for different sets of pieces of geometry. For example, in the command buffer, some shaders (e.g., executed via commands 1311 and 1313) are configured to perform geometry pre - testing on a first set of pieces of geometry. After performing the geometry test, these same shaders (e.g., executed by commands 1311 and 1313) are then configured to perform rendering. After performing rendering on the first set of pieces of geometry, other shaders in the command buffer (e.g., executed via commands 1315 and 1317) are configured to perform geometry pre - testing on a second set of pieces of geometry. After performing the geometry pre - testing, these same shaders (e.g., executed via commands 1315 and 1317) are then configured to perform rendering, and the rendering is performed using these commands on the second set of pieces of geometry. The advantage of this approach is that it becomes possible to dynamically address imbalances between GPUs, for example, by using an asymmetric interleaving of geometry tests throughout the rendering. An example of an asymmetric interleaving of geometry tests was previously introduced in the distribution 102 of FIG. 10.
[0160] Since the interleaving of geometry pre-tests and rendering is done dynamically, the configuration of the GPU (e.g., via register settings or values in RAM) is done implicitly. That is, aspects of the GPU configuration occur outside of the command buffer. For example, the GPU registers may be set to 0 (indicating that a geometry pre-test should be performed) or 1 (indicating that rendering should be performed). Crossing the interleaving of the command buffer and setting this register may be controlled by the GPU based on, for example, the number of objects being processed, the primitives being processed, the imbalance between GPUs, etc. Alternatively, values in RAM can be used. As a result of this external configuration (which means setting externally to the command buffer), when commands (e.g., commands to set state or execute shaders) in a part of command buffer 1300A are executed, based on the GPU state, the results of the commands will be different (e.g., it will be a geometry pre-test versus rendering). That is, the commands in command buffer 1300A may be configured for either geometry pre-test 1391 or rendering 1392. Specifically, a part of command buffer 1300A includes commands for configuring the state of one or more GPUs that execute commands from rendering command buffer 1300A, and commands for executing a shader that performs either a geometry pre-test or rendering depending on the state. For example, commands 1310, 1312, 1314, and 1316 are each used to configure the state of the GPU for the purpose of executing a shader that performs either a geometry pre-test or rendering depending on the state. As shown, command buffer 1310 configures the GPU state so that shader 0 can be executed via command 1311 to perform either a geometry pre-test or rendering of object 0. Also, command buffer 1312 configures the GPU state so that shader 1 can be executed via command 1313 to perform either a geometry pre-test or rendering of object 1.Also, the command buffer 1314 configures the GPU state so that the shader 2 can be executed via the command 1315 to perform either the geometry pre-test or the rendering of the object 2. Further, the command buffer 1316 configures the GPU state so that the shader 3 can be executed via the command 1317 to perform either the geometry pre-test or the rendering of the object 3.
[0161] The geometry pre-test and the rendering may be interleaved for different pieces of the geometry. For illustrative purposes only, the command buffer 1300A may be configured to first execute the geometry pre-test and the rendering of the objects 0 and 1, and then the command buffer 1300A is configured to execute the geometry pre-test and the rendering of the objects 2 and 3 second. Of course, different numbers of pieces of the geometry may be interleaved in different sections. For example, section 1 shows the first traversal through the command buffer 1300A. Based on the GPU state implicitly set as described above, and the GPU state configured by the commands 1310 and 1312, the corresponding shader executes the geometry pre-test. For example, the shader 0 is configured to execute the geometry pre-test for the object 0 (for example, a piece of the geometry) (for example, based on the object shown in FIG. 7B-1), and the shader 1 is configured to execute the geometry pre-test for the object 1. Section 2 shows the second traversal through the command buffer 1300A. Based on the GPU state implicitly set as described above, and the GPU state configured by the commands 1310 and 1312, the corresponding shader executes the rendering. For example, the shader 0 is then configured to execute the rendering of the object 0, and the shader 1 is then configured to execute the rendering of the object 1.
[0162] FIG. 13A shows interleaving the execution of geometry pre-tests and rendering for sets of pieces with different geometries. Specifically, section 3 shows a third partial traversal through command buffer 1300A. Based on the GPU state implicitly set as described above, as well as the GPU state constituted by commands 1314 and 1316, the corresponding shader executes a geometry pre-test. For example, shader 2 (executed via command 1315) executes a geometry test on object 2 (e.g., a piece of geometry) (e.g., based on the object shown in FIG. 7B-1), and shader 3 (executed via command 1317) executes a geometry test on object 3. Section 4 shows a fourth partial traversal through command buffer 1300A. Based on the GPU state implicitly set as described above, as well as the GPU state constituted by commands 1314 and 1316, the corresponding shader executes rendering. For example, shader 2 (executed via command 1315) executes the rendering of object 2, and shader 3 (executed via command 1317) executes the rendering of object 3.
[0163] Note that it should be noted that the hardware context is either retained or recorded (or saved) and read (or restored). For example, the geometry pre-test GPU context at the end of section 1 is required at the beginning of section 3 to perform the geometry pre-test. Also, the rendering GPU context at the end of section 2 is required at the beginning of section 4 for rendering.
[0164] In one embodiment, commands may be skipped or interpreted differently based on the GPU state. For example, certain commands for setting the state (the portions of 1310, 1312, 1314, and 1316) may be skipped based on the GPU state implicitly set as described above. For example, when configuring shader 0 executed via command 1310, if the GPU state that needs to be configured for the geometry test is less than when configuring for the rendering of the geometry, it may be useful to skip setting the unnecessary portions of the GPU state. This is because setting the GPU state can have overhead. To show another example, certain commands for setting the state (the portions of 1310, 1312, 1314, and 1316) may be interpreted differently based on the GPU state implicitly set as described above. For example, when shader 0 executed via command 1310 needs to be configured for a different GPU state for the pre-test of the geometry than for the rendering of the geometry, or when shader 0 executed via command 1310 requires different inputs in the case of the geometry test and the rendering of the geometry.
[0165] In one embodiment, a shader configured for a geometry pre-test does not allocate space in the position and parameter cache as described above. In another embodiment, a single shader is used to perform either the pre-test or the rendering. This can be done in various ways. For example, it can be done via an external hardware state that the shader can check (e.g., as implicitly set as described above), or via the input to the shader (e.g., as set by commands interpreted differently in the first and second passes through the command buffer).
[0166] FIG. 13B is a flowchart illustrating a method for performing graphics processing including interleaving pre - testing and rendering of the geometry of an image frame for different sets of pieces of geometry using separate portions of a corresponding command buffer, according to one embodiment of the present disclosure. As described above, various architectures may include multiple GPUs cooperating to render a single image by performing multi - GPU rendering of the geometry for an application. For example, within one or more cloud gaming servers of a cloud gaming system, or within a stand - alone system (e.g., a personal computer or gaming console including a high - end graphics card having multiple GPUs), and so on.
[0167] Specifically, at 1310, the method includes rendering graphics for an application using multiple GPUs, as described above. At 1320, the method includes dividing the responsibility for rendering the geometry of the graphics among multiple GPUs based on multiple screen regions. Each GPU has a corresponding division of the responsibility known to the multiple GPUs. More specifically, as described above, each GPU is responsible for rendering the geometry within a corresponding set of screen regions of the multiple screen regions. The corresponding set of screen regions includes one or more screen regions. In one embodiment, the screen regions are interleaved (e.g., when the display is divided into sets of screen regions for geometry pre - testing and rendering).
[0168] At 1330, the method includes allocating multiple pieces of the geometry of an image frame to multiple GPUs for geometry testing. Specifically, each of the multiple GPUs is allocated a corresponding portion of the geometry associated with the image frame for the purpose of performing a geometry test. As described above, the allocation of the pieces of geometry may be distributed evenly or unevenly. Each portion may contain one or more pieces of geometry or potentially no pieces of geometry at all.
[0169] At 1340, the method includes interleaving a first set of shaders in a command buffer with a second set of shaders. The shaders are configured to perform both geometry pre - testing and rendering. Specifically, the first set of shaders is configured to perform geometry pre - testing and rendering on a first set of pieces of geometry. Thereafter, the second set of shaders is configured to perform geometry pre - testing and rendering on a second set of pieces of geometry. As described above, the geometry pre - testing generates corresponding information regarding the relationship of each piece of geometry within the first or second set to each of a plurality of screen regions. The plurality of GPUs use the corresponding information to render each piece of geometry within the first or second set. As described above, the GPU state may be set in various ways to perform either geometry pre - testing or rendering. For example, the CPU or GPU can set values in random access memory (RAM). The GPU checks the values in the RAM. In another example, the state may be internal to the GPU. For example, the command buffer is called twice as a subroutine, where the internal GPU state is different between the two subroutine calls.
[0170] The interleaving process will be further described. Specifically, as described above, the first set of shaders in the command buffer is configured to perform a geometry pre-test on the first set of pieces of geometry. The geometry pre-test is performed on the first set of pieces of geometry in a plurality of GPUs to generate first information regarding the relationship of each piece of geometry in the first set to each of the plurality of screen regions. And, as described above, the first set of shaders is configured to perform rendering of the first set of pieces of geometry. Thereafter, the first information is used when rendering the plurality of pieces of geometry in each of the plurality of GPUs (for example, to include completely rendering the first set of pieces of geometry in the corresponding GPU or skipping the rendering of the first set of pieces of geometry). As described above, the information indicates which piece of geometry overlaps with the screen region assigned to the corresponding GPU for object rendering. For example, skipping the rendering of a piece of geometry in the GPU using the information may be done when the information indicates that the piece of geometry does not overlap with any of the screen regions (for example, in the corresponding set) assigned to the GPU for object rendering.
[0171] Then, use the second set of shaders for geometry testing and rendering of the second set of pieces of geometry. Specifically, as described above, the second set of shaders in the command buffer is configured to perform a geometry pre-test on the second set of pieces of geometry. Then, perform the geometry test on the second set of pieces of geometry in a plurality of GPUs to generate second information regarding the relationship of each piece of geometry in the second set to each of the plurality of screen regions. And, as described above, the second set of shaders is configured to perform rendering of the second set of pieces of geometry. Thereafter, perform the rendering of the second set of pieces of geometry in each of the plurality of GPUs using the second information. As described above, the information indicates which piece of geometry overlaps with the screen region (e.g., of the corresponding set) assigned to the corresponding GPU for object rendering.
[0172] Previously, it was described that a plurality of GPUs process geometry in lockstep (i.e., a plurality of GPUs perform a geometry pre-test and a plurality of GPUs perform rendering), but in some embodiments, the GPUs do not explicitly synchronize with each other. For example, while one GPU is rendering a first set of pieces of geometry, a second GPU may be performing a geometry pre-test on a second set of pieces of geometry.
[0173] Figure 14 illustrates components of an example device 1400 that can be used to execute aspects of various embodiments of the present disclosure. For example, in Figure 14, a typical hardware system suitable for performing multi-GPU rendering of geometry for an application is illustrated by performing a pre-test of the geometry against a screen region (which can be interleaved) prior to rendering objects for an image frame, according to embodiments of the present disclosure. The device 1400 illustrated in this block diagram can incorporate or be one of a personal computer, a server computer, a gaming console, a mobile device, or other digital devices, each suitable for executing embodiments of the present invention. The device 1400 includes a central processing unit (CPU) 1402 for executing software applications and optionally an operating system. The CPU 1402 can be composed of one or more homogeneous or heterogeneous processing cores.
[0174] According to various embodiments, the CPU 1402 is one or more general-purpose microprocessors having one or more processing cores. Further embodiments can be implemented using one or more CPUs with a microprocessor architecture specifically adapted for high-parallel and compute-intensive applications (such as media and interactive entertainment applications) configured to perform graphics processing during the execution of a game.
[0175] Memory 1404 stores the applications and data used by CPU 1402 and GPU 1416. Storage device 1406 is a non-volatile storage device and other computer-readable media for applications and data, and may include a fixed disk drive, a removable disk drive, a flash memory device, and a CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other optical storage device, as well as signal transmission and storage media. User input device 1408 transmits user input from one or more users to device 1400. Examples thereof may include a keyboard, a mouse, a joystick, a touch pad, a touch screen, a steel or video recorder / camera, and / or a microphone. Through network interface 1409, device 1400 can communicate with other computer systems via an electronic communication network. Examples of network interface 1409 may include wired or wireless communication via a local area network and a wide area network (e.g., the Internet). Audio processor 1412 is adapted to generate analog or digital audio output from the instructions and / or data provided by CPU 1402, memory 1404, and / or storage device 1406. The components of device 1400 (e.g., CPU 1402, graphics subsystem, e.g., GPU 1416, memory 1404, data storage device 1406, user input device 1408, network interface 1409, and audio processor 1412) are connected via one or more data buses 1422.
[0176] The graphics subsystem 1414 is further connected to the data bus 1422 and the components of the device 1400. The graphics subsystem 1414 includes at least one graphics processing unit (GPU) 1416 and graphics memory 1418. The graphics memory 1418 includes a display memory (e.g., a frame buffer) used to store pixel data for each pixel of the output image. The graphics memory 1418 can be integrated into the same device as the GPU 1416, connected to the GPU 1416 as a separate device, and / or implemented within the memory 1404. Pixel data can be provided directly from the CPU 1402 to the graphics memory 1418. Alternatively, the CPU 1402 provides data and / or instructions that define the desired output image to the GPU 1416. From the desired output image, the GPU 1416 generates pixel data for one or more output images. The data and / or instructions that define the desired output image can be stored in the memory 1404 and / or the graphics memory 1418. In one embodiment, the GPU 1416 includes 3D rendering capabilities for generating pixel data for the output image from instructions and data that define geometry, lighting, shading, texturing, motion, and / or camera parameters for the scene. The GPU 1416 can further include one or more programmable execution units capable of executing shader programs.
[0177] The graphics subsystem 1414 periodically outputs pixel data for an image to be displayed on the display device 1410 or projected by a projection system (not shown) from the graphics memory 1418. The display device 1410 can be any device capable of displaying visual information in response to a signal from the device 1400. For example, a CRT, LCD, plasma, and OLED display. The device 1400 can provide, for example, an analog or digital signal to the display device 1410.
[0178] In other embodiments for optimizing the graphics subsystem 1414, multi-GPU rendering of geometry for an application can be included by pre-testing the geometry against screen regions (which can be interleaved) before rendering objects for an image frame. The graphics subsystem 1414 can be configured as one or more processing devices.
[0179] For example, in one embodiment, the graphics subsystem 1414 may be configured to perform multi-GPU rendering of geometry for an application. Multiple graphics subsystems can execute graphics and / or render pipelines for a single application. That is, the graphics subsystem 1414 includes multiple GPUs used to render each of one or more images of an image or image sequence when executing an application.
[0180] In other embodiments, the graphics subsystem 1414 includes multiple GPU devices. These are combined to perform graphics processing for a single application running on the corresponding CPU. For example, multiple GPUs can perform multi-GPU rendering of geometry for an application by pre-testing the geometry against screen regions (which can be interleaved) before rendering objects for an image frame. In other examples, multiple GPUs can perform frame rendering in an alternating fashion. Here, GPU1 renders the first frame, GPU2 renders the second frame, and this continues until the last GPU is reached, such as in consecutive frame cycles. Then, the first GPU renders the next video frame (e.g., if only two GPUs exist, GPU1 renders the third frame). That is, the GPUs rotate when rendering frames. The rendering operations can overlap. GPU2 may start rendering the second frame before GPU1 finishes rendering the first frame. In another embodiment, different shader operations can be assigned to multiple GPU devices in the rendering and / or graphics pipeline. The master GPU performs the main rendering and composition. For example, in a group including three GPUs, master GPU1 can perform the main rendering (e.g., the first shader operation) and composition of the outputs from slave GPUs 2 and 3. Slave GPU2 can perform the second shader (e.g., fluid effects such as a river) operation, and slave GPU3 can perform the third shader (e.g., particle smoke) operation. Master GPU1 composes the results from each of GPU1, GPU2, and GPU3. Thus, different GPUs can be assigned to perform different shader operations (e.g., waving a flag, wind, smoke, fire, etc.) to render a video frame.In yet another embodiment, each of the three GPUs can be assigned to different objects and / or portions of a scene corresponding to a video frame. In the foregoing embodiments and aspects, these operations can be performed in the same frame period (simultaneously in parallel) or in different frame periods (sequentially in parallel).
[0181] Accordingly, the present disclosure describes a method and system configured for multi-GPU rendering of geometry for an application by pre-testing the geometry against a screen region (which can be interleaved) before rendering the objects for each of one or more respective image frames in an image frame or sequence of image frames when executing the application.
[0182] Of course, the various embodiments defined herein may be combined or assembled into a specific implementation using the various features disclosed herein. Accordingly, the examples provided are merely some possible examples and are not limited to the various embodiments possible by combining the various elements to define even more embodiments. In some examples, some embodiments may include even fewer elements without departing from the spirit of the disclosed embodiments or equivalent embodiments.
[0183] Embodiments of the present disclosure may be executed by various computer system configurations (e.g., handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, etc.). Embodiments of the present disclosure may also be executed in a distributed computing environment where tasks are performed by remote processing devices linked through a wired or wireless network.
[0184] In view of the foregoing embodiments, it is a matter of course that embodiments of the present disclosure can use various computer-implemented operations involving data stored in a computer system. These operations require physical manipulation of physical quantities. Any of the operations described herein that constitute part of the embodiments of the present disclosure are useful machine operations. The embodiments of the present disclosure also relate to devices or apparatuses for performing these operations. The apparatus can be specially configured for the required purpose, or the apparatus can be a general-purpose computer selectively operated or configured by a computer program stored in a computer. Specifically, various general-purpose machines can also be used by a computer program written according to the teachings herein, or it may be more convenient to configure a more specialized apparatus to perform the required operations.
[0185] The present disclosure can also be embodied as computer-readable code on a computer-readable medium. The computer-readable medium can be any data storage device capable of storing data. The data can then be read by a computer system. Examples of computer-readable media include hard drives, network-attached storage (NAS), read-only memory, random access memory, CD-ROM, CD-R, CD-RW, magnetic tape, and other optical and non-optical data storage devices. The computer-readable medium can include a computer-readable tangible expression medium distributed on a network-coupled computer system so that the computer-readable code is stored and executed distributively.
[0186] Although the operations of the present method have been described in a specific order, it is a matter of course that other housekeeping operations may be performed between the operations, or the operations may be adjusted to be performed at slightly different times, or the processing operations may be distributed in a system that allows the processing operations to be performed at various intervals associated with the processing as long as the overlay operation processing is performed in a desired manner.
[0187] Although the foregoing disclosure has been described in some detail for clarity of understanding, it will be apparent that certain variations and modifications can be practiced within the scope of the appended claims. Accordingly, the embodiments should be considered illustrative rather than restrictive, and the embodiments of the present disclosure should not be limited to the details shown herein, but may be modified within the scope of the appended claims and equivalents.
Claims
1. A method for performing graphics processing, comprising: rendering graphics for an application using a plurality of graphics processing units (GPUs); dividing the responsibility for rendering the geometry of the graphics for a plurality of image frames among the plurality of GPUs based on a plurality of screen regions, each GPU having a corresponding division of the responsibility known to the plurality of GPUs; allocating a plurality of geometry pieces of an image frame to the plurality of GPUs during a frame period for a geometry test; setting, in a random access memory during the frame period, a first value that defines a first GPU state for the plurality of GPUs for performing the geometry test, the first value being set externally from the execution of a command buffer; performing a first subroutine call during the frame period to execute the geometry test based on the first value that defines the first GPU state, the subroutine including a plurality of commands in the command buffer; in a first pass of the command buffer during the frame period by each of the plurality of GPUs for processing the subroutine that was first called, each of the plurality of GPUs processes the subroutine that was first called, executes a first command group among the plurality of commands, and configures each of one or more shaders executed by the plurality of GPUs for performing the geometry test on the plurality of geometry pieces of the image frame, and execution of some commands within the first command group is skipped based on the first value, without requiring execution of the geometry test for each of the one or more shaders. In the first pass of the command buffer during the frame period, by executing the remaining commands among the plurality of commands, the one or more shaders executed by the plurality of GPUs perform the geometry test on the plurality of geometry pieces of the image frame, and generate information regarding each geometry piece and its relationship to each of the plurality of screen regions for the image frame. In the execution of the geometry test by executing the remaining commands among the plurality of commands, the interpretation of the subset of the remaining commands is performed using a first interpretation based on the first value that defines the state of the first GPU. In the frame period, a second value that defines a second GPU state for the plurality of GPUs for executing the rendering of the image frame using the plurality of commands in the command buffer is set in the random access memory. The second value is externally set from the execution of the command buffer. In the frame period, a second subroutine call is made to execute the rendering of the image frame based on the second value that defines the second GPU state. In the second pass of the command buffer during the frame period, each of the plurality of GPUs that processes the subroutine called the second time executes the first command group, and configures each of the one or more shaders executed by the plurality of GPUs for executing the rendering of the image frame. The commands in the first command group are executed based on the second value, assuming it is necessary for configuring each of the one or more shaders to execute the rendering of the image frame. In the second pass of the command buffer during the frame period, when executing the remaining commands among the plurality of commands for executing the rendering of the plurality of geometry pieces of the image frame by the plurality of GPUs that render the image frame, the information generated for the plurality of geometry pieces of the image frame is used. The corresponding shader executed by the corresponding GPU is configured to perform the geometry test on the geometry piece of the corresponding object of the image frame in the first pass of the command buffer during the frame period, and the corresponding shader executed by the corresponding GPU is configured to perform the rendering of the geometry piece of the corresponding object of the image frame in the second pass of the command buffer during the frame period. In the execution of the remaining commands among the plurality of commands, the interpretation of the subset of the remaining commands is performed using a second interpretation based on the second value that defines the state of the second GPU. Method.
2. When using the information, if the information indicates that the geometry piece does not overlap with any screen area allocated to the rendering GPU for object rendering, the rendering of the geometry piece in the rendering GPU is skipped, and the rendering GPU is one of the plurality of GPUs. The method according to claim 1.
3. Furthermore, the information is provided as a hint to the rendering GPU, the rendering GPU is one of the plurality of GPUs, the information is considered by the rendering GPU when received before the rendering of the geometry piece, when the information is received after the rendering of the geometry piece has started, the geometry piece is fully rendered in the rendering GPU. The method according to claim 1.
4. Allocate the geometry pieces within the plurality of geometry pieces uniformly or non-uniformly across the plurality of GPUs. Allocate the plurality of geometry pieces such that consecutive geometry pieces are processed by different GPUs. The method according to claim 1.
5. The first GPU performs the geometry pre-test on more geometry pieces than the second GPU, or the first GPU performs the geometry pre-test while the second GPU does not perform any geometry pre-test at all. The method according to claim 1.
6. When executing the commands in the corresponding command buffer according to whether the plurality of GPUs are configured in the first GPU state or the second GPU state, the commands cause the output of the information regarding the geometry piece or the output of vertex positions and parameter information so as to be used in one or more rendering stages later. The corresponding command buffer includes the commands of the subroutine, and when these commands are executed, according to which of the first GPU state or the second GPU state is configured for the one or more shaders, the geometry test is executed or the rendering of the image frame is executed. The method according to claim 1.
7. According to whether the plurality of GPUs are configured in the first GPU state or the second GPU state, the commands affecting the configuration of the GPUs are interpreted in one of two ways. The method according to claim 6.
8. Interleave the generation of first information for the first geometry piece and its relationship to the plurality of screen regions and the rendering of the first geometry piece, and the generation of second information for the second geometry piece and its relationship to the plurality of screen regions and the rendering of the second geometry piece. The method according to claim 1.
9. Is the hardware context retained, or recorded and read? The method according to claim 1.
10. One or more of the plurality of GPUs are parts of a larger GPU configured as a plurality of virtual GPUs. The method according to claim 1.
11. A method for performing graphics processing, comprising: Rendering graphics for an application using a plurality of graphics processing units (GPUs); Dividing the responsibility for the rendering of the geometry of the graphics among the plurality of GPUs based on a plurality of screen regions for a plurality of image frames, each GPU having a corresponding division of the responsibility known to the plurality of GPUs; Allocating a plurality of geometry pieces of an image frame to the plurality of GPUs during a frame period for geometry testing. In each of the plurality of GPUs, a first shader set configured to perform geometry testing and rendering on a first set of geometry pieces of the image frame during the frame period, and a second shader set configured to perform geometry testing and rendering on a second set of geometry pieces of the image frame during the frame period are interleaved, and the geometry testing and the rendering for the first set of geometry pieces executed by the first shader set are fully executed during the frame period before the geometry testing and the rendering for the second set of geometry pieces executed by the second shader set during the frame period. The geometry test for the first set of geometry pieces and the geometry test for the second set of geometry pieces generate correspondence information regarding each geometry piece of the image frame and its relationship with respect to each of the plurality of screen regions. The correspondence information associates the plurality of geometry pieces with respect to the plurality of screen regions of the image frame used by the plurality of GPUs when rendering the plurality of geometry pieces of the image frame by the plurality of GPUs, and the information indicates that each of the plurality of geometry pieces overlaps with any of the plurality of screen regions. A first value is set in the random access memory during the frame period, and a first GPU state for the plurality of GPUs is defined using a plurality of commands in a first pass of the command buffer to execute the geometry test, and the first value is externally set from the execution of the command buffer. A first subroutine call is made during the frame period to execute the geometry test based on the first value defining the first GPU state, the subroutine includes a plurality of commands in the command buffer, and in the execution of the geometry test by executing the plurality of commands, an interpretation of a subset of the plurality of commands is performed using a first interpretation based on the first value defining the state of the first GPU. In the first pass of the command buffer, each of the plurality of GPUs processes the subroutine that was first called during the frame period, and a first command group in the command buffer is executed, and the first shader set is configured to execute the geometry test on the first geometry piece set. Assuming it is not necessary to configure the first shader set to execute the geometry test on the first geometry piece set, the execution of some commands is skipped based on the first value. In the frame period, a second value is set in the random access memory to define a second GPU state for the plurality of GPUs, and the rendering of the first geometry piece set is executed using the plurality of commands in the second pass of the command buffer. The second value is set externally from the execution of the command buffer. In the frame period, a second subroutine call is made to execute the rendering of the image frame based on the second value that defines the second GPU state. In the second pass of the command buffer, each of the plurality of GPUs for the subroutine that was second called executes the plurality of commands in the command buffer, and the first shader set is configured to execute the rendering on the first geometry piece set based on the second value. Assuming it is necessary to configure the first shader set to execute the rendering on the first geometry piece set, the execution of the commands in the first command group is executed based on the second value. In the execution of the plurality of commands for executing the rendering, the interpretation of the subset of the plurality of commands is performed using a second interpretation based on the second value that defines the state of the second GPU. The corresponding shader executed by the corresponding GPU is configured to perform the geometry test on the geometry pieces of the corresponding object of the image frame in the first pass of the command buffer during the frame period, and the corresponding shader executed by the corresponding GPU is configured to perform the rendering of the geometry pieces of the corresponding object of the image frame in the second pass of the command buffer during the frame period. Method. **Claim 12** When the information indicates that the geometry pieces in the first geometry piece set do not overlap with any screen area assigned to the rendering GPU for object rendering, the rendering of the geometry pieces in the first geometry piece set is skipped in the rendering GPU, and the rendering GPU is one of the plurality of GPUs. When the information indicates that the geometry pieces in the second geometry piece set do not overlap with any screen area assigned to the rendering GPU for object rendering, the rendering of the geometry pieces in the second geometry piece set is skipped in the rendering GPU. The method according to claim 11. **Claim 13** In the interleaving, configure the first shader set of the command buffer to perform a geometry test on the first geometry piece set. In the plurality of GPUs, use the first shader set to perform a geometry test on the first geometry piece set and generate first information regarding the relationship between each geometry piece in the first geometry piece set and each of the plurality of screen areas. configure the first shader set to perform the rendering of the first geometry piece set. When the first information indicates that the first geometry piece in the first geometry piece set does not overlap with any screen area assigned to the first rendering GPU for object rendering, the rendering of the first geometry piece is skipped in the first rendering GPU. Configure the second shader set of the command buffer to perform a geometry test on the second geometry piece set, In the plurality of GPUs, use the second shader set to perform a geometry test on the second geometry piece set to generate second information regarding the relationship of each geometry piece in the second geometry piece set to each of the plurality of screen regions, Configure the second shader set to perform rendering of the second geometry piece set, and When the second information indicates that a second piece within a second geometry piece does not overlap with any screen region assigned to a second rendering GPU for object rendering, skip rendering of the second geometry piece in the second rendering GPU, The method according to claim 11.
14. Further, provide the corresponding information as a hint to a rendering GPU, The rendering GPU is one of the plurality of GPUs, The information is considered by the rendering GPU when received before rendering of the corresponding geometry piece, When the information is received after rendering of the corresponding piece of geometry has started, the corresponding piece of geometry is fully rendered in the rendering GPU, The method according to claim 11.
15. Allocate the geometry pieces within the plurality of geometry pieces uniformly or non-uniformly across the entirety of the plurality of GPUs, Allocate the plurality of geometry pieces such that consecutive geometry pieces are processed by different GPUs, The method according to claim 11.
16. One or more of the plurality of GPUs are portions of a larger GPU configured as a plurality of virtual GPUs, The method according to claim 11.
17. A computer system, comprising A processor, and A memory coupled to the processor and having instructions stored thereon, the instructions, when executed by the computer system, cause the computer system to execute a method for performing graphics processing, the method comprising Render graphics for an application using a plurality of graphics processing units (GPUs), divide the responsibility for rendering the geometry of the graphics for a plurality of image frames among the plurality of GPUs based on a plurality of screen regions, each GPU having a corresponding division of the responsibility known to the plurality of GPUs, assign a plurality of geometry pieces of an image frame to the plurality of GPUs during a frame period for geometry testing, set, in random access memory during the frame period, a first value that defines a first GPU state for the plurality of GPUs for performing the geometry test, the first value being set externally from the execution of a command buffer, make a first subroutine call during the frame period for performing the geometry test based on the first value that defines the first GPU state, the subroutine including a plurality of commands in the command buffer, in a first pass of the command buffer during the frame period by each of the plurality of GPUs for processing the subroutine for which the first call was made, each of the plurality of GPUs processes the subroutine for which the first call was made, executes a first command group of the plurality of commands, and configures each of one or more shaders to be executed by the plurality of GPUs for performing the geometry test on the plurality of geometry pieces of the image frame, and execution of some of the commands in the first command group is skipped based on the first value, without being required to perform the geometry test for each of the one or more shaders, In the first pass of the command buffer during the frame period, by executing the remaining commands among the plurality of commands, the one or more shaders executed by the plurality of GPUs perform the geometry test on the plurality of geometry pieces of the image frame, and generate information regarding each geometry piece and its relationship with respect to each of the plurality of screen regions for the image frame. In the execution of the geometry test by executing the remaining commands among the plurality of commands, the interpretation of the subset of the remaining commands is performed using a first interpretation based on the first value that defines the state of the first GPU. During the frame period, a second value that defines a second GPU state for the plurality of GPUs for executing the rendering of the image frame using the plurality of commands in the command buffer is set in the random access memory. The second value is externally set from the execution of the command buffer. During the frame period, a second subroutine call is made to execute the rendering of the image frame based on the second value that defines the second GPU state. In the second pass of the command buffer during the frame period, each of the plurality of GPUs that processes the subroutine called the second time executes the first command group, and configures each of the one or more shaders executed by the plurality of GPUs for executing the rendering of the image frame. The commands in the first command group are executed based on the second value, assuming it is necessary to configure each of the one or more shaders to execute the rendering of the image frame. In the second pass of the command buffer during the frame period, when executing the remaining commands among the plurality of commands for executing the rendering of the plurality of geometry pieces of the image frame by the plurality of GPUs that render the image frame, the information generated for the plurality of geometry pieces of the image frame is used. The corresponding shader executed by the corresponding GPU is configured to perform the geometry test on the geometry piece of the corresponding object of the image frame in the first pass of the command buffer during the frame period, and the corresponding shader executed by the corresponding GPU is configured to perform the rendering of the geometry piece of the corresponding object of the image frame in the second pass of the command buffer during the frame period. In the execution of the remaining commands among the plurality of commands, the interpretation of the subset of the remaining commands is performed using a second interpretation based on the second value that defines the state of the second GPU. A computer system.
18. When using the information in the method, if the information indicates that the geometry piece does not overlap any screen area allocated to the rendering GPU for object rendering, the rendering of the geometry piece in the rendering GPU is skipped, and the rendering GPU is one of the plurality of GPUs. The computer system according to claim 17.
19. The method further includes providing the information as a hint to the rendering GPU, wherein the rendering GPU is one of the plurality of GPUs, the information is considered by the rendering GPU if received before the rendering of the geometry piece, and if the information is received after the rendering of the geometry piece has started, the geometry piece is fully rendered in the rendering GPU. The computer system according to claim 17.
20. In the method, the geometry pieces within the plurality of geometry pieces are allocated uniformly or non-uniformly across the plurality of GPUs. In the method, the plurality of geometry pieces are allocated such that consecutive geometry pieces are processed by different GPUs. The computer system according to claim 17.
21. In the method, the first GPU executes a geometry pretest on a greater number of geometry pieces than the second GPU, or the first GPU executes a geometry pretest while the second GPU does not execute any geometry pretest at all. The computer system according to claim 17.
22. In the method, when executing the commands in the corresponding part of the command buffer according to whether the plurality of GPUs are configured in the first GPU state or the second GPU state, the commands cause an output of the information regarding the geometry piece so as to be used in one or more rendering stages later, or cause an output of vertex positions and parameter information. The corresponding part of the command buffer includes the commands of the subroutine, and when these commands are executed, they execute the geometry test or execute the rendering of the image frame based on which of the first GPU state or the second GPU state is configured for the one or more shaders. The computer system according to claim 17.
23. In the method, depending on whether the plurality of GPUs are configured in the first GPU state or the second GPU state, the commands affecting the configuration of the GPUs are interpreted in one of two ways. The computer system according to claim 22.
24. Interleaving the generation of first information for the first geometry piece and its relationship to the plurality of screen regions and the rendering of the first geometry piece, and the generation of second information for the second geometry piece and its relationship to the plurality of screen regions and the rendering of the second geometry piece. The computer system according to claim 17.
25. In the method, the hardware context is held or recorded and read. The computer system according to claim 17.
26. In the method, one or more of the plurality of GPUs are parts of a larger GPU configured as a plurality of virtual GPUs. The computer system according to claim 17.
Citation Information
Patent Citations
Image processing system and image processing method
JP2008071261A
Displaying compressed supertile images
JP2013541746A
Optimized multi-pass rendering on tile-based architectures
JP2017505476A
Multi-GPU frame rendering
US20190206023A1