Cloud Distributed Graphics Rendering System, Method, Electronic Device and Medium

By sharing computational steps and data across rendering pipelines in a cloud-based distributed rendering system, the system optimizes GPU resource use, reducing rendering costs and improving efficiency.

CN116310026BActive Publication Date: 2025-07-15SHANGHAI BIREN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310153884.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-22
Publication Date
2025-07-15
Estimated Expiration
2043-02-22

AI Technical Summary

Technical Problem

In the prior art, the rendering and computing steps of each game instance in the cloud distributed GPU rendering system are processed independently, resulting in waste of computing resources and high overall rendering costs, especially in multiplayer battle games, which increases the computing burden.

Method used

By sharing the calculation steps and data of multiple graphics rendering pipelines, multiple graphics processing units in the cloud perform preprocessing and shadow lighting preprocessing and other calculations independent of view angles, reducing duplicate calculations, optimizing resource scheduling to reduce overall rendering costs.

Benefits of technology

It effectively reduces the overall cost of cloud rendering systems, improves the utilization rate of computing resources, and reduces rendering time, especially the rendering efficiency of shared scenes in multiplayer games.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310026B_ABST
    Figure CN116310026B_ABST
Patent Text Reader

Abstract

Provided are a cloud distributed graphics rendering system, a cloud distributed graphics rendering method, an electronic device, and a non-transitory storage medium. The system includes: a plurality of graphics processing units in the cloud; wherein one or more of the plurality of graphics processing units run calculation steps that can be shared by a plurality of graphics rendering pipelines to output and / or cache the data obtained from the running for use by the calculation steps of the plurality of graphics rendering pipelines. In this way, by sharing as much calculation and data as possible, the overall rendering cost of the cloud is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer graphics and cloud computing, and more particularly, to a cloud-based distributed graphics rendering system, a cloud-based distributed graphics rendering method, an electronic device, and a non-transitory storage medium. Background Art

[0002] A Graphics Processing Unit (GPU), namely a graphics card, is used to render graphics. Application fields for rendering graphics include: games, advertising, and movies, television, and animation, visual effects in product design, architecture, education, healthcare, smart cities, the metaverse, virtual reality (VR), augmented reality (AR), extended reality (XR), and so on.

[0003] Currently, cloud rendering technology has been developed, that is, cloud native technology based on distributed GPUs on the cloud is used for graphics rendering. Summary of the Invention

[0004] According to one aspect of the present application, there is provided a cloud-based distributed graphics rendering system, including: multiple graphics processing units in the cloud; wherein, one or more of the multiple graphics processing units run calculation steps that can be shared by multiple graphics rendering pipelines to output and / or cache the data obtained from the operation for use by the calculation steps of the multiple graphics rendering pipelines.

[0005] According to another aspect of the present application, there is provided a cloud-based distributed graphics rendering method, including: running, by one or more of the multiple graphics processing units in the cloud, calculation steps that can be shared by multiple graphics rendering pipelines; outputting the data obtained from the operation for use by the calculation steps of the multiple graphics rendering pipelines.

[0006] According to another aspect of the present application, there is provided an electronic device, including: a memory for storing instructions; a processor for reading the instructions in the memory and executing the method according to an embodiment of the present application.

[0007] According to another aspect of the present application, there is provided a non-transitory storage medium having instructions stored thereon, wherein when the instructions are read by a processor, the processor is caused to execute the method according to an embodiment of the present application.

[0008] In this way, by sharing as much calculation and data as possible, the overall rendering cost in the cloud is reduced. Brief Description of the Drawings

[0009] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0010] Figure 1 The block diagram of a cloud system that uses distributed GPUs on the cloud for game graphics rendering in the prior art is shown.

[0011] Figure 2 The block diagram of a cloud distributed graphics rendering system according to an embodiment of the present application is shown.

[0012] Figure 3 The schematic diagram showing whether each processing module in the graphics rendering pipeline according to an embodiment of the present application can be shared is shown.

[0013] Figure 4 The schematic diagram showing whether each calculation step and data of each processing module in the graphics rendering pipeline according to an embodiment of the present application can be shared is shown.

[0014] Figure 5A and 5B The schematic diagram showing the calculation stages and output data that can be shared by multiple graphics rendering pipelines in two types of skinning calculation steps according to an embodiment of the present application is shown.

[0015] Figure 6A and Figure 6B The schematic diagrams showing the view-independent shadow map stage and the view-dependent shadow map stage in shadow and lighting preprocessing according to an embodiment of the present application are shown respectively.

[0016] Figure 6C The schematic diagram showing the lighting preprocessing stage in shadow and lighting preprocessing according to an embodiment of the present application is shown.

[0017] Figure 7A The schematic diagram showing the lighting stage in the opaque lighting calculation step according to an embodiment of the present application is shown.

[0018] Figure 7B The schematic diagram showing the view-independent ambient occlusion stage in the opaque lighting calculation step according to an embodiment of the present application is shown.

[0019] Figure 8 The block diagram of a cloud distributed graphics rendering system according to an embodiment of the present application is shown.

[0020] Figure 9Shows a schematic diagram of the operation stage of an accelerator as a cloud rendering engine according to an embodiment of the present application.

[0021] Figure 10 Shows a flowchart of a cloud-based distributed graphics rendering method according to an embodiment of the present application.

[0022] Figure 11 Shows a block diagram of an exemplary electronic device suitable for implementing an embodiment of the present application.

[0023] Figure 12 Shows a schematic diagram of a non-transitory computer-readable storage medium according to an embodiment of the present application. Detailed Description of the Invention

[0024] Now, specific embodiments of the present application will be described in detail. Examples of the present application are illustrated in the accompanying drawings. Although the present application will be described in conjunction with specific embodiments, it will be understood that it is not intended to limit the present application to the described embodiments. On the contrary, it is intended to cover modifications, variations, and equivalents included within the spirit and scope of the present application as defined by the appended claims. It should be noted that the method steps described herein can be implemented by any functional block or functional arrangement, and any functional block or functional arrangement can be implemented as a physical entity or a logical entity, or a combination of both.

[0025] Graphic rendering needs to complete the following basic functions: a scene infinitely close to reality (using lighting models and physical models), 3D element animation, forward rendering algorithms for actions, film-level post-production special effects processing, and so on. Since the Graphic Processing Unit (GPU) needs to process a large amount of data and operations for graphic acceleration and real-time rendering when rendering graphics, the graphic rendering of a video may take a long time and consume a large amount of computing resources. Currently, cloud rendering technology has been developed, that is, using multiple distributed GPUs based on the cloud (such as cloud native technology) to perform graphic rendering. By using multiple distributed GPUs on the cloud for graphic rendering, the computing pressure on each client to perform graphic rendering with its own hardware is reduced, and the speed of graphic rendering is accelerated. Therefore, it is widely applied in many fields, such as games, advertisements, movies, TV shows, animations, visual effects in product design, architecture, education, medical treatment, smart cities, the metaverse, virtual reality (VR), augmented reality (AR), extended reality (XR), and so on.

[0026] Figure 1 Shows a block diagram of a cloud system 100 for using distributed GPUs on the cloud to perform game graphic rendering in the prior art.

[0027] As Figure 1As shown, there are game instances to be rendered, namely game instance 1 101, game instance 2 102, …… game instance N 103 (where N is a positive integer). The prior art method of using distributed GPUs on the cloud for game graphics rendering is to use GPU 1 107, GPU 2 108, …… GPU N 109 on cloud system 100 to separately process multiple rendering calculation steps (passes) 104, 105, 106 in multiple rendering pipelines for their respective games. For example, GPU 1 107 is responsible for processing multiple rendering calculation steps 104 in multiple rendering pipelines of game instance 1 101, GPU 2 108 is responsible for processing multiple rendering calculation steps 105 in multiple rendering pipelines of game instance 2 102, …… GPU N 109 is responsible for processing multiple rendering calculation steps 106 in multiple rendering pipelines of game instance N 103.

[0028] Both the rendering pipeline and the calculation steps here are concepts at the software level. A rendering pipeline is a series of operations related to lighting and shading, etc. Each rendering pipeline can include multiple rendering calculation steps. An object to be rendered needs to go through multiple rendering calculation steps in multiple rendering pipelines to achieve rendering, and the result of each rendering calculation step will be accumulated into the final rendering result. The data generated by the rendering calculation steps all contain specific information about the scene, such as textures, colors, normals, and depth information. These data will be combined to generate more complex effects, such as shadows, lighting, blur, glow, and other post-processing effects.

[0029] In the prior art, the rendering of each game instance requires completing multiple specific rendering calculation steps of that game instance, and the GPU on the cloud corresponding to that game instance is responsible for processing the multiple specific rendering calculation steps of that game instance.

[0030] However, in the GPU cloud environment of the prior art, for example, multiple game instances of the same game may have or require some identical scenes or data. Specifically, for example, in a multiplayer battle game, some people may play in the same scene, and these people actually share the same scene. If the same scene is rendered once for each person's different game instance, it may result in a large amount of rendering time and computing cost for the calculation steps. Moreover, in the prior art, using a dedicated GPU to handle a specific game instance may also lead to waste of the computing resources of that GPU.

[0031] The embodiments of the present application hope to reduce the overall rendering cost of the cloud by sharing as much computing and data as possible, and also hope to fully utilize the computing resources of the GPU in the cloud through a suitable scheduling mechanism.

[0032] In general, the embodiments of the present application rely on the powerful GPU resources on the cloud to separate the sharable rendering calculation steps in the traditional rendering pipeline and the sharable data in the traditional rendering pipeline. Dedicated or scheduled shared GPU resources can be used to run these sharable rendering calculation steps, and data can be generated once, and then the GPU (or central processing unit (CPU)) can be relied on to synchronously transfer the data to each individual pipeline for the remaining rendering.

[0033] Figure 2 A block diagram of a cloud-based distributed graphics rendering system 200 according to an embodiment of the present application is shown.

[0034] like Figure 2 As shown, the cloud-based distributed graphics rendering system 200 includes: a plurality of graphics processing units 206, 207, ... 208 in the cloud, such as GPU 1, GPU 2, ... GPU N, where N is a positive integer. One or more of the plurality of graphics processing units 206, 207, ... 208 in the cloud (e.g., Figure 2 GPU 1 206, GPU 2 207) shown in FIG. 1 run computational steps that can be shared by multiple graphics rendering pipelines (e.g., Figure 2 The rendering calculation step 204 shown can be shared to output and / or cache the data obtained by the operation for use by other rendering calculation steps.

[0035] Here, one or more of the multiple graphics processing units 206, 207, ... 208 in the cloud can be fixedly set to run the calculation steps that can be shared by multiple graphics rendering pipelines, or can be dynamically scheduled to run the calculation steps that can be shared by multiple graphics rendering pipelines, that is, it can be not fixed as GPU 1 206, GPU 2 207 to run the calculation steps that can be shared by multiple graphics rendering pipelines, but can be dynamically scheduled (for example, according to real-time tasks and load conditions) to run the calculation steps that can be shared by multiple graphics rendering pipelines. This application does not limit this.

[0036] like Figure 2The traditional rendering calculation steps 205 shown may include those calculation steps that cannot be shared by the above-mentioned multiple graphics rendering pipelines. Here, the term "traditional" is intended to emphasize those calculation steps that cannot be shared by the above-mentioned multiple graphics rendering pipelines, so as to distinguish them from the above-mentioned calculation steps that can be shared by multiple graphics rendering pipelines. In addition, the traditional rendering calculation steps 205 may include calculation steps that share the data output by these calculation steps that can be shared by multiple graphics rendering pipelines, that is, for example Figure 2 The GPU 1 206 and GPU 2 207 shown in FIG. 1 run computational steps that can be shared by the above-mentioned multiple graphics rendering pipelines (e.g., Figure 2 The sharable rendering calculation step 204 is shown to output and / or cache the data obtained from the operation for use by the calculation steps in the traditional rendering calculation step 205 that need to share and use the data.

[0037] When configuring a GPU to run a computing step, in one embodiment, it may be considered to pre-fix the configuration by, for example, GPU 1 206 and GPU 2 207 of multiple graphics processing units 206, 207 ... 208 in the cloud to specifically run computing steps that can be shared by multiple graphics rendering pipelines, and to fix the configuration by, for example, GPU N 208 other than GPU 1 206 and GPU 2 207 of multiple graphics processing units 206, 207 ... 208 in the cloud to run other traditional computing steps. In another embodiment, some of the multiple graphics processing units 206, 207 ... 208 in the cloud may be dynamically scheduled (for example, according to real-time tasks and load conditions) to run computing steps that can be shared by multiple graphics rendering pipelines and some GPUs to run other traditional rendering computing steps. Of course, the manner of scheduling multiple graphics processing units in the cloud to run computing steps shared by multiple graphics rendering pipelines and other traditional rendering computing steps is not limited to this.

[0038] In this way, by sharing as much computing and data as possible, the overall rendering cost in the cloud can be reduced.

[0039] In one embodiment, the computing steps that can be shared by multiple graphics rendering pipelines are determined based on whether the computing steps and / or output data of the computing steps can be used by computing steps of multiple graphics rendering pipelines.

[0040] In one embodiment, computation steps that can be shared by multiple graphics rendering pipelines include view independent computation steps.

[0041] For example, in a multiplayer game, some people may play in the same scene. The rendering of this scene is a rendering calculation step independent of each person's perspective. Such a perspective-independent rendering calculation step can be run once by the GPU, and then the output scene rendering data is provided to each person's scene rendering, without the need to render the scene for each person. Similarly, there can be some calculation steps and data that can be shared. Which calculation steps can usually be shared by multiple graphics rendering pipelines will be described in detail below.

[0042] Figure 3 A schematic diagram showing whether each processing module in a graphics rendering pipeline according to an embodiment of the present application can be shared is shown.

[0043] A graphics rendering pipeline generally includes a view-independent pre-processing module 301, a view-dependent pre-processing module 302, a shadow and lighting pre-processing module 303, an opaque lighting module 304, a transparent lighting module 305, a motion vector calculation module 306, and a post-processing module 307.

[0044] The view-independent pre-processing module 301 generally involves calculation steps for view-independent scene environment rendering (Environment Capture) for each person or pre-processing of skinning independent of each person's perspective, such as pre-processing for rendering the same scene in a multiplayer game.

[0045] The view-dependent pre-processing module 302 generally involves calculation steps for pre-processing of view-dependent rendering for each person, such as pre-processing for rendering objects seen from each game player's own perspective.

[0046] The shadow and lighting pre-processing module 303 generally involves calculation steps for rendering lighting and shadows under various light sources.

[0047] The opaque lighting module 304 generally involves calculation steps for rendering the lighting effects of opaque objects. For example, objects closer to the perspective (camera) are rendered first, and then distant objects are rendered, so that occluded objects do not need to be rendered.

[0048] The transparent lighting processing module 305 generally involves the calculation steps for rendering the lighting effects of transparent and semi-transparent objects.

[0049] The motion vector calculation module 306 generally involves calculating the motion vectors of pixels between two frames.

[0050] The post-processing module 307 involves some calculation steps for improving the image quality, such as anti-aliasing, etc.

[0051] In one embodiment, generally, the calculation steps that can be shared by multiple graphics rendering pipelines include the calculation steps that are independent of the viewing angle, such as the rendering of a scene. While the calculation steps related to the viewing angle, such as the rendering of each object, generally cannot be shared by multiple graphics rendering pipelines and need to be calculated separately for each viewing angle.

[0052] In Figures 3 to 7B it, different shades of gray are used to indicate whether the module may be shared by multiple graphics rendering pipelines. For example, a light gray shade indicates that the module can be shared by multiple graphics rendering pipelines, while no shade indicates that the module cannot be shared by multiple graphics rendering pipelines, and a dark gray shade indicates that some (but not all) calculation stages in the module can be shared by multiple graphics rendering pipelines. Among them, generally, the calculation stages independent of the viewing angle can be shared by multiple graphics rendering pipelines, while the calculation stages related to the viewing angle cannot be shared by multiple graphics rendering pipelines and need to be calculated separately for each viewing angle.

[0053] As Figure 3 shown, the calculation steps that can be shared by multiple graphics rendering pipelines include the calculation steps of the preprocessing module 301 independent of the viewing angle, the calculation steps 303 of the shadow and lighting preprocessing module 303, the motion vector calculation steps 306, and a part of the stages in the opaque lighting processing module 304, which will be described in detail later in combination with Figure 4 to.

[0054] Figure 4 shows a schematic diagram of whether the various calculation steps and data of each processing module in the graphics rendering pipeline according to the embodiment of the present application can be shared. For example, a light gray shade indicates that the module can be shared by multiple graphics rendering pipelines, while no shade indicates that the module cannot be shared by multiple graphics rendering pipelines, and a dark gray shade indicates that some (but not all) calculation stages in the module can be shared by multiple graphics rendering pipelines.

[0055] Generally, the calculation stages independent of the viewing angle can be shared by multiple graphics rendering pipelines, while the calculation stages related to the viewing angle cannot be shared by multiple graphics rendering pipelines and need to be calculated separately for each viewing angle.

[0056] As Figure 4As shown, in one embodiment, the computing steps that can be shared by multiple graphics rendering pipelines include a preprocessing computing step 301 that is independent of the viewing angle.

[0057] The computing steps usually involved in the preprocessing module 301 independent of the viewing angle are generally computing steps that can be shared by multiple graphics rendering pipelines, because the preprocessing computing steps related to the viewing-angle-independent scene environment rendering for each person - such as the preprocessing for rendering the same scene in a multiplayer game, etc. - can, for example, after being computed once, provide the output preprocessed data of the scene environment rendering to be shared and used in the scene renderings of multiple individuals, so as to render the scene environment of each person, especially the scene environment rendering of each person playing in the same scene.

[0058] For example, the computing steps independent of the viewing angle include at least one of a skinning computing step 401 and an environment capture computing step 402. Among them, the skinning computing step 401 is run by one or more graphics processing units in the cloud to output skinning vertex data 431 for use in the computing steps of multiple graphics rendering pipelines, and the environment capture computing step 402 is run by one or more graphics processing units in the cloud to output and / or cache cube map data 432 for use in the computing steps of multiple graphics rendering pipelines.

[0059] Skinning refers to covering a skeletal model with skin. Environment capture refers to simulating an all-encompassing environment to obtain something similar to a cube with 6 faces. These two computing steps are both independent of the viewing angle.

[0060] Figure 5A and 5B shows a schematic diagram of the computing stages and output data that can be shared by multiple graphics rendering pipelines in two types of skinning computing steps according to an embodiment of the present application.

[0061] Figure 5A shows a schematic diagram of the computing stages and output data that can be shared by multiple graphics rendering pipelines in a traditional skinning computing step.

[0062] As Figure 5A shown, a traditional skinning computing step includes a Vertex Shader and a Pixel Shader step.

[0063] The Vertex Shader is used to render vertices, such as computing lighting and texture coordinates for vertices, etc. The Pixel Shader is used to render pixels. Generally, the skinned vertex data output by the Vertex Shader step can be used as input data for subsequent stages.

[0064] According to an embodiment of the present application, the skinned vertex data output by the vertex shader step can be shared and used by multiple graphics rendering pipelines. Here, the skinned vertex data is a set of vertices that make up a mesh (skin) rendered by the skeletal model itself. A vertex data consists of a texture coordinate and one or more weights. A vertex position (or coordinate) can be regarded as a weighted average of the weighted positions after matrix transformation.

[0065] The view-independent skinned vertex data represents the texture coordinates and weights rendered by the skeletal model, which can be fully shared and used by other graphics rendering pipelines to generate the skin rendering effect. In this way, the computational resource cost and time cost of other graphics rendering pipelines for recalculating and storing the skinned vertex data are saved.

[0066] Figure 5B A schematic diagram showing the computational stages and output data that can be shared by multiple graphics rendering pipelines in the Ray Tracing Friendly Skinning calculation step.

[0067] In ray tracing, a ray passing through the center of each pixel in the image is traced, and then it is tested whether this ray intersects any geometry in the scene. If an intersection is found, the pixel color can be set to the color of the object intersected by the ray. Since a ray may intersect multiple objects, generally the closest intersection distance needs to be tracked.

[0068] As Figure 5B shown, the Ray Tracing Friendly Skinning calculation step includes running a ComputeShader and obtaining skinned vertex data. The ComputeShader is a shader that can run on the GPU to perform arbitrary calculations. The skinned vertex data can be used as input for the ray tracing shader or acceleration data structure generation.

[0069] Next, the Ray Tracing Friendly Skinning calculation step further includes a vertex shader, a pixel shader, Acceleration Data Structure Generation, RayGeneration / Intersection Shader, and Shadowing, etc. These are all used for tracking the generation, reflection, refraction, and path tracing of rays, as well as hits / intersections.

[0070] Among them, according to an embodiment of the present application, the skinned vertex data obtained by running the ComputeShader can be shared and used by other graphics rendering pipelines. The skinned vertex data can be saved in a buffer so that all other graphics rendering pipeline instances can extract data from this buffer for sharing purposes.

[0071] In this way, the computational resource cost and time cost of other graphics rendering pipelines for recalculating and storing skinned vertex data are saved.

[0072] Of course, the specific stages of the skinning calculation steps are not limited to Figure 5A and Figure 5B the stages shown. This application does not limit the specific stages of the skinning calculation steps, as long as the skinned vertex data obtained in the skinning calculation steps can be shared and used by other graphics rendering pipelines.

[0073] The Environment Capture calculation step 402 is to achieve the reflection and refraction effects in environment mapping and obtain a cube map with the scene environment. Since this cube map is independent of the viewing angle, it can also be shared and used by other graphics rendering pipelines.

[0074] Preprocessing calculation steps related to the viewing angle usually cannot be shared by multiple graphics rendering pipelines and need to be calculated separately for each viewing angle.

[0075] As Figure 4 shown, the preprocessing module 302 related to the viewing angle usually involves calculation steps for preprocessing rendering related to each person's viewing angle, such as objects seen from the viewing angles of individual game players.

[0076] Specifically, the calculation steps of the preprocessing module 302 related to the viewing angle usually include a depth preprocessing calculation step (Depth Pre-pass) 403 for obtaining the depth stencil view 433 of pixels, a G-buffer (geometry buffer) calculation step (G-Buffer Pass) 404 for storing the position, normal, diffuse color, and other useful material parameters corresponding to each pixel, a server-side rendering (Server Side Render, SSR) calculation step 405 for obtaining the reflection data 434, a Hi-Z culling calculation step 406 for rendering only the objects that can be seen within the view frustum to obtain the culling result 435, a screen space ambient occlusion (Screen Space Ambient Occlusion, SSAO) 407 for achieving an approximate ambient occlusion effect, etc. Among them, the G-buffer is a general term for all textures used to store light-related data and used in the final light processing stage. Server-side rendering means rendering on the server side.

[0077] Of course, the above calculation steps are some examples of the calculation steps of view - related pre - processing, not limitations, and may also include more, fewer, or different calculation steps.

[0078] These view - related pre - processing calculation steps generally cannot be shared by multiple graphics rendering pipelines and need to be calculated separately for each view because they are all related to each person's view. The objects seen from different people's views may be different, so it is usually difficult to share the objects seen from one person's view with those seen from other people's views.

[0079] In one embodiment, the calculation steps that can be shared by multiple graphics rendering pipelines include shadow and lighting pre - processing calculation steps. Because the shadow and lighting data obtained from the shadow and lighting pre - processing calculation steps can also be used by other rendering pipeline instances for shadowing and simulating real lighting.

[0080] Specifically, as Figure 4 shown, the shadow and lighting pre - processing calculation steps that can be shared by multiple graphics rendering pipelines include the Character Shadow Maps calculation step 408, the Scene Shadow Maps calculation step 409, the Cluster Lights and Objects calculation step 410, the Character Ambient calculation step 411, and the Cluster Lights and Objects calculation step 412 for the character.

[0081] The Character Shadow Maps refer to attaching a shadow map to the character and making it move with the character to achieve realistic lighting and shadow effects for the character. The Scene Shadow Maps are to attach a shadow map to the scene to achieve realistic lighting and shadow effects for the scene. The Cluster Lights and Objects for the scene refer to clustering the lights and objects in the scene, and the Cluster Lights and Objects for the character refer to clustering the lights and objects of the character. The types of lights include spot lights, point lights, directional lights, etc. The Character Ambient calculation step refers to calculating the ambient light of the character.

[0082] Among them, multiple graphics processing units are configured to run these sharable computing steps to output and / or cache at least one of spot shadow maps, point shadow maps, and directional shadow maps 436 for use by the computing steps of multiple graphics rendering pipelines.

[0083] Specifically, Figure 6A and Figure 6B respectively show schematic diagrams of a view-independent shadow map stage and a view-dependent shadow map stage in shadow and lighting preprocessing according to an embodiment of the present application.

[0084] As Figure 6A shown, the vertex shader, depth test and rasterization, and cascaded shadow map (CSM) stage in the view-independent scene shadow map calculation steps can all be fully shared by multiple rendering pipeline instances. The only problem is how to reasonably select and tile the obtained shadow data for each pipeline instance.

[0085] The depth test refers to simulating the occlusion of distant objects by nearby objects through the depth test. Rasterization refers to using perspective projection to transform the three-dimensional representation of a triangle into a two-dimensional representation to "project" the vertices of the triangle onto the screen. The cascaded shadow map refers to providing a higher-resolution depth texture near the observer and a lower-resolution texture in the distance, which is achieved by dividing the viewing frustum and creating a separate depth map for each partition.

[0086] As Figure 6BAs shown, for view-dependent shadow maps such as per-instance view frustum data calculation stage, view-independent clustered lights, view-independent clustered objects, per-instance view updates stage, vertex shader stage, depth testing, and rasterization stage can be shared by multiple rendering pipeline instances. The tile operations in the cascade sparse shadow maps for all views stage can also be shared by multiple rendering pipeline instances. Per-instance view frustum data refers to calculating the frustum data of each instance's view, where the frustum refers to the truncated pyramid region (view frustum) with a rectangular base that the view can see. View-independent clustered lights and view-independent clustered objects refer to, for example, the clustered lights and clustered objects of the scene respectively. Per-instance view updates refer to the updates of each instance's view. Depth testing refers to simulating that nearby objects occlude distant objects through depth testing. Rasterization refers to using perspective projection to transform the three-dimensional representation of a triangle into a two-dimensional representation to "project" the vertices of the triangle onto the screen.

[0087] In addition, the information calculated in the scene cluster lights and objects calculation step 410, character ambient 411, and character cluster lights and objects calculation step 412 is also shared by multiple rendering pipeline instances.

[0088] Figure 6C The figure shows a schematic diagram of the lighting preprocessing stage in shadow and lighting preprocessing according to an embodiment of the present application.

[0089] In one embodiment, the calculation steps that can be shared by multiple graphics rendering pipelines include a part of the stages in lighting preprocessing.

[0090] Among them, as Figure 6CAs shown, some of the stages in the light preprocessing that can be shared by multiple graphics rendering pipelines include at least one of the Per Instance View Frustum Data stage, Cluster Lights independent of the view, Cluster Objects independent of the view, Diffuse Material Evaluation stage, Per Object Material Cache stage, Per Object Lighting Evaluation stage, Per Object Radiance Cache stage, and Distributed Cache independent of the view. Cache Reconstruction and Sampling cannot be shared by multiple graphics rendering pipelines and needs to be calculated separately.

[0091] Per Instance View Frustum Data refers to the data of the objects visible in the view frustum of each instance. Cluster Lights independent of the view refer to the cluster lights that can be used for the scene, and Cluster Objects independent of the view refer to the cluster objects that can be used for the scene. Diffuse Material Evaluation refers to evaluating the roughness of the surface material of the diffuse material, etc. Per Object Material Cache refers to caching the material information of each evaluated object. Per Object Lighting Evaluation refers to evaluating the lighting information of each object, and Per Object Radiance Cache refers to caching the brightness and chromaticity values of the shading points of each object. Distributed Cache independent of the view refers to the distribution of the cache of material and lighting data independent of the view. Cache Reconstruction and Sampling refers to reconstructing and sampling the cached data.

[0092] That is to say, the clustering of lights and objects in the global scene (i.e., Cluster Lights independent of the view, Cluster Objects independent of the view) can be effectively completed in shared processing. However, the view-dependent cluster lights and cluster objects cannot be shared and need to be calculated separately. The Per Instance View Frustum Data stage, Diffuse Material Evaluation stage, Per Object Material Cache stage, and Per Object Lighting Evaluation stage are independent of the view and can therefore be shared. If object space lighting is enabled, the lighting information of each object can also be pre-calculated as a radiance cache. That is, the Per Object Radiance Cache stage can also be shared. Subsequently, the material and lighting data in the global scene can be distributed to each pipeline instance for view-dependent lighting. That is, the Distributed Cache independent of the view stage can be shared.

[0093] Next, lighting is usually the most complex stage in the rendering pipeline. Opaque lighting refers to the lighting effect of rendering opaque objects, and transparent lighting refers to the lighting effect of rendering transparent or semi-transparent objects.

[0094] As Figure 4 shown, some stages in the lighting process in the opaque lighting calculation step 304 and some stages in the view-independent ambient occlusion are calculation steps that can be shared by multiple graphics rendering pipelines. Generally, view-independent stages can be basically shared by multiple graphics rendering pipelines, while view-dependent stages can basically not be shared by multiple graphics rendering pipelines and need to be calculated separately for each view.

[0095] Figure 7A FIG. shows a schematic diagram of the lighting stage in the opaque lighting calculation step 304 according to an embodiment of the present application.

[0096] As Figure 7A shown, the stages of preparing physically based rendering (PBR) materials, preparing lighting materials, and the mesh shader in the lighting stage cannot be shared by multiple graphics rendering pipelines and need to be calculated separately for each rendering pipeline instance, while the vertex shader stage, tessellation shader stage, geometry shader (GS Shader) stage, and pixel shader stage cannot be shared by multiple graphics rendering pipelines and need to be calculated separately for each rendering pipeline instance.

[0097] Physically based rendering refers to the use of a shading / lighting model modeled based on physical principles and the microfacet theory, and the use of surface parameters measured from reality to accurately represent real-world materials in rendering. Preparing lighting materials refers to preparing the materials used for lighting. The mesh shader refers to shading a mesh, where a mesh is a combination of a series of vertices, edges, and faces that define the shape of an object. The tessellation shader refers to further dividing a surface into smaller sub-surfaces (adding more vertices) to increase the number of triangular faces on the object's surface and shading these subdivided surfaces. The geometry shader refers to shading a primitive, where a primitive refers to a set of vertices that make up a three-dimensional entity. For example, a point in space corresponds to one vertex, a line segment corresponds to two vertices, and a triangle corresponds to three vertices, and these are all primitives.

[0098] Figure 7B FIG. shows a schematic diagram of the view-independent ambient occlusion stage in the opaque lighting calculation step 304 according to an embodiment of the present application.

[0099] Among them, as Figure 7B shown, some stages in the view-independent ambient occlusion include at least one of the Transform Geometry stage, the Per Vertex Ray Marching stage, the Denoising stage, the Generate Cache stage, and the view-independent DistributedCache stage. Cache Reconstruction and Sampling cannot be shared by multiple graphics rendering pipelines and need to be calculated separately.

[0100] Transforming geometry means transforming the geometry of an object. Per vertex ray marching means projecting rays pixel by pixel from the view center in the viewing direction, and based on the principle of reversible light path, inversely finding which objects the light hitting this pixel comes from. It is done by stepping forward along the ray and judging whether a luminous object is hit along the way. Denoising means removing noise (dots) in the rendered image. Generating cache means generating cache data. View-independent distributed cache means distributing cache data independent of the view.

[0101] In this way, multiple graphics processing units are configured to run these shareable calculation steps to output and / or cache Light Data 437, Diffuse Data 438, and Specular Data 439 for use by the calculation steps of multiple graphics rendering pipelines. Among them, the light data reflects the lighting information of the object, the diffuse data reflects the diffuse reflection information of the object, and the specular data reflects the specular reflection information of the object. Diffuse reflection generally occurs when light hits an object with a rough surface. It means that when light hits the surface, some of it is absorbed and the other part is reflected. Specular reflection generally occurs when light hits an object with a very smooth surface. It means that when light hits the surface, almost all of the light is reflected.

[0102] As Figure 4 shown, in one embodiment, the motion vector calculation step 306 is a calculation step that can be shared by multiple graphics rendering pipelines. Among them, multiple graphics processing units are configured to run the step 417 of calculating the motion vector to output and / or cache the motion vector data 440 for use by the calculation steps of multiple graphics rendering pipelines.

[0103] The motion vector refers to the displacement of the same pixel between two frames.

[0104] Motion vectors can be used for motion blur to simulate the blurring effect produced when a fast-moving object is photographed with a slow shutter speed of a film camera.

[0105] As Figure 4 shown, the Fog rendering stage 414 in the transparent lighting process 305, the server-side rendering SSR calculation stage 415, the Forward Objects calculation stage 416, and the Temporal Anti-Aliasing (TAA) stage 418, the Motion Blur stage 419, the Bloom rendering stage 420, and the User Interface (UI) rendering stage in the post-processing 307 cannot be shared by multiple graphics rendering pipelines.

[0106] Fog rendering refers to simulating the effect of fog in the atmosphere. Server-side rendering SSR is performed on the server side. Forward rendering objects refer to rendering the objects illuminated by each light source for each light source. Therefore, depending on the number of light sources in the scene and whether they illuminate the objects, some objects may be rendered multiple times. Temporal anti-aliasing method refers to using temporal information (such as data from historical frames) to improve the anti-aliasing effect, which is used to repair / clean up aliasing in graphics. The temporal anti-aliasing method calculates only one pixel point for any one frame. As multiple frames accumulate, there are multiple sampled pixels, and then the results are mixed to improve the sampling efficiency and achieve the purpose of alleviating aliasing. Motion blur refers to using the motion vector information of pixels between two frames to simulate the blurring effect produced when a fast-moving object is photographed with a slow shutter speed of a film camera. Bloom rendering refers to the effect of the edge of light extending from the edge of a bright area, similar to the optical effect of highlight overflow, where the light from a bright source (such as a flash) appears to leak into surrounding objects. User interface rendering refers to rendering the user interface, such as the user interface operated by game players.

[0107] In this way, by sharing as much calculation and data as possible, the overall rendering cost in the cloud is reduced.

[0108] In the actual operation stage of the GPU cloud, multiple graphics processing units can be scheduled first to run the shared calculation steps. After obtaining and buffering the shared data, then multiple graphics processing units are scheduled to use this shared data to run other calculation steps. In this way, first concentrate on scheduling multiple graphics processing units to calculate all the shared calculation steps at once, and then schedule multiple graphics processing units to calculate other calculation steps. This can reduce the situation where calculation steps wait for each other's execution results because the execution results of the shared calculation steps are more likely to be relied on by other calculation steps. Therefore, some other calculation steps may wait for the execution results of the shared calculation steps.

[0109] Of course, the above scheduling order is not mandatory. Multiple graphics processing units can also be scheduled as a whole to execute overall computing steps including sharable computing steps and other computing steps. For example, after a graphics processing unit finishes executing a sharable computing step, the graphics processing unit can be scheduled to execute the computing steps that share the sharable computing step, or scheduling can be performed without considering the interdependency relationships, as long as the required computing steps are waited for to finish during the computing process. That is, the scheduler can be configured to schedule all multiple graphics processing units to run all sharable computing steps and other computing steps.

[0110] As Figure 4 As shown on the right, assume that instances A, B, … need to be processed. Each instance is preprocessed 441 to obtain data 442 for each instance, and each instance is scheduled 443. Then, first, sharable computing steps are performed, including view-independent processing 444, shadow and lighting preprocessing 447, diffuse reflection processing 450 in opaque lighting, and sharable computing steps and motion vector calculation 451. After obtaining sharable data, including cube texture map 445, spot light source shadow map, point light source shadow map, parallel light source shadow map 448, scene-related (i.e., view-independent) lighting data 449, and buffered motion vectors 452, non-sharable computing steps are performed, that is, some computing steps related to the view, including view-related processing 446, specular reflection processing and non-sharable computing steps 453 in opaque lighting, transparent lighting processing 454, and post-processing 455.

[0111] Thus, next, a scheduling scheme for scheduling the cloud GPU to run computations is introduced.

[0112] In the scheduling scheme of the prior art, in decentralized scheduling, when the cloud is relatively small, unnecessary overhead is added to each node. In centralized scheduling, under the assumption that atomic tasks are very small and can be subdivided, it is not suitable for the case of GPUs.

[0113] Figure 8 A block diagram of a cloud distributed graphics rendering system 200’ according to an embodiment of the present application is shown.

[0114] In one embodiment, in addition to including the same modules as Figure 2 shown, the system 200’ further includes: a scheduler 209 configured to schedule multiple graphics processing units 206, 207 … 208 to run sharable computing steps 204 and / or other computing steps 205.

[0115] In one embodiment, the scheduler is configured to first schedule multiple graphics processing units to run the computable steps that can be shared, and then schedule multiple graphics processing units to run other computable steps.

[0116] In one embodiment, the scheduler 209 is configured to allocate multiple tasks including computable steps 204 and / or other computable steps 205 that can be shared by multiple graphics rendering pipelines to multiple graphics processing units 206, 207... 208 based on the operation time of the multiple graphics processing units and the transmission bandwidth from the scheduler 209 to the multiple graphics processing units 206, 207... 208.

[0117] In one embodiment, the scheduler 209 is configured to allocate multiple tasks to multiple graphics processing units 206, 207... 208 such that for the multiple tasks, the sum of the transmission bandwidths from the scheduler 209 to the multiple graphics processing units 206, 207... 208 and the sum of the operation times of running the multiple tasks on the multiple graphics processing units 206, 207... 208 are minimized.

[0118] Specifically, for example, the scheduler 209 may be configured to allocate computable steps 204 and other computable steps 205 (collectively referred to as tasks) that can be shared by multiple graphics rendering pipelines to multiple graphics processing units 206, 207... 208 by minimizing the value of the following formula:

[0119]

[0120] where n is the total number of tasks, k represents the kth task, and x k represents the graphics processing unit among multiple graphics processing units 206, 207... 208 that is to run the kth task. T(x k ) represents the transmission bandwidth from the scheduler to the graphics processing unit that is to run the kth task, and C(x k ) represents the operation time of the kth task by the graphics processing unit that is to run the kth task.

[0121] By minimizing the sum E of the transmission bandwidths from the scheduler to the corresponding graphics processing units for running all n tasks and the sum of the operation times of the graphics processing units for running all n tasks for all n tasks, and allocating computable steps 204 and multiple graphics rendering pipelines 205 (collectively referred to as tasks) that can be shared by multiple graphics rendering pipelines to multiple graphics processing units 206, 207... 208, the overall operation cost can be minimized.

[0122] In this way, through the above method, the graphics processing unit can be optimally scheduled while sharing as much computation and data as possible to minimize the overall computing cost, thereby further reducing the overall rendering cost.

[0123] The above scheduling method can also be combined with a scheme of first scheduling multiple graphics processing units to run the computable steps that can be shared, and then scheduling multiple graphics processing units to run other computable steps, so that first, multiple graphics processing units are scheduled by the above scheduling method to run the computable steps that can be shared at one time, and then multiple graphics processing units are scheduled by the above scheduling method to run other computable steps. Of course, the order here is not restrictive, and it is also possible to schedule (the same or another) GPU to run another computable step that shares and utilizes the data obtained from this computable step after scheduling one GPU to complete one computable step that can be shared.

[0124] Of course, the scheduling method of the scheduler 209 is not limited to this. Other factors can also be considered, or some factors can be omitted, or scheduling can be performed in a different way of considering factors. The scheduling method is not limited here.

[0125] The system 200’ may further include: an accelerator 210 located in the cloud, configured to accelerate the operations of computable steps that can be shared by multiple graphics rendering pipelines and / or other non - shareable computable steps.

[0126] As Figure 8 shown, the scheduler 209 and the accelerator 210 are located in the cloud.

[0127] In one embodiment, the accelerator is configured to combine the resources required for computable steps that can be shared and / or other non - shareable computable steps, and use the combined resources to perform the operations of computable steps that can be shared and / or other non - shareable computable steps through the graphics processing unit.

[0128] Figure 9 Shows a schematic diagram of the operation stage of the accelerator as a cloud - based rendering engine according to an embodiment of the present application.

[0129] As Figure 9As shown, for example, Instance 1 901 of Game A, Instance 2 902 of Game A, and Instance 1 903 of Game B all need to be rendered, and they have some resource requirements respectively. For example, in Resource Agent 904, Instance 1 of Game A needs Resource 1_1 (i.e., a part of Resource 1), Instance 2 of Game A needs Resource 2_1 (i.e., a part of Resource 2) and Resource 1_2 (i.e., another part of Resource 1), and Instance 1 of Game B needs Resource 2_2 (i.e., another part of Resource 2) and Resource 3. Resource Agent 904 combines these resource requirements, for example, into Resource 1, Resource 2, and Resource 3. Resource Processor 905 allocates Resources 1, 2, and 3 to Renderer 1, Renderer 2, and Renderer 3 in Renderer 906 in the renderer so that the renderer performs Rendering Phase A (e.g., fixing and reusing computable steps that can be shared) and Rendering Phase B (e.g., the rendering pipeline for each instance). Then Renderer 906 sends the results rendered by the renderer back to Instance 1 901 of Game A, Instance 2 902 of Game A, and Instance 1 903 of Game B for display to the user.

[0130] Note that the resource agent, resource processor, and renderer here are all software programs / modules in the adder.

[0131] In this way, while sharing as much computation and data as possible, the resource requirements of each rendering instance can be combined to reduce the overall computing cost, thereby further reducing the overall rendering cost.

[0132] Note that the cloud-native practical scenarios of the distributed GPU based on the cloud according to the embodiments of the present application are not limited to games, the metaverse, digital twins, virtual cities, navigation, scenic stories, etc. The distributed GPU on the cloud can also be further divided into small cloud servers or sub-clouds. For example, the rendering and applications at a certain street intersection in a virtual city can be completed by a small cloud server or sub-cloud. That is to say, the image rendering scenarios can all use the embodiments according to the present application.

[0133] Figure 10 The flowchart of the cloud distributed graphics rendering method 1000 according to the embodiments of the present application is shown.

[0134] As Figure 10 shown, the cloud distributed graphics rendering method 1000 includes: Step 1001, running computable steps that can be shared by multiple graphics rendering pipelines through one or more of the multiple graphics processing units in the cloud; Step 1002, outputting the data obtained from the running for use in the computable steps of the multiple graphics rendering pipelines.

[0135] In this way, by sharing as much computation and data as possible, the overall rendering cost on the cloud is reduced.

[0136] In one embodiment, the computing steps that can be shared by multiple graphics rendering pipelines are determined based on whether the computing steps and / or the output data of the computing steps can be used by the computing steps of multiple graphics rendering pipelines.

[0137] In one embodiment, the computing steps that can be shared by multiple graphics rendering pipelines include computing steps independent of the viewing angle.

[0138] In one embodiment, the computing steps that can be shared by multiple graphics rendering pipelines include preprocessing computing steps independent of the viewing angle.

[0139] In one embodiment, the preprocessing computing steps independent of the viewing angle include at least one of a skinning computing step and an environment capture computing step, where the skinning computing step is run by one or more graphics processing units in the cloud to output and / or cache skinning vertex data for use by the computing steps of multiple graphics rendering pipelines, and the environment capture computing step is run by one or more graphics processing units in the cloud to output and / or cache cube map data for use by the computing steps of multiple graphics rendering pipelines.

[0140] In one embodiment, the computing steps that can be shared by multiple graphics rendering pipelines include shadow and lighting preprocessing computing steps.

[0141] In one embodiment, the shadow and lighting preprocessing computing steps that can be shared by multiple graphics rendering pipelines include at least one of a cascaded shadow map stage, a scene shadow map stage, a scene clustering lights and objects stage, a character environment stage, and a character clustering lights and objects stage in a character shadow map computing step. Among them, a part of the stages in the lighting preprocessing in the shadow and lighting preprocessing computing steps that can be shared by multiple graphics rendering pipelines includes at least one of a per-instance view frustum data stage, a viewing angle-independent clustering light, a viewing angle-independent clustering object, a diffuse material evaluation stage, a per-object material cache stage, a per-object lighting evaluation stage, a per-object irradiance cache stage, and a viewing angle-independent distribution cache stage. Multiple graphics processing units are configured to run the shared computing steps to output and / or cache at least one of spotlight source shadow map data, point light source shadow map data, and directional light source shadow map data for use by the computing steps of multiple graphics rendering pipelines.

[0142] In one embodiment, the computing steps that can be shared by multiple graphics rendering pipelines include at least one of a part of the stages in the lighting processing in the opaque lighting computing step and a part of the stages in the viewing angle-independent ambient occlusion.

[0143] In one embodiment, some stages in the lighting process include at least one of the material stage for preparing physically based rendering, the lighting material stage, and the mesh shader stage. Some stages in view-independent ambient occlusion include at least one of the geometry transformation stage, the per-vertex ray marching stage, the denoising stage, the cache generation stage, and the view-independent distribution cache stage. A plurality of graphics processing units are configured to run shared computing steps to output and / or cache at least one of lighting data, diffuse data, and specular reflection data for use in the computing steps of a plurality of graphics rendering pipelines.

[0144] In one embodiment, the computing steps that can be shared by a plurality of graphics rendering pipelines include motion vector computing steps. A plurality of graphics processing units are configured to run shared computing steps to output and / or cache motion vector data for use in the computing steps of a plurality of graphics rendering pipelines.

[0145] In one embodiment, method 1000 further includes: scheduling a plurality of graphics processing units by a scheduler located in the cloud to run shared computing steps and / or other non-sharable computing steps.

[0146] In one embodiment, the scheduling first schedules a plurality of graphics processing units to run shared computing steps, and then schedules a plurality of graphics processing units to run other computing steps.

[0147] In one embodiment, the scheduling includes: allocating a plurality of tasks including shared computing steps that can be shared by a plurality of graphics rendering pipelines and / or other non-sharable computing steps to a plurality of graphics processing units based on the operation time of the plurality of graphics processing units and the transmission bandwidth from the scheduler to the plurality of graphics processing units.

[0148] In one embodiment, the scheduling includes: allocating shared computing steps that can be shared by a plurality of graphics rendering pipelines and / or other non-sharable computing steps to a plurality of graphics processing units such that for a plurality of tasks, the sum of the transmission bandwidths from the scheduler to the plurality of graphics processing units and the sum of the operation times of running the plurality of tasks on the plurality of graphics processing units are minimized.

[0149] In this way, through the above method, it is possible to optimally schedule graphics processing units while sharing as many computations and data as possible to minimize the overall operation cost, thereby further reducing the overall rendering cost.

[0150] In one embodiment, method 1000 further includes: accelerating the operations of shared computing steps that can be shared by a plurality of graphics rendering pipelines and / or other non-sharable computing steps by an accelerator located in the cloud.

[0151] In one embodiment, accelerating the computing steps that can be shared by multiple graphics rendering pipelines and / or other non-shareable computing steps by an accelerator located in the cloud includes: combining, by the accelerator, the resources required for the shareable computing steps and / or other non-shareable computing steps, so as to perform the operations of the shareable computing steps and / or other non-shareable computing steps by means of a graphics processing unit using the combined resources.

[0152] In this way, while sharing as many computations and data as possible, the resource requirements of each rendering instance can be combined to reduce the overall computing cost, thereby further reducing the overall rendering cost.

[0153] Generally speaking, the embodiments of the present application rely on the powerful GPU resources in the cloud, separate the shareable rendering computing steps in the traditional rendering pipeline, separate the shareable data in the traditional rendering pipeline, can use dedicated shared GPU resources to run these rendering computing steps, and can generate data at one time, and rely on the GPU (or Central Processing Unit (CPU)) to synchronously transmit the data to each individual pipeline for the remaining rendering.

[0154] Figure 11 A block diagram of an exemplary electronic device suitable for implementing the embodiments of the present application is shown.

[0155] The electronic device may include a processor (H1); a storage medium (H2) coupled to the processor (H1) and storing computer-executable instructions therein for performing the steps of the various methods of the embodiments of the present application when executed by the processor.

[0156] The processor (H1) may include, but is not limited to, for example, one or more processors or microprocessors, etc.

[0157] The storage medium (H2) may include, but is not limited to, for example, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (such as hard disks, floppy disks, solid-state drives, removable disks, CD-ROMs, DVD-ROMs, Blu-ray discs, etc.).

[0158] In addition, the electronic device may further include (but is not limited to) a data bus (H3), an input / output (I / O) bus (H4), a display (H5), and input / output devices (H6) (such as a keyboard, a mouse, a speaker, etc.).

[0159] The processor (H1) may communicate with external devices (H5, H6, etc.) via the I / O bus (H4) through a wired or wireless network (not shown).

[0160] The storage medium (H2) may also store at least one computer-executable instruction for performing the respective functions and / or method steps in the embodiments described in the present technology when run by the processor (H1).

[0161] In one embodiment, the at least one computer-executable instruction may also be compiled into or form a software product, where when the multiple computer-executable instructions are run by the processor, they perform the respective functions and / or method steps in the embodiments described in the present technology.

[0162] Figure 12 A schematic diagram of a non-transitory computer-readable storage medium according to an embodiment of the present application is shown.

[0163] As Figure 12 shown, instructions are stored on the computer-readable storage medium 1220, and the instructions are, for example, computer-readable instructions 1210. When the computer-readable instructions 1210 are run by the processor, the respective methods described above can be executed. The computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. For example, the computer-readable storage medium 1220 may be connected to a computing device such as a computer. Then, when the computing device runs the computer-readable instructions 1210 stored on the computer-readable storage medium 1220, the respective methods described above can be performed.

[0164] Of course, the above specific embodiments are only examples and not limitations, and those skilled in the art can combine and combine some steps and devices from the above separately described embodiments according to the concept of the present application to achieve the effects of the present application. Such combined embodiments are also included in the present application, and such combinations are not described one by one here.

[0165] Note that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present application. Additionally, the above-disclosed specific details are only for the purposes of illustration and easy understanding, and not limitations. The above details do not limit the present application to necessarily adopt the above specific details for implementation.

[0166] The block diagrams of devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the phrase "and / or", and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with it.

[0167] The flowchart of steps in this disclosure and the above method descriptions are only illustrative examples and are not intended to require or imply that the steps of each embodiment must be performed in the order given. As those skilled in the art will recognize, the steps in the above embodiments can be performed in any order. Words such as "thereafter", "then", "next", etc. are not intended to limit the order of the steps; these words are only used to guide the reader through the description of these methods. In addition, any reference to a singular element using articles such as "a", "an", or "the" is not to be construed as limiting that element to the singular.

[0168] In addition, the steps and apparatuses in each embodiment herein are not limited to being implemented in a certain embodiment. In fact, according to the concepts of this application, relevant partial steps and partial apparatuses in each embodiment herein can be combined to conceive new embodiments, and these new embodiments are also within the scope of this application.

[0169] Each operation of the methods described above can be performed by any suitable means capable of performing the corresponding functions. Such means can include various hardware and / or software components and / or modules, including but not limited to hardware circuits, application-specific integrated circuits (ASICs), or processors.

[0170] The various illustrative logical blocks, modules, and circuits described can be implemented or performed using a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array signal (FPGA), or other programmable logic devices (PLDs), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor can be a microprocessor, but alternatively, the processor can be any commercially available processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, a microprocessor cooperating with a DSP core, or any other such configuration.

[0171] The steps of the methods or algorithms described in connection with the present disclosure may be directly embodied in hardware, in a software module executed by a processor, or in a combination of the two. The software modules may exist in any form of tangible storage medium. Some examples of storage media that may be used include random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, hard disks, removable disks, CD-ROMs, etc. The storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In an alternative, the storage medium may be integral with the processor. The software modules may be individual instructions or many instructions, and may be distributed over several different code segments, different programs, and across multiple storage media.

[0172] The methods disclosed herein include acts for implementing the described methods. The methods and / or acts may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of acts is specified, the order and / or use of specific acts may be modified without departing from the scope of the claims.

[0173] The above functions may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as instructions on a tangible computer-readable medium. The storage medium may be any available tangible medium accessible by a computer. By way of example, and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other tangible medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks usually reproduce data magnetically, while discs reproduce data optically with lasers.

[0174] Accordingly, the present disclosure may also include a computer program product, where the computer program product may perform the methods, steps, and operations given herein. For example, such a computer program product may be a computer software package, computer code instructions, a computer-readable tangible medium having computer instructions tangibly stored (and / or encoded) thereon that are executable by a processor to perform the operations described herein. The computer program product may include packaging materials.

[0175] Software or instructions may also be transmitted through a transmission medium. For example, transmission media such as coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, or microwave may be used to transmit software from a website, server, or other remote source.

[0176] In addition, modules and / or other suitable means for performing the methods and techniques described herein can be downloaded and / or otherwise obtained by a user terminal and / or a base station as appropriate. For example, such a device can be coupled to a server to facilitate the transfer of means for performing the methods described herein. Alternatively, the various methods described herein can be provided via a storage component (such as RAM, ROM, a physical storage medium such as a CD or a floppy disk) so that a user terminal and / or a base station can obtain the various methods when coupled to the device or provided with the storage component. In addition, any other suitable technique for providing the methods and techniques described herein to a device can be utilized.

[0177] Other examples and implementations are within the scope and spirit of the present disclosure and the appended claims. For example, due to the nature of software, the functions described above can be implemented using software executed by a processor, hardware, firmware, hardwiring, or any combination thereof. The features implementing the functions can also be physically located at various positions, including being implemented at different physical positions for different parts of the distributed functions. Also, as used herein, including in the claims, the "or" used in the listing of items starting with "at least one" indicates a disjunctive listing such that, for example, the listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). In addition, the term "exemplary" does not mean that the examples described are preferred or better than other examples.

[0178] Various changes, substitutions, and alterations to the techniques described herein can be made without departing from the teachings of the technology defined by the appended claims. In addition, the scope of the claims of the present disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Processes, machines, manufactures, compositions of events, means, methods, or acts that currently exist or will be developed that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Accordingly, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.

[0179] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present application. Thus, the present application is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0180] The foregoing description has been presented for purposes of illustration and description. In addition, this description is not intended to limit embodiments of the present application to the forms disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A cloud - distributed graphics rendering system, comprising: Multiple graphics processing units in the cloud; Wherein, one or more of the multiple graphics processing units run computational steps that can be shared by multiple graphics rendering pipelines to output and / or cache the data obtained from the operation for use by the computational steps of the multiple graphics rendering pipelines. Wherein, the computational steps that can be shared by multiple graphics rendering pipelines include pre - processing computational steps independent of the viewing angle. Among them, the pre - processing computational steps independent of the viewing angle include at least one of skinning computational steps and environment capture computational steps. Wherein the skinning computational steps are run by one or more graphics processing units in the cloud to output and / or cache skinning vertex data for use by the computational steps of the multiple graphics rendering pipelines. The environment capture computational steps are run by one or more graphics processing units in the cloud to output and / or cache cube map data for use by the computational steps of the multiple graphics rendering pipelines. Wherein, the computational steps that can be shared by multiple graphics rendering pipelines include shadow and lighting pre - processing computational steps. Wherein, the shadow and lighting pre - processing computational steps that can be shared by multiple graphics rendering pipelines include at least one of the cascaded shadow map stage, scene shadow map stage, scene clustering lights and objects stage, character environment stage, character clustering lights and objects stage in the character shadow map calculation steps. Wherein, a part of the stages in the lighting pre - processing of the shadow and lighting pre - processing computational steps that can be shared by multiple graphics rendering pipelines include each instance view frustum data stage, viewing - angle - independent clustering lights, viewing - angle - independent clustering objects, diffuse material evaluation stage, each object material cache stage, each object lighting evaluation stage, each object irradiance cache stage, viewing - angle - independent distribution cache stage. Wherein the multiple graphics processing units are configured to run the shareable computational steps to output and / or cache at least one of spot light source shadow map data, point light source shadow map data, parallel light source shadow map data for use by the computational steps of the multiple graphics rendering pipelines. Wherein, the computational steps that can be shared by multiple graphics rendering pipelines include at least one of a part of the stages in the lighting processing of the opaque lighting calculation steps and a part of the stages in the viewing - angle - independent ambient occlusion. Wherein a part of the stages in the lighting processing include at least one of the stage of preparing the material for physically - based rendering, the stage of preparing the lighting material, the mesh shader stage, and a part of the stages in the viewing - angle - independent ambient occlusion include at least one of the transform geometry stage, per - vertex ray marching stage, denoising stage, generate cache stage, viewing - angle - independent distribution cache stage. Wherein the multiple graphics processing units are configured to run the shareable computational steps to output and / or cache at least one of lighting data, diffuse data, specular data for use by the computational steps of the multiple graphics rendering pipelines. Wherein, the system further includes a scheduler located in the cloud, configured to schedule the multiple graphics processing units to run the sharable computing steps and / or other non-sharable computing steps. Wherein, the scheduler is configured to allocate multiple tasks including the computing steps sharable by multiple graphics rendering pipelines and / or other non-sharable computing steps to the multiple graphics processing units based on the operation time of the multiple graphics processing units and the transmission bandwidth from the scheduler to the multiple graphics processing units. Wherein, the scheduler is configured to first schedule the multiple graphics processing units to run the sharable computing steps by minimizing the value of the following formula, and then schedule the multiple graphics processing units to run the other computing steps by minimizing the value of the following formula: where n is the total number of tasks, k represents the k-th task, and x k represents the graphics processing unit among the multiple graphics processing units that is to run the k-th task, and T(x k ) represents the transmission bandwidth from the scheduler to the graphics processing unit that is to run the k-th task, and C(x k ) represents the operation time for the graphics processing unit that is to run the k-th task to operate on the k-th task.

2. The system according to claim 1, wherein, The computing steps sharable by multiple graphics rendering pipelines are determined according to whether the computing steps and / or the output data of the computing steps can be used by the computing steps of the multiple graphics rendering pipelines.

3. The system according to claim 1, wherein The computing steps sharable by multiple graphics rendering pipelines include view-independent computing steps.

4. The system according to claim 1 or 2 or 3, wherein The computing steps sharable by multiple graphics rendering pipelines include motion vector computing steps, wherein the multiple graphics processing units are configured to run the sharable computing steps to output and / or cache motion vector data for use by the computing steps of the multiple graphics rendering pipelines.

5. The system according to claim 1, wherein, The scheduler is configured to allocate the multiple tasks to the multiple graphics processing units such that, for the multiple tasks, the sum of the transmission bandwidth from the scheduler to the multiple graphics processing units and the sum of the operation time of running the multiple tasks on the multiple graphics processing units is minimized.

6. The system according to claim 1 or 2 or 3, further comprising: An accelerator located in the cloud, configured to accelerate the operation of the computing steps sharable by multiple graphics rendering pipelines and / or other non-sharable computing steps.

7. The system according to claim 6, wherein, The accelerator is configured to combine the resources required for the sharable computing steps and / or other non-sharable computing steps, and use the combined resources to perform the operation of the sharable computing steps and / or other non-sharable computing steps through the graphics processing unit.

8. A cloud distributed graphics rendering method, comprising: Running, by one or more of the multiple graphics processing units in the cloud, computing steps sharable by multiple graphics rendering pipelines; Outputting the data obtained from the running for use by the computing steps of the multiple graphics rendering pipelines. Wherein, the computing steps sharable by multiple graphics rendering pipelines include view-independent preprocessing computing steps, and wherein the view-independent preprocessing computing steps include at least one of a skinning computing step and an environment capture computing step. Wherein the skinning computing step is run by one or more of the multiple graphics processing units in the cloud to output and / or cache skinning vertex data for use by the computing steps of the multiple graphics rendering pipelines. The environmental capture calculation steps are run by one or more graphics processing units in the cloud to output and / or cache cube map data for use by the calculation steps of the multiple graphics rendering pipelines. Among them, the calculation steps that can be shared by multiple graphics rendering pipelines include shadow and lighting preprocessing calculation steps. Among them, the shadow and lighting preprocessing calculation steps that can be shared by multiple graphics rendering pipelines include at least one of the cascaded shadow map stage, scene shadow map stage, scene clustered lights and objects stage, character environment stage, and character clustered lights and objects stage in the character shadow map calculation steps. Among them, a part of the stages in the lighting preprocessing of the shadow and lighting preprocessing calculation steps that can be shared by multiple graphics rendering pipelines include the per-instance view frustum data stage, view-independent clustered lights, view-independent clustered objects, diffuse material evaluation stage, per-object material cache stage, per-object lighting evaluation stage, per-object irradiance cache stage, and view-independent distribution cache stage. Among them, the multiple graphics processing units are configured to run the sharable calculation steps to output and / or cache at least one of spotlight source shadow map data, point light shadow map data, and directional light shadow map data for use by the calculation steps of the multiple graphics rendering pipelines. Among them, the calculation steps that can be shared by multiple graphics rendering pipelines include at least one of a part of the stages in the lighting processing of the opaque lighting calculation steps and a part of the stages in the view-independent ambient occlusion. Among them, a part of the stages in the lighting processing include at least one of the material stage for preparing physically based rendering, the lighting material preparation stage, and the mesh shader stage. A part of the stages in the view-independent ambient occlusion include at least one of the transform geometry stage, per-vertex ray marching stage, denoising stage, generate cache stage, and view-independent distribution cache stage. Among them, the multiple graphics processing units are configured to run the sharable calculation steps to output and / or cache at least one of lighting data, diffuse data, and specular data for use by the calculation steps of the multiple graphics rendering pipelines. Among them, the method further includes: scheduling the multiple graphics processing units by a scheduler located in the cloud to run the sharable calculation steps and / or other non-sharable calculation steps. Among them, the scheduling includes: based on the operation time of the multiple graphics processing units and the transmission bandwidth from the scheduler to the multiple graphics processing units, allocating multiple tasks including the calculation steps that can be shared by multiple graphics rendering pipelines and / or other non-sharable calculation steps to the multiple graphics processing units. Among them, the scheduling first schedules the multiple graphics processing units to run the sharable calculation steps by minimizing the value of the following formula, and then schedules the multiple graphics processing units to run the other calculation steps by minimizing the value of the following formula: where n is the total number of tasks, k represents the k-th task, and x k represents the graphics processing unit among the multiple graphics processing units that is to run the k-th task. T(x k ) represents the transmission bandwidth from the scheduler to the graphics processing unit that is to run the k-th task, and C(x k ) represents the operation time for the graphics processing unit that is to run the k-th task to operate on the k-th task.

9. The method according to claim 8, wherein, The computing steps that can be shared by multiple graphics rendering pipelines are determined based on whether the computing steps and / or the output data of the computing steps can be used by the computing steps of the multiple graphics rendering pipelines.

10. The method according to claim 8, wherein, The computing steps that can be shared by multiple graphics rendering pipelines include computing steps independent of the viewing angle.

11. The method according to claim 8 or 9 or 10, wherein, The computing steps that can be shared by multiple graphics rendering pipelines include motion vector computing steps, where the multiple graphics processing units are configured to run the sharable computing steps to output and / or cache motion vector data for use by the computing steps of the multiple graphics rendering pipelines.

12. The method according to claim 8, wherein, The scheduling includes: allocating the multiple tasks to the multiple graphics processing units such that for the multiple tasks, the sum of the transmission bandwidths from the scheduler to the multiple graphics processing units and the sum of the operation times of running the multiple tasks on the multiple graphics processing units are minimized.

13. The method according to claim 8 or 9 or 10, further comprising: accelerating the operations of the computing steps that can be shared by multiple graphics rendering pipelines and / or other non - sharable computing steps through an accelerator located in the cloud.

14. The method according to claim 13, wherein The accelerating the operations of the computing steps that can be shared by multiple graphics rendering pipelines and / or other non - sharable computing steps through an accelerator located in the cloud includes: combining, by the accelerator, the resources required for the sharable computing steps and / or other non - sharable computing steps, so as to use the combined resources to perform the operations of the sharable computing steps and / or other non - sharable computing steps through the graphics processing unit.

15. An electronic device, comprising: a memory for storing instructions; a processor for reading the instructions in the memory and executing the method according to any one of claims 8 - 14.

16. A non - transitory storage medium having instructions stored thereon, Among them, wherein when the instructions are read by a processor, the processor is caused to execute the method according to any one of claims 8 - 14.

Citation Information

Patent Citations

  • Cloud-based realtime raytracing

    CN111402377A

  • Distributed task scheduling method, readable medium and electronic equipment

    CN115543590A

  • Graphic Processor

    KR1020160142166A