Method and system for rendering video graphics using scene segmentation

By using world space sampling clusters, programmatic instance culling masks and instance culling clusters on mobile devices, the problem of insufficient memory bandwidth when rendering video graphics on mobile devices is solved, and efficient rendering performance and quality are achieved.

CN120129927APending Publication Date: 2025-06-10创峰科技
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070881.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-21
Filing Date
2023-02-06
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When rendering video graphics on mobile devices, limited memory capacity and bandwidth lead to insufficient rendering performance, making it difficult to achieve high-quality graphics rendering.

Method used

The world-space sampling cluster, programmatic instance culling masks and instance culling clusters are used to reduce the memory bandwidth load and improve rendering performance through these technologies.

Benefits of technology

Implementing instance culling clusters through shared segmentation data significantly improves the performance of specific applications, reduces memory usage, and improves rendering quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120129927A_ABST
    Figure CN120129927A_ABST
Patent Text Reader

Abstract

Systems and methods for rendering graphics. The method includes generating and receiving graphic data from a rasterized 3D scene to determine at least a first object. The primitive data may be generated by a vertex shader using the vertex data of the scene. The provided position vector can define the center of the overall cluster data structure on three axes, and the provided subdivision scalar can divide the structure on the three axes. Using at least the subdivision scalar and the range of the structure, the clusters may be partitioned by bounding boxes. Geometric position data may be mapped to related clusters defined by center cluster positions. A culling mask may be generated for each object using the generated cluster data, and the scene may be rendered using at least the culling mask and the cluster data. Other embodiments include corresponding systems and computer programs configured to perform actions of the above method.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to U.S. Patent Application No. 63 / 418,050, filed on October 21, 2022 (Agency Case No. 1282.INNOPEAK - 1022 - 099 - P), entitled "Scene segmentation for Mobile", the entire content of which is incorporated herein by reference for all purposes. Background Art

[0003] As video graphics standards increase year by year, the resource costs for rendering such video graphics are also constantly climbing. In real - time applications (RTA) such as video games, video conferencing, virtual reality (VR) applications, and extended reality (XR) applications, optimizing these costs is particularly important. In addition, since the use of such real - time applications on mobile devices has become very common, it has become increasingly important to improve the quality of video graphics in mobile applications. However, compared to desktop computers, mobile devices have limited memory capacity and bandwidth, which poses challenges for achieving sufficient rendering performance. Although there are various solutions to address the memory - intensive problems of video graphics rendering, as described below, these solutions are still not entirely satisfactory.

[0004] Therefore, there is a need for new and improved systems and methods for rendering video graphics. Summary of the Invention

[0005] The present invention relates to graphics rendering systems and methods. According to specific embodiments, the present invention provides a method using world - space sampling clusters, procedural instance cull masks, and instance culled clusters. There are other embodiments.

[0006] Embodiments of the present invention can be used in combination with existing systems and processes. For example, the rendering system configuration according to the present invention and its related methods can be widely applied to various systems, including virtual reality (VR) systems, mobile devices, etc. In addition, various techniques according to the present invention can be integrated into existing systems through integrated circuit manufacturing, operating system software, and application programming interfaces (APIs). There are other advantages.

[0007] A system of one or more computers can be configured to perform particular operations or actions by having software, firmware, hardware, or a combination thereof installed on the system, which, when run, cause the system to perform these operations or actions. One or more computer programs can be configured by including instructions that, when executed by a data processing apparatus, cause the apparatus to perform these operations or actions. A general aspect includes a method for rendering graphics on a computer device. The method includes receiving a plurality of graphics data related to a three-dimensional (3D) scene, the plurality of graphics data including a plurality of vertex data related to a plurality of vertices in the 3D scene. The method further includes generating a plurality of screen space primitive data using at least the vertex data and a vertex shader, the plurality of primitive data can include at least barycentric coordinates and triangle indices. The method further includes providing a position vector that defines the center of an overall cluster data structure along three axes. The method further includes a subdivision scalar for uniformly dividing the overall cluster data structure along the three axes. A cluster is created by dividing a bounding box using at least the subdivision scalar and the extent of the overall data structure. The method further includes mapping geometric position data to an associated cluster defined by the cluster center position. The method further includes generating cluster data corresponding to the cluster by executing threads for each cluster by one or more compute shaders. The method further includes generating a culling mask for each object using the cluster data. The method further includes rendering an object represented in a screen space primitive buffer using at least the culling mask and the cluster data by a shader. Other embodiments of this aspect include corresponding computer systems, apparatuses, and computer programs recorded on one or more computer storage devices, respectively configured to perform the actions of the above method.

[0008] Embodiments can include one or more of the following features. The method can include storing Reservoir-Based Spatio-Temporal Importance Sampling (ReSTIR) reservoir data. The method can include storing an ordered list of light sources. The center of the overall grid is independent of the camera position. The center of the overall grid is located at the camera position. The method can include storing a culling mask, which can include a cull-to mask and / or a cull-from mask. The culling mask can be specified based on clusters corresponding to scene segments of uniform size. The culling mask can also be specified based on the material and the distance of each object from the camera. Implementations of the described techniques can include hardware, methods or processes, or computer software on a computer-accessible medium.

[0009] A specific aspect of the present invention includes a system for rendering video graphics. The system includes a memory that may include executable instructions. The system also includes a memory. The system also includes a processor coupled to the memory and the memory, the processor being configured to: generate a plurality of graphics data related to a 3D scene, the plurality of graphics data including a plurality of vertex data related to a plurality of vertices in the 3D scene; generate a plurality of screen space primitive data using at least the plurality of vertex data and a vertex shader, the plurality of primitive data may include at least barycentric coordinates and triangle indices; provide a position vector defining the center of the overall cluster data structure along three axes; provide a subdivision scalar for uniformly dividing the overall cluster data structure along three axes; create clusters by dividing a bounding box using at least the subdivision scalar and the range of the overall data structure; map geometric position data to associated clusters defined by the cluster center positions; generate cluster data corresponding to the clusters by executing threads for each cluster by one or more compute shaders; generate a culling mask for each object using the cluster data; and render the objects represented in the screen space primitive buffer using at least the culling masks and the cluster data by a shader. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, respectively configured to perform the actions of the above method.

[0010] Embodiments may include one or more of the following features. The processor in the system may include a central processing unit (CPU) and a graphics processing unit (GPU), where the GPU is configured. The memory is shared by the CPU and the GPU. The memory may include a frame buffer for storing a first object. The system may include a display configured to display the first object at a refresh rate of at least 24 frames per second. Embodiments of the described technology may include hardware, methods, or processes, or computer software on a computer-accessible medium.

[0011] A specific aspect of the present invention relates to a method for rendering graphics on a computer device. The method includes generating a 3D scene including a first object. The method further includes receiving a plurality of graphics data related to the 3D scene. The method further includes providing a position vector defining the center of the overall cluster data structure on three axes. The method further includes providing a subdivision scalar for dividing the overall cluster data structure along three axes. The method further includes providing a bounding box divided by using at least the subdivision scalar and the range of the overall data structure to create clusters. The method further includes mapping geometric position data to associated clusters defined by the cluster center positions. The method further includes generating cluster data corresponding to the clusters by executing threads for each cluster through one or more compute shaders. The method further includes generating a culling mask for each object using the cluster data. The method further includes rendering the 3D scene by a shader using at least these culling masks and the cluster data. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, respectively configured to perform the actions of the above method.

[0012] Embodiments may include one or more of the following features. In the method, the 3D scene is rasterized to determine at least the first object, and the first object is intersected by the primary rays passing through each pixel. The plurality of graphics data includes a plurality of vertex data related to a plurality of vertices in the 3D scene. The method may include generating a plurality of primitive data using at least the plurality of vertex data and a vertex shader, and the plurality of primitive data may include at least position data. The bounding box is evenly divided. The bounding box of the method is divided according to the horizon in the 3D scene. The method may include obtaining settings for dividing the bounding box. The method may include defining a culling mask based at least on the opacity of the objects in the 3D scene. Embodiments of the described technology may include hardware, methods, or processes, or computer software on a computer-accessible medium.

[0013] It should be recognized that the embodiments of the present invention have many advantages compared with traditional technologies. The present invention provides configurations and methods for a graphics rendering system, which reduce the memory bandwidth load and improve performance by using world space sampling clusters and procedural instance culling masks. In addition, the present invention improves the performance of specific applications by sharing split data to implement instance culling clusters.

[0014] The present invention achieves these and other advantages in the context of known technologies. However, the nature and advantages of the present invention can be further understood by referring to the subsequent text of the specification and the drawings. Description of the Drawings

[0015] Figure 1 is a simplified schematic diagram of a mobile device for rendering video graphics according to an embodiment of the present invention.

[0016] Figure 2 It is a simplified flowchart of a traditional forward pipeline for rendering video graphics.

[0017] Figure 3 It is a simplified flowchart of a traditional hybrid pipeline for rendering video graphics.

[0018] Figures 4A to 4D It is a simplified schematic diagram of a cluster segmentation method according to an embodiment of the present invention.

[0019] Figure 5 It is a simplified flowchart of a graphics rendering method according to an embodiment of the present invention.

[0020] Figure 6 It is a simplified flowchart of a graphics rendering method according to an embodiment of the present invention. Detailed implementation manners

[0021] The present invention relates to a graphics rendering system and method. According to specific embodiments, the present invention provides a method and system that utilize world space sampling clusters, procedural instance culling masks, and instance culling clusters. The present invention can be configured for real-time applications (RTAs), such as video conferencing, video games, virtual reality (VR), and extended reality (XR) applications. There are also other embodiments.

[0022] Mobile and XR applications often need to cope with various limitations in graphics processing. These limitations can affect the visual quality and performance of the applications and pose challenges in providing a smooth and engaging user experience. One limitation is the limited processing power of mobile devices. Compared with desktop computers, most smartphones and tablets are equipped with relatively small and energy-efficient processors that are not designed for intensive graphics processing tasks. Therefore, mobile applications usually need to utilize optimized algorithms and technologies to render graphics efficiently to avoid overloading the device hardware. Another limitation is the limited memory and storage space of mobile devices. Most smartphones and tablets have limited and shared memory, which can restrict the complexity and level of detail of the graphics that can be rendered and may require mobile applications to use efficient data structures and algorithms to minimize the memory and storage space they occupy. A third limitation is the limited battery life of mobile devices. Graphics processing can consume significantly the device's battery, so mobile applications must be designed to minimize the impact on battery life to avoid draining the battery during use. This may involve techniques such as reducing the frame rate, using low-resolution graphics, or disabling certain graphics features when not needed. Another limitation is the limited bandwidth and connectivity of mobile networks. Many mobile applications rely on network connections to access data and resources, but mobile networks can be slow and unreliable, especially in areas with poor coverage. This can affect the performance of graphics-intensive applications and may require the use of techniques such as data compression and caching to minimize the amount of data that needs to be transmitted over the network. To provide a smooth and engaging user experience, mobile applications must be carefully designed, taking full account of these limitations and using optimized algorithms and technologies to render graphics efficiently and effectively.

[0023] It should be recognized that embodiments of the present invention provide techniques related to but not limited to spatial partitioning and forward rendering to improve graphics rendering in mobile devices. Spatial partitioning is a computer graphics technique that improves rendering performance by dividing three-dimensional space into smaller, more manageable regions. In real-time rendering, this is achieved by organizing the objects in a scene into a data structure that allows for efficient spatial queries by traversing only the geometry relevant to the current query. One data structure used for spatial partitioning in forward rendering is the uniform spatial grid. The uniform spatial grid divides three-dimensional space into a regular cell grid, where each cell contains a list of the objects that intersect it. During rendering, the camera's view frustum is used to determine which cells are within the field of view, and only the objects in these cells are added to the list of objects to be rendered. This avoids the need to traverse the entire scene and can significantly reduce the number of objects that need to be considered for rendering. Another data structure commonly used for spatial partitioning in forward rendering is the binary space partitioning (BSP) tree. The BSP tree organizes the objects in a scene into a hierarchical tree structure, where each node represents a plane that divides the space into two regions. During rendering, the camera view frustum is used to determine which nodes are within the field of view, and only these nodes are traversed to find the objects to be rendered. This allows for more efficient culling of objects outside the view frustum and can further reduce the number of objects that need to be considered for rendering. In addition to improving rendering performance, spatial partitioning can also be used for other tasks such as visibility determination, collision detection, and lighting calculations.

[0024] Forward+ rendering is a computer graphics technique used to improve the performance and visual quality of real-time rendering. Forward+ rendering is an extension of the traditional forward rendering method, which renders each object in the scene individually by performing a complete shading for each pixel covered by the object. Forward+ rendering uses minimal preprocessing (prepass) to generate at least a screen-space depth buffer before the full shading pass. This depth buffer is used to reject the execution of more costly main shaders that would render pixels that are occluded due to having a depth value further away than the value in the depth buffer. The Forward+ rendering pipeline can be further extended by using spatial partitioning techniques to divide the three-dimensional space into smaller regions (such as grids or tree structures). The number of light sources in the scene is grouped in the screen-space subdivision part rather than in the world-space subdivision part. A group of pixels (hereinafter referred to as "mesh cells") will contain a certain number of light sources attached to these pixels. During the rendering process, each pixel queries only the light sources associated with its assigned subdivision part, rather than the complete list of light sources in the scene. This allows for efficient culling of objects outside the frustum and reduces the number of objects that need to be considered for rendering. Once the objects in the field of view are determined, a light source culling step is used to generate a list of light sources for each mesh cell, considering only the light sources that affect these objects. This further reduces the number of light sources that need to be processed and can save computational resources. Next, the light source rendering step calculates the lighting for each object. This involves selecting a light source from the list of light sources, evaluating the impact of that light source on the object's surface, accumulating the resulting color, and repeating this process for other light sources as needed. This step can be executed in parallel on the GPU, allowing for efficient processing of multiple light sources in a single pass. Finally, Forward+ rendering uses a shading step to apply the calculated lighting to the objects and generate the final image. This step can also be executed on the GPU, allowing for real-time rendering of the scene.

[0025] Forward+ rendering with a screen-space light list offers many advantages compared to traditional forward rendering. It allows for more efficient culling of objects and lights, which can improve rendering performance and reduce memory usage. It also allows for more lights to be used in the scene, which can improve the visual quality of lighting. Additionally, parallel processing on the GPU allows for real-time rendering of complex scenes. However, this rendering architecture also has some limitations and trade-offs. One drawback is the overhead of building and maintaining a spatial partitioning data structure. For dynamic scenes with changing object positions, the cost can be particularly high and can impact overall rendering performance if not executed efficiently. Another issue is the choice of data structure and cell size. The optimal data structure and cell size will depend on the characteristics of the scene and the desired performance, but finding the right balance can be challenging. Depth preprocessing and light lists are both valuable techniques for improving real-time rendering performance and visual quality. However, these methods are not directly applicable to ray tracing, as ray tracing rendering effects rely on rendering the entire scene. Using a screen-space light list in a path tracer will result in a significant loss of lighting data during rendering, leading to artifacts or other issues during rendering.

[0026] It should be recognized that embodiments of the present invention, as described in further detail below, efficiently implement scene segmentation and forward rendering techniques for mobile applications.

[0027] The following description is intended to enable a person having ordinary skill in the art to make and use the invention and to apply it to the context of a particular application. Various modifications and various uses in different applications will be apparent to those skilled in the art, and the general principles defined herein can be applied to a wide range of embodiments. Therefore, the invention should not be limited to the embodiments described herein, but should be accorded the widest scope in accordance with the principles and novel features disclosed herein.

[0028] In the following detailed description, numerous specific details are set forth in order to provide a more thorough understanding of the present invention. However, it will be apparent to those skilled in the art that the invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the invention.

[0029] Readers of the present invention are cautioned that all documents and literature filed simultaneously with this specification, as well as documents and literature publicly available for inspection together with this specification, are hereby incorporated by reference in their entirety. Unless otherwise expressly stated, all features disclosed in this specification (including any accompanying claims, abstract, and drawings) may be replaced by alternative features having the same, equivalent, or similar purpose. Thus, unless otherwise expressly stated, each feature disclosed herein is merely an example of a generic series of equivalent or similar features.

[0030] Moreover, no element in a claim that is not expressly recited as a "means" for performing a particular function or a "step" for performing a particular function shall be construed to be a "means" or "step" clause as set forth in 35 U.S.C. § 112, ¶ 6. In particular, the use of "step of..." or "act of..." in the claims herein is not intended to invoke the provisions of 35 U.S.C. § 112, ¶ 6.

[0031] Note that, if used, terms such as left, right, front, back, top, bottom, forward, backward, clockwise, and counterclockwise are for convenience only and are not intended to imply any particular fixed orientation. Instead, these are used to reflect the relative position and / or orientation between various parts of an object.

[0032] Figure 1 is a simplified schematic diagram of a mobile device 100 for performing graphics rendering for use scenario segmentation. This figure is for example only and should not unduly limit the scope of the claims. Those of ordinary skill in the art will recognize many variations, alternatives, and modifications.

[0033] As shown, the mobile device 100 may be configured within a housing 110 and may include a camera device 120 (or other image or video capture device), a processor device 130, a memory 140 (e.g., volatile memory storage), and a storage 150 (e.g., permanent memory storage). The camera 120 may be mounted on the housing 110 and configured to capture input images. The input images may be stored in the memory 140, which may include a random-access memory (RAM) device, an image / video buffer device, a frame buffer, etc. Various software, executable instructions, and files may be stored in the storage 150, which may include a read-only memory (ROM), a hard disk, etc. The processor 130 may be coupled to each of the above components and configured to communicate between these components.

[0034] In a specific example, the processor 130 includes a central processing unit (CPU), a network processing unit (NPU), etc. The device 100 may further include a graphics processing unit (GPU) 132, which is connected to at least the processor 130 and the memory 140. In the example, the memory 140 is configured to be shared between the processor 130 (such as the CPU) and the GPU 132, and is configured to store data used when the application is running. Since the memory 140 is shared, it must be used efficiently when using the memory 140. For example, the high memory usage of the GPU 132 may have a negative impact on the system performance.

[0035] The device 100 may further include a user interface 160 and a network interface 170. The user interface 160 may include a display area 162, which is configured to display text, images, videos, rendered graphics, interactive elements, etc. The display 162 may be connected to the GPU 132, and may also be configured to display at a refresh rate of at least 24 frames per second. The display area 162 may include a touch screen display (such as in a mobile device, a tablet computer, etc.). Alternatively, the user interface 160 may further include a touch interface 164, and the touch interface 164 is used to receive user input (such as a keyboard or a keypad in a mobile device, a laptop computer, or other computing devices). The user interface 160 can be used for real-time applications (RTA), such as multimedia streaming, video conferencing, navigation, video games, etc.

[0036] The network interface 170 may be configured to send and receive instructions and files for graphics rendering (such as using Wi-Fi, Bluetooth, Ethernet, etc.). In a specific example, the network interface 170 may be configured to compress or downsample images for transmission or further processing. The network interface 170 may be configured to send one or more images to a server for OCR. The processor 130 may be connected to and configured to communicate between the user interface 160, the network interface 170, and / or other interfaces.

[0037] In the example, the processor 130 and the GPU 132 may be configured to perform steps for rendering video graphics, and these steps may include steps related to executable instructions stored in the memory 150. The processor 130 may be configured to execute application instructions and generate a plurality of graphic data related to a 3D scene including at least a first object. The plurality of graphic data may include a plurality of vertex data related to a plurality of vertices (such as each object) in the 3D scene. The GPU 132 may be configured to generate a plurality of primitive data using at least the plurality of vertex data and a vertex shader. The plurality of primitive data may include at least position data, etc.

[0038] In the example, the GPU 130 can be configured to provide a position vector that defines the center of the overall cluster data structure along three axes. The GPU 132 can also be configured to provide a subdivision scalar that is used to uniformly divide the overall cluster data structure along the three axes. Then, the GPU 132 can be configured to create clusters by dividing a bounding box using at least the subdivision scalar and the extent of the overall data structure. The GPU 132 can be configured to map geometric position data to the associated clusters defined by the cluster center positions. The GPU 132 can also be configured to generate cluster data corresponding to the clusters by executing threads for each cluster via one or more compute shaders. Using the cluster data, the GPU 132 can be configured to generate a culling mask for each object. Additionally, the GPU 132 can be configured to render at least the first object (and any other objects) via a shader using at least the culling mask and the cluster data.

[0039] To execute the described rendering pipeline, the mobile device 100 includes hardware components that need to be configured and optimized to meet the requirements of such a rendering technique. First, the CPU of the device needs to be powerful enough and energy-efficient to handle the computations and data processing required for forward+ rendering. This typically involves using a high-performance CPU with multiple cores and threads, as well as dedicated instructions and techniques such as SIMD and out-of-order execution to maximize performance and efficiency.

[0040] Second, the GPU of the device needs to be able to execute the complex shaders and algorithms used in the described rendering pipeline. This typically involves using a high-performance GPU with a large number of compute units and fast memory bandwidth, as well as supporting advanced graphics APIs and features such as OpenGL ES 3.0 and compute shaders.

[0041] Third, the memory and storage subsystem of the device needs to be large enough and fast enough to support the data structures and textures for storing the spatial grid cells that contain the data required for rendering. This typically involves using high-capacity and high-speed memory and storage technologies such as DDR4 RAM and UFS2.0 or NVMe storage, as well as efficient data structures and algorithms to minimize memory and storage usage.

[0042] Fourth, the display and touchscreen of the device need to have a high resolution and a fast refresh rate to support the visual quality and interactivity of real-time rendering. This typically involves using high-resolution and high-refresh-rate displays, as well as low-latency and high-precision touchscreens to achieve smooth and responsive interactions.

[0043] The battery and power management subsystem of the device needs to be able to support the power requirements of real-time rendering. This typically involves using high-capacity batteries and efficient power management technologies such as smart charging and power-saving modes to ensure that the device can run for a long time without depleting the battery.

[0044] To meet the requirements of the rendering pipeline described herein, the mobile device 100 has powerful and energy-efficient hardware components, including a high-performance CPU, GPU, memory and storage subsystem, display and touch screen, as well as a battery and power management subsystem. These components need to be optimized and configured to support the requirements of forward+ rendering and achieve smooth and attractive visual effects and interactions.

[0045] Other embodiments of the system include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the steps of the method. The details of the method are further discussed with reference to the accompanying drawings.

[0046] Figure 2 is a simplified flowchart of a conventional forward pipeline 200 for rendering video graphics. As shown, the forward pipeline 200 (i.e., includes a vertex shader 210, followed by a fragment shader 220). During the forward pipeline rendering process, the CPU provides the graphic data of the 3D scene (e.g., from memory, storage device, network, etc.) to the graphics card or GPU. In the GPU, the vertex shader 210 transforms the objects in the 3D scene from object space to screen space. This process includes projecting the geometry of the object and decomposing it into vertices, and then transforming these vertices and dividing them into fragments or pixels. In the fragment shader 220, these pixels are shaded (e.g., color, lighting, texture, etc.) and then passed to the display (e.g., the screen of a smartphone, tablet, VR goggles, etc.). In the case of lighting, the rendering effect is processed for each light source, for each vertex and each fragment in the scene.

[0047] Figure 3 is a simplified flowchart of a conventional deferred pipeline 300 for rendering video graphics. Here, the preprocessing process 310 involves receiving the graphic data of the 3D scene from the CPU and generating a G-buffer (G-buffer), which contains the data required for subsequent rendering passes, such as color, depth, normal, etc. The ray tracing reflection pass 320 involves processing the G-buffer data to determine the reflection of the scene, while the ray tracing shadow pass 330 involves processing the G-buffer data to determine the shadow of the scene. Then, the denoising pass 340 removes noise from the ray-traced pixels. In the main shading pass 350, the reflection, shadow, and material evaluation are combined to produce a shaded output with the color of each pixel. In the post-processing pass 360, the shaded output may undergo additional rendering processes, such as color correction, depth of field, etc.

[0048] Compared with the forward pipeline 200, the deferred pipeline 300 reduces the total number of fragments by processing rendering effects based only on unoccluded pixels. This is achieved by breaking the rendering process into multiple stages (i.e., passes), where the color, depth, and normals of objects in the 3D scene are written to separate buffers, and then these buffers are rendered together to produce the final rendered frame. Subsequent passes use the depth values to skip rendering of occluded pixels when executing more complex lighting shaders. The deferred rendering pipeline approach reduces the complexity of any single shader compared to the forward rendering pipeline approach, but multiple rendering passes require greater memory bandwidth, which is particularly problematic in modern mobile architectures with limited and shared memory.

[0049] According to an example, the present invention provides a method and system for world space sampling clusters. The space cluster system is configured with a vector that defines the range of the cluster data structure. The default value is {0,0,0}, using the bounding box of the entire scene and independent of the camera position. If set to a non-zero value, the space cluster is created centered on the camera, and the boundaries correspond to each configured range value.

[0050] Once the boundaries are established, the scene is further divided into smaller clusters based on a user-configured vector that specifies the number of subdivisions on each axis. The default value is {0,0,0}, allowing manual setting of the number of clusters and the boundaries of each cluster (explained further below). Each cluster contains data for accelerating the rendering of points within the cluster. This data is generated using a compute shader that executes threads for each cluster (e.g., Figure 2 and Figure 3 the shader shown in

[0051] Any improved sampling and data that can be used simultaneously should be stored in these clusters. In this example, to support a more general spatial data structure, light-source-specific optimizations are sacrificed, so a list of lights sorted by importance is stored instead of the cut points in the hierarchical light tree. Since ReSTIR integrates seamlessly with light-source importance sampling, ReSTIR reservoirs can also be stored per cluster. Cut-point data for cluster culling masks and environment map visibility can also be stored per cluster.

[0052] Finally, at runtime, the cluster data is loaded once to color the hit point based on the location of the hit point. Depending on the application, if the memory footprint of the hash map and the complexity of implementing an efficient hash scheme are manageable, a hash scheme can be considered for lookup. However, in this case, based on the ranges and subdivisions described earlier, the hit point location is divided into its associated clusters. If the subdivision is {0,0,0}, the cluster ID is derived from the instance culling mask, enabling a clustering method that depends on non-spatial attributes (such as material parameters).

[0053] Embodiments of such a general-purpose spatial clustering system can be used to segment a scene in order to selectively load fragment data and share data within a given fragment. For small scenes with fewer clusters, both the memory footprint and the number of computational schedules required to update the clusters for rendering are small, although they vary with the number of subdivisions. The low minimum cost of this simple spatial clustering system makes it suitable for mobile devices. Compared to screen-space clustering methods, this method introduces a dependency between scene scale and performance, but provides world-space locality that is more meaningful for ray tracing than the projective locality provided by screen-space clustering.

[0054] Furthermore, by avoiding partition optimizations specific to any single sampling optimization, the spatial clusters can be reused for various sampling improvements. Although this compromises the quality of any single sampling optimization, it decouples the construction and update time of the data structure from the number of optimizations implemented. This is particularly important when the grid cells are camera-centered, as they need to be updated every time the camera moves.

[0055] Generally, the culling mask helps reduce the cost of initializing inline ray tracing by reducing the proportion of the acceleration structure that is loaded into memory and traversed. Only the instances in the acceleration structure that match the culling mask are loaded. According to the example, the present invention provides a procedural setting method for instance culling masks based on different heuristics. Five example methods thereof will be discussed below. Since the ray tracing requirements of each scene vary with the scene content, these methods can be experimented with during development and the most suitable method for the scene can be selected by inspection. For a given scene, the selected segmentation method can be computed procedurally at load time to reduce the computational cost at runtime. In the following figures, the coordinate frame used for these descriptions is a coordinate frame with the Y-axis up, the X-axis from right to left on the screen, and the Z-axis from inside the screen outwards.

[0056] In a specific example, there are two important culling mask values, which will be referred to as cull-to and cull-from masks. The cull-to mask is the mask value packed together with the instance data built in the acceleration structure. Only the rays whose mask values are non-zero after a Boolean AND operation with the cull-to mask will intersect with these instances. The cull-from mask is the mask value packed together with the instance data directly accessed by the runtime shader code. The cull-from mask provides the mask value for the rays emitted from the hit points of the corresponding instances. An instance can have different cull-to and cull-from mask values. In the following example, the values will be numbered from 1 to 8, which corresponds to the bits set on an 8-bit culling mask, and all other bits are zero. Of course, there can be other variations, modifications, and alternatives.

[0057] The spatial grid is configured with a range vector and a center vector. By default, these are filled from the scene center and scene range, but the center can also be set to the camera position. For all the following grids, the range is divided by the number of segments on a given axis to determine the boundary positions between the segments. The internal boundaries are placed as calculated, but for grids with external faces, it is assumed that these faces extend to cover the rest of the scene. If the grid is intended to capture only the instances strictly within the range, a Boolean flag (e.g., TraceLocalOnly) can be enabled. This feature is applicable to scenarios where distant objects do not contribute much to the ray tracing effect.

[0058] According to the example, the present invention provides a method for assigning scene instances to segments based only on instance transformation (i.e., only spatial partitioning). The instances overlapping with a given segment will be assigned to that segment. The instances overlapping with multiple segments will be assigned to all the overlapping segments by setting the corresponding bits of these segments to 1. For only spatial partitioning, the cull-to mask and the cull-from mask are always equal.

[0059] Figure 4A is a simplified schematic diagram of the only spatial equal partitioning method according to an embodiment of the present invention. The most straightforward way to partition a scene is to divide the scene into equal parts. This is exactly the method adopted by the world space clusters: given a list of instances, calculate the boundaries of the scene and then divide it into equal-sized volumes. For an 8-bit instance culling mask, the number of volumes must be limited to eight. As shown in illustration 401, the most straightforward way is to use two layers of four segments (a uniform grid with a segment size of 2x2x2), defined by an axis-aligned bounding box (AABB). The AABB allows for efficient comparison of the boundaries with each instance for assignment to segments.

[0060] However, most game levels (e.g., RPG or FPS levels) are restricted to movement on a single plane, and the scene data is arranged more closely in the Y direction, so there is little need to subdivide objects according to their Y positions. Figure 4BIt is a simplified schematic diagram of the spatial single-plane segmentation method according to an embodiment of the present invention. As shown in Fig. 402, the number of segments along the Z-axis or X-axis is doubled (2D grid, segment size is 2x1x4). This axis can be selected according to the larger range of the total scene bounding box.

[0061] Finally, for scenes with large areas of transmissive and reflective elements, such as water bodies, the scene will be divided using the X-Z plane. However, these scenes usually have more scene geometries above or below the water surface. Figure 4C It is a simplified schematic diagram of the spatial asymmetric segmentation method according to an embodiment of the present invention. As shown in Fig. 403, only two half segments are assigned to the Y division with fewer instances, and the remaining six segments are assigned to the Y division with more instances (X-Z division grid, 2x1 segments above and 3x2 segments below).

[0062] In a specific example, this method can be further refined by setting the Y division to the value of the instance with the largest X-Z boundary square. Those of ordinary skill in the art will recognize other variations, modifications, and alternatives of these spatial-only segmentation methods (e.g., dividing along different axes, different numbers and ratios of divisions, etc.).

[0063] All the grids described above can be directly centered on the camera position and generate sizes according to the configured range. However, another grid view specifically for camera-centered rendering is also proposed here. Therefore, according to the example, the present invention provides a camera-centered segmentation method.

[0064] Generally, objects closer to the camera are more prominent and thus must be rendered with higher detail than objects in the distance. The camera-centered grid takes advantage of this and distributes grid cells according to the camera's transformation. Figure 4D It is a simplified schematic diagram of the camera-centered segmentation method according to an embodiment of the present invention. As shown in Fig. 404, four segments are assigned to cover the configurable maximum radius around the camera (labeled 1 to 4), and the remaining four segments (labeled 5 to 8) divide all the remaining instances in the scene into four quadrants.

[0065] Fragment boundaries may produce visible artifacts in indirect lighting, especially on mirror-like surfaces, which may show sharp cutoffs in reflections between instances assigned to different segments. In the example, the grid is aligned at a 45-degree angle with the camera's local space in the X-Z plane to reduce the visibility of these cutoff artifacts. This requires maintaining a separate "rotated view matrix" for transforming instances from world space to the rotated camera local space. Note that the bounding box of each instance must also be correctly transformed. Then, AABB comparison can be performed normally in the rotated view space.

[0066] In the example, to avoid complicating the comparison of the outer segments 5-8 of non-box shapes, all instances are first compared to the configured radius of the inner segment. Any instance that overlaps this radius can be added to a list of instances that will only be tested for inclusion within the inner segment, and all other instances can be added to another list of instances that will only be tested for inclusion within the outer segment. This eliminates the need to test the outer boundary during the boundary test for both lists, simplifying the boundary test to a quadrant test.

[0067] According to the example, the present invention provides material-based scene segmentation methods. These methods utilize the fact that some rays only need to intersect certain types of materials. For example, subsurface rays only need to intersect materials with subsurface scattering, while shadow rays must ignore emissive materials in order to correctly produce shadows. These methods also consider that the spatial properties of instances can differently affect the relevance of indirect illumination with different materials. For example, a perfectly smooth specular surface can be expected to reflect distant objects, while a rough object is typically only affected by the indirect illumination of nearby objects.

[0068] In a specific example, the material categories can include thick subsurface, light source, alpha cutoff, smooth transmission, smooth reflection, rough object, center, large projection scale, etc. Determining the material type can include a logical check of properties such as subsurface scattering, index of refraction (IOR), emission, opacity, roughness, distance from the center, center view size, etc. Additionally, each material category can separately define cull-to and cull-from values as described previously.

[0069] Based on the spatial properties of each instance, a center category and a large projection scale category are assigned. Instances that have already been assigned to other material culling groups can still be considered for assignment to these two culling groups. They can be set relative to the currently active camera or a user-configured view projection matrix. The latter may be useful when the scene content distribution is such that the ray tracing effect is only significant at static positions of interest.

[0070] The center category is calculated based on the view space distance of each instance relative to the origin. When considering the large projection scale category, all instances assigned to the center category are excluded.

[0071] To calculate the projected scale of an instance, multiply the camera's projection matrix by a modified view matrix that rotates the instance to the position (0, 0, d) in the modified view space when the distance between the camera and the instance is d. Then, after projecting each instance to the center of the view plane, the projected extent of each instance's bounding box can be used as a surrogate metric for the importance of the instance in the indirect rendering appearance. If the projected X or Y extent exceeds a user-configured threshold, the instance is added to a specified category (e.g., category 8 for an 8-bit mask).

[0072] The bandwidth consumption of the ray tracing process (e.g., rayQuery initialization) is the main bottleneck for mobile ray tracing. For example, by using an instance culling mask to initialize rayQuery with only 1 / 8 of the scene geometry, the bandwidth consumption is significantly reduced. This achieved a speed improvement of 15 - 20 fps on a mobile device when rendering a test scene with 35,000 triangles at a resolution of 1697x760. The instance culling mask for this test scene was manually configured based on the appearance of the scene, specifically based on the layout and materials of the scene geometry.

[0073] The method for programmatically setting these culling masks replaces the cumbersome manual configuration step with a simple mode selector, allowing the user to select the segmentation mode that best suits their scene by inspection. Additionally, embodiments of this method are the first to specify instance masks based on the scale or position of the instance in the scene. Moreover, the programmatic setting of instance culling masks does not require additional manual adjustment by the user.

[0074] According to an example, the present invention provides a method for sharing scene segmentation data between world space clusters and an instance culling system. In the example, both the spatial sampling clusters and the programmatic instance culling masks are managed by an overall scene segmentation manager. These two functions can be selectively linked, sacrificing the resolution of the spatial clusters and the flexibility of the instance culling masks in exchange for performance improvements brought about by greater reuse.

[0075] When linked, for an 8-bit mask, the bit width of the instance culling mask limits the maximum number of clusters in the spatial cluster to eight. At this lower resolution, the spatial clusters are less able to capture high-frequency lighting differences, so for scenes with hundreds of small light sources with different characteristics, the optimization based on spatial clusters will no longer bring much improvement. However, most modern games are first-person or third-person games where the camera view is centered on the player character, and most scene objects are similar in scale to the player character, and such scenes can be well covered with just eight clusters.

[0076] On the other hand, instance culling masks may only be able to assign categories using only spatial methods. Attempting to construct a bounding box from all the rough objects in a scene will almost certainly result in bounding boxes that overlap with other bounding boxes and span a large enough volume that it is not possible to share the available data within that volume.

[0077] As previously mentioned, the procedurally subdivided portions of the scene boundary can be covered with precomputed bounds. When a spatial sampling cluster is linked to a procedural instance culling mask, the procedural instance culling mask class can generate a bounding box for the active mesh pattern and pass it to the spatial cluster class. Then, since the spatial cluster has the same culling category as the instance, the bit offset of the culling-to value of the instance being shaded can be directly used to determine which spatial cluster contains the relevant data for improving shading. Since each instance can only be associated with one spatial cluster, this limits the procedural instance culling mask such that the culling-to value for each instance must be "one-hot", i.e., only one non-zero bit. Therefore, the previous method can be extended to check for overlaps between the instance and all candidate categories and only assign the instance to the category with the largest overlap.

[0078] Sharing the scene segmentation data between the world space clusters and the instance culling system reduces the overhead of maintaining both methods separately, since the scene data structure only needs to be computed or updated once and can then be shared between the two methods. This sacrifices the resolution of the spatial clusters (and thus the quality of scenes with high-frequency direct illumination) and can only be used with the only-spatial method of the procedural instance culling mask. Nevertheless, the camera-centric scheme for instance culling clusters achieves acceptable rendering quality for typical first-person and third-person game content and has improved performance.

[0079] Figure 5 is a simplified flowchart of a method 500 for rendering graphics on a computer device according to an embodiment of the present invention. This figure is for illustration only and should not unduly limit the scope of the claims. Those of ordinary skill in the art will recognize many variations, alternatives, and modifications. For example, one or more steps can be added, deleted, repeated, replaced, modified, rearranged, and / or overlapped, and these should not limit the scope of the claims.

[0080] According to an example, method 500 can be executed on a rendering system, such as Figure 1System 100 in. More specifically, the processor of the system can be configured to perform the operations of method 500 by executable code stored in the system memory (such as permanent memory). As shown, method 500 may include step 502 of receiving a plurality of graphics data related to a three-dimensional (3D) scene, which is rasterized to determine at least a first object (or all object instances in the scene), and the object will be intersected by a primary ray passing through a plurality of screen space pixels. This may include determining all data required for the first object intersected by the rays passing through each pixel in the viewport, where the plurality of graphics data includes a plurality of vertex data related to a plurality of vertices in the 3D scene. In an example, the method includes generating a plurality of screen space primitive data using at least the plurality of vertex data. The plurality of primitive data includes at least barycentric coordinates and triangle indices, but may also include other position data.

[0081] In step 504, the method includes providing a position vector defining the center of the overall cluster data structure on three axes. In a specific example, the center of the overall cluster data structure is independent of the camera position, located at the camera position, or at other positions. In step 506, the method includes providing a subdivision scalar for uniformly dividing the overall cluster data structure along three axes. In step 508, the method includes dividing the bounding box by using at least the subdivision scalar and the range of the overall data structure to create clusters. Dividing the bounding box may include any of the splitting techniques discussed above and their variations.

[0082] In step 510, the method includes mapping the geometric position data to the associated cluster defined by the cluster center position. In step 512, the method includes generating cluster data corresponding to the cluster by executing threads for each cluster through one or more compute shaders. In step 514, the method includes generating a culling mask for each object using the cluster data. In a specific example, the culling mask includes a culling-to mask and / or a culling-from mask. As described above, the culling mask can be specified based on clusters corresponding to equal scene segments, or based on the material and the distance of each object from the camera.

[0083] In step 516, the method includes rendering at least the first object (or all objects represented in the screen space primitive buffer) using at least the culling mask and the cluster data. The shader can be configured in a computer device, such as Figure 1 the device 100 shown in, and may include, for example, Figure 2 and Figure 3 the pipeline configuration shown in. In addition, the method may include storing ReSTIR reservoir data, an ordered list of light sources, or the like.

[0084] Figure 6FIG. 500 is a simplified flowchart of a method of rendering graphics on a computer device according to an embodiment of the present invention. This figure is for illustration only and should not unduly limit the scope of the claims. Those of ordinary skill in the art will recognize many variations, alternatives, and modifications. For example, one or more steps may be added, deleted, repeated, replaced, modified, rearranged, and / or overlapped, and this should not limit the scope of the claims.

[0085] According to an example, method 600 may be executed on a rendering system, such as Figure 1 system 100 in. More specifically, the processor of the system may be configured to perform the operations of method 500 by executable code stored in the memory (permanent storage) of the system. As shown, method 500 may include step 602 of generating a three-dimensional (3D) scene that includes a first object (or includes all object instances in the scene). In a specific example, the 3D scene includes a first object intersected by a primary ray.

[0086] In step 604, the method includes receiving a plurality of graphics data related to the 3D scene. In a specific example, the plurality of graphics data includes a plurality of vertex data related to a plurality of vertices in the 3D scene. In an example, the method further includes generating a plurality of primitive data using at least the plurality of vertex data and a vertex shader, and the plurality of primitive data includes at least position data.

[0087] In step 606, the method includes providing a position vector that defines the center of the overall cluster data structure along three axes. In step 608, the method includes providing a subdivision scalar for dividing the overall cluster data structure along three axes. In step 610, the method includes dividing a bounding box by using at least the subdivision scalar and the range of the overall data structure to create clusters. In a specific example, the bounding box is evenly divided, divided according to the horizon in the 3D scene, or the like and combinations thereof. In an example, the method further includes obtaining settings for dividing the bounding box. Dividing the bounding box may include any of the partitioning techniques discussed above and variations thereof.

[0088] In step 612, the method includes mapping geometric position data to an associated cluster defined by the cluster center position. In step 614, the method includes generating cluster data corresponding to the cluster by executing threads for each cluster through one or more compute shaders. In step 616, the method includes generating a culling mask for each object (e.g., the first object) using the cluster data. As described above, the method may further include defining the culling mask at least according to the opacity of the objects in the 3D scene.

[0089] In step 618, the method includes rendering the 3D scene through a shader using at least the culling mask and the cluster data. As described above, the shader may be configured in a computer device, such as Figure 1the device 100 shown therein, and may include, for example Figure 2 and Figure 3 the pipeline configuration shown therein.

[0090] Although the foregoing is a complete description of a particular embodiment, various modifications, alternative constructions, and equivalents may be used. Accordingly, the above description and drawings should not be regarded as limiting the scope of the invention, which is defined by the appended claims.

Claims

1. A method for rendering graphics on a computer device, the method comprises: Receiving a plurality of graphics data related to a three-dimensional (3D) scene, the 3D scene being rasterized to determine at least a first object, the first object being intersected by a primary ray passing through each of a plurality of screen space pixels, the plurality of graphics data including a plurality of vertex data related to a plurality of vertices in the 3D scene; Generating a plurality of screen space primitive data using at least the plurality of vertex data and a vertex shader, the plurality of primitive data including at least barycentric coordinates and triangle indices; Providing a position vector that defines the center of an overall cluster data structure on three axes; Providing a subdivision scalar for uniformly dividing the overall cluster data structure along the three axes; Creating clusters by dividing a bounding box using at least the subdivision scalar and the range of the overall data structure; Mapping geometric position data to associated clusters defined by the position of the cluster centers; Executing threads for each cluster in the clusters through one or more compute shaders to generate cluster data corresponding to the clusters; Generating a culling mask for each object using the cluster data; and Rendering the first object through a shader using at least the culling mask and the cluster data.

2. The method according to claim 1, further comprising storing ReSTIR reservoir data.

3. The method according to claim 1, further comprising storing an ordered light source list.

4. The method according to claim 1, wherein, the center of the overall cluster data structure is independent of the camera position.

5. The method according to claim 1, wherein, the center of the overall cluster data structure is located at the camera position.

6. The method according to claim 1, wherein, the culling mask includes a cull-to mask and / or a cull-from mask.

7. The method according to claim 1, wherein, the culling mask is specified based on clusters corresponding to equal scene segments.

8. The method according to claim 1, wherein, the culling mask is specified based on the material and the distance of each object relative to the camera.

9. A system for rendering video graphics, the system comprises: A memory including executable instructions; A memory; and A processor connected to the memory and the memory, the processor being configured to: Generate a plurality of graphics data related to a three-dimensional (3D) scene, the 3D scene being rasterized to determine at least a first object, the first object being intersected by a primary ray passing through each of a plurality of screen space pixels, the plurality of graphics data including a plurality of vertex data related to a plurality of vertices in the 3D scene; Generate a plurality of screen space primitive data using at least the plurality of vertex data and a vertex shader, the plurality of primitive data including at least barycentric coordinates and triangle indices; Provide a position vector that defines the center of an overall cluster data structure on three axes; Provide a subdivision scalar for uniformly dividing the overall cluster data structure along the three axes; Create clusters by dividing a bounding box using at least the subdivision scalar and the range of the overall data structure; Map the geometric location data to an associated cluster defined by the location of the cluster center; Execute threads for each cluster in the cluster by one or more compute shaders to generate cluster data corresponding to the cluster; Generate a culling mask for each object using the cluster data; And Render the first object through a shader using at least the culling mask and the cluster data.

10. The system according to claim 9, Wherein, The processor includes a central processing unit (CPU) and a graphics processing unit (GPU).

11. The system according to claim 10, Wherein, The memory is shared by the CPU and the GPU.

12. The system according to claim 9, Wherein, The memory includes a frame buffer for storing the first object.

13. The system according to claim 9, further comprising a display configured to display the first object at a refresh rate of at least 24 frames per second.

14. A method for rendering graphics on a computer device, the method Comprises: Generate a three-dimensional (3D) scene, the 3D scene including a first object; Receive a plurality of graphics data related to the 3D scene; Provide a position vector that defines the center of the overall cluster data structure on three axes; Provide a subdivision scalar for dividing the overall cluster data structure along the three axes; Create clusters by dividing a bounding box using at least the subdivision scalar and the range of the overall data structure; Map the geometric location data to an associated cluster defined by the location of the cluster center; Execute threads for each cluster in the cluster by one or more compute shaders to generate cluster data corresponding to the cluster; Generate a culling mask for each object using the cluster data; And Render the 3D scene through a shader using at least the culling mask and the cluster data.

15. The method according to claim 14, Wherein, The 3D scene is rasterized to determine at least a first object that will be intersected by a primary ray passing through each of a plurality of screen space pixels, and the plurality of graphics data includes a plurality of vertex data related to a plurality of vertices in the 3D scene.

16. The method according to claim 15, further comprising generating a plurality of screen space primitive data using at least the plurality of vertex data and a vertex shader, the plurality of primitive data including at least barycentric coordinates and triangle indices.

17. The method according to claim 14, Wherein, The bounding box is evenly divided.

18. The method according to claim 14, Wherein, The bounding box is divided according to the horizon in the 3D scene.

19. The method according to claim 14, further comprising obtaining settings for dividing the bounding box.

20. The method according to claim 14, further comprising defining the culling mask based at least on the opacity of the objects in the 3D scene.